oddsrail
Oddsrail is a self-hosted MCP server for AI agents to research, check, simulate, and execute trades on Polymarket and Kalshi, with local signing and operator guardrails.
Search and fetch markets, orderbooks, price history, trades, positions, and resolution criteria across Polymarket and Kalshi.
Find and compare the same event across venues, quote book-walked trade costs, and audit settlement divergence.
Run signals: overshoot/fade detection and UMA dispute-risk scoring.
Run deterministic pre-trade checks and fractional-Kelly position sizing.
Place, cancel, and kill-switch orders; dry-run by default, live with local env config and signing.
Track order status, fills, positions, balance, and open orders.
Paper trade against the live book with P&L and reset.
Gaslessly split, merge, and redeem positions via Polymarket relayer when configured.
Stream realtime book events for up to 60 seconds.
Inspect builder attribution, leaderboard stats, server mode, reachability, and geoblock verdict.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@oddsrailfind Polymarket World Cup final markets and run overshoot on the favorite"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
oddsrail
Give your AI agent Polymarket tools, with local signing and spending limits.
oddsrail is an open-source MCP server. It gives Claude, Cursor or any MCP
client live market data, costed fills and order routing on Polymarket.
Operator-set guards enforce order limits. A separate deterministic
check_order tool reviews the market, outcome and price before submission;
ask your agent to call it before each order. It cannot guarantee that an
agent will choose or execute a sound strategy.
That is a real check_order result on a live market, not a mock-up. The agent
meant NO and built a YES order. Without the check it would have taken the
opposite side of the trade, and nothing would have told it.
Review before submitting.
check_orderchecks the market, side, price and venue minimum. Operator guardrails run in the order path; calling the separate review tool remains the agent's responsibility.Non-custodial. Your keys stay on your machine. Everything starts in dry-run, with a paper ledger filled against the live order book.
Free, at 0 bps. oddsrail adds nothing to your trade; see how it is funded.
Read the OddsRail documentation for setup, wallets, trading limits and developer notes.
Claude Desktop: install without Python
Download OddsRail 0.19.0 for Claude Desktop. Open the file in Claude Desktop, or install it from Settings, Extensions. A current Claude Desktop version with MCPB/uv support supplies the local Python runtime. The extension starts with Paper trading ticked and needs no wallet key to explore markets or simulate orders.
For real orders, use an existing funded Polymarket account with trading approvals. Enter its signer private key only in the extension's settings, never in a chat or the website, and enter the trading account's public wallet address. Claude Desktop stores sensitive settings using its secure storage; the local OddsRail process receives the key for signing. Review your account, limits and proposed order before deliberately unticking Paper trading.
The defaults are $25 per order and $100 of submitted order notional per local server session. Restarting the extension resets the session budget. These are not daily loss or account-wide limits. Venue fees may apply.
The extension gives Claude local trading tools. Website strategy drafts can be handed to Claude for review; installing the extension does not import those rules or start an autonomous background agent. Keep Claude Desktop and your computer running while using the tools. Hosted website trading is a limited pilot: a Deposit Wallet session key plus an owner allowlist on the server, quoting one market at a time.
Related MCP server: telekash-mcp-server
Developer quickstart
Python 3.11+ required.
pip install oddsrailclaude mcp add --transport stdio oddsrail -- oddsrailOr from a clone, without installing:
python3 -m venv .venv && .venv/bin/pip install -r requirements.txtclaude mcp add --transport stdio oddsrail -- /abs/path/to/oddsrail/.venv/bin/python -m oddsrail.serverThen ask the agent: "search markets about the World Cup final and run the overshoot signal on the favorite".
Install in one step
Client | How |
Claude Desktop (local tools, default simulation) | Download the extension. No Python installation needed. |
Claude web or desktop, nothing to install (hosted, paper trading) | Settings, Connectors, Add custom connector, URL |
Claude Code (hosted, paper trading) |
|
Claude Code (plugin, with the four workflow skills) |
|
Claude Code (server only) |
|
Any agent that reads skills |
|
Cursor | |
VS Code | |
Anything else that speaks MCP over stdio |
|
The plugin and the one-click links launch the server with uvx, so they
need uv on the machine. Without uv, pip install oddsrail gives you an oddsrail command to point any client at.
Everything starts in dry-run.
The four skills (skills/*/SKILL.md) are generated from the server's own
MCP prompts by scripts/gen_skills.py, and a test fails if they drift, so a
skill and the prompt it mirrors can never disagree.
How oddsrail compares
Verified against each alternative directly (their repos, live endpoints, and registry entries, September 2026), not from their marketing:
oddsrail | raw venue APIs | pmxt | Simmer | Polymarket agent-skills | |
What it is | self-hosted MCP server | the venues themselves | unified API + SDK + MCP, "CCXT for prediction markets" | agent trading platform + SDK + MCP | markdown skill docs for agents |
Custody | non-custodial; keys never leave your machine | yours | hosted mode: "PMXT handles custody, signing infrastructure"; self-hosted mode: your keys | self-custody, local signing | yours (documentation only) |
Attribution you control | yes: | n/a | not documented | not documented | documents builder headers for your own code |
Cost to the trader | 0 bps, free tools | free | hosted pricing not in the README | not documented | free |
Open source | MIT, full source | n/a | MIT, ~2.1k stars | not stated | docs; license not stated |
Operator guardrails | notional caps, open-order cap, allowed markets; enforced pre-request, in dry-run too | none | not documented | per-trade limits, daily caps, stop-loss/take-profit, kill switch | none |
Paper trading | dry-run fills against the live book, P&L | none | not documented | virtual $SIM sandbox, then graduate to real money | none |
Book-walked cost, settlement audit, jurisdiction-classified failures, dated venue-quirk notes | yes, all four | no | not documented | not documented | quirks partly documented |
Realtime |
| websocket, yours to wire | not documented in the README | not documented | websocket documented |
Verified 2026-09-02 from each project's own README or docs (pmxt: github.com/pmxt-dev/pmxt; Simmer: docs.simmer.markets; agent-skills: github.com/Polymarket/agent-skills). "Not documented" means exactly that, not "absent". Re-check before quoting; these projects move.
The wedge, in one line: pmxt is the reference for trading everywhere; Simmer is the reference for an agent economy with a sandbox and a reputation layer; oddsrail is the reference for trading correctly, non-custodially, with attribution you own.
Where the others are honestly ahead: pmxt offers broader venue coverage for trading and data, with hosted convenience and a community many times ours. Simmer has a virtual-balance sandbox, stop-loss and take-profit rails we do not have, a public reasoning/reputation layer, and a strategy-skills marketplace. Polymarket's agent-skills is the venue's own documentation and covers bridging and deposits, which oddsrail does not.
The raw Polymarket API has the endpoints. It also models rejections as
ok:false return values, orders its books worst-first, ships a trades
endpoint that returns the market's public tape, and enforces an
undocumented $1 minimum notional. oddsrail exists because we hit every one of
those and encoded the fix.
Free to use, and free of fees
oddsrail ships with a project builder code
registered at 0 bps, so orders routed through it are attributed without
adding a single basis point to anyone's trade. The project's income is a share
of Polymarket's weekly builder reward pool, paid by Polymarket's own program,
not by you. Running your own builder profile instead is one environment
variable (ODDSRAIL_BUILDER_CODE), and server_info always tells you which
code is in use. No fee tiers, no paywalled tools, no account required.
Hosted: nothing to install
mcp.oddsrail.app runs the same server as a remote MCP endpoint with
accounts, so an agent inside Claude can use it without a machine of its own.
Add the URL as a custom connector (Pro, Max, Team and Enterprise plans), sign
in with your email when Claude asks, and every call from then on carries
your account.
What the hosted server is, in one breath: Polymarket market data, the signal
tools, check_order, and paper trading with a $1,000 virtual bankroll per
account, filled against the live book. What it is not: a place where money
moves. It holds no wallet keys and executes no real order. Account-scoped
tools such as
open_orders and the gasless relayer tools are absent, because on a shared
server they would describe nobody's account. Twenty-four tools remain: the
public-data and paper tools plus arena_register, arena_unregister and
arena_status, which put the account's paper ledger on the public board.
Live trading from Claude stays self-hosted: pip install oddsrail with your own
key, and the same place_order posts real orders when you set
ODDSRAIL_DRY_RUN=0. The hosted runner is a separate service.
The paper ledger you build up in Claude is yours to reset with
paper_reset; nothing else about the account exists. Privacy policy:
oddsrail.app/privacy. This repository holds
the MCP server; oddsrail/hosted.py lists what the hosted profile removes.
The hosted service and the website are developed in a separate, private
repository.
Builder page and competition
Build your agent by choosing strategies, markets and risk limits. The final step saves the draft and opens Claude Desktop setup, with the installer and a strategy brief that preserves your settings.
The brief asks Claude for read-only review. It does not install a strategy, start an autonomous runner or enforce the builder's requested daily-loss, exposure or market rules. Local live trading needs separate account setup, operator guards and explicit authorization. Hosted agent activation needs a session key for a Deposit Wallet and, during the pilot, an allowlisted owner.
The agent competition is coming soon. Entries are not open. The proposed format uses equal starting capital and a shared set of markets; dates and final rules will be published before entries open.
How attribution works (CLOB V2, verified Aug 2026)
Get your builder code (a bytes32) at polymarket.com → Settings → Builders. Set your fee rates there: taker up to 100 bps, maker up to 50 bps, additive on top of platform fees, settled to your builder wallet.
export ODDSRAIL_BUILDER_CODE=0x...where the server runs.Every order any agent routes through
place_orderhas the code placed in the V2 order struct'sbuilderfield before signing, so attribution is on-chain, visible in everyOrderFilledevent on CTF Exchange V2.Verify with the
builder_statstool (public builder-trades endpoint + leaderboard).
If you skip this, orders carry the bundled oddsrail builder code
(0xa576c5ce…, registered at 0 bps maker / 0 bps taker), costing you nothing
and funding the project. If you set your own, yours wins; the default is a
default, not a lock-in.
The oddsrail builder profile is Verified in Polymarket's builder program (2026-09-02), and Polymarket's builder team confirmed builder-code attribution as the right pattern for a self-hosted, non-custodial tool: no keys ship with the product, and the code is attached and signed by the operator's own wallet.
Environment variables
Variable | Default | Meaning |
|
|
|
| project default | Your bytes32 builder code. Overrides the bundled project default so attribution (and any reward-pool share) accrues to you instead. |
| unset | Operator wallet key; required only for real trading. Never leaves this machine. |
| unset | Proxy/deposit wallet address, if the account uses one. |
| unset | Your own Relayer API key (polymarket.com → Settings → Relayer API keys), for gasless |
| unset | The address the relayer key was issued for. Both halves are required; without them the gasless tools send nothing. |
| unset | Guardrail: max USDC notional per order. Enforced before any request, in dry-run too. |
| unset | Guardrail: max cumulative notional of live orders submitted by this server process. |
| unset | Guardrail: max resting orders on the account (live; checked against the venue before placing). |
| unset | Guardrail: comma-separated allowed market identifiers. For Polymarket, use exact outcome token IDs. Anything else is refused. |
|
| Paper-trade dry-run Polymarket orders against the live book. |
|
| Where the paper ledger lives. One local JSON file. |
|
| Starting paper cash in USDC. |
Status
Tests: the 0.18.1 release and its test-only follow-up passed 1,356 Python tests on Python 3.11, 3.12 and 3.13, plus 119 JavaScript tests. Coverage includes spending limits, concurrent submissions, order reconciliation, book walking, sizing, dry-run behavior, wallet authentication and the builder handoff. See the testing guide for the locked environment and commands.
Release verification: the installer launch command installed the published PyPI package and completed a real MCP session. Default and unexpected mode settings stayed in simulation. No funded orders were placed during that release check; installation through Claude Desktop's interface and funded trading still need testing.
Earlier Polymarket integration checks, including attributed orders and gasless position management, are recorded in the live verification notes.
Where this works
Two different things can stop oddsrail from trading, and they have opposite remedies. One is a venue restriction, enforced at the order. The other is a network filter, which breaks the connection itself.
Polymarket restrictions. Polymarket publishes its restricted-jurisdiction list as an API reference: https://docs.polymarket.com/api-reference/geoblock. There are three tiers. OFAC-sanctioned jurisdictions (Iran, Syria, Cuba, North Korea, and the Crimea, Donetsk and Luhansk regions of Ukraine) are blocked on both the frontend and the API, with no new orders and no closing of existing positions. A longer second tier is close-only on both the frontend and the API: existing positions can be closed, new ones cannot be opened. It includes the United States, the United Kingdom, France, Germany, Italy, Poland, Slovakia, Belgium, Singapore, Australia, New Zealand, Brazil, Russia, Taiwan, Thailand and the Canadian provinces of Ontario, Quebec, British Columbia and Alberta. A third group, Ireland, Japan, Malta (sports only) and the Netherlands, is close-only on Polymarket's frontend, with the API explicitly not restricted.
Note the shape of that failure: it lands on the order, not the connection. Public reads answer normally, so oddsrail will look like it is working right up until an order is rejected. Verified against Polymarket's documentation on 2026-08-31; Polymarket updates the list without notice, so read the URL rather than this paragraph.
The United States. polymarket.com, the venue oddsrail talks to, is close-only for the US. Polymarket separately operates Polymarket US (polymarket.us), run by QCX LLC as a CFTC-regulated Designated Contract Market. oddsrail does not support it. It is a different API host, a different authentication model (API-key headers rather than EIP-712 wallet signatures), a different SDK and a different funding rail. A polymarket.us account and its keys will not work with this server.
Network filters. Separately from any venue rule, a national filter can block the domains outright. Turkey does this: Polymarket does not restrict Turkey, but Turkish ISPs block polymarket.com. That is a connectivity problem, not an eligibility one, and it looks different: DNS failures, TLS errors, resets, or an ISP interstitial page served where JSON was expected. oddsrail classifies both shapes and tells the calling agent which one it hit.
Eligibility is the operator's, not the tool's. oddsrail is self-hosted
and non-custodial, which is a real advantage and also means you hold the
account and you make the venue's representations; there is no intermediary
making them for you. Polymarket's trading flow requires an attestation that
you are not a U.S. person, are not located in a restricted jurisdiction, and
are not "using a VPN or other measures to circumvent or attempt to
circumvent" restrictions, and states that Polymarket reserves the right to
put a non-compliant wallet in close-only mode. server_info reports
Polymarket's geoblock verdict for this machine's IP, but a technical probe is
not a compliance
check: the terms bind on residence, citizenship and incorporation, not on
egress IP. Read the terms; if any of this matters to you, get your own legal
advice. Nothing here is legal advice.
The signal logic, the MCP layer and the whole test suite run fine offline regardless.
Guardrails: limits the agent cannot argue with
Anyone handing keys to an agent wants three things first: a cap on one order, a cap on a session, and a fence around which markets it may touch. All three are operator-set environment variables (table above), enforced before any request goes out, in dry-run as well as live, so the agent meets the fence in rehearsal. A refusal is a structured answer that names the rule, the limit and the request:
{"accepted": false, "blocked_by": "guardrail", "rule": "max_order_notional",
"limit": 25.0, "requested": 99.5, "note": "refused by an operator-set guardrail ... Nothing was sent."}The session counter lives in the server process; restarting it resets the
budget, which is the operator's call. server_info reports the active limits
and how much of the session budget is used.
Paper trading: dry-run with a memory
By default, every dry-run Polymarket order is filled against the live
order book, walked within the limit price; whatever does not fill rests as a
paper order and fills later if the market crosses it. paper_positions
reports cash, positions at current marks, realized and unrealized P&L and the
resting paper orders; paper_reset starts over. The ledger is one local JSON
file. Be clear about what this is: fills assume no queue position, no latency,
no market impact and no fees, so paper results are an upper bound on the same
strategy live.
Realtime: watch the book move
watch_book(token_id, seconds, max_events) subscribes to a token's realtime
stream and returns the events that arrived (book snapshot, then price changes
and trades), bounded to at most 60 seconds so an agent cannot hang a session
on a quiet market. Use it after get_orderbook when the decision depends on
the book moving, not just where it is.
If the stream fails with CERTIFICATE_VERIFY_FAILED while the REST tools
work, your Python has no CA bundle (common with python.org macOS installs).
oddsrail classifies that as local_tls and tells the agent the fix: run
Install Certificates.command from the Python folder in /Applications, or
set SSL_CERT_FILE to the path printed by python -m certifi.
Gasless position management (relayer)
Three tools move collateral without paying gas, through Polymarket's relayer:
split_position (USDC → a full YES+NO set), merge_positions (matching
YES+NO → USDC, or max), and redeem_positions (a resolved market's winning
shares → USDC). All three respect dry-run and return the relayer transaction
id and hash plus the terminal outcome.
They use your own Relayer API key, created at polymarket.com → Settings →
Relayer API keys and exported as POLYMARKET_RELAYER_API_KEY +
POLYMARKET_RELAYER_API_KEY_ADDRESS. That is the pattern Polymarket's builder
team recommends for a self-hosted tool: no builder secret ships with oddsrail,
and each operator authenticates the relayer as themselves. Relayer limits are
per builder tier: 100 requests/day unverified, 10,000 verified. Without the
key the tools return a structured "not configured" answer and send nothing;
they never fall back to a gas-paying broadcast from the signer.
Exercised live (2026-09-02): a 1 USDC split and the matching merge went
through the relayer from this code, gasless, on the maintainer's test account
with its own Relayer API key. Relayer ids and Polygon transaction hashes are
in docs/live-proof.md. redeem_positions is still
unproven live: it needs a resolved market with winning shares, which that
account has not held yet. redeemable_positions lists what the configured
wallet could redeem or merge right now, and the settle_resolved prompt
chains the two.
Market discovery and trade costs
find_markets(query, venues="polymarket"): searches Polymarket and returns normalized market IDs, titles, outcome prices, bid/ask prices, spread, volume and close time. Set the venue explicitly for this workflow.quote_cost("polymarket", market_id, side, size): walks the order book for the requested share size. Returns average fill price, slippage, notional, levels consumed, whether the size is fillable, and the market's fee schedule where published.
Order lifecycle & discovery
order_status(order_id): resting / partially_filled / filled / gone, with size_matched. The answer an agent needs after place_order.my_fills(),my_positions(): the operator's executions and holdings, no address juggling. (Fills come from the Data API activity feed; the SDK's list_account_trades returns the market's public tape and is not used.)cancel_all_orders(): kill switch, flattens every resting order at once.resolution_criteria(venue, market_id)returns the full resolution contract: what resolves YES, who resolves it, from which sources. Read it before trusting a price.closing_soon(hours, venues="polymarket"): Polymarket markets closing within N hours, where activity concentrates.
Workflow prompts
MCP prompts show up in clients as ready-made workflows, and they encode the order of operations that keeps an agent out of trouble; the sequencing is the expertise, which a flat tool list cannot convey.
/find_fade_setup(query, bankroll): signal → book → cost → resolution → size → dry-run, with the rejection criteria at each step/daily_review: positions, resting orders, fills, closing-soon, attribution
Pre-trade checks and sizing
check_order(venue, market_id, side, price, size, intent): the last step beforeplace_order. Deterministic checks of the proposed order against the operator's own words and the live market: does the market exist and accept orders, do the intent's words match the market and the YES/NO side, is the price sane against the book, is the size above Polymarket's $1 minimum and inside the guardrails, is there liquidity within the limit, is a resolution source named. Returnsok/caution/blockwith the evidence per check and a one-line read-back. No second model judges anything; nothing is sent.position_size(bankroll_usd, price, fair_value): fractional-Kelly sizing, capped, refusing negative-edge bets, returning its own assumptions.
Polymarket tool highlights
search_markets,get_market,get_orderbook,price_history,get_positions: read-only, no keysovershoot_signal, premium: fresh panic-jump detection + this market's historical reversion tendency (ported from the polymarket-wc analyzer)dispute_risk, premium: transparent 0–100 heuristic for contested (UMA-dispute-prone) resolutionsplace_order,cancel_order,open_orders: trading, dry-run by default.priceis a probability in (0,1),sizeis in SHARES, and the exchange enforces a $1 minimum notional on marketable orders. Trading tools carrydestructiveHintannotations so clients can gate them.builder_stats: attribution verification + public builder leaderboardfind_markets,quote_cost: market discovery and trade costs (above)server_info: server mode, configured credentials and operator guardrails
Stack notes
Official unified SDK
polymarket-client(0.6.x):AsyncPublicClientfor data,AsyncSecureClient.place_limit_order(..., builder_code=...)for attributed orders. The legacypy-clob-clientis archived and cannot attach builder codes. Do not use it.MCP SDK 2.0:
MCPServerfrommcp.server.mcpserver(the oldmcp.server.fastmcp.FastMCPimport is gone in 2.x).x402 (planned): the official
x402PyPI package (v2.20+) can wrap MCP tools directly (x402.mcp, payment rides in tool-call_meta), but its MCP helpers currently target mcp 1.x, so integrating means pinningmcp>=1.28,<2or waiting for the 2.x-compatible release. Mainnet settlement needs a facilitator (Coinbase CDP: 1,000 free settlements/mo, then $0.001). Keep free tiers of both signals so registries can index the server.
Who this is for
Polymarket's public builder leaderboard shows what a single operator routing
their own flow is worth. Pulled 2026-08-31 via this server's own
builder_stats tool. Re-run it, the numbers move:
weekly volume | |
#1 (traderline) | $7.70M |
median of top 25 | $533K |
entry to top 25 | $140K |
The instructive rows are the small ones: MagicMarkets routes $901K/week with a single active user; Jupiter $515K with one; Sharkbetting $1.15M with two. Those are bot operators routing their own flow, which is exactly who this is built for.
Roadmap
Live smoke test from an unblocked network: done 2026-08-23, all tools passRegister builder code (polymarket.com → Settings → Builders), set fees to 0 bps at launch, export
ODDSRAIL_BUILDER_CODE; first attributed order on a tiny sizex402 paid wrapping for the two signals once the mcp-2.x conflict clears
Available Tools
42 toolsattribution_ledgerARead-onlyIdempotentInspect
Attribution ledger for the builder code in use: every trade carrying it, aggregated per Sunday-start week and per wallet, with the maintainer's own wallets split out into an honest 'external' line. Built from Polymarket's own public builder feed, so any figure it reports can be checked at data-api.polymarket.com/v1/builders/leaderboard.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false, so the description does not need to restate safety. It adds meaningful behavioral context: data is aggregated per Sunday-start week and per wallet, maintainer wallets are split into an 'external' line, and the source is Polymarket's public builder feed, making every figure externally verifiable at the given URL. This goes beyond the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no filler. The first sentence front-loads what the ledger contains and the aggregation granularity, while the second provides source and verifiability. Every clause earns its place, and the density is appropriate for a zero-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-argument, read-only, idempotent tool, the description is nearly complete: it states scope, aggregation, wallet treatment, source, and a verification endpoint. The only minor gap is the absence of any explicit output format or field listing, since there is no output schema to carry that information, but the ledger concept and contents are described well enough for an agent to understand what the tool returns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters of any kind, so the baseline is 4: the parameter semantics dimension is essentially not applicableable, and the description correctly spends no effort describing parameters. It does not need to compensate for any schema coverage gap because there are no parameters to document.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as an attribution ledger for a specific builder code, stating what it contains: every trade carrying the code, aggregated by Sunday-start week and wallet, with maintainer wallets split out. It is more than a tautology and gives a concrete resource, though it does not explicitly differentiate itself from the closely related builder_stats sibling or name any alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description tells what the ledger is and how it is built, but it never states when to use this tool versus the many sibling market, order, or stats tools. There is no explicit when-to-use, when-not-to-use, or alternative routing guidance, leaving the agent to infer the appropriate context from the tool's purpose alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
builder_statsARead-onlyIdempotentInspect
Builder attribution stats: the public builder leaderboard, and (if ODDSRAIL_BUILDER_CODE is set) matched trades attributed to this operator's code.
| Name | Required | Description | Default |
|---|---|---|---|
| time_period | No | WEEK |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, and non-destructive behavior. The description adds useful context beyond those annotations: the leaderboard is public, and the response varies based on an environment variable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence that front-loads the tool's purpose and packs the conditional behavior into a parenthetical. There is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The core selection logic is present, and the tool is simple with one optional parameter. However, the description omits the response shape and any guidance on time_period values, which matters because no output schema is provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, time_period, has no schema description and is not mentioned in the tool description. With 0% schema description coverage, the description should compensate by explaining accepted values or the effect of the parameter, but it does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a concrete deliverable: the public builder leaderboard plus operator-matched trades when an environment variable is set. It lacks a strong imperative verb but is specific enough to distinguish the tool's core data resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides a clear conditional context by stating that matched trades are only included if ODDSRAIL_BUILDER_CODE is set. It does not explicitly name alternatives or exclusions, but the intended use case is reasonably clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cancel_all_ordersADestructiveInspect
KILL SWITCH — cancel every resting Polymarket order on the operator account at once. Use when exposure must go to zero fast. Respects dry-run.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is known. The description adds genuinely non-obvious behavior beyond annotations: the 'KILL SWITCH' all-at-once semantics and the 'Respects dry-run' caveat. No contradiction with annotations — the destructive framing aligns with destructiveHint=true.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short clauses — what it does, when to use it, and the dry-run caveat — with zero filler. The verb and resource are front-loaded before the usage guidance, and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with annotations covering the destructive/read-only profile, the description covers what it does, when to use it, and a key behavioral caveat. It could be a 5 with an explicit pointer to single-order cancel siblings, but nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters (empty schema), so the baseline of 4 applies — there is no parameter meaning for the description to add or clarify. The description correctly focuses on behavior rather than invented parameter detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource — 'cancel every resting Polymarket order on the operator account' — with a clear scope ('at once'). The bulk 'every...at once' framing distinguishes it from single-order siblings like kalshi_cancel_order and cancel_order without needing to open their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use when exposure must go to zero fast' provides an explicit trigger condition for the emergency bulk-cancel scenario. However, it doesn't name the alternative for single-order cancellation (cancel_order / kalshi_cancel_order) or state when not to use it, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cancel_orderBDestructiveInspect
Cancel an open order by id (respects dry-run).
| Name | Required | Description | Default |
|---|---|---|---|
| order_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the description isn't required to restate those. It adds the 'respects dry-run' nuance, which is useful, but it doesn't disclose behavior for invalid IDs, already-canceled orders, or error responses. With annotations covering the destructive nature, the additional context earns a middle score.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, direct sentence that front-loads the action and resource. No filler or redundancy. It earns full marks for conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive operation with only one parameter and no output schema, the description is too thin. It lacks critical context about the dry-run behavior (what does 'respects' mean?), error handling, and whether the operation is reversible. While annotations cover some safety aspects, the agent is left without enough information to call this tool confidently, especially with many similar siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the bare order_id parameter. It says 'by id', which clarifies the parameter's role, and 'open order' suggests the ID must refer to an open order. However, it doesn't specify the ID format, where to obtain it, or what happens if the ID is invalid. Minimal value added.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (cancel) and resource (an open order by id). It distinguishes from cancel_all_orders by implying a single order, but it doesn't differentiate from the sibling kalshi_cancel_order. Still, the core purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool vs alternatives like cancel_all_orders or kalshi_cancel_order. The phrase 'open order' implies a condition, but there's no statement about when not to use it or which sibling to choose. Usage is only implied, not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_orderARead-onlyIdempotentInspect
CHECK BEFORE YOU PLACE: deterministic verification of a proposed order against the operator's intent, the live market and the guardrails. Pass the operator's own words as intent. Returns ok / caution / block with the evidence per check (market exists and is open, intent matches the market and the YES/NO side, price is sane vs the book, size and $1 minimum, guardrails, liquidity, resolution source) plus a one-line read-back. No model judges anything; nothing is sent.
| Name | Required | Description | Default |
|---|---|---|---|
| side | Yes | ||
| size | Yes | ||
| price | Yes | ||
| venue | Yes | ||
| intent | No | ||
| outcome | No | ||
| market_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds valuable disclosure beyond that: it states the check is deterministic, that 'No model judges anything', and that 'nothing is sent'. It also describes the ok/caution/block result and per-check evidence, giving the agent a clear behavioral model.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the purpose and usage context, and every sentence contributes. The long parenthetical list of checks is dense but informative, with no filler. It could be improved with bullets, but the structure is still effective.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 7 parameters and no output schema, the description carries substantial burden, and it largely delivers: it enumerates the return verdicts, the evidence checks, the role of intent, and the no-side-effects guarantee. The main gaps are the semantics of `venue` and `outcome`, and how an agent should source `market_id`, though sibling market tools likely fill that need.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explicitly explains `intent` and indirectly gives meaning to `side` ('YES/NO side'), `price` ('sane vs the book'), `size` ('$1 minimum'), and `market_id` ('market exists and is open'). However, `venue` and `outcome` are never defined, and no exact value constraints or formats are provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description leads with 'CHECK BEFORE YOU PLACE: deterministic verification of a proposed order against the operator's intent, the live market and the guardrails.' This names a specific verb and resource, and clearly distinguishes the tool from order placement or market lookup. It also states the return verdicts, making its role unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'CHECK BEFORE YOU PLACE' explicitly signals when the tool should be used, and 'Pass the operator's own words as intent' gives a concrete invocation instruction. It does not explicitly list sibling alternatives or exclusions, but the before-placing context is clear enough to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
closing_soonARead-onlyIdempotentInspect
Markets closing within N hours on either venue, by volume — where trading activity concentrates.
| Name | Required | Description | Default |
|---|---|---|---|
| hours | No | ||
| limit | No | ||
| venues | No | both |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, so safety is covered. The description adds useful behavioral context beyond annotations: it scopes results to both venues and indicates ordering by volume, giving the agent a clear model of what the tool returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. Every phrase conveys meaning, and the dash clause adds useful context about why volume matters without bloating the description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple query tool with rich annotations and only three optional parameters, the description covers the core semantics: time window, venue scope, and ordering. The limit parameter is left implicit, and there is no output schema, but the defaults in the schema mitigate the gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does clarify 'hours' via 'within N hours' and 'venues' via 'either venue', but 'limit' is never explained and no parameter names are explicitly connected to their meanings. The 'by volume' clause partially clarifies ordering but not the limit parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the resource ('markets closing within N hours') and the selection/sort criterion ('by volume'). It is specific enough to distinguish the tool from broad search tools, though it does not explicitly name a sibling or state 'list/find' as the verb.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context is implied: use this when you want markets that are closing soon and sorted by trading activity. However, there is no explicit guidance about when not to use it or which sibling tools (e.g., find_markets, search_markets) are better suited for other discovery needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_venuesARead-onlyIdempotentInspect
Find markets that may be the SAME event on both Polymarket and Kalshi. NOT an arbitrage scanner: matches are candidates from title similarity plus a close-date check, and a price difference between two candidates is not profit. Identical wording does not mean identical resolution criteria — read both, and run quote_cost on each leg, before acting.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| min_similarity | No | ||
| max_close_days_apart | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and non-destructive behavior. The description adds useful behavioral context beyond this: the tool returns candidates based on title similarity and close-date, not verified matches, and identical wording does not imply identical resolution criteria. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all purposeful and front-loaded: the primary purpose first, followed by a critical negative warning and an actionable caveat. There is no fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a multi-venue comparison tool with no output schema and zero parameter descriptions, the description conveys the algorithm, the risk, and the recommended next steps. It is mostly sufficient, though it could more explicitly describe the expected output and each parameter's role.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It hints at the meaning of min_similarity and max_close_days_apart via 'title similarity plus a close-date check,' but it does not explicitly explain the query parameter or the limit parameter. This is partial compensation, not full.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Find markets that may be the SAME event on both Polymarket and Kalshi.' It also distinguishes itself from an arbitrage scanner, which helps differentiate it from related trading/quoting tools like quote_cost.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says what the tool is NOT ('NOT an arbitrage scanner') and gives a clear condition for correct use: matches are candidates, price differences are not profit, and users should read both resolution criteria and run quote_cost on each leg before acting. This provides actionable when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dispute_riskARead-onlyIdempotentInspect
PREMIUM SIGNAL — dispute-risk triage. Scores 0-100 how likely a market's resolution gets contested (UMA dispute risk) with transparent reasons. Takes the market slug, the Gamma id, or a CLOB token id (the market_id a search returns).
| Name | Required | Description | Default |
|---|---|---|---|
| id_or_slug | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral context by specifying the output scale (0-100) and the presence of 'transparent reasons', which an agent would not know from the schema or annotations alone. It does not contradict any annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no fluff. The key purpose and output scale are front-loaded, and the input flexibility is stated in the second sentence. Every phrase adds value, and the 'PREMIUM SIGNAL' label is a compact attention marker rather than a distraction.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, read-only tool without an output schema, the description is quite complete: it explains what the tool does, the input format, and the kind of output (score plus reasons). It omits details about the exact response structure or any preconditions, but given the tool's simplicity and annotation coverage, the remaining gap is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the full burden of explaining the parameter. It does this well by stating that the single parameter can be a market slug, Gamma id, or CLOB token id, and clarifies the last as 'the market_id a search returns'. This gives the agent concrete sense of what values to pass, though it could include an explicit example format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('scores 0-100 how likely a market's resolution gets contested') and a resource ('dispute-risk triage'), making the tool's purpose immediately clear. It distinguishes itself from sibling tools by focusing on UMA dispute risk, which no other sibling mentions. The 'PREMIUM SIGNAL' label adds context without obscuring the core function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance on when to use this tool versus alternatives or when not to use it. It implies it can be used after a search (by mentioning the 'market_id a search returns') but does not compare with related tools like resolution_criteria or settlement_audit. The agent is left to infer the appropriate context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_marketsARead-onlyIdempotentInspect
Search BOTH Polymarket and Kalshi at once and return one normalised shape per market: venue, market_id (the id that venue's place_order takes), title, yes/no price as probabilities in (0,1), best bid/ask, spread, 24h volume, close time. Use this instead of the per-venue search tools when you do not already know the venue.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| venues | No | both |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, covering the safety profile. The description adds value beyond that by disclosing the cross-venue aggregation behavior and, importantly, that market_id is 'the id that venue's place_order takes' — a contract detail that directly affects downstream tool selection. It doesn't discuss rate limits or edge cases, but the annotated safety profile lowers the burden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, each earning its place: the first front-loads the core action and return contract, the second delivers the routing rule. The inline field enumeration is lengthy but justified given there is no output schema. Slightly more compact phrasing was possible, but nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search tool with strong annotations and no output schema, the description compensates well by enumerating the return fields and the place_order id contract. However, the total absence of parameter semantics (query, limit, venues) and no mention of pagination or result-limiting behavior leave meaningful gaps for an agent attempting correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden of explaining query, limit, and venues — but it never names or explains any of them. The 'BOTH Polymarket and Kalshi' phrasing hints at the venues parameter's purpose and the default suggests valid values, but there is no specification of query syntax, limit behavior, or venues value options. The description does not compensate for the absent schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Search BOTH Polymarket and Kalshi at once', which immediately distinguishes it from per-venue siblings like kalshi_search_markets and search_markets. It further strengthens clarity by enumerating the full normalized return shape, so an agent knows exactly what it receives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The final sentence gives an explicit routing rule: 'Use this instead of the per-venue search tools when you do not already know the venue.' This names the alternative tool class and the exact deciding condition, and by implication states when not to use it (when the venue is known). No inference is required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_marketARead-onlyIdempotentInspect
Get one market's details. Accepts the market slug, the Gamma market id, or either of its CLOB token ids (the market_id that find_markets and search_markets return), so a token id from a search can be passed straight in.
| Name | Required | Description | Default |
|---|---|---|---|
| id_or_slug | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (read-only, idempotent, non-destructive), so the bar is lower. The description adds context about accepted identifier forms, which is useful, but it does not disclose what 'details' includes or how the response is shaped. This is adequate but not rich behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the core action front-loaded and the input semantics unfolded efficiently. No filler, redundancy, or unnecessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, read-only tool, the description fully covers the parameter semantics, and annotations cover behavioral safety. The only real omission is the unspecified content of 'details' since there is no output schema, leaving the agent to infer the response shape — a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, leaving the description to fully explain the id_or_slug parameter. It does so excellently by enumerating the three accepted forms — market slug, Gamma market id, and CLOB token ids — and tying the last one to find_markets/search_markets output.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Get') and resource ('one market's details'), clearly distinguishing it as a single-market lookup versus list-returning tools like find_markets and search_markets. The mention of accepted identifier forms adds precision without confusion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context by explaining that token ids from find_markets and search_markets can be passed straight in, implying when to use this tool. However, it never explicitly names alternatives like get_orderbook or states when not to use this tool, so the guidance is implied rather than fully explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_orderbookBRead-onlyIdempotentInspect
Get the live orderbook (bids/asks) for a CLOB token id.
| Name | Required | Description | Default |
|---|---|---|---|
| token_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the description is not required to cover safety. It adds that the data is 'live' and includes bids/asks, which gives some behavioral context about the response. However, it doesn't disclose any rate limits, pagination, or data staleness behavior, so it's only moderate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one sentence, front-loads the verb and resource, and contains no unnecessary words. It is an excellent example of conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (1 param, no output schema) and that annotations cover safety, the description is generally adequate. But it lacks guidance on the expected return structure beyond 'bids/asks' and fails to mention any limitations or caveats of live data. It also doesn't help disambiguate from sibling orderbook tools, so there are clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate for the token_id parameter. It adds that the parameter is a 'CLOB token id', clarifying its domain. However, it doesn't explain the format, where to find the id, or how it relates to other tool parameters. This is minimal compensation, so a 2 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool gets the live orderbook for a CLOB token id, specifying the resource (orderbook) and the key qualifier (live, CLOB token id). However, it does not differentiate from sibling tools like kalshi_get_orderbook or watch_book, so it scores a 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives, such as kalshi_get_orderbook or price_history. There are no conditions, exclusions, or alternative suggestions. The only implied usage is when an orderbook is needed, but that is not explicit enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_positionsCRead-onlyIdempotentInspect
Current positions for a wallet address.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| address | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds no behavioral context beyond what annotations already provide. readOnlyHint, openWorldHint, idempotentHint, and destructiveHint are all set, and the description simply restates the core function without explaining details like pagination, limit behavior, or what 'current' means. It does not contradict annotations but contributes no extra transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that is easy to parse and front-loaded with the core purpose. However, it is so brief that it sacrifices necessary detail, but that is more a completeness issue than a structure issue. The structure itself is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the large number of sibling tools that deal with positions, the lack of any disambiguation or mention of the venue (e.g., Kalshi, paper, etc.) makes this definition incomplete. There is no output schema, so the description should clarify return format or limit semantics, but it does not. An agent may call the wrong tool without further context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions 'wallet address' which clarifies the address parameter, but it does not explain the 'limit' parameter at all. The agent is left guessing what limit controls, such as max number of positions returned. This is a significant gap given the 0% coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states that the tool returns current positions for a wallet address, which is a clear resource and scope. However, it does not differentiate from sibling tools like my_positions, kalshi_positions, or paper_positions, all of which may also return positions. The purpose is clear but lacks sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. There is no mention of exclusions, preferred use cases, or why one would choose this over my_positions or kalshi_positions. The single sentence gives no context for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kalshi_balanceARead-onlyIdempotentInspect
Kalshi account balance (needs the operator's API key).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds the API key requirement, which is useful operational context beyond the structured data. It does not describe return format or units, but for a simple zero-parameter read this is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler. It front-loads the resource being accessed and immediately states the key requirement. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only tool, the description is nearly sufficient. It states the resource and the authentication requirement, while annotations cover the safety profile. The exact meaning of 'balance' (e.g., cash vs. available equity) is left open, but this is acceptable given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema carries no burden and the description does not need to explain parameters. The baseline for a parameter-less tool is 4, and there is nothing to deduct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource (Kalshi account balance) and the implied operation is retrieval. It lacks a specific verb like 'get' or 'fetch', but the noun phrase is unambiguous and not a tautology. However, it doesn't differentiate from the sibling 'my_balance', which could be confused with this tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear prerequisite ('needs the operator's API key'), implying it should be used when the operator's credentials are available. It does not mention any alternative tools or state when not to use it, so guidance is only implied. This is adequate but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kalshi_cancel_orderADestructiveInspect
Cancel a Kalshi order by id (respects dry-run).
| Name | Required | Description | Default |
|---|---|---|---|
| order_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds the 'respects dry-run' behavior, which is not present in the annotations (destructiveHint=true, readOnlyHint=false). This is valuable context beyond the structured data. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with zero filler. It states the action and key behavior efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with no output schema, the description is adequate. It covers the action and a key behavioral nuance (dry-run), though it omits potential error scenarios or success outcomes, which are not critical for this simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description mentions 'by id' which maps to the order_id parameter, but does not elaborate on format or semantics. Schema coverage is 0%, so the description partially compensates but leaves room for clarification about what the id refers to.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Cancel'), the resource ('Kalshi order'), and the identifier ('by id'). It differentiates from the generic 'cancel_order' sibling by explicitly naming Kalshi, and from 'cancel_all_orders' by specifying a single order.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly implies when to use it (to cancel a specific Kalshi order) and implicitly distinguishes from cancel-all. However, it does not explicitly mention alternatives or when not to use it, but the context is clear enough for an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kalshi_get_marketBRead-onlyIdempotentInspect
Get one Kalshi market by ticker.
| Name | Required | Description | Default |
|---|---|---|---|
| ticker | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnly=true, idempotent=true, and non-destructive behavior. The description adds minimal scoping detail by saying the lookup is by ticker and returns a single market, but it does not disclose edge-case behavior like unknown-ticker handling or market availability. This is acceptable for a simple read-only getter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence contains exactly the information needed to identify the operation. There is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, read-only lookup tool with supportive annotations, the description is largely sufficient to invoke the tool correctly. It does not specify the return shape or ticker format, but those are minor gaps given the low complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only restates that `ticker` is the lookup key. It provides no ticker format, example, or clarification of what qualifies as a valid Kalshi ticker, so it does not compensate for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Get'), a specific resource ('one Kalshi market'), and the lookup key ('by ticker'), making the core operation unambiguous. It does not differentiate from the sibling `get_market`, but the behavior itself is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives such as `kalshi_search_markets`, `find_markets`, or the similarly named `get_market`. The only implied signal is that the user already has a specific ticker.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kalshi_get_orderbookARead-onlyIdempotentInspect
Kalshi orderbook for a ticker, normalised to a YES-book bid/ask view (Kalshi publishes bid ladders only; asks are derived as 1 - NO bid). Raw ladders included.
| Name | Required | Description | Default |
|---|---|---|---|
| depth | No | ||
| ticker | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint false. The description adds substantial behavioral detail beyond that: normalization to a YES-book bid/ask view, the fact that Kalshi publishes bid ladders only, asks derived as 1 - NO bid, and raw ladders included. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no filler. It front-loads the core purpose and then adds the critical normalization detail. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description explains the core normalization behavior well, but with no output schema it does not describe the return structure, and it omits depth semantics. Given the strong annotations and read-only nature, it is adequate but leaves meaningful gaps for an agent invoking the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies the ticker parameter ('for a ticker') but says nothing about depth, its default, or how it affects the orderbook. This leaves an important parameter semantically undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource ('Kalshi orderbook for a ticker') and adds valuable specificity about the normalized YES-book view and derived asks. However, it does not explicitly differentiate this from sibling orderbook tools like watch_book or the generic get_orderbook.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance about when to use this tool versus alternatives, and it does not mention any exclusions or sibling tools. The only implied usage context is that it is Kalshi-specific.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kalshi_get_tradesBRead-onlyIdempotentInspect
Recent public trades for a Kalshi ticker.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| ticker | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, establishing the safe read-only behavior. The description adds useful context that the trades are 'public' and 'recent,' but it does not describe return format, pagination, rate limits, or any other operational behavior beyond what the annotations already cover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no wasted words, with the core resource front-loaded. It is appropriately short for a simple tool, though it omits enough detail that the structure alone does not carry the full explanatory burden.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a simple two-parameter read-only tool with rich annotations, and the description states the essential entity being fetched. However, with no output schema and no parameter-level guidance, the description leaves reasonable gaps around return values and limit semantics that an agent would need for precise invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only adds meaning for the ticker parameter: it identifies trades 'for a Kalshi ticker.' The limit parameter is not explained at all, leaving its semantics and default behavior undocumented outside the schema's default value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource ('recent public trades for a Kalshi ticker') and specifies the action of retrieving them. It is clear enough to distinguish from most siblings like kalshi_balance or kalshi_positions, but it does not explicitly name or differentiate against similar data-access tools such as my_fills or price_history.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus alternatives such as my_fills, price_history, or kalshi_get_orderbook. The intended use case is vaguely implied by the name, but the description does not state exclusions, prerequisites, or context for choosing this tool over its siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kalshi_open_ordersDRead-onlyIdempotentInspect
Kalshi resting orders (needs the operator's API key).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds a useful requirement (operator's API key) that is not in the annotations. However, it does not disclose other behavior such as pagination, return format, or error conditions, so it adds minimal context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely short (a single phrase) and could be considered concise, but it is under-specified. It lacks a clear subject-verb structure and does not front-load key information. It reads more like a label than a functional description, so it is not effectively structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with one optional parameter and no output schema, the description should at least clarify what the tool returns and how the limit works. It does neither, only stating the resource and a requirement. The description is incomplete for an agent to correctly understand the tool's behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, and the description does not mention the 'limit' parameter at all. With no parameter semantics provided, an agent has no guidance on what limit controls (e.g., number of orders returned, pagination size). The description fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Kalshi resting orders' is a noun phrase without a verb, so it does not explicitly state whether the tool lists, retrieves, or manages resting orders. It also fails to distinguish this from sibling tools like kalshi_positions or open_orders, which likely have similar purposes. The only action implied is 'get', but it is not stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It only mentions the API key requirement, which is a prerequisite, not a usage condition. There is no comparison to siblings or indication of specific scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kalshi_place_orderBDestructiveInspect
Place a Kalshi limit order. State it naturally: outcome yes|no, action buy|sell, price = probability of THAT outcome in (0,1). Translated to Kalshi's YES-book bid/ask internally. DRY-RUN by default.
| Name | Required | Description | Default |
|---|---|---|---|
| count | Yes | ||
| price | Yes | ||
| action | Yes | ||
| ticker | Yes | ||
| outcome | Yes | ||
| time_in_force | No | good_till_cancelled |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructive and non-idempotent behavior. The description adds valuable context: the order is a limit order, prices are probabilities, the request is internally translated to Kalshi's YES-book, and it is DRY-RUN by default. It does not disclose the full side effects of a live order or how dry-run is toggled, but the added behavioral detail is meaningful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded, with each sentence adding a distinct piece of information: the action, the natural-language mapping, the internal translation, and the safety default. The phrase 'State it naturally' is slightly vague, but overall the structure is efficient and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a trading tool with six parameters, no output schema, and no schema-level descriptions, the description omits key context: the response format, how to actually place a live order versus dry-run, the impact on positions or balance, and the meaning of time_in_force. It is enough to make a basic call but not fully complete for safe autonomous use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate. It explains outcome (yes|no), action (buy|sell), and price (probability of that outcome in 0,1), which is genuinely helpful. However, it leaves count, ticker, and time_in_force entirely unexplained, leaving a meaningful gap for a six-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Place a Kalshi limit order.' It also clarifies the meaning of outcome and action, which helps an agent understand what the tool does. It does not explicitly differentiate itself from the many sibling order-management tools, but the core purpose is unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives natural-language usage syntax but offers no guidance on when to use this tool versus alternatives like kalshi_cancel_order, place_order, or check_order. It also does not explain whether the dry-run behavior can be overridden or when a real order should be placed, leaving important selection and timing context implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kalshi_positionsDRead-onlyIdempotentInspect
Kalshi positions (needs the operator's API key).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds only that it needs the operator's API key, which is a useful auth context but minimal. It doesn't disclose pagination, rate limits, or what data is returned, so it adds little beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short sentence, but it is under-specified rather than concise. It omits critical information and does not front-load any useful semantic content. Conciseness is only a virtue when the content is sufficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (many siblings, no output schema, one undocumented parameter), the description is drastically incomplete. It doesn't state what the return value is, how results are ordered or filtered, or when to prefer this tool over alternatives. An agent cannot reliably invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one optional parameter 'limit' with default 50, but schema description coverage is 0% and the description gives no explanation of what 'limit' controls (e.g., number of positions returned). The agent cannot infer the parameter's meaning or constraints from either source.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is a noun phrase 'Kalshi positions' with no verb, essentially restating the tool name. It doesn't specify what the tool does (e.g., list, retrieve, fetch) and doesn't distinguish it from numerous siblings like my_positions, get_positions, or paper_positions. The added phrase about needing the API key is a prerequisite, not a purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus the many similar position-related tools (my_positions, get_positions, paper_positions). It doesn't mention any context, prerequisites beyond the API key, or exclusions. The agent is left to guess which of the ~50 siblings is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kalshi_search_marketsARead-onlyIdempotentInspect
Search Kalshi markets. Kalshi has no text-search endpoint, so this pages open markets and filters on title/ticker; auto-generated MVE combo shards are excluded. Prices are dollar strings, not cents.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | No | ||
| min_volume | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds valuable behavioral details: the absence of a native text-search endpoint, the exclusion of MVE combo shards, and that prices are dollar strings (not cents). No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose, followed by the workaround and pricing detail. Every sentence earns its place with no filler. The structure is highly efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 3 parameters, no output schema, and annotations present. The description covers the purpose and the paging/filtering mechanism but omits parameter semantics and doesn't describe the return structure (e.g., list shape, fields, pagination behavior). While not entirely inadequate, an agent would still be guessing about limits and output format, so it's not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for explaining parameters. It only hints at the 'query' parameter through 'filters on title/ticker', but it does not explain 'limit' or 'min_volume' at all. The burden on the description is high, and it fails to provide sufficient parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (search) and resource (Kalshi markets), and adds specific distinctions: it pages open markets, filters on title/ticker, excludes MVE combo shards, and notes dollar-string pricing. This is not a tautology and sets it apart from generic 'search_markets' siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the lack of a text-search endpoint and the workaround of paging/filtering, which implies usage context. However, it does not explicitly say when to choose this over alternatives like 'find_markets' or 'search_markets', nor does it give exclusions. Guidance is only implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
merge_positionsADestructiveInspect
Merge matching YES+NO shares back into USDC, GASLESS via the relayer. amount is in shares, or 'max' for the largest balanced amount held. Needs the operator's relayer key. DRY-RUN by default.
| Name | Required | Description | Default |
|---|---|---|---|
| amount | No | max | |
| condition_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true and readOnlyHint=false, but the description adds critical context: 'GASLESS via the relayer', 'DRY-RUN by default', and the need for the operator's relayer key. This goes beyond annotations, highlighting that the tool is a state-changing operation and explaining a default safety mechanism. No contradiction found.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a compact block of three sentences, with the core action and key qualifiers (gasless, dry-run) front-loaded. Every sentence adds value, though it could be slightly better structured with bullet points. No fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter tool with no output schema and destructive annotations, the description covers the main purpose, parameter semantics, and a key default. However, it omits details like what happens on merge (e.g., conversion rate, fees), error cases, or the exact behavior of 'max'. Given the complexity is moderate, it's short of fully complete but not grossly inadequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must define parameters. It explains 'amount' as being in shares or 'max', and implicitly condition_id is the identifier for the position. This adds meaning beyond the schema, but it doesn't fully clarify the format of condition_id or how 'max' is determined. Given the low coverage, this is adequate but not comprehensive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'merge' and the resource 'YES+NO shares back into USDC', and mentions 'GASLESS via the relayer'. It distinguishes from siblings like split_position and redeem_positions, though not explicitly. It is specific enough for an agent to understand the primary action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when the user wants to merge positions, but it does not explicitly state when to use this vs alternatives. It mentions 'DRY-RUN by default' and requires the operator's relayer key, which gives context. However, no exclusion or alternative tools are mentioned, so it's implied but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
my_balanceARead-onlyIdempotentInspect
Authenticated Polymarket cash and allowance snapshot for an existing configured wallet. Cash is separate from positions value; available-to-trade remains unknown until outstanding orders and unsettled fills are reconciled. Requires local account credentials. In simulation use paper_positions.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive behavior. The description adds valuable context beyond that: authentication requirement, the distinction between cash and position value, and the caveat that available-to-trade is not directly reported until order reconciliation. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense, non-redundant sentences. The primary purpose is front-loaded, followed by an important caveat and a routing instruction. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 0 parameters and no output schema, the description reasonably covers purpose, relevant caveats, and usage context. It could be more explicit about the return format or units, but it is sufficiently complete for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has 0 parameters, and the description still adds semantic context about what the snapshot represents and its limitations. With no parameters to document, this baseline is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a clear resource (Polymarket cash/allowance for an authenticated configured wallet) and a specific operation (snapshot). It distinguishes itself from sibling tools like kalshi_balance and paper_positions, and clarifies cash is separate from positions value.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context: it is for real authenticated Polymarket accounts requiring local credentialsiard, and explicitly names paper_positions as the simulation alternative. While it does not enumerate all alternatives, it gives sufficient directional guidance for an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
my_fillsBRead-onlyIdempotentInspect
Recent executions (fills) for the operator wallet — confirms what actually traded, with tx hashes.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is clear. The description adds useful output context by mentioning tx hashes and confirming actual executions, but it does not disclose ordering, pagination, or whether the list is limited by time or only the 'limit' parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence that front-loads the resource and scope, then adds the key differentiators (confirms actual trades, includes tx hashes). Every phrase earns its place with no redundant content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list with one optional parameter, the description gives enough to call it correctly: scope is the operator wallet, output is recent fills with tx hashes. It could be more explicit about the meaning of 'recent' or how limit applies, but the schema covers the limit default and the annotations cover safety.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not mention the 'limit' parameter at all. The parameter is simple and self-explanatory from its name and default value, but the description fails to compensate for the low coverage as required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource as fills/executions and scopes it to the operator wallet, distinguishing it from sibling tools that cover orders, positions, or balances. It lacks an explicit verb like 'list' or 'get,' but 'Recent executions' conveys the retrieval action well enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'confirms what actually traded' implies the intended use case of verifying completed trades, which is useful context. However, it does not explicitly state when to use this tool instead of related siblings like open_orders, my_positions, or kalshi_get_trades.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
my_positionsARead-onlyIdempotentInspect
The operator's current Polymarket positions (uses the configured wallet; no address needed).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile with readOnlyHint, idempotentHint, openWorldHint, and destructiveHint false. The description adds useful context about the configured wallet and that no address parameter is needed, but it does not disclose what 'current' includes or how the limit behaves. Given the annotations, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence states the core purpose and adds a useful parenthetical clarification. Every word contributes meaning; there is no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with one optional parameter and no output schema, the description conveys source, scope, and wallet dependency. However, it leaves 'current' somewhat ambiguous (open vs. settled positions) and does not explain the limit parameter, so a slightly fuller description would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one optional 'limit' parameter with a default of 50 but no description, and schema coverage is 0%. The description does not mention 'limit' or explain its effect, so it adds no parameter semantics and fails to compensate for the missing schema description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly names the resource ('current Polymarket positions') and the ownership scope ('operator's', 'configured wallet'). 'No address needed' sharpens the boundary and helps distinguish it from venue-specific or address-based position tools in the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for when to use the tool: when the agent needs the configured wallet's current Polymarket positions. It does not explicitly list alternatives or when-not-to-use conditions, but the wallet/venue framing makes the usage context obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_ordersARead-onlyIdempotentInspect
List the operator wallet's open orders.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, so the safety profile is clear. The description adds the 'operator wallet' scoping detail but does not disclose additional behavioral traits like pagination, ordering, or response contents.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, direct sentence with no filler. Every word contributes meaning, and the key object ('operator wallet's open orders') is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only list operation, the description is largely complete: it names the action and the scope. However, the lack of any clarification relative to the similarly named sibling tool and the absence of output format details are minor gaps given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero parameters and 100% schema coverage, there is nothing for the description to add about parameter meanings. The baseline of 4 is appropriate given the absence of parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('List') and the resource ('the operator wallet's open orders'). However, it does not differentiate from the similarly-named sibling kalshi_open_orders, leaving some ambiguity about scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives like kalshi_open_orders, check_order, or cancel_all_orders. The description implies it is for listing, but provides no explicit usage context or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
order_statusARead-onlyIdempotentInspect
Status of one Polymarket order: resting, partially filled, filled, or gone. The lifecycle answer an agent needs after place_order.
| Name | Required | Description | Default |
|---|---|---|---|
| order_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only/idempotent behavior, and the description adds the concrete outcome set (resting, partially filled, filled, gone), telling the agent what the call reports. It also conveys that this is a point-in-time lifecycle check. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one compact sentence that front-loads the essential information, the possible statuses, with no filler. Every word contributes to understanding the tool's purpose and output semantics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read-only tool with no output schema, the description supplies the return-value semantics and the context for use. It doesn't spell out the exact response shape or field names, but that is not required for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is only one parameter, order_id, which is self-evident, and the description's 'one Polymarket order' reinforces that the ID identifies the order to inspect. The schema provides no description, but with a single standard ID field little additional semantic burden falls on the tool description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the operation as retrieving the status of a single Polymarket order and lists the possible states, so an agent knows what it does. It doesn't explicitly differentiate from sibling tools like check_order, but naming 'one Polymarket order' and the lifecycle context distinguishes it from list-oriented tools like open_orders.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly frames usage as 'after place_order,' giving the temporal context for when to call it. It doesn't state when not to use alternatives such as check_order or open_orders, but the single-order lifecycle focus makes the intended use clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
overshoot_signalBRead-onlyIdempotentInspect
PREMIUM SIGNAL — overshoot/fade detector. Analyzes a token's recent price series for fresh panic jumps and reports whether a fade setup is active plus this market's historical reversion tendency.
| Name | Required | Description | Default |
|---|---|---|---|
| hours | No | ||
| token_id | Yes | ||
| threshold | No | ||
| lookback_s | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, open-world, idempotent, and non-destructive traits. The description adds value beyond those: it flags the tool as a 'PREMIUM SIGNAL' (suggesting access gating), clarifies it operates on a time-constrained recent window, and specifies the two computed outputs. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two well-organized sentences: the front-loaded label ('PREMIUM SIGNAL — overshoot/fade detector') conveys the tool class and gated status, and the second sentence carries the input/output detail. No wasted filler, though 'PREMIUM SIGNAL' is slightly marketing-toned.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool with no output schema and 0% schema coverage, the description gives the gist of input and output but not enough precision: the return format for 'fade setup active' and 'reversion tendency' is unspecified, and parameter semantics remain ambiguous. Adequate for a basic call but gaps remain for correct parameter tuning.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must explain the parameters but only does so loosely: token_id maps to 'a token's', hours/lookback_s map to 'recent price series', and threshold maps conceptually to 'panic jumps'. It never states what threshold is compared against, how hours relates to lookback_s, or what the defaults (6, 0.05, 60) mean.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (a token's recent price series), the specific action (analyzes for fresh panic jumps), and the two outputs (whether a fade setup is active, historical reversion tendency). It is clearly distinguishable from sibling order/position/balance tools by function, though it does not explicitly name any alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the purpose: an agent would call this to detect overshoot/fade setups after a panic move. However, it never states when not to use it or names alternatives such as price_history for raw series or dispute_risk for risk assessment, leaving the choice to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
paper_positionsARead-onlyIdempotentInspect
Paper-trading portfolio for dry-run: cash, positions at current marks, realized and unrealized P&L, resting paper orders (filled here if the market has crossed them). Dry-run Polymarket orders are papered against the live book automatically. Simulated: no queue, no impact, no fees, so results are an upper bound.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint, the description discloses important simulation behavior: resting paper orders fill if crossed, dry-run orders are papered against the live book automatically, and there is no queue, impact, or fees, making results an upper bound. This meaningfully adds context the annotations do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences cover the portfolio contents, the auto-papering mechanics, and the simulation caveats. Every sentence adds value, and the core purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter, read-only snapshot tool with no output schema, the description is complete: it states what the portfolio contains, how orders are treated, and the key simulation limitations. An agent can invoke this tool correctly without further clarification.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and schema coverage is 100%, so there is no parameter burden for the description to carry. The baseline of 4 for no-parameter tools applies; nothing additional is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly defines this as a paper-trading portfolio for dry-run and enumerates its contents: cash, positions at current marks, realized/unrealized P&L, and resting paper orders. This makes the tool's purpose unmistakable and differentiates it from live-position siblings like my_positions, get_positions, and kalshi_positions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly scopes the tool to dry-run paper trading, which is clear context for when to use it. However, it does not explicitly name alternatives or state 'use this instead of live position tools,' so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
paper_resetADestructiveInspect
Reset the paper-trading ledger to its starting bankroll (ODDSRAIL_PAPER_BANKROLL, default 1000 USDC). Deletes simulated fills and positions; touches nothing real.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as destructive, but the description adds valuable specifics: it deletes simulated fills and positions and affects nothing real. It also discloses the bankroll source and default value, going beyond the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The action, target, side effects, and safety boundary are all front-loaded and clearly stated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter destructive tool, the description covers the action, target, default value, what is deleted, and what is untouched. No output schema exists, but the absence of return-value details does not hinder correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters, so there is nothing to document beyond the empty schema. The description adds useful context about the bankroll constant, but parameter semantics is trivially satisfied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Reset the paper-trading ledger' with a clear end state (starting bankroll). It distinguishes itself from the many live-market siblings by explicitly saying it operates on simulated data and 'touches nothing real.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is clear: reset the paper-trading ledger to its initial bankroll. It does not name alternative tools or state when-not-to-use, but the paper-trading scope and safety boundary make the context unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
place_orderADestructiveInspect
Place a limit order. DRY-RUN by default: returns the order it would post. Real trading needs ODDSRAIL_DRY_RUN=0 and POLYMARKET_PRIVATE_KEY. The operator's builder code is signed into the order. price is the implied probability in (0,1); size is in SHARES (notional = price * size), and the exchange enforces a $1 minimum notional on marketable orders. post_only=True rejects rather than crosses the book.
| Name | Required | Description | Default |
|---|---|---|---|
| side | Yes | ||
| size | Yes | ||
| price | Yes | ||
| token_id | Yes | ||
| post_only | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations that already mark this as destructive and non-read-only, the description discloses dry-run default behavior, signed builder code, post_only rejection semantics, and the exchange's $1 minimum notional. This materially enriches the agent's understanding of side effects and execution behavior, with no contradiction of the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences pack the essential behavioral and parameter information with no filler. The dry-run warning is front-loaded, and the parameter clarifications are grouped logically, making the description easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the key runtime behavior, prerequisites, order semantics, and constraints for a trading tool. It lacks an explicit return shape or error behavior, but the statement 'returns the order it would post' gives a sufficient baseline, and annotations already signal risk.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must carry the burden. It does well for price (implied probability in (0,1)), size (shares), notional calculation, and post_only, but leaves the required 'side' and 'token_id' parameters unexplained. This is a clear gap for such a consequential tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Place a limit order.' It immediately establishes scope by mentioning Polymarket's private key, which distinguishes it from the Kalshi-specific sibling 'kalshi_place_order.' The dry-run default further sharpens what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states the default dry-run behavior and the exact environment variables required for real trading, giving the agent concrete conditions for when actual side effects occur. It does not explicitly name alternatives or say when not to use the tool, but the context is sufficient for most routing decisions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
position_sizeARead-onlyIdempotentInspect
Fractional-Kelly position size for a binary contract, given your bankroll, the market price, and YOUR fair value estimate. Caps at a fraction of full Kelly and refuses negative-edge bets. Returns its assumptions — subtract quote_cost before trusting the number.
| Name | Required | Description | Default |
|---|---|---|---|
| price | Yes | ||
| fair_value | Yes | ||
| bankroll_usd | Yes | ||
| max_fraction_of_kelly | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered. The description adds genuinely useful behavior beyond that: it caps at a fraction of full Kelly, refuses negative-edge bets, returns its assumptions, and warns that quote_cost must be subtracted. These are non-obvious traits an agent could not infer from the schema or annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with zero waste, each earning its place: the first states the core purpose, the second discloses capping and refusal behavior, and the third delivers the output caveat. The most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple calculator with no output schema and no parameter descriptions, the description covers inputs, algorithm, refusal behavior, and a return-value caveat. The only notable gap is the exact shape of the returned 'assumptions' and the primary numeric output, which the agent must infer rather than read.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must carry the semantic burden, and it largely does: it maps bankroll, market price, and the user's fair value estimate to the three required parameters, emphasizing that fair_value is 'YOUR' subjective estimate as opposed to the market price. It does not clarify units or the exact expected range for price/fair_value, and max_fraction_of_kelly is only alluded to via 'Caps at a fraction of full Kelly.'
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it computes a 'Fractional-Kelly position size for a binary contract' from three named inputs. It clearly distinguishes itself from sibling list/order/balance tools by being a sizing calculation rather than a position query or order mutation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies its role in a pre-trade workflow, especially with 'subtract quote_cost before trusting the number,' which connects it to a sibling tool. However, it never explicitly states when to use this versus alternatives (e.g., 'use before placing an order') or when not to use it, leaving the usage context implicit rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
price_historyBRead-onlyIdempotentInspect
Recent price history for a CLOB token id: hours back, at fidelity_minutes resolution.
| Name | Required | Description | Default |
|---|---|---|---|
| hours | No | ||
| token_id | Yes | ||
| fidelity_minutes | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds minor context (the temporal scope and resolution) but does not explain return format, pagination, or potential data latency. With annotations carrying the main behavioral disclosure, this is adequate but thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the purpose and immediately links the parameters. There is no redundant filler or repetition of the tool name. Every word contributes to the meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (3 parameters, one required) and annotations cover read-only and idempotent behavior. Still, with no output schema, the description does not state what the response looks like (e.g., an array of price points, timestamps, or units). This leaves some ambiguity, but the phrase 'price history at fidelity_minutes resolution' hints at aggregated time series data, so it is not severely incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does relate the parameters to the query: 'hours back' clarifies the 'hours' parameter, 'fidelity_minutes resolution' clarifies fidelity_minutes, and 'CLOB token id' identifies token_id. However, this is only a minimal restatement of the parameter names and does not add significant detail about formats, constraints, or allowed values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource ('recent price history') and its scope ('CLOB token id', 'hours back', 'fidelity_minutes resolution'). It is distinct from sibling tools like get_orderbook or get_trades because it specifically mentions price history, though it lacks an explicit verb such as 'gets' or 'retrieves'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives. It does not mention any exclusions, prerequisites, or refer to sibling tools such as kalshi_get_trades or get_market. An agent is left to infer that price history is for historical data rather than live quotes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
quote_costARead-onlyIdempotentInspect
What a given size would ACTUALLY cost, by walking the order book rather than reading the top level. Returns average fill price, slippage vs best, notional, and the levels consumed — plus the venue's fee schedule where it publishes one. Call this before sizing any trade, and on both legs before acting on a cross-venue gap.
| Name | Required | Description | Default |
|---|---|---|---|
| side | Yes | ||
| size | Yes | ||
| venue | Yes | ||
| market_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, and the description adds meaningful behavior beyond that: it walks the order book (mechanism), returns average fill price/slippage/notional/levels consumed (output contract), and conditions the fee-schedule portion on venue publication ('where it publishes one'). Nothing contradicts the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: purpose contrast, return contract, and usage timing. The core differentiator is front-loaded in the first clause, and there is no filler or repetition of what the schema already shows.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating return values (average fill price, slippage, notional, levels consumed, conditional fee schedule), and the annotations carry the safety profile. Remaining gaps — what happens when size exceeds book depth and exact parameter value formats — are not disclosed, though they are partially inferable from sibling tools like get_market and get_orderbook.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the four parameters carry only names and types. The description partially compensates: 'size' is clarified as the hypothetical order size, and 'both legs before acting on a cross-venue gap' implies venue/market_id identify each leg. However, it never specifies side values (buy/sell) or the expected formats of venue/market_id, leaving the compensation incomplete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The lead clause 'What a given size would ACTUALLY cost, by walking the order book rather than reading the top level' names a specific verb (estimate true cost), a resource (order book depth), and an execution method. The explicit 'rather than reading the top level' contrast differentiates it from siblings like get_orderbook or watch_book, so an agent can tell it apart without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Two explicit directives are given: 'Call this before sizing any trade, and on both legs before acting on a cross-venue gap.' This is clear, concrete when-to-use context tied to real workflows. It stops short of naming the specific alternative tool for top-level reads or stating a when-not condition, so it earns a 4 rather than a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
redeemable_positionsARead-onlyIdempotentInspect
What the operator can turn back into USDC now: redeemable (resolved, winning) positions and mergeable (hold both YES and NO) positions, with the ids the gasless tools take. Uses the configured wallet.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly, idempotent, openWorld, and non-destructive behavior. The description adds useful behavioral context beyond those: it tells the agent that the configured wallet is used, that the positions are resolved/winning, and that the returned IDs are compatible with the gasless tools. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the main purpose and then adds defining detail. It contains no filler, though the 'What ... now:' construction is slightly awkward and could be phrased more directly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple listing tool with rich annotations, the description is largely complete: it names the resource, defines redeemable and mergeable, identifies the wallet scope, and notes the gasless-tool ID compatibility. However, it does not describe the return shape or the effect of the limit parameter, which would make it fully self-contained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention the single limit parameter at all. The parameter is optional and has a default, so the bare name 'limit' carries some meaning, but the description fails to compensate for the low schema coverage as required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states what the operator can do: turn redeemable and mergeable positions back into USDC, and that the tool returns the IDs needed for gasless operations. It is specific enough to distinguish this listing tool from generic position tools like kalshi_positions or my_positions, and from the action tools redeem_positions and merge_positions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use this to find positions that are ready for redemption or merging, and the IDs are meant for the gasless tools. It does not explicitly name an alternative or state when not to use it, but the specialized scope is implied well enough for an agent to route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
redeem_positionsADestructiveInspect
Redeem the winning shares of a RESOLVED market for USDC, GASLESS via the relayer. Pass exactly one of condition_id or market_id. Needs the operator's relayer key. DRY-RUN by default.
| Name | Required | Description | Default |
|---|---|---|---|
| market_id | No | ||
| condition_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already signal destructive and non-read-only behavior, and the description adds valuable context beyond that: 'GASLESS via the relayer', 'Needs the operator's relayer key', and 'DRY-RUN by default'. These traits are not visible in the schema or annotations, and the description does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short sentences with no filler. The purpose is front-loaded, followed immediately by the parameter rule, then the key requirement and default behavior. Every sentence carries operational value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, parameter exclusivity, the required relayer key, and dry-run default, which is a solid base. However, with no output schema, it does not hint at what the tool returns or how success is signaled, and it does not explain how to move from a dry run to an actual redemption. This leaves an agent uncertain about the full execution flow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It directly addresses the most important parameter ambiguity by saying 'Pass exactly one of condition_id or market_id', which is crucial because both parameters are optional in the schema. It does not provide deeper format or example details, but for two self-explanatory ID parameters this is reasonably sufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Redeem the winning shares'), a specific resource (RESOLVED market), and the output (USDC), with the gasless relayer detail. This clearly distinguishes it from sibling tools like redeemable_positions, which would be about identifying redeemable positions rather than executing redemption.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly conveys when to use the tool: when a market is RESOLVED and you hold winning shares. It also gives an important operating condition ('Pass exactly one of condition_id or market_id') and flags dry-run behavior. However, it does not explicitly name alternatives or state when not to use it, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolution_criteriaARead-onlyIdempotentInspect
READ BEFORE TRUSTING A PRICE: the full resolution contract for a market — what exactly resolves YES, who resolves it, from which sources. venue is 'polymarket' (pass the slug, the Gamma id, or the market_id a search returned, which is a CLOB token id) or 'kalshi' (pass the ticker).
| Name | Required | Description | Default |
|---|---|---|---|
| venue | Yes | ||
| market_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive behavior. The description adds useful behavioral context beyond annotations: it reveals the returned information's scope and the accepted identifier formats for each venue. This sufficiently informs an agent about the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact but information-dense, front-loading urgency and purpose before diving into parameter details. The long sentence with parentheticals is slightly dense but every clause earns its place and no unnecessary words appear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read-only lookup tool with no output schema, the description is complete: it tells the agent what information will be returned, how to identify the market on each venue, and is backed by annotations covering safety. Nothing needed to call the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must carry the parameter semantics, and it does. It explicitly defines 'venue' as either 'polymarket' or 'kalshi' and explains what 'market_id' should be for each: slug, Gamma id, or CLOB token id for Polymarket; ticker for Kalshi.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool's function: retrieving the full resolution contract for a market, and specifies what that contract contains (what resolves YES, who resolves it, from which sources). This is specific enough to distinguish it from generic market lookup tools like get_market or search_markets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The opening phrase 'READ BEFORE TRUSTING A PRICE' strongly signals when to use the tool, and the description explains how to adapt the input for Polymarket versus Kalshi. It does not explicitly name alternatives or state when not to use it, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_marketsARead-onlyIdempotentInspect
Search Polymarket markets by text (Gamma public-search under the hood); empty query lists open markets. Returns token ids, prices, metrics, resolution info.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only/idempotent behavior, and the description adds meaningful context: the empty-query behavior and what the tool returns (token ids, prices, metrics, resolution info). It does not discuss pagination or rate limits, but those are minor for an annotated read-only search tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences front-load the core behavior and return information with no repetition of schema or annotations. Every clause adds information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only search tool with two defaulted parameters and no output schema, the description covers the action, the empty-query edge case, and the return content. The only notable omission is explicit guidance on how limit behaves, though its default is present in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must carry parameter meaning but only addresses query via 'by text' and the empty-query behavior. The limit parameter is not explained at all, leaving a meaningful gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Search Polymarket markets by text,' with the venue name distinguishing it from sibling tools like kalshi_search_markets. The additional 'empty query lists open markets' behavior further pins down the tool's purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it—for text search of Polymarket markets, and for Kalshi there is a sibling with that name—but it never states when not to use it or names an alternative such as find_markets for cross-venue searches.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
server_infoARead-onlyIdempotentInspect
Server status: dry-run state, attribution config, enabled capabilities, venue reachability from this machine, and Polymarket's geoblock verdict for this machine's IP (advisory — not a compliance check). Call this before the first order of a session.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as read-only, idempotent, non-destructive, and open-world. The description adds useful context beyond those hints: it lists exactly what status categories are included and clarifies that the Polymarket geoblock verdict is advisory only, not a compliance check. This helps the agent interpret the result correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: it opens with 'Server status', lists the key components in a scannable manner, adds an important caveat, and closes with a clear instruction on when to call it. No sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only status tool with no output schema, the description is complete. It tells the agent what information will be available, flags the advisory nature of the geoblock check, and states when to invoke the tool. There are no missing inputs or ambiguous actions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is nothing for the description to explain about inputs. The rubric's baseline for zero-parameter tools is 4, and the description appropriately focuses on behavior rather than parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource as server status and enumerates the specific components included (dry-run state, attribution config, enabled capabilities, venue reachability, geoblock verdict). It does not use an explicit verb like 'get' or 'retrieve', and it does not explicitly compare itself to sibling tools, though its uniqueness among the trading tools is evident.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit temporal guidance: 'Call this before the first order of a session.' This tells the agent when to use the tool. It does not mention when not to use it or name alternatives, but no sibling appears to offer the same server-status functionality.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
settlement_auditARead-onlyIdempotentInspect
Settlement-divergence audit for a cross-venue pair, on LIVE data (no pre-curated pair list). Compares close times, resolution sources, UMA dispute status and market structure, and returns ok / caution / block with reasons. polymarket_id takes the slug, the Gamma id or the market_id a search returned (a CLOB token id); kalshi_ticker takes the ticker. Run this before treating any cross-venue price difference as an edge.
| Name | Required | Description | Default |
|---|---|---|---|
| notional_usd | No | ||
| kalshi_ticker | Yes | ||
| polymarket_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this read-only, open-world, idempotent, and non-destructive; the description adds behavior beyond the annotations by specifying live-data operation, the comparison dimensions, and the three-way verdict format with reasons, plus input-id aliasing for polymarket_id. No contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the core purpose, then output, then parameter semantics, then usage trigger. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, inputs (for the two required params), output categories, and when to run. Gaps: notional_usd's meaning and a more precise return shape, though the verdict categories give a workable contract. Overall sufficient for an audit tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has zero description coverage, but the description explains that polymarket_id accepts a slug, Gamma id, or CLOB token id from search, and kalshi_ticker accepts the ticker. However, notional_usd (optional, default 0) is never described, leaving its purpose unclear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (audit) and resource (settlement divergence for a cross-venue pair), lists the exact comparisons (close times, resolution sources, UMA dispute status, market structure) and the output categories (ok/caution/block). This distinguishes it from siblings like compare_venues or dispute_risk and from search tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The final sentence gives an explicit trigger: 'Run this before treating any cross-venue price difference as an edge.' The 'LIVE data (no pre-curated pair list)' note sets expectations about applicability. No explicit alternatives are named or excluded, so not a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
split_positionADestructiveInspect
Split USDC collateral into a full YES+NO share set for one market, GASLESS via Polymarket's relayer. amount_usdc is collateral in USDC (e.g. 25). Needs POLYMARKET_RELAYER_API_KEY and POLYMARKET_RELAYER_API_KEY_ADDRESS (polymarket.com -> Settings -> Relayer API keys). DRY-RUN by default. Never falls back to a gas-paying broadcast.
| Name | Required | Description | Default |
|---|---|---|---|
| amount_usdc | Yes | ||
| condition_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the destructiveHint and readOnlyHint annotations, the description adds critical behavioral detail: execution is gasless, it never falls back to a gas-paying broadcast, and dry-run is the default mode. This tells an agent exactly what side effects to expect and how failures or fallbacks will behave.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences each carry distinct information: the action, the parameter format, the required keys, and the dry-run/no-fallback safety behavior. There is no fluff, and the most important operational detail is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers key operational context such as API keys, gasless execution, and dry-run behavior. However, with no output schema and no explanation of the required condition_id parameter, the agent must guess that parameter's meaning and what the dry-run invocation returns. It is adequate but has clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain the parameters itself. It explains amount_usdc with an example ('collateral in USDC e.g. 25') but says nothing at all about condition_id, which is a required parameter. This leaves a significant semantic gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a precise action ('Split USDC collateral into a full YES+NO share set') and ties it to a specific venue and mechanism (Polymarket's gasless relayer). This clearly distinguishes it from sibling tools like merge_positions and redeem_positions without needing to inspect their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives strong context for correct use: it operates on one market, requires specific API keys, and is dry-run by default. It does not explicitly name alternatives or state when not to use this tool, but the conditions for using it are concrete and unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
watch_bookARead-onlyIdempotentInspect
Stream one Polymarket token's realtime market events (book snapshot, price changes, trades) for up to seconds (1-60) or max_events, then return them. Use after get_orderbook when you need to see the book MOVE before acting; a quiet market may deliver only the initial snapshot.
| Name | Required | Description | Default |
|---|---|---|---|
| seconds | No | ||
| token_id | Yes | ||
| max_events | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only, idempotent, non-destructive nature, and the description adds behavioral detail: streaming is bounded by seconds (1-60) or max_events, and the tool returns the collected events afterward. It also warns about quiet-market behavior, which is useful beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences deliver resource, event types, termination conditions, usage sequencing, and an expectation-setting caveat with no filler. The most important operational facts are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only streaming tool with a small schema and no output schema, the description covers what events are returned, what terminates the stream, and how it fits into the get_orderbook workflow. It could specify the exact return shape, but the event types and bounded-stream semantics are sufficient for invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Despite 0% schema description coverage, the description clarifies token_id as the target Polymarket token and explains how seconds (1-60) and max_events act as termination limits. It does not restate defaults, but the schema already provides those values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific verb ('stream') and resource ('one Polymarket token's realtime market events') and enumerates event types (book snapshot, price changes, trades). It clearly distinguishes itself from get_orderbook by emphasizing book movement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use it after get_orderbook when the agent needs to see the book move before acting, and sets expectations that a quiet market may only return the initial snapshot. This provides both a sequencing cue and a condition for when the tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
42 tool updates
v0.18.1- Added
attribution_ledger - Changed
builder_stats1 field changed- changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "builder_statsOutput", - "type": "object" -}New value: +null
- Added
cancel_all_orders - Changed
cancel_order1 field changed- changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "cancel_orderOutput", - "type": "object" -}New value: +null
- Added
check_order - Added
closing_soon - Added
compare_venues - Changed
dispute_risk1 field changed- changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "dispute_riskOutput", - "type": "object" -}New value: +null
- Added
find_markets - Changed
get_market1 field changed- changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "get_marketOutput", - "type": "object" -}New value: +null
- Changed
get_orderbook1 field changed- changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "get_orderbookOutput", - "type": "object" -}New value: +null
- Changed
get_positions1 field changed- changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "get_positionsOutput", - "type": "object" -}New value: +null
- Changed
kalshi_balance1 field changed- changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "kalshi_balanceOutput", - "type": "object" -}New value: +null
- Changed
kalshi_cancel_order1 field changed- changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "kalshi_cancel_orderOutput", - "type": "object" -}New value: +null
- Changed
kalshi_get_market1 field changed- changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "kalshi_get_marketOutput", - "type": "object" -}New value: +null
- Changed
kalshi_get_orderbook1 field changed- changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "kalshi_get_orderbookOutput", - "type": "object" -}New value: +null
- Changed
kalshi_get_trades1 field changed- changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "kalshi_get_tradesOutput", - "type": "object" -}New value: +null
- Changed
kalshi_open_orders1 field changed- changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "kalshi_open_ordersOutput", - "type": "object" -}New value: +null
- Changed
kalshi_place_order1 field changed- changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "kalshi_place_orderOutput", - "type": "object" -}New value: +null
- Changed
kalshi_positions1 field changed- changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "kalshi_positionsOutput", - "type": "object" -}New value: +null
- Changed
kalshi_search_markets1 field changed- changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "kalshi_search_marketsOutput", - "type": "object" -}New value: +null
- Added
merge_positions - Added
my_balance - Added
my_fills - Added
my_positions - Changed
open_orders1 field changed- changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "open_ordersOutput", - "type": "object" -}New value: +null
- Added
order_status - Changed
overshoot_signal1 field changed- changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "overshoot_signalOutput", - "type": "object" -}New value: +null
- Added
paper_positions - Added
paper_reset - Changed
place_order3 fields changed- removed
Input schema / properties / order_typeRemoved value: -{ - "default": "GTC", - "title": "Order Type", - "type": "string" -} - added
Input schema / properties / post_onlyAdded value: +{ + "default": false, + "title": "Post Only", + "type": "boolean" +} - changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "place_orderOutput", - "type": "object" -}New value: +null
- Added
position_size - Changed
price_history1 field changed- changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "price_historyOutput", - "type": "object" -}New value: +null
- Added
quote_cost - Added
redeem_positions - Added
redeemable_positions - Added
resolution_criteria - Changed
search_markets1 field changed- changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "search_marketsOutput", - "type": "object" -}New value: +null
- Changed
server_info1 field changed- changed
Output schema / (root)Previous value: -{ - "properties": { - "result": { - "title": "Result", - "type": "string" - } - }, - "required": [ - "result" - ], - "title": "server_infoOutput", - "type": "object" -}New value: +null
- Added
settlement_audit - Added
split_position - Added
watch_book
21 tool updates
v0.3.0- First observed
builder_stats - First observed
cancel_order - First observed
dispute_risk - First observed
get_market - First observed
get_orderbook - First observed
get_positions - First observed
kalshi_balance - First observed
kalshi_cancel_order - First observed
kalshi_get_market - First observed
kalshi_get_orderbook - First observed
kalshi_get_trades - First observed
kalshi_open_orders - First observed
kalshi_place_order - First observed
kalshi_positions - First observed
kalshi_search_markets - First observed
open_orders - First observed
overshoot_signal - First observed
place_order - First observed
price_history - First observed
search_markets - First observed
server_info
TDQS
Scored across 42 tools
The venue prefixes (kalshi_ vs no prefix) and action nouns keep most tools distinct. Only mild overlaps exist—notably find_markets vs search_markets vs kalshi_search_markets, and my_positions vs get_positions—but the descriptions clarify when to use each. No tools are truly indistinguishable.
Names are consistently lowercase snake_case and mostly follow a predictable venue_action_noun pattern, e.g. kalshi_place_order, get_orderbook, cancel_all_orders. A few noun-style names like server_info, builder_stats, and position_size break the verb-first feel, but there is no chaotic style mixing.
42 tools is well beyond the 25+ threshold for 'too many' and feels heavy even for a two-venue trading platform. Many tools could be consolidated or grouped, such as search variants, position variants, and several signal/analysis utilities. The breadth is real, but the count creates selection overhead.
The toolkit is broad: market discovery, order placement/cancellation, balances/positions, paper trading, redemption, resolution checks, builder attribution, and cross-venue settlement audits are all covered. Minor gaps include Kalshi-specific order status/fill history and a dedicated Polymarket trade history tool, so agents may need to work around those edges.
Maintenance
Related MCP Connectors
Calibrated world model for AI agents. 40 tools: world state, markets, trading. Kalshi + Polymarket.
Your agent needs the crowd's number — live odds on elections, policy, macro prints and sport, from the venues where people put money behind the opinion. **What you can ask for** • "What are the current odds on this event?" • "List the open markets on this topic across both venues." • "Show recent trades and how the price moved." • "What is the implied probability now versus a week ago?" **How to use it** Point any MCP client at https://mcp.aisa.one/prediction-market-data/mcp and sign in with OAuth — there is no key to create or paste. 5 tools: Kalshi markets and trades, Polymarket markets, events and activity. **Why this rather than the source** Both venues in one shape, so the same question can be priced twice. **It is also a door to the rest** The same login reaches 26 sources and 580+ operations. Read the odds here, then ask the same agent for the market data or the news behind them — without adding a second server. **What it costs** Finding and inspecting an operation is free. Running one is billed per call at API prices, with no seat and no monthly minimum, and every call takes max_price_usd so an agent cannot overspend by accident. **Where else it reaches** https://mcp.aisa.one/finance/mcp for equities, crypto and prediction markets in one place.
Polymarket + Hyperliquid + macro for AI agents. 38 tools, signal backtest, SSE streaming. Free tier.
Hosted MCP for Kalshi prediction markets: search, odds, order books, settlement rules, and trading.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceReal-time prediction market intelligence for AI agents. Query Polymarket and Kalshi markets, wallet profiles, smart money leaderboards, social pulse signals, price candlesticks, and orderbook data — 13 agents, one MCP connection. Powered by 1.1TB+ of historical data.MIT
- AlicenseAqualityDmaintenancePrediction market probability oracle for AI agents. 26 tools across 500+ live markets from Kalshi and Polymarket. Cross-source arbitrage detection, structured TPF signals, Kelly Criterion sizing, agent performance tracking, and webhook alerts.960 npm1MIT
- AlicenseAqualityFmaintenanceProvides prediction market intelligence, research, and strategy signals for platforms like Kalshi, Polymarket, and Robinhood. It enables AI assistants to perform market screening, arbitrage detection, and deep causal analysis to support informed trading decisions.271MIT
- AlicenseNot gradedqualityBmaintenanceAn MCP server and Python toolkit that provides AI agents with real-time tools for Polymarket prediction markets, including liquidity scanning, arbitrage detection, and slippage estimation. It also offers advanced wallet intelligence, portfolio risk calculation, and probabilistic reasoning to enhance market analysis and strategy.94 PyPI1MIT