polymarket-mcp-server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@polymarket-mcp-servershow me the order book for the next Fed rate cut"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
polymarket-mcp-server
A read-only MCP server for Polymarket. It lets an MCP client (Claude Code, Claude Desktop, …) query markets, order books, price history, wallet positions, and market holders — using Polymarket's three public, unauthenticated APIs.
No trading. No private keys. No signer. No authenticated CLOB endpoints. It only uses the SDK's public read-only client; viem/ethers are not installed (note: @polymarket/client depends on ox, so npm install pulls some crypto primitives transitively, unused here). The market/data tools are readOnlyHint: true; the paper-trading + calibration tools only write a local, gitignored data/ ledger of simulated bets — never real money, no order is ever placed.
What it can do — 23 tools
13 read-only market/data tools, 5 simulated paper-trading tools (local ledger, no real money), and 5 analytics/automation tools.
Tool | API | Purpose |
| Gamma | Find markets by keyword, or list top markets by volume. |
| Gamma | Full detail of one market by slug, incl. CLOB token ids per outcome. |
| Gamma | Events grouping related markets (a whole tournament/election). |
| CLOB (SDK) | Live best bid/ask, spread, mid, depth — is an edge executable? |
| CLOB (SDK) | Historical implied probability over a window, with summary stats. |
| CLOB (SDK) | Quick midpoint / buy / sell / spread / last-trade for one outcome. |
| Data | Recent trades — a market's tape, or a wallet's trade history. |
| Data | Open interest (capital at stake) for one or more markets. |
| Data | Global trader ranking by realized P&L or volume. |
| Data | Open positions + P&L for a public wallet. |
| Data | A wallet's resolved positions with realized P&L. |
| Data | A wallet's full activity feed (trades, splits, merges, …). |
| Data | Largest holders ("whales") of each outcome in a market. |
| local | SIMULATED paper trading — record bets to a local ledger (fills at ask + taker fee), mark-to-market, auto-score at resolution, Brier vs the market. Opening is capped by the available virtual bankroll. No real money, no order placed. |
| local | Rewrite the simulated ledger: drop one bet, or wipe it and start a fresh virtual bankroll ( |
| CLOB (SDK) | Scan for baskets priced under $1 (Yes+No, or a negRisk event) net of fees — the one edge a read-only bot can detect. Detection ≠ capture. |
| mixed | One-call deep look: detail + quote + whales + price history + trade tape. |
| mixed | Snapshot markets now; when they resolve, score a reliability curve + Brier — is the market well-priced? |
| mixed | A "morning briefing": paper status + arbitrage scan + top movers + calibration status. |
The CLOB read tools run on the official @polymarket/client SDK (public read-only client — no keys, no wallet); Gamma and Data run on native fetch. The paper-trading and calibration tools write only a local, gitignored data/ ledger — never real money.
Related MCP server: polymarket-mcp
The identifier model (read this once)
Polymarket uses several ids and it's easy to get lost:
A binary market has a human
slug, aconditionId, and twoclobTokenIds— one per outcome (Yes / No).The order book and price history key on the
token_id(a ~77-digit number), not the slug.
Flow: slug → polymarket_get_market gives you the clobTokenIds → CLOB tools use a token_id.
For convenience, the CLOB tools (get_orderbook, get_price_history, get_quote) accept either:
token_iddirectly (preferred if you have it), orslug+outcome(e.g."Yes") — the server resolves the token id for you via Gamma.
Prerequisites
Node.js ≥ 24 (required by the official Polymarket SDK; Gamma/Data still use the built-in
fetch). Check withnode --version.Windows:
winget install OpenJS.NodeJS.LTSor download from nodejs.org.
Install & build
npm install
npm run build # compiles TypeScript to dist/Try it with the MCP Inspector
npm run inspector # builds, then opens the MCP Inspector against dist/index.jsThen call e.g. polymarket_search_markets with { "query": "bitcoin" }.
Register in Claude Code
This repo already ships a project-scoped .mcp.json with a relative path, so it works on any machine once you've run npm install && npm run build from the project root:
{
"mcpServers": {
"polymarket": {
"command": "node",
"args": ["dist/index.js"]
}
}
}The relative dist/index.js resolves from the project root (where Claude Code is launched). If your client needs an absolute path, substitute the full path to dist/index.js on that machine. No secrets or environment variables are required — everything it touches is public.
Install from a clone
git clone <your-repo-url>
cd polymarket-mcp
npm install
npm run buildThen reload MCP / restart Claude Code (project .mcp.json servers need a one-time approval in an interactive session). Requires Node ≥ 24 on the new machine.
Usage examples
"What Polymarket markets are there about the Fed?" →
polymarket_search_markets { query: "Fed" }"What's the real spread on Argentina to win the World Cup?" →
polymarket_get_orderbook { slug: "will-argentina-win-the-2026-fifa-world-cup-245", outcome: "Yes" }"How has that market moved this month?" →
polymarket_get_price_history { slug: "…", outcome: "Yes", interval: "1m" }"How is wallet 0xabc… doing?" →
polymarket_get_positions { wallet: "0xabc…" }"Who are the whales on the Yes side?" →
polymarket_get_market_holders { slug: "…" }"Give me the full picture on this market" →
polymarket_xray { slug: "…", outcome: "Yes" }"Any arbitrage in the top events?" →
polymarket_find_arbitrage { scan_top: 10 }"Track a simulated bet" →
polymarket_paper_open { slug: "…", outcome: "No", size_usdc: 50, p_estimate: 0.6 }, thenpolymarket_paper_status"Start the paper lab over" →
polymarket_paper_reset { confirm: true, bankroll: 500 }(erases every simulated bet — there is no undo)"My morning briefing" →
polymarket_daily_digest
Scheduling the daily digest (Windows)
The server is a local stdio process, so a cloud routine can't reach it. To get the digest unattended, use Windows Task Scheduler to run a headless Claude command with its "Start in" set to the project root:
claude -p "Run polymarket_daily_digest and print it" --allowedTools "mcp__polymarket__polymarket_daily_digest" --output-format text >> "%USERPROFILE%\polymarket-digest.log"Requires: the PC on/awake at the trigger time; Claude Code installed and authenticated for
non-interactive use; the task's Start in = the project folder (so .mcp.json's relative
dist/index.js resolves); and the one tool pre-allowed (a scheduled run can't answer permission
prompts). Each run consumes Claude usage. (/loop also works, but only while a session stays open;
cloud /schedule does not reach a local stdio server.)
Design notes
APIs / transport: Gamma
gamma-api.polymarket.comand Datadata-api.polymarket.comover nativefetch; the CLOB reads (order book, price history, quote) go through the official@polymarket/clientpublic client (createPublicClient(), no keys). Seesrc/clobSdk.ts.Rate limits: Gamma allows ~60 req/min unauthenticated. The client backs off exponentially (honoring
Retry-After) on 429/5xx and retries transient network errors.Context discipline: responses default to concise Markdown; pass
response_format: "json"for full structured data. Every tool also returns machine-readablestructuredContent.JSON-string fields: Gamma returns
outcomes,outcomePrices, andclobTokenIdsas JSON-encoded strings; the client parses and zips them so each outcome is paired with its token id and implied price.
Project layout
src/
index.ts # entry: McpServer + stdio transport
constants.ts # base URLs, limits
types.ts # raw API + normalized types
client.ts # PolymarketClient: Gamma/Data fetch+backoff, error mapping, slug→token; CLOB via SDK
clobSdk.ts # official @polymarket/client public client + pure response mappers (CLOB reads)
arbMath.ts # pure arbitrage math (edge + depth-walked executable size)
calibrationMath.ts # pure calibration math (buckets, reliability curve, Brier, ECE)
paperFees.ts # taker-fee model + fractional-Kelly + Brier (pure)
paperStore.ts / calibrationStore.ts / dataDir.ts # local JSON ledgers (gitignored data/)
schemas.ts # Zod input/output schemas
format.ts # formatting + result/error helpers
tools/ # one file per tool (23)Scope / non-goals
Read-only by design. Order placement, cancellation, approvals, or anything requiring a signature/API key is intentionally out of scope — the official SDK is used only via its public read-only client (createPublicClient()), which pulls no wallet/crypto libraries. Global trader rankings are available via polymarket_get_leaderboard; polymarket_get_market_holders covers per-market whale visibility.
License
MIT
Screenshots
How it was built
Most of the code was written by Claude Code. I set the scope, made the design decisions and tested the behaviour.
Available Tools
23 toolspolymarket_calibration_reportCalibration reportA
Score the calibration snapshots: resolve any that have settled since, then compute a reliability curve (predicted vs observed per decile) + Brier score + ECE. Tells you whether the market's prices are honest. Needs elapsed time for markets to resolve, so it's sparse early on.
Args:
min_resolved (number, default 1): note if fewer than this have resolved.
Returns: { totalSnapshots, resolvedCount, pendingCount, brier, ece, curve:[{bucket,label,n,predictedMean,observedFreq}] }.
| Name | Required | Description | Default |
|---|---|---|---|
| min_resolved | No | Minimum resolved snapshots before computing metrics (default 1). | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| ece | Yes | |
| brier | Yes | |
| curve | Yes | |
| pendingCount | Yes | |
| resolvedCount | Yes | |
| totalSnapshots | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, idempotentHint=false, destructiveHint=false; the description independently discloses the write behavior ('resolve any that have settled since'), which is what makes this non-read-only and non-idempotent, and adds the sparsity caveat. It adds real context the annotations alone don't convey, though it doesn't say whether resolution is persisted or how it affects counts.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose in the first sentence, then scoped Args and Returns sections — easy to scan. However, the Args and Returns blocks substantially restate the input schema and the existing output schema, which is mild redundancy rather than wasted length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter, zero-required analytics tool with both an output schema and full annotation coverage, the description is close to complete: purpose, timing caveat, parameter meaning, and result shape are all covered. The Returns block is redundant with the output schema, and the relationship to the snapshot-creation sibling is the one real omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are documented inline; baseline is 3. The description's gloss on min_resolved ('note if fewer than this have resolved') adds a nuance about it acting as a warning threshold rather than a hard gate, but that reading is slightly at odds with the schema's 'before computing metrics' and is not resolved.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific set of verbs and resources: resolve settled snapshots, then compute a reliability curve (predicted vs observed per decile), Brier score, and ECE, with the framing 'tells you whether the market's prices are honest.' That is far more than a restatement of the name. It does not name the sibling polymarket_calibration_snapshot, so the create-vs-score boundary is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides one genuine usage condition — 'Needs elapsed time for markets to resolve, so it's sparse early on' — which tells the agent when results are worth reading. But there is no explicit when-to-use/when-not guidance and no routing to or from polymarket_calibration_snapshot, which is the obvious complementary tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_calibration_snapshotSnapshot markets for calibrationA
Record a SNAPSHOT of current open markets' implied probabilities (one per outcome) into a local file, so that when they resolve, polymarket_calibration_report can measure whether the market is well-calibrated (do 30% markets resolve Yes ~30% of the time?). Forward-only. Writes a local file; no real money.
Args:
query (string, optional): snapshot markets about a keyword; omit for top markets by volume.
limit (1-50, default 25), min_volume (default 0, filters dead markets).
Returns: { recorded, skipped, totalSnapshots, byBucket }. Re-running the same day won't double-count.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max markets to snapshot (default 25). | |
| query | No | Keyword to snapshot markets about; omit for top markets by volume. | |
| min_volume | No | Skip markets below this total volume (filters dead markets). | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| skipped | Yes | |
| byBucket | Yes | |
| recorded | Yes | |
| totalSnapshots | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag readOnlyHint=false and destructiveHint=false; the description adds real value by clarifying 'Writes a local file; no real money,' the forward-only nature, and the dedupe behavior 'Re-running the same day won't double-count.' Note this dedupe claim sits in mild tension with idempotentHint=false, though it is scoped to same-day re-runs rather than a flat contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, followed by a compact args list and a returns line. Dense but well-organized; the parenthetical calibration explanation earns its place by clarifying intent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, usage, args, return shape, idempotency, and the downstream report relationship. Since an output schema exists, return values need not be re-explained, and the description is nearly complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters. The description restates query/limit/min_volume with light usage color ('filters dead markets', 'omit for top markets by volume') but omits response_format, adding only marginal meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Record a SNAPSHOT of current open markets' implied probabilities ... into a local file.' It also explicitly distinguishes itself from the sibling it feeds, polymarket_calibration_report, so an agent understands the relationship without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear when-to-use context: snapshot now so the later report can measure calibration, and the key constraint 'Forward-only.' It names the downstream tool but does not give explicit when-not-to-use guidance or alternatives for capturing market state.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_daily_digestPolymarket daily digestARead-onlyIdempotent
A one-shot "morning briefing": your paper-trading status, a small arbitrage scan, the day's top movers, and your calibration status — all in one call. Read-only vs Polymarket (reads the local ledgers). Meant to be run each morning (see the README for scheduling it on Windows).
Args:
arb_scan_top (1-15, default 5): top events to scan for arbitrage.
movers_limit (1-15, default 5): how many top markets to rank by 24h move.
| Name | Required | Description | Default |
|---|---|---|---|
| arb_scan_top | No | Top-N events to scan for arbitrage (default 5). | |
| movers_limit | No | How many top markets to check for 24h moves (default 5). | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| arbs | Yes | |
| paper | Yes | |
| movers | Yes | |
| calibration | Yes | |
| generatedIso | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds useful context beyond that: it reads local ledgers rather than hitting Polymarket, and points to the README for scheduling, though it doesn't discuss partial-failure behavior or cost of the embedded scans.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core value proposition and tightly written, with the arg list kept brief. The parenthetical about Windows scheduling is mild extra weight but plausibly useful for the intended daily-use workflow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need no explanation, and annotations cover the safety/behavior profile, so the description is largely complete. It could better note how the composite behaves if one sub-scan fails, but nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters with types, ranges and defaults. The description restates arb_scan_top and movers_limit but adds no new meaning and omits response_format entirely, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific composite resource and enumerates exactly the four payloads it returns (paper-trading status, arbitrage scan, top movers, calibration status) in one call, which cleanly separates it from the granular siblings like paper_status or find_arbitrage. An agent can tell it is an aggregator without inspecting any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Meant to be run each morning" gives a clear when-to-use context, and "all in one call" implies the alternative of calling the individual tools separately. It stops short of explicitly naming the sibling tools it replaces or conditions under which the granular calls are preferred, so it is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_find_arbitrageFind Polymarket arbitrageARead-onlyIdempotent
Scan for risk-free pricing gaps: a binary market whose best-ask(YES)+best-ask(NO) < $1, or a negRisk event basket whose Σ best-ask(legs) < $1. Fee-aware (reports gross AND net-of-fee edge) and depth-aware (estimates executable size). This is the one edge a read-only bot can DETECT precisely — but detection is not capture.
Args:
event_slug (string, optional) OR scan_top (1-25): scan one event, or sweep the top-N events by 24h volume.
min_edge_pct (number, default 0.5): only flag arbs whose NET edge% ≥ this.
category ('crypto'|'sports'|'politics'|'macro'|'geopolitics'|'other'): taker-fee assumption (geopolitics free).
max_legs (number, default 12): skip baskets larger than this.
Returns: { scanned, feePerShare, category, count, arbs:[{ type, title, legs, grossEdge/Pct, netEdge/Pct, executable:{sets,notionalUsd,estNetProfitUsd,bottleneckLeg} }] }.
| Name | Required | Description | Default |
|---|---|---|---|
| category | No | Taker-fee assumption bucket (geopolitics = free). | other |
| max_legs | No | Skip baskets larger than this many legs (default 12). | |
| scan_top | No | Sweep the top-N events by 24h volume. Provide this OR event_slug. | |
| event_slug | No | Scan one event's negRisk basket + its binary markets. Provide this OR scan_top. | |
| min_edge_pct | No | Only flag arbs whose NET edge (after fees) ≥ this % (default 0.5). | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| arbs | Yes | |
| count | Yes | |
| scanned | Yes | |
| category | Yes | |
| feePerShare | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuinely new behavioral context: it is fee-aware (gross AND net-of-fee edge), depth-aware (estimates executable size, bottleneck leg), and warns that detection does not equal capture. That is meaningful disclosure beyond structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the detection condition, then cleanly split into Args and Returns sections. The one prose flourish ('the one edge a read-only bot can DETECT precisely') earns its place as routing context, though the Returns block partly duplicates the output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, zero-required-param scan tool with an output schema and full annotation coverage, the description supplies both argument semantics and the shape of the returned edge objects (gross/net edge, executable size, bottleneck leg). Nothing an agent needs to invoke or interpret it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all six parameters, including the fee bucket, the OR relationship, and defaults. The description largely restates these in prose (event_slug/scan_top, min_edge_pct default 0.5, category, max_legs default 12) without adding syntax or edge-case semantics beyond what the schema says. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Scan for risk-free pricing gaps') and then defines the exact mathematical condition it detects: best-ask(YES)+best-ask(NO) < $1 or a negRisk basket whose Σ best-ask(legs) < $1. This is so specific that it self-distinguishes from every sibling, which are generic getters or paper-trading tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives the operative choice condition explicitly: 'scan one event, or sweep the top-N events by 24h volume', mirrored by the event_slug OR scan_top params. The closing 'detection is not capture' implicitly points the agent at the paper-trading siblings for execution, but it never names them or states exclusions, so it falls short of a full when/when-not routing rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_get_closed_positionsGet Polymarket closed positionsARead-onlyIdempotent
Get a wallet's CLOSED (resolved) positions with REALIZED profit/loss — the settled history behind polymarket_get_positions (which shows only open positions).
Args:
wallet (string): 0x… address (42 chars).
limit (number): max closed positions, 1-100 (default 25).
response_format ('markdown' | 'json'): default 'markdown'.
Returns: { wallet, count, totalRealizedPnl, positions:[{ title, slug, outcome, conditionId, avgPrice, totalBought, realizedPnl, curPrice }] }. An unknown wallet returns 0 positions.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max closed positions (1-100, default 25). | |
| wallet | Yes | Public wallet address (0x…, 42 chars). | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | Yes | |
| wallet | Yes | |
| positions | Yes | |
| totalRealizedPnl | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuinely useful behavior: it returns realized P/L totals and that an unknown wallet returns 0 positions rather than erroring — valuable edge-case disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the key differentiator (closed/resolved vs open) before the Args/Returns blocks. Efficient overall, though the Args list duplicates the schema slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists, the return-values detail is optional but present, and the edge-case note on unknown wallets closes the remaining ambiguity. Nothing needed to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the Args section largely restates the schema (wallet 42 chars, limit 1-100 default 25, response_format). It adds marginal meaning beyond the structured fields, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Get a wallet's CLOSED (resolved) positions') and immediately scopes it with realized P/L. It explicitly distinguishes itself from the sibling polymarket_get_positions, which 'shows only open positions'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly routes the agent: use this for settled/resolved history, use polymarket_get_positions for open positions. It does not cover exclusions (e.g., when trades or user_activity would be better), so it falls short of fully explicit when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_get_eventsGet Polymarket eventsARead-onlyIdempotent
Get Polymarket EVENTS, which group related markets under one question (e.g. a tournament or election, with a market per team/candidate).
Use this when a question spans many outcomes: one event carries all its markets. Pass a slug for one event, or omit it for the top events by 24h volume.
Args:
slug (string, optional): event slug for one event with all its markets. If omitted, returns top events.
limit (number): max events when listing, 1-20 (default 5).
active_only (boolean): only open/unresolved events (default true).
response_format ('markdown' | 'json'): default 'markdown'.
Returns: { count, events:[{ slug, title, startDate, endDate, active, closed, volume, liquidity, volume24hr, negRisk, marketCount, markets:[{ question, slug, conditionId, outcomes, volume, closed }] }] }. Errors: "No event found for slug '…'" if the slug is wrong.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | No | Event slug for a single event with all its markets. If omitted, returns top events by 24h volume. | |
| limit | No | Max events when listing (1-20, default 5). | |
| active_only | No | Only open / unresolved events (default true). | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | Yes | |
| events | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnly, idempotent, non-destructive, openWorld). The description adds real behavioral context on top: default listing behavior (top events by 24h volume when slug is omitted) and the exact error string returned for a bad slug. It doesn't discuss rate limits or pagination, keeping it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the resource definition, then usage, args, returns, and errors in a clean scannable structure. The Returns block partially duplicates the existing output schema and the Args block repeats the schema, so a small amount of redundancy keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter read tool with annotations and an output schema, the description covers purpose, usage, defaults, and error behavior. An agent has everything needed to select and invoke it correctly without further inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters, including the limit range and response_format enum. The Args block in the description restates these without adding syntax or format detail beyond the schema, so the baseline of 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb+resource (get Polymarket EVENTS) and immediately defines the abstraction: events group related markets under one question. This conceptual definition is exactly what separates it from siblings like polymarket_get_market and polymarket_search_markets, which operate on individual markets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance ('Use this when a question spans many outcomes') and explains the two invocation modes (pass a slug for one event, omit it for top events by volume). It stops short of naming an alternative tool or stating when-not to use it, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_get_leaderboardGet Polymarket trader leaderboardARead-onlyIdempotent
Get the trader LEADERBOARD — top wallets ranked by realized profit (PnL) or traded volume over a time window.
Args:
order_by ('PNL' | 'VOL'): rank by profit or volume (default 'PNL').
time_period ('DAY' | 'WEEK' | 'MONTH' | 'ALL'): window (default 'MONTH').
category (string, optional): e.g. OVERALL, POLITICS, SPORTS, CRYPTO, CULTURE, ECONOMICS, TECH, FINANCE.
limit (number): top N traders, 1-50 (default 10).
response_format ('markdown' | 'json'): default 'markdown'.
Returns: { orderBy, timePeriod, count, entries:[{ rank, wallet, name, pnl, volume }] }. This is a global ranking (unlike polymarket_get_market_holders, which is per-market).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Top N traders (1-50, default 10). | |
| category | No | Optional category, e.g. OVERALL, POLITICS, SPORTS, CRYPTO, CULTURE, ECONOMICS, TECH, FINANCE. | |
| order_by | No | Rank by profit ('PNL') or volume ('VOL'). Default 'PNL'. | PNL |
| time_period | No | Window: DAY, WEEK, MONTH or ALL (default MONTH). | MONTH |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | Yes | |
| entries | Yes | |
| orderBy | Yes | |
| timePeriod | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, openWorld), so the burden is lighter. The description adds the ranking semantics (realized PnL vs. volume) and the global-vs-market scoping behavior, though it says nothing about rate limits, auth, or pagination.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose sentence, then a clean Args block with defaults and a Returns line. No wasted prose; each line earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, yet the description still summarizes the return shape for convenience. Combined with full parameter documentation and annotation-backed safety, an agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented in the schema. The description largely restates the same defaults, enums, and category examples, adding little meaning beyond the structured fields. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get the trader LEADERBOARD') plus the exact ranking basis (realized profit or traded volume) and time window. It explicitly differentiates itself from the sibling polymarket_get_market_holders by noting this is a global ranking versus a per-market one.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly routes the agent: use this for global wallet rankings, use polymarket_get_market_holders for per-market holders. The condition that selects this tool (global vs. per-market) is stated rather than left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_get_marketGet one Polymarket marketARead-onlyIdempotent
Fetch the full detail of a single market by its slug, including the CLOB token ids for each outcome.
Use this to follow a specific market you already know, or after a search to get the exact token ids needed by the order book / price history tools.
Args:
slug (string): the market slug (from polymarket_search_markets).
response_format ('markdown' | 'json'): default 'markdown'.
Returns: { market: { slug, question, conditionId, description, startDate, endDate, negRisk, enableOrderBook, acceptingOrders, active, closed, volume, liquidity, volume24hr, eventSlug, outcomes:[{outcome, tokenId, price, impliedProbabilityPct}] } }.
Errors: "No market found for slug '…'" if the slug is wrong — verify it with polymarket_search_markets.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | Yes | The market slug, e.g. 'will-argentina-win-the-2026-fifa-world-cup-245'. Get it from polymarket_search_markets. | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| market | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so safety is covered. The description adds value beyond that with an explicit error string ('No market found for slug …') and the returned field set, though it does not discuss rate limits or pagination (none apparently needed for a single-item fetch).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then Args, Returns and Errors in a scannable structure. The verbose Returns enumeration partially duplicates the existing output schema, which is the one place words are spent that structured data already covers.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read tool with an output schema present, the description supplies everything needed: required arg source, format option, error condition and recovery path. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, and the description goes slightly beyond by explaining the provenance of slug (from polymarket_search_markets) and the markdown-vs-json tradeoff. The added meaning is marginal but real, particularly the link between slug lookup and downstream token ids.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Fetch the full detail of a single market by its slug') and names the distinguishing payload ('CLOB token ids for each outcome'), which separates it cleanly from polymarket_search_markets. An agent can tell what it retrieves without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives two explicit triggering conditions: following a market you already know, or following a search to obtain the exact token ids that the order book / price history tools require. It also routes the agent to polymarket_search_markets for slug recovery, so the when-to-use and alternative are both covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_get_market_holdersGet Polymarket market holders (whales)ARead-onlyIdempotent
Get the largest holders of each outcome in a market — the "whales" with the most exposure.
Useful to see who is positioned on each side and how concentrated a market is. Identify the market by slug (preferred, so outcomes get named) OR conditionId.
Args:
slug (string, optional): market slug. Provide this OR condition_id.
condition_id (string, optional): market conditionId (0x…). Provide this OR slug.
limit (number): top holders per outcome, 1-20 (default 5).
response_format ('markdown' | 'json'): default 'markdown'.
Returns: { conditionId, slug, question, outcomes:[{ outcome, tokenId, holders:[{ rank, name, wallet, amountShares }] }] }.
Notes: amounts are share counts (each resolves to $1 if that outcome wins). Anonymous holders show a shortened wallet. This is per-market holder ranking, not a global leaderboard. Errors: "Provide either 'condition_id' or 'slug'"; "No market found for slug…".
| Name | Required | Description | Default |
|---|---|---|---|
| slug | No | Market slug. Provide this OR 'condition_id'. | |
| limit | No | Top holders per outcome (1-20, default 5). | |
| condition_id | No | Market conditionId (0x…). Provide this OR 'slug'. | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| slug | Yes | |
| outcomes | Yes | |
| question | Yes | |
| conditionId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), so the description is free to add substantive context: amounts are share counts that resolve to $1 per winning share, anonymous holders appear as shortened wallets, and it names the exact error strings. This is meaningful behavior beyond structured fields, though it stops short of describing ordering/pagination guarantees.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose sentence, then cleanly sectioned into Args, Returns, Notes, Errors. Every block is relevant, though the Returns section duplicates what the existing output schema already provides.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the one-of slug/condition_id requirement, limit bounds, format options, the share-count semantics, and the failure modes. With annotations carrying safety and an output schema present, nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3. The description exceeds it by explaining why slug is preferred ('so outcomes get named') and by enumerating each parameter with meaning (limit = top holders per outcome, response_format tradeoff), adding rationale the schema does not carry.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (Get) plus a precise resource (largest holders of each outcome, the 'whales'). It explicitly scopes itself against the leaderboard concept ('not a global leaderboard'), letting an agent separate it from polymarket_get_leaderboard without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the analytical purpose ('see who is positioned on each side and how concentrated a market is') and gives a clear selector rule (slug preferred so outcomes get named, OR condition_id), plus a contrast with the global leaderboard. It lacks explicit when-not guidance relative to siblings like polymarket_get_market or polymarket_get_positions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_get_open_interestGet Polymarket open interestARead-onlyIdempotent
Get the OPEN INTEREST (total value of outstanding positions) for one or more markets — a measure of how much capital is currently at stake.
Identify markets by conditionId(s), or pass a slug (resolved to its conditionId).
Args:
condition_ids (string[], optional): one or more market conditionIds (0x…). Provide this OR slug.
slug (string, optional): market slug (resolved to its conditionId). Provide this OR condition_ids.
response_format ('markdown' | 'json'): default 'markdown'.
Returns: { count, totalValue, openInterest:[{ market, value }] }. Errors: "Provide either 'condition_ids' or 'slug'."; "No market found for slug…".
| Name | Required | Description | Default |
|---|---|---|---|
| slug | No | Market slug (resolved to its conditionId). Provide this OR condition_ids. | |
| condition_ids | No | One or more market conditionIds (0x…). Provide this OR slug. | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | Yes | |
| totalValue | Yes | |
| openInterest | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and open-world behavior, so the safety profile is covered. The description adds useful context beyond the annotations: the semantic meaning of open interest, the mutually exclusive input requirement, and the exact validation error messages an agent may encounter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the concept definition before args/returns/errors, and every section is compact. The Args list duplicates the schema descriptions almost verbatim, which is the only redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers concept, input modes, output format, return shape, and error cases; an output schema exists so the return-shape mention is a bonus rather than a necessity. The only gap is the absence of guidance on when to prefer this over related market-data siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already documents the OR-relationship for slug/condition_ids and the response_format enum, so the Args block largely restates structured data. It adds minor value by repeating the exact error strings and clarifying that slug is resolved to a conditionId, but the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Get the OPEN INTEREST ... for one or more markets') and immediately parenthesizes the concept ('total value of outstanding positions'), which disambiguates it from siblings like get_market_holders, get_orderbook, and get_positions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states how to identify the target markets (condition_ids OR slug) and that slug is resolved to a conditionId, but it gives no when-to-use/when-not guidance relative to sibling tools such as get_market_holders or get_market, leaving the agent to infer the use case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_get_orderbookGet Polymarket order bookARead-onlyIdempotent
Get the live order book for one market outcome: best bid/ask, spread, mid price, implied probability, and top-of-book depth.
This is where you check whether an edge is real: the implied probability (mid) is the market's estimate, while best ask is what you'd actually pay to buy and best bid what you'd receive to sell. A wide spread or thin depth means the "price" you saw elsewhere may not be executable.
Identify the outcome either by token_id (preferred) OR by slug + outcome.
Args:
token_id (string, optional): the ~77-digit CLOB token id for one outcome.
slug (string, optional) + outcome (string, optional): e.g. slug + "Yes". Outcome defaults to the first.
depth (number): price levels per side to include, 1-50 (default 10).
response_format ('markdown' | 'json'): default 'markdown'.
Returns: { tokenId, outcome, slug, question, bestBid, bestAsk, mid, spread, impliedProbabilityPct, spreadPct, totalBidSize, totalAskSize, bids:[{price,size}], asks:[{price,size}], timestamp }.
Errors: "No market found for slug…" (bad slug); an empty side is reported as bestBid/bestAsk = null.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | No | Market slug — alternative to token_id. Combine with 'outcome'. | |
| depth | No | How many price levels per side to include (1-50, default 10). | |
| outcome | No | Outcome name to select when using 'slug' (e.g. 'Yes' or 'No'). Defaults to the first outcome. | |
| token_id | No | CLOB token id (the ~77-digit id for ONE outcome). Preferred when you already have it. | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| mid | Yes | |
| asks | Yes | |
| bids | Yes | |
| slug | Yes | |
| spread | Yes | |
| bestAsk | Yes | |
| bestBid | Yes | |
| outcome | Yes | |
| tokenId | Yes | |
| question | Yes | |
| spreadPct | Yes | |
| timestamp | Yes | |
| totalAskSize | Yes | |
| totalBidSize | Yes | |
| impliedProbabilityPct | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive/openWorld, so the safety profile is covered. The description adds genuine behavioral value beyond that: the error case ('No market found for slug…') and the edge case that an empty side surfaces as bestBid/bestAsk = null, which an agent cannot infer from annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is front-loaded with the purpose, then interpretation, then args/returns/errors in labeled blocks, so an agent can scan it quickly. It is slightly longer than strictly necessary, with some sentences (the economic interpretation) edging toward prose rather than specification.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Return fields are enumerated and an output schema also exists, errors and null-side behavior are documented, and all five parameters are described. Nothing an agent needs to invoke it correctly appears to be missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description still adds value by stating the token_id vs. slug+outcome selection rule ('either... OR') and marking token_id as preferred, a relationship the flat schema (no oneOf) does not express; the other params largely restate the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence names a specific verb and resource ('Get the live order book for one market outcome') and enumerates exactly what it contains (best bid/ask, spread, mid, implied probability, depth). The follow-up sentence implicitly differentiates it from price/quote siblings by framing it as the place where a quoted 'price' is validated as executable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear use context: 'This is where you check whether an edge is real' and explains the interpretation of mid vs. best ask/bid and what a wide spread or thin depth implies. It does not explicitly name alternatives (e.g. get_quote, get_price_history) or state when not to use this tool, so it falls short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_get_positionsGet Polymarket wallet positionsARead-onlyIdempotent
Get the open positions and P&L for a public wallet address ("how is this wallet doing").
Returns each position's entry price, current price, current value and profit/loss, plus the wallet's total portfolio value. Read-only and public — anyone's positions are visible by address.
Args:
wallet (string): 0x… address (42 chars). The on-chain proxy wallet, not a username.
limit (number): max positions, 1-100 (default 25).
min_value (number): hide positions worth less than this USD (default 1, filters dust).
redeemable_only (boolean): only resolved positions ready to redeem (default false).
response_format ('markdown' | 'json'): default 'markdown'.
Returns: { wallet, portfolioValue, openPositions, totalCurrentValue, totalCashPnl, positions:[{ title, slug, outcome, conditionId, tokenId, size, avgPrice, curPrice, initialValue, currentValue, cashPnl, percentPnl, realizedPnl, redeemable, endDate }] }.
Errors: an invalid address is rejected before any request; an unknown wallet returns 0 positions.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max positions to return (1-100, default 25). | |
| wallet | Yes | Public wallet address (0x…, 42 chars). This is the on-chain proxy wallet, not a username. | |
| min_value | No | Hide positions whose current value is below this USD amount (default $1, filters dust). | |
| redeemable_only | No | Only show resolved positions that can be redeemed (default false). | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| wallet | Yes | |
| positions | Yes | |
| totalCashPnl | Yes | |
| openPositions | Yes | |
| portfolioValue | Yes | |
| totalCurrentValue | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, openWorld, non-destructive), and the description adds genuinely useful context: public by address, invalid address rejected before any request, unknown wallet returns 0 positions, and dust filtering defaults. It doesn't cover pagination or rate limits, so it's short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose, then cleanly sectioned into Args and Returns. It is efficient, though the Args and Returns blocks largely duplicate the schema and output schema, costing a little redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, all parameters, error conditions, and default behaviors; an output schema exists, so the extra Returns block is optional but harmless. Nothing needed to invoke this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter including defaults, ranges and the proxy-vs-username caveat. The Args section essentially restates the schema without adding new semantics, which is the baseline 3 case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (get open positions and P&L for a wallet) plus the practical intent ("how is this wallet doing"). The word "open" implicitly distinguishes it from the sibling polymarket_get_closed_positions, and "public wallet address" separates it from activity/market tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly frames the use case (checking how an arbitrary public wallet is doing) and notes read-only/public visibility, which tells the agent it is safe to call on any address. However, it never names an alternative (e.g. get_closed_positions or get_user_activity) or states when NOT to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_get_price_historyGet Polymarket price historyARead-onlyIdempotent
Get the historical price (implied probability over time) for one market outcome, plus summary stats.
Use this to see how a market moved and where you might have entered or exited. Identify the outcome by token_id (preferred) OR slug + outcome.
Args:
token_id (string, optional): the ~77-digit CLOB token id.
slug (string, optional) + outcome (string, optional): e.g. slug + "Yes".
interval ('1h'|'6h'|'1d'|'1w'|'1m'|'max'): look-back window (default '1w').
fidelity (number, optional): minutes between points.
max_points (number): downsample to at most this many points (default 150).
response_format ('markdown' | 'json'): default 'markdown'.
Returns: { tokenId, outcome, interval, count, truncated, summary:{first,last,min,max,changeAbs,changePct,from,to}, points:[{t, iso, p}] }. Prices are 0-1 (probability). 'count' is the number of returned points; 'truncated' indicates downsampling.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | No | Market slug — alternative to token_id. Combine with 'outcome'. | |
| outcome | No | Outcome name to select when using 'slug' (e.g. 'Yes' or 'No'). Defaults to the first outcome. | |
| fidelity | No | Resolution in minutes between points (optional; the API picks a sensible default). | |
| interval | No | Look-back window: 1h, 6h, 1d, 1w, 1m or max (default 1w). | 1w |
| token_id | No | CLOB token id (the ~77-digit id for ONE outcome). Preferred when you already have it. | |
| max_points | No | Downsample the series to at most this many points (default 150). | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | Yes | |
| points | Yes | |
| outcome | Yes | |
| summary | Yes | |
| tokenId | Yes | |
| interval | Yes | |
| truncated | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations already covering safety (readOnly, idempotent, destructive=false, openWorld), the description adds meaningful behavioral context: prices are 0-1 probabilities, 'count' is the number of returned points, and 'truncated' signals downsampling. This goes beyond the annotations, though it leans on the return section rather than describing operational traits like rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well front-loaded: purpose, then usage, then Args, then Returns. The Args block repeats most of the 100%-covered schema, which is somewhat wasteful, but the layout is scannable and each section stays compact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter read tool, the definition is complete: an output schema exists and the description still summarizes the return shape and downsampling semantics, annotations cover the safety profile, and all parameters (including defaults and the identification alternatives) are accounted for. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds genuine meaning the schema lacks: it frames token_id vs slug+outcome as alternatives (the schema marks all params optional with no such relationship) and explicitly calls token_id 'preferred'. The remaining args largely restate schema text and defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Get'), a precise resource ('historical price ... for one market outcome') and an additional deliverable ('plus summary stats'), with the parenthetical 'implied probability over time' clarifying the domain. This is clearly distinguished from siblings like get_market or get_quote that deal with current state rather than a time series.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states a concrete use case ('see how a market moved and where you might have entered or exited'), which gives clear context for when to reach for it. However, it names no alternatives and offers no when-not guidance versus related tools such as get_market or get_quote.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_get_quoteGet a quick Polymarket quoteARead-onlyIdempotent
Get a lightweight live quote for one market outcome — midpoint, best buy/sell price, spread and last trade — WITHOUT pulling the full order book.
Faster than polymarket_get_orderbook when you just want "what's this worth right now". Identify the outcome by token_id (preferred) OR slug + outcome.
Args:
token_id (string, optional): the ~77-digit CLOB token id.
slug (string, optional) + outcome (string, optional): e.g. slug + "Yes". Outcome defaults to the first.
response_format ('markdown' | 'json'): default 'markdown'.
Returns: { tokenId, outcome, slug, question, midpoint, buyPrice, sellPrice, spread, spreadPct, impliedProbabilityPct, lastTradePrice, lastTradeSide }. Prices are 0-1 (midpoint ≈ implied probability; buyPrice is what you'd pay, sellPrice what you'd receive). Any value the book can't provide is null. For full depth, use polymarket_get_orderbook.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | No | Market slug — alternative to token_id. Combine with 'outcome'. | |
| outcome | No | Outcome name to select when using 'slug' (e.g. 'Yes' or 'No'). Defaults to the first outcome. | |
| token_id | No | CLOB token id (the ~77-digit id for ONE outcome). Preferred when you already have it. | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| slug | Yes | |
| spread | Yes | |
| outcome | Yes | |
| tokenId | Yes | |
| buyPrice | Yes | |
| midpoint | Yes | |
| question | Yes | |
| sellPrice | Yes | |
| spreadPct | Yes | |
| lastTradeSide | Yes | |
| lastTradePrice | Yes | |
| impliedProbabilityPct | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety bar is low. The description adds genuinely useful behavioral context beyond that: prices are 0-1 with midpoint≈probability, buyPrice vs sellPrice meaning, and that any value the book can't provide returns null.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and the sibling contrast before the arg/return detail, and each section earns its place. Slightly long because the Args and Returns blocks largely mirror the schema and output schema, which is mildly redundant but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values need not be explained, yet the definition still provides them along with value semantics and the null-on-missing rule. For a low-complexity quote tool, nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters including the 'preferred' note on token_id, the outcome default, and the response_format enum. The description restates these with minor added framing (token_id=preferred, outcome defaults to first), landing at the baseline rather than adding substantial new meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('get a lightweight live quote for one market outcome') and explicitly enumerates the returned fields (midpoint, best buy/sell, spread, last trade). It distinguishes itself from polymarket_get_orderbook by scope, so an agent can tell the two apart without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative and the condition that selects it: 'Faster than polymarket_get_orderbook when you just want what's this worth right now,' and the reverse routing 'For full depth, use polymarket_get_orderbook.' When-to-use and when-to-use-something-else are both covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_get_tradesGet Polymarket tradesARead-onlyIdempotent
Get recent trades — either a MARKET's tape (who traded, at what price) or a WALLET's trade history.
Provide market (a conditionId) for the market tape, OR wallet for that address's trades.
Args:
market (string, optional): market conditionId (0x…). Provide this OR wallet.
wallet (string, optional): 0x… address. Provide this OR market.
taker_only (boolean): only taker trades (default true); false adds maker fills for the full tape.
limit (number): max trades, 1-100 (default 20).
response_format ('markdown' | 'json'): default 'markdown'.
Returns: { market, user, count, trades:[{ wallet, name, side, outcome, price, size, usdcSize, timestamp, iso, conditionId, title, slug, transactionHash }] }. Errors: "Provide either 'market' or 'wallet'." if neither is given.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max trades to return (1-100, default 20). | |
| market | No | Market conditionId (0x…) for that market's recent trade tape. Provide this OR wallet. | |
| wallet | No | Wallet address for that wallet's trade history. Provide this OR market. | |
| taker_only | No | Only taker trades (default true); false adds maker fills for the full tape. | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| user | Yes | |
| count | Yes | |
| market | Yes | |
| trades | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds useful behavioral context: the taker-only default and what false adds (maker fills), the limit range, defaults, and the exact error string.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose sentence followed by a structured Args list and Returns/Errors sections, easy to scan. The Args section duplicates the schema somewhat, which costs a bit of efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists, return values needn't be re-explained, yet the description still sketches them alongside documented error behavior, defaults, and the mutual-exclusivity rule. Nothing an agent needs to invoke correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter including the 'OR' constraint and taker_only. The description's Args section largely restates the schema, adding only marginal clarifying value, which anchors it at the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get recent trades') and immediately disambiguates the two modes — a market's tape vs a wallet's trade history. This distinguishes it from siblings like get_user_activity and get_positions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the selection rule: 'Provide market for the market tape, OR wallet for that address's trades,' and the error case when neither is given. It does not explicitly name a sibling alternative for trade history vs activity, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_get_user_activityGet Polymarket wallet activityARead-onlyIdempotent
Get a wallet's full on-chain ACTIVITY feed — beyond open positions: trades plus splits, merges, redeems, rewards, deposits and withdrawals.
Complements polymarket_get_positions (current holdings) and polymarket_get_trades (just trades) with the complete history.
Args:
wallet (string): 0x… address (42 chars).
type (string[], optional): filter to types like TRADE, SPLIT, MERGE, REDEEM, REWARD, DEPOSIT, WITHDRAWAL (default: all).
limit (number): max entries to request, 1-500 (default 50). At most 200 are returned.
response_format ('markdown' | 'json'): default 'markdown'.
Returns: { wallet, count, truncated, activities:[{ type, timestamp, iso, side, outcome, size, usdcSize, price, conditionId, title, slug, transactionHash }] }. The 'truncated' flag is true when more entries existed than were returned. An unknown wallet returns 0 activities.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Filter to these activity types (default: all). | |
| limit | No | Max activity entries to request (1-500, default 50). At most 200 are returned; past that the response is capped and flagged with `truncated`. | |
| wallet | Yes | Public wallet address (0x…, 42 chars). | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | Yes | |
| wallet | Yes | |
| truncated | Yes | |
| activities | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, openWorld), so the bar is lower. The description still adds genuinely useful behavior beyond that: the 200-entry return cap with a 'truncated' flag, and that an unknown wallet returns 0 activities rather than erroring.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core scope and sibling differentiation, then uses clean Args/Returns sections. Every line carries information and nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read tool with full annotations and an output schema, the description covers scope, filtering, pagination/cap behavior, and error-case behavior (unknown wallet). Nothing an agent needs in order to call it correctly is missing, though the inline Returns block is mildly redundant given the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters, and baseline 3 applies. The description largely repeats the schema, and its type list (TRADE, SPLIT, MERGE, REDEEM, REWARD, DEPOSIT, WITHDRAWAL) is actually a subset of the schema enum, which also includes CONVERSION, MAKER_REBATE, TAKER_REBATE, REFERRAL_REWARD, and YIELD.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Get a wallet's full on-chain ACTIVITY feed') and explicitly scopes it beyond the two closest siblings, naming them: 'Complements polymarket_get_positions (current holdings) and polymarket_get_trades (just trades) with the complete history.' An agent can tell exactly what this returns versus the alternatives without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context and names the alternatives: use get_positions for current holdings and get_trades for trades only, and this tool when the complete history is needed. It stops short of explicit when-not guidance or prerequisites, so it falls just below the top bar.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_paper_closeClose a SIMULATED (paper) bet earlyA
Manually close an OPEN paper bet at the current price (models selling before the market resolves). Fills at the executable BID minus the taker fee. No real money.
Args:
id (string): the paper bet id (from polymarket_paper_status).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The paper bet id to close (from polymarket_paper_status). | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| bet | Yes | |
| bankroll | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare not-read-only, not-idempotent, not-destructive, closed-world. The description adds genuinely non-obvious mechanics: fills at the executable BID minus the taker fee, and that no real money is involved. That pricing/fee disclosure goes beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and fill mechanics in two tight sentences, then a short args note. Minor redundancy in repeating the id parameter already covered by the schema, but no wasted prose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and annotations cover the safety profile. The description supplies the fill model and the no-real-money caveat, which is sufficient for a simulated-close tool, though it omits error behavior when the bet is not OPEN.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both id and response_format are documented in the schema itself; the description merely restates the id source. Baseline 3 applies since the schema carries the parameter burden and the description adds no new syntax or format detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (close), resource (OPEN paper bet), and mechanism (at the current price, selling before resolution). It is clearly distinguishable from the sibling paper_open, paper_status, paper_delete, and paper_reset tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies the precondition that the bet must be OPEN and directs the agent to polymarket_paper_status for the id, which is useful context. However, it never says when to choose this over polymarket_paper_delete or polymarket_paper_reset, leaving the alternative selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_paper_deleteDelete a SIMULATED (paper) betADestructive
Permanently remove one bet from the local paper-trading ledger. Works on any status (open, won, lost, closed).
This rewrites simulated history — a deleted resolved bet stops counting towards realized P&L, ROI, hit rate and Brier, so your track record will change. To stop tracking a bet while keeping an honest record, use polymarket_paper_close instead. No real money is involved either way.
Args:
id (string): the paper bet id (from polymarket_paper_status).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The paper bet id to delete permanently (from polymarket_paper_status). | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| bet | Yes | |
| bankroll | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, but the description goes well beyond them: it warns the action 'rewrites simulated history' and enumerates the downstream effects on realized P&L, ROI, hit rate and Brier, and clarifies no real money is involved. That is exactly the consequence-level context an agent needs before an irreversible write and cannot get from the annotation block.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The destructive warning and the alternative route are front-loaded and each sentence does work. The trailing 'Args' block restates the schema's id description, which is mild redundancy rather than bloat, so it falls just short of a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need no explanation, and the description already covers scope, permitted statuses, side effects, and the safer sibling alternative. Nothing needed to invoke this destructive tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baselines sit at 3; the description only repeats that id comes from polymarket_paper_status, which the schema already states, and never mentions the response_format parameter. It adds no syntax, format, or constraint detail beyond the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb ('Permanently remove') and a precise scope ('one bet from the local paper-trading ledger'), and the status clause ('any status: open, won, lost, closed') bounds it. It is clearly distinguishable from polymarket_paper_close and polymarket_paper_reset without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit when-not and names the alternative: 'To stop tracking a bet while keeping an honest record, use polymarket_paper_close instead.' Combined with 'Works on any status,' the agent has both the selection condition and the sibling route spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_paper_openOpen a SIMULATED (paper) Polymarket betA
Record a SIMULATED bet in the local paper-trading ledger. No real money, no order is placed — this only writes a local file to track a hypothesis and score it later.
Fills honestly at the executable ASK plus the category taker fee (Polymarket charges takers since 2026-03-30), so the simulated P&L isn't a fantasy. Prefer slug + outcome so the conditionId is captured for auto-scoring at resolution.
Args:
token_id (string, optional) OR slug (string) + outcome (string).
size_usdc (number): simulated stake (virtual money).
p_estimate (number 0-1, optional): your probability for this outcome (for Brier scoring + the Kelly hint).
category ('crypto'|'sports'|'politics'|'macro'|'geopolitics'|'other'): taker-fee bucket.
note (string, optional): the hypothesis you're testing.
Returns: the recorded bet + a fractional-Kelly sizing hint (full & half) + the virtual bankroll.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | Optional note, e.g. the hypothesis being tested. | |
| slug | No | Market slug — alternative to token_id. Combine with 'outcome'. | |
| outcome | No | Outcome name to select when using 'slug' (e.g. 'Yes' or 'No'). Defaults to the first outcome. | |
| category | No | Category for the taker-fee model (crypto highest, geopolitics free). | other |
| token_id | No | CLOB token id (the ~77-digit id for ONE outcome). Preferred when you already have it. | |
| size_usdc | Yes | SIMULATED stake in USDC (virtual money — no real trade is placed). | |
| p_estimate | No | Your probability estimate 0-1 for this outcome (used later for Brier/calibration scoring). | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| bet | Yes | |
| bankroll | Yes | |
| kellyFull | Yes | |
| kellyHalf | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations already declaring the mutation/safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false), the description adds real behavioral detail beyond them: no order is placed, only a local file is written, fills occur at the executable ASK plus the category taker fee (dated 2026-03-30), and auto-scoring is set up via conditionId. It stops short of disclosing duplicate-token_id behavior or overwrite semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is reasonably front-loaded, but the Args block duplicates the already-complete schema and the simulation caveat is stated twice ("No real money, no order is placed" and "only writes a local file"), which exceeds what earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no elaboration, and for a local 8-parameter write tool the description covers the key facts (simulated nature, fill/fee model, preferred identifier pair). Minor gaps remain, such as duplicate-bet handling and whether repeated calls accumulate or overwrite.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter's meaning; the description's Args section largely restates that. The only genuine additions are the token_id-vs-slug+outcome precedence and the auto-scoring rationale for preferring slug+outcome, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Record a SIMULATED bet in the local paper-trading ledger") and immediately scopes it against reality ("No real money, no order is placed"), which cleanly distinguishes it from sibling read tools and from paper_close/paper_delete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description conveys the context (track a hypothesis and score it later) and gives argument-level preference ("Prefer slug + outcome..."), but it never states explicit when-to-use/when-not conditions nor names an alternative sibling to route to, so the usage guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_paper_resetReset the SIMULATED paper-trading ledgerADestructiveIdempotent
Wipe the local paper-trading ledger and start a fresh virtual bankroll. Erases every simulated bet, including resolved ones — realized P&L, ROI, hit rate and Brier all go back to zero. There is no undo.
Requires confirm=true; without it the call is rejected so the ledger can't be wiped by accident. No real money is involved — this only rewrites a local file.
Args:
confirm (boolean): must be true to proceed.
bankroll (number, optional): new virtual starting bankroll (default 1000).
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | No | Must be true to proceed — this erases every SIMULATED bet in the local ledger. | |
| bankroll | No | New virtual starting bankroll (default 1000). Virtual money — nothing real is moved. | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| bankroll | Yes | |
| deletedCount | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (destructiveHint=true, idempotentHint=true) by disclosing there is no undo, exactly what is destroyed (simulated bets including resolved ones, P&L, ROI, hit rate, Brier), the confirm gate that prevents accidental wipes, and that only a local file is rewritten. This is rich, decision-relevant behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the destructive warning and the no-undo consequence before any parameter detail, which is the right ordering. The trailing Args list somewhat duplicates an already-complete schema, so there is minor redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no prose, and the description covers the destructive semantics, confirmation requirement, and scope of erasure. Nothing an agent needs to invoke this safely is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is already 100%, and the Args block largely restates what the schema documents (confirm must be true, bankroll default 1000). It adds no new syntax, bounds, or interaction detail beyond the structured fields, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (wipe/reset) and resource (local SIMULATED paper-trading ledger), plus the effect (fresh virtual bankroll). It is clearly distinguishable from siblings like polymarket_paper_delete, which would remove individual bets, since this erases the entire ledger.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: this is a destructive full reset that requires confirm=true and involves no real money. However it does not explicitly name when to prefer polymarket_paper_delete or polymarket_paper_close instead, so the alternative-routing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_paper_statusPaper-trading status & scoringA
Show the SIMULATED paper-trading ledger: open bets marked-to-market, resolved bets scored automatically (a resolved market pays the winning outcome $1), and totals — realized/unrealized P&L net of fees, ROI, hit rate, and Brier scores (yours vs the market's implied probability at entry) on your probability estimates. No real money anywhere.
Args:
response_format ('markdown' | 'json'): default 'markdown'.
| Name | Required | Description | Default |
|---|---|---|---|
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| bets | Yes | |
| roiPct | Yes | |
| hitRate | Yes | |
| bankroll | Yes | |
| ourBrier | Yes | |
| openCount | Yes | |
| marketBrier | Yes | |
| realizedPnl | Yes | |
| totalStaked | Yes | |
| resolvedCount | Yes | |
| unrealizedPnl | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint=false, idempotentHint=false, openWorldHint=true, destructiveHint=false), the description discloses the scoring model ('a resolved market pays the winning outcome $1'), that P&L is net of fees, and that Brier scores compare user estimates against market implied probability at entry. It tacitly signals a side effect ('resolved bets scored automatically') consistent with readOnlyHint=false, though it never states that calling this tool may settle/mutate ledger state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The lead sentence is dense but front-loaded with the core purpose and scope, and the metrics list earns its place. The separate Args section duplicates the schema's response_format documentation, which is minor waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described, and the description still explains the scoring rules and metric definitions an agent needs to interpret results. The only gap is an explicit statement that the call may trigger settlement writes, which matters given readOnlyHint=false.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single enum parameter is fully documented in the schema, so the baseline is 3. The description's Args block merely repeats the enum values and default already present in the structured schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Show') and resource ('SIMULATED paper-trading ledger') and enumerates exactly what is displayed: open bets marked-to-market, resolved bets, totals. The word SIMULATED and 'No real money anywhere' cleanly separate it from the real-position siblings like polymarket_get_positions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the resource name — an agent can infer this is the read/report view of the paper_* family. There is no explicit statement of when to call it (e.g. after paper_open or paper_close) nor any named alternative for users wanting real positions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_search_marketsSearch Polymarket marketsARead-onlyIdempotent
Discover Polymarket markets by keyword, or list the highest-volume markets.
Use this first to find a market's slug and its current implied probability, then feed the slug into the other tools.
Args:
query (string, optional): keyword(s) like 'bitcoin' or 'election'. If omitted, returns the top markets by volume.
limit (number): max markets, 1-50 (default 10).
active_only (boolean): only open/unresolved markets (default true).
response_format ('markdown' | 'json'): default 'markdown'.
Returns: { count, query, markets: [{ slug, question, conditionId, outcomes:[{outcome, tokenId, price, impliedProbabilityPct}], volume, liquidity, volume24hr, endDate, active, closed, acceptingOrders, eventSlug }] }.
Examples:
"What markets are there about the Fed?" -> query="Fed".
"Show me the biggest markets right now" -> no query. Note: 'price' is the implied probability of that outcome (0-1). For an executable price/spread, use polymarket_get_orderbook.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max markets to return (1-50, default 10). | |
| query | No | Keyword(s) to search, e.g. 'bitcoin' or 'election'. If omitted, returns the top markets by volume. | |
| active_only | No | Only include markets that are open / not yet resolved (default true). | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | Yes | |
| query | Yes | |
| markets | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/openWorldHint, so the safety profile is covered. The description adds real semantic context beyond that: price is the implied probability (0-1), not an executable quote, and it directs to the orderbook for tradeable prices. It does not elaborate on pagination or rate limits, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and usage, then args and examples. The 'Returns' block largely duplicates the existing output schema and is the one section that does not fully earn its place, but overall structure is tight and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given four optional params, 100% schema coverage, and an output schema, the definition covers everything an agent needs: mode selection, slug handoff, price interpretation, and format choice. No material gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description adds value with concrete query examples ('bitcoin', 'election'), the omitted-query fallback behavior, and a worked example ('Fed'), which exceeds the schema's field-level text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (discover/search) and resource (Polymarket markets), plus the dual mode (keyword search vs. top-by-volume listing). The line 'Use this first to find a market's slug' explicitly positions it against siblings like polymarket_get_market and polymarket_get_orderbook.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit sequencing guidance ('use this first ... then feed the slug into the other tools') and an explicit alternative with its trigger condition ('For an executable price/spread, use polymarket_get_orderbook'). This tells the agent when to use it and when to route elsewhere.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
polymarket_xrayX-ray a Polymarket marketARead-onlyIdempotent
The one-shot deep look at a market: detail + executable quote + whales + recent price history + recent trade tape, all in one call (several API calls bundled). Identify by token_id OR slug + outcome.
Args:
token_id (string, optional) OR slug (string) + outcome (string).
history_interval ('1h'|'6h'|'1d'|'1w'|'1m'|'max'): default '1d'.
trade_limit (1-50, default 10), holder_limit (1-20, default 5).
Returns: { tokenId, outcome, slug, question, detail, quote, whales, history, tape } (sections null-tolerant).
| Name | Required | Description | Default |
|---|---|---|---|
| slug | No | Market slug — alternative to token_id. Combine with 'outcome'. | |
| outcome | No | Outcome name to select when using 'slug' (e.g. 'Yes' or 'No'). Defaults to the first outcome. | |
| token_id | No | CLOB token id (the ~77-digit id for ONE outcome). Preferred when you already have it. | |
| trade_limit | No | Recent trades to include (default 10). | |
| holder_limit | No | Top holders per outcome (default 5). | |
| response_format | No | Output format: 'markdown' (concise, human-readable; default) or 'json' (full structured data). | markdown |
| history_interval | No | Price-history window (default 1d). | 1d |
Output Schema
| Name | Required | Description |
|---|---|---|
| slug | Yes | |
| tape | Yes | |
| quote | Yes | |
| detail | Yes | |
| whales | Yes | |
| history | Yes | |
| outcome | Yes | |
| tokenId | Yes | |
| question | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnly, idempotent, openWorld, non-destructive). The description adds meaningful behavioral context: sections are 'null-tolerant', the call bundles several API calls, and defaults are declared. This goes beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then structured Args/Returns. Efficient, though it partially duplicates the schema's own parameter listing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, yet the description still summarizes the return shape and notes null tolerance. Combined with the Args block and annotations, the agent has everything needed to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema documents all 7 parameters including enums and ranges. The description restates defaults and the token_id/slug+outcome relationship but adds little beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('one-shot deep look') and resource ('a market'), and enumerates exactly what it bundles: detail, quote, whales, history, tape. This clearly distinguishes it from the sibling atomic tools like polymarket_get_market or polymarket_get_quote—it's the composite call.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the composite nature ('several API calls bundled') and the two identification paths (token_id OR slug+outcome), which implicitly tells the agent when to prefer it over single-purpose siblings. It doesn't explicitly say 'use this instead of calling the individual tools', so 4 rather than 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
23 tool updates
v1.0.0- First observed
polymarket_calibration_report - First observed
polymarket_calibration_snapshot - First observed
polymarket_daily_digest - First observed
polymarket_find_arbitrage - First observed
polymarket_get_closed_positions - First observed
polymarket_get_events - First observed
polymarket_get_leaderboard - First observed
polymarket_get_market - First observed
polymarket_get_market_holders - First observed
polymarket_get_open_interest - First observed
polymarket_get_orderbook - First observed
polymarket_get_positions - First observed
polymarket_get_price_history - First observed
polymarket_get_quote - First observed
polymarket_get_trades - First observed
polymarket_get_user_activity - First observed
polymarket_paper_close - First observed
polymarket_paper_delete - First observed
polymarket_paper_open - First observed
polymarket_paper_reset - First observed
polymarket_paper_status - First observed
polymarket_search_markets - First observed
polymarket_xray
TDQS
Scored across 23 tools
Most tools target clearly distinct resources and actions (market vs event vs orderbook vs positions). Composites like polymarket_xray and polymarket_daily_digest bundle existing tools but are explicitly described as shortcuts. Slight potential confusion between polymarket_get_quote and polymarket_get_orderbook, but descriptions clarify.
All tools share the polymarket_ prefix and snake_case formatting. Most follow a verb_noun pattern (search_markets, get_market, find_arbitrage), but there are mixed verb styles (get_ vs search_ vs paper_ vs xray) and a few noun-first names like calibration_snapshot.
At 23 tools, the set is above the typical 3–15 sweet spot and feels heavy for a read-only/paper-trading server. While each tool has a plausible role, several could be consolidated (e.g. paper status/close/delete/reset, calibration snapshot/report) without losing core functionality.
The surface covers market discovery, detailed data, user analytics, paper trading, arbitrage detection, calibration, and a daily digest. Missing real order execution, order cancellation, and real-time streaming, but those appear intentionally out of scope given the explicit no-real-money design.
Maintenance
Related MCP Connectors
Read-only MCP server for live Polymarket, Kalshi, Limitless odds; Manifold sentiment.
Polymarket MCP — prediction-market data via Gamma + CLOB public APIs.
Hosted MCP for Kalshi prediction markets: search, odds, order books, settlement rules, and trading.
Live prices, perps, prediction markets and a paper trading desk over one MCP.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceA Model Context Protocol (MCP) server for Polymarket prediction markets, providing real-time market data, prices, and AI-powered analysis tools for Claude Desktop integration.48MIT
- FlicenseAqualityBmaintenanceAI-agent ready FastMCP server for Polymarket market discovery, wallet analytics, and public CLOB data, providing a read-only interface for querying markets, wallets, and order books.22-
- AlicenseAqualityCmaintenanceA read-only MCP server exposing Polymarket's public prediction-market data. Search markets, read live odds and order books, pull historical probability time-series, and inspect public wallet positions.1424 PyPIMIT
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol server that enables seamless integration with Polymarket, providing tools to search markets, fetch events, analyze leaderboards, query user activity, and more via MCP-compatible clients like Claude.30 npmMIT