Skip to main content
Glama

Oracle3

Oracle3 is an open-source trading engine and MCP server for prediction markets. It maps the logical relations between event contracts, finds prices that break the axioms of probability after each venue's fees, and trades them live on Kalshi, Polymarket and Solana, or on paper, under pre-trade risk limits.

Tests PyPI License DOI

Trades live on Kalshi, Polymarket and Solana. oracle3 live run executes with the same engine that runs paper trading, behind pre-trade risk limits and a kill switch. AI agents plug in through a 13-tool MCP server.

At a glance

Venues

Kalshi, Polymarket and Solana (DFlow)

Execution

Live and paper on one engine, with pre-trade risk limits, a kill switch and Jito bundle submission on Solana

Relations checked

implication, exclusivity, complement, same event across venues, event sum

Costs

Each market's own fee schedule from the venue API (Kalshi taker 0.07·M·C·P·(1−P); Polymarket taker rate·C·p·(1−p))

Strategies

6 constraint-based, 2 statistical-arbitrage, 2 model-driven

Agent interfaces

MCP server with 13 tools, JSON CLI, 6 agent skills, Python API

Tests

600+, with ruff, mypy and codespell in CI

Install

pip install oracle3

License

Apache-2.0; the original U Lab portions are MIT (see NOTICE)

Related MCP server: pmxt-mcp

What problem does it solve?

Contracts on related outcomes are tied together by probability. If A implies B, then P(A) ≤ P(B). If A and B cannot both happen, P(A) + P(B) ≤ 1. The outcomes of one event sum to one. Quoted prices break these bounds, within a venue and across venues, and a basket of contracts that pays a known amount in every state can then be bought for less than that amount.

The gaps are small, and both venues charge taker fees that scale with p(1 − p). Whether a gap is worth anything depends on the fee on every leg of the basket. Oracle3 does three things with that:

  1. Relations. It records which markets are related and how (implication, exclusivity, complement, same event, event sum).

  2. Checks. For each relation it finds the cheapest basket at executable prices, prices every leg under that market's own fee schedule, and reports the edge before and after fees.

  3. Execution. It trades the baskets that survive, live or on paper, under position, drawdown and exposure limits, with a kill switch.

How do I run it?

pip install oracle3

# Find markets (JSON output for scripts and agents)
oracle3 market search --exchange kalshi --query "fed" --json
oracle3 market search --exchange polymarket --query "fed decision" --json

# Start the MCP server over stdio
oracle3 mcp

# Trade live on Kalshi (key via KALSHI_API_KEY_ID and KALSHI_PRIVATE_KEY_PATH)
oracle3 live run --exchange kalshi --monitor \
  --strategy-ref oracle3.strategy.contrib.implication_arb_strategy:ImplicationArbStrategy

From Python:

from oracle3.arbitrage import Quote, check_constraint
from oracle3.fees import KalshiSchedule

# A implies B, but A is bid at 0.60 while B is offered at 0.55.
result = check_constraint(
    "implication",
    [Quote("A", yes_bid=0.60, schedule=KalshiSchedule()),
     Quote("B", yes_ask=0.55, schedule=KalshiSchedule())],
)
best = result.best
print(best.description, best.gross_edge, best.fees, best.net_edge)
# NO on A + YES on B 0.05 0.0342 0.0158

Full CLI reference: documentation.

How do AI agents use it?

MCP server

Add it to any MCP client. For Claude Code:

claude mcp add oracle3 -- uvx oracle3 mcp

For Claude Desktop, Cursor and other clients that read an mcpServers block:

{
  "mcpServers": {
    "oracle3": { "command": "uvx", "args": ["oracle3", "mcp"] }
  }
}

Tool

What it does

Side effects

search_markets

Keyword search on Kalshi or Polymarket; Kalshi series listing

read-only

get_market

Prices, volume, close time and resolution rules

read-only

get_orderbook

Both sides of the book, best level first

read-only

get_quote

Best bid and ask on YES and NO, with the market's fee schedule

read-only

check_constraint_live

Fetch quotes and fee schedules, then check a relation

read-only

check_constraint

Check a relation on quotes you supply

none

trading_fee

Fee for one fill under a venue schedule

none

fair_value

Probability implied by a price under the Wang transform

none

list_relation_types

The supported relations and their bounds

none

list_relations

Relations saved locally by the research CLI

reads a local file

paper_order

Buy in a local paper ledger, filling against the live book with fees

writes a local file

paper_portfolio

Cash, positions and fills in the paper ledger

reads a local file

paper_reset

Erase the paper ledger (requires confirm=true)

writes a local file

Real-money execution stays in the CLI: agents research and paper-trade through MCP, and a human signs off on live orders.

If your client ran oracle3 1.2.0, which failed to start with mcp 2.x, refresh uv's cached copy once with uvx --refresh oracle3 mcp.

Agent skills

skills/ (mirrored in .claude/skills/ for Claude Code) holds step-by-step instructions for agents:

Skill

Use it to

pm-constraint-arbitrage

Check related markets for a fee-surviving violation with the MCP tools

pm-data-discovery

Find markets and save research samples

pm-quant-strategy-authoring

Write a tunable QuantStrategy

pm-agent-strategy-authoring

Write an LLM- or tool-driven AgentStrategy

pm-paper-trade-ops

Run, monitor and archive paper trading

pm-live-trade-ops

Live trading, only with explicit user approval

JSON CLI

Every market, paper and trade command, and every research command except research memory, accepts --json. A running engine can be paused, resumed, inspected and stopped from another process with oracle3 trade pause|resume|state|stop --json. See AGENTS.md for which commands are read-only.

What do fees do to the edge?

Both venues charge taker fees proportional to p(1 − p). A two-leg taker basket with both legs near 0.50 has to clear these violations per contract before any edge is left:

Venues

Break-even violation

Kalshi + Kalshi

3.50¢

Kalshi + Polymarket (rate 0.05)

3.00¢

Kalshi + Polymarket (rate 0.04)

2.75¢

Polymarket + Polymarket (rate 0.04)

2.00¢

Buying every outcome of an n-way event costs k(1 − Σp²) per contract, which approaches 7¢ on Kalshi as outcomes multiply. The derivation, the tables and the sources are in Do prediction-market arbitrage edges survive fees?; python scripts/fee_frontier.py reproduces every number.

How is it tested?

  • oracle3.fees reproduces Kalshi's published fee table and Polymarket's documented fee example.

  • oracle3.arbitrage is unit-tested for every relation, including mixed-venue baskets and missing quotes.

  • The MCP server is tested against mocked venue APIs, run against the live public APIs, and checked in CI on both major versions of the MCP SDK.

  • The pricing engine uses the coefficients from the companion working paper (SSRN 6468338), checked against its replication package.

Roadmap

  1. Price every strategy signal with the venue fee schedules in oracle3.fees.

  2. Measure how often and how deeply live violations clear the fee hurdle, per relation and venue pair.

  3. Wire SpreadExecutor, multi-leg execution with LIFO unwind on partial fills, into the multi-leg strategies.

  4. Publish a pre-registered forward track record with timestamped daily snapshots.

How is it built?

graph TD
    R[Relation store<br/>implication · exclusivity · complement · same event · event sum] --> C[Constraint checker<br/>oracle3.arbitrage + oracle3.fees]
    Q[Venue data<br/>Kalshi · Polymarket public APIs] --> C
    C --> S[Strategy layer<br/>6 constraint-based · 2 statistical · 2 model-driven · LLM agents]
    P[Pricing engine<br/>Wang transform, calibrated in Yang 2026] --> S
    S --> E[Trading engine<br/>risk manager · position tracker · kill switch]
    E --> T[Paper trader]
    E --> L[Live traders<br/>CLI only]
    C --> M[MCP server<br/>read-only tools + paper ledger]
    Q --> M

Relations and venue quotes feed the constraint checker, which prices every basket under each market's fee schedule. Strategies consume those checks and the pricing engine's fair values and send orders through a trading engine that enforces risk limits. The MCP server exposes the data, the checker and a separate paper ledger to agents; live traders are reachable only from the CLI.

Constraint-based strategies, each enforcing one probability bound:

Strategy

Bound

Cross-market

Same event, same price across venues

Exclusivity

P(A) + P(B) ≤ 1 for mutually exclusive events

Implication

P(A) ≤ P(B) when A implies B

Conditional

P(A | B) within derived bounds

Event sum

Σ P(outcome) = 1 within an event

Structural

P(A) = β·P(B) + α from a fitted relation

Statistical arbitrage: cointegration spread, lead-lag. Model-driven: fair-value divergence and premium decay, using the pricing model below.

Pricing model. Fair values come from the Wang transform p_mkt = Φ(Φ⁻¹(p) + λ), with λ estimated on 291,309 resolved contracts in the companion working paper, Yang (2026), Pricing Prediction Markets: Incomplete Markets, Selection Rules, and Calibration Wedges (SSRN 6468338). The model and its estimates are documented there.

How can I collaborate?

How do I cite it?

Citation metadata is in CITATION.cff, and every release is archived on Zenodo (DOI 10.5281/zenodo.20062548). For the pricing model, cite the working paper (SSRN 6468338).

Origin and attribution

Oracle3 began as ulab-uiuc/oracle3, developed by Yicheng Yang and Haofei Yu at U Lab (University of Illinois Urbana-Champaign) under the MIT License, and it bundles the coinjure package from the same lab. The strategy, pricing, risk, dashboard, and test layers in this repository were added on top of that base; see NOTICE for the retained license text.

License

Apache 2.0; see LICENSE. Portions from the original U Lab code remain under the MIT License reproduced in NOTICE.

This software is for research and education. Trading involves financial risk.

Available Tools

13 tools
check_constraintA
Read-onlyIdempotent

Check a no-arbitrage relation on quotes you supply (offline).

Quote order: implication (A, B) means A implies B; complement and same_event take (A, B);
exclusivity and event_sum take every outcome.
ParametersJSON Schema
NameRequiredDescriptionDefault
makerNo
quotesYes
relationYes
contractsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly, idempotent, and non-open-world behavior, and the description adds meaningfully beyond them by clarifying that checking is performed offline on supplied quotes and by specifying the arity/semantics of each relation type. It does not describe the result payload, but with annotations carrying the safety profile and an output schema present, this is solid added context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with no filler. The core purpose is front-loaded and the per-relation quote-ordering rule, the highest-risk ambiguity, is spelled out compactly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given five relation types with differing arities and a nested QuoteInput, the description handles the central ambiguity (which relations take two quotes vs. all outcomes) well. However it omits what 'offline' means operationally (no network fetch, fees must be supplied), and how maker and contracts affect the check, leaving gaps an agent cannot resolve from the schema alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage on top-level parameters is 0%, so the schema is silent on relation, quotes, maker, and contracts. The description compensates partially by explaining quote ordering per relation ('implication (A, B) means A implies B; complement and same_event take (A, B); exclusivity and event_sum take every outcome'), but leaves maker, contracts, and the fee/venue fields of QuoteInput undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and object ('Check a no-arbitrage relation on quotes you supply') and adds the '(offline)' qualifier, which implicitly separates it from the sibling check_constraint_live. It does not name that sibling explicitly, so differentiation is inferred rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The '(offline)' marker implies this is the variant to use when quotes are supplied by hand rather than fetched live, but the description never names check_constraint_live or states the condition that selects one over the other. Usage is implied, not prescribed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_constraint_liveC
Read-only

Fetch current quotes and fee schedules for the markets, then check the relation.

ParametersJSON Schema
NameRequiredDescriptionDefault
makerNo
marketsYes
relationYes
contractsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the agent knows this is a non-mutating call that reaches outside the system. The description adds real value by disclosing that it performs live network fetches of quotes and fee schedules before evaluating, implying latency and market-dependent results. It does not, however, mention rate limits, freshness/staleness of quotes, or any failure behavior when a market is unavailable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that follows the tool's actual execution order (fetch, then check). No filler, though it is arguably under-specified rather than concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, but with 0% parameter coverage the description does not compensate for the undocumented relation/maker/contracts semantics, and it omits the live-vs-check_constraint distinction that is the tool's entire reason for existing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across four parameters. The description only loosely gestures at 'markets' and 'the relation'; it says nothing about the relation enum values (implication, exclusivity, complement, same_event, event_sum), what 'maker' toggles, or what 'contracts' controls. For a tool whose core input is an opaque relation type, this is a meaningful gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a compound verb sequence — fetch live quotes and fee schedules, then check the relation — which is more than a tautology but still vague about what 'check the relation' actually produces. Critically, it never distinguishes this tool from the obvious sibling check_constraint (the likely non-live variant), so an agent cannot route between them from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no exclusions, and no mention of the sibling check_constraint, which appears to be the offline counterpart. The only implicit signal is the word 'live' in the name and the fetching language in the description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fair_valueA
Read-onlyIdempotent

Probability implied by a market price under the Wang transform p_mkt = Phi(Phi^-1(p) + lam).

The default lam = 0.183 is the pooled estimate in Yang (2026), SSRN 6468338; it pools
real-money and play-money venues, so treat it as illustrative.
ParametersJSON Schema
NameRequiredDescriptionDefault
lamNo
market_priceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint=false, so the safety profile is covered. The description adds genuinely new behavioral context: the exact transform applied, the model's provenance (Yang 2026, SSRN 6468338), and a limitation on the default coefficient (pooled across real-money and play-money venues, treat as illustrative). That caveat is the kind of domain nuance annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the formula so the core semantics are immediate. The citation and pooled-estimate caveat take space but earn it by qualifying the default argument. Nothing redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the formula plus the lam caveat cover the mechanics. What remains missing is usage context – when an agent should reach for this over the sibling market/quote tools – which for a small calculation utility is a noticeable but not fatal gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the parameter burden. It explains lam (default 0.183, its source, and the pooled-estimate caveat) and makes market_price's role clear via the formula p_mkt = Phi(Phi^-1(p) + lam). This largely compensates for the empty schema, though it doesn't spell out units or valid ranges.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific computational verb (probability implied by a market price) and names the exact model (Wang transform) with a formula, so the agent knows this is a math utility rather than a market lookup. It does not explicitly contrast itself with siblings like get_quote or search_markets, which keeps it just shy of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to call this versus alternatives (e.g., get_quote for raw prices vs this for a risk-adjusted probability). The only contextual note – that the default lam is illustrative because it pools venue types – is a caveat on the default value, not usage guidance. An agent must infer the use case from the formula alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_marketB
Read-only

Market details: title, best prices, volume, close time and (Polymarket) fee schedule and outcomes.

ParametersJSON Schema
NameRequiredDescriptionDefault
venueYes
market_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so safety is covered. The description adds mild behavioral context by indicating that fee schedule and outcomes are venue-dependent, specifically noting Polymarket, but does not discuss authentication, rate limits, or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence that lists the key returned fields without waste. It is appropriately sized for a simple retrieval tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema exists, so return values need not be exhaustively explained, and annotations cover the read-only/open-world profile. However, the description leaves input parameter semantics and sibling-tool routing unclear, making it minimally adequate rather than fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description does not explain the required venue enum or the market_id format. It lists returned fields rather than input parameters, so it fails to compensate for the missing parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource as market details and lists the specific fields returned, including best prices, volume, close time, fee schedule, and outcomes. It does not explicitly differentiate itself from sibling tools like search_markets, get_orderbook, or get_quote, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no when-to-use guidance, prerequisites, or alternatives to sibling tools such as search_markets or get_quote. Usage is only implied by the noun phrase 'Market details' and the required venue/market_id parameters.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_orderbookB
Read-only

Order book for both sides as [price, size] levels, best first.

ParametersJSON Schema
NameRequiredDescriptionDefault
depthNo
venueYes
market_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds the useful behavioral detail that levels are returned best-first and cover both sides, but it says nothing about depth truncation, pagination, or how the two venues differ.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence that front-loads the resource and the return shape, with no filler. Everything stated earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although an output schema exists so return values need not be re-explained, the description leaves all three parameters undocumented at 0% schema coverage. For a per-venue market-data tool, the meaning of depth and venue-specific behavior are gaps an agent would need to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden and fails: it never mentions venue, market_id, or the depth parameter's default of 10 and what it controls. Only the venue enum in the schema gives any parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific resource (the order book), its shape ([price, size] levels for both sides), and ordering (best first), so an agent knows exactly what data comes back. It does not differentiate itself from the sibling get_quote, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus get_quote or get_market, and no mention of prerequisites or context. The agent must infer that this is the depth-of-book tool while get_quote is the top-of-book tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_quoteB
Read-only

Best bid and ask on YES and NO plus the fee schedule the venue reports for this market.

ParametersJSON Schema
NameRequiredDescriptionDefault
venueYes
market_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so safety profile is covered. The description adds useful content disclosure (it returns fee schedule data alongside prices) but says nothing about freshness, rate limits, or error conditions for unavailable markets.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, and the substantive content (prices + fees) is stated immediately. It is efficient, though it omits structure such as a clear action verb that would make it even tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is not required here, and this is a simple read-only two-parameter tool. The description adequately conveys what is fetched; the only real gap is parameter semantics, which is minor for such a small signature.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two required parameters, so the description carries the burden. It hints at market scoping with 'this market' and venue-specific reporting with 'the venue reports', but never explains the market_id format, ID conventions per venue, or what happens with an unknown venue/market pair.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The sentence names a specific resource (best bid/ask on YES and NO plus the venue fee schedule for a market), which is concrete enough to distinguish it from a plain market lookup. However, it lacks an explicit action verb and never contrasts itself with the closest sibling, get_orderbook, leaving the agent to infer the difference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or when-not-to-use guidance is given, and no alternative is named. The agent must guess whether this replaces get_orderbook or complements it, and nothing says when a quote is preferable to fair_value.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_relationsC
Read-only

Relations saved locally by the oracle3 research CLI (~/.oracle3/relations.json).

ParametersJSON Schema
NameRequiredDescriptionDefault
statusNo
market_idNo
spread_typeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint and openWorldHint annotations already cover the safety and world-access profile. The description adds useful context by identifying the local persistence file (~/.oracle3/relations.json), but does not describe filtering behavior, response shape, or other operational traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short and has no filler, but it is an incomplete sentence fragment rather than a front-loaded statement of purpose. Brevity is achieved at the cost of clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even accounting for the existing output schema and annotations, the description omits parameter semantics and usage guidance. An agent cannot tell how to use the optional filters or why it should choose this tool over list_relation_types.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description provides no information about the three parameters (status, market_id, spread_type). With low coverage, the description must compensate, but it adds no meaning beyond the parameter names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a noun phrase that identifies the resource and where it is stored, but it never states the action (e.g., 'lists relations'). The tool name carries the verb, so the description itself adds little clarity about what the tool actually does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives such as list_relation_types. The only contextual clue is that relations are saved locally by the oracle3 CLI, which is not enough to route an agent reliably.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_relation_typesB
Read-onlyIdempotent

Supported relations and the probability bound each one enforces.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=false, covering the safe read-only profile. The description adds domain context about probability bounds but no extra behavioral details such as pagination, auth, or side effects; with annotations doing the heavy lifting, this is minimally adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single front-loaded sentence with no filler, which is concise. It reads as a sentence fragment without a verb, so it is slightly less structured than an ideal action description, but it wastes no space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema exists, so the description need not explain return values, and the read-only annotations cover safety. However, because a sibling list_relations exists, the description leaves ambiguity about scope and when to choose this tool; that is a meaningful gap for a listing tool even if calling it is straightforward.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics to document; the baseline for a no-param tool is 4. The description does not need to compensate for any schema gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource and the key data it exposes (supported relations and their probability bounds), so an agent can tell what the tool returns. It lacks an explicit verb and does not distinguish the tool from the sibling list_relations, but the content is specific enough to avoid being vague.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, prerequisites, or alternative guidance is provided. The agent must infer that this differs from list_relations and when to pick it, so usage guidance is essentially absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

paper_orderB

Buy in the local paper ledger, filling against the live displayed book with venue fees. Never trades for real.

ParametersJSON Schema
NameRequiredDescriptionDefault
sideYes
venueYes
contractsYes
market_idYes
limit_priceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=true. The description adds genuinely non-obvious behavior: fills occur against the live displayed book, venue fees are applied, and no real trade results — details not derivable from the annotations or schema. It omits what happens on insufficient paper balance or whether the order persists across paper_reset.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler, and the ledger/simulation nature is front-loaded. The opening 'Buy in the local paper ledger' is slightly compressed and the punchy 'Never trades for real' sits well as a closing constraint.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and annotations cover the safety profile. Still, for a 5-required-parameter mutation tool with 0% schema coverage, the description leaves the agent without parameter-level guidance or error/state behavior, so it is only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for five required parameters (venue, market_id, side, contracts, limit_price), so the description carries the full burden — yet it says nothing about units, price format, contract counts, or how limit_price interacts with the displayed book beyond a vague 'filling against' phrase. Only a loose hint about limit pricing is present.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action on a specific resource: submitting an order into a local paper (simulated) ledger, filled against the live book with venue fees. It distinguishes itself from the paper_portfolio/paper_reset siblings by being the order-placing tool, though it never explicitly names a sibling or contrasts with real-trading tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: 'Never trades for real' signals this is only for simulated trading, which is a useful scoping cue, but the description never says when to choose this over alternatives or what preconditions (e.g. a funded paper account) apply.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

paper_portfolioB
Read-only

Cash, positions and fill count in the local paper ledger.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds only that the ledger is local/paper, which is useful domain context but does not disclose additional behavioral traits such as auth needs, rate limits, or freshness of the data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence with no filler and the key output fields are front-loaded. It is appropriately sized for a zero-parameter read tool, though the fragment style lacks an explicit verb.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple zero-parameter read tool with an output schema and annotations already covering read-only and local scope, the description is largely sufficient. It could still be improved by stating the retrieval action and when to use it, but no critical detail is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the description is not required to document any inputs. Baseline 4 applies per the rubric for zero-parameter tools.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource and its contents (cash, positions, fill count in the local paper ledger), making clear what information the tool exposes. However, it uses a noun phrase rather than a verb+resource construction and does not differentiate from siblings such as paper_order or paper_reset.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. It does not mention paper_order, paper_reset, or any condition for selecting this portfolio view, leaving usage entirely to inference from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

paper_resetA
Destructive

Erase the paper ledger and restore starting cash. Requires confirm=true.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is covered structurally. The description adds value beyond that by naming exactly what is destroyed ('the paper ledger') and what is restored ('starting cash'), plus the confirm gate — meaningful context for an irreversible operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, effect front-loaded and the precondition second. Nothing extraneous.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and annotations cover the destructive profile. What is erased, what is restored, and the confirm requirement are all stated. Minor gap: no note on irreversibility or partial-reset behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter has 0% schema description coverage and a bare boolean default. The description compensates by explaining that confirm=true is required, which the schema does not convey. It could go further on what happens with confirm=false, hence not a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource: 'Erase the paper ledger and restore starting cash.' That clearly distinguishes it from sibling tools like paper_order and paper_portfolio, though it never names them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the precondition 'Requires confirm=true', which is real usage guidance, but says nothing about when a caller should reset versus inspect or rebuild state, and no alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_marketsB
Read-only

Find open markets by keyword. On Kalshi, series_ticker (e.g. KXFEDDECISION) lists one series.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryNo
venueYes
series_tickerNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint, so the safety profile is covered. The description adds a useful behavioral fact ('open' markets only, and Kalshi series semantics) but says nothing about pagination, the limit default, or whether results are ranked — modest added value over the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the core purpose front-loaded and a venue-specific caveat second. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. But with 0% parameter coverage in the schema, the description should compensate for limit/venue and does not, leaving meaningful gaps for a 4-parameter search tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden for four parameters. It clarifies series_ticker for Kalshi and that query is a keyword, but leaves limit and the venue enum values (kalshi/polymarket) unexplained — only half the parameters get any semantic help.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Find open markets by keyword'), which cleanly separates it from single-item siblings like get_market and get_orderbook. It does not explicitly name those siblings, so it falls short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives one concrete usage hint — using series_ticker on Kalshi to list a series — which implies when the tool is useful. However, it never states when to prefer this over get_market or other lookup siblings, nor any exclusions, leaving the agent to infer selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trading_feeC
Read-onlyIdempotent

Fee for one fill under the venue's published schedule.

ParametersJSON Schema
NameRequiredDescriptionDefault
makerNo
priceYes
venueYes
contractsYes
polymarket_rateNo
kalshi_maker_feesNo
kalshi_multiplierNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and openWorldHint=false, so the safety profile is covered. The description adds a modest behavioral clue — fees come from the venue's published schedule rather than being negotiated — but says nothing about how the override parameters (polymarket_rate, kalshi_multiplier, kalshi_maker_fees) change the result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, so nothing is wasted. It is arguably under-sized for a 7-parameter tool, but that is a completeness problem rather than a conciseness one.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, but the tool has 7 parameters, venue-specific override semantics, and zero schema documentation. For a fee-calculation tool whose correctness depends on which knobs are set, the description is far too thin to let an agent invoke it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 7 parameters, so the description must carry the load and it does not: maker, polimarket_rate, kalshi_maker_fees and kalshi_multiplier are all unexplained. Only 'contracts'/'price' are obliquely implied by 'one fill', leaving the venue-specific knobs opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Fee for one fill under the venue's published schedule' names a specific computational resource and constrains its scope to a single fill evaluated against a venue schedule. It is clear what the tool returns, but it never distinguishes itself from siblings like fair_value or the paper_* tools that also produce per-unit economics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of alternatives such as fair_value when the agent needs valuation rather than fees. The agent is left to infer the calling context from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 13 tool updatesv1.2.1
    • First observedcheck_constraint
    • First observedcheck_constraint_live
    • First observedfair_value
    • First observedget_market
    • First observedget_orderbook
    • First observedget_quote
    • First observedlist_relation_types
    • First observedlist_relations
    • First observedpaper_order
    • First observedpaper_portfolio
    • First observedpaper_reset
    • First observedsearch_markets
    • First observedtrading_fee

TDQS

B3.2/5.0

Scored across 13 tools

Disambiguation4/5

Most tools target clearly distinct operations (search vs. detail vs. orderbook vs. quote), and check_constraint/check_constraint_live are cleanly split by the _live suffix. The main confusion risk is list_relations vs. list_relation_types, which sound nearly identical despite one returning static supported types and the other saved local relations; trading_fee also overlaps slightly with get_quote's embedded fee schedule.

Naming Consistency4/5

A clear verb_noun convention dominates (search_markets, get_market, list_relations, check_constraint, paper_order). Minor deviations are trading_fee and fair_value, which are noun phrases rather than verb-led, but the overall pattern is predictable and readable.

Tool Count5/5

13 tools is well within the healthy range and each earns its place across discovery, market data, constraint analysis, valuation, and paper trading. No redundant or filler tools bloat the set.

Completeness4/5

The surface covers the full research lifecycle: finding markets, fetching quotes/orderbooks, checking arbitrage relations offline and live, fair-value pricing, and simulated execution with portfolio management. Real order placement and market history/resolution data are absent, but paper-only trading appears intentional, leaving only minor gaps.

Maintenance

ActivityMaintained
ResponsivenessSlow

Related MCP Connectors

Related MCP Servers

  • A
    license
    C
    quality
    C
    maintenance
    MCP server that provides a unified prediction market API for multiple venues like Polymarket and Kalshi, allowing AI agents to discover markets, fetch order books, and execute trades through a single interface.
    32
    262 npm
    8
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    A read-only MCP server exposing Polymarket's public prediction-market data. Search markets, read live odds and order books, pull historical probability time-series, and inspect public wallet positions.
    14
    24 PyPI
    MIT