Skip to main content
Glama
CoinRithm

CoinRithm/coinrithm-agent-trading

Official

Per-venue market-price calibration

pm_data_calibration
Read-only

Score and compare prediction-market venue price calibration: match each outcome's price about 24h before resolution to the realized result, then view calibration error and sample sizes.

Instructions

Free public per-venue market-price calibration scorecard. The primary scored lane compares the venue price for each outcome at one complete-book snapshot selected nearest 24h before resolution within the inclusive 20-28h window against the realised result. calibrationError is event-weighted Expected Calibration Error (0-1, lower is better) within comparable samples; sampleSize counts scored events. This measures market-price calibration, not provider or agent forecast skill, profitability, or a continuous 24h history. Venues below minSample (currently 30 scored events) appear in pending. The additive finalPrice and ownCapture lanes use different timing bases and are not interchangeable with the primary scored lane. Cite CoinRithm's methodology and excluded counts when comparing venues. No API key required.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
okYes
bodyNo
httpStatusYes
ledgerStatusNo
ledgerEventIdNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed5 schema fields changedv0.1.23
    • removedOutput schema / properties / body / description
      Removed value: -"Parsed CoinRithm response body, or raw text when the response is not JSON."
    • removedOutput schema / properties / httpStatus / description
      Removed value: -"HTTP status returned by CoinRithm, or 0 for network errors."
    • removedOutput schema / properties / ledgerEventId / description
      Removed value: -"Private AgentActionEvent id returned by /api/agent/*, when present."
    • removedOutput schema / properties / ledgerStatus / description
      Removed value: -"Ledger write status header returned by CoinRithm, when present."
    • removedOutput schema / properties / ok / description
      Removed value: -"True when CoinRithm returned a successful 2xx response."
  2. Addedv0.1.14

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover safety (readOnlyHint=true, destructiveHint=false), and the description adds meaningful disclosure beyond them: it is free/public, requires no API key, uses a single snapshot window (20-28h, nearest 24h), enforces a minSample threshold, and warns that lanes are not interchangeable. It does not discuss pagination or rate limits, but the added operational context is substantial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the first sentence, and subsequent sentences are dense but each adds a caveat (sample threshold, lane non-interchangeability, citation requirement). It is on the long side and could be tightened, but every clause carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description needn't detail return values, yet it helpfully defines the key fields (calibrationError scale/direction, sampleSize) and the pending/threshold behavior. Combined with full annotation coverage, this is complete enough for an agent to call and interpret the tool, with minor room to state the underlying venue list or exclusions more concretely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. The schema is empty at 100% coverage and the description correctly provides no parameter guidance, instead clarifying output semantics (calibrationError definition, sampleSize meaning).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('per-venue market-price calibration scorecard') and explicitly scopes what it measures versus what it does not (provider/agent forecast skill, profitability, continuous 24h history). That scoping helps separate it from other pm_data_* tools, but no sibling is named directly, so differentiation is implied rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives real usage context: venues below minSample appear in `pending`, the finalPrice/ownCapture lanes are not interchangeable with the primary lane, and results should be cited with methodology and excluded counts. However, it never states when to reach for this tool versus alternatives like pm_data_overview or pm_data_canonical, so usage is inferred rather than routed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.