Skip to main content
Glama

JYOTINT Sealed Forecasts

get_luck_test

Read-onlyIdempotent

The corpus-level 'could this record be luck?' significance test, computed AGAINST the record: EVERY graded call clustered into independent events (correlated calls share one event; live counts ship in the response), strict scoring (one NEAR fails the whole event), luck-prior floored at a coin flip per event. Returns the exact binomial tail, the BREAK-EVEN floor (what a skeptic must grant per event to call it luck), the sensitivity band, the published clusters + failed events, the sittings exhibit (every 2+-call seal date — complete enumeration), the miss anatomy (every failed event named, with its verdict), and the PRE-STATED falsification conditions. Caveats ship in the same object — quote them with the numbers. Measures improbability-of-luck, never calibration skill (the aggregate Brier's base-rate tie stays disclosed).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), and the description adds substantial context beyond that: clustering rules, strict one-NEAR-fails scoring, the luck-prior floor, and the fact that caveats ship in the same response object and must be quoted with the numbers. It omits any note on cost, latency, or invocation constraints, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded, which is good, but the body is a single sprawling run-on paragraph packing in break-even floors, sensitivity bands, sittings exhibits, and miss anatomy with heavy capitalization and coined jargon. Every clause does carry content, so it is not filler, but the structure makes it hard to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With zero parameters and no output schema, the description carries the full burden of explaining returns, and it does so thoroughly — binomial tail, break-even floor, sensitivity band, clusters, failed events, sittings exhibit, falsification conditions, and shipped caveats. It is nearly complete for a no-arg analytical tool, missing only invocation-level details like auth or rate limits.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics to document and the baseline of 4 applies. The description correctly avoids inventing parameters and instead spends its budget on output semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb and resource — the corpus-level 'could this record be luck?' significance test computed against the record — and explicitly distinguishes it from calibration measurement, which routes agents away from the get_calibration_and_integrity sibling. However, the dense jargon-heavy framing buries the plain 'what it does' statement, so it is clear but not instantly scannable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The closing clause 'Measures improbability-of-luck, never calibration skill' implicitly tells the agent to use get_calibration_and_integrity for calibration, which is a useful routing hint. But there is no explicit when-to-use guidance, no prerequisites, and no statement of when this tool is the wrong choice beyond that one negative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources