Skip to main content
Glama

JYOTINT Sealed Forecasts

get_regrade_kit

Read-onlyIdempotent

The grade-it-yourself kit: inputs to recompute the record's Brier (calibration), named-mechanism specificity, AND Information Yield under YOUR OWN verdicts — plus the one-step stress-test recipes (harsh-verdicts, externally-adjudicated-only, estimative-worst-case, …). Each call carries its verbatim claim/outcome, the operator's p + verdict to override, and the surprise_bits / 1-in-N inputs. A base rate scores 0 on specificity and 0 bits on IY. Pass an optional id for one call's row; omit for the recipes + usage + count.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
idNoOptional advisory id for one call's regrade row.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), so the bar is lower, and the description adds real operational context: what inputs accompany each call, the base-rate scoring rule (0 specificity, 0 IY bits), and how output shape changes with/without id. What is missing is any note on authorization prerequisites or output size for the recipe set.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The content is compact for how much it covers, but it is one run-on block saturated with em dashes, capitalized jargon, and an ellipsis list that obscures the priority order. The most decision-relevant fact for an agent (id vs no-id behavior) is buried at the very end rather than front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read tool with no output schema, the description does the needed work: it lists the returned artifacts, the stress-test recipes, and the alternate mode when id is omitted. Only minor gaps remain, such as how large the recipe payload is or what the verdict-override input format looks like.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and there is only one optional string param, so the baseline is already 3. The description adds meaning beyond the schema by explaining that omitting id yields the recipes plus usage and count, i.e. it clarifies the semantic consequence of the parameter rather than just restating it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a concrete deliverable (a 'grade-it-yourself kit') and enumerates exactly what it returns: Brier/calibration inputs, specificity, Information Yield under the caller's own verdicts, plus stress-test recipes. It implicitly differentiates from siblings get_calibration_and_integrity and get_information_yield via the 'YOUR OWN verdicts' override framing, though it never names those siblings explicitly. Heavy domain jargon makes the purpose slower to grasp than it needs to be.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear parameter-branching guidance ('Pass an optional id for one call's row; omit for the recipes + usage + count'), which tells the agent how the two invocation modes differ. However, it never states when to reach for this tool instead of get_calibration_and_integrity or get_information_yield, and the 'under YOUR OWN verdicts' distinction is left for the agent to infer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources