Skip to main content
Glama

groundtruth_scorecard

Read-onlyIdempotent

How often calls stamped under this API key were right. Groups them by band and by what the coin actually did, resolved from the same data the card reads. A call whose coin has not resolved counts as PENDING, never as a win -- report pending alongside resolved, because the denominator is the point. Also returns whether the hash chain is intact; if it is not, say so rather than quoting the numbers.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
limitNorows to read, 1-500, default 200

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, idempotent, non-destructive behavior, and the description adds real value beyond that: unresolved calls are classified PENDING and never as wins, the denominator matters, and the tool also reports hash-chain integrity with an explicit instruction not to quote numbers if the chain is broken. That is substantive behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the first sentence, but the prose wanders: 'resolved from the same data the card reads' and 'because the denominator is the point' are flavor that a terse tool description could drop. Four clauses for a one-parameter tool is heavier than necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does the work of explaining the return shape (bands, resolved vs PENDING, hash-chain status) and how to report it. That is close to complete, though the absence of any mention of the limit's effect on completeness of results leaves a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single 'limit' parameter at 100% schema coverage, so the schema already carries the semantics (rows to read, 1-500, default 200). The description adds nothing about pagination or how limit interacts with band grouping, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the resource (accuracy of calls stamped under the API key) and what it returns (groupings by band and by coin outcome), which is specific enough to distinguish it from a raw scoreboard. However, it never names or contrasts itself with the closest sibling, groundtruth_scoreboard, so an agent must infer the difference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives output-handling guidance ('report pending alongside resolved', 'if the chain is not intact, say so rather than quoting the numbers') but no when-to-use condition or alternative routing among the many sibling tools. Usage is implied rather than specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources