Skip to main content
Glama

record_outcome

Attach ground truth to a past evaluate() call to track calibration and improve future decisions.

Instructions

Attach ground truth to a past evaluate() call for calibration tracking. outcome: boolean applied to all answers, or {questionId: boolean}.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
noteNo
outcomeYes
decision_idYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full behavioral burden. It reveals the outcome shape but says nothing about whether the decision_id must already exist, whether an outcome can be overwritten, idempotency, permissions, or what a failed attach produces — significant gaps for a mutation-style call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler, with the core action and the sibling coupling front-loaded before the parameter detail. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 3 undocumented parameters, the description covers the highest-value ambiguity (outcome format) but omits the note parameter, error/return behavior, and prerequisites. Just enough to call it, not enough to call it confidently in edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the 'outcome' property is untyped ({}), so the description's explanation of the two accepted shapes (boolean, or {questionId: boolean}) is genuinely valuable. However, it leaves 'decision_id' (only indirectly implied as coming from evaluate) and the 'note' parameter entirely undocumented, so it only partially compensates.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific action ('attach ground truth') tied to a concrete upstream artifact ('a past evaluate() call'), so an agent knows this is the write-back companion to evaluate. It is clearly distinct from decision_stats, though it never explicitly contrasts itself with either sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for calibration tracking' implies the use case and the reference to a 'past evaluate() call' implies ordering, but there is no explicit when-to-use/when-not guidance or named alternative. Adequate but inference is required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools