Skip to main content
Glama

record_result

Log a measurable outcome and link it to the agent's confidence snapshot for grading against actual results. Bind it to a prediction ID from a prior check-in.

Instructions

Record a measurable outcome and pair it with this agent's EISV snapshot so verdicts can be graded against what really happened. Pass the prediction_id from a sync_state reply to bind the outcome to that check-in's confidence; it is consumed on first use and TTL-bound (an hour by default). Needs a bound or explicit agent_id, and refuses under strict identity from an ephemeral session. Provenance cannot be self-attested here: verification_source is forced and provenance keys in detail are stripped. Use store_finding for durable knowledge. outcome_event is the canonical twin; this name adds a digest envelope and keeps the raw payload under raw_governance only with response_mode='full' or include_semantics=true, or when the write returned no outcome_id. EISV field definitions: describe_tool(tool_name='check_working_state').

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
detailNoType-specific metadata (e.g., mark_count, test_name, error_message)
is_badNoWhether this is a negative outcome. Inferred from type if omitted.
agent_idNoAgent ID. Falls back to session-bound agent_id if omitted.
confidenceNoAgent confidence at outcome time (0-1). Looked up from last check-in if omitted.
session_idNoOptional session id; falls back to client_session_id and then to context.
outcome_typeYesType of outcome event
outcome_scoreNoQuality score 0.0 (worst) to 1.0 (best). Inferred from type if omitted.
prediction_idNoTactical prediction id from a prior process_agent_update response.
response_modeNo'full' is an alias for include_semantics=true; 'lite' (default) returns the small snapshot.
decision_actionNoThe decision the agent took (e.g. 'proceed', 'pause').
continuity_tokenNoSame-process rebind proof only; never a cross-process resume.
client_session_idNoBinding id for calls in this process; not a cross-process proof.
include_semanticsNoIf true, the response's eisv_snapshot carries the full EISV ontology (state_semantics role table + hierarchy).
verification_sourceNoProvenance of this outcome.agent_reported_tool_result

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed3 schema fields changedv3.1.0
    • changedInput schema / description
      Previous value: -"Parameters for outcome_event"New value: +"Outcome record parameters."
    • changedInput schema / properties / client_session_id / description
      Previous value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof."
    • changedInput schema / properties / continuity_token / description
      Previous value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume."
  2. First observed

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavior beyond the annotations: prediction_id is consumed on first use and TTL-bound (~1 hour), a bound or explicit agent_id is needed, it refuses under strict identity from an ephemeral session, and provenance cannot be self-attested (verification_source forced, provenance keys stripped). This is exactly the write-safety and state-mutation context an agent needs, and it is consistent with idempotentHint=false and destructiveHint=false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose, then layers binding, TTL, identity, and provenance rules in a compact sequence where each clause carries real information. It is dense and occasionally hard to parse (the outcome_event/raw_governance sentence), keeping it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Appropriately complete for a 14-parameter mutation tool with no output schema: it covers binding lifecycle, identity requirements, provenance constraints, response-mode effects, and points to describe_tool for EISV field definitions. Nothing critical to invoking it correctly appears missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is 3; the description earns above baseline by explaining why prediction_id exists (bind to a check-in's confidence), how response_mode/include_semantics gate the digest envelope and raw payload, and the forced verification_source. One minor wrinkle: description says prediction_id comes from a sync_state reply while the schema says process_agent_update response.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Record a measurable outcome') plus the reason it exists ('pair it with this agent's EISV snapshot so verdicts can be graded'). It explicitly differentiates from siblings by naming 'store_finding for durable knowledge' and 'outcome_event is the canonical twin', so an agent can distinguish it without opening another schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear routing context: pass prediction_id from a sync_state reply, use store_finding for durable knowledge, and it clarifies the relationship to the outcome_event twin. It stops short of a crisp 'use this when X, use outcome_event when Y' rule, leaving the digest-envelope vs raw-payload distinction somewhat implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.