Skip to main content
Glama

record_observation

Capture measured metrics and variance for a completed trial, tagged by spatiotemporal region, to document reproducible evidence and reject all-zero variance.

Instructions

Record an observation (data item about a quality).

metrics and variance may be sent as JSON-encoded strings.

Only completed trials can be observed — a failed trial produced no measurement; its failure lives in its status, executor_output, and artifacts, not here. All-zero variance is rejected (n identical outcomes are one effective measurement, not reproducibility). If nothing is concludable, close the programme 'abandoned' rather than fabricating variance.

Enforcement: commitment 7 — reproducibility is the price of admission. Rejects a single point estimate with no variance.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
metricsYesMeasured values {metric_name: float}; may be a JSON-encoded string.
trial_idYesID of the target trial.
varianceYesPer-seed variance {metric: float} — all-zero rejected (commitment 7); may be JSON-encoded.
spatiotemporal_regionYesWhere/when the observation was produced (run footprint tag).

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
statusNo
observation_idNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.28

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it does disclose important behavioral rules: zero variance is rejected, single point estimates without variance are rejected, and it cites an enforcement policy ('commitment 7'). It does not cover permissions/authorization requirements or reversibility, which are meaningful gaps for a mutation tool with no annotation support, so it is a strong 4 rather than a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded, followed by constraints, a fallback action, and the enforcement note. Each sentence carries information, though phrasing like 'All-zero variance is rejected (n identical outcomes are one effective measurement, not reproducibility)' is somewhat dense. Efficient overall with minimal waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because an output schema exists, return values need not be described, and the description appropriately focuses on preconditions and rejection rules. It also contextualizes the failure case against trial status/executor_output/artifacts. The main missing piece is authorization or permission requirements for this write operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description's note that metrics and variance 'may be sent as JSON-encoded strings' merely repeats what the schema already states in both parameter descriptions. It adds no format, range, or semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The definition opens with a specific verb+resource: 'Record an observation', clarified as 'a data item about a quality'. This is clearly distinguishable from siblings like run_trial and get_trial_status by the fact that it records measurements against a completed trial. It stops short of explicitly contrasting itself with the nearest sibling tools, so it earns a 4 rather than a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage conditions are explicit and cover both when and when-not: only completed trials may be observed, failed trials' information 'lives in its status, executor_output, and artifacts, not here', and if nothing is concludable the agent is routed to close_programme 'abandoned'. This is precisely the when/when-not/alternative structure that scores top marks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.