Skip to main content
Glama

ZeroWidth

Read the flow trace behind one scored item

caliper_evals_runs_item_execution
Read-only

What the flow actually DID on one run item: every step in order (what it said, which tools it called with what arguments, what came back) and the final answer, plus status, duration, and cost. Read this when a low score needs explaining beyond the judge's reasoning — a wrong tool call or an empty tool result is usually the cause, and the fix is different from a prompt fix. Pass the run item's id (NOT itemId) from caliper_evals_runs_get. Items whose outputs were supplied from outside (external evals) have no trace and return not_found. Steps are capped for transport; the run page in Caliper has the full record.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
runIdYesRun id.
evalIdYesEval the run belongs to.
runItemIdYesThe run item's `id` from caliper_evals_runs_get (not its dataset itemId).
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a non-destructive read (readOnlyHint=true, destructiveHint=false, openWorldHint=false), and the description adds real behavioral context beyond that: not_found for external-eval items, transport-level step capping, and a pointer to the full record on the Caliper run page. It does not discuss auth or pagination semantics for the capped step list, so it falls short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, all load-bearing: what it returns, when to reach for it, and the two failure/limitation caveats. It is front-loaded with the payload description. Minor redundancy between the schema's runItemId note and the description's id-vs-itemId warning costs it a point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the burden of describing returns — and it does, enumerating ordered steps (utterances, tool calls with arguments, tool results), final answer, status, duration and cost. Combined with the not_found and truncation caveats, an agent has everything needed to call it and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters, including the note that runItemId is the item's `id` rather than its dataset itemId. The description's "pass the run item's `id` (NOT `itemId`) from caliper_evals_runs_get" largely restates that schema text, adding only reinforcement and the sibling that produces the value. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it returns what the flow actually DID on one run item — ordered steps, tool calls with arguments, tool results, final answer, plus status, duration and cost. This is clearly distinguishable from caliper_evals_runs_get (which yields scores/metadata) and caliper_traces_get, so an agent can pick it without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit trigger ("Read this when a low score needs explaining beyond the judge's reasoning") and explains the diagnostic value (a wrong tool call or empty result is usually the cause, and the fix differs from a prompt fix). It also names a when-not case: externally supplied outputs have no trace and return not_found.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources