Skip to main content
Glama

ZeroWidth

List an eval's runs (status + scores)

caliper_evals_runs_list
Read-only

Recent runs for one eval, newest first: status (PENDING/RUNNING/SCORING/DONE/FAILED), overall score once DONE, label, and timestamps. THE CHECK-BACK for caliper_evals_run: when the user asks how the run went, read this — cite the run id and score, and compare against the PREVIOUS run's score for the delta. Still don't poll in a loop; check when the user asks or when reporting.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
evalIdYesEval whose runs to list.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover the safe-read profile (readOnlyHint, destructiveHint=false). The description goes well beyond them: newest-first ordering, the PENDING/RUNNING/SCORING/DONE/FAILED lifecycle, that a score exists only once DONE, and an anti-polling behavioral constraint. That is substantial behavioral context layered on top of the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the resource and return fields, then the routing guidance. Dense but every clause carries information; the em-dash-heavy final sentence runs slightly long but nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description correctly enumerates the returned fields (status, score once DONE, label, timestamps, ordering) and pairs that with the delta-comparison workflow. Nothing an agent needs to call and interpret this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters, so evalId and workspace are already fully documented. The description adds only the single-eval scoping ('for one eval') and does not elaborate on the workspace/permission nuances beyond what the schema states. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Recent runs for one eval') plus scope ('newest first') and the exact fields returned (status, score, label, timestamps). An agent can immediately distinguish this from caliper_evals_get and caliper_evals_runs_get without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly frames itself as 'THE CHECK-BACK for caliper_evals_run', names the triggering condition ('when the user asks how the run went'), and adds a when-not rule ('don't poll in a loop; check when the user asks or when reporting'). This is exactly the when/when-not/related-tool guidance the dimension asks for.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources