Skip to main content
Glama

ZeroWidth Caliper

Stop a running Caliper eval run

caliper_evals_runs_cancel
Destructive

Stops a run that is still PENDING / RUNNING / SCORING — the brake on a run that's spending more than expected or was started by mistake. The queue stops at once; an item already handed to the flow finishes on its own timeout. The run settles as FAILED with a 'cancelled' reason and keeps the items it completed. A run that already finished returns run_not_live. May return needs_confirmation.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
runIdYesRun id, from caliper_evals_run or caliper_evals_runs_list.
evalIdYesEval the run belongs to.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the destructiveHint=true annotation by disclosing exactly how destruction is scoped: the queue stops immediately, an in-flight item finishes on its own timeout, the run settles as FAILED with a 'cancelled' reason, and completed items are retained. It also surfaces the run_not_live failure mode and a possible needs_confirmation response that requires a follow-up call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the operation and its eligible states, then layers consequence, error, and confirmation behavior in short sentences with no filler. Every sentence carries information an agent needs before calling.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, but the description compensates by covering the terminal state of the run, what is preserved, and both the error and confirmation paths. Combined with destructiveHint annotations and fully documented params, nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so runId, evalId, workspace, and approvalId are all documented in the schema; the description adds only the indirect hint that needs_confirmation implies a second call carrying an approvalId. Baseline 3 is appropriate given the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Stops') and resource ('a run') with an explicit eligible-state set (PENDING / RUNNING / SCORING), which distinguishes it from read-only siblings like caliper_evals_runs_get and caliper_evals_runs_list. An agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete motivating contexts — a run spending more than expected, or started by mistake — and an implicit exclusion via 'A run that already finished returns run_not_live.' It does not name an alternative tool for the already-finished case, but no sibling offers a competing cancellation path, so little is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources