Skip to main content
Glama

ZeroWidth

Delete a Caliper eval

caliper_evals_delete
Destructive

Deletes an eval and stops its schedule. Run history is kept but no longer reachable from the eval. May return needs_confirmation.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
evalIdYesEval id, from caliper_evals_list.
workspaceNoWorkspace slug. Personal tokens with no default workspace MUST pass this; tokens with a default can override per call. Ignored for workspace API keys.
approvalIdNoApproval id from a prior needs_confirmation response. Omit on the first call.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is known, but the description adds real value beyond them: the schedule is stopped, run history becomes unreachable, and the call may return needs_confirmation. It stops short of stating auth/permission requirements or whether deletion is reversible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each carrying distinct information (what is deleted, what happens to history, what a confirmation response implies), with the primary action front-loaded. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter destructive tool with no output schema, the description covers the salient outcomes an agent needs: the eval is gone, its schedule stops, history is orphaned, and a confirmation round-trip may be required. Only the permission/auth prerequisites are absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description earns above that by explaining the needs_confirmation/approvalId retry cycle, which gives the approvalId parameter meaning the schema alone only partially conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Deletes an eval') and immediately names the collateral effect ('stops its schedule'). It reads clearly against caliper_evals_update/run/get, though it does not explicitly differentiate itself from the closer sibling caliper_evals_runs_cancel.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance, no statement of prerequisites, and no pointer to alternatives (e.g., cancel a run vs. delete the eval). The only usage-adjacent signal is the approval flow implied by the closing sentence.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources