Skip to main content
Glama

Run evaluations

run_evaluations
Destructive

Grades up to 200 finished agent sessions using the agent's evaluation setup; results are queued for asynchronous polling until completed or failed.

Instructions

Grades up to 200 of the agent's finished sessions with its own evaluation configuration. Grading is asynchronous: each accepted session gets a result with status: queued; poll it with GET /agents/{agent_id}/evaluations/{evaluation_id} until it is completed or failed. A new result replaces the previous result for that session.

Sessions are skipped, not rejected, when they are unfinished, incognito, or not owned by this agent (ineligible), or already have a queued or running result (in_flight, with the existing result_id). The caller is charged one credit per queued session. Set dry_run: true to see the cost and skips without queuing anything.

Requires edit access on the agent and a plan with evaluations enabled. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
confirmNoSet true only when the user asked for exactly this action.
dry_runNoReport cost and skipped sessions without queuing.
payloadNoComplete JSON request body instead of body flags. Preserves current endpoint fields and values.
agent_idYesID of the agent that owns the sessions.
session_idsNoSessions to grade. Duplicates are rejected.
payload_fileNoRegular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv3.0.0
    • changedInput schema / properties / confirm / description
      Previous value: -"Must be true for this exact requested account change, agent/flow execution, upload or deletion."New value: +"Set true only when the user asked for exactly this action."
  2. First observedv2.0.1

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (destructiveHint/openWorldHint only flag risk). It discloses asynchronous queuing with status: queued, the polling endpoint to reach completed/failed, replacement of prior results, skip-not-reject semantics with ineligible/in_flight reasons and result_id, per-session credit cost, and the confirmation/anti-resubmission rule. This is unusually complete behavioral disclosure for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the verb, resource and cap in the first sentence, then layers async behavior, skip semantics, cost, and prerequisites in short paragraphs. Slightly dense across four paragraphs, and the closing confirmation boilerplate is the least information-dense part, but every block carries operational meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent mutation with no output schema, the description supplies the missing return information (queued status plus polling path), the cost model, permission prerequisites, and partial-success behavior. An agent has everything needed to invoke it correctly and safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so a 3 is the baseline, but the description adds genuine semantics: dry_run reports cost and skips without queuing, session_ids are skipped for unfinished/incognito/not-owned/already-in-flight sessions, and duplicates are rejected. This explains parameter consequences rather than restating the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (grades), the exact resource (the agent's finished sessions using its own evaluation configuration), and the hard cap (up to 200). It is clearly distinguishable from siblings like list_evaluations, retrieve_evaluation, get_evaluation_metrics and update_evaluation_config, which configure or read rather than execute grading.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives strong operating context: prerequisites (edit access on the agent, a plan with evaluations enabled), a recommended dry_run: true preview path, and an explicit confirmation requirement. It does not, however, name a sibling alternative for related tasks (e.g. configuring or reading evaluations), so the routing guidance is contextual rather than comparative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools