Skip to main content
Glama

List evaluation results

list_organization_evaluation_results
Read-onlyIdempotent

List paginated results for one evaluation across every agent it grades, newest first. Filter by grade, status, agent, session, or date to track a run to completion.

Instructions

Cursor-paginated results for one evaluation across every agent it grades, newest first. Each session appears once with its latest result; queued and in-progress results are included so a run can be followed to completion. Read operation.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
gradeNoFilter by grade.
cursorNoPagination cursor from a previous response's `next_cursor`.
statusNoFilter by status.
accountNoNamed private Gumloop account; selects private credentials and user/team identity.
agent_idNoOnly results for this agent.
page_sizeNoItems per page (1-100).
session_idNoOnly results for this session.
created_afterNoOnly results created at or after this time. RFC 3339 with an explicit offset (for example `2026-09-01T00:00:00Z`).
evaluation_idYesID of the organization evaluation.
created_beforeNoOnly results created before this time. RFC 3339 with an explicit offset.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv2.0.1

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations carry the safety profile (readOnlyHint, idempotentHint, openWorldHint, destructiveHint=false), so the bar is lower. The description still adds real behavioral value beyond them: newest-first ordering, session deduplication, and the inclusion of queued/in-progress results that make it suitable for polling a run.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with zero padding, front-loading the cursor-paginated scope and ordering before the dedup rule. 'Read operation.' is a useful one-word safety restatement at the end.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Complete enough given a 10-param schema with 100% coverage and annotations covering the safety profile. The only gap is the absence of any routing hint to the sibling singleton result tool, which would help in a list of ~90 siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% across all 10 params, including enums for grade and status and filter semantics. The description adds no parameter-level detail beyond what the schema already states, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (list) + resource (evaluation results for one evaluation) with explicit scope: 'across every agent it grades, newest first'. Names the ordering and the deduplication rule ('each session appears once with its latest result'), which distinguishes it from get_organization_evaluation_result (singular) in the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the name and content, and the description notes that queued/in-progress results are included 'so a run can be followed to completion' — a hint about when to use it. But there is no explicit when-to-use vs get_organization_evaluation_result / get_organization_evaluation_metrics, and no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools