Skip to main content
Glama

run_status

Check the progress of a benchmark run or list recent runs by omitting run_id, with an optional limit for how many to show.

Instructions

Progress of one run, or a list of recent runs if run_id is omitted.

Args: run_id: the run to check. limit: how many recent runs to list when run_id is omitted.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
limitNo
run_idNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose a genuine behavioral nuance — omitting run_id switches from a single-run check to a recent-runs listing — which is useful. However, it says nothing about read-only nature, permissions, pagination, or truncation limits, leaving meaningful behavioral gaps for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, front-loaded with the primary behavior before the argument details. The Args block is slightly mechanical and restates run_id ('the run to check'), but nothing is padded or wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A read-only status tool with an output schema, so return-value documentation is not required. Both parameters and both operating modes are explained. Minor gaps remain around error behavior for an invalid run_id and whether the recent-runs list is bounded, but the definition is sufficient to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it largely does: run_id is 'the run to check' and limit is scoped to only apply 'when run_id is omitted.' That conditional coupling of limit to run_id is not derivable from the schema, which only shows types and defaults. It omits the default value (15) and acceptable ranges, keeping it from a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource and operation: 'Progress of one run, or a list of recent runs.' The dual behavior is spelled out up front, so an agent can tell it apart from siblings like start_run, resume_run, get_report, and compare_runs, which all do something different. It does not explicitly name a sibling, so it falls short of the top tier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The condition selecting each mode is explicit: provide run_id to check one run, omit it to list recent runs. This is real routing guidance rather than implied usage. There are no exclusions or prerequisites (e.g., auth, valid state), so it stops at 4.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.