Skip to main content
Glama

Read or finalize an agent test timeline

read_agent_test_report

Read calls, timestamps, operations and resulting state. Optional finalization freezes the run. A linked repeat includes the previous frozen report while retained. No score, certification or judgment of the agent.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
run_idYesRun identifier from start_agent_test.
finalizeNoTrue freezes the run and prevents further actions. Omitted means read only.
access_tokenYesSecret short-lived run token. Omitted from reports. MCP client history may retain this argument.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • addedInput schema / required
      Added value: +[
      +  "run_id",
      +  "access_token"
      +]
  2. First observed

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes beyond the annotations by disclosing what the report contains, that optional finalization 'freezes the run,' that a linked repeat includes the previous frozen report while retained, and that the tool produces no judgment of the agent. This meaningfully explains side effects and non-behaviors beyond the readOnlyHint=false and destructiveHint=false annotations. There is no contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded, opening with the primary read action and following with finalization, linked-repeat behavior, and a clarifying note about no scoring. Each sentence earns its place, though 'while retained' is slightly vague.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with three simple parameters, no nested objects, and no output schema, the description covers the main content of the report, the finalization side effect, and the absence of evaluative judgment. It does not describe exact response formatting or error cases, but the absence is not critical given the low complexity and strong schema coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and each parameter is already well documented: run_id is the identifier from start_agent_test, finalize controls freezing, and access_token is a secret short-lived token omitted from reports. The description adds little beyond what the schema provides, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: 'Read calls, timestamps, operations and resulting state' and 'finalization freezes the run.' It gives enough specificity about what the report contains and the optional finalize behavior. It does not explicitly name a sibling alternative like inspect_test_state, but the phrasing is distinct enough for an agent to identify this as the report-reading/finalizing tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool: when you need to read or finalize an agent test timeline, and the 'No score, certification or judgment' sentence gives a meaningful exclusion. However, it does not explicitly state when to prefer this over inspect_test_state or other siblings, nor does it describe conditions such as run completion or token validity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources