Skip to main content
Glama

replay_session

Replays a recorded Session against a pipeline version, including a draft, into a new Session to test pipeline changes on real sessions without touching the source or deployed pipeline.

Instructions

Replays a recorded Session against a pipeline version, a draft included, into a new Session.

Use it to check a pipeline change on real Sessions: replay the Sessions behind the change against the draft version, then try_evaluator the relevant Evaluators on each source Session and on its replayed_session_id, and compare. Each replay runs the pipeline on the customer's models, so say what it costs and get a yes before replaying. The source Session and the deployed pipeline are untouched.

FIRST_USER_MESSAGE replays the first turn only; ALL_USER_MESSAGES replays every turn of a chat in order. Leave it unset to use the server's default (the first turn), and say so when reporting. If status is still CREATED or STARTED the wait ran out: poll get_session_replay. :param pipeline_name: Name of the pipeline. :param session_id: The Session to replay: a search_session_id, or a single-turn run's query_id. :param pipeline_version_id: The pipeline version to replay against, e.g. the current draft's. :param replay_mode: FIRST_USER_MESSAGE or ALL_USER_MESSAGES; unset uses the server's default. :param wait_seconds: How long to wait for the replay before handing back the id to poll. :returns: The replay report with the new Session's id, or an error message.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
session_idYes
replay_modeNo
wait_secondsNo
pipeline_nameYes
pipeline_version_idYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.1.29

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does so: it discloses that the source Session and deployed pipeline are untouched, that each replay runs on the customer's models and costs money requiring a yes, and the exact semantics of FIRST_USER_MESSAGE vs ALL_USER_MESSAGES with the unset default. This is unusually complete behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose and the cost/workflow guidance before the enumerated parameter block, and every sentence carries signal. It is on the long side and the param block partially restates schema fields, but nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a mutation-ish 5-param tool with no annotations and no output schema, the description covers cost, consent, scope, defaults, polling fallback, and return value, leaving nothing an agent needs to call it correctly unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and it does, documenting all five parameters including session_id accepting a search_session_id or query_id, pipeline_version_id as e.g. the current draft, replay_mode's enum semantics and default, and wait_seconds' waiting behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (replays) and resource (a recorded Session against a pipeline version, draft included, into a new Session). An agent can immediately distinguish it from siblings like get_session_replay or run_component.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use (check a pipeline change on real Sessions), a concrete workflow (replay the Sessions behind the change against the draft, then try_evaluator and compare), an alternative for polling (get_session_replay) if status is CREATED/STARTED, and a consent prerequisite before spending customer models.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools