Skip to main content
Glama

Get workflow run traces

get_workflow_run_traces
Read-only

Get the full per-step execution trace of a run by run id (paged). Each step lists its blocks with the input each consumed and the output it produced, plus any error_message, and the variable scope captured at that step. Use this to debug why an expression or block produced the wrong value. Large captured values are shortened, with a marker saying so: a scope entry may point at the block output beside it rather than repeat it (the value is in the same trace). On UNATTENDED runs only — a live endpoint serve, a schedule falling due — text copied verbatim from the definition (a literal expression) is also replaced by its path there; read it with get_workflow_version. Those runs are shortened harder overall too, so to capture a value in full, re-run it yourself (run_workflow / preview_dynamic_endpoint / run_schedule_now) — runs you trigger keep everything, literals included.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
run_idYesthe run id returned by run_workflow
page_sizeNomaximum number of step traces to return in this page
page_tokenNotoken from a previous response's next_page_token to fetch the next page

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
stepsYes
next_page_tokenYes

Schema Changelog

Changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. First observed

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, and nothing contradicts them. Beyond that, the description discloses significant non-obvious behavior: pagination, shortening of large captured values with a marker, scope entries that point at adjacent block outputs instead of repeating values, and that UNATTENDED runs replace literal expressions with their path and are shortened harder overall. These are exactly the surprises an agent needs before invoking the tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose in the first sentence, then flows logically through trace structure, shortening behavior, and unattended-run caveats. It is long, but each clause carries distinct information — nothing is filler. It is slightly dense with parenthetical examples, but the complexity of the behavior justifies the length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists and safety annotations are present, the description covers all remaining needs: the content of each step trace, paging behavior, the shortening edge cases, the unattended-run difference, and the concrete workaround for full capture. Nothing an agent needs to correctly invoke or interpret the tool is left unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: run_id, page_size, and page_token each have meaningful schema descriptions, including the next_page_token contract for paging. The description only reinforces this ('by run id', 'paged') without adding format or syntax detail, so the baseline 3 applies since the schema carries the burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb and resource: 'Get the full per-step execution trace of a run by run id (paged)'. This clearly distinguishes it from sibling read tools like get_workflow_run (run-level metadata) and list_runs (run listing), and the rest of the text specifies exactly what a trace contains: blocks, inputs, outputs, error_message, and variable scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use guidance is given: 'Use this to debug why an expression or block produced the wrong value.' It also provides when-not guidance and names concrete alternatives: for literal values on unattended runs use get_workflow_version, and to capture full values re-run via run_workflow / preview_dynamic_endpoint / run_schedule_now. The routing is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

A4/5.0
Disambiguation5/5

Every resource family follows the same verb+noun pattern and each tool name uniquely identifies a resource-action pair (create_app vs create_app_version vs update_app vs publish_app). Closest overlaps like analyze_resource vs get_resource_graph and patch_datafile vs update_datafile are explicitly differentiated by their descriptions, so misselection risk is low despite the scale.

Naming Consistency5/5

Names are almost uniformly verb_noun snake_case with a consistent lifecycle vocabulary: create/get/update/delete/list/publish/unpublish/version. Minor outliers like whoami and run_schedule_now are idiomatic and do not break the predictability of the set.

Tool Count1/5

At 93 tools this far exceeds the calibration's 50+ extreme-mismatch case. The count is inflated by repeating create/get/update/delete/version/publish/unpublish across ten resource families; even though each family is systematic, the combined surface is very hard for an agent to navigate and keep in context.

Completeness4/5

Core CRUD/publish/version lifecycles are present for apps, workflows, endpoints, schedules, schemas, datafiles, and api templates, and dependency analysis is well covered. However, secret creation/updating, asset upload, custom-domain deletion, and version-range enumeration for several resource types are absent or left to the external dashboard, so agents hit a few manual dead ends.

Resources