Skip to main content
Glama

task_execution_trace

Get a task's execution timeline to see how it actually went, optionally including checkpoints and derived stats such as duration and step count.

Instructions

Get a task's execution timeline — plain, or with checkpoints + stats.

Answers "how did this task actually go"; include_stats=True adds the derived summary on top of the timeline.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
task_idYesTask ID
include_statsNoFalse (default) — timeline only (memo records + task lifecycle events, chronological). True — adds `checkpoints` (decision/summary points only) and `stats` (duration, step count, subtask count, memo-type breakdown).

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv1.11.2
    • addedInput schema / properties / include_stats
      Added value: +{
      +  "default": false,
      +  "description": "False (default) — timeline only (memo records + task\nlifecycle events, chronological). True — adds `checkpoints`\n(decision/summary points only) and `stats` (duration, step count,\nsubtask count, memo-type breakdown).",
      +  "type": "boolean"
      +}
  2. First observedv1.9.0

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral burden. 'Get' implies a non-mutating read, and it discloses that include_stats layers derived summary on top of the timeline, which is useful. It does not state read-only safety or that the timeline is chronological beyond what the schema already says, so it is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero waste, with the plain-vs-stats distinction front-loaded before the 'how did this task go' framing. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and both parameters are fully documented in the schema. For a simple two-parameter read tool the description is nearly complete; only the lack of sibling routing keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, including a detailed description of include_stats and the exact fields it adds, so the schema already does the heavy lifting. The description restates the include_stats effect without adding syntax or format detail beyond the schema; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get a task's execution timeline') with a clear scope qualifier (plain vs. checkpoints+stats). It does not name which sibling to use instead (e.g., task_memo_read or diagnose_task_failure), so an agent gets a clear purpose but no sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase "Answers 'how did this task actually go'" implies the diagnostic use case, which is helpful context. However, there is no explicit when-not guidance and no routing to the nearby alternatives (diagnose_task_failure, failure_analysis, task_memo_read), leaving the agent to infer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools