Skip to main content
Glama
Pawangunjkar

agent-trace-mcp

by Pawangunjkar

Agent Trace MCP Server

Open-source MCP server owned by Pawan Gunjkar (pawangunjkar@gmail.com · GitHub). MIT licensed.

This is the trace log for agents, not for Java services. Prometheus and Loki show the commerce suite. This server shows which agent called which MCP tool, how long it took, whether the tool failed, and how many tokens the turn used.

Sibling servers: observability-mcp, github-mcp, deps-mcp.

Project information

Item

Value

Package

pawangunjkar-agent-trace-mcp

Runtime

Python 3.10+, FastMCP, stdio

Store

Local JSONL, default ~/.agent-trace-mcp/traces.jsonl

Write

trace_record appends one span

Reads

runs, one run, failed tools, token totals, slowest spans

Pass the same run_id for every tool in one user turn. Keep secrets out of args_preview. The preview is truncated to 300 characters.

Related MCP server: MCP Agent Trace

Architecture

flowchart TB
  subgraph L1["Layer 1 — Agent"]
    AG["LangGraph CommerceAgent"]
    HUB["McpHub.call"]
  end

  subgraph L2["Layer 2 — MCP"]
    SRV["agent-trace-mcp"]
  end

  subgraph L3["Layer 3 — Store"]
    FILE["traces.jsonl"]
  end

  subgraph L4["Layer 4 — Questions"]
    RUNS["trace_list_runs"]
    FAIL["trace_failed_tools"]
    COST["trace_token_cost"]
  end

  AG --> HUB
  HUB -->|"trace_record"| SRV
  SRV --> FILE
  FILE --> RUNS
  FILE --> FAIL
  FILE --> COST
flowchart LR
  START["User turn"] --> REC["trace_record per tool"]
  REC --> FILE["JSONL span"]
  FILE --> ASK["Which checkout step failed?"]
  ASK --> GET["trace_get_run"]

Tools

Tool

What it does

trace_record

Append agent, tool, latency, ok, error, tokens

trace_list_runs

Recent runs with span, error, and token counts

trace_get_run

Every span for one run_id

trace_failed_tools

Spans with ok=false

trace_token_cost

Sum tokens for one run or the whole file

trace_slowest_tools

Highest latency_ms

trace_status

File path and totals

Cursor

{
  "mcpServers": {
    "agent-trace": {
      "command": "uv",
      "args": ["run", "--directory", "C:/AI_Workspaces/Anti_Workspace/agent-trace-mcp", "server.py"],
      "env": {
        "TRACE_FILE": "C:/AI_Workspaces/Anti_Workspace/agent-trace-mcp/traces.jsonl"
      }
    }
  }
}

After a commerce-agent scenario, record one span per MCP call. Example: agent checkout, tool order-orchestrator.place_order, ok=false, error text from the tool JSON.

Available Tools

7 tools
trace_failed_toolsC

Return recent spans that failed.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure, and it discloses almost nothing beyond the core read operation. It does not define what 'failed' means (errors, exceptions, timeouts, cancellations), whether 'recent' implies a bounded time window, what ordering is used, or whether there is pagination. No contradictions exist, but the behavioral surface is essentially bare.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single six-word sentence with zero filler, and the verb 'Return' is front-loaded. It is efficient for what it says, though brevity here edges toward under-specification rather than disciplined conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return structure is covered elsewhere, and the tool has only one optional parameter. But the two key semantic terms, 'recent' and 'failed', are left undefined: the agent cannot know the recency window or what failure criterion determines inclusion, which materially affects correct interpretation of results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description should compensate but does not mention the single `limit` parameter at all. However, the parameter is low-risk: its name 'limit' and integer type with default 20 make the meaning (cap the number of returned spans) nearly self-evident, so an agent is unlikely to misuse it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Return'), a specific resource ('recent spans'), and a filter ('that failed'). This distinguishes it from siblings like trace_token_cost and trace_slowest_tools, which target cost and latency respectively. However, it does not fully differentiate from all siblings, since 'failed' is not defined and trace_list_runs or trace_get_run could overlap in some interpretations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the six siblings. No alternatives are named, no conditions are given (e.g., 'use trace_slowest_tools for latency debugging'), and an agent debugging a run has no stated basis for choosing 'failed spans' over 'list runs' or 'get run'. All usage context must be inferred from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trace_get_runB

Return every span for one run id.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral burden. It clearly identifies a read operation ('Return') and states scope ('every span' for 'one run id'), implying no mutation or pagination. However, it does not mention required permissions, error behavior, ordering, or cardinality beyond 'every span'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single seven-word sentence that front-loads the action and object. Every word contributes to the core meaning and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a single required parameter and the presence of an output schema, the description is mostly sufficient for a simple retrieval tool. It misses only richer usage context such as how a run_id is created (via trace_record) or which sibling would be better for status/token/failure data.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage and only names run_id with title 'Run Id.' The description merely repeats 'one run id' and adds no format, example, or guidance on where to obtain a valid run_id. Since there is no enum and no schema description, the description should have compensated but did not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear action and resource: 'Return every span for one run id.' This distinguishes the tool from siblings like trace_list_runs, since it requires a run_id and returns spans rather than listing runs, but it does not explicitly name or contrast any sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to choose this tool over trace_status, trace_list_runs, or trace_failed_tools. The only signal is the run_id requirement, which implies the agent must already have a run, but no explicit preconditions or alternatives are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trace_list_runsB

List recent agent runs with span counts, errors, and token totals.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. 'List' clearly implies a non-destructive read operation, so the safety profile is inferable. However, the description doesn't disclose what 'recent' means (time window), the default ordering, or pagination behavior — details an agent needs to interpret results correctly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with zero waste. The verb and resource appear first, followed by the key data points returned. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one optional parameter and a provided output schema (which covers the return format), the description is reasonably complete. The main gaps are behavioral: no definition of the 'recent' window, no ordering semantics, and no explanation of how limit interacts with the result set. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description never mentions the sole parameter 'limit'. While 'limit' is somewhat self-explanatory from its name and default of 20, the description must compensate at this low coverage level and adds nothing about how the limit constrains results (e.g., most recent N, or first N in chronological order).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('List recent agent runs') and enumerates the returned data ('span counts, errors, and token totals'). This clearly distinguishes it from specialized siblings like trace_failed_tools, trace_token_cost, and trace_slowest_tools. It doesn't explicitly contrast with trace_get_run, but the general listing role is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus its six siblings. The description never mentions alternatives or conditions (e.g., 'for a single run's details, use trace_get_run' or 'for cost breakdowns, use trace_token_cost'). Given the large sibling set, this is a significant gap in routing an agent to the right tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trace_recordA

Append one agent or MCP tool span. Reuse run_id for every span in the same user turn.

ParametersJSON Schema
NameRequiredDescriptionDefault
okNo
toolYes
agentYes
errorNo
run_idNo
tokens_inNo
latency_msNo
tokens_outNo
args_previewNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It clearly discloses that the tool appends/writes a span (a mutating action) and explains run_id grouping, but it does not discuss persistence semantics, whether run_id must already exist, or error behavior. This is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, with the core action front-loaded and the run_id guidance immediately after. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple enough that the purpose and grouping rule may suffice for a basic call, and an output schema exists to cover return values. However, with no annotations and no parameter descriptions for several fields, an agent still has to infer several semantics and the expected lifecycle around run_id.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the central grouping parameter (run_id) and indicates that agent and tool identify the span, but leaves ok, error, tokens_in/out, latency_ms, and args_preview to inference from their names. This is a minimal viable explanation rather than complete guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and object — 'Append one agent or MCP tool span' — so an agent knows exactly what operation is performed. This write-oriented purpose is clearly distinct from the read/aggregation sibling tools (trace_list_runs, trace_get_run, trace_failed_tools, trace_status, etc.).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear usage context: call this per span, and reuse run_id across all spans in the same user turn. It does not explicitly name alternatives or when-not-to-use conditions, but the sibling set makes the read-vs-write split obvious.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trace_slowest_toolsB

Return the slowest spans by latency_ms.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only says 'return,' which implies a read operation, but it does not disclose the scope of spans queried, the sort direction implied by 'slowest,' or the effect of the default limit. These behavioral details are left to inference.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single efficient sentence with no wasted words and the core purpose is front-loaded. It borders on under-specification, but for a simple one-parameter tool the brevity is acceptable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema covers return values and there is only one optional parameter, so complexity is low. However, the description omits scope (which spans are considered) and the ordering semantics of 'slowest,' leaving gaps that an agent would need to guess.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description does not mention the 'limit' parameter at all. The burden falls on the description to compensate for the schema's lack of documentation, and it fails to do so, even though the single parameter is fairly self-explanatory from its name and default value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('return'), a well-defined resource ('slowest spans'), and an explicit criterion ('by latency_ms'). This clearly distinguishes it from siblings like trace_failed_tools (failed tools) and trace_token_cost (cost), so an agent can tell them apart at a glance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives such as trace_failed_tools or trace_token_cost, and there is no mention of scope (e.g., within a run vs across all runs) or prerequisites. The when-to-use context is entirely absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trace_statusA

Show the JSONL path and how many spans are stored.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It indicates a read-only 'show' action and what it returns, but it does not note prerequisites, error behavior, or whether any state is modified. Still, for a status tool, the core behavior is reasonably clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, tightly worded sentence with no filler. It front-loads the key information and every word contributes to understanding what the tool does.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that the tool has no parameters and an output schema exists, the description covers the essential purpose. It could mention when the status is relevant (e.g., requiring a prior trace), but nothing critical is missing for basic invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is no parameter semantics to document. The description correctly focuses on the tool's output, and the baseline of 4 applies because parameter documentation is not needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Show') and names precise resources: the JSONL path and the span count. This clearly distinguishes it from sibling tools like trace_record or trace_list_runs, though it does not explicitly contrast itself with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied: call this when you need to know where spans are stored and how many exist. However, there is no explicit when-to-use statement or mention of alternatives/exclusions, so the agent must infer the usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trace_token_costA

Sum input and output tokens. Omit run_id to sum the whole file.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden. It communicates that this is an aggregation (sum) rather than a mutation and explains the file-wide versus run-scoped behavior triggered by omitting run_id. It does not discuss output shape, but an output schema is present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core behavior and followed by the only parameter nuance. No filler or repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-optional-parameter aggregation tool with an output schema, the description provides the purpose, the scope modes, and the behavior of the only parameter. Nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and run_id only has a title and default, so the description must compensate. 'Omit run_id to sum the whole file' meaningfully explains the parameter's optional scope behavior, though it could say more explicitly that providing run_id sums only that run.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Sum'), a resource ('input and output tokens'), and a scope option (whole file vs. run). It does not explicitly contrast with sibling tools, but the computation it describes is unique among the trace_* siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies the tool is for token-count totals and gives a useful scope condition ('Omit run_id to sum the whole file'). It does not state when to prefer this over sibling tools such as trace_get_run or trace_status, so the guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv1.0.0
    • First observedtrace_failed_tools
    • First observedtrace_get_run
    • First observedtrace_list_runs
    • First observedtrace_record
    • First observedtrace_slowest_tools
    • First observedtrace_status
    • First observedtrace_token_cost

TDQS

A3.7/5.0

Scored across 7 tools

Disambiguation5/5

Each tool has a distinct purpose: recording, listing runs, retrieving a run, filtering failures, computing token costs, finding slowest tools, and checking status. No overlaps or ambiguity.

Naming Consistency5/5

All tools follow the consistent pattern of 'trace_' prefix plus a descriptive verb or noun phrase, using snake_case throughout. This makes the tool set predictable and easy to navigate.

Tool Count5/5

With 7 tools, the server is well-scoped for tracing functionality—covering recording, querying, and analysis without excess or deficiency.

Completeness4/5

The surface covers core tracing needs: record, list, get, filter failures, cost, performance, and status. Minor gaps like clearing or deleting runs exist but are not critical for typical usage.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    B
    maintenance
    Records all MCP tool interactions in a centralized ledger, enabling developers to trace, replay, inspect, and audit AI agent workflows.
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables recording and analyzing AI agent execution traces, including event logging, metric computation, loop detection, and JSON export for debugging agent behavior.
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables live tracing and observability of MCP tool calls, persisting execution traces to MongoDB and streaming them in real time to a dashboard.
    -