agent-trace-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@agent-trace-mcpshow failed tool calls from the last agent run"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Agent Trace MCP Server
Open-source MCP server owned by Pawan Gunjkar (pawangunjkar@gmail.com · GitHub). MIT licensed.
This is the trace log for agents, not for Java services. Prometheus and Loki show the commerce suite. This server shows which agent called which MCP tool, how long it took, whether the tool failed, and how many tokens the turn used.
Sibling servers: observability-mcp, github-mcp, deps-mcp.
Project information
Item | Value |
Package |
|
Runtime | Python 3.10+, FastMCP, stdio |
Store | Local JSONL, default |
Write |
|
Reads | runs, one run, failed tools, token totals, slowest spans |
Pass the same run_id for every tool in one user turn. Keep secrets out of args_preview. The preview is truncated to 300 characters.
Related MCP server: MCP Agent Trace
Architecture
flowchart TB
subgraph L1["Layer 1 — Agent"]
AG["LangGraph CommerceAgent"]
HUB["McpHub.call"]
end
subgraph L2["Layer 2 — MCP"]
SRV["agent-trace-mcp"]
end
subgraph L3["Layer 3 — Store"]
FILE["traces.jsonl"]
end
subgraph L4["Layer 4 — Questions"]
RUNS["trace_list_runs"]
FAIL["trace_failed_tools"]
COST["trace_token_cost"]
end
AG --> HUB
HUB -->|"trace_record"| SRV
SRV --> FILE
FILE --> RUNS
FILE --> FAIL
FILE --> COSTflowchart LR
START["User turn"] --> REC["trace_record per tool"]
REC --> FILE["JSONL span"]
FILE --> ASK["Which checkout step failed?"]
ASK --> GET["trace_get_run"]Tools
Tool | What it does |
| Append agent, tool, latency, ok, error, tokens |
| Recent runs with span, error, and token counts |
| Every span for one |
| Spans with |
| Sum tokens for one run or the whole file |
| Highest |
| File path and totals |
Cursor
{
"mcpServers": {
"agent-trace": {
"command": "uv",
"args": ["run", "--directory", "C:/AI_Workspaces/Anti_Workspace/agent-trace-mcp", "server.py"],
"env": {
"TRACE_FILE": "C:/AI_Workspaces/Anti_Workspace/agent-trace-mcp/traces.jsonl"
}
}
}
}After a commerce-agent scenario, record one span per MCP call. Example: agent checkout, tool order-orchestrator.place_order, ok=false, error text from the tool JSON.
Available Tools
7 toolstrace_failed_toolsC
Return recent spans that failed.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it discloses almost nothing beyond the core read operation. It does not define what 'failed' means (errors, exceptions, timeouts, cancellations), whether 'recent' implies a bounded time window, what ordering is used, or whether there is pagination. No contradictions exist, but the behavioral surface is essentially bare.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single six-word sentence with zero filler, and the verb 'Return' is front-loaded. It is efficient for what it says, though brevity here edges toward under-specification rather than disciplined conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return structure is covered elsewhere, and the tool has only one optional parameter. But the two key semantic terms, 'recent' and 'failed', are left undefined: the agent cannot know the recency window or what failure criterion determines inclusion, which materially affects correct interpretation of results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description should compensate but does not mention the single `limit` parameter at all. However, the parameter is low-risk: its name 'limit' and integer type with default 20 make the meaning (cap the number of returned spans) nearly self-evident, so an agent is unlikely to misuse it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Return'), a specific resource ('recent spans'), and a filter ('that failed'). This distinguishes it from siblings like trace_token_cost and trace_slowest_tools, which target cost and latency respectively. However, it does not fully differentiate from all siblings, since 'failed' is not defined and trace_list_runs or trace_get_run could overlap in some interpretations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus the six siblings. No alternatives are named, no conditions are given (e.g., 'use trace_slowest_tools for latency debugging'), and an agent debugging a run has no stated basis for choosing 'failed spans' over 'list runs' or 'get run'. All usage context must be inferred from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trace_get_runB
Return every span for one run id.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the behavioral burden. It clearly identifies a read operation ('Return') and states scope ('every span' for 'one run id'), implying no mutation or pagination. However, it does not mention required permissions, error behavior, ordering, or cardinality beyond 'every span'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single seven-word sentence that front-loads the action and object. Every word contributes to the core meaning and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a single required parameter and the presence of an output schema, the description is mostly sufficient for a simple retrieval tool. It misses only richer usage context such as how a run_id is created (via trace_record) or which sibling would be better for status/token/failure data.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage and only names run_id with title 'Run Id.' The description merely repeats 'one run id' and adds no format, example, or guidance on where to obtain a valid run_id. Since there is no enum and no schema description, the description should have compensated but did not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action and resource: 'Return every span for one run id.' This distinguishes the tool from siblings like trace_list_runs, since it requires a run_id and returns spans rather than listing runs, but it does not explicitly name or contrast any sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to choose this tool over trace_status, trace_list_runs, or trace_failed_tools. The only signal is the run_id requirement, which implies the agent must already have a run, but no explicit preconditions or alternatives are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trace_list_runsB
List recent agent runs with span counts, errors, and token totals.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. 'List' clearly implies a non-destructive read operation, so the safety profile is inferable. However, the description doesn't disclose what 'recent' means (time window), the default ordering, or pagination behavior — details an agent needs to interpret results correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with zero waste. The verb and resource appear first, followed by the key data points returned. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one optional parameter and a provided output schema (which covers the return format), the description is reasonably complete. The main gaps are behavioral: no definition of the 'recent' window, no ordering semantics, and no explanation of how limit interacts with the result set. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions the sole parameter 'limit'. While 'limit' is somewhat self-explanatory from its name and default of 20, the description must compensate at this low coverage level and adds nothing about how the limit constrains results (e.g., most recent N, or first N in chronological order).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List recent agent runs') and enumerates the returned data ('span counts, errors, and token totals'). This clearly distinguishes it from specialized siblings like trace_failed_tools, trace_token_cost, and trace_slowest_tools. It doesn't explicitly contrast with trace_get_run, but the general listing role is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus its six siblings. The description never mentions alternatives or conditions (e.g., 'for a single run's details, use trace_get_run' or 'for cost breakdowns, use trace_token_cost'). Given the large sibling set, this is a significant gap in routing an agent to the right tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trace_recordA
Append one agent or MCP tool span. Reuse run_id for every span in the same user turn.
| Name | Required | Description | Default |
|---|---|---|---|
| ok | No | ||
| tool | Yes | ||
| agent | Yes | ||
| error | No | ||
| run_id | No | ||
| tokens_in | No | ||
| latency_ms | No | ||
| tokens_out | No | ||
| args_preview | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It clearly discloses that the tool appends/writes a span (a mutating action) and explains run_id grouping, but it does not discuss persistence semantics, whether run_id must already exist, or error behavior. This is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, with the core action front-loaded and the run_id guidance immediately after. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple enough that the purpose and grouping rule may suffice for a basic call, and an output schema exists to cover return values. However, with no annotations and no parameter descriptions for several fields, an agent still has to infer several semantics and the expected lifecycle around run_id.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the central grouping parameter (run_id) and indicates that agent and tool identify the span, but leaves ok, error, tokens_in/out, latency_ms, and args_preview to inference from their names. This is a minimal viable explanation rather than complete guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and object — 'Append one agent or MCP tool span' — so an agent knows exactly what operation is performed. This write-oriented purpose is clearly distinct from the read/aggregation sibling tools (trace_list_runs, trace_get_run, trace_failed_tools, trace_status, etc.).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear usage context: call this per span, and reuse run_id across all spans in the same user turn. It does not explicitly name alternatives or when-not-to-use conditions, but the sibling set makes the read-vs-write split obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trace_slowest_toolsB
Return the slowest spans by latency_ms.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It only says 'return,' which implies a read operation, but it does not disclose the scope of spans queried, the sort direction implied by 'slowest,' or the effect of the default limit. These behavioral details are left to inference.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence with no wasted words and the core purpose is front-loaded. It borders on under-specification, but for a simple one-parameter tool the brevity is acceptable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values and there is only one optional parameter, so complexity is low. However, the description omits scope (which spans are considered) and the ordering semantics of 'slowest,' leaving gaps that an agent would need to guess.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not mention the 'limit' parameter at all. The burden falls on the description to compensate for the schema's lack of documentation, and it fails to do so, even though the single parameter is fairly self-explanatory from its name and default value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('return'), a well-defined resource ('slowest spans'), and an explicit criterion ('by latency_ms'). This clearly distinguishes it from siblings like trace_failed_tools (failed tools) and trace_token_cost (cost), so an agent can tell them apart at a glance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives such as trace_failed_tools or trace_token_cost, and there is no mention of scope (e.g., within a run vs across all runs) or prerequisites. The when-to-use context is entirely absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trace_statusA
Show the JSONL path and how many spans are stored.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It indicates a read-only 'show' action and what it returns, but it does not note prerequisites, error behavior, or whether any state is modified. Still, for a status tool, the core behavior is reasonably clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly worded sentence with no filler. It front-loads the key information and every word contributes to understanding what the tool does.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that the tool has no parameters and an output schema exists, the description covers the essential purpose. It could mention when the status is relevant (e.g., requiring a prior trace), but nothing critical is missing for basic invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is no parameter semantics to document. The description correctly focuses on the tool's output, and the baseline of 4 applies because parameter documentation is not needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Show') and names precise resources: the JSONL path and the span count. This clearly distinguishes it from sibling tools like trace_record or trace_list_runs, though it does not explicitly contrast itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied: call this when you need to know where spans are stored and how many exist. However, there is no explicit when-to-use statement or mention of alternatives/exclusions, so the agent must infer the usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trace_token_costA
Sum input and output tokens. Omit run_id to sum the whole file.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. It communicates that this is an aggregation (sum) rather than a mutation and explains the file-wide versus run-scoped behavior triggered by omitting run_id. It does not discuss output shape, but an output schema is present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core behavior and followed by the only parameter nuance. No filler or repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-optional-parameter aggregation tool with an output schema, the description provides the purpose, the scope modes, and the behavior of the only parameter. Nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and run_id only has a title and default, so the description must compensate. 'Omit run_id to sum the whole file' meaningfully explains the parameter's optional scope behavior, though it could say more explicitly that providing run_id sums only that run.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Sum'), a resource ('input and output tokens'), and a scope option (whole file vs. run). It does not explicitly contrast with sibling tools, but the computation it describes is unique among the trace_* siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies the tool is for token-count totals and gives a useful scope condition ('Omit run_id to sum the whole file'). It does not state when to prefer this over sibling tools such as trace_get_run or trace_status, so the guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v1.0.0- First observed
trace_failed_tools - First observed
trace_get_run - First observed
trace_list_runs - First observed
trace_record - First observed
trace_slowest_tools - First observed
trace_status - First observed
trace_token_cost
TDQS
Scored across 7 tools
Each tool has a distinct purpose: recording, listing runs, retrieving a run, filtering failures, computing token costs, finding slowest tools, and checking status. No overlaps or ambiguity.
All tools follow the consistent pattern of 'trace_' prefix plus a descriptive verb or noun phrase, using snake_case throughout. This makes the tool set predictable and easy to navigate.
With 7 tools, the server is well-scoped for tracing functionality—covering recording, querying, and analysis without excess or deficiency.
The surface covers core tracing needs: record, list, get, filter failures, cost, performance, and status. Minor gaps like clearing or deleting runs exist but are not critical for typical usage.
Maintenance
Related MCP Connectors
Analytics for MCP servers. Query your tool calls, first-call success, retries and schema cost.
- SpanlyOAuthcom.spanly
MCP observability. Query live traffic, errors, duration, and alerts from your AI agent.
AI agent observability for production traces, natural-language insights, and improvement loops.
Hash-chained HMAC-signed audit log MCP for A2A (agent-to-agent) calls. Every tool-call, agent-ha...
Related MCP Servers
- FlicenseNot gradedqualityBmaintenanceRecords all MCP tool interactions in a centralized ledger, enabling developers to trace, replay, inspect, and audit AI agent workflows.-
- AlicenseNot gradedqualityCmaintenanceEnables recording and analyzing AI agent execution traces, including event logging, metric computation, loop detection, and JSON export for debugging agent behavior.MIT
- AlicenseNot gradedqualityDmaintenanceProvides unified AI agent observability including tracing, cost tracking, performance monitoring, anomaly detection, and audit trails via MCP.56 npmMIT
- FlicenseNot gradedqualityCmaintenanceEnables live tracing and observability of MCP tool calls, persisting execution traces to MongoDB and streaming them in real time to a dashboard.-