mcp-agent-trace-inspector
Allows exporting traces in OpenTelemetry OTLP JSON span format for use with OpenTelemetry-compatible observability backends.
Allows sending alert notifications to Slack when configured alert rules (latency, error rate, or cost) are triggered.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-agent-trace-inspectorShow me the summary of my last trace"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP Agent Trace Inspector
npm mcp-agent-trace-inspector package
Local-first, MCP-native observability for agent workflows. Every tool call, prompt transformation, latency, and token count is recorded in a local SQLite database — no cloud account, no API key, no traces leaving your machine. Built specifically for MCP rather than bolted onto a generic LLM proxy.
Tool reference | Configuration | Contributing | Troubleshooting | Design principles
Key features
Tool call tracing: Captures inputs, outputs, latency, and token usage for every step in a workflow.
Persistent storage: Traces survive session restarts; stored locally in SQLite with no external dependencies.
HTML dashboard: Generates a self-contained single-file dashboard with an interactive step timeline.
Token cost estimation: Calculates USD cost per trace using a configurable model pricing table — no API calls required.
Trace comparison: Diff two traces side by side to measure the impact of prompt or tool changes.
Low overhead: Adds less than 5ms per step; never becomes the bottleneck.
Related MCP server: MCP Agent Trace
Why this over LangSmith / AgentOps?
mcp-agent-trace-inspector | LangSmith / AgentOps | |
Data location | Local SQLite — never leaves your machine | Cloud-hosted; traces sent to external servers |
Setup |
| Account signup, API key, SDK instrumentation |
MCP-aware | Native — records tool calls as first-class steps | Generic LLM proxy; MCP structure is opaque |
Run diffs | Built-in | Separate paid feature or manual export |
Cost estimation | Offline tiktoken + configurable pricing table | Requires live API traffic through their proxy |
Overhead | <5ms per step | Network round-trip per event |
If your traces contain sensitive tool outputs, proprietary prompts, or data that must stay on-device, this is the right tool. If you need cross-team trace sharing or a managed SaaS, use LangSmith.
Disclaimers
mcp-agent-trace-inspector stores tool call inputs and outputs locally in a SQLite database. Traces may contain sensitive information passed to or returned from your tools. Review trace contents before sharing dashboard exports. Traces are not automatically transmitted; optional alert webhooks are available.
Requirements
Node.js v22.5.0 or newer.
npm.
Getting started
Add the following config to your MCP client:
{
"mcpServers": {
"trace-inspector": {
"command": "npx",
"args": ["-y", "mcp-agent-trace-inspector@latest"]
}
}
}To set a custom storage path:
{
"mcpServers": {
"trace-inspector": {
"command": "npx",
"args": [
"-y",
"mcp-agent-trace-inspector@latest",
"--db=~/traces/my-project.db"
]
}
}
}MCP Client configuration
Amp · Claude Code · Cline · Cursor · VS Code · Windsurf · Zed
Your first prompt
Enter the following in your MCP client to verify everything is working:
Start a trace called "test-run", then list the files in the current directory, then end the trace and show me the summary.Your client should return a summary showing step count, total tokens, and latency.
Tools
Trace lifecycle (3 tools)
trace_start— begin a new trace; returns atrace_idfor subsequent callstrace_step— record one tool call step (inputs, outputs, optional token count and latency)trace_end— mark a trace as completed
Inspection (4 tools)
list_traces— list stored traces with names, statuses, and timestampsget_trace_summary— token totals, step count, latency, and cost estimate for a tracecompare_traces— diff two traces side by side (step counts, tokens, latency)extract_reasoning_chain— extract only reasoning/thinking steps from a trace
Export (3 tools)
export_dashboard— generate a self-contained single-file HTML dashboard with latency waterfallexport_otel— export one or all traces in OpenTelemetry OTLP JSON span formatexport_compliance_log— export the compliance audit log as JSON or CSV, with optional date range filtering
Operations (3 tools)
configure_alerts— configure alert rules on latency, error rate, or cost; fire to Slack or generic webhooksset_retention_policy— set how many days to keep traces (in-memory; must be called beforeapply_retention)apply_retention— archive traces older than the configured threshold; delete traces past 2x the threshold
Configuration
--db / --db-path
Path to the SQLite database file used to store traces.
Type: string
Default: ~/.mcp/traces.db
--retention-days
Automatically delete traces older than N days. Set to 0 to disable.
Type: number
Default: 0
--pricing-table
Path to a JSON file containing custom model pricing ($/1K tokens). Overrides the built-in table.
Type: string
--no-token-count
Disable tiktoken-based token counting. Traces will omit token usage metrics.
Type: boolean
Default: false
Pass flags via the args property in your JSON config:
{
"mcpServers": {
"trace-inspector": {
"command": "npx",
"args": ["-y", "mcp-agent-trace-inspector@latest", "--retention-days=30"]
}
}
}Design principles
Append-only traces: Steps are immutable once recorded. Trust requires integrity.
Local-first: All core functionality works without a network connection.
Portable dashboards: HTML exports are always single-file; no server required to view them.
Verification
Before publishing a new version, verify the server with MCP Inspector to confirm all tools are exposed correctly and the protocol handshake succeeds.
Interactive UI (opens browser):
npm run build && npm run inspectCLI mode (scripted / CI-friendly):
# List all tools
npx @modelcontextprotocol/inspector --cli node dist/index.js --method tools/list
# List resources and prompts
npx @modelcontextprotocol/inspector --cli node dist/index.js --method resources/list
npx @modelcontextprotocol/inspector --cli node dist/index.js --method prompts/list
# Call a tool (example — replace with a relevant read-only tool for this plugin)
npx @modelcontextprotocol/inspector --cli node dist/index.js \
--method tools/call --tool-name list_traces
# Call a tool with arguments
npx @modelcontextprotocol/inspector --cli node dist/index.js \
--method tools/call --tool-name list_traces --tool-arg key=valueRun before publishing to catch regressions in tool registration and runtime startup.
Contributing
See CONTRIBUTING.md for full contribution guidelines.
npm install && npm testMCP Registry & Marketplace
This plugin is available on:
Search for mcp-agent-trace-inspector.
Available Tools
13 toolsapply_retentionADestructive
Apply the current retention policy: archives traces older than retention_days and deletes archived traces past 2× threshold. Example: {}
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, so the destructive nature is covered structurally; the description adds real value by naming the concrete consequences (archival threshold, deletion past 2× threshold). It stops short of stating irreversibility, whether deletion is permanent, or any permission requirements, which is the remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the action and its effects front-loaded. The trailing 'Example: {}' earns little for a zero-parameter tool and is arguably filler, but it is harmless and does signal that no arguments are required.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-argument, no-output tool whose annotations already flag destructive behavior, the description supplies the essential policy semantics needed to call it safely. It could be marginally more complete by clarifying that the 2× threshold is relative to retention_days, but nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The schema is an empty object and the description's 'Example: {}' reinforces that no input is needed; there is simply nothing further to document.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Apply the current retention policy') and then spells out both effects: archiving traces older than retention_days and deleting archived traces past 2× threshold. The phrase 'current retention policy' implicitly distinguishes it from the sibling set_retention_policy, so an agent can tell which one mutates policy versus executes it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: it applies an already-configured policy, so it presumably follows set_retention_policy. There is no explicit statement of when to call it, when not to, or what alternative exists if no policy has been set. Adequate but leaves the agent to infer ordering and prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_tracesARead-only
Compare two traces side by side: step counts, tokens, latency, and step-by-step differences. Example: { "trace_id_a": "abc-123", "trace_id_b": "def-456" }
| Name | Required | Description | Default |
|---|---|---|---|
| trace_id_a | Yes | ID of the first trace (baseline). Example: "abc-123" | |
| trace_id_b | Yes | ID of the second trace (to compare against). Example: "def-456" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds useful context by stating what the comparison surfaces (step counts, tokens, latency, step diffs), which is output content rather than behavioral detail – there is no mention of auth needs, limits, or edge cases like mismatched trace formats.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core capability, followed by a compact call example. Slightly redundant since the schema already provides the same example IDs, but nothing is wasted and the structure is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only two-parameter comparison tool, the definition covers what is compared and how the IDs pair, and annotations carry the safety profile. With no output schema, the enumerated comparison dimensions adequately tell the agent what to expect back; only edge-case behavior is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already documents both parameters with examples and the baseline/against semantics. The description's inline example reinforces the pairing but adds no meaning beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (compare) and resource (two traces), and enumerates the dimensions compared: step counts, tokens, latency, and step-by-step differences. This clearly differentiates it from get_trace_summary and list_traces, which inspect single or many traces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The pairing semantics are implied by naming trace_id_a as baseline and trace_id_b as the comparison target, but there is no explicit statement of when to reach for this tool versus get_trace_summary or extract_reasoning_chain, and no exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
configure_alertsC
Configure alert rules for latency, error rate, or cost thresholds. Example: { "rules": [{"type":"latency","threshold":1000}], "slack_webhook": "https://hooks.slack.com/..." }
| Name | Required | Description | Default |
|---|---|---|---|
| rules | Yes | Array of alert rules to configure. | |
| webhook_url | No | Optional generic webhook URL for JSON POST notifications. | |
| slack_webhook | No | Optional Slack webhook URL for Slack block notifications. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false), so the bar is lower, but the description adds nothing about behavior: it does not say whether supplied rules replace or append to existing alert rules, whether configuration is idempotent, or what permissions/webhook validity are required. For a mutation tool this is a notable gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the purpose, followed by a compact example. No filler, though the example consumes more space than the explanatory text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With full schema coverage and no output schema, the description is adequate but incomplete for a mutation tool: it omits replace-vs-merge semantics and any indication of success/error behavior, which an agent would need to call it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all three parameters including the type enum and threshold. The example adds a concrete composite shape for 'rules', which is mildly useful, but it shows only slack_webhook and never mentions webhook_url, so it does not extend much beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Configure alert rules') plus the three supported threshold categories (latency, error rate, cost). This clearly separates it from siblings like set_retention_policy or trace_start, though it does not explicitly name a contrasting sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description never says when to use this tool, what prerequisite setup is needed, or how it relates to the other configuration/monitoring siblings. The provided example implies a call shape but offers no when/when-not guidance or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_compliance_logBRead-only
Export the compliance audit log as JSON or CSV. Example: { "format": "json" } or { "format": "csv", "from_date": "2025-01-01", "to_date": "2025-12-31" }
| Name | Required | Description | Default |
|---|---|---|---|
| format | Yes | Output format: "json" or "csv". | |
| to_date | No | Optional ISO date string to filter entries to (inclusive). | |
| from_date | No | Optional ISO date string to filter entries from (inclusive). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered by structured data. The description adds nothing behavioral beyond that: no mention of auth/permission requirements, result size limits, pagination, or whether output is streamed versus returned inline. For a read-only export, the bar is lower, but no extra context is actually contributed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the core action front-loaded and examples kept to the point. The examples are mildly redundant with the enum already in the schema, costing a little density, but nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter read-only tool with full schema coverage this is close to sufficient, but there is no output schema and the description never says what an export returns (inline JSON/CSV payload versus a file/dataset reference), which is the one detail an agent needs before calling it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so format, from_date, and to_date are already documented with enum values and ISO-date semantics. The description's examples restate the same information without adding syntax or edge-case detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Export) and resource (compliance audit log) with the output formats enumerated, which cleanly separates it from siblings like export_dashboard and export_otel. It stops short of explicitly naming those siblings or scoping conditions, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The examples imply the intended usage (ad-hoc audit-log export with optional date filtering), but there is no explicit when-to-use guidance, no exclusions, and no routing away from the other export tools. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_dashboardARead-only
Generate a self-contained HTML dashboard for a trace with error highlighting and latency waterfall visualization. Returns the full HTML as a string. Example: { "trace_id": "abc-123" }
| Name | Required | Description | Default |
|---|---|---|---|
| trace_id | Yes | ID of the trace to export. Example: "abc-123-def-456" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered, and the description adds genuinely new behavior: the tool returns the complete HTML as a string (important, since there is no output schema) and the HTML is self-contained. It does not mention size limits, latency, or whether large traces are truncated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences plus a short example, front-loaded with the primary purpose and return type. Slight redundancy between the inline example and the schema example costs it a point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, read-only export tool with no output schema, the description supplies the missing return-value context (full HTML string, self-contained). Nothing essential for correct invocation is absent, though noting output size or truncation behavior would complete it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter with 100% schema description coverage; the schema already documents trace_id and its example. The description's example payload repeats, rather than extends, what the schema provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (generate) and resource (self-contained HTML dashboard) plus the two visual features it contains (error highlighting, latency waterfall). This distinguishes it cleanly from siblings like export_otel, export_compliance_log, and get_trace_summary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the trace-scoped purpose, but there is no explicit when-to-use vs. when-not guidance or routing to alternatives such as export_otel for machine-readable output or get_trace_summary for a quick view. An agent must infer the choice from the purpose line alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_otelARead-only
Export trace(s) in OpenTelemetry OTLP JSON span format. Omit trace_id to export all traces. Example: { "format": "json" } or { "trace_id": "abc-123", "format": "json" }
| Name | Required | Description | Default |
|---|---|---|---|
| format | Yes | Export format. Currently only "json" is supported. | |
| trace_id | No | Optional trace ID to export. Omit to export all traces. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered without description help. The description adds the bulk-export behavior when trace_id is omitted, but says nothing about size limits, streaming, or what happens with an unknown trace_id.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the core purpose front-loaded and the omit-trace_id rule stated before the example. The example is mildly redundant given the schema's enum and could be trimmed, but it is not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter export with a single supported format, the description covers purpose, the optional-filter behavior, and the wire format. With no output schema, it would ideally mention the shape of the exported payload, but the OTLP JSON mention largely covers that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented, and the enum on format is self-explanatory. The description's example reinforces usage but adds no format semantics beyond what the schema states; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (export) and resource (traces) plus the output serialization format (OpenTelemetry OTLP JSON span format), which is far more informative than the title. It does not, however, differentiate itself from siblings like export_compliance_log or export_dashboard.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives one useful scoping rule — omit trace_id to export all traces — which is genuine when-to-use guidance for parameter selection. It offers no guidance about when to prefer this tool over the other export_* siblings or any prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_reasoning_chainARead-only
Extract only the reasoning/thinking steps from a trace. Returns steps whose tool_name matches reasoning patterns (reason, think, plan, reflect, analyz, consider) or whose input/output content includes the word "think". Example: { "trace_id": "abc-123" }
| Name | Required | Description | Default |
|---|---|---|---|
| trace_id | Yes | ID of the trace to extract reasoning from. Example: "abc-123-def-456" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds valuable behavioral detail by specifying the exact filtering heuristic: steps whose tool_name matches reasoning patterns or whose content includes 'think'. This goes beyond the annotations, though it does not cover return format or pagination.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, followed by the matching logic and an example. Mostly efficient, but the example duplicates the schema's own example, which is slightly redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only tool with no output schema, the description is largely complete: it explains what is returned and the filtering rules. It could still mention response shape or ordering, but nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single trace_id parameter, including its own example. The description's example is redundant with the schema and adds no extra semantic meaning beyond what is already documented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Extract) and resource (reasoning/thinking steps from a trace). Clear purpose, but does not distinguish itself from siblings like get_trace_summary or trace_step, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance or comparison to alternatives. It states what it extracts but not when to prefer it over other trace-reading tools in the sibling set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_trace_summaryARead-only
Retrieve a summary of a trace including step count, total tokens, total latency, cost estimate, and reasoning chain detection. Example: { "trace_id": "abc-123", "model": "claude-sonnet-4-6" }
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Model name for cost estimation. Example: "claude-sonnet-4-6". Defaults to "claude-sonnet-4-6". | |
| trace_id | Yes | ID of the trace to summarize. Example: "abc-123-def-456" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds genuine value by disclosing the returned fields (step count, tokens, latency, cost estimate, reasoning chain detection), which is important since there is no output schema. No auth or rate-limit context, but with annotations present this is solid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the return contents front-loaded and an inline example. The example partly duplicates information already in the schema, but it is short and aids invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter, read-only summary tool with no output schema, the description does the important work of listing the returned metrics, and the schema documents both params. Missing only guidance on when to prefer it over sibling analysis tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (trace_id and model, including its default) are already fully documented in the schema. The description only restates them in an example, adding no meaning beyond the structured fields, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (retrieve) and resource (summary of a trace) and enumerates exactly what the summary contains: step count, tokens, latency, cost estimate, and reasoning chain detection. It is clear what the tool returns, though it does not explicitly distinguish itself from siblings like list_traces, compare_traces, or extract_reasoning_chain.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives such as list_traces or compare_traces, nor any prerequisites or exclusions. Usage is only implied by the tool name and the example payload.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_tracesARead-only
List all stored traces with their names, statuses, and timestamps. Example: { "limit": 10 } to get the 10 most recent traces.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of traces to return (most recent first). Example: 20. Omit for all. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds the returned field set and 'most recent first' ordering, but the ordering is already in the schema and no further behavior (rate limits, auth) is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the core purpose followed by a usage example. Slightly redundant with the schema's own example but still efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A simple read-only list tool with no output schema; the description adequately conveys the returned content and default ordering. Completeness would improve with brief alternative routing, but nothing essential to calling it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a single documented 'limit' parameter, so the schema does the heavy lifting. The description's example merely restates the limit value without adding new semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List all stored traces') and enumerates the returned fields (names, statuses, timestamps). It is clear but does not explicitly distinguish itself from siblings like get_trace_summary or compare_traces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The example ({ "limit": 10 }) implies usage but gives no explicit when-to-use guidance or conditions selecting this over sibling tools such as get_trace_summary or export_otel.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_retention_policyB
Set the trace retention policy (how many days to keep traces before archiving). Example: { "retention_days": 30 }
| Name | Required | Description | Default |
|---|---|---|---|
| retention_days | Yes | Number of days to retain traces before archiving. Must be > 0. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and destructiveHint=false, so the write/non-destructive profile is already covered. The description earns partial credit by disclosing what actually happens to expired traces ('before archiving'), which clarifies that data is archived rather than deleted. It says nothing about permissions, reversibility, or effect on traces already past the threshold.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence plus a minimal example, with the core action front-loaded and no filler. Every element carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-required-parameter tool with a fully documented schema and no output schema, the description is nearly sufficient. The remaining gap is the missing distinction from 'apply_retention', which an agent selecting between siblings would need.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already states the unit, meaning, and the '> 0' constraint more precisely than the description does. The example object adds usage syntax but no semantic detail beyond the schema's baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Set the trace retention policy') and clarifies the resource with a parenthetical definition of retention. It does not differentiate itself from the near-identical sibling 'apply_retention', so an agent cannot tell them apart from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives, which matters because the sibling list contains 'apply_retention', a plausible-looking duplicate. The example payload shows invocation syntax but gives no conditional guidance or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trace_endA
Mark a trace as completed. No further steps should be added after this call. Example: { "trace_id": "abc-123" }
| Name | Required | Description | Default |
|---|---|---|---|
| trace_id | Yes | ID of the trace to complete. Example: "abc-123-def-456" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=false, so the agent knows this is a non-destructive mutation. The description adds the terminal-finality trait, but says nothing about idempotency or what happens if the trace is already ended or does not exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences plus an example, with the core action front-loaded. The example partially duplicates the schema example, which is slightly redundant but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter terminal-marker tool with no output schema, the description covers the essential action and the key sequencing rule. Error behavior and idempotency are the only notable omissions, which are minor at this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single trace_id parameter, so the schema already carries the semantics and the description's inline example duplicates the schema's own example. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and object ('Mark a trace as completed'), which is unambiguous about the operation performed. It does not explicitly contrast itself with siblings like trace_start or trace_step, but the start/step/end naming family makes the role inferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The ordering constraint 'No further steps should be added after this call' is genuine timing guidance, telling the agent this is terminal. However, it names no alternatives or conditions for when to leave a trace open instead, so usage is only implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trace_startA
Begin a new agent workflow trace. Returns a trace_id to use with subsequent calls. Example: { "name": "My Search Workflow" }. Pass "auto" or leave name blank for an auto-generated name.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Human-readable name for this trace/workflow run. Example: "Product Search Workflow 2024-01-15". Pass "auto" to auto-generate. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, so the agent already knows this creates state; the description usefully adds that a trace_id is returned and that names can be auto-generated. It does not disclose persistence, lifecycle constraints, or limits on concurrent traces, so it adds moderate value beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with purpose and the return value, with no filler. The inline example object is slightly redundant given the schema but does not bloat the text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter state-creating tool with no output schema, the description covers purpose, return value, and the auto-naming option, and annotations cover the safety profile. Missing only lifecycle-adjacent context such as whether traces must be explicitly ended.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents the single name parameter fully; baseline is 3. The description largely restates the schema's 'auto' behavior, and its 'leave name blank' suggestion sits awkwardly against a parameter marked required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Begin a new agent workflow trace') and clarifies the return value (a trace_id). The start-of-lifecycle role is clear against siblings like trace_step and trace_end, but no sibling is named explicitly, so differentiation is inferred rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The line 'Returns a trace_id to use with subsequent calls' implies the tool is used first, before trace_step/trace_end, but it never states when not to use it or names an alternative. Guidance is present but must be inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trace_stepB
Record a single step within a trace. Captures tool name, input, output, optional token count and latency. Example: { "trace_id": "abc-123", "tool_name": "web_search", "input": {"query": "hello"}, "output": {"results": ["a","b"]}, "token_count": 50, "latency_ms": 320 }
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes | Input payload passed to the tool. Example: { "query": "weather in Paris" } | |
| output | Yes | Output/result returned by the tool. Example: { "result": "Sunny, 22°C" }. Add "error" field or "isError": true to flag errors. | |
| trace_id | Yes | ID of the trace to record this step under. Example: "abc-123-def-456" | |
| tool_name | Yes | Name of the tool or action that was executed. Example: "web_search" or "llm_call" | |
| latency_ms | No | Optional wall-clock latency of this step in milliseconds. Example: 450 | |
| token_count | No | Optional number of tokens consumed in this step. Example: 150 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations establish this as a non-readOnly, non-destructive write, so the safety profile is covered. The description adds that token_count and latency_ms are optional, but says nothing about idempotency, whether steps append to an existing trace, prerequisite trace existence, or how errors are surfaced besides the schema-level 'isError' note.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences front-load the purpose followed by a concrete example payload; every element earns its place. Slight redundancy in restating the fields the example already shows.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a nested-object write tool with no output schema, the description covers the call shape well but omits lifecycle context (trace must exist, ordering relative to trace_start/trace_end) and response behavior, leaving gaps an agent would need to resolve elsewhere.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with per-field descriptions and examples, so the schema already carries the parameter meaning. The description's inline example reinforces the shape but adds no semantics beyond the schema, which is the expected baseline when structured data does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Record a single step within a trace') and enumerates what is captured, which lets an agent separate it from trace_start/trace_end and the read/export siblings. It stops short of explicitly contrasting itself with those siblings, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the phrasing 'within a trace' and the example, suggesting it is called per step during a trace, but there is no explicit when-to-use guidance, no statement that a trace must already exist via trace_start, and no mention of the trace_end alternative. The agent must infer ordering from sibling names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
13 tool updates
v1.0.0- First observed
apply_retention - First observed
compare_traces - First observed
configure_alerts - First observed
export_compliance_log - First observed
export_dashboard - First observed
export_otel - First observed
extract_reasoning_chain - First observed
get_trace_summary - First observed
list_traces - First observed
set_retention_policy - First observed
trace_end - First observed
trace_start - First observed
trace_step
TDQS
Scored across 13 tools
The core lifecycle tools (trace_start, trace_step, trace_end, get_trace_summary, list_traces) are clearly distinct, and the retention pair (set_retention_policy vs apply_retention) is differentiated by config-vs-execute semantics. The three export tools (export_compliance_log, export_dashboard, export_otel) are the only mild risk, but their descriptions and output formats clearly separate them.
Most tools follow a consistent verb_noun snake_case convention (configure_alerts, set_retention_policy, export_compliance_log, get_trace_summary, list_traces). The trace_start/trace_step/trace_end trio uses noun_verb ordering, a minor deviation but internally consistent.
13 tools is well-scoped for a trace inspector, covering lifecycle, analysis, export, and governance. Each tool earns its place with no redundant entries.
The trace lifecycle (start/step/end/summary/list), comparison, reasoning extraction, retention, alerts, compliance, and multiple export formats are all covered. Minor gaps exist such as no direct trace deletion or per-step retrieval/filtering, but agents can work around these.
Maintenance
Related MCP Connectors
AI agent observability for production traces, natural-language insights, and improvement loops.
- memnodeOAuthdev.memnode
Persistent, inspectable memory for AI agents with lineage, correction, and a hosted MCP endpoint.
- SpanlyOAuthcom.spanly
MCP observability. Query live traffic, errors, duration, and alerts from your AI agent.
Private-by-default, local-first memory/context/task orchestrator for MCP apps and agents.
Related MCP Servers
- AlicenseAqualityBmaintenanceMCP-native agent evaluation and observability server. Log traces, evaluate output quality with 12 built-in rules (PII detection, prompt injection, cost thresholds), and track agent costs. Real-time dashboard, OTel-compatible spans. Self-hosted, MIT licensed.9649 npm9MIT
- AlicenseNot gradedqualityCmaintenanceEnables recording and analyzing AI agent execution traces, including event logging, metric computation, loop detection, and JSON export for debugging agent behavior.MIT
- AlicenseAqualityBmaintenanceMCP server for AI agent observability, providing trace and span logging, search, latency/tokens/cost metrics, and anomaly detection using an in-memory buffer.625 npmMIT
- AlicenseNot gradedqualityFmaintenanceLocal-first, auditable memory for AI agents. Provides durable context for MCP hosts with SQLite storage, CLI, and MCP tools for memory management.2Apache 2.0