Skip to main content
Glama
dbsectrainer

mcp-agent-trace-inspector

by dbsectrainer

MCP Agent Trace Inspector

npm mcp-agent-trace-inspector package

Local-first, MCP-native observability for agent workflows. Every tool call, prompt transformation, latency, and token count is recorded in a local SQLite database — no cloud account, no API key, no traces leaving your machine. Built specifically for MCP rather than bolted onto a generic LLM proxy.

Tool reference | Configuration | Contributing | Troubleshooting | Design principles

Key features

  • Tool call tracing: Captures inputs, outputs, latency, and token usage for every step in a workflow.

  • Persistent storage: Traces survive session restarts; stored locally in SQLite with no external dependencies.

  • HTML dashboard: Generates a self-contained single-file dashboard with an interactive step timeline.

  • Token cost estimation: Calculates USD cost per trace using a configurable model pricing table — no API calls required.

  • Trace comparison: Diff two traces side by side to measure the impact of prompt or tool changes.

  • Low overhead: Adds less than 5ms per step; never becomes the bottleneck.

Related MCP server: MCP Agent Trace

Why this over LangSmith / AgentOps?

mcp-agent-trace-inspector

LangSmith / AgentOps

Data location

Local SQLite — never leaves your machine

Cloud-hosted; traces sent to external servers

Setup

npx one-liner, zero config

Account signup, API key, SDK instrumentation

MCP-aware

Native — records tool calls as first-class steps

Generic LLM proxy; MCP structure is opaque

Run diffs

Built-in compare_traces diff

Separate paid feature or manual export

Cost estimation

Offline tiktoken + configurable pricing table

Requires live API traffic through their proxy

Overhead

<5ms per step

Network round-trip per event

If your traces contain sensitive tool outputs, proprietary prompts, or data that must stay on-device, this is the right tool. If you need cross-team trace sharing or a managed SaaS, use LangSmith.

Disclaimers

mcp-agent-trace-inspector stores tool call inputs and outputs locally in a SQLite database. Traces may contain sensitive information passed to or returned from your tools. Review trace contents before sharing dashboard exports. Traces are not automatically transmitted; optional alert webhooks are available.

Requirements

  • Node.js v22.5.0 or newer.

  • npm.

Getting started

Add the following config to your MCP client:

{
  "mcpServers": {
    "trace-inspector": {
      "command": "npx",
      "args": ["-y", "mcp-agent-trace-inspector@latest"]
    }
  }
}

To set a custom storage path:

{
  "mcpServers": {
    "trace-inspector": {
      "command": "npx",
      "args": [
        "-y",
        "mcp-agent-trace-inspector@latest",
        "--db=~/traces/my-project.db"
      ]
    }
  }
}

MCP Client configuration

Amp · Claude Code · Cline · Cursor · VS Code · Windsurf · Zed

Your first prompt

Enter the following in your MCP client to verify everything is working:

Start a trace called "test-run", then list the files in the current directory, then end the trace and show me the summary.

Your client should return a summary showing step count, total tokens, and latency.

Tools

Trace lifecycle (3 tools)

  • trace_start — begin a new trace; returns a trace_id for subsequent calls

  • trace_step — record one tool call step (inputs, outputs, optional token count and latency)

  • trace_end — mark a trace as completed

Inspection (4 tools)

  • list_traces — list stored traces with names, statuses, and timestamps

  • get_trace_summary — token totals, step count, latency, and cost estimate for a trace

  • compare_traces — diff two traces side by side (step counts, tokens, latency)

  • extract_reasoning_chain — extract only reasoning/thinking steps from a trace

Export (3 tools)

  • export_dashboard — generate a self-contained single-file HTML dashboard with latency waterfall

  • export_otel — export one or all traces in OpenTelemetry OTLP JSON span format

  • export_compliance_log — export the compliance audit log as JSON or CSV, with optional date range filtering

Operations (3 tools)

  • configure_alerts — configure alert rules on latency, error rate, or cost; fire to Slack or generic webhooks

  • set_retention_policy — set how many days to keep traces (in-memory; must be called before apply_retention)

  • apply_retention — archive traces older than the configured threshold; delete traces past 2x the threshold

Configuration

--db / --db-path

Path to the SQLite database file used to store traces.

Type: string Default: ~/.mcp/traces.db

--retention-days

Automatically delete traces older than N days. Set to 0 to disable.

Type: number Default: 0

--pricing-table

Path to a JSON file containing custom model pricing ($/1K tokens). Overrides the built-in table.

Type: string

--no-token-count

Disable tiktoken-based token counting. Traces will omit token usage metrics.

Type: boolean Default: false

Pass flags via the args property in your JSON config:

{
  "mcpServers": {
    "trace-inspector": {
      "command": "npx",
      "args": ["-y", "mcp-agent-trace-inspector@latest", "--retention-days=30"]
    }
  }
}

Design principles

  • Append-only traces: Steps are immutable once recorded. Trust requires integrity.

  • Local-first: All core functionality works without a network connection.

  • Portable dashboards: HTML exports are always single-file; no server required to view them.

Verification

Before publishing a new version, verify the server with MCP Inspector to confirm all tools are exposed correctly and the protocol handshake succeeds.

Interactive UI (opens browser):

npm run build && npm run inspect

CLI mode (scripted / CI-friendly):

# List all tools
npx @modelcontextprotocol/inspector --cli node dist/index.js --method tools/list

# List resources and prompts
npx @modelcontextprotocol/inspector --cli node dist/index.js --method resources/list
npx @modelcontextprotocol/inspector --cli node dist/index.js --method prompts/list

# Call a tool (example — replace with a relevant read-only tool for this plugin)
npx @modelcontextprotocol/inspector --cli node dist/index.js \
  --method tools/call --tool-name list_traces

# Call a tool with arguments
npx @modelcontextprotocol/inspector --cli node dist/index.js \
  --method tools/call --tool-name list_traces --tool-arg key=value

Run before publishing to catch regressions in tool registration and runtime startup.

Contributing

See CONTRIBUTING.md for full contribution guidelines.

npm install && npm test

MCP Registry & Marketplace

This plugin is available on:

Search for mcp-agent-trace-inspector.

Available Tools

13 tools
apply_retentionA
Destructive

Apply the current retention policy: archives traces older than retention_days and deletes archived traces past 2× threshold. Example: {}

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, so the destructive nature is covered structurally; the description adds real value by naming the concrete consequences (archival threshold, deletion past 2× threshold). It stops short of stating irreversibility, whether deletion is permanent, or any permission requirements, which is the remaining gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with the action and its effects front-loaded. The trailing 'Example: {}' earns little for a zero-parameter tool and is arguably filler, but it is harmless and does signal that no arguments are required.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-argument, no-output tool whose annotations already flag destructive behavior, the description supplies the essential policy semantics needed to call it safely. It could be marginally more complete by clarifying that the 2× threshold is relative to retention_days, but nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. The schema is an empty object and the description's 'Example: {}' reinforces that no input is needed; there is simply nothing further to document.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Apply the current retention policy') and then spells out both effects: archiving traces older than retention_days and deleting archived traces past 2× threshold. The phrase 'current retention policy' implicitly distinguishes it from the sibling set_retention_policy, so an agent can tell which one mutates policy versus executes it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: it applies an already-configured policy, so it presumably follows set_retention_policy. There is no explicit statement of when to call it, when not to, or what alternative exists if no policy has been set. Adequate but leaves the agent to infer ordering and prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compare_tracesA
Read-only

Compare two traces side by side: step counts, tokens, latency, and step-by-step differences. Example: { "trace_id_a": "abc-123", "trace_id_b": "def-456" }

ParametersJSON Schema
NameRequiredDescriptionDefault
trace_id_aYesID of the first trace (baseline). Example: "abc-123"
trace_id_bYesID of the second trace (to compare against). Example: "def-456"

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds useful context by stating what the comparison surfaces (step counts, tokens, latency, step diffs), which is output content rather than behavioral detail – there is no mention of auth needs, limits, or edge cases like mismatched trace formats.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core capability, followed by a compact call example. Slightly redundant since the schema already provides the same example IDs, but nothing is wasted and the structure is easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only two-parameter comparison tool, the definition covers what is compared and how the IDs pair, and annotations carry the safety profile. With no output schema, the enumerated comparison dimensions adequately tell the agent what to expect back; only edge-case behavior is unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already documents both parameters with examples and the baseline/against semantics. The description's inline example reinforces the pairing but adds no meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (compare) and resource (two traces), and enumerates the dimensions compared: step counts, tokens, latency, and step-by-step differences. This clearly differentiates it from get_trace_summary and list_traces, which inspect single or many traces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The pairing semantics are implied by naming trace_id_a as baseline and trace_id_b as the comparison target, but there is no explicit statement of when to reach for this tool versus get_trace_summary or extract_reasoning_chain, and no exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

configure_alertsC

Configure alert rules for latency, error rate, or cost thresholds. Example: { "rules": [{"type":"latency","threshold":1000}], "slack_webhook": "https://hooks.slack.com/..." }

ParametersJSON Schema
NameRequiredDescriptionDefault
rulesYesArray of alert rules to configure.
webhook_urlNoOptional generic webhook URL for JSON POST notifications.
slack_webhookNoOptional Slack webhook URL for Slack block notifications.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false), so the bar is lower, but the description adds nothing about behavior: it does not say whether supplied rules replace or append to existing alert rules, whether configuration is idempotent, or what permissions/webhook validity are required. For a mutation tool this is a notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the purpose, followed by a compact example. No filler, though the example consumes more space than the explanatory text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With full schema coverage and no output schema, the description is adequate but incomplete for a mutation tool: it omits replace-vs-merge semantics and any indication of success/error behavior, which an agent would need to call it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all three parameters including the type enum and threshold. The example adds a concrete composite shape for 'rules', which is mildly useful, but it shows only slack_webhook and never mentions webhook_url, so it does not extend much beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Configure alert rules') plus the three supported threshold categories (latency, error rate, cost). This clearly separates it from siblings like set_retention_policy or trace_start, though it does not explicitly name a contrasting sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description never says when to use this tool, what prerequisite setup is needed, or how it relates to the other configuration/monitoring siblings. The provided example implies a call shape but offers no when/when-not guidance or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_compliance_logB
Read-only

Export the compliance audit log as JSON or CSV. Example: { "format": "json" } or { "format": "csv", "from_date": "2025-01-01", "to_date": "2025-12-31" }

ParametersJSON Schema
NameRequiredDescriptionDefault
formatYesOutput format: "json" or "csv".
to_dateNoOptional ISO date string to filter entries to (inclusive).
from_dateNoOptional ISO date string to filter entries from (inclusive).

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered by structured data. The description adds nothing behavioral beyond that: no mention of auth/permission requirements, result size limits, pagination, or whether output is streamed versus returned inline. For a read-only export, the bar is lower, but no extra context is actually contributed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with the core action front-loaded and examples kept to the point. The examples are mildly redundant with the enum already in the schema, costing a little density, but nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple three-parameter read-only tool with full schema coverage this is close to sufficient, but there is no output schema and the description never says what an export returns (inline JSON/CSV payload versus a file/dataset reference), which is the one detail an agent needs before calling it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so format, from_date, and to_date are already documented with enum values and ISO-date semantics. The description's examples restate the same information without adding syntax or edge-case detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Export) and resource (compliance audit log) with the output formats enumerated, which cleanly separates it from siblings like export_dashboard and export_otel. It stops short of explicitly naming those siblings or scoping conditions, so it is clear but not fully differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The examples imply the intended usage (ad-hoc audit-log export with optional date filtering), but there is no explicit when-to-use guidance, no exclusions, and no routing away from the other export tools. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_dashboardA
Read-only

Generate a self-contained HTML dashboard for a trace with error highlighting and latency waterfall visualization. Returns the full HTML as a string. Example: { "trace_id": "abc-123" }

ParametersJSON Schema
NameRequiredDescriptionDefault
trace_idYesID of the trace to export. Example: "abc-123-def-456"

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered, and the description adds genuinely new behavior: the tool returns the complete HTML as a string (important, since there is no output schema) and the HTML is self-contained. It does not mention size limits, latency, or whether large traces are truncated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences plus a short example, front-loaded with the primary purpose and return type. Slight redundancy between the inline example and the schema example costs it a point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, read-only export tool with no output schema, the description supplies the missing return-value context (full HTML string, self-contained). Nothing essential for correct invocation is absent, though noting output size or truncation behavior would complete it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter with 100% schema description coverage; the schema already documents trace_id and its example. The description's example payload repeats, rather than extends, what the schema provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (generate) and resource (self-contained HTML dashboard) plus the two visual features it contains (error highlighting, latency waterfall). This distinguishes it cleanly from siblings like export_otel, export_compliance_log, and get_trace_summary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the trace-scoped purpose, but there is no explicit when-to-use vs. when-not guidance or routing to alternatives such as export_otel for machine-readable output or get_trace_summary for a quick view. An agent must infer the choice from the purpose line alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_otelA
Read-only

Export trace(s) in OpenTelemetry OTLP JSON span format. Omit trace_id to export all traces. Example: { "format": "json" } or { "trace_id": "abc-123", "format": "json" }

ParametersJSON Schema
NameRequiredDescriptionDefault
formatYesExport format. Currently only "json" is supported.
trace_idNoOptional trace ID to export. Omit to export all traces.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered without description help. The description adds the bulk-export behavior when trace_id is omitted, but says nothing about size limits, streaming, or what happens with an unknown trace_id.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with the core purpose front-loaded and the omit-trace_id rule stated before the example. The example is mildly redundant given the schema's enum and could be trimmed, but it is not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter export with a single supported format, the description covers purpose, the optional-filter behavior, and the wire format. With no output schema, it would ideally mention the shape of the exported payload, but the OTLP JSON mention largely covers that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented, and the enum on format is self-explanatory. The description's example reinforces usage but adds no format semantics beyond what the schema states; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (export) and resource (traces) plus the output serialization format (OpenTelemetry OTLP JSON span format), which is far more informative than the title. It does not, however, differentiate itself from siblings like export_compliance_log or export_dashboard.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives one useful scoping rule — omit trace_id to export all traces — which is genuine when-to-use guidance for parameter selection. It offers no guidance about when to prefer this tool over the other export_* siblings or any prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extract_reasoning_chainA
Read-only

Extract only the reasoning/thinking steps from a trace. Returns steps whose tool_name matches reasoning patterns (reason, think, plan, reflect, analyz, consider) or whose input/output content includes the word "think". Example: { "trace_id": "abc-123" }

ParametersJSON Schema
NameRequiredDescriptionDefault
trace_idYesID of the trace to extract reasoning from. Example: "abc-123-def-456"

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds valuable behavioral detail by specifying the exact filtering heuristic: steps whose tool_name matches reasoning patterns or whose content includes 'think'. This goes beyond the annotations, though it does not cover return format or pagination.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, followed by the matching logic and an example. Mostly efficient, but the example duplicates the schema's own example, which is slightly redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read-only tool with no output schema, the description is largely complete: it explains what is returned and the filtering rules. It could still mention response shape or ordering, but nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single trace_id parameter, including its own example. The description's example is redundant with the schema and adds no extra semantic meaning beyond what is already documented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Extract) and resource (reasoning/thinking steps from a trace). Clear purpose, but does not distinguish itself from siblings like get_trace_summary or trace_step, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance or comparison to alternatives. It states what it extracts but not when to prefer it over other trace-reading tools in the sibling set.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_trace_summaryA
Read-only

Retrieve a summary of a trace including step count, total tokens, total latency, cost estimate, and reasoning chain detection. Example: { "trace_id": "abc-123", "model": "claude-sonnet-4-6" }

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoModel name for cost estimation. Example: "claude-sonnet-4-6". Defaults to "claude-sonnet-4-6".
trace_idYesID of the trace to summarize. Example: "abc-123-def-456"

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds genuine value by disclosing the returned fields (step count, tokens, latency, cost estimate, reasoning chain detection), which is important since there is no output schema. No auth or rate-limit context, but with annotations present this is solid.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with the return contents front-loaded and an inline example. The example partly duplicates information already in the schema, but it is short and aids invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter, read-only summary tool with no output schema, the description does the important work of listing the returned metrics, and the schema documents both params. Missing only guidance on when to prefer it over sibling analysis tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (trace_id and model, including its default) are already fully documented in the schema. The description only restates them in an example, adding no meaning beyond the structured fields, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (retrieve) and resource (summary of a trace) and enumerates exactly what the summary contains: step count, tokens, latency, cost estimate, and reasoning chain detection. It is clear what the tool returns, though it does not explicitly distinguish itself from siblings like list_traces, compare_traces, or extract_reasoning_chain.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives such as list_traces or compare_traces, nor any prerequisites or exclusions. Usage is only implied by the tool name and the example payload.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_tracesA
Read-only

List all stored traces with their names, statuses, and timestamps. Example: { "limit": 10 } to get the 10 most recent traces.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of traces to return (most recent first). Example: 20. Omit for all.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds the returned field set and 'most recent first' ordering, but the ordering is already in the schema and no further behavior (rate limits, auth) is disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, front-loaded with the core purpose followed by a usage example. Slightly redundant with the schema's own example but still efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A simple read-only list tool with no output schema; the description adequately conveys the returned content and default ordering. Completeness would improve with brief alternative routing, but nothing essential to calling it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a single documented 'limit' parameter, so the schema does the heavy lifting. The description's example merely restates the limit value without adding new semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List all stored traces') and enumerates the returned fields (names, statuses, timestamps). It is clear but does not explicitly distinguish itself from siblings like get_trace_summary or compare_traces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The example ({ "limit": 10 }) implies usage but gives no explicit when-to-use guidance or conditions selecting this over sibling tools such as get_trace_summary or export_otel.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_retention_policyB

Set the trace retention policy (how many days to keep traces before archiving). Example: { "retention_days": 30 }

ParametersJSON Schema
NameRequiredDescriptionDefault
retention_daysYesNumber of days to retain traces before archiving. Must be > 0.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and destructiveHint=false, so the write/non-destructive profile is already covered. The description earns partial credit by disclosing what actually happens to expired traces ('before archiving'), which clarifies that data is archived rather than deleted. It says nothing about permissions, reversibility, or effect on traces already past the threshold.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence plus a minimal example, with the core action front-loaded and no filler. Every element carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-required-parameter tool with a fully documented schema and no output schema, the description is nearly sufficient. The remaining gap is the missing distinction from 'apply_retention', which an agent selecting between siblings would need.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema already states the unit, meaning, and the '> 0' constraint more precisely than the description does. The example object adds usage syntax but no semantic detail beyond the schema's baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Set the trace retention policy') and clarifies the resource with a parenthetical definition of retention. It does not differentiate itself from the near-identical sibling 'apply_retention', so an agent cannot tell them apart from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives, which matters because the sibling list contains 'apply_retention', a plausible-looking duplicate. The example payload shows invocation syntax but gives no conditional guidance or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trace_endA

Mark a trace as completed. No further steps should be added after this call. Example: { "trace_id": "abc-123" }

ParametersJSON Schema
NameRequiredDescriptionDefault
trace_idYesID of the trace to complete. Example: "abc-123-def-456"

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false, so the agent knows this is a non-destructive mutation. The description adds the terminal-finality trait, but says nothing about idempotency or what happens if the trace is already ended or does not exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences plus an example, with the core action front-loaded. The example partially duplicates the schema example, which is slightly redundant but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter terminal-marker tool with no output schema, the description covers the essential action and the key sequencing rule. Error behavior and idempotency are the only notable omissions, which are minor at this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single trace_id parameter, so the schema already carries the semantics and the description's inline example duplicates the schema's own example. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and object ('Mark a trace as completed'), which is unambiguous about the operation performed. It does not explicitly contrast itself with siblings like trace_start or trace_step, but the start/step/end naming family makes the role inferable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The ordering constraint 'No further steps should be added after this call' is genuine timing guidance, telling the agent this is terminal. However, it names no alternatives or conditions for when to leave a trace open instead, so usage is only implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trace_startA

Begin a new agent workflow trace. Returns a trace_id to use with subsequent calls. Example: { "name": "My Search Workflow" }. Pass "auto" or leave name blank for an auto-generated name.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesHuman-readable name for this trace/workflow run. Example: "Product Search Workflow 2024-01-15". Pass "auto" to auto-generate.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, so the agent already knows this creates state; the description usefully adds that a trace_id is returned and that names can be auto-generated. It does not disclose persistence, lifecycle constraints, or limits on concurrent traces, so it adds moderate value beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with purpose and the return value, with no filler. The inline example object is slightly redundant given the schema but does not bloat the text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter state-creating tool with no output schema, the description covers purpose, return value, and the auto-naming option, and annotations cover the safety profile. Missing only lifecycle-adjacent context such as whether traces must be explicitly ended.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents the single name parameter fully; baseline is 3. The description largely restates the schema's 'auto' behavior, and its 'leave name blank' suggestion sits awkwardly against a parameter marked required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Begin a new agent workflow trace') and clarifies the return value (a trace_id). The start-of-lifecycle role is clear against siblings like trace_step and trace_end, but no sibling is named explicitly, so differentiation is inferred rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The line 'Returns a trace_id to use with subsequent calls' implies the tool is used first, before trace_step/trace_end, but it never states when not to use it or names an alternative. Guidance is present but must be inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trace_stepB

Record a single step within a trace. Captures tool name, input, output, optional token count and latency. Example: { "trace_id": "abc-123", "tool_name": "web_search", "input": {"query": "hello"}, "output": {"results": ["a","b"]}, "token_count": 50, "latency_ms": 320 }

ParametersJSON Schema
NameRequiredDescriptionDefault
inputYesInput payload passed to the tool. Example: { "query": "weather in Paris" }
outputYesOutput/result returned by the tool. Example: { "result": "Sunny, 22°C" }. Add "error" field or "isError": true to flag errors.
trace_idYesID of the trace to record this step under. Example: "abc-123-def-456"
tool_nameYesName of the tool or action that was executed. Example: "web_search" or "llm_call"
latency_msNoOptional wall-clock latency of this step in milliseconds. Example: 450
token_countNoOptional number of tokens consumed in this step. Example: 150

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations establish this as a non-readOnly, non-destructive write, so the safety profile is covered. The description adds that token_count and latency_ms are optional, but says nothing about idempotency, whether steps append to an existing trace, prerequisite trace existence, or how errors are surfaced besides the schema-level 'isError' note.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences front-load the purpose followed by a concrete example payload; every element earns its place. Slight redundancy in restating the fields the example already shows.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a nested-object write tool with no output schema, the description covers the call shape well but omits lifecycle context (trace must exist, ordering relative to trace_start/trace_end) and response behavior, leaving gaps an agent would need to resolve elsewhere.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with per-field descriptions and examples, so the schema already carries the parameter meaning. The description's inline example reinforces the shape but adds no semantics beyond the schema, which is the expected baseline when structured data does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Record a single step within a trace') and enumerates what is captured, which lets an agent separate it from trace_start/trace_end and the read/export siblings. It stops short of explicitly contrasting itself with those siblings, so it is clear but not fully differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the phrasing 'within a trace' and the example, suggesting it is called per step during a trace, but there is no explicit when-to-use guidance, no statement that a trace must already exist via trace_start, and no mention of the trace_end alternative. The agent must infer ordering from sibling names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 13 tool updatesv1.0.0
    • First observedapply_retention
    • First observedcompare_traces
    • First observedconfigure_alerts
    • First observedexport_compliance_log
    • First observedexport_dashboard
    • First observedexport_otel
    • First observedextract_reasoning_chain
    • First observedget_trace_summary
    • First observedlist_traces
    • First observedset_retention_policy
    • First observedtrace_end
    • First observedtrace_start
    • First observedtrace_step

TDQS

A3.6/5.0

Scored across 13 tools

Disambiguation4/5

The core lifecycle tools (trace_start, trace_step, trace_end, get_trace_summary, list_traces) are clearly distinct, and the retention pair (set_retention_policy vs apply_retention) is differentiated by config-vs-execute semantics. The three export tools (export_compliance_log, export_dashboard, export_otel) are the only mild risk, but their descriptions and output formats clearly separate them.

Naming Consistency4/5

Most tools follow a consistent verb_noun snake_case convention (configure_alerts, set_retention_policy, export_compliance_log, get_trace_summary, list_traces). The trace_start/trace_step/trace_end trio uses noun_verb ordering, a minor deviation but internally consistent.

Tool Count5/5

13 tools is well-scoped for a trace inspector, covering lifecycle, analysis, export, and governance. Each tool earns its place with no redundant entries.

Completeness4/5

The trace lifecycle (start/step/end/summary/list), comparison, reasoning extraction, retention, alerts, compliance, and multiple export formats are all covered. Minor gaps exist such as no direct trace deletion or per-step retrieval/filtering, but agents can work around these.

Maintenance

ActivityMaintained
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    MCP-native agent evaluation and observability server. Log traces, evaluate output quality with 12 built-in rules (PII detection, prompt injection, cost thresholds), and track agent costs. Real-time dashboard, OTel-compatible spans. Self-hosted, MIT licensed.
    9
    649 npm
    9
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables recording and analyzing AI agent execution traces, including event logging, metric computation, loop detection, and JSON export for debugging agent behavior.
    MIT
  • A
    license
    Not graded
    quality
    F
    maintenance
    Local-first, auditable memory for AI agents. Provides durable context for MCP hosts with SQLite storage, CLI, and MCP tools for memory management.
    2
    Apache 2.0