tracehub-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| LOG_LEVEL | No | Logging level: DEBUG, INFO, WARNING, ERROR | INFO |
| BACKEND_URL | No | Backend API endpoint (required) | |
| BACKEND_TYPE | No | Backend type: jaeger, tempo, traceloop, or datadog | jaeger |
| BACKEND_API_KEY | No | API key (required for Traceloop and Datadog) | |
| BACKEND_APP_KEY | No | Application key (required for Datadog, in addition to API key) | |
| BACKEND_TIMEOUT | No | Request timeout in seconds | 30 |
| MAX_TRACES_PER_QUERY | No | Maximum traces to return per query (1-1000) | 100 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
| logging | {} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| search_tracesB | Search for OpenTelemetry traces with filters. Supports both simple parameters and advanced generic filter system. |
| get_traceA | Get complete trace details by trace ID. Returns all spans with attributes, including parsed Opentelemetry data for LLM operations. |
| triage_traceA | Synthesize a likely-root-cause diagnosis for a trace, instead of returning raw trace data for the caller to re-derive one from every time. Computes a critical path (the "Last Finishing Child" chain actually responsible for the trace's total latency), ranks spans by self-time (latency contribution net of children, top 10), and - when the trace contains an error anywhere under any root span - identifies the deepest error span in the trace's error chain as the likely root cause. Falls back to the highest self-time span as a pure-latency diagnosis when no error is present. Deterministic (no LLM call); works against any configured backend. |
| correlate_traceA | Try to find the corresponding trace in a second, independently- configured backend (e.g. a Datadog trace and its downstream Sentry error, joined) - given a trace_id known to the primary backend. Tries a direct trace_id match in the secondary backend first
(confidence "high"); if that fails, falls back to a time-window +
service-name-overlap heuristic search (confidence "low"). This is a
best-effort correlation, not a guaranteed join - see the always-present
|
| get_llm_usageB | Get aggregated LLM usage metrics (token counts) for a time period. Provides breakdowns by model and service. |
| list_servicesA | List all available services in the OpenTelemetry backend. Returns: JSON string with list of services |
| find_errorsB | Find traces with errors. Including detailed error messages, stack traces, and LLM-specific error information. |
| list_llm_modelsA | List all LLM models being used with usage statistics. Discovers what models are deployed and tracks their usage patterns. |
| get_llm_model_statsA | Get detailed performance statistics for a specific LLM model. Analyzes request count, latency percentiles (p50, p95, p99), token usage statistics, error rates, and finish reason distributions. |
| list_sessionsA | List conversations/sessions grouped by gen_ai.conversation.id. Groups spans that carry the gen_ai.conversation.id attribute (a real, cross-industry OTel semantic convention for session/conversation grouping) to surface per-conversation span counts, token usage, and time bounds - useful for understanding multi-turn conversation activity. |
| get_session_statsA | Get detailed statistics for a single conversation/session. Analyzes span count, distinct services, time bounds, LLM request/success/ error counts, latency percentiles, and token usage for every span sharing the given gen_ai.conversation.id. |
| compare_time_windowsA | Compare aggregated LLM usage metrics between two time windows. Runs the same usage aggregation for both ranges and returns the delta - useful for "this week vs last week" or "before/after a deploy" style comparisons of request/token counts. |
| investigate_cost_spikeA | Investigate an LLM cost spike: compare a recent window against a baseline and rank which models/services contributed most to the change. On-request/pull-based analysis, not a push alert - mirrors SigNoz's own "investigate telemetry cost" skill. Call this when you suspect (or want to check for) a cost increase, rather than polling get_llm_usage by hand. |
| investigate_error_spikeA | Investigate an error-rate spike: compare a recent window against a baseline and rank which services/models/error types contributed most. is_spike requires both an absolute error-count floor and a relative rate-multiplier to hold, so a tiny sample (e.g. 1 error becoming 2) doesn't read as a spike. |
| get_prompt_version_statsA | Get aggregated performance stats grouped by prompt name and version. Groups spans by gen_ai.prompt.name + gen_ai.prompt.version, mirroring Langfuse's shipped per-prompt Metrics tab. Real-world adoption of these two attributes is still thin, so this tool may often return an empty list until more instrumentations populate them. |
| get_llm_expensive_tracesA | Find traces with highest LLM token usage. Useful for cost optimization and identifying inefficient prompts. |
| get_llm_slow_tracesA | Find slowest LLM traces by duration. Useful for performance optimization and identifying latency bottlenecks. |
| search_spans_toolA | Search for individual OpenTelemetry spans with optional filters. Unlike search_traces, this returns individual spans rather than grouped traces, which is useful for analyzing specific operations or finding spans with certain characteristics (e.g., LLM tool calls with traceloop.span.kind == tool). |
| list_llm_tools_toolA | List all LLM tools being used by identifying traceloop.span.kind == tool. Discovers which tools/functions LLM applications are calling, grouped by tool name with usage statistics. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 19 tools
Most tools target clearly distinct resources or analytical tasks—trace lookup/search, LLM usage, sessions, models, prompts, and spike investigation are separated. A few tools overlap in spirit (search_traces vs find_errors, compare_time_windows vs investigate_cost_spike), but the descriptions provide enough differentiation to avoid serious misselection.
The overwhelming majority follow a clean snake_case verb_noun pattern (list_services, get_trace, investigate_cost_spike, get_llm_model_stats). The deviations are search_spans_tool and list_llm_tools_tool, whose redundant or awkward _tool suffix breaks the otherwise predictable convention.
Nineteen tools is on the heavy side and sits in the 16-25 range where a toolset starts to feel bulky. The breadth is somewhat justified by the combination of general tracing, LLM analytics, session analysis, and spike investigation, but several specialized investigation tools could arguably be folded into fewer general-purpose tools.
The surface covers trace search/retrieval, root-cause reasoning, error discovery, LLM usage/model stats, session grouping, prompt metrics, and cost/error spike analysis, leaving few obvious dead ends. Notable minor gaps are per-service detailed performance metrics and dependency/map-style analysis tools.