Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
LOG_LEVELNoLogging level: DEBUG, INFO, WARNING, ERRORINFO
BACKEND_URLNoBackend API endpoint (required)
BACKEND_TYPENoBackend type: jaeger, tempo, traceloop, or datadogjaeger
BACKEND_API_KEYNoAPI key (required for Traceloop and Datadog)
BACKEND_APP_KEYNoApplication key (required for Datadog, in addition to API key)
BACKEND_TIMEOUTNoRequest timeout in seconds30
MAX_TRACES_PER_QUERYNoMaximum traces to return per query (1-1000)100

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": true
}
logging
{}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}

Tools

Functions exposed to the LLM to take actions

NameDescription
search_tracesB

Search for OpenTelemetry traces with filters.

Supports both simple parameters and advanced generic filter system.

get_traceA

Get complete trace details by trace ID.

Returns all spans with attributes, including parsed Opentelemetry data for LLM operations.

triage_traceA

Synthesize a likely-root-cause diagnosis for a trace, instead of returning raw trace data for the caller to re-derive one from every time.

Computes a critical path (the "Last Finishing Child" chain actually responsible for the trace's total latency), ranks spans by self-time (latency contribution net of children, top 10), and - when the trace contains an error anywhere under any root span - identifies the deepest error span in the trace's error chain as the likely root cause. Falls back to the highest self-time span as a pure-latency diagnosis when no error is present. Deterministic (no LLM call); works against any configured backend.

correlate_traceA

Try to find the corresponding trace in a second, independently- configured backend (e.g. a Datadog trace and its downstream Sentry error, joined) - given a trace_id known to the primary backend.

Tries a direct trace_id match in the secondary backend first (confidence "high"); if that fails, falls back to a time-window + service-name-overlap heuristic search (confidence "low"). This is a best-effort correlation, not a guaranteed join - see the always-present limitations in the result for why. Requires a secondary backend to be configured via SECONDARY_BACKEND_TYPE/SECONDARY_BACKEND_URL (and any backend-specific fields) environment variables; raises a clear error otherwise.

get_llm_usageB

Get aggregated LLM usage metrics (token counts) for a time period.

Provides breakdowns by model and service.

list_servicesA

List all available services in the OpenTelemetry backend.

Returns: JSON string with list of services

find_errorsB

Find traces with errors.

Including detailed error messages, stack traces, and LLM-specific error information.

list_llm_modelsA

List all LLM models being used with usage statistics.

Discovers what models are deployed and tracks their usage patterns.

get_llm_model_statsA

Get detailed performance statistics for a specific LLM model.

Analyzes request count, latency percentiles (p50, p95, p99), token usage statistics, error rates, and finish reason distributions.

list_sessionsA

List conversations/sessions grouped by gen_ai.conversation.id.

Groups spans that carry the gen_ai.conversation.id attribute (a real, cross-industry OTel semantic convention for session/conversation grouping) to surface per-conversation span counts, token usage, and time bounds - useful for understanding multi-turn conversation activity.

get_session_statsA

Get detailed statistics for a single conversation/session.

Analyzes span count, distinct services, time bounds, LLM request/success/ error counts, latency percentiles, and token usage for every span sharing the given gen_ai.conversation.id.

compare_time_windowsA

Compare aggregated LLM usage metrics between two time windows.

Runs the same usage aggregation for both ranges and returns the delta - useful for "this week vs last week" or "before/after a deploy" style comparisons of request/token counts.

investigate_cost_spikeA

Investigate an LLM cost spike: compare a recent window against a baseline and rank which models/services contributed most to the change.

On-request/pull-based analysis, not a push alert - mirrors SigNoz's own "investigate telemetry cost" skill. Call this when you suspect (or want to check for) a cost increase, rather than polling get_llm_usage by hand.

investigate_error_spikeA

Investigate an error-rate spike: compare a recent window against a baseline and rank which services/models/error types contributed most.

is_spike requires both an absolute error-count floor and a relative rate-multiplier to hold, so a tiny sample (e.g. 1 error becoming 2) doesn't read as a spike.

get_prompt_version_statsA

Get aggregated performance stats grouped by prompt name and version.

Groups spans by gen_ai.prompt.name + gen_ai.prompt.version, mirroring Langfuse's shipped per-prompt Metrics tab. Real-world adoption of these two attributes is still thin, so this tool may often return an empty list until more instrumentations populate them.

get_llm_expensive_tracesA

Find traces with highest LLM token usage.

Useful for cost optimization and identifying inefficient prompts.

get_llm_slow_tracesA

Find slowest LLM traces by duration.

Useful for performance optimization and identifying latency bottlenecks.

search_spans_toolA

Search for individual OpenTelemetry spans with optional filters.

Unlike search_traces, this returns individual spans rather than grouped traces, which is useful for analyzing specific operations or finding spans with certain characteristics (e.g., LLM tool calls with traceloop.span.kind == tool).

list_llm_tools_toolA

List all LLM tools being used by identifying traceloop.span.kind == tool.

Discovers which tools/functions LLM applications are calling, grouped by tool name with usage statistics.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A3.7/5.0

Scored across 19 tools

Disambiguation4/5

Most tools target clearly distinct resources or analytical tasks—trace lookup/search, LLM usage, sessions, models, prompts, and spike investigation are separated. A few tools overlap in spirit (search_traces vs find_errors, compare_time_windows vs investigate_cost_spike), but the descriptions provide enough differentiation to avoid serious misselection.

Naming Consistency4/5

The overwhelming majority follow a clean snake_case verb_noun pattern (list_services, get_trace, investigate_cost_spike, get_llm_model_stats). The deviations are search_spans_tool and list_llm_tools_tool, whose redundant or awkward _tool suffix breaks the otherwise predictable convention.

Tool Count3/5

Nineteen tools is on the heavy side and sits in the 16-25 range where a toolset starts to feel bulky. The breadth is somewhat justified by the combination of general tracing, LLM analytics, session analysis, and spike investigation, but several specialized investigation tools could arguably be folded into fewer general-purpose tools.

Completeness4/5

The surface covers trace search/retrieval, root-cause reasoning, error discovery, LLM usage/model stats, session grouping, prompt metrics, and cost/error spike analysis, leaving few obvious dead ends. Notable minor gaps are per-service detailed performance metrics and dependency/map-style analysis tools.

Maintenance

ActivityMaintained
ResponsivenessNo issues