agent-trace-intelligence
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| JUDGE_MODEL | No | Model to use for judging. Options include 'azure_ai/claude-opus-4-6', 'azure_ai/gpt-4.1', 'gpt-4o-mini', 'claude-haiku-4-5-20251001' | azure_ai/claude-opus-4-6 |
| OPENAI_API_KEY | No | Your OpenAI API key (required for OpenAI-based judge model) | |
| AZURE_AI_API_KEY | No | Your Azure AI Foundry API key (required for Azure-based judge model) | |
| AZURE_AI_API_BASE | No | Your Azure AI Foundry API base URL (required for Azure-based judge model) |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| judge_traceB | Diagnoses an agent trace — identifies root causes of failure, scores performance across four dimensions, and suggests a concrete fix. Returns verdict, grade, and plain-English explanation. |
| trace_breakdownB | Step-by-step scoring of every agent decision. Flags issues like redundant tool calls, reasoning gaps, and goal drift. |
| efficiency_scoreA | Deterministic efficiency analysis of token usage, tool redundancy, and latency. No API key required — runs instantly. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 3 tools
The tools are largely distinct: judge_trace provides an overall verdict, trace_breakdown offers step-by-step scoring, and efficiency_score focuses on efficiency metrics. However, judge_trace and trace_breakdown both provide performance scores, which could cause some confusion about which to use for a given task.
The naming pattern is inconsistent: judge_trace follows a verb_noun convention while trace_breakdown and efficiency_score use noun_noun. Although all names use snake_case, the lack of a consistent pattern makes it harder to predict tool names.
With only three tools, the server is well-scoped for its niche purpose of agent trace analysis. Each tool addresses a distinct aspect (overall judgment, detailed breakdown, and efficiency), so none feels redundant.
The server covers the core analysis lifecycle: holistic diagnosis, step-by-step scoring, and efficiency measurement. Minor gaps exist, such as the lack of tools for comparing multiple traces or retrieving raw trace data, but these are not critical to the primary function.