Skip to main content
Glama
harinarayn

agent-trace-intelligence

by harinarayn

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
JUDGE_MODELNoModel to use for judging. Options include 'azure_ai/claude-opus-4-6', 'azure_ai/gpt-4.1', 'gpt-4o-mini', 'claude-haiku-4-5-20251001'azure_ai/claude-opus-4-6
OPENAI_API_KEYNoYour OpenAI API key (required for OpenAI-based judge model)
AZURE_AI_API_KEYNoYour Azure AI Foundry API key (required for Azure-based judge model)
AZURE_AI_API_BASENoYour Azure AI Foundry API base URL (required for Azure-based judge model)

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
judge_traceB

Diagnoses an agent trace — identifies root causes of failure, scores performance across four dimensions, and suggests a concrete fix. Returns verdict, grade, and plain-English explanation.

trace_breakdownB

Step-by-step scoring of every agent decision. Flags issues like redundant tool calls, reasoning gaps, and goal drift.

efficiency_scoreA

Deterministic efficiency analysis of token usage, tool redundancy, and latency. No API key required — runs instantly.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A3.5/5.0

Scored across 3 tools

Disambiguation4/5

The tools are largely distinct: judge_trace provides an overall verdict, trace_breakdown offers step-by-step scoring, and efficiency_score focuses on efficiency metrics. However, judge_trace and trace_breakdown both provide performance scores, which could cause some confusion about which to use for a given task.

Naming Consistency2/5

The naming pattern is inconsistent: judge_trace follows a verb_noun convention while trace_breakdown and efficiency_score use noun_noun. Although all names use snake_case, the lack of a consistent pattern makes it harder to predict tool names.

Tool Count5/5

With only three tools, the server is well-scoped for its niche purpose of agent trace analysis. Each tool addresses a distinct aspect (overall judgment, detailed breakdown, and efficiency), so none feels redundant.

Completeness4/5

The server covers the core analysis lifecycle: holistic diagnosis, step-by-step scoring, and efficiency measurement. Minor gaps exist, such as the lack of tools for comparing multiple traces or retrieving raw trace data, but these are not critical to the primary function.

Maintenance

ActivityInactive
ResponsivenessNo issues