agent-eval
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| list_sessionsA | List session transcripts under a root directory. Returns one summary per session: id, working directory, tool-call count, human turn count and cost. Use it to find the session worth analysing. Args: root: Directory to search. Defaults to the standard transcript location. limit: Maximum number of sessions to return. |
| analyze_sessionA | Analyse one session transcript end to end. Returns tool usage and failure rates, detected loops, test-run outcomes, permission posture, cost, and a short list of findings worth a human's attention. Prompt text and file contents are never included. Args: path: Path to a .jsonl session transcript. loop_threshold: How many repeats of one action count as a loop. |
| find_loopsA | Return only the repeated stateful actions in a session. A loop is the same command or edit issued at least Args: path: Path to a .jsonl session transcript. threshold: Repeats of one action that count as a loop. |
| cost_reportA | Aggregate cost across every session under a root. Returns total spend, total lines changed, and spend per 100 lines changed. Args: root: Directory to search. Defaults to the standard transcript location. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 4 tools
The tools are mostly distinct: list_sessions, analyze_session, find_loops, and cost_report each have clear responsibilities. find_loops overlaps somewhat with analyze_session since analyze_session already reports detected loops, but the focused purpose keeps them distinguishable.
Three tools follow a verb_noun pattern (list_sessions, analyze_session, find_loops), but cost_report breaks the pattern by leading with a noun rather than an action verb. This is a minor inconsistency in an otherwise readable set.
Four tools is well-scoped for a session-analysis server. Each tool earns its place by covering listing, deep analysis, loop detection, and cost aggregation without unnecessary bloat.
The tool surface covers the core workflow: enumerate available sessions, analyze an individual session, inspect loops specifically, and aggregate costs. No obvious necessary operation is missing for the stated domain.