A
licenseNot graded
qualityB
maintenanceEnables agents to run evals mid-task, providing tools to validate suites, execute model comparisons, and inspect calibration and win-rate reports.
MIT
No user-submitted related servers found.
Scored across 18 tools
Each tool targets a distinct resource (agent, run, test case, test set, metric, persona) with specific CRUD operations, leaving no ambiguity between tools.
All tool names follow a consistent verb_noun pattern (create_, get_, list_, update_), making it easy to predict tool purpose from the name.
18 tools cover the complete lifecycle of evaluating AI agents (agents, personas, test sets, runs, metrics) without being overly numerous or sparse.
Core workflows are covered, but missing delete operations for agents, test cases, and test sets, as well as update for test sets, leave minor gaps in the surface.