Agent Evaluation and Red-Team Platform
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Agent Evaluation and Red-Team Platformrun the support-core eval on the hardened agent and gate the release"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Agent Evaluation and Red-Team Platform
Evaluate an agent in a synthetic support sandbox, preserve its traces, score them with deterministic rules, and stop at a release gate when human review is required. The platform checks factuality, tools, permissions, injection resistance, data disclosure, citations, calibration, recovery, latency, estimated cost and repeatability. It never deploys an agent.
Version 0.2.0 is a local-first Python/MCP application with an optional authenticated HTTP server. For a quick review, open the browser demo, read the project case study, or try the offline command below.

Play the 2:30 walkthrough ? Inspect scorer limitations ? Read operational evidence
Try the offline demo
Use Python 3.12 or 3.14:
git clone https://github.com/ahines99/agent-eval-redteam.git
cd agent-eval-redteam
python -m pip install uv==0.12.18
python -m uv sync --frozen
python -m uv run --frozen agent-eval demoThe demo is offline, uses an in-memory database and needs no API key. It runs the current
35-case support-core@1.2.0 suite with three baseline repeats and five failure probes:
Control | Expected result |
Hardened reference | 35/35 cases pass; eligible |
Flaky candidate | Citation regressions; waits for review; demo records rejection |
Naive reference | Critical failures; blocked; override refused |
It also shows self-approval refusal, an ad-hoc timeout probe, regression monitoring and an evidence-linked report. The browser demo makes the output readable; the original terminal recording preserves real execution timing. The narration script explains each step. To retain results:
python -m uv run --frozen agent-eval demo --db sqlite:///./data/demo.db
python -m uv run --frozen agent-eval report RUN_ID --db sqlite:///./data/demo.dbReplace RUN_ID with a printed run id. The normal demo exits after printing its report.
Related MCP server: mcp-testing-tools
Use MCP or contribute
For an executable client-driven example:
python -m uv run --frozen python scripts/mcp_walkthrough.pyThe MCP quickstart includes desktop-client configuration,
Windows/POSIX environment examples and the health → registry → run → report sequence.
agent-eval serve is a long-running stdio server: it waits for an MCP client and does
not print an interactive prompt. A desktop client normally starts that process itself.
To run the development checks:
python -m uv sync --frozen --all-extras
python -m uv run --frozen pytest --cov --cov-report=term-missing
python -m uv run --frozen ruff check src tests migrations scripts
python -m uv run --frozen mypy srcSee contributing, changes, security reporting and the audit resolution. Installation metadata allows Python 3.12+, but the verification record determines which interpreter/OS combinations have actually been exercised; it does not establish support for future versions.
How it works
MCP client / CLI
|
Typed tools, resources, prompts -- HTTP identity/scopes when configured
|
Domain services: policies, deterministic scoring, statistics
|
Eight-step workflow: execution leases, checkpoints, release review
|
Sandbox + agent adapters SQL repository
Scripted controls / optional Claude SQLite / PostgreSQLThe workflow is Register system → Load eval suite → Run baseline → Inject failures → Score traces → Compare versions/models → Gate release → Monitor regressions. Checkpoint artifacts, audit events and pause state commit together. Execution leases prevent overlapping workers from writing the same run; persisted traces are reused on resume. A process dying after a provider response but before the trace is saved can still cause that uncommitted call to be repeated.
Scoring requires the complete expected trace manifest and verifies stored evidence.
The release baseline is the latest accepted run of the same agent name with matching
suite version, suite content, sandbox world, scorer and gate-policy identity. A supplied
baseline_run_id creates a separate informational comparison. A previously blocked
agent version remains blocked across suites.
MCP surface
Boundary | Tools |
Registry |
|
Evaluation |
|
Evidence |
|
Failure injection |
|
Policy |
|
Diagnostics |
|
Resources: project://policies, suites://{suite_id}/{version}, runs://{run_id}/report,
runs://{run_id}/audit. Prompts: review_run, triage_failures, plan_redteam.
The four skills provide evaluation, security, tool-use and reliability procedures.
Agents, suites and access
scripted agents are deterministic controls. The claude adapter implements a live
tool-use loop, validates model/pricing configuration and classifies provider failures;
its API behavior is covered by fake clients and a real Sonnet 5 validation run.
The first live run passed 1/10 cases under the existing strict rules and paused for review;
it recorded no critical security findings and about $0.104 in estimated token cost.
Read the result and measurement limits.
Install the claude extra and configure Anthropic credentials only for an authorized,
budgeted live run. Cost thresholds score completed traces; they are not a hard spend cap.
Published support-core@1.0.0 (31 cases) and @1.1.0 (35 cases) remain available and their
JSON files are unchanged. @1.2.0 strengthens content, retrieval, success and clarification
expectations. The shared world and scorer changed, so old and new runtime results are
not silently treated as comparable. Security authorization is derived from case content
and rechecked immediately before each case adapter invocation.
Stdio trusts access to the local process. HTTP requires AGENT_EVAL_AUTH_FILE: hashed
service tokens map to server-derived actors, scopes and separate tenant databases.
agent-eval serve --transport streamable-http binds to loopback port 8000; remote use
needs a TLS proxy. Missing configuration fails closed. See deployment
for setup, token handling, migrations, container commands and backups.
Why this is not just a chatbot
SQL stores runs, evidence, findings, decisions and audit events independently of conversations.
Versioned deterministic scorers produce reproducible verdicts from declared expectations.
Permission, authorization, evidence and gate checks execute in application code.
Critical failures cannot be overridden; review decisions require a different actor and a reason.
Regression tests exercise controls, scoring counterexamples, recovery, persistence, authenticated transport and migration behavior.
Delivery and limits
The repository includes locked dependencies, CI definitions, a PostgreSQL integration job, Docker/Compose, Alembic migrations, an MIT license and opt-in OpenTelemetry export. Local verification includes tests, typing/lint, package build, an installed-wheel demo, SQLite migrations, actual stdio restart and authenticated HTTP socket tests. A real Linux Docker build and behavioral container smoke passed, and PostgreSQL 17.11 passed the backend contract checks. These local results do not establish remote CI success or an externally deployed service. Check GitHub Actions and the verification record for release-specific evidence.
The sandbox models one fictional retailer. Phrase checks and PII recognition are bounded heuristics, not general semantic judges or complete data-loss prevention. Wilson intervals describe this case set; it is not a sample of production traffic. Hashes detect corruption but are not signatures against a database administrator. Signing remains deferred.
Admission limits bound suite size and work; per-database active-run and per-process HTTP limits reduce overload. They do not impose an aggregate provider billing limit, retention policy or organization-wide quota. Real-data deployments need explicit data handling, backup, retention and encryption decisions.
Details: architecture, data contracts, threat model, deployment.
Current verification and completion status, CI runs and releases.
This server cannot be deployed
Maintenance
Related MCP Connectors
Test the voice agents you run: scored transcripts, pass/fail verdicts, latency and WER metrics.
- OkareoOAuthcom.okareo
Simulation, evaluation and monitoring for voice agents.
Remote MCP for Gemini upgrade evals, prompt regressions, output diffs, and eval receipts.
MCP server for building and testing AI agents with multi-model experimentation and insights.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceEnables LLM evaluation and observability by uploading documents, building test sets, running RAG pipelines, and automatically scoring answers for groundedness, hallucination risk, retrieval quality, latency, and cost, with tools exposed to MCP-compatible clients.1MIT
- AlicenseAqualityDmaintenanceProvides testing and quality assurance tools for AI agents via MCP, enabling generation of test cases, mock data, API mocks, coverage analysis, and assertions.540 npmMIT
- AlicenseBqualityCmaintenanceEnables deterministic security testing of AI agents that use tools by serving synthetic MCP environments with poisoned data, fake secrets, and privileged actions. Records agent tool calls and evaluates security invariants (e.g., canary leaks, forbidden access, approval binding) without an LLM judge or real systems.8MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to test and gate their own code by listing scenarios, running scenario tests, and applying statistical release gates inline via MCP.Apache 2.0