Skip to main content
Glama
wuyilun

AI Incident Lab MCP Server

by wuyilun

AI Incident Lab

简体中文 · Architecture · Contributing · Security

An observable playground for testing how AI agents diagnose and recover from incidents through MCP.

Inject a fault, watch its impact spread through the service topology, and follow an agent from alert reception to evidence collection, diagnosis, SOP execution, and verified recovery. Compare agents using recorded outcomes and replay each experiment.

flowchart LR
    Fault[Fault injection] --> Environment[Simulated services]
    Environment --> Alert[Alert]
    Alert --> Agent[Registered Agent]
    Agent --> MCP[MCP tools + diagnostic skills]
    MCP --> SOP[Safe SOP action]
    SOP --> Environment
    Environment --> Evaluation[Independent recovery evaluation]

Quick start

Requires Docker with Docker Compose:

cp .env.example .env
docker compose up --build

Open http://localhost:5173, choose the Reference Agent, and inject a fault. No API key is needed. API documentation is available at http://localhost:8000/docs.

The Reference Agent is a deterministic baseline. To use an LLM, configure LLM_API_KEY, LLM_BASE_URL, and LLM_MODEL in your local .env, restart the services, and select an LLM Agent. The provider must support Chat Completions function tools.

Related MCP server: SREGym Incident Replay Gate

What you can explore

  • Environment: ten simulated services and twelve dependencies, with node health, causal fault propagation, and fault markers visible to the operator.

  • Agent execution: MCP calls, structured evidence, diagnosis, action decisions, recovery verification, and downloadable reports.

  • Timeline and replay: metrics annotated with injection, alert, agent intervention, repair, and recovery events.

  • Comparison: agent rankings based on measured recovery, diagnosis, safety, and scenario coverage.

Fault

Injection target

Connection leak

Worker

CPU saturation

API

Slow queries

Database

Consumer backlog

Worker

Service outage

Queue

Shared connection pool exhaustion

Database proxy

Services are stateful simulations, not real Redis/database workloads. This is a local experimentation platform; the control API is not authenticated for production use. Compose binds published ports to localhost.

Connect your agent

  1. Register an external agent in the UI and save its one-time token locally.

  2. Configure MCP at http://localhost:5173/mcp using Authorization: Bearer <token>.

  3. Call check_connection, then poll receive_alert for assignments.

  4. Use MCP observations, diagnostic skills, and SOP tools to diagnose, repair, and verify the incident.

For a working reference client, set INCIDENTLAB_AGENT_TOKEN in your ignored .env and run:

uv run python -m apps.agent.direct --url http://localhost:5173/mcp --once

See the MCP integration guide for the tool contract and agent instructions. Tokens are scoped to registered agents; mutations and reports require the assigned run_id. Ground truth is kept outside agent observations.

Develop and verify

Requires Python 3.12+, uv, and Node.js 22.12+:

make install
make dev
make test
make lint
cd apps/frontend
npx playwright install chromium
npm run test:e2e

CI runs Python tests, static checks, frontend builds, browser tests, Compose smoke checks, and a full-history secret scan. Browser traces are disabled because registration responses contain credentials. Local databases, logs, screenshots, and environment files are excluded from Git.

Project layout

Directory

Responsibility

apps/control_api/

Control API, agent registry, orchestration

apps/agent/

Isolated reference and LLM agents, direct MCP client

apps/frontend/

Environment, execution timeline, replay, leaderboard

lab_mcp/

MCP observation and action interface

simulator/

Causal environment and fault simulation

evaluator/

Independent outcome evaluation

packages/

Shared contracts and utilities

scripts/

Development and verification entry points

tests/

Unit and integration tests

docs/

Design, extension guides, verification records

Chinese usage and API reference · Environment extension guide · Engineering specification · Verification records

A single active environment and one control worker keep experiments reproducible and implementation small. SQLite persists events and results; interrupted runs are marked failed when the control process restarts. LLM results depend on the selected provider and model. Action approval mode blocks automated actions; an approval UI is not implemented.

Contributing

See CONTRIBUTING.md before opening a pull request. The documentation and contribution layout takes inspiration from FastAPI and Chaos Mesh.

Licensed under MIT.

Related MCP Connectors

Related MCP Servers