AI Incident Lab MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@AI Incident Lab MCP Serverreceive the next incident alert and gather diagnostic evidence"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
AI Incident Lab
简体中文 · Architecture · Contributing · Security
An observable playground for testing how AI agents diagnose and recover from incidents through MCP.
Inject a fault, watch its impact spread through the service topology, and follow an agent from alert reception to evidence collection, diagnosis, SOP execution, and verified recovery. Compare agents using recorded outcomes and replay each experiment.
flowchart LR
Fault[Fault injection] --> Environment[Simulated services]
Environment --> Alert[Alert]
Alert --> Agent[Registered Agent]
Agent --> MCP[MCP tools + diagnostic skills]
MCP --> SOP[Safe SOP action]
SOP --> Environment
Environment --> Evaluation[Independent recovery evaluation]Quick start
Requires Docker with Docker Compose:
cp .env.example .env
docker compose up --buildOpen http://localhost:5173, choose the Reference Agent, and inject a fault. No API key is needed. API documentation is available at http://localhost:8000/docs.
The Reference Agent is a deterministic baseline. To use an LLM, configure LLM_API_KEY, LLM_BASE_URL, and LLM_MODEL in your local .env, restart the services, and select an LLM Agent. The provider must support Chat Completions function tools.
Related MCP server: SREGym Incident Replay Gate
What you can explore
Environment: ten simulated services and twelve dependencies, with node health, causal fault propagation, and fault markers visible to the operator.
Agent execution: MCP calls, structured evidence, diagnosis, action decisions, recovery verification, and downloadable reports.
Timeline and replay: metrics annotated with injection, alert, agent intervention, repair, and recovery events.
Comparison: agent rankings based on measured recovery, diagnosis, safety, and scenario coverage.
Fault | Injection target |
Connection leak | Worker |
CPU saturation | API |
Slow queries | Database |
Consumer backlog | Worker |
Service outage | Queue |
Shared connection pool exhaustion | Database proxy |
Services are stateful simulations, not real Redis/database workloads. This is a local experimentation platform; the control API is not authenticated for production use. Compose binds published ports to localhost.
Connect your agent
Register an external agent in the UI and save its one-time token locally.
Configure MCP at
http://localhost:5173/mcpusingAuthorization: Bearer <token>.Call
check_connection, then pollreceive_alertfor assignments.Use MCP observations, diagnostic skills, and SOP tools to diagnose, repair, and verify the incident.
For a working reference client, set INCIDENTLAB_AGENT_TOKEN in your ignored .env and run:
uv run python -m apps.agent.direct --url http://localhost:5173/mcp --onceSee the MCP integration guide for the tool contract and agent instructions. Tokens are scoped to registered agents; mutations and reports require the assigned run_id. Ground truth is kept outside agent observations.
Develop and verify
Requires Python 3.12+, uv, and Node.js 22.12+:
make install
make devmake test
make lint
cd apps/frontend
npx playwright install chromium
npm run test:e2eCI runs Python tests, static checks, frontend builds, browser tests, Compose smoke checks, and a full-history secret scan. Browser traces are disabled because registration responses contain credentials. Local databases, logs, screenshots, and environment files are excluded from Git.
Project layout
Directory | Responsibility |
| Control API, agent registry, orchestration |
| Isolated reference and LLM agents, direct MCP client |
| Environment, execution timeline, replay, leaderboard |
| MCP observation and action interface |
| Causal environment and fault simulation |
| Independent outcome evaluation |
| Shared contracts and utilities |
| Development and verification entry points |
| Unit and integration tests |
| Design, extension guides, verification records |
Chinese usage and API reference · Environment extension guide · Engineering specification · Verification records
A single active environment and one control worker keep experiments reproducible and implementation small. SQLite persists events and results; interrupted runs are marked failed when the control process restarts. LLM results depend on the selected provider and model. Action approval mode blocks automated actions; an approval UI is not implemented.
Contributing
See CONTRIBUTING.md before opening a pull request. The documentation and contribution layout takes inspiration from FastAPI and Chaos Mesh.
Licensed under MIT.
This server cannot be deployed
Maintenance
Related MCP Connectors
Public MCP digital twin with synthetic systems and an agent firewall. No customer data.
- mttrlyOAuthcom.mttrly
AI-powered incident management and server monitoring via MCP.
Build, validate, and manage API simulations in WireMock Cloud from MCP-compatible AI agents.
MCP server for building and testing AI agents with multi-model experimentation and insights.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceEnables AI agents to investigate and manipulate live network simulations via MCP, including protocol debugging and fault injection.-
- FlicenseNot gradedqualityDmaintenanceProve your SRE agent can resolve incidents before production by replaying synthetic incidents and scoring agent performance via MCP tools.-
- FlicenseNot gradedqualityBmaintenanceAn MCP-native AI incident response system that empowers agents to investigate production incidents, collect evidence, hypothesize root causes, and drive controlled remediation and recovery verification.1-
- AlicenseNot gradedqualityCmaintenanceEnables AI agents and software engineers to investigate and resolve realistic distributed-system incidents using safe MCP-based diagnostic tools for code, logs, metrics, traces, database, and deployments.Apache 2.0