P12 SRE Ops MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@P12 SRE Ops MCP Servercheck SLO status for payment service"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
P12 · SRE Ops MCP Server
Custom MCP (Model Context Protocol) server exposing SRE tools as LLM-callable actions. Capstone project of the Staff SRE · AI Engineer Portfolio.
Where things live
What | Where |
MCP server code | This repo ( |
Interactive demo | |
Tool examples + audit log |
Related MCP server: Datadog MCP Server
Tools exposed
Tool | Description |
| SLO burn rate + error budget per service |
| Fetch runbook steps for an alert |
| Recent firing alerts by severity |
| Create incident record |
| Timeline of events for an incident |
| LLM-generated incident summary |
Use with Claude Desktop
{
"mcpServers": {
"sre-ops": {
"command": "python",
"args": ["-m", "src.mcp_server"],
"cwd": "/path/to/p12-sre-mcp"
}
}
}Then ask Claude: "What's burning right now?" — it calls query_alerts +
get_slo_status automatically and synthesizes a response.
SRE additions
Every tool call audit logged (timestamp, latency, success — postmortem-ready)
Tool latency tracked — SLO: p95 < 500ms
Graceful fallback when tools fail
JSON-RPC 2.0 protocol compliance
Both stdio (Claude Desktop) and HTTP (HF Space demo) transports
Run locally
git clone https://github.com/amarshiv86/p12-sre-mcp
cd p12-sre-mcp
pip install -r requirements.txt
# Run tests
pytest tests/ -v
# Start MCP server (stdio transport)
python -m src.mcp_server
# Test with manual JSON-RPC call
echo '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}' | python -m src.mcp_serverProject structure
p12-sre-mcp/
├── src/
│ ├── tools.py # 6 SRE tools + audit logging + timed_tool decorator
│ └── mcp_server.py # JSON-RPC 2.0 stdio transport
├── tests/
│ └── test_mcp.py # 20 tests (tools + MCP protocol)
├── hf_space/
│ ├── app.py # Gradio demo — scenario runner + manual tool call
│ ├── README.md # sdk_version: 5.29.0
│ └── requirements.txt
├── data/
│ ├── raw/tool_call_examples.jsonl
│ └── processed/sample_audit_log.json
├── .github/workflows/
│ ├── ci.yml
│ ├── deploy-hf-space.yml
│ └── deploy-hf-dataset.yml
└── requirements.txtStack
MCP Protocol · JSON-RPC 2.0 · Qwen2.5-0.5B · Gradio 5 · Audit logging · GitHub Actions
This server cannot be deployed
Maintenance
Related MCP Connectors
AI agent run monitoring with incident replay and SLA receipts.
Provides capabilities that let LLM agents perform a range of infrastructure management tasks.
Agentic CI operations for build inspection, failure diagnosis, and runner troubleshooting.
MCP-native AI SRE: ask what's broken in production, get a reviewed GitHub fix PR.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables AI-driven incident response by connecting Claude to monitoring tools like Prometheus, Grafana, Loki, PagerDuty, and Slack for automated investigation and runbook generation.9-
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to interact with Datadog's observability platform via natural language, covering metrics, logs, APM, monitors, dashboards, incidents, and infrastructure.670 npm1MIT
- FlicenseNot gradedqualityDmaintenanceProvides New Relic observability tools for AI assistants, enabling discovery, data access, alerting, incident response, and performance analytics via natural language queries.-
- AlicenseNot gradedqualityBmaintenanceEnables autonomous SRE incident investigation by allowing users to describe incidents in natural language. The agent follows a governed state machine to gather read-only evidence and produce grounded conclusions.MIT