MCP server for autonomous MLOps incident response, enabling drift detection, deployment history analysis, and human-approved rollback execution via gated tools.
Enables deterministic triage of OpenShift alerts by executing YAML-defined runbooks against an upstream OCP MCP server, collecting auditable evidence bundles for LLM interpretation.
A multi-agent MCP server that turns LLMs into an autonomous incident-response copilot, enabling rapid investigation, correlation, and remediation of production incidents.
MCP server for enterprise operations, enabling Sentry error triage, Linear issue sync, and Slack incident response with tools for parsing stacktraces, dispatching alerts, and generating postmortems.
Enables LLM agents to perform SRE reliability calculations like error budgets and burn rates using deterministic tools, integrating with Prometheus and Loki for real data.
Enables incident responders to answer who owns a service, who to escalate to and when, which runbook steps to try first, and whether the current escalation step is on time, all from a local file with no external APIs or network calls.