A Kubernetes diagnostic agent that provides on-demand root cause analysis and human-in-the-loop remediation via Slack, using LLM reasoning with OPA-bounded security controls.
Gives an LLM agent a fixed, typed set of on-call tools over a self-hosted GitOps Kubernetes platform, reading metrics, logs, alerts and deploy history from Prometheus, Loki, Alertmanager and GitHub without any cluster write access. Rollbacks and config changes are proposed as one-shot ids that require explicit human confirmation before becoming pull requests against the Git repository, with both steps recorded in an audit log.
Enables AI assistants to safely interact with Kubernetes clusters through scoped, redacted tools, with credentials kept local and write operations requiring human approval.
An MCP server exposing Kubernetes-style diagnostic tools to an LLM agent, with a safety approval gate for destructive actions, all backed by a mock cluster for local testing.
A multi-agent MCP server that turns LLMs into an autonomous incident-response copilot, enabling rapid investigation, correlation, and remediation of production incidents.