Kube Triage
Provides read-only multi-cluster Kubernetes triage and health correlation, using deployment inventory, metric evidence, firing alerts, and reusable troubleshooting skills. It generates context-explicit Kubernetes commands for operators to review and run locally without connecting to the Kubernetes API or executing kubectl.
Reads deployment inventory from PostgreSQL using parameterized SELECT-only queries, correlating inventory freshness and deployment metadata for health assessment. Requires a read-only runtime role with SELECT access on the deployments table only.
Queries Prometheus metrics and firing alerts through approved instant and bounded range queries, correlating metric presence, freshness, and alert state with Kubernetes deployment health. Supports both fixture mode and HTTP Prometheus with enforced timeouts and response-size limits.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Kube Triagewhy is payments-api degraded in cluster-a?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Kube Triage
A read-only multi-cluster Kubernetes operations server for Model Context Protocol (MCP). Kube Triage correlates deployment inventory, Prometheus evidence, firing alerts, and reusable Kubernetes troubleshooting procedures.
The repository is complete and runnable without a Kubernetes cluster. The default demo uses a seeded SQLite inventory and deterministic Prometheus fixtures for two clusters. It exposes the same MCP tools, resources, prompts, response envelopes, and health correlation used by HTTP Prometheus and PostgreSQL deployments.
Safety Boundary
The MCP runtime never:
Executes
kubectl, shell, SSH, cloud CLI, or application CLI commands.Reads kubeconfig or connects to a Kubernetes API.
Creates, patches, edits, scales, restarts, or deletes workloads.
Executes database writes or changes Prometheus configuration.
Returns credentials, connection strings, raw environment variables, or Secrets.
Commands in skills are display-only. Every Kubernetes command includes
--context <context> and must be reviewed and run by the operator in an
authenticated local terminal. The demo seed script writes only the local SQLite
file during setup and is not imported or exposed by the MCP server.
Related MCP server: opsagent
Architecture
flowchart LR
USER[Operator] --> CLIENT[MCP AI Client]
CLIENT <-->|Streamable HTTP or stdio| SERVER[Kubernetes Operations MCP Server]
SERVER --> DB[(Read-only inventory)]
SERVER --> PROM[Fixture or HTTP Prometheus]
SERVER --> SKILLS[Validated skill registry]
CLIENT -->|Displays context-explicit command| USER
USER -->|Reviews and runs locally| CLI[Authenticated CLI]
CLI --> K8S[(Selected cluster)]Quick Start Without a Cluster
Prerequisites: Python 3.12 and PowerShell.
py -3.12 -m venv .venv
.\.venv\Scripts\python.exe -m pip install -e ".[dev]"
.\.venv\Scripts\python.exe scripts\seed_inventory.py
.\.venv\Scripts\python.exe -m kube_triage --transport streamable-httpThe server listens on:
MCP:
http://127.0.0.1:8080/mcpLiveness:
http://127.0.0.1:8080/health
The default database URL uses SQLite URI read-only mode. Running the seed script again resets only the local demo inventory.
Demo Scenarios
Target | Expected state | Evidence |
|
| 2/3 replicas, 7 restarts, warning alert |
|
| 0/2 replicas, critical alert |
|
| stale and missing metrics |
|
| 2/2 replicas, no alert or pressure signal |
The repeated payments/payments-api name proves that cluster_name remains part
of every dynamic target and prevents cross-cluster collisions.
VS Code and MCP Inspector
.vscode/mcp.json configures a local stdio server using the
workspace virtual environment. Use the VS Code MCP Servers view to start or debug
kube-triage.
For Streamable HTTP inspection, start the server and connect MCP Inspector to
http://127.0.0.1:8080/mcp:
npx -y @modelcontextprotocol/inspectorMCP Surface
Tools
Tool | Purpose |
| Find at most two skills and matching sanitized known issues. |
| List all valid skills and the malformed-file skip count. |
| Read one skill by stable ID for clients without resource support. |
| List inventory with exact cluster, namespace, owner, or app filters. |
| Resolve one complete cluster/namespace/deployment target. |
| Run one approved instant metric query. |
| Run one approved bounded metric timeline query. |
| Read firing alerts for an explicit cluster and optional target. |
| Correlate inventory, metrics, alerts, freshness, and skills. |
Every tool returns a ToolResultEnvelope with ok, data, error, and
observed_at. Expected failures identify validation, database, prometheus,
knowledge, or correlation as the responsible component.
Resources
policy://kube-triage/read-onlycatalog://kube-triage/metricscatalog://kube-triage/inventory-schemaskill://kube-triage/{category}/{name}for all six loaded skills
Prompts
deployment_health_checkinvestigate_deploymentcluster_health_summaryinvestigate_unknown_issue
Prompts tell the MCP client which tools to call. They do not collect data or run commands themselves.
Health Correlation
deployment_health evaluates:
Inventory freshness and desired replicas.
Live desired versus available replicas.
One-hour restart increase.
CPU usage versus the inventory CPU request.
Memory working set versus the inventory memory request.
Warning and critical firing alerts.
Metric presence and freshness.
State precedence is deterministic:
criticalwhen no desired replica is available or a critical alert fires.unknownwhen required noncritical evidence is missing or stale.degradedfor replica mismatch, restart/resource thresholds, or warnings.healthyonly when all required evidence is fresh and no issue is present.
Facts, hypotheses, matched skills, alerts, and missing evidence remain separate in the response.
Approved Prometheus Contract
Kube Triage does not accept arbitrary PromQL. ApprovedMetric permits:
desired_replicasavailable_replicasrestarts_last_1hcpu_usage_coresmemory_working_set_bytestarget_up
Deployment queries always include cluster, namespace, and deployment selectors. Range requests enforce a maximum duration and minimum step. HTTP requests enforce a timeout and response-size limit. The expected normalized recording rules are:
deployment:container_cpu_usage:rate5mdeployment:container_memory_working_set_bytes:sumdeployment:pod_restarts:increase1h
Each rule must expose cluster, namespace, and deployment labels.
Connect Real Read-Only Data
Copy .env.example to .env for local development, then override
only the values needed by the target environment.
For PostgreSQL:
DATABASE_URL=postgresql+psycopg://inventory_reader:...@db.example/inventoryGrant the runtime role CONNECT, schema USAGE, and SELECT on the deployments
table only. Do not grant insert, update, delete, DDL, or ownership privileges.
All SQLAlchemy queries in the runtime are parameterized SELECT statements.
For Prometheus:
PROMETHEUS_MODE=http
PROMETHEUS_URL=https://prometheus.example
PROMETHEUS_CLUSTER_LABEL=clusterOptional basic-auth values are consumed as secrets and are never included in MCP resources, tool responses, or logs. The server still requires no kubeconfig or Kubernetes credentials.
Configuration
Variable | Default | Purpose |
|
| Advertised server name. |
|
| HTTP bind address. |
| read-only demo SQLite URI | Inventory connection. |
|
|
|
| empty | Required only in HTTP mode. |
|
| HTTP request timeout. |
|
| Maximum timeline duration. |
|
| Minimum timeline step. |
|
| Maximum HTTP response body. |
|
| Metric freshness threshold. |
|
| Inventory freshness threshold. |
|
| Restart degradation threshold. |
|
| CPU request ratio threshold. |
|
| Memory request ratio threshold. |
|
| Markdown skill root. |
|
| Sanitized issue catalog. |
|
| Structured stderr log level. |
Docker
docker compose up --buildThe image runs as a non-root user, has all Linux capabilities dropped by Compose, uses a read-only container filesystem, and binds port 8080 to localhost. The demo database is seeded at build time and opened read-only at runtime.
Development
.\.venv\Scripts\python.exe -m ruff check .
.\.venv\Scripts\python.exe -m pytest
.\.venv\Scripts\python.exe -m buildTests use the official SDK's in-memory client transport. They cover tool/resource/ prompt discovery, structured protocol results, all four health states, exact and natural-language skill discovery, multi-cluster isolation, Prometheus bounds, SELECT-only query paths, and the absence of subprocess/Kubernetes client imports.
Project Layout
src/kube_triage/ MCP server, adapters, tools, resources, prompts
skill-registry/ Validated Markdown investigation procedures
known-issues/ Sanitized reusable failure patterns
fixtures/ Deterministic Prometheus demo evidence
scripts/ Setup-only local inventory seeding
tests/ Unit, protocol, and safety testsThe server uses the official MCP Python SDK v1 maintenance line pinned in pyproject.toml. The corresponding SDK references are recorded in .github/copilot-instructions.md.
This server cannot be deployed
Maintenance
Related MCP Connectors
Collision-first MCP diagnostics for MCP, API, webhook and agent failures. Returns structured evidence, bounded repair guidance, agent compatibility checks and safe recovery routing.
The Cortex MCP server provides read-only access to real-time engineering context from the Cortex developer portal, allowing AI coding assistants to answer natural language questions about your organization's catalog (microservices, libraries, domains, teams, infrastructure), scorecards (engineering standards and best practices), initiatives (goals and deadlines), and Engineering Intelligence metrics. It includes tools for querying documentation, tracking personal entities, and accessing AI-assisted insights across the entire Cortex ecosystem.
The Google GKE MCP server is a managed Model Context Protocol server that provides AI applications with tools to manage Google Kubernetes Engine (GKE) clusters and Kubernetes resources. It exposes a structured, discoverable interface that allows AI agents to interact with GKE and Kubernetes APIs, enabling them to inspect cluster configurations, retrieve Kubernetes resource YAMLs, monitor operations like cluster upgrades, diagnose issues, and optimize costs—all without needing to parse text output or use complex kubectl commands.
Governed data discovery, exact queries, decisions, simulations, and runtime utilities over MCP.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceIn-cluster MCP server for read-only diagnostics of Grafana, Prometheus, Alertmanager, and Loki, enabling metric queries, alert listings, and log queries through Grafana datasource proxies with Kubernetes RBAC authentication.-
- AlicenseNot gradedqualityBmaintenanceA read-only MCP server that exposes Kubernetes cluster telemetry tools (pods, events, logs, metrics, ArgoCD syncs) with automatic redaction for incident triage and hypothesis ranking.MIT
- FlicenseAqualityCmaintenanceEnables diagnosing Kubernetes clusters without creating or changing cluster objects, with tools to inspect health, workloads, events, storage, RBAC, custom resources, logs, and configuration.18-
- AlicenseNot gradedqualityCmaintenanceEnables read-only Kubernetes cluster diagnostics through MCP tools, allowing a local LLM to inspect pods and events and debug issues via natural language queries.MIT