Skip to main content
Glama

Kube Triage

A read-only multi-cluster Kubernetes operations server for Model Context Protocol (MCP). Kube Triage correlates deployment inventory, Prometheus evidence, firing alerts, and reusable Kubernetes troubleshooting procedures.

The repository is complete and runnable without a Kubernetes cluster. The default demo uses a seeded SQLite inventory and deterministic Prometheus fixtures for two clusters. It exposes the same MCP tools, resources, prompts, response envelopes, and health correlation used by HTTP Prometheus and PostgreSQL deployments.

Safety Boundary

The MCP runtime never:

  • Executes kubectl, shell, SSH, cloud CLI, or application CLI commands.

  • Reads kubeconfig or connects to a Kubernetes API.

  • Creates, patches, edits, scales, restarts, or deletes workloads.

  • Executes database writes or changes Prometheus configuration.

  • Returns credentials, connection strings, raw environment variables, or Secrets.

Commands in skills are display-only. Every Kubernetes command includes --context <context> and must be reviewed and run by the operator in an authenticated local terminal. The demo seed script writes only the local SQLite file during setup and is not imported or exposed by the MCP server.

Related MCP server: opsagent

Architecture

flowchart LR
    USER[Operator] --> CLIENT[MCP AI Client]
    CLIENT <-->|Streamable HTTP or stdio| SERVER[Kubernetes Operations MCP Server]
    SERVER --> DB[(Read-only inventory)]
    SERVER --> PROM[Fixture or HTTP Prometheus]
    SERVER --> SKILLS[Validated skill registry]
    CLIENT -->|Displays context-explicit command| USER
    USER -->|Reviews and runs locally| CLI[Authenticated CLI]
    CLI --> K8S[(Selected cluster)]

Quick Start Without a Cluster

Prerequisites: Python 3.12 and PowerShell.

py -3.12 -m venv .venv
.\.venv\Scripts\python.exe -m pip install -e ".[dev]"
.\.venv\Scripts\python.exe scripts\seed_inventory.py
.\.venv\Scripts\python.exe -m kube_triage --transport streamable-http

The server listens on:

  • MCP: http://127.0.0.1:8080/mcp

  • Liveness: http://127.0.0.1:8080/health

The default database URL uses SQLite URI read-only mode. Running the seed script again resets only the local demo inventory.

Demo Scenarios

Target

Expected state

Evidence

cluster-a/payments/payments-api

degraded

2/3 replicas, 7 restarts, warning alert

cluster-a/checkout/checkout-api

critical

0/2 replicas, critical alert

cluster-a/platform/event-consumer

unknown

stale and missing metrics

cluster-b/payments/payments-api

healthy

2/2 replicas, no alert or pressure signal

The repeated payments/payments-api name proves that cluster_name remains part of every dynamic target and prevents cross-cluster collisions.

VS Code and MCP Inspector

.vscode/mcp.json configures a local stdio server using the workspace virtual environment. Use the VS Code MCP Servers view to start or debug kube-triage.

For Streamable HTTP inspection, start the server and connect MCP Inspector to http://127.0.0.1:8080/mcp:

npx -y @modelcontextprotocol/inspector

MCP Surface

Tools

Tool

Purpose

search_skills

Find at most two skills and matching sanitized known issues.

get_skill_index

List all valid skills and the malformed-file skip count.

read_skill

Read one skill by stable ID for clients without resource support.

list_deployments

List inventory with exact cluster, namespace, owner, or app filters.

get_deployment

Resolve one complete cluster/namespace/deployment target.

prom_query

Run one approved instant metric query.

prom_query_range

Run one approved bounded metric timeline query.

get_firing_alerts

Read firing alerts for an explicit cluster and optional target.

deployment_health

Correlate inventory, metrics, alerts, freshness, and skills.

Every tool returns a ToolResultEnvelope with ok, data, error, and observed_at. Expected failures identify validation, database, prometheus, knowledge, or correlation as the responsible component.

Resources

  • policy://kube-triage/read-only

  • catalog://kube-triage/metrics

  • catalog://kube-triage/inventory-schema

  • skill://kube-triage/{category}/{name} for all six loaded skills

Prompts

  • deployment_health_check

  • investigate_deployment

  • cluster_health_summary

  • investigate_unknown_issue

Prompts tell the MCP client which tools to call. They do not collect data or run commands themselves.

Health Correlation

deployment_health evaluates:

  • Inventory freshness and desired replicas.

  • Live desired versus available replicas.

  • One-hour restart increase.

  • CPU usage versus the inventory CPU request.

  • Memory working set versus the inventory memory request.

  • Warning and critical firing alerts.

  • Metric presence and freshness.

State precedence is deterministic:

  1. critical when no desired replica is available or a critical alert fires.

  2. unknown when required noncritical evidence is missing or stale.

  3. degraded for replica mismatch, restart/resource thresholds, or warnings.

  4. healthy only when all required evidence is fresh and no issue is present.

Facts, hypotheses, matched skills, alerts, and missing evidence remain separate in the response.

Approved Prometheus Contract

Kube Triage does not accept arbitrary PromQL. ApprovedMetric permits:

  • desired_replicas

  • available_replicas

  • restarts_last_1h

  • cpu_usage_cores

  • memory_working_set_bytes

  • target_up

Deployment queries always include cluster, namespace, and deployment selectors. Range requests enforce a maximum duration and minimum step. HTTP requests enforce a timeout and response-size limit. The expected normalized recording rules are:

  • deployment:container_cpu_usage:rate5m

  • deployment:container_memory_working_set_bytes:sum

  • deployment:pod_restarts:increase1h

Each rule must expose cluster, namespace, and deployment labels.

Connect Real Read-Only Data

Copy .env.example to .env for local development, then override only the values needed by the target environment.

For PostgreSQL:

DATABASE_URL=postgresql+psycopg://inventory_reader:...@db.example/inventory

Grant the runtime role CONNECT, schema USAGE, and SELECT on the deployments table only. Do not grant insert, update, delete, DDL, or ownership privileges. All SQLAlchemy queries in the runtime are parameterized SELECT statements.

For Prometheus:

PROMETHEUS_MODE=http
PROMETHEUS_URL=https://prometheus.example
PROMETHEUS_CLUSTER_LABEL=cluster

Optional basic-auth values are consumed as secrets and are never included in MCP resources, tool responses, or logs. The server still requires no kubeconfig or Kubernetes credentials.

Configuration

Variable

Default

Purpose

MCP_SERVER_NAME

kube-triage

Advertised server name.

MCP_HOST / MCP_PORT

127.0.0.1 / 8080

HTTP bind address.

DATABASE_URL

read-only demo SQLite URI

Inventory connection.

PROMETHEUS_MODE

fixture

fixture or http.

PROMETHEUS_URL

empty

Required only in HTTP mode.

PROMETHEUS_TIMEOUT_SECONDS

15

HTTP request timeout.

PROMETHEUS_MAX_RANGE_HOURS

24

Maximum timeline duration.

PROMETHEUS_MIN_STEP_SECONDS

30

Minimum timeline step.

PROMETHEUS_MAX_RESPONSE_BYTES

2000000

Maximum HTTP response body.

METRIC_STALE_AFTER_SECONDS

300

Metric freshness threshold.

INVENTORY_STALE_AFTER_HOURS

168

Inventory freshness threshold.

HEALTH_RESTART_WARNING

5

Restart degradation threshold.

HEALTH_CPU_WARNING_RATIO

0.9

CPU request ratio threshold.

HEALTH_MEMORY_WARNING_RATIO

0.9

Memory request ratio threshold.

SKILL_REGISTRY_DIR

./skill-registry

Markdown skill root.

KNOWN_ISSUES_FILE

./known-issues/catalog.yaml

Sanitized issue catalog.

LOG_LEVEL

INFO

Structured stderr log level.

Docker

docker compose up --build

The image runs as a non-root user, has all Linux capabilities dropped by Compose, uses a read-only container filesystem, and binds port 8080 to localhost. The demo database is seeded at build time and opened read-only at runtime.

Development

.\.venv\Scripts\python.exe -m ruff check .
.\.venv\Scripts\python.exe -m pytest
.\.venv\Scripts\python.exe -m build

Tests use the official SDK's in-memory client transport. They cover tool/resource/ prompt discovery, structured protocol results, all four health states, exact and natural-language skill discovery, multi-cluster isolation, Prometheus bounds, SELECT-only query paths, and the absence of subprocess/Kubernetes client imports.

Project Layout

src/kube_triage/        MCP server, adapters, tools, resources, prompts
skill-registry/         Validated Markdown investigation procedures
known-issues/           Sanitized reusable failure patterns
fixtures/               Deterministic Prometheus demo evidence
scripts/                Setup-only local inventory seeding
tests/                  Unit, protocol, and safety tests

The server uses the official MCP Python SDK v1 maintenance line pinned in pyproject.toml. The corresponding SDK references are recorded in .github/copilot-instructions.md.

Related MCP Connectors

  • Collision-first MCP diagnostics for MCP, API, webhook and agent failures. Returns structured evidence, bounded repair guidance, agent compatibility checks and safe recovery routing.

  • The Cortex MCP server provides read-only access to real-time engineering context from the Cortex developer portal, allowing AI coding assistants to answer natural language questions about your organization's catalog (microservices, libraries, domains, teams, infrastructure), scorecards (engineering standards and best practices), initiatives (goals and deadlines), and Engineering Intelligence metrics. It includes tools for querying documentation, tracking personal entities, and accessing AI-assisted insights across the entire Cortex ecosystem.

  • The Google GKE MCP server is a managed Model Context Protocol server that provides AI applications with tools to manage Google Kubernetes Engine (GKE) clusters and Kubernetes resources. It exposes a structured, discoverable interface that allows AI agents to interact with GKE and Kubernetes APIs, enabling them to inspect cluster configurations, retrieve Kubernetes resource YAMLs, monitor operations like cluster upgrades, diagnose issues, and optimize costs—all without needing to parse text output or use complex kubectl commands.

  • Governed data discovery, exact queries, decisions, simulations, and runtime utilities over MCP.

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    In-cluster MCP server for read-only diagnostics of Grafana, Prometheus, Alertmanager, and Loki, enabling metric queries, alert listings, and log queries through Grafana datasource proxies with Kubernetes RBAC authentication.
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    A read-only MCP server that exposes Kubernetes cluster telemetry tools (pods, events, logs, metrics, ArgoCD syncs) with automatic redaction for incident triage and hypothesis ranking.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables read-only Kubernetes cluster diagnostics through MCP tools, allowing a local LLM to inspect pods and events and debug issues via natural language queries.
    MIT