kubeaid-mcp
# KubeAid MCP
An [MCP](https://modelcontextprotocol.io) server for
[KubeAid](https://github.com/Obmondo/kubeaid)-managed Kubernetes clusters.
Plug it into Claude Code, Claude Desktop, Cursor or any MCP-compatible client
and ask questions like *"which apps are out of sync?"*, *"why is
kube-prometheus degraded?"* or *"sync the cert-manager app"* — the model
answers by calling this server's tools, which talk to your cluster through
your own kubeconfig credentials.
Built in TypeScript on the official
[`@modelcontextprotocol/sdk`](https://github.com/modelcontextprotocol/typescript-sdk)
and [`@kubernetes/client-node`](https://github.com/kubernetes-client/javascript).
**Author:** Deep Poharkar
## Why KubeAid-specific?
KubeAid runs clusters the GitOps way: every application is an ArgoCD
`Application` generated from your kubeaid-config repo, and **auto-sync is
disabled by default** — drift between Git and the cluster is normal and waits
for a deliberate sync. A generic Kubernetes assistant doesn't understand
that workflow. This server does:
- The ArgoCD `Application` CR is a first-class citizen: sync status, health,
failing conditions and drifted resources are one tool call away.
- Syncing is modeled as the explicit "deploy now" decision it is in KubeAid —
gated behind an opt-in flag, with `dry_run` support, and never allowed on
contexts you mark as protected.
- The bundled prompts encode KubeAid operational habits: review drift before
syncing, fix things in the kubeaid-config repo rather than hand-editing the
cluster.
## Tools
Read tools — always registered:
| Tool | What it does |
|------|--------------|
| `list_contexts` | Kubeconfig contexts the server can target, with default and write-protection flags. |
| `list_argocd_apps` | Every KubeAid-managed ArgoCD app with sync status, health, revision and auto-sync flag. Supports `only_problems`. |
| `describe_argocd_app` | One app in depth: sources, error conditions, drifted/unhealthy resources, last sync result, deploy history. |
| `list_namespaces` | Namespaces with status and age. |
| `list_nodes` | Nodes with Ready status, roles, kubelet version, internal IP. |
| `list_pods` | Pods with kubectl-style derived status (CrashLoopBackOff etc.), ready count, restarts, node, age. |
| `describe_pod` | Per-container states and last-crash details, conditions, and the pod's recent events. |
| `get_pod_logs` | Tail container logs; `previous: true` reads the crashed instance's logs. |
| `list_deployments` | Deployments with ready/up-to-date/available counts. |
| `get_events` | Recent events, optionally warnings-only, per namespace or cluster-wide. |
| `describe_resource` | Any resource by apiVersion/kind/name — including CRDs like `SealedSecret` or `Certificate`. |
Write tools — registered only when `KUBEAID_MCP_ALLOW_WRITES=true`:
| Tool | What it does |
|------|--------------|
| `sync_argocd_app` | Trigger an ArgoCD sync (the KubeAid "deploy now" action). Supports `revision`, `prune`, `dry_run`. |
| `scale_deployment` | Set a deployment's replica count. |
| `rollout_restart` | Rolling-restart a deployment/statefulset/daemonset. |
| `delete_pod` | Delete a stuck pod so its controller replaces it. |
Every tool takes an optional `context` argument to target any cluster in your
kubeconfig from a single server process.
## Prompts
Prompts appear as slash commands / menu items in the client and walk the
model through a workflow using the tools above:
| Prompt | Workflow |
|--------|----------|
| `diagnose_app` | Root-cause a Degraded/OutOfSync app: conditions → resources → pod logs → recommendation. |
| `sync_review` | List all drifted apps, assess the risk of each pending sync, then ask before syncing anything. |
| `cluster_health_check` | Sweep nodes, apps, pods and warning events; report findings by severity. |
## Quick start
Requires Node.js 20+.
```bash
git clone <your-repo-url> kubeaid-mcp
cd kubeaid-mcp
npm install
npm run build
```
### Claude Code
```bash
claude mcp add kubeaid -- node "$(pwd)/dist/index.js"
```
Read-only by default. To enable writes while keeping production untouchable:
```bash
claude mcp add kubeaid \
-e KUBEAID_MCP_ALLOW_WRITES=true \
-e KUBEAID_MCP_PROTECTED_CONTEXTS=prod-cluster-1,prod-cluster-2 \
-- node "$(pwd)/dist/index.js"
```
### Claude Desktop
Enable **Settings → Developer → Local MCP servers**, then add to the config
file that page opens:
```json
{
"mcpServers": {
"kubeaid": {
"command": "node",
"args": ["/absolute/path/to/kubeaid-mcp/dist/index.js"],
"env": {
"KUBEAID_MCP_ALLOW_WRITES": "true",
"KUBEAID_MCP_PROTECTED_CONTEXTS": "prod-cluster-1"
}
}
}
}
```
Fully quit and reopen the app afterwards.
## Configuration
All configuration is environment variables, read once at startup:
| Variable | Default | Meaning |
|----------|---------|---------|
| `KUBECONFIG` | client-node default lookup (`~/.kube/config`) | Kubeconfig path. |
| `KUBEAID_MCP_CONTEXT` | follows current-context live | Pin a fixed default context. When unset, `kubectl config use-context` retargets the server mid-session; per-call `context` arguments always win. |
| `KUBEAID_MCP_ARGOCD_NAMESPACE` | `argocd` | Namespace holding ArgoCD's `Application` CRs. |
| `KUBEAID_MCP_ALLOW_WRITES` | `false` | Register the mutating tools at all. |
| `KUBEAID_MCP_PROTECTED_CONTEXTS` | none | Comma-separated contexts that refuse every mutating call, even with writes enabled. |
## Safety model
Three independent layers:
1. **Capability gating** — with writes disabled (the default), mutating tools
are never registered, so the model cannot even see them.
2. **Protected contexts** — every mutating handler re-checks the target
context against `KUBEAID_MCP_PROTECTED_CONTEXTS` and refuses protected
ones, whatever the model asks for.
3. **Client confirmation** — tools carry MCP `readOnlyHint`/`destructiveHint`
annotations, so well-behaved clients prompt you before risky calls.
The server holds no credentials of its own; it can only do what your
kubeconfig identity is already authorized to do via RBAC.
## Development
```bash
npm run typecheck # strict TS, no emit
npm run build # compile to dist/
npm run smoke # spawn the server, MCP handshake, list tools + prompts
KUBEAID_MCP_ALLOW_WRITES=true npm run smoke # verify write tools register
```
The smoke test needs no cluster — it drives the raw JSON-RPC protocol over
stdio.
## Roadmap
- `check_chart_updates`: compare deployed chart versions against the upstream
KubeAid `argocd-helm-charts` directory.
- Sealed-secrets helpers: list `SealedSecret` status, flag failed unseals.
- Prometheus/Alertmanager tools: surface firing alerts from the
kube-prometheus stack KubeAid ships.
## Acknowledgements
- [KubeAid](https://github.com/Obmondo/kubeaid) by Obmondo — the cluster
management stack this server is built for.
- [Argo CD](https://argo-cd.readthedocs.io/) — the GitOps engine underneath
KubeAid.
## License
MIT © 2026 Deep Poharkar
TDQS
Scored across 11 tools
Each tool targets a distinct resource/action - nodes, events, ArgoCD apps, pods, deployments, generic resource fetch, contexts, namespaces - with no overlapping purposes. The list/describe pairs for pods and ArgoCD apps are clearly separated by summary vs. detail, and describe_resource is explicitly generic.
Naming follows a consistent verb_noun pattern: list_* for summary listings, describe_* for detailed views, and get_pod_logs as a specific action. Even though get_events uses 'get' instead of 'list', the pattern is predictable and readable, so deviation is minor and not confusing.
With 11 tools, the set is well-scoped for a Kubernetes diagnostic server. Each tool addresses a common troubleshooting need (node health, pod issues, logs, ArgoCD sync, deployments) without excess or missing core operations.
The surface covers the main diagnostic workflows: listing and describing nodes, pods, deployments, and ArgoCD apps; retrieving events and logs; fetching any resource via describe_resource; and listing contexts/namespaces. For a read-only troubleshooting tool, this is comprehensive with no critical dead ends.