Skip to main content
Glama

Why this exists

Rancher is how real fleets run Kubernetes — and it speaks two APIs (the legacy Norman /v3 plane and the modern Steve /v1 plane), varies by version, and wraps every cluster behind its own proxy. Pointing a generic Kubernetes MCP server at it misses everything Rancher-specific; pointing an agent at raw kubectl gives up auditability, guardrails, and the management-plane view entirely.

MCP Rancher is built for that reality:

  • Capability-aware, not version-naive. It detects what each connected Rancher actually supports instead of assuming. One binary spans 2.6.5 → 2.9.3 with the same tool surface.

  • Multi-instance first. Lab, staging, prod — configure them all; mark prod read_only: true and every mutation is refused at the config layer, before any guard even has to fire.

  • Nothing is out of reach. Curated tools cover the common 95%; the generic engine reaches every resource either API plane exposes — even types nobody wrote a tool for yet.

Related MCP server: Rancher MCP Server

The tool surface

206 tools: 178 read-only · 28 writes · 5 destructive — counted from the registry itself, not by hand. docs/tool-manifest.json is generated from the live FastMCP registry (make tool-manifest) and a CI gate fails the build if it ever drifts from the code. Per-tool descriptions, safety annotations, and parameters all live there; the narrative registry with slice tracking is docs/tool-catalog.md.

Layer

What it does

Examples

Discovery & schema

Explore what any instance can do

rancher_server_version, rancher_norman_schema_list, rancher_capability_domain_list

Generic engine

CRUD + actions + links + watch on any resource, both planes

rancher_steve_resource_list, rancher_norman_resource_action_invoke, rancher_steve_resource_watch

Curated reads

Typed, shaped responses across ~25 domains

rancher_pods_list, rancher_deployments_list, rancher_longhorn_volumes_list, rancher_policy_reports_list

Curated writes

Guarded mutations

rancher_deployment_scale, rancher_deployment_restart, rancher_cron_job_suspend, rancher_node_cordon, rancher_secret_create

Operator rollups

One-call triage

rancher_cluster_health_check, rancher_find_failing_pods, rancher_find_stalled_rollouts, rancher_project_health_summary

Domains covered: clusters & nodes · projects & namespaces · workloads · pods & services · storage · networking · config & secrets (values masked) · certificates (keys masked) · RBAC · auth & identity · apps & catalogs · logging pipeline · Prometheus monitoring · policy reports · CIS compliance · backup operator · etcd backups · Longhorn · Fleet · provisioning · settings & features · alerts & notifiers.

All 206 stay exposed by default — every tool schema is deferred behind Claude Code's own search, so a small default would only help other hosts at the good host's expense. A constrained host (a small local model, a tight context budget) can opt into a smaller surface via RANCHER_TOOLSETS; see Toolset profiles under Configuration.

Quick start

Requirements

Install & run

# From source
git clone https://github.com/rex/mcp-rancher.git
cd mcp-rancher
make setup                 # deps, .env scaffold, pre-commit hooks
cp .env.example .env       # set RANCHER_URL + RANCHER_TOKEN
make dev                   # run the MCP server (stdio)

Once published to PyPI, it's one line: uvx rancher-mcp.

Claude Code

claude mcp add rancher \
  -e RANCHER_URL=https://rancher.example.com \
  -e RANCHER_TOKEN=token-xxxxx:yyyyyyyyy \
  -- uv run --directory /path/to/mcp-rancher rancher-mcp

Claude Desktop

{
  "mcpServers": {
    "rancher": {
      "command": "uv",
      "args": ["run", "--directory", "/path/to/mcp-rancher", "rancher-mcp"],
      "env": {
        "RANCHER_URL": "https://rancher.example.com",
        "RANCHER_TOKEN": "token-xxxxx:yyyyyyyyy"
      }
    }
  }
}

Multiple instances

RANCHER_INSTANCES_JSON='{
  "production": {"url": "https://rancher.prod.example.com",    "token": "token-a:xxx", "verify_ssl": true, "read_only": true},
  "lab":        {"url": "https://rancher.lab.example.com",     "token": "token-b:yyy", "verify_ssl": false, "read_only": false}
}'
RANCHER_DEFAULT_INSTANCE=production

Every tool takes an optional instance argument. Instances flagged read_only: true refuse all mutations at the settings layer.

Architecture

flowchart LR
    A[MCP client<br/>Claude Code · Claude Desktop · any] -- stdio --> S

    subgraph S[rancher-mcp]
        direction TB
        L1[Discovery & schema<br/>planes · schemas · capabilities]
        L2[Generic engine<br/>any resource · both planes<br/>CRUD · actions · links · watch]
        L3[Curated tools<br/>typed models · shaped output<br/>next-step hints]
        G[Safety layer<br/>read-only guard · confirmation phrases<br/>audit log · rate limit · masking]
        L1 --> L2 --> L3
        L3 --> G
        L2 --> G
    end

    G -- Norman /v3 --> R1[(Rancher<br/>instance A)]
    G -- Steve /v1 + k8s proxy --> R1
    G -- Norman + Steve --> R2[(Rancher<br/>instance B)]

Three layers, deliberately separate: discovery tells you what an instance can do, the generic engine can touch anything it exposes, and curated tools make the common paths typed, shaped, and self-describing (every response carries suggested_next_steps). Most curated tools are generated from YAML descriptors (catalog/curated_tools/) with a drift gate — the editorial decisions live in descriptors, not boilerplate.

Safety model

Built for the day an agent is pointed at the cluster that pays your salary:

Guard

Behavior

Read-only instances

read_only: true refuses every mutation for that instance, before tool logic runs

Destructive confirmation

Deletes require an explicit typed phrase (e.g. "delete steve namespace foo") — no phrase, no delete

Tool annotations

Every tool declares readOnlyHint / destructiveHint / idempotentHint, so clients can gate UX on them

Audit log

Every mutation emits a structured event="audit" record — tool, operation, plane, instance, resource, outcome. Argument names only; values never logged

Rate limiting

Token-bucket on writes (default 60/min) — a runaway loop can't machine-gun your API

Secret & key masking

Secret values and certificate private keys are structurally absent from curated responses (reveal is an explicit generic-tool opt-in)

Structured errors

Guard rejections return typed error_code envelopes agents can branch on — never raw strings

Compatibility

Primary target

Rancher 2.9.3 (production-validated)

Compatibility floor

Rancher 2.6.5 (kept green via capability detection)

API planes

Norman /v3 + Steve /v1 (+ per-cluster Kubernetes proxy)

Transport

stdio

Capability detection bridges version differences at runtime — no version-pinned builds, no "works on my Rancher." Both targets are exercised by the same test suite, and read paths have been validated live against both a 2.6.5 lab and a 2.9.3 production fleet (validation report).

Configuration

Variable

Default

Purpose

RANCHER_URL

Rancher server URL (single-instance mode)

RANCHER_TOKEN

API token (token-xxxxx:yyyyyyyyy)

RANCHER_VERIFY_SSL

true

TLS verification

RANCHER_INSTANCES_JSON

Multi-instance config (see above)

RANCHER_DEFAULT_INSTANCE

first defined

Instance used when a tool call names none

RANCHER_MCP_SERVER_NAME

rancher-mcp

Server identity announced to clients

RANCHER_MCP_SERVER_DESCRIPTION

built-in

Server description announced to clients

RANCHER_MCP_WRITE_RATE_LIMIT_PER_MIN

60

Write rate limit (0 disables)

RANCHER_TOOLSETS

all

Comma-separated toolset profile(s) exposed at startup — see below

RANCHER_TOOLS

Comma-separated tool names force-included on top of the selected profile(s)

RANCHER_EXCLUDE_TOOLS

Comma-separated tool names removed, applied last — always wins over RANCHER_TOOLS

Toolset profiles

The default is all: every tool stays exposed. Claude Code, the primary host, defers every tool schema behind its own search, so a small default would only help other hosts at the cost of making the good host worse — this is a deliberate choice, not an oversight.

RANCHER_TOOLSETS opts a constrained host (a small local model, or a context budget) into a smaller surface. Values are either a family name — one per src/rancher_mcp/tools/ module (storage, workloads, pods_services, rbac, …; see rancher_mcp.toolsets.FAMILY_REGISTRARS for the full list) — or the cross-family core profile: a ~32-tool triage/orientation set (rancher_find_*, the health/summary rollups, core list/get pairs, and the generic rancher_{steve,norman}_resource_{list,get} escape hatches). core cuts the tools/list payload from 206 tools / ~401 KB to 32 tools / ~78 KB (~80% smaller). Profiles compose: RANCHER_TOOLSETS=core,storage gives the triage set plus everything storage-related.

RANCHER_TOOLSETS=core                        # small triage surface
RANCHER_TOOLSETS=core,storage,workloads      # triage + two full families
RANCHER_TOOLS=rancher_secret_get             # add one extra tool on top
RANCHER_EXCLUDE_TOOLS=rancher_secret_create  # remove one, wins over the above

An unknown profile name fails loudly at startup rather than silently producing an empty or shrunken surface. Calling a real tool that exists but isn't in the active profile returns a structured TOOLSET_NOT_ENABLED error naming the tool, the toolset that would enable it, and the env var to set — never a bare "unknown tool", which would make a disabled tool indistinguishable from a typo.

Project status

Shipping and stable for read, triage, and guarded write operations. Honest ledger of what's beyond that:

  • Destructive workflows (node drain, etcd/backup restore, cert rotation, cluster upgrade/delete) are roadmap — deliberately staged after real-world read-path mileage. The generic engine + confirmation guard already covers these cases for operators who need them today.

  • The Alertmanager routes/silences surface needs an in-cluster API integration and is deferred.

  • The full per-version compatibility matrix (Track G) is in progress; the first live validation run covers the read matrix on both targets.

Work is tracked to the tool level: docs/tool-catalog.md (every tool has a row, every gap a slice ID) and ROADMAP.md.

Development

make help               # every target, documented
make validate           # codegen drift + manifest drift + architecture + lint + typecheck + tests
make tool-manifest      # regenerate docs/tool-manifest.json from the registry
make lab-up             # local Rancher 2.6.5 lab (kind + helm), fully scripted
make integration-current # isolated Rancher 2.14.3 end-to-end test run
make live-read-matrix   # read-only validation probes against configured instances
make mock-rancher       # fixture-backed mock Rancher for provider-config testing
  • Local lab — a self-contained Rancher 2.6.5 on kind with a simulated downstream cluster; repo-local kubeconfigs, never touches your machine state.

  • Current integration lab — Rancher 2.14.3 on separate Kind clusters and port 9443; run it serially with make integration-current to avoid overlapping Docker resource demand with the legacy lab.

  • Contract fixtures — sanitized captures from live Rancher committed under tests/fixtures/; respx pins the HTTP boundary in tests.

  • Codegen — curated tools are emitted from catalog/curated_tools/*.yml descriptors; make check-codegen fails on drift.

  • Gates — ruff, pyright strict, 624 tests with coverage floor, architecture line-limits, module-shape checks, secret scanning. All fail closed, all wired into pre-commit.

Stack: Python 3.12 · FastMCP · httpx · Pydantic v2 · structlog · uv

Security

See SECURITY.md for the threat model, token guidance, and how to report vulnerabilities.

License

MIT

A
license - permissive license
Not graded
quality - not tested
A
maintenance

Maintenance

UpdatingMaintainers
UpdatingResponse time
6dRelease cycle
8Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides MCP multi-cluster Kubernetes management and operations. It can be integrated as an SDK into your own project and includes nearly 50 built-in tools covering common DevOps and development scenarios. Supports both standard and CRD resources.
    149
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    An MCP server exposing scoped, read-only enterprise operations tools with fail-closed credential handling. It returns opaque approval IDs for mutations and requires a separate operator approval command to release one-time capabilities.
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Enables MCP-capable clients to inspect and manage Keycloak realms, users, clients, roles, and groups with layered security modes, realm allowlisting, protected realms, delete gating, dry-run, and audit logging.
    8
    345
    1
    MIT

View all related MCP servers

Related MCP Connectors

  • Managed Keycloak from any MCP client: clusters, realms, apps, SSO, users, domains, audit events.

  • Access Kernel's cloud-based browsers and app actions via MCP (remote HTTP + OAuth).

  • Remote MCP for A2A caller identity, scope policy, verdict receipts, and audit history.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/rex/mcp-rancher'

If you have feedback or need assistance with the MCP directory API, please join our Discord server