Skip to main content
Glama

agentctl

K8s-native agent/model workload operator with SAM mesh registration and MCP ops plane.

Replaces process-spawning orchestrators (vllm-orchestrator) with declarative K8s resources. Every operation is exposed as an MCP tool for LangGraph/LangChain workflow composition.

Architecture

LangGraph agent ──MCP──▶ agentctl ops plane (:8091)
                              │
                    ┌─────────┼─────────┐
                    ▼         ▼         ▼
              workload_*   fleet_*   sam_*
              tools        tools     tools
                    │         │         │
                    ▼         ▼         ▼
              K8s API    K8s API   SAM node
              (CRDs)    (pods)    (discovery)

Related MCP server: Hatchet MCP

Install

pip install agentctl
# or
uv pip install agentctl

Install (cluster)

kubectl apply -f https://github.com/moreWax/agentctl/releases/latest/download/install.yaml

This installs the CRD, the operator, and the MCP ops plane into agentctl-system. The ops plane is reachable in-cluster at http://agentctl-mcp.agentctl-system.svc.cluster.local:8091/mcp (stateless Streamable HTTP).

Usage

Deploy the example workloads

kubectl apply -f examples/qwen3.8-27b.yaml        # Qwen3.8-27B, TP2 on 2 GPUs
kubectl apply -f examples/qwen3.8-27b-fp8.yaml    # FP8 variant, single GPU
kubectl apply -f examples/litellm-gateway.yaml    # LiteLLM proxy gateway

From source (development)

uv venv && uv pip install -e ".[dev]"
kubectl apply -f crd/agentctl-crd.yaml
agentctl serve          # operator (kopf controllers)
agentctl mcp --port 8091  # MCP ops-plane server (stateless Streamable HTTP)

Deploy a model workload

kubectl apply -f examples/qwen3.8-27b.yaml        # Qwen3.8-27B, TP2 on 2 GPUs
kubectl apply -f examples/qwen3.8-27b-fp8.yaml    # FP8 variant, single GPU
kubectl apply -f examples/litellm-gateway.yaml     # LiteLLM proxy gateway

The LiteLLM gateway routes OpenAI-compatible requests to the vLLM backends. Point your agents at http://litellm-gateway:4000/v1 and use model names like qwen3.8-27b or qwen3.8-27b-fp8.

MCP tools (for LangGraph / agent composition)

The ops plane exposes 3 action-routed tools — a small surface that keeps agent context cheap and tool selection trivial, while all 20 capabilities stay available as actions. Everything is stateless, returns structured JSON, and gates destructive actions behind approved=true.

fleet(action=...) — cluster-wide views (read-only)

Action

Params

Description

status

Aggregate: workloads by phase, GPU totals

gpus

Per-node capacity / allocated / free + consumers

pods

namespace, limit

Per-pod: node, boot phase, role, sleeping, GPUs, restarts

footprint

gpu_count

Dry-run placement — call before boot

health

namespace

Endpoint readiness per service

sam_services

type

SAM mesh discovery

workload(action=...) — AgentWorkload CRUD

Action

Params

Description

list

namespace ("all")

Workloads with phase, GPUs, pin state

status

name

Full detail incl. per-pod boot phases

logs

name, tail

Recent container logs per pod

scale

name, replicas, force

0 parks; refuses pinned scale-down unless forced

delete

name, approved=true

Delete workload + pods

deploy

model_id, image, gpu_count, args, sam_enabled

Create (idempotent)

undeploy

model_id

Remove

model(action=...) — model lifecycle & actuation

Action

Params

Description

boot

model_id, image, gpu_count, args, wait_seconds

Deploy + optional wait → outcome: ready/failed/timeout

boot_status

model_id, tail_logs

Pod phases + vLLM markers (LoadingWeightsSizingKVCacheCapturingGraphsServing) + %

sleep

model_id

Suspend, free GPUs. FMA: seconds-level (fma-suspend); native: full stop

wake

model_id, wait_seconds

Resume; FMA wake takes seconds

kill_pod

pod_name, approved=true

Force-delete hung pod; Deployment recreates it

pin / unpin

model_id

Pin blocks scale-down/sleep unless forced

Agent boot-wait pattern (LangGraph loop):

model(action="boot", model_id="qwen38-27b", image="vllm/vllm-openai:v0.26.0", gpu_count=2)
while True:
    s = model(action="boot_status", model_id="qwen38-27b")
    if s["ready"]:  break          # proceed to inference
    if s["failed"]: ...            # inspect s["pods"], kill_pod + retry
    time.sleep(10)                 # poll again

Errors are structured: {"error": "unknown_action", "valid_actions": [...]}, {"error": "missing_params", "required": [...]}, or {"error": <cause>, "detail": ...}.

Development

uv venv && uv pip install -e ".[dev]"
pytest tests/ -q

License

MIT

Maintenance

ActivityMaintained
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables management of Mistral AI agents through MCP tools including creating, listing, searching, viewing details, and deleting agents. Integrates with Mistral API to provide agent management capabilities through natural language interactions.
    1
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables running durable, traceable AI agents via LangGraph through a universal MCP interface, integrating with Hatchet for orchestration, logging, and retries. Provides tools for knowledge management (ingestion, RAG) and Kubernetes operations (diagnosis, auto-fix).
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI-powered Kubernetes management using natural language, supporting kubectl, Helm, diagnostics, and port forwarding via MCP protocol.
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/moreWax/agentctl'

If you have feedback or need assistance with the MCP directory API, please join our Discord server