clawops
This MCP server lets AI agents safely deploy, operate, monitor, and troubleshoot self-hosted OpenClaw stacks on AWS, GCP, Azure, or local VMs through typed MCP tools.
Inspect infrastructure: get stack status, run diagnostics (doctor), monitor gateway health, list stacks, and poll long-running tasks
Read logs & config: tail gateway logs, fetch config values, validate deployed config, and list OpenClaw agents
Provision and deploy: initialize stacks, generate deploy plans, apply plans, and run
upto provision or update a stackModify configuration: set or unset remote OpenClaw config keys and restart the gateway daemon
Destroy resources: tear down entire stacks with explicit confirmation and dry-run support
Harden and secure: apply SSH/UFW/fail2ban hardening, join Tailscale networks, and close public access via private-only plans
Run end-to-end workflows: deploy OpenClaw from intent to healthy gateway, or diagnose/recover an unhealthy stack
Operate safely: read-only mode, destructive-tool elicitation, audit logs, and no cloud credential storage
Allows selecting Amazon Bedrock as the LLM provider for the OpenClaw instance, supporting models available through Bedrock.
Adds Discord as a chat integration, allowing the OpenClaw gateway to interact via Discord bots.
Allows selecting Ollama as a local LLM provider for the OpenClaw instance.
Allows selecting OpenAI's models as the LLM provider for the OpenClaw instance.
Adds Slack as a chat integration, allowing the OpenClaw gateway to interact via Slack apps.
Adds Telegram as a chat integration, allowing the OpenClaw gateway to interact via Telegram bots.
Adds WhatsApp as a chat integration, allowing the OpenClaw gateway to interact via WhatsApp bots.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@clawopsdeploy a new OpenClaw stack on AWS"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
clawops
MCP-native infrastructure ops for OpenClaw, with read-only mode, destructive-action confirmation, and audit logs built in.
clawops is a CLI and MCP server for deploying and operating self-hosted OpenClaw instances. Provision on AWS, GCP, Azure, or any Linux VM, then manage day-to-day operations from the terminal, or let Claude Code and Cursor drive them through typed MCP tools with explicit safety controls.
What's new: 2.0.1 and 2.0 release notes below, and the full history in CHANGELOG.md.
Who this is for
OpenClaw users who want the simplest path to self-hosting across cloud or local VMs, with reliable deploy, status checks, logs, backups, and upgrades in a single CLI.
Claude Code / Cursor / MCP users looking for a real-world reference implementation of safe infrastructure operations through MCP. Typed tool schemas, read-only mode, destructive-action confirmation, and audit logs.
Self-hosted AI and local-first developers who want to run their own AI assistant without committing to Kubernetes, a managed SaaS platform, or a single cloud provider.
Related MCP server: Infraveil MCP server
What clawops does
Provisions and tears down OpenClaw infrastructure on AWS, GCP, Azure, and local VMs using the Pulumi Automation API. You do not install Pulumi; clawops installs the CLI it needs into
~/.clawops/.pulumi-clion first use.Manages day-to-day operations: status, logs, SSH, tunnels, config, agents, gateway, backups.
Exposes every operation as a typed MCP tool so AI agents can drive ops safely.
Enforces a plan → review → apply discipline for cloud deployments.
Emits JSON output everywhere (
--json) for scripting and automation.Never stores cloud credentials. Reads them from your environment's existing CLI profiles.
What clawops does not do
No high availability or clustering. Optimized for single-node deployments.
No Kubernetes. It deploys to VMs, not container orchestration platforms.
No OpenClaw skill/agent authoring. clawops manages infrastructure; what runs on it is up to you and OpenClaw.
No TLS or domain automation (yet). Bring your own reverse proxy or see
docs/limitations.mdfor the manual path.No credential storage. Cloud credentials must be configured in your environment before using clawops. They are never written to
~/.clawops/config.json.No native Windows. WSL2 is fully supported; see
docs/support-matrix.md.
Quick Start
npm install -g @clawops/cli
clawops setupclawops setup is an interactive wizard that gets OpenClaw running in about 2 minutes. It
handles everything in one flow, no config files to write by hand, no commands to memorize.
What the wizard does
Step 1. Choose a deployment target
Pick an existing server you can SSH into (Linux or macOS), or a new cloud VM on AWS, GCP, or Azure. Cloud deployments walk you through authenticating with the provider CLI if you aren't already signed in.
Step 2. Pick an LLM provider
Choose from Anthropic, OpenAI, Amazon Bedrock, Ollama, or others. The wizard prompts for your
API key and saves it locally (in ~/.clawops/secrets/, chmod 600), it is never sent anywhere
except to OpenClaw on the target host when the config is applied.
Step 3. Add chat integrations (optional)
Select any combination of Discord, Telegram, Slack, WhatsApp, or Teams. The wizard collects each integration's bot token the same way as the API key. Paste it in, reference an env var, or point to a file.
Step 4. Wire your AI editor
Select which AI apps should have access to clawops. Claude Desktop, Claude Code, Cursor, Windsurf, VS Code, and Zed are all supported. The wizard writes an MCP server entry into each app's config file using the absolute binary path so the app can launch it independently.
Step 5. Deploy
The wizard bootstraps OpenClaw on the target host over SSH (installs Docker, pulls the image, starts the container), applies your LLM and integration config, generates a gateway auth token, and prints a direct dashboard URL:
✔ All done! OpenClaw is running.
ℹ Open dashboard: http://192.168.1.50:18789?token=<your-token>
ℹ Token saved to ~/.clawops/secrets/GATEWAY_TOKEN_my-stackPrerequisites: Node.js ≥ 22, an SSH key, and either an SSH-reachable Linux/macOS host or a
cloud account with CLI credentials configured (aws configure, gcloud auth login, or az login).
You do not need Pulumi. The first cloud deployment installs the CLI clawops drives into
~/.clawops/.pulumi-cli and says so while it does.
For a full narrated walkthrough with example output, see docs/demo-script.md.
Manual setup, existing server
If you prefer step-by-step control, or are adding clawops to an already-running deployment:
npm install -g @clawops/cli
clawops doctor # verify environment
clawops init --provider local --host 192.168.1.50 --user ubuntu --key-path ~/.ssh/id_ed25519
clawops up # installs Docker + OpenClaw over SSH
clawops statusSee docs/examples/local-vm.md for SSH prerequisites, firewall
setup, and troubleshooting.
Manual setup, cloud (AWS)
npm install -g @clawops/cli
# Requires AWS credentials in your environment (AWS_PROFILE or ~/.aws/credentials)
clawops init --provider aws
# Edit ~/.clawops/config.json: set stateUrl to your S3 bucket
clawops plan --provider aws --stack default --ssh-cidr auto --out /tmp/plan.json
clawops apply /tmp/plan.json--ssh-cidr auto allows SSH from this machine's public IP, resolved while the plan is
generated and written into it. Without it the plan allows no ingress at all and nothing,
including clawops, will be able to connect.
Connect an AI editor
The setup wizard handles this automatically (Step 4). To wire or re-wire editors at any time:
clawops mcp installThis opens the same interactive checkbox used in the wizard, select Claude Desktop, Claude Code, Cursor, Windsurf, VS Code, or Zed and clawops writes the MCP entry into each app's config using the correct absolute binary path.
To add the entry manually instead, paste this into your editor's MCP config:
{
"mcpServers": {
"clawops": {
"command": "npx",
"args": ["-y", "@clawops/cli", "mcp", "serve", "--read-only"]
}
}
}That form needs nothing on $PATH and is what a directory or an installer will copy. If you
would rather point at the binary you already have, use its absolute path — the output of
which clawops — with the same arguments:
{
"mcpServers": {
"clawops": {
"command": "/path/to/clawops",
"args": ["mcp", "serve", "--read-only"]
}
}
}Either way, pass the arguments. mcp serve is what speaks the protocol, and an explicit config
is one that still reads clearly a year later. clawops does not strand a client that omits them:
run with no command at all and a pipe on stdin — how every MCP client starts a server — and it
starts mcp serve, saying so on stderr. Typed at a terminal, clawops still prints help.
Config file locations:
App | Path |
Claude Desktop (macOS) |
|
Claude Desktop (Linux) |
|
Claude Code |
|
Cursor |
|
Windsurf |
|
VS Code (macOS) |
|
VS Code (Linux) |
|
Zed |
|
Start with --read-only. It enables status, logs, config reads, and diagnostics while
blocking mutations. Remove it only after reviewing
docs/security/mcp-safety.md.
Destructive tools (clawops_destroy, clawops_up, clawops_config_set, etc.) require explicit
confirmation before executing, they will never run silently.
For HTTP mode setup see docs/mcp/.
Day-to-day operations
clawops status # Stack outputs: IP, gateway URL, SSH info
clawops logs -f # Tail OpenClaw logs over SSH
clawops ssh # Interactive SSH session
clawops ssh --command "docker ps"
clawops config get maxAgents
clawops config set maxAgents 8
clawops tunnel # Port-forward gateway UI to localhost
clawops destroy --yes # Destroy cloud-provider stack
clawops down --yes # Destroy local-provider stackCommands
Command | Description |
| First-run wizard: guided LLM, integrations, and deploy-plan generation |
| Register a stack in |
| Provision or update stack ( |
| Destroy local-provider stack (requires |
| Destroy cloud-provider stack with confirmation prompt ( |
| Show stack outputs: IP, gateway URL, region, provisioned time |
| Generate a deploy-plan JSON artifact (dry-run safe). |
| Apply a previously reviewed plan file ( |
| Interactive SSH session or run a remote command |
| Stream OpenClaw logs ( |
| Local port-forward to gateway UI over SSH |
| Get/set remote OpenClaw config values ( |
| List OpenClaw agents, or stream one agent's logs |
| Restart the OpenClaw gateway service |
| Create and restore OpenClaw state backups ( |
| List named stacks and their state |
| Check the local machine; with |
| Manage secrets: |
| Live dashboard: gateway health, container stats, log tail, stack picker |
| Start the embedded MCP server (stdio, or HTTP with |
| Interactively wire clawops into AI editors |
| Wire the gateway's AI as an MCP client of clawops (verifies the connection before saving) |
| List all commands and global flags |
| Apply security hardening to a deployed stack (SSH, UFW, fail2ban, unattended-upgrades, Docker socket; AWS: SG audit, SSM check, Flow Logs, GuardDuty). |
| Open a pre-filled GitHub issue with system context from |
Full flag reference: clawops <command> --help
Plan → Apply workflow
For non-local providers, clawops enforces a review-before-apply discipline:
# 1. Generate a plan: runs `pulumi preview` internally, produces JSON
# --ssh-cidr decides who may connect. `auto` means this machine; omit it and nobody can.
clawops plan --provider aws --region us-east-1 --ssh-cidr auto --out /tmp/plan.json
# 2. Review plan.json: the `diff` field shows projected changes at plan-generation time
cat /tmp/plan.json | jq .diff
# 3. Apply: reads and validates the plan file, then runs `pulumi up`
clawops apply /tmp/plan.json
# Without --yes, apply prompts: "Continue? (y/N)"
clawops apply /tmp/plan.json --yes # skip prompt in automationThe plan JSON conforms to spec/deploy-plan.schema.json (AJV-validated) and captures reviewed
intent: provider, region, instance type, CIDR ranges, and OpenClaw version. apply re-runs
pulumi up using those parameters against the current live state, it does not replay a locked
execution artifact. Review and apply in the same session to minimize drift risk.
See docs/plan-apply.md for full semantics, drift guidance, and the safe CI pattern.
MCP server
clawops ships an embedded MCP server. Claude Code, Cursor, and any MCP-compatible agent can drive deployments without leaving the chat interface.
Wire your editor
clawops mcp install # interactive checkbox: writes config for selected appsThe wizard resolves the absolute binary path automatically so app launchers can find clawops
without inheriting your shell's PATH. See Connect an AI editor above
for manual config paths.
Wire the gateway AI
The OpenClaw gateway runs its own AI agent. Once wired, that agent can call clawops directly instead of guessing at infrastructure state:
clawops mcp wire --stack prod # write MCP client entry into gateway config + restartRequires OpenClaw ≥ 2026.4 on the gateway. The clawops setup wizard offers this step
automatically after a successful deploy.
Stdio mode (Claude Code / Cursor / VS Code)
Start the server manually or confirm your config is correct:
clawops mcp serve --read-only # safe for first evaluation
clawops mcp serve # full mode: enables provisioning, config write, ssh execHTTP mode (remote / multi-client)
clawops mcp serve --http 3333 --bind 127.0.0.1
# MCP HTTP server listening on 127.0.0.1:3333Do not bind to a non-loopback address without additional authentication controls in front of it.
Available tools
Tool | Toolset | Description |
| cli | Show stack outputs (what is deployed, not whether it works) |
| cli | Run diagnostics: local prerequisites, and with a stack, remote health |
| cli | Tail OpenClaw logs |
| cli | Sample gateway and host metrics |
| admin | List all stacks and their state |
| cli | Read a remote config value |
| cli | List running agents |
| cli | Provision or update a stack |
| cli | Destroy a stack (elicits confirmation) |
| cli | Apply a plan file |
| cli | Generate a deploy plan |
| cli | Write a remote config value |
| cli | Remove a remote config key |
| cli | Validate the deployed config against the OpenClaw schema |
| cli | Restart the gateway (elicits confirmation) |
| cli | Apply hardening modules; join or leave a tailnet (elicits confirmation) |
| cli | Register a stack and write |
| workflow | End-to-end deploy: plan → confirm → apply → status |
| workflow | Diagnostic workflow for an unhealthy stack |
| cli | Poll a long-running task |
Tools in the read toolset are also available in --read-only mode; the table's Toolset column shows the primary toolset. All other toolsets require full mode.
Destructive tools require explicit confirmation (elicitation) unless yes: true is passed.
See docs/security/tool-risk-matrix.md for the full risk
classification of every tool.
Configuration
Config lives at ~/.clawops/config.json (override with $CLAWOPS_HOME).
{
"version": 1,
"defaults": {
"provider": "aws",
"stack": "default"
},
"stacks": {
"default": {
"provider": "aws",
"region": "us-east-1",
"stateUrl": "s3://my-clawops-state"
}
},
"ssh": {
"keyPath": "~/.clawops/id_ed25519",
"knownHostsPath": "~/.clawops/known_hosts"
}
}Cloud credentials are never stored in config. Clawops reads them from the environment:
Provider | Credential source |
AWS |
|
GCP |
|
Azure |
|
Local | SSH host + key configured in |
Known limitations
See docs/limitations.md for the full list. Key points:
Single-node deployments only, not a high-availability or clustering platform.
clawops applyis not an immutable plan execution. Seedocs/plan-apply.md.No TLS/domain automation in the current release.
MCP tools execute privileged operations, use
--read-onlyfor first evaluation.
Architecture
clawops
├── src/cli/ citty-based commands (one file per verb)
├── src/config/ ~/.clawops/config.json management
├── src/providers/ Cloud adapters (AWS, GCP, Azure, local)
│ ├── aws/ Pulumi inline program + ProviderAdapter
│ ├── gcp/
│ ├── azure/
│ └── local/ SSH bootstrap (no Pulumi)
├── src/pulumi/ Pulumi Automation API wrapper + output helpers
├── src/transport/ SSH client (ssh2) + connection pool + tunnels
├── src/mcp/ MCP server, tool handlers, progress tracking
├── src/plan/ Maker plan generation, AJV validation, apply
├── src/output/ ASCII table, spinner, JSON, human-readable output
├── src/errors/ Typed error hierarchy with exit codes
└── spec/ Machine-readable ground truth (JSON Schema, YAML)Key design decisions:
Pulumi Automation API: the user installs no Pulumi. Clawops installs the CLI the API drives into
~/.clawops/.pulumi-cli, pinned to the bundled SDK, without editing$PATH(ADR 0010); Pulumi home is sandboxed to~/.clawops/.pulumi; stack programs are inline TypeScript closuresState in cloud blob storage: GCS (
gs://), S3 (s3://), Azure Blob, no local state files, nopulumi.yamlSSH via
ssh2: never shells out to/usr/bin/ssh; TOFU host verification against~/.clawops/known_hosts; connection pool with 5-min idle TTLPlan → apply discipline: every non-local deployment goes through
generatePlan()→ review →applyPlan(); destructive changes always require human review of the plan JSONMCP-first: every CLI operation has a typed MCP tool; schemas generated from
spec/mcp-tools.yaml; all destructive tools use elicitation
See docs/architecture.md for a full narrative, and docs/decisions/ for ADRs.
Cloud provider stacks
Each cloud provider is an inline Pulumi program that creates the resources below. All three share the same outputs (publicIp, gatewayUrl, sshHost, sshPort, sshUser) consumed by the SSH and config-overlay layers.
AWS
flowchart LR
subgraph NET["Networking"]
VPC["VPC (10.0.0.0/16)"]
IGW[Internet Gateway]
SUBNET["Subnet (10.0.1.0/24)"]
RT[Route Table]
SG["Security Group (ports 22, 18789)"]
end
subgraph IAM["IAM"]
ROLE[IAM Role]
SSM[SSM Policy Attachment]
BED["Bedrock Policy Attachment (optional)"]
IP[Instance Profile]
end
subgraph COMPUTE["Compute"]
KP[EC2 Key Pair]
EC2["EC2 Instance (Ubuntu 22.04, IMDSv2)"]
EIP[Elastic IP]
endGCP
flowchart LR
subgraph NET["Networking"]
NW[VPC Network]
SN["Subnetwork (10.0.0.0/24)"]
FW1["Firewall: SSH port 22 (conditional)"]
FW2["Firewall: Gateway port 18789 (conditional)"]
ADDR[Static External IP]
end
subgraph COMPUTE["Compute"]
VM["Compute Instance (Debian 12, 20 GB)"]
endAzure
flowchart LR
RG[Resource Group]
subgraph NET["Networking"]
VNET["Virtual Network (10.0.0.0/16)"]
SUBNET["Subnet (10.0.1.0/24)"]
NSG["Network Security Group (ports 22, 18789)"]
PIP["Public IP Address (Static)"]
NIC[Network Interface]
end
subgraph COMPUTE["Compute"]
VM["VM (Ubuntu 22.04, managed identity)"]
end
subgraph KV["Key Vault (optional)"]
VAULT["Key Vault (RBAC, name max 24 chars)"]
RA["Role Assignment (Secrets User)"]
SECRET["Secret: gateway-token"]
endDevelopment
Setup
git clone https://github.com/dfridkin/clawops.git
cd clawops
# Node 22+ required; use nvm: nvm use
pnpm install
pnpm dev doctor # verify toolchainScripts
pnpm dev # run CLI from src/ via tsx
pnpm build # tsup → dist/
pnpm test # vitest (1977 tests, ~13s)
pnpm test:changed # vitest --changed (fast edit loop)
pnpm test:integration # Docker-based SSH integration tests
pnpm test:e2e:local # local provider bootstrap for real, in a systemd container
pnpm typecheck # tsc --noEmit
pnpm lint # eslint src/ tests/ scripts/ (--max-warnings=0)
pnpm gen:schemas # regenerate src/providers/types.ts + src/mcp/tools/_generated.ts
pnpm gen:schemas --check # CI guard: committed generated files match spec
pnpm graph # local coupling report (--base <ref> for this branch's delta)
pnpm verify:pack # install the packed tarball elsewhere and run it (CI gate)
pnpm sync:server-json # write package.json's version into server.json
pnpm changeset # record a release note before mergingProject layout
Path | Purpose |
| Machine-readable ground truth: JSON Schema, YAML. Treat as source of truth. |
| Full technical specification (milestones, rules, schemas) |
| 25 normative rules (R1–R25) referenced throughout the codebase |
| Narrative system overview |
| Plan/apply semantics, drift guidance, CI pattern |
| CI integration guide: OIDC, env vars, plan → apply in CI |
| MCP safety model, tool risk matrix, redaction, audit logs |
| Per-provider capability matrix |
| Architecture Decision Records |
| Invokable procedures: |
| Path-scoped lint rules loaded by Claude Code |
Code generation
Two files are generated from spec/ and must not be hand-edited:
src/providers/types.ts.ProviderAdapterinterface fromspec/providers.schema.jsonsrc/mcp/tools/_generated.ts. Zod schemas and type exports fromspec/mcp-tools.yaml
Run pnpm gen:schemas after modifying either spec file. CI enforces this with --check.
Adding a provider
Use the /add-provider skill in Claude Code, or follow src/providers/CLAUDE.md. Every adapter must satisfy ProviderAdapter in src/providers/types.ts. Do not relax the schema to fit the adapter.
Adding an MCP tool
Use the /mcp-tool skill. The skill adds the tool to spec/mcp-tools.yaml, runs pnpm gen:schemas, creates the handler in src/mcp/tools/<toolset>/<name>.ts, and wires it into the registry. All four annotation hints (readOnlyHint, destructiveHint, idempotentHint, openWorldHint) are required on every tool.
Conventional commits
feat(scope): description
fix(scope): description
docs / refactor / chore / test / perf / ciUse pnpm changeset to record a release note before merging a feat or fix.
What's new in 2.1
Private networking, hardening on every cloud, and two fixes to commands that could not start. 2.1.1 follows with the fixes below it.
Reach a stack over your tailnet
clawops harden --tailscaleinstalls Tailscale on a stack, joins it to your tailnet asclawops-<stack>, and reports the address it was given.The same command then moves clawops onto that address, but only after opening an SSH session to it — against host keys pinned over the public connection it already trusts.
The Tailscale auth key comes from
clawops secret set TAILSCALE_AUTH_KEY, and reaches the host over the SSH data channel. It never appears in a command line, a process list or a log.clawops plan --private-only→clawops applycloses public SSH and gateway access on a stack reached over its tailnet. Both refuse unless that address answers SSH at that moment (ADR 0013).clawops harden --tailscale-reverttakes a host off the tailnet and returns clawops to its public address. On a private-only stack it refuses, and prints the commands that reopen SSH.clawops destroyforgets the host keys for both addresses of a stack on its tailnet, instead of leaving the public one pinned for an instance that no longer exists.
Hardening covers all three clouds
Azure: NSG audit, disk encryption, Defender for Cloud and JIT VM access, all check-only.
GCP: VPC firewall audit, Shielded VM and OS Login, all check-only.
GCP instances boot with Secure Boot on. Existing stacks get it as an update that keeps the boot disk and all OpenClaw state.
A check that could not run reports as skipped, naming what was missing, rather than as a pass.
Plans say what they will disturb
clawops plancounts and lists resources that would be replaced. It used to summarise a preview that would destroy the instance and its boot disk as "0 to create, 0 to update".A plan that changes a live deployment warns before you apply it: a replacement names what goes with it and points at
clawops backup create; an update says the gateway goes down.
Fixes
clawops mcp servecould not start at all when installed from npm — it died on import before emitting any protocol, so every MCP client got nothing.pnpm verify:packnow speaks MCP to the packed tarball, so this class of failure cannot ship again.server.json, the MCP registry manifest, is versioned with the package rather than rewritten at publish time. The committed file had read1.7.3against a published2.0.2.The published package carries its license, keywords and issue tracker, so it is findable on npm and its listing is complete.
2.1.1
clawopsstarted with no command and a pipe on stdin serves MCP instead of printing help. That is how every MCP client starts a server, and how directories that infer a run command start one; several were getting the help text and reporting the server as broken. Typed at a terminal,clawopsstill prints help, and so doesclawops | less.The MCP config example is
npx -y @clawops/cli mcp serve, which runs as written. It used to say/path/to/clawops, which nothing could run and no directory could copy.The description is inside the MCP registry's 100-character limit. The one 2.1.0 shipped was 110, and the registry refused it with a 422 after npm had already published — so 2.1.0 reached npm and not the registry.
clawops hardensays a stack is not deployed, or that its state could not be read, instead of passing alongcode: -2and a subprocess dump.The MCP registry backfill registers the version that was released rather than the one the release tooling is preparing, so an entry that falls behind can actually be repaired.
The Docker image builds. It had never been built, and did not:
npm pack --pack-destinationdoes not create its destination.pnpm verify:dockerbuilds it and speaks MCP to the running container, both through its entrypoint and as a bare binary, in CI.
What's new in 2.0.1
A patch release, and a large one: in 2.0.0 no cloud deploy succeeded by any path. Every item
below is a fix or an addition in 2.0.1. The reasoning behind each one is in its commit message,
and the decisions that came out of them are in docs/decisions/.
Deploying to a cloud
clawops plan→clawops applyprovisions a cloud stack and deploys OpenClaw onto it.clawops updeploys to AWS, GCP and Azure, running the same path asplan→apply.clawops installs the Pulumi CLI it needs into
~/.clawops/.pulumi-cli, or uses a compatible one already on$PATH(ADR 0010).clawops creates and stores the passphrase its state backend requires (ADR 0011).
clawops plantakes--ssh-cidr,--gateway-cidrand--publish-gateway, andapplypasses them to the cloud firewall.autoresolves this machine's address.clawops planstops, and names the cause, when it cannot open the state backend.--instance-typetakes a clawops alias (micro–gpu) or a machine type your cloud names itself, and the plan records the concrete type.Deploys pin the account they were planned against:
gcp:projecton GCP,azure-native:subscriptionIdon Azure.
Checking the account before you spend
clawops doctor --provider <cloud>checks one cloud's credentials and account setup, with or without a stack.--instance-typepoints the size check at the size you are deploying.AWS. The account the credentials resolve to, the state bucket, and whether the instance type is offered in the region.
GCP. The project, the APIs a deploy needs, and the state bucket.
Azure. The subscription, the resource providers, the VM size, and the azblob credentials Pulumi authenticates with.
clawops setupruns the same checks and offers to fix what it safely can, enabling an API, creating a state bucket with versioning on and public access blocked, naming the change before making it.A check clawops could not perform reports as a warning naming the error, rather than as a pass or a failure.
Azure accepts your
az login; a service principal is no longer required.
Naming, config and setup
clawops names the state backend after the account it is deploying into, instead of asking you for a name or writing a placeholder (ADR 0012).
A name you type instead is checked against the rules of the cloud that has to accept it.
clawops initkeeps the stacks already in your config.clawops initgenerates an SSH key that clawops can read. If you raninitbefore this release,clawops doctorwill tell you whether yours is usable.gcloud config set projectis honoured.The setup wizard writes model configuration that OpenClaw accepts, and installs the plugin your chosen provider needs.
Amazon Bedrock works: the right transport, and an inference profile resolved against your deployment region and recorded in the plan. Needs
bedrock:ListInferenceProfiles.
While a deploy is running
applywaits for SSH, then waits for the gateway to answer, before reporting success.applyreports progress as it goes instead of going quiet for minutes.A deploy that times out prints what the host was doing, from its bootstrap log.
A host still installing Docker is treated as still booting rather than as a failed deploy.
Day-two commands
clawops logsreads from the gateway on AWS.doctor --stack,ssh,logs,gateway,configandagentswork against a freshly deployed stack.clawops doctorvalidates cloud credentials.clawops tells a refused Docker socket from a missing container, and says which it found.
clawops destroyforgets the instance's host key, so redeploying onto an address the cloud has recycled no longer fails verification.
Documentation
The GCP guide names the credential source clawops actually reads, and describes 2.0 firewall behaviour.
The smoke-test plan covers 2.0, and
pnpm test:cloud aws|gcp|azureruns it against a real deployment and destroys it afterwards.
What's new in 2.0
clawops 2.x targets OpenClaw >= 2026.9.2. The 1.x line continues for OpenClaw
<= 2026.7.1-2 under the legacy dist-tag until 2027-03-31:
npm install -g @clawops/cli # 2.x
npm install -g @clawops/cli@legacy # 1.x maintenancePin the tag in CI. latest moves to 2.x, so an unpinned pipeline will change lines.
CHANGELOG.md carries the full history; this section covers what changed
about how clawops behaves.
Your deployment keeps its state
OpenClaw 2.0 stores sessions, transcripts and credentials in SQLite. clawops mounted no
state at all, so every restart destroyed them, and a restart is what gateway restart,
gateway update and config set all do.
One host directory (/var/lib/clawops/openclaw) is now bind-mounted at OpenClaw's own
default location, holding the config, the database and any provider plugins. Existing
deployments migrate on the next up/apply.
clawops up / clawops apply
flowchart TD
A["clawops plan"] --> B{"config valid<br/>against OpenClaw schema?"}
B -- no --> B1["refuse: plan is still<br/>a file you can edit"]
B -- yes --> C["clawops apply"]
C --> D{"OpenClaw version<br/>in supported range?"}
D -- no --> D1["refuse: names<br/>@clawops/cli@legacy"]
D -- yes --> E["provision host"]
E --> F["state dir, owned 1000:1000<br/>migrate any pre-2.0 config"]
F --> G["write config<br/>validated before writing"]
G --> H["install provider plugins<br/>while egress exists"]
H --> I["start gateway"]
I --> J{"/startupz says started?"}
J -- no --> J1["fail with the reason"]
J -- yes --> K{"configured providers<br/>all loaded?"}
K -- no --> K1["warn: healthy gateway,<br/>missing model backend"]
K -- yes --> L["done"]Three of those steps are new, and each exists because the old flow could report success
while something was wrong: the config was never validated before being written, provider
plugins were left to be fetched at boot (or silently missing on a deny-all host), and
"started" was inferred from docker run exiting 0.
clawops gateway update
Previously: pull, run, report success. docker run exiting 0 means the container was
created, and the container it replaced is already gone.
flowchart TD
A["clawops gateway update X"] --> B{"X in supported range?"}
B -- no --> B1["refuse before pulling"]
B -- yes --> C["docker pull X"]
C --> D["snapshot state database"]
D -- cannot snapshot --> D1["refuse: no rollback point"]
D --> E{"target release understands<br/>this schema?"}
E -- no --> E1["refuse: downgrade across<br/>a schema boundary"]
E -- yes --> F["swap container"]
F --> G{"/startupz says started?"}
G -- yes --> H["done"]
G -- no --> I["one-shot doctor --fix<br/>in a throwaway container"]
I --> J["re-run, re-gate"]
J -- started --> K["done: reported as repaired"]
J -- still not --> L["roll back to previous image"]
L -- started --> M["rolled back, reason reported"]
L -- still not --> N["failed: snapshot path named"]The snapshot is not only a rollback point: database preflight refuses a live database
because the schema version sits in the WAL until checkpointed, so the consolidated snapshot
is what makes the compatibility check possible at all.
clawops gateway restart
A restart changes neither the deployed version nor who can reach the gateway. Both are read back from the running container rather than guessed:
flowchart LR
A["gateway restart"] --> B["read current image"]
B -- no container --> B1["refuse: nothing to reuse.<br/>latest and stable point at 2.0"]
B --> C["read current publish scope"]
C --> D["recreate with the same<br/>version and reachability"]
D --> E{"/startupz says started?"}
E -- no --> E1["fail with the reason"]
E -- yes --> F["done"]Migrating an existing 1.x deployment
flowchart TD
A["clawops migrate"] --> B{"1.x container running?"}
B -- no --> B1["nothing to rescue: state was<br/>already lost to an earlier restart"]
B -- yes --> C["verified backup, inside the running container"]
C -- "backup fails" --> C1["refused: nothing touched"]
C --> D["extract state from the RUNNING container"]
D --> E["chown 1000:1000"]
E --> F["stop and remove 1.x"]
F --> G["synthesise a valid 2.0 config"]
G --> H["start 2.0 with the state directory"]
H --> I{"/startupz started?"}
I -- "no: schema still migrating" --> J["restart once"]
J --> K{"started?"}
K -- no --> K1["failed: points at the backup"]
K --> L["report"]
I -- yes --> L
L --> M["what carried over,<br/>device identity, config to review"]Two things about that shape are not obvious, and both came from running a real migration:
State is extracted from the running container. All 1.x state lived inside it, clawops mounted none, so stopping first destroys what the migration came to save.
The config is synthesised, not carried forward. 1.x never had one that applied; the file clawops mounted was read by nothing. Your old settings are reported as intent to review, never applied blindly. Their channel blocks would not validate against 2.0 anyway.
The gateway also needs two starts: the first performs the state-schema migration and reports
it as pending. migrate waits for the second rather than declaring success early.
If you ran gateway restart, gateway update or config set on a clawops before 2.0, your
state is already gone, nothing was mounted to survive the container replacement. migrate
says so plainly rather than pretending to rescue it.
clawops backup restore works again, and never in place
v1.7.5 made restore fail with an explanation, because the OpenClaw it supported had no restore subcommand to call. 2.0 does, and clawops delegates to it:
flowchart TD
A["clawops backup restore --file X"] --> B["upload archive to the host"]
B --> C["openclaw backup restore --target <staging>"]
C -- "target not empty" --> C1["refused by OpenClaw"]
C --> D["archive verified, expanded<br/>into a fresh directory"]
D --> E["warnings printed verbatim<br/>time travel, channel relink,<br/>approvals, plugins"]
E --> F["nothing activated"]
F --> G["you stop the gateway, swap the<br/>state dir, restart, re-apply"]clawops does not extract archives itself and does not restore in place. The final step is
manual on purpose, and re-applying matters: the archive does not carry plugin
node_modules, so a restored deployment starts without its model providers, looking
healthy while doing it.
The archive is a credential. It carries the state database, mcp_oauth_stores,
secret_store_entries, worker_environment_credentials, device_auth_tokens, unencrypted.
clawops now writes it 0600 locally; it previously used the default 0644.
Model providers that need a plugin are installed for you
OpenClaw 2.0 made model providers install-gated plugins. Twenty-four ship in the image,
anthropic, openai, google, ollama, openrouter among them, but not all of them.
Configuring one that is not bundled, without installing it, produces a gateway that starts,
reports healthy, and has no model backend.
clawops installs what your config needs, pinned to an exact version, during apply:
Resolving clawhub:@openclaw/deepseek-provider@2026.9.2…
Downloading plugin @openclaw/deepseek-provider@2026.9.2 from ClawHub…
Installed plugin: deepseekThis adds an outbound dependency the 1.x line did not have: clawhub.ai. It is needed
while apply is running, not at boot. Deliberately, so a failure reaches the person running
the command rather than a locked-down host at 3am. Blocked, it looks like this:
fetch failed | getaddrinfo EAI_AGAIN clawhub.ai | EAI_AGAINclawops checks the installed provider IDs afterwards and will not call the deploy finished while a configured provider is missing. Required outbound access lists every destination and when it is needed.
Chat channels are installed for you too
Every channel in OpenClaw 2.0 is an install-gated plugin. clawops apply installs the ones
your config names, during the deploy while egress exists, and then asks the gateway whether
they are really installed:
[clawops] warning: the gateway is running, but these configured channels are not installed:
discord. They will never connect.It has to ask. openclaw channels add. The obvious command, returns success even when the
plugin install fails, so clawops uses openclaw plugins install and verifies against
channels list --all --json.
Channel plugins are pinned to the supported runtime. The current latest does not install on
it: plugin "discord" requires plugin API >=2026.9.3, but this OpenClaw runtime exposes 2026.9.2. The same drift that forced version pins on model providers.
Telegram needs nothing installed: it ships in the image.
Bad config is caught before it is written
Config is validated against OpenClaw's own schema, captured from the image, not
hand-written, before anything is sent to the host, and again before a write replaces a
working file. clawops plan refuses a plan whose config the gateway would reject, while the
plan is still a file you can edit.
A rejected config is kept at <path>.rejected.<timestamp> and the live one is left alone, so
a validation failure never costs you what you were trying to write.
One rule is clawops's own: gateway.mode is optional in the schema and mandatory in
practice. A config without it passes openclaw config validate and then exits 78.
Containers are hardened
The gateway runs with --cap-drop=ALL, --security-opt no-new-privileges, --init and
--pids-limit 512. State is owned numerically by 1000:1000, matching the container's user
rather than a host account that may not have that uid.
The version pin is enforced everywhere it can change
doctor, plan, up and apply refuse an OpenClaw release outside the supported range, and
gateway restart reuses the version already deployed rather than resolving a moving tag. A
restart changes neither the version nor who can reach it.
The gateway is no longer exposed to your network
The container publishes on 127.0.0.1:18789 instead of 0.0.0.0:18789. Reach it with
clawops tunnel or a reverse proxy on the host.
Previously the wizard set allowedGatewayCidrs from the CIDR you gave for SSH, so a
plaintext HTTP dashboard. Token in the URL. Was opened to your whole shell-access network
as a side effect of one unrelated answer. To bind all interfaces deliberately, set
network.publishGateway: "all".
You must act if a client or reverse proxy on another machine reaches the gateway
directly, or external monitoring hits /health. A proxy on the host is unaffected; one in a
container on the host needs --network host.
Health checks can actually fail
The gateway serves its Control UI on a catch-all route, so any unmatched path answers 200 with HTML:
/healthz 200 application/json {"ok":true,"status":"live"}
/health-typo 200 text/html <!doctype html>…clawops probed with curl -fsS … >/dev/null, which succeeds on a typo. It proved something
was listening on the port, not that the gateway was healthy. Probes now read the response
body, and the restart gate uses /startupz rather than liveness, after a restart the
process listens long before startup finishes.
clawops mcp wire actually wires something now
It has never worked, not on 2.0, not on any 1.x release. It wrote gateway.mcpClients,
which is not a key OpenClaw has: checked against the config schemas of 2026.4.5,
2026.7.1-2 and 2026.9.2. The real key is top-level mcp.servers. And the entry it wrote
was command: "clawops" over stdio, which spawns inside the gateway container, where
clawops is not installed and nothing installs it.
On 1.x nothing validated the write, so clawops stored a key nothing read, restarted your gateway, and reported: "The gateway's AI can now run clawops commands." It could not.
flowchart TD
A["clawops mcp wire"] --> B["openclaw mcp add --transport streamable-http"]
B --> C{"gateway connects<br/>to the URL?"}
C -- no --> C1["probe fails, nothing saved,<br/>clawops prints the reason"]
C -- yes --> D["saved to mcp.servers.clawops"]
D --> E["openclaw mcp reload"]It delegates to openclaw mcp add now, which probes the server before saving, so
"wired" means the gateway connected, not that a file was written.
You have to run the server yourself. clawops is not installed on the gateway host:
clawops mcp serve --http 18790 --bind 0.0.0.0 --token "$(openssl rand -hex 16)"
clawops mcp wire --stack prod --token <same token>Installing clawops on the gateway host is a deliberate follow-up, not part of 2.0: it puts
deployment credentials on the deployed box, and the gateway's AI is reachable from every
channel it is connected to. See docs/security/threat-model.md T11.
clawops mcp serve --http serves more than one client, and asks who you are
Two bugs, found by testing against a real gateway rather than a mock.
It built one transport for the whole process, so the first client to connect claimed it
and every later one. A second editor, a reconnect, the gateway's own probe, was answered
"Server already initialized". HTTP mode is the multi-client mode.
It had no authentication, while exposing every tool including clawops_destroy. It now
takes a bearer token, compares it in constant time, and refuses to bind anywhere but loopback
without one.
The firewall follows the deployment
flowchart TD
A["clawops plan"] --> B{"publishGateway?"}
B -- "loopback (default)" --> C{"allowedGatewayCidrs empty?"}
C -- no --> C1["refuse: those rules would admit<br/>traffic to a closed port"]
C -- yes --> D["SSH rules only"]
B -- all --> E["SSH rules + gateway rules<br/>on spec.network.gatewayPort"]
D --> F["clawops harden"]
E --> F
F --> G["read the container's port bindings"]
G --> H{"published to the network?"}
H -- no --> H1["ufw: SSH only"]
H -- yes --> H2["ufw: SSH + the published port"]Three security controls were doing the opposite of what they say.
clawops harden opened the gateway port on every deployment. The ufw module ran
ufw allow 18789/tcp unconditionally. Since the gateway publishes on 127.0.0.1, that
opened a port nothing was listening on. A hardening step widening the firewall past what the
deployment exposes. It now reads the running container's port bindings and adds the rule only
when the gateway is really published, on whatever port it is published on.
The AWS security-group audit exempted the two ports it exists to check. Ports 22 and
18789 were on an "expected" list, so a group opening SSH or the gateway to 0.0.0.0/0 came
back as "No unexpected open ingress rules found". It also never read IPv6 rules, so ::/0
was invisible.
The setup wizard defaulted SSH access to 0.0.0.0/0. Pressing Enter opened SSH to the
whole internet, on the path most first-time users take. It offers your own IP as a /32 now,
and when that cannot be detected it offers no default and requires an answer.
clawops plan could not express any of it, and apply never passed any of it to Pulumi.
Both are fixed in 2.0.1. See the list at the top of this section.
The gateway port comes from the plan
"network": {
"allowedSshCidrs": ["203.0.113.4/32"],
"allowedGatewayCidrs": [],
"publishGateway": "loopback",
"gatewayPort": 9443
}One value now reaches the security-group rules, the container publish flag, the default
gateway.port and the gateway URL. It was a constant redeclared in eleven places, so
changing it meant finding all of them, and missing one produced a container publishing one
port, a gateway listening on another, and a firewall opening a third.
Local deployments use clawops up --gateway-port 9443.
clawops doctor answers whether it works, and says so in its exit code
flowchart TD
A["clawops doctor"] --> B["local: Node, Pulumi CLI + home,<br/>config, SSH key, credentials"]
B --> C{"--stack given?"}
C -- no --> Z["report"]
C -- yes --> D["container state"]
D --> E["deployed OpenClaw version"]
E --> F["probe /startupz<br/>and read the body"]
F --> G["published scope, disk,<br/>log rotation, hardening drift"]
G --> Z
Z --> Y{"any check failed?"}
Y -- no --> Y1["exit 0"]
Y -- yes --> Y2["exit 1"]Three changes:
It asks the gateway. doctor used to read docker inspect's healthcheck field, which
the OpenClaw image does not set, so it reported "no healthcheck configured" and moved on. A
running container means the process started, not that it serves. It now probes /startupz
and reads the body.
It exits 1 when something failed. Only an old Node.js used to do that; an unreadable SSH
key or an unsupported gateway exited 0, so a CI step running clawops doctor read a broken
deployment as success. Warnings still exit 0, a fresh machine with no stacks is
unconfigured, not broken.
It is an MCP tool. clawops_doctor returns the same report as structured data, so an
agent that hits a failure can find out why. It reports only; it never runs openclaw doctor --fix. --json gives the CLI the same report.
clawops agents list stops inventing an empty list
The command ended in || echo '[]', so a stopped container, a gateway still starting, or a
Docker permission error all produced "No agents running.", a wrong answer rather than an
error. It now fails, and says which.
Day-two commands work on AWS
gateway restart, logs, monitor, backup, agents, config set and doctor's
container checks were all broken on AWS: clawops connects as ubuntu, but provisioning
only put clawops in the docker group, so every Docker command failed with permission denied. GCP and Azure connect as clawops, so only AWS was affected.
Removed
clawops agents restart and the clawops_agents_restart MCP tool. OpenClaw 2.0 has no
per-agent restart, only gateway restart and daemon restart, both of which interrupt
every agent on the host. Use clawops gateway restart, or stay on @clawops/cli@legacy.
clawops agents list and clawops agents logs are unaffected.
Milestones
Milestone | Status | What ships |
M0: Scaffold | ✅ | Tooling, CI, stubs, generated types |
M1: GCP MVP | ✅ |
|
M2: Remote Mgmt | ✅ |
|
M3: AWS + Azure | ✅ | AWS EC2 + Azure VM adapters; |
M4: Local VM | ✅ | Local adapter (SSH bootstrap, no Pulumi); |
M5: MCP Layer | ✅ |
|
M6: Plan/Apply | ✅ |
|
M7: v1.0 Polish | ✅ | Full |
See docs/roadmap.md for the public roadmap and upcoming work.
License
MPL-2.0, see LICENSE.
Available Tools
20 toolsclawops_agents_listList OpenClaw AgentsARead-onlyIdempotent
List agents currently registered on the remote OpenClaw gateway.
Use when: the user wants to see which agents are running, debug agent routing, or count active workspaces.
Do NOT use when: the user wants one agent's logs — no tool exposes those; tell
the user to run clawops agents logs <name>.
| Name | Required | Description | Default |
|---|---|---|---|
| stackName | No | Which stack's agents to list. Omitted = the default stack in ~/.clawops/config.json |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds useful context that the listing targets the remote gateway and reflects 'currently registered' agents, which is meaningful beyond the annotations, though it does not detail output shape or pagination.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: a one-line action, a use-when block, and a do-not-use block with a concrete alternative. Every sentence adds decision-relevant value, and the core action is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, read-only, optional-parameter list tool, the description is complete: it states the action, the intended use cases, the exclusion, and the fallback for the excluded case. No output schema is present, but the nature of a listing operation makes return expectations straightforward.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the stackName parameter is fully documented in the schema itself. The description adds no parameter-specific guidance, but the baseline of 3 applies because the schema already carries the semantic weight.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with a specific verb and resource: 'List agents currently registered on the remote OpenClaw gateway.' It clearly identifies what the tool does and differentiates it from sibling tools about stacks, status, config, and logs by focusing on agent registration.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use the tool ('see which agents are running, debug agent routing, or count active workspaces') and when not to use it (one agent's logs), even providing the fallback command `clawops agents logs <name>`. This leaves no ambiguity for an agent deciding between tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clawops_applyApply Maker PlanADestructive
Apply a previously-generated Maker plan (deploy-plan.schema.json). Deterministic — the plan describes exactly what will be created.
Use when: the user has a plan file path and wants to apply it. This is the agent-friendly path for any deploy/destroy operation.
Do NOT use when: there's no plan file — generate one first with clawops_plan.
| Name | Required | Description | Default |
|---|---|---|---|
| yes | No | Skip the confirmation and apply immediately. Only when the user has already approved this specific plan | |
| planPath | Yes | Absolute path to plan JSON (R7) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide destructiveHint=trueaint and idempotentHint=false, so the safety profile is partially covered. The description adds valuable behavioral context by labeling the operation 'Deterministic' and clarifying that the plan describes exactly what will be created, while also noting it is the agent-friendly path for deploy/destroy operations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the core purpose, and uses short, scannable 'Use when' and 'Do NOT use when' sections. Every sentence earns its place with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no output schemachers, the description covers the key context: when to use it, when not to use it, the deterministic behavior, and the relationship to plan generation. It doesn't spell out the return value or confirmation flow, but annotations and the schema cover the critical safety and parameter aspects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already fully documented in the schema. The description doesn't add syntax or additional meaning beyond pointing to the plan file and the concept of applying it, which is sufficient but not above the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Apply') and resource ('previously-generated Maker plan'), and explicitly links it to deploy-plan.schema.json. It also differentiates from clawops_plan by specifying 'previously-generated' and is clear about the deterministic nature of the action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit 'Use when' and 'Do NOT use when' guidance, including the exact alternative (clawops_plan) and the condition that selects it ('no plan file'). This leaves no ambiguity for an agent deciding between tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clawops_config_getGet OpenClaw Config ValueARead-onlyIdempotent
Read a configuration value from the remote OpenClaw gateway.
Use when: the user wants to inspect current OpenClaw config (e.g., which model provider is active, which channels are enabled).
Do NOT use when: the user wants to change the config — use clawops_config_set instead.
| Name | Required | Description | Default |
|---|---|---|---|
| key | No | Dot-path config key, e.g., gateway.auth.mode. Omit to dump the full config. | |
| stackName | No | Which stack's gateway config to read. Omitted = the default stack in ~/.clawops/config.json |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile with readOnlyHint, idempotentHint, and destructiveHint false, so the description does not need to restate it. It adds useful context about the remote gateway and the full-dump behavior when key is omitted, but does not disclose return format or error behavior. This is adequate but not rich beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well structured: one action sentence followed by clear use and non-use conditions. Every sentence earns its place, and the primary purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a simple read-only tool with two optional parameters, no output schema, and rich annotations. The schema plus description fully specify how to invoke it, and sibling routing is explicit. Nothing required for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with key and stackName already documented including the dot-path example and the default stack path. The tool description adds no extra parameter meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Read') and resource ('configuration value from the remote OpenClaw gateway'), and gives concrete inspection examples such as active model provider and enabled channels. This clearly distinguishes it from mutation siblings like clawops_config_set.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly provides 'Use when' and 'Do NOT use when' conditions, naming clawops_config_set as the alternative for changing config. An agent can reliably decide between read and mutate operations without further inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clawops_config_setSet OpenClaw Config ValueADestructiveIdempotent
Set a configuration value on the remote OpenClaw gateway. Optionally restarts the gateway after.
Use when: the user wants to change OpenClaw configuration (model provider, channel auth, gateway port).
Do NOT use when: the user is reading config — use clawops_config_get. Do NOT bulk-edit; one key at a time so changes are auditable (R21).
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Dotted path into the gateway config, e.g. models.provider or channels.slack.enabled | |
| value | Yes | Stringified; JSON for complex values | |
| restart | No | Restart the gateway so the change takes effect. Without it the value is written and the running gateway keeps the old one | |
| stackName | No | Which stack's gateway config to write. Omitted = the default stack in ~/.clawops/config.json |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already set readOnlyHint=false, destructiveHint=true, idempotentHint=true, and openWorldHint=true, so the mutation/safety profile is covered. The description adds genuine behavioral detail: optional gateway restart, the fact that the running gateway keeps the old value without restart, and the one-key-at-a-time auditability rule. It does not discuss failure modes or permissions, but nothing contradicts the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three compact sentences plus two clear usage lines, with the core action and optional restart stated first and every sentence earning its place. There is no redundant repetition of the title or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter mutation tool with full schema coverage and safety annotations, the description supplies the missing decision context: when to use it, which sibling to use instead, and the no-bulk-edit constraint. No output schema is present, but the write behavior and restart semantics are adequately specified, so an agent can invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with each parameter already documented (dotted key path, stringified JSON value, restart default/effect, stackName default). The description contributes context through example config areas, but the schema carries the parameter-semantic burden and the description does not add much beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Set') with a clear resource ('configuration value on the remote OpenClaw gateway') and a scope-limiting optional behavior (restart). It also names clawops_config_get as the reading alternative, so the tool's identity is not confused with siblings like clawops_config_unset or clawops_config_get.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use examples ('model provider, channel auth, gateway port') and an explicit when-not-to-use with the alternative tool ('reading config — use clawops_config_get'). It also prohibits bulk-edit usage, which is a concrete exclusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clawops_config_unsetUnset OpenClaw Config KeyADestructiveIdempotent
Remove a configuration key from the remote OpenClaw gateway config, reverting it to the OpenClaw default.
Use when: the user wants to delete a config key entirely (e.g., remove a channel, clear an override).
Do NOT use when: the user wants to set the key to a new value — use clawops_config_set instead.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Dot-path config key to remove | |
| restart | No | Restart the gateway so the removal takes effect. Without it the key is removed and the running gateway keeps the old value | |
| stackName | No | Which stack's gateway config to write. Omitted = the default stack in ~/.clawops/config.json |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=true. The description adds genuinely useful context: the removal reverts to the OpenClaw default and, as implied by the restart parameter, does not affect the running gateway until a restart. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tight and front-loaded: the core action and outcome come first, followed by structured use/don't-use guidance. Every sentence earns its place; no filler or redundant restatement of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 3-parameter mutation tool, the schema covers all parameters, annotations cover destructive/idempotent behavior, and the description supplies alternative routing. There is no output schema, so no return-value expectation exists. A minor gap is that the description does not explicitly flag irreversibility or failure modes, but overall it is sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (key, restart, stackName) is already documented with meaningful descriptions. The tool description adds no parameter-level semantics beyond what the schema provides, warranting the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Remove') and resource ('configuration key from the remote OpenClaw gateway config'), and spells out the result ('reverting it to the OpenClaw default'). This clearly distinguishes it from sibling tools like clawops_config_set and clawops_config_get without needing to inspect schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'Use when' and 'Do NOT use when' sections give unambiguous routing criteria and name the exact alternative (clawops_config_set) for the contrasting case. An agent can decide with zero inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clawops_config_validateValidate OpenClaw ConfigARead-onlyIdempotent
Validate the remote OpenClaw gateway config against the known schema. Checks for structural errors (wrong types, unknown top-level keys) that would cause OpenClaw to fail on startup.
Use when: the user wants to verify config before restarting the gateway, or after editing openclaw.json manually.
Do NOT use when: the user wants to change config — use clawops_config_set.
| Name | Required | Description | Default |
|---|---|---|---|
| stackName | No | Which stack's deployed config to validate. Omitted = the default stack in ~/.clawops/config.json |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds context about what the validation actually checks (structural errors that would cause startup failure) and that it targets the remote gateway config. This is useful beyond annotations, though it doesn't detail the return format or whether it reports pass/fail. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured: the first sentence states the core purpose, the second explains what it checks, and the third/fourth provide usage conditions. Every sentence earns its place, with no redundant information. It's front-loaded with the most important information and stays within a reasonable length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a validation tool with one optional parameter and readOnly/idempotent annotations, the description covers the essential context: what it does, what it checks, and when to use it. The only gap is that it doesn't describe the output format (e.g., pass/fail, error report). However, with no output schema and annotations indicating a safe operation, the agent likely needs to know the result format to interpret the response. Still, the description is largely complete for invocation decisions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter stackName is fully described in the schema (100% coverage), and the description adds no additional parameter details. Since the schema already explains the parameter and its default behavior, the baseline of 3 applies. The description does not add value here beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool validates the OpenClaw gateway config against a known schema, checking for specific structural errors (wrong types, unknown top-level keys). It distinguishes itself from the config-changing sibling (clawops_config_set) by naming the alternative and its use case. The verb 'validate' and resource 'remote OpenClaw gateway config' are precise and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'Use when' and 'Do NOT use when' sections give concrete scenarios: verifying config before restarting or after manual edits, and explicitly routing config changes to clawops_config_set. This is the gold standard for usage guidance – the agent knows exactly when to invoke this tool and when to choose a sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clawops_destroyDestroy Stack (DESTRUCTIVE)ADestructiveIdempotent
Destroy a clawops stack. Removes ALL provisioned resources. Triggers elicitation confirmation showing the resource diff before execution (R19). Cannot be undone.
Use when: the user explicitly asks to destroy, tear down, delete, or remove a stack. Always confirm the stack name.
Do NOT use when: the user wants to stop the gateway temporarily — that is not a destroy, and no tool stops a gateway; clawops_gateway_restart is the only gateway tool. Do NOT use when in doubt about which stack; list first with clawops_stacks_list.
| Name | Required | Description | Default |
|---|---|---|---|
| yes | No | Skip elicitation; CI/scripted use only | |
| stackName | Yes | The stack to destroy. Required, and worth repeating back to the user before calling: this deletes every resource it provisioned |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is known. The description adds valuable specifics: triggers elicitation confirmation with a resource diff (R19), irreversibility ('Cannot be undone'), and the instruction to always confirm the stack name. It does not mention idempotency behavior (despite idempotentHint=true), but this is a minor omission.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections: the core action and consequences, then 'Use when' and 'Do NOT use when' blocks. Every sentence carries meaning—there is no filler. The key warning about irreversibility is front-loaded, and the guidance is organized for quick parsing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive action with no output schema, the description covers all essential context: what it does, when to use, when not to use, safety confirmations, and the irreversibility. It also addresses common pitfalls (confusing with gateway restart, unsure stack). Nothing critical is missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% – both 'stackName' and 'yes' have descriptions in the schema. The tool description adds a reiteration that stackName is 'worth repeating back to the user' due to destructiveness, but this is largely a reinforcement of the schema's warning rather than new meaning. Given the baseline of 3 for high coverage, this is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Destroy') and resource ('a clawops stack'), and explicitly notes it 'Removes ALL provisioned resources,' making the purpose unambiguous. It also differentiates from siblings by warning against using it for gateway stop, reinforcing that it is exclusively for destruction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit 'Use when' conditions (user asks to destroy/tear down/delete/remove) and 'Do NOT use when' conditions (temporary gateway stop, uncertainty about stack), naming alternatives like clawops_gateway_restart and clawops_stacks_list. This is exemplary guidance for tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clawops_doctorRun DiagnosticsARead-onlyIdempotent
Run clawops's diagnostics and return the report: Node and Pulumi runtime, config, SSH key and known_hosts, cloud credentials per configured provider, and the supported OpenClaw range. With stackName, also contacts the host for container state, the deployed OpenClaw version, a real gateway health probe, whether the port is published to the internet, disk usage on the state directory, log rotation, and hardening drift.
Use when: something is not working and you do not yet know what; before any deploy, upgrade or migration; or to find out which OpenClaw version a gateway is actually running.
Do NOT use when: you already know the problem and want to fix it. This tool only
reports — it changes nothing, and never runs openclaw doctor --fix. No tool
repairs a gateway: clawops gateway update is CLI-only, so tell the user to run
it themselves, after clawops backup create.
Every check carries a status: fail (something is wrong that clawops can name),
warn (worth knowing, not broken), info (did not apply). ok is false only when
something failed — a fresh machine with no stacks is full of warnings and healthy.
| Name | Required | Description | Default |
|---|---|---|---|
| stackName | No | Stack to include remote checks for. Without it, only the local machine is checked — no SSH connection is made. | |
| failuresOnly | No | Return only failing and warning checks. Passing checks are counted, not listed. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations that declare read-only and idempotent hints, the description states 'it changes nothing, and never runs `openclaw doctor --fix`,' reinforcing the non-destructive behavior. It also explains the status semantics (fail/warn/info) and the meaning of `ok`, which is crucial for interpreting the report, and discloses that with `stackName` it contacts the host remotely.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is organized into three distinct paragraphs: scope, usage criteria, and output semantics. Every sentence adds necessary information (what, when, when not, what it changes, how to read results) without fluff, and it is front-loaded with the core diagnostic purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no required parameters and no output schema, so the description carries the burden of explaining returns. It lists all check categories, describes the status values, and clarifies the meaning of `ok` and warnings, giving an agent enough context to invoke the tool and interpret results. The usage restrictions and behavioral notes fill gaps that annotations alone would not cover.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
While the schema already documents both parameters with 100% coverage, the description adds semantic detail by listing the additional remote checks triggered by `stackName` (container state, deployed version, gateway probe, port exposure, disk usage, log rotation, hardening drift) and explains the status categories that `failuresOnly` filters on. This enriches the meaning beyond the schema's brief descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Run clawops's diagnostics and return the report,' then enumerates the exact checks performed (Node/Pulumi runtime, config, SSH key, cloud credentials, etc.). It clearly differentiates from siblings by framing itself as the 'do not know what's wrong' diagnostic tool, distinct from `clawops_status`, `clawops_harden`, or `clawops_monitor`.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description includes explicit 'Use when' and 'Do NOT use when' sections, giving the agent clear decision criteria: use it for unknown problems, before deploy/upgrade/migration, or to discover the actual gateway version. It explicitly says not to use when the problem is already known and wants a fix, and it points out that no tool repairs a gateway, directing the user to CLI commands instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clawops_gateway_restartRestart Gateway DaemonADestructiveIdempotent
Restart the OpenClaw gateway daemon on the remote instance.
Use when: a gateway-wide config change requires reload, or the gateway is reported as unresponsive.
Note: OpenClaw 2.0 removed per-agent restart; gateway restart is the only
restart it offers, and it affects every agent on the host. Brief downtime (~10s).
| Name | Required | Description | Default |
|---|---|---|---|
| stackName | No | Which stack's gateway to restart. Omitted = the default stack in ~/.clawops/config.json |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds valuable context beyond annotations: 'it affects every agent on the host' and 'Brief downtime (~10s).' Annotations already signal destructive and non-read-only, but the specific scope and duration are new behavioral disclosures. It also highlights the version-specific limitation (OpenClaw 2.0), which helps set expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is composed of three short sentences, each carrying necessary information: the action, the trigger conditions, and the global impact/downtime. There is no fluff, but it could be slightly tightened; however, each sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with a single optional parameter and no output schema, the description adequately covers the purpose, when to use it, impact scope, and downtime. It doesn't explicitly mention how to confirm the restart succeeded (e.g., via clawops_status), but that is a minor omission for such a simple restart operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides a full description of stackName, including the default-stack behavior when omitted. The tool description does not add any extra meaning to the parameter, so the baseline 3 for high schema coverage applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Restart the OpenClaw gateway daemon on the remote instance.' This clearly distinguishes it from sibling tools like clawops_up or clawops_apply, which are about deployment/configuration rather than restarting. The 'Use when' clause further reinforces the intended purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit trigger conditions: 'gateway-wide config change requires reload' or 'gateway is reported as unresponsive.' It also notes that per-agent restart is unavailable, ruling out that alternative. However, it doesn't name another sibling tool to use for different scenarios, only states that this is the only restart available.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clawops_hardenHarden StackADestructiveIdempotent
Apply security hardening to a deployed stack: SSH, UFW, fail2ban, unattended-upgrades, the Docker socket, and per-cloud checks. Optionally join the stack to a Tailscale network and reach it there instead of over the public internet.
Use when: the user asks to harden, secure, or lock down a stack; asks what the hardening report says (with dryRun: true, which changes nothing); or asks to put a stack on their tailnet.
tailscale: true installs Tailscale, joins the tailnet as clawops-, and
then moves clawops onto that address — but only after opening an SSH session to
it, against host keys pinned over the connection already trusted. If that fails,
nothing is recorded and the public address stays in use. The auth key comes from
clawops secret set TAILSCALE_AUTH_KEY and is never passed through this tool
(R6).
tailscaleRevert: true takes the host back off the tailnet, over its public address. On a private-only stack it refuses and names the plan/apply commands that reopen SSH first — relay them rather than trying to work around it.
To close the public ports afterwards, plan with privateOnly: true and apply that plan; this tool does not change firewall rules.
Do NOT use when: the user wants to know whether a stack is healthy — that is clawops_doctor. Do NOT pass tailscale and tailscaleRevert together.
| Name | Required | Description | Default |
|---|---|---|---|
| yes | No | Skip elicitation; CI/scripted use only | |
| dryRun | No | Report the current state, change nothing | |
| options | No | Comma-separated module IDs; default is every defaultOn module for the provider | |
| stackName | No | Which stack to harden. Omitted = the default stack in ~/.clawops/config.json | |
| tailscale | No | Join the tailnet, verify this machine reaches the host there, then use that address | |
| tailscaleRevert | No | Leave the tailnet and go back to the public address |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes far beyond the annotations. It discloses the rollback/failure path for tailscale: true ('If that fails, nothing is recorded and the public address stays in use'), the refusal behavior for tailscaleRevert on private-only stacks, that it 'does not change firewall rules', and that the auth key is never passed through the tool (R6). Consistent with readOnlyHint=false and destructiveHint=true — no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded in the first sentence, and the paragraphs are logically grouped (usage triggers, tailscale, revert, exclusions). It is long, but the Tailscale integration is genuinely intricate and every sentence earns its place — no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Comprehensive for a 6-param, security-sensitive tool: covers scope, when-to-use, failure modes, reversals, and non-goals (firewall, health checks). There is no output schema, and the description doesn't spell out the return format, but it does reference 'the hardening report' under dryRun. Minor gap only.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds genuine value beyond the schema for tailscale and tailscaleRevert (host-key pinning, failure semantics, private-only refusal), pushing it to a 4. The yes, options, and stackName params are fully covered by the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb-resource pair ('Apply security hardening to a deployed stack') and enumerates concrete modules (SSH, UFW, fail2ban, unattended-upgrades, Docker socket, per-cloud checks). It also names the sibling it is not ('that is clawops_doctor'), so an agent can disambiguate without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit 'Use when' triggers (harden/secure/lock down, ask for the hardening report, put a stack on the tailnet) and explicit 'Do NOT use when' exclusions (health checks → clawops_doctor), plus a hard constraint ('Do NOT pass tailscale and tailscaleRevert together'). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clawops_initInitialise clawops ConfigA
Register a stack and write ~/.clawops/config.json: the provider, the state backend, the region, and an SSH key pair generated if one is not already there. Nothing is provisioned and nothing is charged; this only creates local configuration.
Use when: any other clawops tool reports that there is no config, or the user wants to add a second stack alongside the ones they have. This is the first call on a machine that has never run clawops — a fresh container, a new laptop, a sandbox.
Adding a stack is additive and safe: existing stacks are kept. Overwriting one needs force: true, because changing a state backend orphans the Pulumi state it points at — the infrastructure stays up and clawops can no longer see or destroy it.
For aws, gcp and azure, stateUrl can be omitted and clawops names a bucket from the account it can see; that needs cloud credentials in the environment, so in a sandbox pass stateUrl explicitly. The local provider needs host instead, and no cloud account at all.
Do NOT use when: the user wants to deploy — that is clawops_up, after this. Credentials are never passed here: clawops reads them from the environment (R6).
| Name | Required | Description | Default |
|---|---|---|---|
| host | No | [local only] Hostname or IP of the machine to manage. Required when provider is local | |
| force | No | Overwrite a stack that already exists. Refused without it, because replacing a state backend orphans the state it points at | |
| region | No | Cloud region in the provider's own spelling. Omitted = us-east-1 (aws), us-central1 (gcp), eastus (azure) | |
| sshPort | No | [local only] SSH port. Omitted = 22 | |
| sshUser | No | [local only] SSH login user. Omitted = root | |
| provider | Yes | Which cloud this stack deploys to, or 'local' for a machine you already have | |
| stateUrl | No | Where Pulumi state lives, e.g. s3://bucket/clawops, gs://bucket/clawops, azblob://container. Omitted = clawops names one from the cloud account it can see, which needs credentials | |
| stackName | No | Name for the stack, used by every later call. Omitted = "default" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial behavior beyond annotations: SSH key pair generation, additive safety of stacks, the force:true overwrite path orphaning Pulumi state, credential handling from environment, and provider-specific bucket naming behavior. The annotations only say read-only/destructive/idempotent flags, so this context is genuinely valuable and consistent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average, but every sentence carries operational information: safety, prerequisites, provider nuances, and routing to the right sibling. It is front-loaded with the core purpose and organized into scannable paragraphs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter initialization tool with no output schema, the description covers purpose, safety, prerequisites, provider-specific parameter rules, exclusions, and sibling routing. An agent has everything it needs to decide when to call it and how to configure it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds inter-parameter meaning beyond the schema: omitting stateUrl makes clawops derive a bucket needing credentials, local requires host instead of cloud account, and force:true is tied to the state-backend orphaning warning. This is a clear improvement over the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Register a stack and write ~/.clawops/config.json', and names exactly what it creates. It clearly distinguishes itself from deployment tools by stating 'Nothing is provisioned and nothing is charged; this only creates local configuration.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit use conditions: when no config exists, when adding a second stack, or on a fresh machine. It also names an exclusion and alternative: 'Do NOT use when: the user wants to deploy — that is clawops_up, after this.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clawops_logs_tailTail Gateway LogsARead-onlyIdempotent
Tail recent gateway logs from a clawops-managed instance.
Use when: the user wants to investigate recent activity or errors, or asks "what's been happening" on the gateway.
Do NOT use when: the user wants real-time streaming logs (those are not
well-suited to tool calls; suggest the user run clawops logs -f directly
in their terminal). Do NOT use for instance-level system logs: no tool runs
arbitrary remote commands, so tell the user to run clawops ssh --command 'journalctl ...' themselves.
| Name | Required | Description | Default |
|---|---|---|---|
| sinceMin | No | Lines since N minutes ago | |
| stackName | No | Which stack's gateway to read. Omitted = the default stack in ~/.clawops/config.json | |
| tailLines | No | How many of the most recent lines to return. Keep it small; output is trimmed to 8KB regardless |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as readOnly, idempotent, and non-destructive, so the safety profile is covered. The description adds useful behavioral context by clarifying this is not real-time streaming and that output is trimmed to 8KB regardless of requested line count.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The opening sentence states the core purpose, followed by tightly organized 'Use when' and 'Do NOT use when' sections. Every sentence is decision-relevant and there is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only log tool with three optional parameters, the description provides all needed context: what resource it targets, what cases are out of scope, and how users can accomplish those out-of-scope tasks. The output trimming behavior is also disclosed, which is sufficient given the simple log-return expectation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameters are already well documented in the schema and the tool description does not need to compensate. The description adds a small caveat about keeping tailLines small, but it does not add significant meaning beyond the structured schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Tail recent gateway logs from a clawops-managed instance.' It clearly distinguishes this from instance-level system logs and real-time streaming, and none of the sibling tools cover log tailing, so an agent can identify its role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit 'Use when' conditions ('investigate recent activity or errors') and explicit 'Do NOT use when' conditions with concrete alternatives: suggest `clawops logs -f` for streaming and `clawops ssh --command 'journalctl ...'` for system logs. This is model behavior for routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clawops_monitorMonitor Stack HealthARead-onlyIdempotent
Take a live snapshot of a running clawops stack: gateway health, container status, resource usage (CPU, memory, disk), and recent log lines.
Use when: the user wants to know if the gateway is running, how much memory or CPU it is using, what recent log activity looks like, or wants a quick health overview richer than clawops_status.
Do NOT use when: the user wants real-time streaming logs (use
clawops_logs_tail or suggest clawops logs -f). Do NOT use for
configuration queries (use clawops_config_get).
| Name | Required | Description | Default |
|---|---|---|---|
| stackName | No | Stack name. Defaults to active stack. | |
| tailLines | No | Log lines to include in snapshot. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful non-schema behavior: it is a live snapshot rather than streaming, it is scoped to a running stack, and it is richer than clawops_status. This meaningfully supplements the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly scoped and front-loads the core action and contents, then provides crisp usage guidance. Every sentence earns its place, and the Do NOT use section prevents misuse without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only snapshot tool with two optional, fully documented parameters, the description covers invocation scope, output contents, relationship to siblings, and exclusion cases. It is complete enough for an agent to select and call the tool correctly without additional information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both parameters already documented (stackName defaults to active stack; tailLines defaults to 5). The description's mention of 'recent log lines' aligns with tailLines but does not add parameter-level depth beyond the schema. Baseline 3 is appropriate when the schema carries the semantic load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Take a live snapshot') and resource ('running clawops stack'), and enumerates exactly what it covers: gateway health, container status, resource usage, and recent log lines. It also differentiates itself from clawops_status by positioning itself as a richer health overview.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit 'Use when' conditions and explicit 'Do NOT use when' exclusions with named alternatives (clawops_logs_tail, clawops_config_get). An agent can unambiguously select this tool versus siblings based on stated user intent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clawops_planGenerate Deploy PlanARead-onlyIdempotent
Generate a Maker deploy plan (does NOT apply). Plan is JSON conforming to deploy-plan.schema.json — review before applying.
Use when: the user wants to see what would be created before committing, or you (the agent) need a reviewable artifact for the user to approve.
Do NOT use when: the user has explicitly asked to deploy and you already have their approval — go directly to clawops_up.
| Name | Required | Description | Default |
|---|---|---|---|
| region | No | Cloud region, in the provider's own spelling (us-east-1, us-central1, eastus). Omitted = the region recorded for the stack | |
| outPath | No | Absolute path to write plan; if omitted, plan returned inline | |
| sshCidr | No | CIDR(s) allowed to reach SSH, comma-separated, or 'auto' for this machine. Omitted = none, and nothing will be able to connect | |
| provider | No | Cloud to plan against. Omitted = the provider recorded for the stack. 'local' has no plan/apply path and is refused | |
| stackName | No | Which stack the plan is for. Omitted = the default stack in ~/.clawops/config.json | |
| gatewayCidr | No | CIDR(s) allowed to reach the gateway port, or 'auto'. Requires publishGateway=all | |
| privateOnly | No | Close public SSH and gateway access; reach the stack over its tailnet. Requires a verified tailnet address (clawops_harden with tailscale), and refuses unless that address answers SSH now | |
| instanceType | No | A clawops alias (micro|small|medium|large|gpu) or a machine type the cloud names itself, e.g. t3.small | |
| publishGateway | No | Which interface the gateway binds. loopback (default) keeps it off the network | |
| openclawVersion | No | semver, or 'stable'/'dev' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds meaningful behavior beyond that: 'does NOT apply', output is JSON conforming to an external schema, and the artifact is meant for review before application. It does not disclose minor operational details like auth requirements, but those are less critical for a read-only plan tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the core purpose, and each sentence earns its place. The 'Use when' and 'Do NOT use when' structure is scannable and gives the agent direct decision rules without wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-optional-parameter planning tool, the schema covers all parameters and the description covers output format ('JSON conforming to deploy-plan.schema.json') and the decision boundary for when to use it. It is slightly less self-contained because the referenced schema is not inline, and the distinction between clawops_up and clawops_apply is not elaborated, but nothing essential is missing for correct selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and every parameter already has a detailed description, including defaults, omitted behavior, and refusal cases. The tool description itself does not add parameter-level meaning, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Generate a Maker deploy plan' and immediately clarifies the key distinction that it 'does NOT apply'. This differentiates it from apply/up siblings without requiring the agent to inspect any other tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use it ('when the user wants to see what would be created before committing') and when not to use it ('user has explicitly asked to deploy... go directly to clawops_up'). This is model guidance with a named alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clawops_stacks_listList StacksARead-onlyIdempotent
List all clawops-managed stacks across all configured providers.
Use when: the user wants an overview of their deployments, asks "what stacks do I have", or wants to compare stacks before an operation.
Do NOT use when: the user named a specific stack — use clawops_status directly with that name.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already convey readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds behavioral scope: it enumerates managed stacks across all configured providers and has no filtering. It doesn't describe output format or pagination, but for a read-only list with zero parameters that is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: one sentence states what the tool does, followed by expanded use-when and do-not-use guidance. Every sentence earns its place, and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only, non-destructive list tool with no output schema, the description is complete. It tells the agent what the tool returns at a conceptual level ('overview of deployments'), when to use it, and which sibling to route to when a specific stack is named. No critical context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters and schema description coverage is 100%, so the schema already fully covers this dimension. The description reinforces the no-parameter, unfiltered nature by saying 'all' stacks 'across all configured providers,' which helps an agent understand that no arguments are needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'List all clawops-managed stacks across all configured providers.' It also differentiates itself from the sibling clawops_status by explicitly saying that a named specific stack should go to clawops_status instead. This gives an agent a precise, unambiguous purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use the tool ('user wants an overview of their deployments', 'what stacks do I have', 'compare stacks before an operation') and when not to use it ('user named a specific stack'), and names the alternative tool (clawops_status). This satisfies the when/when-not/alternatives requirement fully.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clawops_statusGet Stack StatusARead-onlyIdempotent
Get what is deployed for a clawops-managed stack: public IP, gateway URL, SSH user, or that nothing is deployed yet. Reads stack outputs; does not contact the host.
Use when: the user asks what exists for a stack, or where to reach it.
Do NOT use when: the user asks whether the gateway is actually WORKING — that needs the host, so use clawops_doctor. Also not for live logs (use clawops_logs_tail) or config values (use clawops_config_get).
| Name | Required | Description | Default |
|---|---|---|---|
| stackName | No | Stack name. Defaults to active stack from config. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false. The description adds valuable behavioral context beyond annotations: 'Reads stack outputs; does not contact the host,' which clarifies why this is not suitable for liveness checks. It also discloses the possible empty result ('nothing is deployed yet').
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then gives crisp usage boundaries. Every sentence earns its place, and the use/do-not-use sections are structured for quick agent scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-optional-parameter read-only tool with no output schema, the description provides enough to select and invoke correctly: what it returns, what it does not do, and when to choose an alternative. Nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: the only parameter, stackName, is fully described in the schema including the default behavior. The description adds no additional parameter details, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Get') with a clear resource ('what is deployed for a clawops-managed stack') and enumerates concrete outputs: public IP, gateway URL, SSH user, or nothing deployed. It also distinguishes itself from siblings by explicitly noting it reads stack outputs and does not contact the host.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is an explicit 'Use when' clause and a 'Do NOT use when' clause with named alternatives (clawops_doctor, clawops_logs_tail, clawops_config_get). This gives an agent unambiguous routing guidance for selecting this tool over siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clawops_task_statusGet Task StatusARead-onlyIdempotent
Poll the status of a long-running clawops task (returned by clawops_up, clawops_destroy, clawops_apply, etc.). Per R12 streaming model.
Use when: the user is waiting on a long-running deploy/destroy and wants progress, OR you need to check whether a previously-started operation finished.
Do NOT use when: there is no active taskId — start the operation first.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes | The taskId returned by a long-running tool such as clawops_up, clawops_apply or clawops_destroy |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds that it's a polling mechanism 'Per R12 streaming model' and clarifies it checks progress/completion. This goes slightly beyond annotations, though it doesn't detail polling intervals or response format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured, with purpose stated first and usage conditions clearly separated. Every sentence earns its place; there is no redundancy or padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple poll tool with one parameter and annotations covering safety, the description provides usage context, prerequisites, and the type of task. It mentions the streaming model but doesn't describe the return format; however, the agent can infer it returns status. This is adequate for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% — the taskId parameter is fully described as the ID returned by long-running tools. The description reinforces this but doesn't add new meaning beyond the schema. With full schema coverage, the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Poll') and resource ('long-running clawops task'), and names the exact tools (clawops_up, clawops_destroy, clawops_apply) that return the taskId. This clearly differentiates it from siblings like clawops_status (which likely reports overall system state) and other list tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit 'Use when' conditions (user waiting on a long-running deploy/destroy, or checking if an operation finished) and a 'Do NOT use when' condition (no active taskId — start the operation first). This leaves no ambiguity about when to invoke it versus starting a new operation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clawops_upProvision and Deploy StackADestructiveIdempotent
Provision and deploy a clawops stack. Idempotent — re-running with no spec change produces no diff. Long-running (median 3min, p99 8min) so emits progress notifications per R12.
Use when: the user explicitly asks to deploy, provision, create, or "spin up" a stack. Always after the user has reviewed a plan (clawops_plan first when in doubt).
Do NOT use when: the user has not yet generated a plan and is in exploratory/discovery mode — use clawops_plan first. Do NOT use for an existing stack you only need to update; refresh first.
| Name | Required | Description | Default |
|---|---|---|---|
| dryRun | No | Show what would be created and change nothing. Use this first when the user has not yet approved a spend | |
| region | No | Cloud region to deploy into, in the provider's own spelling (us-east-1, us-central1, eastus). Omitted = the region recorded for the stack | |
| sshCidr | No | CIDR(s) allowed to reach SSH, comma-separated, or 'auto' for the caller's own address. Omitted means none, and nothing will be able to connect — including every clawops day-two command. | |
| provider | No | Defaults to provider configured for this stack | |
| stackName | No | Name for the stack to provision, and the name every later command refers to it by. Omitted = the default stack in ~/.clawops/config.json | |
| gatewayCidr | No | CIDR(s) allowed to reach the gateway port. Requires publishGateway=all. | |
| instanceType | No | A clawops size (micro|small|medium|large|gpu) or a provider-native machine type. Not an enum: Azure offers SKU families per subscription, and an account offered none of the five sizes clawops names would otherwise have no way to deploy. | small |
| publishGateway | No | Which interface the gateway binds. 'all' serves plaintext HTTP. | |
| openclawVersion | No | semver or 'stable'/'dev' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide idempotentHint and destructiveHint, but the description adds concrete behavior: idempotency detail ('no diff'), long-running timing (median 3min, p99 8min), and progress notifications. This is valuable context beyond the annotations, though it does not explicitly describe side effects on spec changes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and front-loaded: core function first, then idempotency and timing, then usage guidance. Every sentence adds value, and it is not verbose despite covering multiple aspects.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter provisioning tool with no output schema, the description covers usage, alternatives, timing, and idempotency. It does not describe the return value, but given the lack of output schema and the presence of progress notifications, this is acceptable and likely sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 9 parameters. The description does not add parameter-specific meaning beyond referencing plan usage; it meets the baseline but does not elevate it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Provision and deploy a clawops stack') and differentiates from siblings by explicitly naming clawops_plan for planning and refresh for updates. It is unambiguous about what this tool does versus alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit 'Use when' and 'Do NOT use when' conditions, naming alternatives (clawops_plan, refresh) and the decision criteria (explicit deploy request, plan reviewed, exploratory mode). This is exemplary guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clawops_workflow_deploy_appDeploy OpenClaw (End-to-End Workflow)ADestructiveIdempotent
Single-tool workflow that takes a user from "I want to deploy OpenClaw to " to a verified, healthy gateway. Internally: plan → user confirms (elicitation) → up → wait for healthy → return URL.
Use when: the user expresses end-to-end deployment intent ("deploy to AWS", "spin up an OpenClaw on GCP for me").
Do NOT use when: the user is mid-deployment and only needs one step (e.g., they already have a plan; use clawops_apply). Do NOT use for destroying or updating — separate workflows.
| Name | Required | Description | Default |
|---|---|---|---|
| region | No | Cloud region, in the provider's own spelling (us-east-1, us-central1, eastus). Omitted = the provider's default | |
| provider | Yes | Cloud to deploy to. Omitted = the default provider in ~/.clawops/config.json | |
| stackName | No | Name for the new stack. Omitted = the default stack name in ~/.clawops/config.json | default |
| instanceType | No | Machine size: a clawops alias (micro|small|medium|large|gpu) or a type the cloud names itself, e.g. t3.small | small |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds valuable context beyond the annotations: it reveals the internal sequencing ('plan → user confirms (elicitation) → up → wait for healthy → return URL') and the completion condition. It does not explicitly discuss destructive side effects, but the annotations already carry destructiveHint: true, so the description is not misleading and the annotations cover that risk.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly structured: the workflow is front-loaded in the first sentence, followed by crisp use-case and exclusion rules. There is no filler or repetition of schema content, and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex workflow tool with no output schema, the description covers the overall lifecycle, user confirmation requirement, health verification, and the returned URL. It leaves minor gaps around failure behavior and output formatting, but these are secondary given the schema and annotations already cover parameter defaults and safety signals.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema documents all four parameters clearly with defaults, enums, and provider-specific spelling notes. The description adds no new parameter-level detail, which is acceptable given the schema already carries the semantic load; a baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('deploy OpenClaw') and a clear scope: an end-to-end workflow from intent to a 'verified, healthy gateway.' It also distinguishes itself from siblings by calling out that it is a 'Single-tool workflow' and explicitly marking 'clawops_apply' as the one-step alternative, so an agent can identify it without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit 'Use when' and 'Do NOT use when' guidance with concrete user phrasing examples ('deploy to AWS', 'spin up an OpenClaw on GCP for me'). It names the alternative tool (clawops_apply) and excludes destroy/update cases, leaving no ambiguity about when to select this workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clawops_workflow_recoverRecover/Diagnose StackARead-onlyIdempotent
Diagnostic workflow for an unhealthy stack. Internally: status check → gateway logs → agent logs → systemd service status → produces a structured diagnostic report with suggested remediation.
Use when: the user reports any "not working" symptom and you don't know where to start. Best entry point for troubleshooting.
Do NOT use when: the user has already identified the problem and asks for a specific fix.
| Name | Required | Description | Default |
|---|---|---|---|
| stackName | No | Which stack to diagnose. Omitted = the default stack in ~/.clawops/config.json |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is established. The description adds meaningful behavioral detail by listing the internal steps and the output format, showing that the tool reads logs and service status rather than mutating state. It does not explain partial-failure behavior or permission requirements, but this is not a major gap given the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the core purpose and internal workflow come first, followed by crisp use and do-not-use conditions. Every sentence contributes; there is no filler or redundant restating of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one optional parameter and no output schema, the description covers the essential context: when to trigger it, what steps it performs, what it returns, and when to avoid it. The only remaining parameter detail is already captured by the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers 100% of the parameter documentation: stackName is described as 'Which stack to diagnose' with a clear default behavior. The description adds no parameter-specific detail beyond the context of diagnosing a stack, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a diagnostic workflow for an unhealthy stack and details its internal pipeline: status check, gateway logs, agent logs, systemd service status, then a structured diagnostic report with remediation. This multi-step framing distinguishes it from simpler sibling tools like clawops_status or clawops_logs_tail, so the agent knows this is a diagnostic entry point, not a single-status command.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit 'Use when' and 'Do NOT use when' conditions: it is the best entry point when a 'not working' symptom is reported without a known cause, and it should not be used when a specific fix is already requested. It stops just short of naming the alternative sibling tools that would handle those specific cases, so it provides clear context but not complete routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
19 tool updates
v2.1.3- Changed
clawops_agents_list1 field changed- added
Input schema / properties / stackName / descriptionAdded value: +"Which stack's agents to list. Omitted = the default stack in ~/.clawops/config.json"
- Changed
clawops_apply2 fields changed- added
Input schema / properties / planPath / descriptionAdded value: +"Absolute path to plan JSON (R7)" - added
Input schema / properties / yes / descriptionAdded value: +"Skip the confirmation and apply immediately. Only when the user has already approved this specific plan"
- Changed
clawops_config_get2 fields changed- added
Input schema / properties / key / descriptionAdded value: +"Dot-path config key, e.g., gateway.auth.mode. Omit to dump the full config." - added
Input schema / properties / stackName / descriptionAdded value: +"Which stack's gateway config to read. Omitted = the default stack in ~/.clawops/config.json"
- Changed
clawops_config_set4 fields changed- added
Input schema / properties / key / descriptionAdded value: +"Dotted path into the gateway config, e.g. models.provider or channels.slack.enabled" - added
Input schema / properties / restart / descriptionAdded value: +"Restart the gateway so the change takes effect. Without it the value is written and the running gateway keeps the old one" - added
Input schema / properties / stackName / descriptionAdded value: +"Which stack's gateway config to write. Omitted = the default stack in ~/.clawops/config.json" - added
Input schema / properties / value / descriptionAdded value: +"Stringified; JSON for complex values"
- Changed
clawops_config_unset3 fields changed- added
Input schema / properties / key / descriptionAdded value: +"Dot-path config key to remove" - added
Input schema / properties / restart / descriptionAdded value: +"Restart the gateway so the removal takes effect. Without it the key is removed and the running gateway keeps the old value" - added
Input schema / properties / stackName / descriptionAdded value: +"Which stack's gateway config to write. Omitted = the default stack in ~/.clawops/config.json"
- Changed
clawops_config_validate1 field changed- added
Input schema / properties / stackName / descriptionAdded value: +"Which stack's deployed config to validate. Omitted = the default stack in ~/.clawops/config.json"
- Changed
clawops_destroy2 fields changed- added
Input schema / properties / stackName / descriptionAdded value: +"The stack to destroy. Required, and worth repeating back to the user before calling: this deletes every resource it provisioned" - added
Input schema / properties / yes / descriptionAdded value: +"Skip elicitation; CI/scripted use only"
- Changed
clawops_doctor2 fields changed- added
Input schema / properties / failuresOnly / descriptionAdded value: +"Return only failing and warning checks. Passing checks are counted, not listed." - added
Input schema / properties / stackName / descriptionAdded value: +"Stack to include remote checks for. Without it, only the local machine is\nchecked — no SSH connection is made.\n"
- Changed
clawops_gateway_restart1 field changed- added
Input schema / properties / stackName / descriptionAdded value: +"Which stack's gateway to restart. Omitted = the default stack in ~/.clawops/config.json"
- Changed
clawops_harden6 fields changed- added
Input schema / properties / dryRun / descriptionAdded value: +"Report the current state, change nothing" - added
Input schema / properties / options / descriptionAdded value: +"Comma-separated module IDs; default is every defaultOn module for the provider" - added
Input schema / properties / stackName / descriptionAdded value: +"Which stack to harden. Omitted = the default stack in ~/.clawops/config.json" - added
Input schema / properties / tailscale / descriptionAdded value: +"Join the tailnet, verify this machine reaches the host there, then use that address" - added
Input schema / properties / tailscaleRevert / descriptionAdded value: +"Leave the tailnet and go back to the public address" - added
Input schema / properties / yes / descriptionAdded value: +"Skip elicitation; CI/scripted use only"
- Added
clawops_init - Changed
clawops_logs_tail3 fields changed- added
Input schema / properties / sinceMin / descriptionAdded value: +"Lines since N minutes ago" - added
Input schema / properties / stackName / descriptionAdded value: +"Which stack's gateway to read. Omitted = the default stack in ~/.clawops/config.json" - added
Input schema / properties / tailLines / descriptionAdded value: +"How many of the most recent lines to return. Keep it small; output is trimmed to 8KB regardless"
- Changed
clawops_monitor2 fields changed- added
Input schema / properties / stackName / descriptionAdded value: +"Stack name. Defaults to active stack." - added
Input schema / properties / tailLines / descriptionAdded value: +"Log lines to include in snapshot."
- Changed
clawops_plan10 fields changed- added
Input schema / properties / gatewayCidr / descriptionAdded value: +"CIDR(s) allowed to reach the gateway port, or 'auto'. Requires publishGateway=all" - added
Input schema / properties / instanceType / descriptionAdded value: +"A clawops alias (micro|small|medium|large|gpu) or a machine type the cloud names itself, e.g. t3.small" - added
Input schema / properties / openclawVersion / descriptionAdded value: +"semver, or 'stable'/'dev'" - added
Input schema / properties / outPath / descriptionAdded value: +"Absolute path to write plan; if omitted, plan returned inline" - added
Input schema / properties / privateOnly / descriptionAdded value: +"Close public SSH and gateway access; reach the stack over its tailnet. Requires a verified tailnet address (clawops_harden with tailscale), and refuses unless that address answers SSH now" - added
Input schema / properties / provider / descriptionAdded value: +"Cloud to plan against. Omitted = the provider recorded for the stack. 'local' has no plan/apply path and is refused" - added
Input schema / properties / publishGateway / descriptionAdded value: +"Which interface the gateway binds. loopback (default) keeps it off the network" - added
Input schema / properties / region / descriptionAdded value: +"Cloud region, in the provider's own spelling (us-east-1, us-central1, eastus). Omitted = the region recorded for the stack" - added
Input schema / properties / sshCidr / descriptionAdded value: +"CIDR(s) allowed to reach SSH, comma-separated, or 'auto' for this machine. Omitted = none, and nothing will be able to connect" - added
Input schema / properties / stackName / descriptionAdded value: +"Which stack the plan is for. Omitted = the default stack in ~/.clawops/config.json"
- Changed
clawops_status1 field changed- added
Input schema / properties / stackName / descriptionAdded value: +"Stack name. Defaults to active stack from config."
- Changed
clawops_task_status1 field changed- added
Input schema / properties / taskId / descriptionAdded value: +"The taskId returned by a long-running tool such as clawops_up, clawops_apply or clawops_destroy"
- Changed
clawops_up9 fields changed- added
Input schema / properties / dryRun / descriptionAdded value: +"Show what would be created and change nothing. Use this first when the user has not yet approved a spend" - added
Input schema / properties / gatewayCidr / descriptionAdded value: +"CIDR(s) allowed to reach the gateway port. Requires publishGateway=all." - added
Input schema / properties / instanceType / descriptionAdded value: +"A clawops size (micro|small|medium|large|gpu) or a provider-native machine type.\nNot an enum: Azure offers SKU families per subscription, and an account offered\nnone of the five sizes clawops names would otherwise have no way to deploy.\n" - added
Input schema / properties / openclawVersion / descriptionAdded value: +"semver or 'stable'/'dev'" - added
Input schema / properties / provider / descriptionAdded value: +"Defaults to provider configured for this stack" - added
Input schema / properties / publishGateway / descriptionAdded value: +"Which interface the gateway binds. 'all' serves plaintext HTTP." - added
Input schema / properties / region / descriptionAdded value: +"Cloud region to deploy into, in the provider's own spelling (us-east-1, us-central1, eastus). Omitted = the region recorded for the stack" - added
Input schema / properties / sshCidr / descriptionAdded value: +"CIDR(s) allowed to reach SSH, comma-separated, or 'auto' for the caller's own\naddress. Omitted means none, and nothing will be able to connect — including\nevery clawops day-two command.\n" - added
Input schema / properties / stackName / descriptionAdded value: +"Name for the stack to provision, and the name every later command refers to it by. Omitted = the default stack in ~/.clawops/config.json"
- Changed
clawops_workflow_deploy_app4 fields changed- added
Input schema / properties / instanceType / descriptionAdded value: +"Machine size: a clawops alias (micro|small|medium|large|gpu) or a type the cloud names itself, e.g. t3.small" - added
Input schema / properties / provider / descriptionAdded value: +"Cloud to deploy to. Omitted = the default provider in ~/.clawops/config.json" - added
Input schema / properties / region / descriptionAdded value: +"Cloud region, in the provider's own spelling (us-east-1, us-central1, eastus). Omitted = the provider's default" - added
Input schema / properties / stackName / descriptionAdded value: +"Name for the new stack. Omitted = the default stack name in ~/.clawops/config.json"
- Changed
clawops_workflow_recover1 field changed- added
Input schema / properties / stackName / descriptionAdded value: +"Which stack to diagnose. Omitted = the default stack in ~/.clawops/config.json"
19 tool updates
v2.1.1- First observed
clawops_agents_list - First observed
clawops_apply - First observed
clawops_config_get - First observed
clawops_config_set - First observed
clawops_config_unset - First observed
clawops_config_validate - First observed
clawops_destroy - First observed
clawops_doctor - First observed
clawops_gateway_restart - First observed
clawops_harden - First observed
clawops_logs_tail - First observed
clawops_monitor - First observed
clawops_plan - First observed
clawops_stacks_list - First observed
clawops_status - First observed
clawops_task_status - First observed
clawops_up - First observed
clawops_workflow_deploy_app - First observed
clawops_workflow_recover
TDQS
Scored across 20 tools
Most tools target distinct actions on stacks, config, the gateway, or workflows, and the use/don't-use notes are unusually explicit. The main risk is the observability cluster (status, doctor, monitor, logs_tail, workflow_recover), where an agent may need to think before choosing, but the descriptions mostly separate them.
All tool names share the clawops_ prefix, use snake_case, and many follow a resource_action shape such as stacks_list, config_get, or gateway_restart. The pattern is weakened by action-only names like up, plan, apply, doctor, and status, but the overall style is predictable and readable.
20 tools is on the heavy side for a single server, especially with two workflow-wrapping tools and multiple observability tools that partly overlap. That said, the lifecycle, config, and diagnostic groups are each justified, so it is not unmanageably bloated.
The surface covers stack lifecycle (init/plan/apply/up/destroy), config CRUD and validation, gateway restart, hardening, health checks, logs, and status. Obvious gaps like stack refresh, backup, streaming logs, and gateway update are explicitly deferred to the CLI, which agents can work around by telling the user to run those commands.
Maintenance
Related MCP Connectors
Governed MCP gateway: one endpoint for your tools, with credential custody and audit log.
MCP server for AI agents to plan, verify, and deploy Cloudflare-native apps.
Devopness MCP server for DevOps happiness! Empower AI Agents to deploy apps and infra, to any cloud.
MCP-first control plane for ProAgentStore agents and private instances.
Related MCP Servers
- AlicenseBqualityDmaintenanceAn MCP server that enables Claude to manage infrastructure across Kubernetes, Docker, Prometheus, and Terraform through natural language. It provides over 42 specialized tools with a safety-first design, including risk-based command classification and audit logging.43MIT

Infraveil MCP serverofficial
AlicenseAqualityBmaintenanceA hardened, self-hosted MCP server that lets AI agents query and govern an Infraveil control plane in-loop, reading state and filing deploy/remediation requests that require human approval.7AGPL 3.0- AlicenseNot gradedqualityDmaintenanceEnables infrastructure operations including provisioning, configuration, monitoring, compliance auditing, and auto-remediation through natural language, using Terraform and Ansible tools exposed over MCP.MIT
- AlicenseNot gradedqualityDmaintenanceAn AI-DevOps MCP server that gives LLMs read-only-by-default access to Kubernetes clusters, Prometheus metrics, and GitHub Actions, enabling natural language queries about infrastructure status and safe write operations with previews.MIT