Skip to main content
Glama
brew install Higangssh/homebutler/homebutler     # or: curl -fsSL https://raw.githubusercontent.com/Higangssh/homebutler/main/install.sh | sh
homebutler report                                # first run saves a baseline; the second tells you what moved

Or as a container, reaching your machines over SSH — docs/docker.md: docker run -v ~/.config/homebutler:/config -v ~/.ssh:/root/.ssh:ro ghcr.io/higangssh/homebutler

Section rules, labels, and severities are colour-coded in a terminal. Colour is dropped automatically when output is piped, redirected, or run from cron.

That is the whole idea. Most homelab tools show you a graph of right now, and leave "does this matter?" to you. HomeButler remembers what your server looked like last time, decides what is worth saying, and says it — six containers before and six after is not "no change" when one of them is a different container.

Reading a change

Every line is three columns: what kind of change, what it happened to, and what exactly happened. The kind is one of eight words, and it is the same word in --json, so an agent branches on it without reading prose:

Kind

Means

You would see it after

gone

it was there last time and is not now

docker rm, a service stopping, a port closing

new

it was not there last time and is now

starting anything

replaced

same name, different thing underneath

docker compose up -d — the container is recreated, so the name and the count are unchanged

image

same container, different image

pulling a new tag

state

same container, running where it was stopped, or the reverse

a crash, or bringing something back up

port

same port, a different process answering on it

one service taking over another's port

disk

a mount moved by more than half a gigabyte

anything that writes

skipped

the comparison could not be made

Docker was down when either snapshot was taken

replaced is the one the rest of this exists for. A container recreated under the same name leaves every count identical, which is why a report that compares counts — as this one did before 0.26.0 — answers "no significant changes" while the thing you were running has been swapped out underneath you.

skipped is the second: homebutler says it could not compare rather than reporting nothing changed. An all-clear it cannot stand behind is worse than no answer.

The header names the snapshot being compared against, so "what changed" is never ambiguous about the window it covers.

📖 What earns a line, and what is deliberately suppressed →

HomeButler helps you answer the boring but painful questions every homelab eventually creates:

  • What is running on my server right now?

  • Which container owns this port?

  • Why did this service restart at 3 AM?

  • Is my backup actually restorable?

  • Can I install this self-hosted app without hand-writing another compose file?

  • Can I let an AI assistant inspect my server without handing it a full SSH shell?

No daemon required, no database. Run it from a terminal, cron, a web dashboard you start when you want one, or an AI agent.

The design goal is simple: give humans and agents a narrow, structured interface to the server. HomeButler returns readable summaries and JSON instead of asking you to trust a black-box shell session.

Quick Start

# One-line install (auto-detects OS/arch)
curl -fsSL https://raw.githubusercontent.com/Higangssh/homebutler/main/install.sh | sh

# Or via Homebrew
brew install Higangssh/homebutler/homebutler

# Or as a container — the hub that reaches your machines over SSH
docker run -d -p 8080:8080 -v ~/.config/homebutler:/config -v ~/.ssh:/root/.ssh:ro \
  ghcr.io/higangssh/homebutler serve --host 0.0.0.0 --token "$(openssl rand -hex 16)"

# Interactive setup — add your servers in seconds
homebutler init

Use it right away:

homebutler status                    # CPU, memory, disk, uptime
homebutler docker list               # running containers
homebutler inventory scan            # containers + ports + topology
homebutler report                    # butler-style health report + change summary
homebutler install uptime-kuma       # deploy a self-hosted app
homebutler backup drill uptime-kuma  # verify a backup actually restores
homebutler watch tui                 # terminal dashboard
homebutler serve                     # web dashboard at http://localhost:8080

Machine-readable output is available everywhere:

homebutler status --json
homebutler inventory scan --json
homebutler report --json

Related MCP server: Plutus MCP Server

What it does

  • Install apps — deploy Uptime Kuma, Jellyfin, Pi-hole, Gitea, Portainer, and more with one command

  • Map your server — see containers, exposed ports, system ports, and service topology

  • Run a doctor check — diagnose resource pressure, stopped containers, public ports, backup hygiene, notifications, report baseline readiness, and configured Proxmox endpoint reachability

  • Catch crashes — save logs before/after Docker, systemd, or PM2 restarts and detect flapping loops

  • Verify backups — boot backups in isolated containers before you trust them

  • See a Proxmox cluster — nodes, QEMU and LXC guests, storage, task status, and honest dashboard freshness, with power actions that name their target explicitly

  • Use it anywhere — CLI, JSON, web dashboard, or MCP for AI agents without giving them SSH

Why homebutler?

Self-hosting is not hard because one docker compose up is hard. It is hard because the maintenance never ends: ports collide, containers restart silently, backups look fine until restore day, and every server becomes a slightly different snowflake.

HomeButler is a small operations toolkit for that messy middle.

Alongside what you already run

Keep Uptime Kuma for is it up, Beszel or Netdata for the graph, Dozzle for the logs. HomeButler answers the question none of them ask: what is different from last time, and does it matter?

  • No agent on the machines it watches. The binary is put there once with homebutler deploy and runs only when asked, over SSH — no daemon, no open port, nothing listening between runs.

  • The judgement is a written rule, not a model. What earns a line is in docs/report.md, and the same input gives the same report — no AI required, no account, no paid tier.

  • It checks that a backup comes back. A backup you have never restored is a folder. backup drill boots the archive as a second copy of the app, on a network and port of its own, and requires it to answer an HTTP health check before calling the archive good — then removes everything it made. A failed drill exits non-zero, so backup drill --all belongs in the same cron entry as backup. docs/backup.md has what it does and what it does not prove.

    🔐 Integrity: ✅ tar valid (8 files)
    🚀 Boot: ✅ container started in 0s
    🌐 Health: ✅ HTTP 200 on port 60405
    
    ✅ DRILL PASSED

Core workflows

🧾 Butler Report

homebutler report
homebutler report --keep 7      # retain only the latest 7 snapshots
homebutler report --no-save     # preview without writing a snapshot

report gives you a concise butler-style summary of your homelab: current health, warnings, notable changes since the previous snapshot, and suggested next commands. On the first run, HomeButler creates a baseline under ~/.homebutler/reports/snapshots/; later runs compare against the latest snapshot. Old snapshots are pruned automatically (--keep 30 by default) so reports do not grow forever.

🩺 Doctor Check

homebutler doctor
homebutler doctor --strict          # non-zero exit if warnings/failures are found
homebutler doctor --json            # automation / MCP friendly

doctor is a read-only preflight for the problems homelab users usually discover too late: high disk or memory usage, stopped containers, public bind ports, stale or missing backups, missing notifications, whether report has a baseline for change detection, and whether each configured Proxmox endpoint is reachable with the token it has. Every finding names the next command to run, so --strict makes it usable from cron or CI — including a Proxmox host that is unreachable or rebooting.

Every finding carries a category, and --json carries the same word, so a caller filters on it without reading the title:

Category

What it is about

system

CPU, memory, disk

docker

containers that are stopped or unhealthy

exposure

ports listening on every interface

backup

backups missing, stale, never drilled, or growing without a retention limit

report

whether there is a baseline to compare against

watch

targets listed with nothing polling them

notifications

channels configured, or never tested

proxmox

an endpoint that cannot be reached with the token it has

config

the config file itself — permissions, keys homebutler does not know

incomplete

a collector did not answer, so this diagnosis is partial

overall

nothing specific: the all-clear

incomplete is the one worth handling. It means the answer you are reading is missing a section, which is different from that section being fine.

🗂 Config Validation

homebutler config validate
homebutler config validate --strict   # exit non-zero on warnings too
homebutler config validate --json

config validate reads your config without starting anything and tells you which file was used, which of the four resolution rules picked it, and what homebutler actually made of each section. It exists because the two ways config goes wrong are both silent: a key homebutler does not recognise is dropped without a word, and a --config path that does not exist falls back to built-in defaults rather than failing.

Sections
   ✓ servers     2 servers (homelab, nas)
   · notify      not set
   ✓ alerts      cpu 95% · memory 85% · disk 90%

Findings
   ⚠️ Line 5: field notifiy not found in the homebutler config
      → Did you mean "notify"? Unrecognised keys are ignored silently.

📦 One-Command App Install

$ homebutler install list
📦 Available apps (15):

  uptime-kuma          A self-hosted monitoring tool
  vaultwarden          Lightweight Bitwarden-compatible password manager
  jellyfin             Free software media system for streaming movies, TV, and music
  pi-hole              Network-wide ad blocking via DNS filtering
  …

homebutler install uptime-kuma — Deploy self-hosted apps in seconds. Pre-checks Docker, ports, and duplicates. Generates docker-compose.yml automatically. See all available apps →

🗺️ Inventory & Topology

homebutler inventory scan
homebutler inventory show --filter exposed
homebutler inventory export --format mermaid
homebutler --json inventory scan

inventory scan gives you a quick map of what is running on a server: system health, Docker containers, app ports, and system ports. Docker-published ports are connected back to the container that owns them, so local forwarding details like Colima/Lima stay understandable.

🏠 Home Network
   Server  homelab (192.168.1.10)
   Summary ✅ 1 running · ⚪ 1 stopped · 🌍 2 public ports · 🔒 4 local ports

📦 Containers (2)
   ├─ ⚪ vaultwarden · not started
   │  └─ image vaultwarden/server:latest
   └─ ✅ api-server · running
      ├─ image my-api:latest
      └─ exposes :8080 → 8080/tcp

🌐 App Ports (1)
   └─ 🌍 :8080/tcp · api-server

To answer "what is reachable from outside my machine/network?" without reading the whole tree, filter the scan to exposed ports only:

homebutler inventory scan --filter exposed
🏠 Home Network
   Server  homelab

🌐 Exposed Ports
   ├─ :8080/tcp · api-server
   └─ :8443/tcp · dashboard

Only ports listening on all interfaces (0.0.0.0, ::, *) are shown. Anything bound to a specific address is hidden, including loopback and LAN addresses. Unsupported filter values return an error, as does combining --filter with --json; the default inventory scan output is unchanged.

Use Mermaid export when you want a diagram for GitHub, Obsidian, docs, or an AI assistant:

graph TD
  home["🏠 Home Network"] --> homelab["🖥 homelab<br/>192.168.1.10"]
  homelab --> c1["📦 api-server<br/>running"]
  homelab --> p1["🌍 :8080/tcp<br/>api-server"]
  c1 -. exposes .-> p1

Demo

🌐 Web Dashboard

homebutler serve — A web dashboard embedded in the single binary via go:embed. Servers, Docker containers, open ports, alerts and Wake-on-LAN from any browser. With --token it also edits the config: thresholds, notification channels, servers, Proxmox endpoints and wake devices.

The Report tab, where you actually read it

The screen somebody opens after a notification. What needs attention first, then what moved with the kind word each change carries, then the snapshot being compared against and how old it is. A comparison that could not be made says so rather than reading as an all-clear.

Loading it saves nothing — a page that took a snapshot every time it was opened would prune the baseline you wanted to compare against. Saving one is a button.

  • Server Overview — See all servers at a glance with color-coded status (green = online, red = offline)

  • System Metrics — CPU, memory, disk usage with progress bars and color thresholds

  • Docker Containers — Running/stopped status with friendly labels ("Running · 4d", "Stopped · 6h ago")

  • Top Processes — Top processes sorted by CPU/memory with zombie detection

  • Resource Warnings — Visual CPU, memory, and disk thresholds in the dashboard

  • Network Ports — Open ports with process names and bind addresses

  • Wake-on-LAN — One-click wake buttons for configured devices

  • Server Switching — Dropdown to switch between local and remote servers

  • Zero dependencies — No Node.js runtime needed. Frontend is compiled into the Go binary at build time

homebutler serve              # Start on port 8080
homebutler serve --port 3000  # Custom port
homebutler serve --demo       # Demo mode with realistic sample data
homebutler serve install      # Let the host keep it running

🔄 Process Restart Watch

Your container crashed at 3 AM — but why? homebutler watch catches it the moment it happens, saves the dying logs, figures out the cause, and tells you if it's happening over and over.

Supported backends: Docker (real-time event stream) · systemd (polling) · PM2 (polling)

Step 1: Add targets to watch

homebutler watch add nginx              # Interactive: choose Docker / systemd / PM2
homebutler watch add --kind docker nginx          # or specify directly
homebutler watch add --kind systemd nginx.service
homebutler watch add --kind pm2 my-api
homebutler watch list                   # See what you're watching

Step 2: Start monitoring

homebutler watch start                  # Foreground, Ctrl+C to stop
homebutler watch start --interval 10s   # Custom poll interval (default 30s)
homebutler watch install     # register it with systemd or launchd
homebutler watch installed   # is it registered?
homebutler watch uninstall

watch install hands the loop to whatever supervises the host — a systemd user unit on Linux, a launchd agent on macOS — so monitoring survives logout and reboot. Both are user-level and neither is a preference: on Linux the watch list lives in your home directory, so a root unit would find an empty list; on macOS Docker Desktop only runs inside a logged-in session, so a LaunchDaemon would poll a daemon that is not there. On Linux a user unit stops at logout unless lingering is on; watch install checks, and if it is off tells you to run loginctl enable-linger $USER (with sudo only if your system refuses it).

watch start is the monitoring process. It watches the containers and services on the watch list for restarts, checks CPU, memory and disk against your thresholds, and runs any remediation rules you have configured — one process, one set of notification providers. alerts --watch still exists and does the threshold half on its own.

Every endpoint under proxmox: in your config is polled too: unreachable or ACL-filtered endpoints and any guest listed under that endpoint's guests: report one incident when the problem starts and one recovery incident when it clears. A guest not listed there is observational only — watch start never alerts on it, deliberately stopped or not. See Proxmox setup → for the guests: field.

When a crash is detected, you'll see:

[03:14:22] INCIDENT: nginx (incident nginx-20260410-031422.581-7a2124)
  Crash: OOM — process killed by SIGKILL (oom, confidence: high)
  ⚠ FLAPPING: acute (3 restarts in short window)

Step 3: Investigate

homebutler watch history                # List all incidents
homebutler watch show <incident-id>     # Full details

watch show output includes:

  • Pre-death logs — what the process printed right before it died

  • Post-restart logs — what happened after the restart

  • Crash analysis — category (oom / panic / segfault / timeout / dependency / error), reason, confidence level, matched log patterns

  • Flapping status — if the process is stuck in a crash loop

Crash Analysis

Every incident is automatically analyzed using exit codes and log patterns:

Signal

Exit Code

Meaning

SIGKILL

137

OOM Killer or forced kill

SIGSEGV

139

Segmentation fault (memory corruption)

SIGTERM

143

Graceful shutdown request

—

1

Application error

—

0

Clean exit (may be intentional restart)

Log patterns like panic:, Out of memory, Connection refused, FATAL, and timeout are matched automatically to help identify the root cause.

Flapping Detection

Detects when a process is stuck in a restart loop (e.g., crash → restart → crash again):

  • Acute — 3+ restarts within 10 minutes (something is broken right now)

  • Chronic — 5+ restarts within 24 hours (slow recurring issue)

Flapping incidents are tagged [FLAPPING] in history and highlighted in watch show.

Notifications (optional, off by default)

Notifications are disabled by default, which is useful for air-gapped or closed networks where everything runs locally.

A minimal example in ~/.config/homebutler/config.yaml:

notify:
  telegram:
    bot_token: "your-bot-token"
    chat_id: "your-chat-id"

watch:
  enabled: true
  notify_on: flapping
  cooldown: 5m
  flapping:
    short_window: 10m
    short_threshold: 3
    long_window: 24h
    long_threshold: 5
  retention:
    max_incidents: 200

alerts:
  cpu: 90
  memory: 85
  disk: 90
  rules:
    - name: cpu-spike
      metric: cpu
      threshold: 90
      action: notify

    - name: elsa-monitor-down
      metric: container
      kind: systemd          # docker (default) | systemd | pm2
      watch: [lh-elsa-monitor.service]
      action: restart

Restarting things that are not containers

action: restart restarts Docker containers unless the rule says otherwise. kind: systemd or kind: pm2 points it at a service or a PM2 app instead.

The kind is written on the rule rather than looked up from the watch list, so restarting a host service is something you asked for in the config. It also means every rule written before kind existed keeps meaning exactly what it meant.

Two things worth knowing before using it:

systemctl restart needs root or a polkit rule. Running homebutler unprivileged, a systemd restart will be refused, reported as failed, and warned about when alerts --watch starts rather than when the rule first fires.

A target that is flapping is not restarted. Restarting something already in a restart loop feeds the loop, and most systemd units carry Restart=always, so homebutler restarting them fights systemd's own backoff. The thresholds are the watch.flapping ones above, and the skip is reported rather than counted as either success or failure. This applies to Docker targets too.

Legacy ~/.homebutler/watch/config.json is still read as a fallback for watch-specific settings, and legacy alerts.yaml notify/webhook provider settings are still accepted for older setups.

  • watch.enabled: true — allow watch notifications

  • watch.notify_on: flapping — notify only when repeated restart loops are detected

  • watch.notify_on: incident — notify on every incident

  • watch.notify_on: all — notify on both incidents and flapping

  • watch.notify_on: off — disable watch notifications without removing provider config

  • watch.cooldown: 5m — suppress duplicate notifications for the same event fingerprint during the cooldown window

  • watch.flapping — optional advanced tuning for restart-loop detection

  • watch.retention.max_incidents: 200 — how many incidents to keep on disk, newest first. The directory grows fastest exactly when a service is restarting in a loop. Set -1 to keep everything.

    Each incident keeps up to 100 captured log lines per side, and at most 64 KB of them. Line counts alone do not bound a file: one stack trace or JSON document on a single line is arbitrarily long, and a container being OOM-killed is exactly the one likely to write one. A log that does not fit keeps its end — the last thing a process said is what explains why it stopped — and says how much was dropped.

These settings can also be written under a watch.notify: block, which is the canonical form:

watch:
  notify:
    enabled: true
    notify_on: flapping
    cooldown: 5m
  flapping:
    short_window: 10m

Both spellings are read, so either layout works. If a file contains both, the notify: block wins and homebutler config validate says so.

Manage targets

homebutler watch remove nginx           # Stop watching
homebutler watch check                  # One-shot check (no continuous monitoring)

🧊 Proxmox VE

homebutler proxmox status
homebutler proxmox guests --status running
homebutler proxmox guest shutdown --node pve1 --type lxc --vmid 105 --confirm
homebutler proxmox task UPID:pve1:... --node pve1

A Proxmox endpoint is its own kind of target, configured under proxmox: with an API token rather than SSH, so it does not join the --server or --all fan-out. TLS verification stays on: trust comes from a pinned SHA-256 fingerprint, then a CA file, and only then an explicit insecure fallback.

Reads are plain. Power actions are not: every one of them takes an explicit endpoint, node, guest type and VMID, and refuses to run without --confirm, which is checked before any credential is read. They also need their own action_token_id (plus action_token or action_token_file) configured on the endpoint — the read token alone will not start, reboot, or shut down a guest; see Proxmox setup → for creating that second token. shutdown asks the guest to shut down cleanly — it is not Proxmox's hard stop, which cuts power and can leave a filesystem behind it. A successful action reports the task it submitted, not that the guest finished; proxmox task answers that separately.

proxmox script prints the install command for a Community Script pinned to one commit, along with a warning that the script is not reviewed by homebutler and runs as root. It never fetches or runs it — see #62 for why that line is where it is.

📖 Proxmox setup, tokens, and TLS →

🖥️ TUI Dashboard

homebutler watch tui — A terminal-based dashboard powered by Bubble Tea. Monitors all configured servers with real-time updates, color-coded resource bars, and Docker container status. No browser needed.

🧠 AI-Powered Management (MCP)

Use natural language when you want automation. MCP clients can call homebutler tools to check server status, list Docker containers, inspect ports, or run operational workflows. See screenshots & setup →

App Install

Deploy self-hosted apps with a single command. Each app runs via docker compose with automatic pre-checks, health verification, and clean lifecycle management.

# List available apps
homebutler install list

# Install (default port)
homebutler install uptime-kuma

# Install with custom port
homebutler install uptime-kuma --port 8080

# Install jellyfin with media directory
homebutler install jellyfin --media /mnt/movies

# Check status
homebutler install status uptime-kuma

# Stop (data preserved)
homebutler install uninstall uptime-kuma

# Stop + delete everything
homebutler install purge uptime-kuma

How it works

~/.homebutler/apps/
  └── uptime-kuma/
       ├── docker-compose.yml   ← auto-generated, editable
       └── data/                ← persistent data (bind mount)
  • Pre-checks — Verifies docker is installed/running, port is available, no duplicate containers

  • Compose-based — Each app gets its own docker-compose.yml you can inspect and customize

  • Data safety — uninstall stops containers but keeps your data; purge removes everything

  • Cross-platform — Auto-detects docker socket (default, colima, podman)

Available apps

App

Default Port

Description

Notes

uptime-kuma

3001

Self-hosted monitoring tool

plex

32400

Plex Media Server

--media /path to mount media dir

vaultwarden

8080

Bitwarden-compatible password manager

filebrowser

8081

Web-based file manager

it-tools

8082

Developer utilities (JSON, Base64, Hash, etc.)

gitea

3002

Lightweight self-hosted Git service

jellyfin

8096

Media system (movies, TV, music)

--media /path to mount media dir

homepage

3010

Modern homelab dashboard

stirling-pdf

8083

All-in-one PDF tool (merge, split, convert, OCR)

speedtest-tracker

8084

Internet speed test with historical graphs

mealie

9925

Recipe manager and meal planner

pi-hole

8088

DNS ad blocking

⚠️ Uses port 53 (DNS), NET_ADMIN capability

adguard-home

3000

DNS ad blocker and privacy

⚠️ Uses port 53 (DNS)

portainer

9443

Docker management GUI

⚠️ Mounts Docker socket (HTTPS)

nginx-proxy-manager

81

Reverse proxy with SSL and web UI

⚠️ Uses ports 80/443

App-specific options

# Jellyfin: mount your media library
homebutler install jellyfin --media /mnt/movies

# Pi-hole / AdGuard: DNS ad blocking (port 53 required)
homebutler install pi-hole
# ⚠️ If port 53 is in use (Linux): sudo systemctl disable --now systemd-resolved

# Portainer: Docker GUI (mounts docker socket)
homebutler install portainer
# Access via HTTPS: https://localhost:9443

# Nginx Proxy Manager: reverse proxy
homebutler install nginx-proxy-manager
# Default login: admin@example.com / changeme (change immediately!)

# Any app: custom port
homebutler install <app> --port 9999

Safety checks

  • Port conflict detection — Checks if the port is already in use before install

  • DNS mutual exclusion — Warns if pi-hole and adguard-home are both installed

  • Docker socket warning — Alerts when an app requires Docker socket access (portainer)

  • OS-specific guidance — Linux gets systemd-resolved fix, macOS gets lsof command

  • Post-install tips — DNS setup, HTTPS access, default credential warnings

Want more apps? Open an issue or see Contributing.

Usage

homebutler <command> [flags]

Commands:
  status              System status (CPU, memory, disk, uptime)
  doctor              Diagnose health, exposure, backups, and readiness
  config validate     Check the config file and report what is ignored
  docker list         List running containers
  install <app>       Install a self-hosted app (docker compose)
  alerts              Show current alert status
  watch tui           TUI dashboard (monitors all configured servers)
  watch add/list/remove  Manage watched containers
  watch check/start   One-shot or continuous restart detection
  watch history/show  Browse restart history
  proxmox status      Proxmox VE cluster, nodes, guests, and storage
  serve               Web dashboard (browser-based, go:embed)

Flags:
  --json              JSON output (default: human-readable)
  --verbose, -v       Show detailed error information
  --server <name>     Run on a specific remote server
  --all               Run on all configured servers in parallel
  --port <number>     Port for serve command (default: 8080)
  --config <path>     Config file (auto-detected, see Configuration)

Run homebutler --help for all commands.

Commands:
  init                Interactive setup wizard
  config validate     Check the config file and report what is ignored
  status              System status (CPU, memory, disk, uptime)
  doctor              Diagnose health, exposure, backups, and readiness
  watch tui           TUI dashboard (monitors all configured servers)
  watch add <name>    Add container to restart watch list
  watch list          Show watched containers
  watch remove <name> Remove container from watch list
  watch check         One-shot restart check
  watch start         Continuous monitoring: restarts, thresholds, rules
  watch install       Register watch with systemd or launchd
  watch installed     Report whether it is registered
  watch uninstall     Remove the service unit
  watch history       List restart history (alias: incidents)
  watch show <id>     Show restart details with logs
  serve               Web dashboard (browser-based, go:embed)
  docker list         List running containers
  docker restart <n>  Restart a container
  docker stop <n>     Stop a container
  docker logs <n>     Show container logs
  docker top <n>      Show processes running inside a container
  docker inspect <n>  Show image, state, ports, mounts, networks, health
  report              What changed since the last snapshot
  inventory scan      Map containers, ports, and topology
  inventory show      Same as scan (--filter exposed narrows it)
  inventory export    Export the map (--format mermaid)
  proxmox status      Proxmox VE cluster, nodes, guests, storage
  proxmox guests      List QEMU and LXC guests
  proxmox node <n>    Node detail
  proxmox guest ...   start / shutdown / reboot (needs --confirm)
  proxmox task <upid> Task status for an action already submitted
  proxmox tasks       Recent tasks on a node
  proxmox script      Community Script install commands (prints, never runs)
  notify test         Send a test notification through configured providers
  wake <name>         Send Wake-on-LAN packet
  ports               List open ports with process info
  ps                  Show top processes (alias: processes)
  ps --sort mem       Sort by memory instead of CPU
  ps --limit 20       Show top 20 (default: 10, 0 = all)
  network scan        Discover devices on LAN
  alerts              Show current alert status
  alerts --watch      Thresholds only (watch start covers these too)
  trust <server>      Register SSH host key (TOFU)
  backup              Backup Docker volumes, compose files, and env
  backup list         List existing backups
  backup drill <app>  Verify backup restores correctly (isolated)
  backup drill --all  Verify all apps in backup
  restore <archive>   Restore from a backup archive
  upgrade             Upgrade local + all remote servers to latest
  deploy              Install homebutler on remote servers
  install <app>       Install a self-hosted app (docker compose)
  install list        List available apps
  install status <a>  Check installed app status
  install uninstall   Stop app (keep data)
  install purge       Stop app + delete all data
  mcp                 Start MCP server (JSON-RPC over stdio)
  version             Print version

Flags:
  --json              JSON output (default: human-readable)
  --verbose, -v       Show detailed error information
  --server <name>     Run on a specific remote server
  --all               Run on all configured servers in parallel
  --port <number>     Port for serve command (default: 8080)
  --demo              Run serve with realistic demo data
  --watch             Continuous monitoring mode (alerts command)
  --interval <dur>    Watch interval, e.g. 30s, 1m (default: 30s)
  --config <path>     Config file (auto-detected, see Configuration)
  --local             Upgrade only the local binary (skip remote servers)
  --local <path>      Use local binary for deploy (air-gapped)
  --service <name>    Target a specific Docker service (backup/restore)
  --allow-bind <path> Host path a restore may write a bind mount to (repeatable)
  --endpoint <name>   Proxmox endpoint from config (optional if only one)
  --confirm           Required for a Proxmox guest power action
  --to <path>         Custom backup destination directory
  --archive <path>    Specific backup archive for drill
  --all               Verify all supported apps (backup drill)

homebutler serve starts an embedded web dashboard — no Node.js, no Docker, no extra dependencies.

homebutler serve                # http://localhost:8080
homebutler serve --port 3000    # custom port
homebutler serve --demo         # demo mode with sample data
homebutler serve install        # hand it to launchd or systemd, so it survives logout
homebutler serve uninstall      # and take it back

The installed unit records the address and never the token — that comes from web.token in the config file, because --token is visible in ps to every user on the machine.

📖 Web dashboard details →

Backup & Restore

One-command Docker backup — volumes, compose files, and env variables.

homebutler backup                          # backup everything
homebutler backup --service jellyfin       # specific service
homebutler backup --to /mnt/nas/backups/   # custom destination
homebutler backup list                     # list backups
homebutler restore ./backup.tar.gz         # restore

⚠️ Database services should be paused before backup for data consistency.

📖 Full backup documentation → — how it works, archive structure, security notes.

Alert Thresholds (Advanced)

alerts still exists for CPU, memory, and disk threshold checks, but it is an advanced flow and not the recommended first step for new users.

homebutler alerts --watch                  # default: 30s interval
homebutler alerts --watch --interval 10s   # check every 10 seconds
homebutler alerts history                  # view alert history
homebutler notify test                     # test your notification channels

Default thresholds: CPU 90%, Memory 85%, Disk 90%. Start with watch, then add alerts only if you specifically want threshold-based checks.

🔍 Backup Drill

"Having a backup" and "being able to restore" are different things.

Backup Drill boots your backup in an isolated Docker environment and verifies the app actually responds — like a fire drill for your data. A failed drill exits non-zero.

homebutler backup drill uptime-kuma        # verify one app
homebutler backup drill --all              # verify all apps
homebutler backup drill --json             # machine-readable output
homebutler backup drill --archive ./file   # use a specific backup

Full walkthrough, including what a passing drill does not prove: docs/backup.md.

What happens:

  1. Finds the latest backup archive

  2. Verifies archive integrity (tar validation)

  3. Creates an isolated Docker network + random port

  4. Boots the app from backup data

  5. Runs an HTTP health check

  6. Reports pass/fail and cleans up everything

🔍 Backup Drill — uptime-kuma

  📦 Backup: ~/.homebutler/backups/backup_2026-04-04_1711.tar.gz
  📏 Size: 18.6 MB
  🔐 Integrity: ✅ tar valid (8 files)

  🚀 Boot: ✅ container started in 0s
  🌐 Health: ✅ HTTP 200 on port 58574
  ⏱️  Total: 2s

  ✅ DRILL PASSED

Zero risk — runs in a completely isolated environment. Your running services are never touched.

Supports health checks for: nginx-proxy-manager, vaultwarden, uptime-kuma, pi-hole, gitea, jellyfin, plex, portainer, homepage, adguard-home.

Configuration

homebutler init    # interactive setup wizard

📖 What report compares → — what earns a line, what is deliberately suppressed, and why.

📖 Configuration details → — config file locations, watch/notify options, and advanced alert thresholds.

Multi-server

Manage multiple servers from a single machine over SSH.

homebutler status --server rpi     # query specific server
homebutler status --all            # query all in parallel
homebutler deploy --server rpi     # install on remote server
homebutler upgrade                 # upgrade all servers

📖 Multi-server setup → — SSH auth, config examples, deploy & upgrade.

MCP Server

Built-in MCP server — an agent gets the same report as typed JSON, kind, target and detail rather than a sentence to parse, and every tool is classed read, write or destructive.

{
  "mcpServers": {
    "homebutler": {
      "command": "npx",
      "args": ["-y", "homebutler@latest"]
    }
  }
}

Works with Claude Desktop, ChatGPT, Cursor, Windsurf, and any MCP client.

📖 MCP server setup → — supported clients, available tools, agent skills.

Installation

brew install Higangssh/homebutler/homebutler

Automatically installs to PATH. Works on macOS and Linux.

One-line Install

curl -fsSL https://raw.githubusercontent.com/Higangssh/homebutler/main/install.sh | sh

Auto-detects OS/architecture, downloads the latest release, and installs to PATH.

npm (MCP server)

npm install -g homebutler

Downloads the Go binary automatically. This package publishes one command, homebutler-mcp, and it starts the MCP server — npx -y homebutler@latest launches that, not the CLI. For the command line, use one of the other installs above.

Go Install

go install github.com/Higangssh/homebutler@latest

Build from Source

git clone https://github.com/Higangssh/homebutler.git
cd homebutler
make build-all

make build-all compiles the web dashboard into the binary and needs Node installed. make build skips it — the CLI is complete either way, and homebutler serve then says the dashboard is missing and names the two ways to get it.

Uninstall

rm $(which homebutler)           # Remove binary
rm -rf ~/.config/homebutler      # Remove config (optional)

Architecture

The butler watches, decides what is worth saying, and says it — to you or to an agent.

Something changes → report names it and ranks it → a notification, the dashboard, or an MCP call carries the same answer.

        ┌─────────┐  ┌─────────┐  ┌─────────┐
        │   CLI   │  │   MCP   │  │   Web   │
        │ stdout  │  │  stdio  │  │  :8080  │
        └────┬────┘  └────┬────┘  └────┬────┘
             └────────────┼────────────┘
                          ▼
                   internal/*
         system · docker · ports · network
         wake · alerts · remote (SSH)

Three interfaces, one core:

Interface

Transport

Use case

CLI

Shell stdout/stderr

Terminal, scripts, AI agents via exec

MCP

JSON-RPC over stdio

Claude Desktop, ChatGPT, Cursor, any MCP client

Web

HTTP (go:embed)

Browser dashboard, on-demand with homebutler serve

All three call the same internal/ packages — no code duplication.

An agent reads the same report a person does, and gets the parts it has to branch on rather than a screen it has to interpret:

report = json.loads(subprocess.run(
    ["homebutler", "report", "--json"], capture_output=True, text=True).stdout)

for change in report["notable_changes"]:
    if change["kind"] == "replaced":         # the container came back as something else
        redeploy(change["target"])
    else:
        log(change["text"])                  # "port: :8080/tcp — nginx → caddy"

Every line carries text as well, so a caller that only wants to print it does not have to put the sentence back together. A run that found nothing returns "notable_changes": [], so the loop above does nothing on a quiet day.

Nothing above this is homebutler's business: an MCP client, a chat bot, a cron line or a person at a terminal all reach the same answer.

No ports opened by default. CLI and MCP use stdin/stdout only. The web dashboard is opt-in (homebutler serve, binds 127.0.0.1).

Contributing

Contributions welcome! Please open an issue first to discuss what you'd like to change. CONTRIBUTING.md covers what homebutler accepts and what a new target has to prove.

Security

Found a vulnerability? Report it privately through the Security tab rather than a public issue. SECURITY.md covers what is in scope and what to expect.

License

MIT

Available Tools

44 tools
alertsA

Check resource alerts for CPU, memory, and disk usage against configured thresholds. It reads and compares; nothing is sent anywhere and nothing is recorded, which is what a running watcher does instead

ParametersJSON Schema
NameRequiredDescriptionDefault
serverNoRemote server name from config (optional, runs locally if omitted)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It explicitly discloses that the tool only reads and compares and has no side effects: nothing is sent anywhere and nothing is recorded. It does not describe output shape or error behavior, but for a read-only status check this is strong transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler; every clause contributes purpose, scope, or a differentiating behavioral trait. The safety-relevant non-side-effect property is clearly stated and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter read-only check, the description covers what it does and its behavioral boundary. It does not spell out the exact return format, but the name 'alerts' and the verb 'check' make the result type reasonably predictable. No significant gaps remain for deciding whether to call it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already fully documents the single optional parameter, including the local-run behavior when omitted. The description adds no parameter-specific detail, but the high schema coverage makes the baseline of 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (check), a clear resource (resource alerts for CPU, memory, and disk usage), and a qualifier (against configured thresholds). The contrast with watcher behavior and the phrase 'nothing is recorded' also separates it from recording/history tools like alerts_history and watch tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies and supports the use case: current threshold status check, not notifications or recording. It explicitly says it reads and compares and that sending/recording would be a running watcher's job, which helps an agent choose this over watcher tools. It does not explicitly name alternatives like alerts_history, but the guidance is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

alerts_historyA

Show recorded alert and remediation history. Entries are only written while a watcher is running, so an empty list means nothing was recording rather than nothing went wrong

ParametersJSON Schema
NameRequiredDescriptionDefault
serverNoRemote server name from config (optional, runs locally if omitted)

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It adds a valuable caveat about recording semantics and the meaning of an empty result, which goes beyond a basic summary. It does not describe output formatting or authorization, but for a simple read-only history tool this is adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused sentence that front-loads the core purpose and then adds a high-value interpretive caveat. There is no redundant or filler content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity tool with one optional, schema-documented parameter and no output schema, the description covers the essential interpretive context. The main missing element is an explicit statement that this is a read-only operation, though the verb 'Show' strongly implies it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single optional 'server' parameter is already fully documented in the schema. The description adds no additional parameter semantics, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb-resource pair, 'Show recorded alert and remediation history,' which clearly identifies the tool's function. It is distinct from sibling 'alerts' by emphasizing recorded history, though it does not explicitly name or contrast itself with related tools like 'watch_history.'

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides important contextual guidance: history is only written while a watcher is running, and an empty list means nothing was recording rather than nothing went wrong. This clarifies when results may be misleading, though it does not explicitly direct the agent to alternatives like 'watch_history' or 'alerts.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

backup_createA

Create a Docker compose backup archive for all services or one service. Volumes are read while the containers run, so a database mid-write can land inconsistent; an archive is not evidence it restores, which is what backup_drill answers

ParametersJSON Schema
NameRequiredDescriptionDefault
toNoCustom backup destination directory (optional)
serverNoRemote server name from config (optional, runs locally if omitted)
excludeNoHost paths not to archive, covering bind mounts at or under each one. What was excluded is written into the archive manifest, so a later restore can tell a mount that was left out from one that never existed (optional)
serviceNoSpecific service to back up (optional)

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses a meaningful operational trait: volumes are read while containers run, so mid-write database states can be inconsistent, and it warns that the archive alone is not proof of restore success. It omits return behavior and permission requirements, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no wasted words. The first front-loads the action and scope, and the second adds the critical caveat and sibling pointer efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The main purpose, runtime caveat, and verification sibling are covered. However, with no output schema, the description does not say what the tool returns or where the archive is written by default, nor does it mention backup_restore for restoring archives. For a tool with four optional parameters and no annotations, this is a moderate gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline of 3 applies. The description adds only slight support for the service parameter via 'all services or one service' and does not otherwise add semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Create a Docker compose backup archive,' and states scope explicitly ('all services or one service'). It also distinguishes itself from backup_drill by noting that an archive is not evidence of restorability, which is what backup_drill answers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly indicates when to use this tool: to create a Docker compose backup. It also points to backup_drill as the alternative for verifying restorability. It does not explicitly mention backup_restore or state an exclusion condition, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

backup_drillA

Verify a backup by booting an app in an isolated Docker environment and checking that it responds. A second copy runs beside the live one on a network and port of its own, and everything it made is removed either way. A pass means the archive is not corrupt and the app starts on it, not that every row is there

ParametersJSON Schema
NameRequiredDescriptionDefault
allNoDrill all supported apps in the backup
appNoApp/service to drill (required unless all=true)
serverNoRemote server name from config (optional, runs locally if omitted)
archiveNoSpecific backup archive to verify (optional)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses that the tool runs an isolated Docker environment, creates a second copy, removes everything it made 'either way' (cleanup regardless of outcome), and clarifies the verification scope ('not that every row is there'). This is strong behavioral transparency for a verification tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with no wasted words. It front-loads the core purpose, then adds the isolation/cleanup behavior, then clarifies the verification scope. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for a verification tool with no output schema. It explains the method, the cleanup behavior, and the meaning of a pass. It does not describe the return format or how to interpret failure, but the absence of an output schema and the simplicity of the tool make this a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters. The description adds context about the overall behavior but does not add specific parameter-level meaning beyond what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Verify a backup'), a resource ('backup'), and a method ('booting an app in an isolated Docker environment and checking that it responds'). It clearly distinguishes this from backup_create, backup_list, and backup_restore siblings by focusing on verification rather than creation, listing, or restoration.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains what the tool does and its scope ('a second copy runs beside the live one on a network and port of its own'), and clarifies what a pass means. It does not explicitly name alternative tools or state when not to use it, but the context makes the verification use case clear. The 'required unless all=true' parameter note adds usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

backup_listA

List existing backup archives in the configured backup directory. It reads names, sizes and dates — that an archive is here says nothing about whether it restores, which is what backup_drill answers

ParametersJSON Schema
NameRequiredDescriptionDefault
serverNoRemote server name from config (optional, runs locally if omitted)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It transparently states this is a read operation ('reads names, sizes and dates') and warns that existence does not imply restorability. This goes beyond a simple listing description by surfacing a key limitation the agent might otherwise assume away.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded: the core action is stated immediately, followed by a targeted clarification about what the data does and does not mean. Both sentences earn their place with no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one optional parameter, no required parameters, and no output schema, the description covers the essential context: what is listed, what fields are shown, a key limitation, and a pointer to the relevant sibling tool. It could mention the exact return format, but the description already conveys the core information needed for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides 100% parameter description coverage for the single optional 'server' parameter, stating it is an optional remote server name from config. The tool description does not add parameter-level detail beyond what the schema already provides, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly names the verb 'List', the resource 'existing backup archives', and the location 'the configured backup directory'. It also differentiates from the sibling backup_drill by explicitly stating that presence of an archive does not imply restorability, so an agent can distinguish this tool from related backup tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool: when you need to see backup names, sizes, and dates. It also provides a meaningful exclusion by directing restoration validation to backup_drill. However, it does not explicitly discuss when to use backup_create or backup_restore, leaving some selection guidance implicit rather than fully explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

backup_restoreA

Restore Docker volumes from a backup archive, overwriting the data the app is running on. Destructive: confirm intent before calling. Bind mounts declared by the archive are always refused here, because an agent has no way to name a host path it may write to

ParametersJSON Schema
NameRequiredDescriptionDefault
serverNoRemote server name from config (optional, runs locally if omitted)
archiveYesBackup archive path to restore
serviceNoSpecific service to restore (optional)

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure and does it well: it states that live app data is overwritten, warns that the operation is destructive, and explains that bind mounts from archives are always refused with a clear reason. This goes well beyond a generic restore description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences front-load the core operation, warn about destructive effects, and state the bind-mount constraint with rationale. Every sentence earns its place and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive restore tool with no annotations and no output schema, the description provides the key cautions and constraints an agent needs to call it safely. It could more explicitly state how the optional service parameter affects restore scope or what happens on failure, but overall it is sufficiently complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All three parameters are already described in the schema, so the schema coverage baseline of 3 applies. The description adds minimal per-parameter meaning beyond identifying the archive as the restore source; server and service semantics are not elaborated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: restore Docker volumes from a backup archive, and explicitly calls out the overwrite behavior. This clearly differentiates it from sibling backup_create/backup_drill/backup_list tools, which concern creation, testing, or listing rather than restoration.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description communicates that destructive action requires confirmation, which is important usage guidance, but it never names alternatives or states when to use this tool versus backup_create/backup_drill. The restore context is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

config_validateA

Check the config file this server is running on: which file was used, which rule selected it, what was read from each section, and anything wrong or silently ignored

ParametersJSON Schema
NameRequiredDescriptionDefault
strictNoTreat warnings as failures in the passed field (default: false)

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Even without annotations, the description conveys the tool's scope and behavior: it inspects config, reports per-section reads, surfaces issues, and flags silently ignored items. This carries the behavioral burden well for a read-only validation tool, though it doesn't explicitly state that no changes are made or what the return shape looks like.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that packs in useful specifics without rambling. Its list of checks is moderately long but each element earns its place by clarifying what the user will learn from the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the main intent and the key outputs, which is helpful given there is no output schema. It doesn't fully explain the return format or whether results are printed or returned, but the listed content is comprehensive enough for an agent to know what it's getting.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The one parameter 'strict' is fully described in the input schema as treating warnings as failures, so the schema covers its semantics completely. The description doesn't add additional parameter context, landing at the baseline 3 for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Check'), names the resource ('the config file this server is running on'), and details the exact output categories: which file, rule, section reads, errors, and silently ignored items. This provides a crisp, unambiguous purpose that distinguishes it from generic diagnostic siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the scenario clear: use this when you need to validate or inspect the server's effective config. It doesn't explicitly name alternatives or exclude other diagnostic tools, but the very specific scope gives a solid sense of when it applies.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

docker_inspectA

Summarize a Docker container's image, state, restart policy, ports, mounts, networks, and health. Environment variable values are never included

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesContainer name to summarize
serverNoRemote server name from config (optional, runs locally if omitted)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. 'Summarize' implies a non-mutating read operation, and it explicitly discloses that environment variable values are never included, which is a useful privacy/limitation signal. It does not cover error behavior or prerequisites, but the key transparency is present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence with a dense, well-scoped list of result facets and a high-value exclusion about environment variables. Every word earns its place, and the most important use information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with no output schema, the description communicates what the result will cover and a key limitation. It could explicitly note that this is read-only/non-mutating, but 'Summarize' conveys that sufficiently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters ('name' and 'server') are already fully documented. The description adds no additional parameter-level meaning, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Summarize') and resource ('Docker container'), and enumerates the exact facets returned: image, state, restart policy, ports, mounts, networks, and health. This clearly distinguishes it from siblings like docker_list, docker_logs, docker_stats, and docker_top.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The listed facets make the usage context clear: use when you need a container's configuration/status summary. It does not explicitly name alternatives or say when not to use it, so it stops short of a 5, but the intended use is strongly implied by the description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

docker_listA

List Docker containers with their status, image, and ports. Stopped containers are included, so a name appearing here is not a name that is running — read state

ParametersJSON Schema
NameRequiredDescriptionDefault
serverNoRemote server name from config (optional, runs locally if omitted)

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden, and it does disclose an important gotcha: stopped containers are included, so a name appearing here is not necessarily a running one. It also implies a read-only listing operation, though it doesn't explicitly state absence of side effects or describe output structure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler: the first states the core purpose and returned fields, and the second adds a valuable behavioral caveat. The warning is front-loaded near the purpose it clarifies.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with one optional parameter fully documented in the schema, the description covers what is returned, that stopped containers are included, and the important interpretive caveat. Nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already has 100% coverage of the single 'server' parameter, explaining it is optional and runs locally if omitted. The description adds no parameter-level meaning, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'List Docker containers' and names the returned fields—status, image, and ports. It also distinguishes itself from inspect/stats/top siblings by framing this as the container-level listing operation that includes stopped containers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit comparison to siblings like docker_inspect, docker_stats, or docker_logs, so an agent must infer when to use this tool. The warning that stopped containers are included implies a listing purpose, but it never says 'use docker_inspect for a single container' or 'use docker_stats for metrics.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

docker_logsA

Get logs from a Docker container: the last lines only, 50 by default, and it returns rather than following the stream

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesContainer name to get logs from
linesNoNumber of log lines to return (default: 50)
serverNoRemote server name from config (optional, runs locally if omitted)

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does contribute two traits not present in the schema: the result is a bounded tail ('last lines only') and it returns instead of blocking on a stream. It omits permissions, behavior for stopped/unknown containers, and error semantics, so a 4 rather than a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with a colon-delimited clause; every element (resource, scope, default, non-streaming behavior) earns its place and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter read tool with no output schema and no annotations, the description covers what the tool does and how the return differs from a streaming call. It leaves minor gaps around error cases and log stream contents (stdout/stderr, timestamps), which keeps it below 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (name, lines, server) are already documented, including the 50-line default. The description restates the default line count but adds no syntax, format, or interaction guidance beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get logs from a Docker container') and even scopes what subset is returned. It is clearly distinct from log-adjacent siblings like docker_stats or docker_top by the nature of the resource, but it never names or differentiates itself from them explicitly, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Returns rather than following the stream' implies the usage context (a one-shot snapshot instead of a live tail), which is the only when-to-use signal offered. There is no explicit statement of when to prefer this over docker_top, docker_inspect, or any other diagnostics sibling, and no prerequisites are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

docker_restartA

Restart a Docker container by name. It goes down and comes back, and the result says the restart command succeeded, not that the app inside is serving again. Read the logs first if you do not know why it needs restarting

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesContainer name to restart
serverNoRemote server name from config (optional, runs locally if omitted)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does well by explaining that the container 'goes down and comes back' and that a successful restart command does not mean the app inside is serving again. This is a meaningful behavioral caveat beyond what the schema conveys, although it does not discuss permissions, auth, or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place: the first states the action, the second clarifies the success semantics, and the third provides operational guidance. The most important information is front-loaded, and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with no output schema, the description is fairly complete. It covers the invocation target, the behavioral consequence of restart, and the meaning of the result. It does not describe possible error conditions or prerequisites, but the schema covers parameter semantics and the core caveat is stated clearly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema fully documents both parameters (name and server), so the baseline is 3. The description adds only that the operation targets a container 'by name,' which aligns with the required parameter but contributes no additional meaning beyond the schema. It does not elaborate on the optional server parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Restart a Docker container by name.' It clearly distinguishes the action from sibling tools like docker_stop, docker_list, and docker_logs by framing it as a restart rather than a stop, list, or log retrieval, and the 'by name' qualifier anchors the primary parameter.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The final sentence gives concrete guidance: 'Read the logs first if you do not know why it needs restarting.' This implicitly points to a sibling tool (docker_logs) and tells the agent when to gather more context before acting. It lacks an explicit 'use tool X instead' statement or a when-not-to-use condition, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

docker_statsA

Get resource usage statistics (CPU, memory, network, block I/O) for all running Docker containers

ParametersJSON Schema
NameRequiredDescriptionDefault
serverNoRemote server name from config (optional, runs locally if omitted)

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavioral traits. It only states that it gets stats, without mentioning side effects (none), behavior when no containers are running, or output format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that conveys all necessary information without any extraneous words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity and single parameter, the description is mostly complete, though it could benefit from indicating the output format or real-time nature.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter (server) has 100% schema coverage, and the description adds no additional information beyond the schema's explanation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (Get), the resource (resource usage statistics), and the scope (all running Docker containers), distinguishing it from sibling tools like docker_list and docker_logs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for monitoring live stats, but does not explicitly state when to use this tool over alternatives or mention any prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

docker_stopA

Stop a Docker container by name. Nothing here starts it again: there is no start tool, so the operator brings it back themselves

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesContainer name to stop
serverNoRemote server name from config (optional, runs locally if omitted)

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the disclosure burden. It goes beyond the schema by warning that the container is not started again and that the operator must bring it back manually. This is meaningful behavioral context for a state-changing operation, though it could also mention effects on running processes or return behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, no redundancy, and the core operation is front-loaded. The caution about starting the container earns its place as essential operational context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The definition is sufficient for a simple two-parameter stop operation: the schema documents parameters, and the description covers the most important non-obvious consequence (no automated restart). It does not describe return values, but the absence of an output schema makes this less critical.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: both 'name' and 'server' already have descriptions in the input schema. The description only reinforces 'by name' and adds no new parameter-level semantics, so a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Stop a Docker container by name.' This clearly identifies the operation and immediately distinguishes it from related container tools like docker_restart and docker_logs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied ('Stop a Docker container'), and the warning that no start tool exists provides some context, but the description does not explicitly compare against alternatives such as docker_restart or state when stopping is preferred. No exclusions or explicit when-not-to-use guidance is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

docker_topA

List the processes running inside a Docker container, read from the host. Read-only: no exec, no TTY

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesContainer name to inspect
serverNoRemote server name from config (optional, runs locally if omitted)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the safety burden. It explicitly discloses Read-only, no exec, and no TTY, and clarifies the operation runs from the host, which is important behavioral context. It doesn't describe failure modes, but for a simple read-only list the essential traits are covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact sentence conveys the action, scope, and safety profile without wasted words. The purpose is front-loaded and the behavioral qualifier is placed right after.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read-only list tool, the description plus full schema coverage is sufficient to invoke it correctly. It doesn't specify expected output formatting or runtime prerequisites, but those are not critical gaps for a top/process-listing tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already describes name and server with 100% coverage, so the baseline is 3. The description adds only 'Docker container' in prose and does not offer extra semantics such as how server selection affects the host read, so no credit above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'List the processes running inside a Docker container.' It adds 'read from the host,' which distinguishes it from tools that exec into containers or inspect configuration, and from docker_list/docker_stats siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The statement 'read from the host' and 'no exec, no TTY' gives clear context: use this for safe, host-side process inspection, and not for interactive or mutating actions. It doesn't name alternatives explicitly, but the context and exclusions make the intended choice clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

doctorB

Run a read-only diagnosis for resource pressure, stopped containers, public ports, backup hygiene, notifications, and report baseline readiness

ParametersJSON Schema
NameRequiredDescriptionDefault
serverNoRemote server name from config (optional, runs locally if omitted)
backup_max_age_hoursNoWarn when the latest backup is older than this many hours (default: 168)

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the key safety trait: "read-only". It also scopes the work to specific check categories. However it says nothing about runtime, permissions required, or side effects (e.g., does the notification check send anything?), which matters for a multi-system diagnostic.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with dense, non-redundant enumeration — each list item corresponds to a real check category. No filler, though the run-on list of six items is slightly dense rather than cleanly structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a diagnostic tool with no output schema and no annotations, the description tells the agent what is inspected but nothing about the shape of the results — warnings vs pass/fail, per-check vs summary, or the baseline-readiness verdict. It is adequate but leaves the return semantics to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented, setting the baseline at 3. The description adds only an indirect link between the '"backup hygiene"' check and the backup_max_age_hours threshold; it contributes no format or default detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("Run a read-only diagnosis") and enumerates exactly what it inspects: resource pressure, stopped containers, public ports, backup hygiene, notifications, baseline readiness. An agent can tell it is an aggregate health-check distinct from atomic siblings like docker_stats or open_ports. It stops short of naming a sibling it is not, so not a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or alternative routing. The agent is not told how this differs in practice from report, system_status, or config_validate, all of which sit in the sibling list. It only implies a broad check implicitly through the enumerated scope.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

install_appB

Install a self-hosted app via docker compose. Pre-checks docker, ports, and duplicates automatically, and a refusal comes back as a result with the reasons rather than as an error

ParametersJSON Schema
NameRequiredDescriptionDefault
appYesApp name (e.g. uptime-kuma, vaultwarden)
portNoCustom host port (optional, uses default if omitted)

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose useful behavior: automatic pre-checks for docker, ports, and duplicates, and that refusals are returned as results rather than errors. However, it says nothing about what success produces (container started? status returned?) or idempotency, leaving a mutation tool partially opaque.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two front-loaded sentences: purpose first, behavior second, with zero filler. Every clause earns its place and the key action is stated in the opening words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with no output schema and no annotations, the description covers the main action and the notable refusal-as-result behavior, which is the most important edge case an agent needs. It falls short only on success semantics and explicit prerequisites.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both 'app' and 'port' are already documented with examples and defaults. The description adds no syntax or format detail beyond the schema, so the baseline of 3 applies; the mention of port pre-checks loosely relates but adds no parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (install), resource (self-hosted app), and mechanism (via docker compose), which contrasts naturally with siblings like install_uninstall, install_purge, and install_status. It does not explicitly name an alternative, but the action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description never says when to use this tool versus install_list/install_status/install_uninstall, nor does it state prerequisites or exclusions. Usage is only implied by the verb 'install'. The pre-check sentence describes mechanics, not selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

install_listA

List available self-hosted apps that can be installed. The catalogue is compiled into the binary, so this reaches no network and answers the same on any machine

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses two key behaviors: the catalogue is compiled into the binary, so this reaches no network, and it answers the same on any machine. This adds valuable context beyond 'list'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences: purpose first, then a single behavioral qualification. Every sentence earns its place with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only listing tool, it covers purpose and distinguishes its offline behavior. It does not detail the output format, but 'List...' sufficiently implies a list of available apps; still, with no output schema, a note on what fields are returned would make it fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters and schema coverage is 100%, so no parameter documentation is required. The description correctly focuses on behavior rather than parameter details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List'), a clear resource ('available self-hosted apps that can be installed'), and distinguishes from siblings like install_app and install_uninstall. The offline/deterministic note reinforces the catalogue semantics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies the tool is for discovering installable apps before choosing one, but does not explicitly state when to prefer it over siblings or when not to use it. No alternative tool is named, so guidance is only implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

install_purgeA

Stop an installed app and delete all data including containers, config, and volumes. Nothing here restores it and no backup is taken first: take one before calling if the data matters

ParametersJSON Schema
NameRequiredDescriptionDefault
appYesApp name

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does exceptionally well. It discloses destructive behavior, enumerates what is deleted (containers, config, volumes), states irreversibility, and explicitly notes no automatic backup is taken.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two efficient sentences front-load the core action and then add a critical safety warning. Every phrase earns its place, with no redundant filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter destructive operation with no output schema, the description is complete: it states the action, scope of deletion, irreversibility, and the backup prerequisite. Nothing essential is missing for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with the 'app' parameter already described as 'App name'. The description adds minimal extra meaning by noting the app must be an installed app, but it does not need to compensate for any schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Stop', 'delete') and a concrete resource ('installed app' and all data including containers, config, volumes). Clearly distinguishes itself from sibling tools like install_uninstall by emphasizing irreversible data deletion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context for when to use the tool: when complete removal of an app and its data is intended. It warns to take a backup first if data matters, but does not explicitly name or exclude alternatives such as install_uninstall.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

install_statusA

Check the status of an installed app, as its containers report it. An app homebutler did not install is not known here

ParametersJSON Schema
NameRequiredDescriptionDefault
appYesApp name

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral transparency burden. It discloses that status comes from containers and that unrecognized apps are not known, which is useful. However, it does not explain possible status values, failure modes, or whether the operation is purely read-only, leaving gaps for an agent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences and every sentence earns its place: the first states the core function, the second clarifies scope and edge behavior. It is concise and front-loaded, though the second sentence is slightly awkwardly phrased.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter status tool with no output schema, the description gives enough context: what is checked, where the data comes from, and when the app may be unknown. It could be more complete by describing the response format, but the basic usage and behavior are adequately covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema only describes 'app' as 'App name', but the description adds meaning by clarifying the app must be one installed by homebutler. This helps the agent understand valid parameter values and what happens for unknown apps, going beyond the bare schema definition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Check'), a resource ('status of an installed app'), and a data source ('as its containers report it'). It further differentiates from install_list and similar siblings by explicitly limiting scope to apps homebutler installed. This leaves no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool: for installed apps known to homebutler, not for arbitrary containers. However, it does not explicitly mention alternatives such as docker_inspect or install_list, nor does it state when not to use it beyond the scope restriction. The usage guidance is mostly implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

install_uninstallA

Stop an installed app and remove its containers. The app directory and its volumes stay on disk; install_purge is the one that deletes them

ParametersJSON Schema
NameRequiredDescriptionDefault
appYesApp name

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses what is removed (containers), what is preserved (app directory and volumes), and points to install_purge for destructive deletion. This is genuinely informative, though it does not address reversibility or side effects beyond removal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one information-dense sentence with no filler. The key behavioral distinction is front-loaded, and the alternative tool is mentioned without unnecessary elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter operation, the description provides enough to understand the action, what persists, and which sibling covers the destructive case. It leaves out return values or error behavior, but that is not critical given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the only parameter 'app' is already described as 'App name'. The description adds little parameter-level detail, but none is needed for this simple one-parameter tool, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific behavior: stop an installed app and remove its containers. It also distinguishes this tool from install_purge by noting that the app directory and volumes remain on disk, making the scope of the operation unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives strong context by naming install_purge as the alternative that deletes app directories and volumes, implying this tool is for cleanup while preserving data. It does not explicitly say 'use this when...' but the contrast is clear enough for an agent to choose correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inventory_exportB

Export server inventory/topology as a Mermaid diagram locally, or JSON locally/remotely

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNoExport format: mermaid (default, local) or json
serverNoRemote server name from config (optional; remote supports format=json)

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description is the sole source for behavioral traits. It does not disclose whether the operation is read-only, whether it modifies server state, authentication needs, or file effects. Only states export format and destination.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence with front-loaded verb 'Export', no wasted words. Efficiently conveys purpose and key options.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Lacks output schema and does not explain what the tool returns (e.g., file path, content, or stdout). No behavioral details like side effects or prerequisites. With 100% schema coverage but no output description, it is minimally complete but could be richer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema already covers both parameters with clear descriptions (format: mermaid/json, server: remote server name). Tool description adds only a summary of capabilities (e.g., 'remote supports format=json'), but does not introduce new semantic detail beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the tool exports server inventory/topology, with two specific formats (Mermaid diagram and JSON) and local vs remote options. Distinguishes from siblings like inventory_scan which is for scanning.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool over alternatives such as inventory_scan or other export-like tools. The description only implies usage through the verb 'Export', lacking context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inventory_scanB

Collect server inventory/topology including system status, Docker containers, app ports, and system ports

ParametersJSON Schema
NameRequiredDescriptionDefault
serverNoRemote server name from config (optional, runs locally if omitted)

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided; description does not disclose whether the tool is read-only, has side effects, or requires special permissions. Only lists output categories without behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence covering all key points without redundancy. Could be slightly more structured (e.g., separate lines for each component) but is acceptably concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given one optional parameter and no output schema, the description adequately defines the tool's scope. However, lacks usage context and behavioral details, which are needed for full completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Single parameter 'server' with schema description 'Remote server name from config (optional, runs locally if omitted)' adds meaning beyond the schema, clarifying config dependency and default behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states 'collect server inventory/topology' with specific items (system status, Docker containers, ports). Distinguishes from siblings like system_status and docker_list by combining multiple data sources, but 'topology' is vague.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives like system_status, docker_list, or network_scan. Does not indicate prerequisites or when to prefer inventory_scan.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

network_scanA

Scan the local network to discover devices (IP, MAC, hostname). It probes every address on the subnet and takes up to 30 seconds, so it is an answer to a question somebody asked rather than a way to begin

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Since no annotations are provided, the description carries the full burden of behavioral disclosure. It clearly states that the tool probes every address on the subnet and that execution may take up to 30 seconds, which are useful runtime behaviors beyond what the schema reveals. It does not explicitly state whether it is read-only, but the described behavior of probing and discovering devices strongly implies a non-destructive network scan.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. It front-loads the core action and result, then adds the important performance characteristic and usage framing. Every sentence adds value, and there is no redundant or tautological content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters and no output schema, the description provides the essential information an agent needs: what is scanned, what will be discovered, how long it may take, and when it should be used. It could optionally mention that the result is a list of discovered devices, but the stated outcome '(IP, MAC, hostname)' makes that reasonably clear. Overall, it is complete enough for a zero-argument tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters)SkipProcessing... The schema description coverage is 100% vacuously, so there is no parameter information for the description to supplement. Per the baseline for zero-parameter tools, a score of 4 is appropriate because the description focuses on what the tool returns and how it behaves rather than parameter details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Scan'), a clear resource ('the local network'), and the expected output ('devices (IP, MAC, hostname)'). It is immediately distinguishable from sibling tools like open_ports or system_status because it targets network device discovery rather than system or port inspection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit context on when to use it: it is an answer to a question somebody asked, not a way to begin a session or workflow. It also sets expectations about duration ('up to 30 seconds'), which helps an agent decide whether to invoke it. It does not name specific alternative tools for exclusion, but the guidance is sufficient for a zero-parameter, one-off scan.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

notify_testA

Send one test notification through every configured channel and report which ones arrived. A real message goes out to each, so anyone reading those channels sees it

ParametersJSON Schema
NameRequiredDescriptionDefault
serverNoRemote server name from config (optional, runs locally if omitted)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses the key behavioral side effect ('A real message goes out to each, so anyone reading those channels sees it'), which is crucial for an agent to anticipate impact. It also indicates it reports arrival status. Although annotations are absent, this description covers the primary behavioral nuance.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The main action and outcome are front-loaded, and the side-effect warning follows. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one optional parameter and no output schema, the description provides sufficient context: it explains the purpose, the real-message side effect, and that it reports arrival. It does not detail the report format or error handling, but these are minor given the simple nature of the operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema fully documents the single optional parameter (server) with its description 'Remote server name from config (optional, runs locally if omitted)'. The tool description adds no further detail about this parameter, so with 100% schema coverage, a baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (send one test notification), the scope (every configured channel), and the outcome (report which ones arrived). It is distinct from siblings like alerts or system_status, which do not test notification channels.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool (to test notification channels) but does not explicitly state alternatives or exclusions. However, since no sibling tool performs a similar function, the lack of explicit guidance is not a significant gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

open_portsA

List open network ports with associated process information. The process behind a port is not always readable without privilege, and missing_process says so rather than leaving the field quietly empty

ParametersJSON Schema
NameRequiredDescriptionDefault
serverNoRemote server name from config (optional, runs locally if omitted)

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description explicitly discloses an important behavioral nuance: process information may require privileges and is represented by missing_process instead of being silently omitted. This goes beyond a bare list operation and prepares the agent for a realistic failure mode. No annotations are present, so this disclosure carries useful weight.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no wasted words. The primary action is front-loaded, and the second sentence adds a valuable edge-case clarification without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, read-only listing tool with no output schema, the description covers the core behavior and a significant output nuance. It could be more complete by sketching the typical return fields or noting any prerequisites, but it still gives an agent enough context to invoke the tool confidently and interpret missing_process.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter, server, is fully documented in the schema with 'Remote server name from config (optional, runs locally if omitted).' Since schema coverage is 100%, the description does not need to add much. It adds no extra parameter-level meaning, matching the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'List open network ports with associated process information.' It clearly identifies what the tool does and what it returns. However, it does not explicitly distinguish itself from sibling tools such as network_scan or processes, so it misses the full differentiation credit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention when local or remote execution is preferred, nor does it explain exclusions such as when to use processes or network_scan instead. The only context is the optional server parameter in the schema, which is not repeated or expanded on.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

processesB

List the top processes by CPU or memory, with a total count and any zombies broken out separately

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoNumber of processes to return (default: 10, 0 for all)
serverNoRemote server name from config (optional, runs locally if omitted)
sort_byNoSort by cpu (default) or mem

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the disclosure burden. It usefully describes the return shape (a total count and a separate zombie breakout), which implies a read-only snapshot, but says nothing about permissions, cost, or whether the remote 'server' path changes behavior. Adequate but incomplete for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that names the action, the resource, the sort options, and the notable output features with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless-required, read-only listing tool with no output schema, the description plus schema cover what an agent needs to invoke it correctly, including the zombie/count return detail. Only the remote 'server' behavior is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters are already fully documented (defaults, enum-ish values for sort_by, remote-vs-local semantics for server). The description echoes the CPU/memory sorting concept but adds no detail beyond the schema, making 3 the correct baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('processes') plus scope qualifiers ('top ... by CPU or memory'). It distinguishes the sort dimension, though it doesn't contrast itself with the nearest sibling docker_top, which is why it falls short of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description never says when to use this tool versus siblings like docker_top or system_status, nor any prerequisites. Usage must be inferred entirely from the name and parameter list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

proxmox_guest_rebootB

Reboot one explicitly targeted Proxmox guest after confirmation and return the accepted task UPID

ParametersJSON Schema
NameRequiredDescriptionDefault
nodeYesProxmox node name
typeYesGuest type: qemu or lxc
vmidYesProxmox guest VMID from 1 through 999999999
confirmYesMust be true to confirm the explicit guest action target
endpointYesExplicit Proxmox endpoint name from config

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the full behavioral burden. It discloses the confirmation requirement and the return of a task UPID, but fails to warn that reboot is disruptive (downtime, service interruption) or state prerequisites such as the guest being running or required permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that states the action, the constraint, and the return value without any wasted words. Appropriate size for the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Schema is complete and the description supplies the return value (accepted task UPID) and confirmation condition. However, for a destructive reboot with no annotations or output schema, it should also disclose safety/downtime or prerequisite information; those gaps remain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are fully documented in the schema. The description adds only the notion of explicit targeting and confirmation, which maps to the confirm parameter without adding syntax or format details beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Reboot one explicitly targeted Proxmox guest'. The scope (one guest, explicitly targeted) is clear and distinguishes it from a bulk operation, though it does not name sibling alternatives like proxmox_guest_shutdown or proxmox_guest_start.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies usage context: reboot a guest after confirmation and only when explicitly targeted. No when-not conditions or explicit alternatives to shutdown/start are provided, leaving the agent to infer routing from the verb alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

proxmox_guestsA

List Proxmox QEMU and LXC guests, optionally filtered by node, status, or type

ParametersJSON Schema
NameRequiredDescriptionDefault
nodeNoOnly guests on this Proxmox node (optional)
typeNoOnly guests of this type: qemu or lxc (optional)
statusNoOnly guests with this status, such as running or stopped (optional)
endpointNoProxmox endpoint name from config (optional when exactly one is configured)

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. 'List' signals a non-mutating read operation and the optional filters add useful behavior, but the description does not disclose default behavior when no filters are provided, authentication expectations, pagination, or response shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the action and resource, then lists filters compactly. Every word earns its place; there is no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The operation is simple, the schema fully documents all parameters, and the description clearly states the resource and optional filters. It could be slightly richer about the default all-guests behavior and endpoint selection, but those are largely inferable from the schema and the list verb.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are already documented in the input schema. The description mostly restates 'node, status, or type' and adds no meaning beyond the schema, which is the baseline 3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and a specific resource ('Proxmox QEMU and LXC guests'), and it names the optional filtering dimensions (node, status, type). This clearly distinguishes it from sibling tools like proxmox_node, proxmox_tasks, and proxmox_status, which target different resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use for enumerating guests rather than nodes, tasks, or Docker containers, but it does not explicitly state when to prefer this tool over siblings or when a filter should be applied. Usage context is present only implicitly through the resource name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

proxmox_guest_shutdownA

Gracefully shut down one explicitly targeted Proxmox guest after confirmation and return the accepted task UPID

ParametersJSON Schema
NameRequiredDescriptionDefault
nodeYesProxmox node name
typeYesGuest type: qemu or lxc
vmidYesProxmox guest VMID from 1 through 999999999
confirmYesMust be true to confirm the explicit guest action target
endpointYesExplicit Proxmox endpoint name from config

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the burden: it discloses the graceful-shutdown semantics, the mandatory confirmation gate, and that the call returns an accepted task UPID (an async task, not a synchronous result). It omits permission/auth requirements and any timeout or force-fallback behavior, so it falls short of fully rich disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the action, scope, precondition, and return value with no wasted words. Optimal for a mutation tool of this complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations and no output schema, the description covers the essentials an agent needs: target scope, confirmation precondition, and the async UPID return. It lacks auth prerequisites and error/timeout expectations, which is a minor gap rather than a blocking one.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (node, type, vmid, confirm, endpoint) is already documented in the schema, establishing the baseline of 3. The description reinforces the confirmation semantics of 'confirm' but adds no syntax or format detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (gracefully shut down) and resource (one Proxmox guest), plus the confirmation requirement and return value. The 'gracefully' qualifier and 'one explicitly targeted guest' scope make it clearly distinct from sibling proxmox_guest_reboot and proxmox_guest_start without needing their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'After confirmation' implies the confirm gate and the single-target scope implies this is a targeted stop, but it does not explicitly say when to choose shutdown over reboot or how to handle a hanging guest (force stop). Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

proxmox_guest_startA

Start one explicitly targeted Proxmox guest after confirmation and return the accepted task UPID

ParametersJSON Schema
NameRequiredDescriptionDefault
nodeYesProxmox node name
typeYesGuest type: qemu or lxc
vmidYesProxmox guest VMID from 1 through 999999999
confirmYesMust be true to confirm the explicit guest action target
endpointYesExplicit Proxmox endpoint name from config

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses two meaningful traits: the action is asynchronous (returns an accepted task UPID rather than a final result) and it is gated on explicit confirmation. It does not cover permissions, idempotency, or error/refusal behavior on an already-running guest.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the verb and target come first, followed by the confirmation gate and return value. Efficient, though slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-required-parameter async action with no output schema, the description covers the operation, its preconditions, and its return contract (task UPID), which is enough for an agent to call it correctly. Permission/auth expectations and behavior on an already-running guest remain unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all five parameters including the confirm gate and enum for type are already documented. The description only reinforces the 'targeted single guest' and 'confirmation' semantics, adding little beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (start) and resource (Proxmox guest) with a scope constraint ('one explicitly targeted'), so the agent knows this acts on a single identified guest. It doesn't explicitly distinguish itself from proxmox_guest_shutdown/reboot in text, but the verb and name make the boundary obvious.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'After confirmation' and 'explicitly targeted' imply preconditions (the confirm flag, a specific vmid), which is useful implied usage. However, there is no explicit statement of when to use this versus the shutdown/reboot siblings or what confirmation workflow is expected.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

proxmox_nodeC

Get detailed Proxmox node status

ParametersJSON Schema
NameRequiredDescriptionDefault
nodeYesProxmox node name
endpointNoProxmox endpoint name from config (optional when exactly one is configured)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It implies a read-only operation through 'Get', but does not explain return shape, pagination, authentication requirements, failure behavior, or what 'detailed status' includes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise, front-loaded sentence with no redundant words. It states the action and resource directly, which is appropriate for a simple status-retrieval tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of annotations, missing output schema, and the existence of closely related sibling tools, this description is too sparse. It does not clarify when to prefer this tool, what 'detailed' status includes, or how the optional endpoint parameter behaves.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents both parameters. The description adds no additional meaning about node or endpoint beyond what the schema provides, so a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource ('Get detailed Proxmox node status'), making the tool's core purpose clear. It does not explicitly differentiate itself from sibling tools like proxmox_status, though the 'node' subject creates reasonable distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus alternatives such as proxmox_status or proxmox_guests. There are no exclusions, prerequisites, or contextual hints beyond the description itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

proxmox_script_commandA

Render the pinned install command for one Proxmox VE Community Script. Never fetches or runs it; the caller reviews and runs it themselves on the Proxmox host

ParametersJSON Schema
NameRequiredDescriptionDefault
slugYesScript slug from proxmox_script_list, such as docker

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly discloses that the tool never fetches or runs the script and that the caller must review and run it on the Proxmox host. This is meaningful safety-relevant behavior, though it does not describe error cases or output format in depth.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one compact sentence that front-loads the action and immediately follows with the critical non-execution guarantee. Every word contributes value, with no repetition of the name or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a single well-documented parameter, no output schema, and no annotations, the description is complete: it states what is rendered, that it is not executed, and that the caller runs it on the Proxmox host. An agent has enough context to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the slug parameter is already documented as 'Script slug from proxmox_script_list, such as docker.' The description adds no additional parameter meaning beyond naming 'one Proxmox VE Community Script,' so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Render') and names the exact resource ('the pinned install command for one Proxmox VE Community Script'). It also distinguishes the tool from any execution behavior by stating it 'Never fetches or runs it,' making it clear this is a command-rendering tool, not a script runner.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: this tool only renders a command and does not fetch or execute it, so an agent knows not to use it when execution is needed. It does not explicitly name alternatives like proxmox_script_list, but the boundary with running commands is clearly implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

proxmox_script_listA

List the curated Proxmox VE Community Scripts catalog (community-scripts/ProxmoxVE)

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the disclosure burden. The verb 'List' and the word 'catalog' imply a read-only operation, which is useful, but the description does not mention whether it fetches live data from the community repository, whether network access is required, or what fields the returned catalog contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that states the action and the source without wasted words. Every element contributes to the agent's understanding of what the tool returns and from where.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple zero-parameter list tool with no annotations and no output schema, the description is mostly complete: it identifies the resource and source. It could additionally mention the output shape, but the listing nature of the tool makes the missing detail a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, and the schema coverage is 100%, so there are no parameter semantics to document. The description does not need to add parameter meaning, and the baseline for a no-parameter tool is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and a clear resource ('curated Proxmox VE Community Scripts catalog'), and names the upstream source repo. This makes it immediately distinguishable from sibling tools like proxmox_script_command, which implies execution rather than listing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied: call this when the user wants to see the available community Proxmox scripts. However, there is no explicit when-to-use versus alternatives, such as 'use proxmox_script_command to run one of these scripts,' so guidance is only implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

proxmox_statusB

Get Proxmox VE version, cluster status, and resources

ParametersJSON Schema
NameRequiredDescriptionDefault
endpointNoProxmox endpoint name from config (optional when exactly one is configured)

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the transparency burden. 'Get' clearly implies a read-only operation, and the listed outputs give a basic sense of what the call returns. However, it does not disclose endpoint-resolution behavior, potential failures when no endpoint is configured, or any operational side effects, leaving some behavior implicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, direct sentence with no filler and the main action front-loaded. It earns its place by naming three concrete output categories, though it could have been slightly richer with sibling guidance without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool with one well-documented optional parameter, the description covers the basics. However, there is no output schema to clarify return structure, and the description does not clarify when to prefer proxmox_status over closely related siblings like proxmox_node or system_status, leaving moderate ambiguity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already documents the single optional parameter effectively, including when it can be omitted. The description adds no additional parameter-level meaning, but because schema description coverage is 100%, the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a clear action verb ('Get') and a specific resource ('Proxmox VE'), then enumerates the output categories: version, cluster status, and resources. It is understandable on its own, but it does not explicitly distinguish itself from siblings like proxmox_node or system_status, so it stops short of a perfect score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool rather than alternatives such as proxmox_node, proxmox_guests, or system_status. There is no mention of scenarios, exclusions, or a preferred alternative, leaving the agent to infer usage solely from the name and general wording.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

proxmox_tasksA

Get the 50 most recent Proxmox tasks for a node

ParametersJSON Schema
NameRequiredDescriptionDefault
nodeYesProxmox node name
endpointNoProxmox endpoint name from config (optional when exactly one is configured)

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that this is a retrieval operation and that only 50 tasks are returned, but it does not state permissions, error behavior, sorting guarantees, or confirm read-only semantics beyond the verb 'Get'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that communicates the key resource, scope, and result limit with no filler. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read tool with fully documented parameters and no output schema, this is largely sufficient. It could mention endpoint selection or the shape of the returned task list, but the core behavior and parameters are clear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters already have meaningful descriptions in the schema. The tool description adds no additional parameter-level meaning, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Get'), a specific resource ('Proxmox tasks'), a scope ('for a node'), and a hard limit (50 most recent). This clearly differentiates it from siblings like proxmox_status and proxmox_node without requiring schema inspection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use when recent Proxmox task history for a node is needed, but it gives no explicit when-to-use or when-not-to-use guidance, no exclusions, and no reference to alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

proxmox_task_statusA

Inspect one asynchronous Proxmox task by node and opaque UPID

ParametersJSON Schema
NameRequiredDescriptionDefault
nodeYesProxmox node name
upidYesOpaque Proxmox task UPID
endpointYesExplicit Proxmox endpoint name from config

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states the operation is an inspection of an asynchronous task, which implies a read-only status check, but it does not mention whether the task may still be running, what happens for an invalid or missing UPID, or whether the call polls until completion.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler or repeated information. It conveys the action, resource, and identifying parameters efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-task inspection tool, the description covers the essential action and parameters, and the schema fully documents the three required inputs. It could be slightly richer by noting what the returned status looks like or that the task may be in-progress, but overall an agent can reasonably determine how to call it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already documents node, upid, and endpoint. The description adds context that the UPID is opaque and that the task is asynchronous, but it does not meaningfully elaborate on the parameters beyond the schema. A baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Inspect') and resource ('one asynchronous Proxmox task'), and clearly scopes the operation by node and UPID. It also distinguishes this from the plural sibling `proxmox_tasks`, which presumably lists tasks rather than inspecting a single one.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this should be used when a single task's status is needed by node and UPID, particularly since `proxmox_tasks` likely handles listing. However, there is no explicit statement of when to prefer this over siblings or what circumstances would make it the wrong choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reportA

Generate a butler-style health report with snapshot comparison, warnings, notable changes, and suggested actions. It saves a snapshot unless no_save is set, which moves the window every later comparison is measured from — so a loop that calls this leaves nothing to compare against

ParametersJSON Schema
NameRequiredDescriptionDefault
keepNoNumber of snapshots to retain (default: 30)
serverNoRemote server name from config (optional, runs locally if omitted)
no_saveNoPreview without writing a snapshot

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the write side effect (saves a snapshot by default), the stateful consequence (moving the window future comparisons are measured from), and the preview escape hatch. It omits return format and any permission/remoteness caveats, which keeps it below a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with purpose and then the critical side-effect caveat. The second sentence is dense but every clause (default save, window shift, loop consequence) earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-param, no-annotation, no-output-schema tool, the description covers purpose, contents, and the key stateful side effect. An agent can call it correctly; only the exact output shape and remote-server behavior are left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema's terse 'Preview without writing a snapshot' by explaining that no_save affects the comparison window. keep and server remain undocumented in prose, so it is not exhaustive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('generate') and resource ('butler-style health report') and enumerates its contents (snapshot comparison, warnings, notable changes, suggested actions). This distinguishes it reasonably from generic siblings like system_status or doctor, though it never names those alternatives explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operational context for no_save, warning that saving moves the comparison window and that looping leaves nothing to compare against — a genuine when-not-use condition. It stops short of routing the agent to a specific sibling for other use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

system_statusA

Get system status including CPU, memory, disk usage, and uptime, for the machine this binary runs on. Inside a container that is the container, not the host underneath it

ParametersJSON Schema
NameRequiredDescriptionDefault
serverNoRemote server name from config (optional, runs locally if omitted)

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It usefully discloses that results reflect the container, not the host, when run inside one. But it omits the return format, whether the call can fail or block, and any permission implications — the enumerated metrics are the only behavioral detail, which is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero waste: the first states purpose and scope in one pass, and the second adds the single most decision-relevant caveat (container vs. host). The key scope constraint is front-loaded, and nothing extraneous is included.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, one-optional-parameter read tool with no output schema, the description is largely sufficient — it names the metrics returned and the execution scope. The only real gap is that no output shape is described anywhere, but given the low complexity, the level of detail is near-complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema fully documents the 'server' parameter as an optional remote target that runs locally if omitted. The description's 'machine this binary runs on' phrasing reinforces this but adds no new syntax or format detail beyond the schema, warranting the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get') with a clear resource ('system status') and enumerates the concrete metrics (CPU, memory, disk usage, uptime). It also pins the scope precisely to 'the machine this binary runs on', which distinguishes it from container-level tools like docker_stats in the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description conveys clear context about scope — it targets the local/host machine and clarifies the container boundary ('Inside a container that is the container, not the host underneath it'). However, it never names an alternative like docker_stats or proxmox_status, or states when NOT to use it, so an agent must infer routing from sibling names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wakeA

Send a Wake-on-LAN magic packet to wake a machine. The packet is fire-and-forget: a successful result means it was sent, not that anything woke up

ParametersJSON Schema
NameRequiredDescriptionDefault
targetYesMAC address or configured device name
broadcastNoBroadcast address (default: 255.255.255.255)

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly states the fire-and-forget nature and clarifies that a successful result only means the packet was sent, not that the machine woke. This is valuable, though it omits details like network requirements or auth/rate-limit behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two compact sentences that front-load the core purpose and immediately follow with the critical behavioral caveat. Every word earns its place; there is no fluff or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with no output schema, the description gives the essential behavior and result semantics. It could add a note about return values or when a result might be misleading, but the fire-and-forget clarification already addresses the most important gap an agent would face.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters are already described in the schema. The description adds no further meaning beyond what the schema provides, such as accepted formats for 'target' or examples for 'broadcast', so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Send a Wake-on-LAN magic packet') with a clear resource ('a machine') and even a key mechanism caveat. It is unambiguous and easily distinguished from sibling tools like proxmox_guest_start or docker_restart, none of which perform power-on via magic packets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains what the tool does but gives no explicit guidance on when to use it versus alternatives. It does not mention when to prefer this over powering on via a hypervisor (e.g., proxmox_guest_start), nor does it note any prerequisites like the target being on the same LAN.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

watch_addA

Add a Docker container, systemd unit, or PM2 app to the watch list. This writes the list and nothing begins watching it — a supervisor has to be installed separately, which only the operator can do. Adding a target twice is reported as added=false rather than an error

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoWhat it is: docker, systemd, or pm2 (default docker)
unitNoActual unit or app name when it differs from the name above (optional)
serverNoRemote server name from config (optional, runs locally if omitted)
containerYesContainer, unit, or app name to watch

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly says the operation writes the list, does not start watching anything, requires a separately installed supervisor, and returns added=false for duplicates instead of throwing an error. This is exemplary behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the main action and then uses two short sentences to convey critical side effects and duplicate behavior. Every sentence adds necessary information and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the key operational caveats: list writes, no immediate watching, supervisor requirement, and duplicate behavior. It does not fully describe the success response shape, but the duplicate note implies an added boolean, which is enough for an agent to interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline applies. The description reinforces the meaning of kind by listing docker, systemd, and pm2, but it does not add new parameter semantics beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Add a Docker container, systemd unit, or PM2 app to the watch list.' This clearly distinguishes it from sibling tools like watch_remove, watch_check, and watch_list, and it mentions the three supported kinds explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the tool's core use case obvious (adding an item to the watch list) but does not explicitly state when not to use it or name alternatives like watch_remove for deletions. Usage is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

watch_checkA

Run a one-shot restart check on watched targets and report restarts detected since the last check. Only docker targets can be inspected this way; systemd and pm2 targets are reported as skipped rather than assumed healthy

ParametersJSON Schema
NameRequiredDescriptionDefault
serverNoRemote server name from config (optional, runs locally if omitted)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations supplied, the description carries the behavioral burden. It discloses that this is a one-shot operation, that it is based on the last check, and that unsupported target types are reported as skipped rather than silently treated as healthy. This is useful transparency, though it could go slightly further in describing the actual return shape or state it depends on.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, and the main purpose is front-loaded. The second sentence contributes an operational nuance end important caveat. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one optional parameter and a single critical caveat (docker-only), the description covers the essential domain. It lacks only details about the output/format of the restart report and what 'watched targets' means concretely, but it is sufficient for an agent to decide whether to invoke it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the sole optional parameter server is already documented as 'Remote server name from config (optional, runs locally if omitted)'. The tool description does not add anything beyond this schema. The baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('run a one-shot restart check') and a specific resource ('watched targets'), and clarifies the result as 'restarts detected since the last check'. It also differentiates this from restart actions such as docker_restart by using 'check' rather than an action verb.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context and a concrete exclusion: only docker targets can be inspected, while systemd and pm2 targets are reported as skipped. This is strong usage guidance, though it stops short of naming alternative tools or explicitly saying when to choose a sibling instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

watch_historyB

List recorded restart incidents, newest first. Captured logs are excluded unless include_logs is set, because every incident carries a hundred lines of output twice over

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMost recent N incidents (default: 10, 0 for all)
serverNoRemote server name from config (optional, runs locally if omitted)
containerNoOnly incidents for this target (optional)
include_logsNoInclude the logs captured before and after each restart (default: false)

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral burden. It adds useful facts: newest-first ordering and default log exclusion with a payload-size rationale, but omits read-only nature, auth implications, and whether the server parameter changes execution context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, front-loaded with the core purpose and key default. The rationale clause about log volume earns its place by explaining a non-obvious default.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-optional-filter list tool with no output schema, the description covers ordering and log volume well. It leaves some gaps around server/container filter behavior and read-only execution, but schema coverage fills most parameter detail.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds the reason behind include_logs' default, but the limit, server, and container filters are explained only in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('recorded restart incidents'), plus ordering ('newest first'). The resource is distinct enough from most siblings, but it does not explicitly contrast itself with watch_list or alerts_history.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or when-not-to-use guidance, and no alternatives are named. The only conditional instruction concerns include_logs, not selecting this tool over watch_list, watch_check, or alerts_history.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

watch_listA

List the targets being watched, with their kind and what the last check recorded

ParametersJSON Schema
NameRequiredDescriptionDefault
serverNoRemote server name from config (optional, runs locally if omitted)

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full behavioral burden. 'List' signals a non-mutating read operation, and the output is described well enough for an agent to set expectations. It does not mention optional server scope or edge cases like empty results, but beyond the schema that is minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence covers the action, the target, and the returned data without any filler or repetition. Every part contributes to the purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With one optional parameter, no output schema, and a simple list behavior, the description and schema together are almost fully sufficient. The only missing detail is explicit handling of the optional server scope (e.g., that it queries remote vs local watch lists), but the schema already conveys that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, server, is 100% described in the input schema (remote server from config, optional, local if omitted), so the description has no obligation to repeat it. The description adds nothing beyond the schema, matching the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('List') and resource ('targets being watched'), and further clarifies the output as 'kind and what the last check recorded.' This makes the tool's purpose immediately distinct from the infrastructure-focused siblings like system_status and the history-focused watch_history.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: use this when you want the current watch list and its last recorded result. However, there is no explicit guidance about when to choose watch_list over watch_history or watch_check, or any when-not-to-use note.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

watch_removeA

Remove a target from the watch list, leaving its recorded incidents in place. Only the list changes: whatever was supervising it keeps running until the operator stops it

ParametersJSON Schema
NameRequiredDescriptionDefault
serverNoRemote server name from config (optional, runs locally if omitted)
containerYesName to stop watching

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly states that recorded incidents are left in place and that the supervising process continues to run until the operator stops it. This is meaningful non-obvious behavior. However, it does not mention potential error behaviors (e.g., if the container is not currently watched) or any required permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. It front-loads the primary action ('Remove a target from the watch list') and then immediately adds the key behavioral nuance. Every word earns its place, and it is appropriately sized for a simple operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description does not explain the return value, but for a removal operation this is usually a confirmation. It also does not mention prerequisites like whether the container must already be watched, but the sibling tools and context make this inferable. The description covers the main behavior and side effects, which is adequate for its simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (server and container) are already described in the schema. The tool description adds no extra detail about parameter values, formats, or interplay. Since coverage is high, a baseline of 3 is appropriate and the description does not add value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('remove') and resource ('a target from the watch list'), and clarifies scope by noting recorded incidents are preserved. It clearly distinguishes from siblings like watch_add and watch_list by making the removal action unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly defines usage by contrasting with the add operation: it removes from the list without stopping supervision. It explains the side effect that the watcher keeps running, which is relevant when deciding whether to use this tool versus watch_stop or stopping the supervisor. However, it does not explicitly name alternatives or give conditions like 'use when you need to stop watching a container but retain its history'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 9 tool updatesv0.41.0
    • Changeddocker_logs1 field changed
      • changedInput schema / properties / lines / type
        Previous value: -"string"New value: +"integer"
    • Changeddoctor1 field changed
      • changedInput schema / properties / backup_max_age_hours / type
        Previous value: -"number"New value: +"integer"
    • Changedinstall_app1 field changed
      • changedInput schema / properties / port / type
        Previous value: -"string"New value: +"integer"
    • Changedprocesses1 field changed
      • changedInput schema / properties / limit / type
        Previous value: -"string"New value: +"integer"
    • Changedproxmox_guest_reboot1 field changed
      • addedInput schema / properties / type / enum
        Added value: +[
        +  "qemu",
        +  "lxc"
        +]
    • Changedproxmox_guest_shutdown1 field changed
      • addedInput schema / properties / type / enum
        Added value: +[
        +  "qemu",
        +  "lxc"
        +]
    • Changedproxmox_guest_start1 field changed
      • addedInput schema / properties / type / enum
        Added value: +[
        +  "qemu",
        +  "lxc"
        +]
    • Changedreport1 field changed
      • changedInput schema / properties / keep / type
        Previous value: -"number"New value: +"integer"
    • Changedwatch_history1 field changed
      • changedInput schema / properties / limit / type
        Previous value: -"string"New value: +"integer"
  2. 1 tool updatev0.39.0
    • Changedbackup_create1 field changed
      • addedInput schema / properties / exclude
        Added value: +{
        +  "description": "Host paths not to archive, covering bind mounts at or under each one. What was excluded is written into the archive manifest, so a later restore can tell a mount that was left out from one that never existed (optional)",
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
  3. 1 tool updatev0.35.1
    • Addednotify_test
  4. 3 tool updatesv0.31.0
    • Addedalerts_history
    • Addedwatch_add
    • Addedwatch_remove
  5. 6 tool updatesv0.23.0
    • Addedproxmox_guest_reboot
    • Addedproxmox_guest_shutdown
    • Addedproxmox_guest_start
    • Addedproxmox_script_command
    • Addedproxmox_script_list
    • Addedproxmox_task_status
  6. 6 tool updatesv0.22.1
    • Addeddocker_inspect
    • Addeddocker_top
    • Addedproxmox_guests
    • Addedproxmox_node
    • Addedproxmox_status
    • Addedproxmox_tasks
  7. 4 tool updatesv0.22.0
    • Addedconfig_validate
    • Addedprocesses
    • Addedwatch_history
    • Addedwatch_list
  8. 1 tool updatev0.21.2
    • Addedwatch_check
  9. 1 tool updatev0.19.0
    • Addeddoctor

TDQS

A3.6/5.0

Scored across 44 tools

Disambiguation4/5

Most tools have sharply distinct purposes, and descriptions explicitly cross-reference siblings (install_uninstall vs install_purge, backup_list vs backup_drill) to steer selection. However a few boundaries blur, notably doctor/report (both produce health diagnoses), install_status vs install_list, and system_status vs docker_stats vs inventory_scan, which could be confused.

Naming Consistency4/5

Names follow a largely predictable prefix_verb/noun scheme grouped by subsystem (docker_*, install_*, backup_*, watch_*, proxmox_*), each readable and consistent within its group. Minor deviations exist in bare noun tools (alerts, doctor, report, processes, wake) that break the otherwise uniform convention.

Tool Count3/5

At 44 tools this is heavy, and the surface spans many subsystems (docker, proxmox, backups, installs, watching, network, reporting) so the total is large rather than redundant. Most tools earn their place, but the sheer count raises discovery and misselection cost.

Completeness4/5

Coverage is broad and deep: full docker control, backup lifecycle including restore and verification, install lifecycle, proxmox guest management, and monitoring/watching. The one obvious dead end is that docker_stop has no docker_start (acknowledged in the description), leaving containers needing an out-of-band start.

Maintenance

ActivityActive
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    MCP servers for managing homelab infrastructure. Monitor Docker/Podman containers, Ollama AI models, Pi-hole DNS, Unifi networks, and Ansible inventory.
    42
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    MCP server for homelab diagnostics + auto-update pipeline, managing Docker hosts via SSH with read-only diagnostics, image-drift visibility, and automated update execution with rollback.
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    A single MCP server that gives an AI assistant comprehensive access to manage a homelab, including SSH, Docker, Proxmox, Synology, Cloudflare, and more, with 85 tools and a centralized configuration.
    MIT