VPS-Guardian-MCP
VPS Guardian MCP is a safety-checked MCP server that lets AI agents observe, diagnose, and carefully mutate a Linux VPS through named tools instead of raw shell access.
System monitoring: full health snapshot (CPU, RAM/swap, disk, network, uptime), top processes, OOM events, kernel/hardware errors, disk-usage hotspots.
Services & logs: check systemd unit status, list failed units, read/filter service logs (systemd or Docker).
Docker: list containers, read container logs, live CPU/memory/IO stats.
Networking & security: open ports, UFW rules, failed SSH logins, Fail2ban status, SSH config audit, SSL/TLS certificate expiry, DNS health, outbound connectivity/latency benchmarks.
Files & configs: read, list, and atomically write (with timestamped backup) files in whitelisted config/web directories; test Nginx configs and enumerate virtual hosts.
Scheduling & updates: audit cron jobs and systemd timers, check OS security patches and Guardian's own version.
Backups: create compressed tar.gz archives of permitted site/config directories into an isolated repository.
Recovery (whitelisted mutations): restart services, kill runaway processes, prune Docker cache, clean logs/package cache, apply security updates, self-update — gated by
read-only/controlled/unrestrictedsafety modes with confirmation tokens.
Provides Docker container health monitoring (detecting exited, dead, restarting, or unhealthy containers), container-specific log retrieval, and recovery actions such as cleaning Docker caches.
Monitors Linux VPS system health metrics including CPU usage, load averages, RAM and swap usage, and root disk utilization.
Reads Nginx service logs and supports safely restarting the Nginx web server via systemctl or Docker.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@VPS-Guardian-MCPCheck VPS health and Docker container status"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
VPS Guardian MCP
VPS Guardian is a secure Model Context Protocol server for AI agents that work with Linux VPSs. It replaces an unrestricted “run this command and paste the result” loop with named, structured and safety-checked operations.
An agent can inspect a workload, collect bounded diagnostics, preview the impact of a change, and request an exact confirmation for a mutation. The server never exposes a general-purpose shell tool.
Explore the project: capabilities · agent workflows · security model · tool catalogue · release notes
Quick start
VPS Guardian has two parts:
The Python MCP server runs on the VPS.
The small npm launcher runs on the computer where Codex, Claude, Cursor or another AI client is installed. It opens an SSH stdio connection and never uploads the private key.
What you need
A Linux VPS reachable via SSH.
An SSH key for that VPS.
Python 3.10+ on the VPS and Node.js 16+ on the AI client's computer.
A verified SSH host key. Password-based SSH is intentionally unsupported by the launcher.
Optional: local MCP access panel
The English-language panel runs on your computer, not on the VPS or this project's website. Its main screen manages agent permissions: pause/resume MCP calls, cap the safety mode, allow individual tools and set project roots. A collapsed server overview provides optional read-only monitoring. The npm package remains the MCP launcher; the panel comes with the Python package.
With uv installed on your computer:
uvx --from vps-guardian-mcp==0.32.1 vps-guardian-panelWithout uv, install in a local virtual environment. Windows PowerShell:
py -m venv .guardian-panel
.\.guardian-panel\Scripts\python.exe -m pip install vps-guardian-mcp==0.32.1
.\.guardian-panel\Scripts\python.exe -m src.local_panelLinux/macOS:
python3 -m venv .guardian-panel
.guardian-panel/bin/pip install vps-guardian-mcp==0.32.1
.guardian-panel/bin/vps-guardian-panelThe browser opens automatically. Enter the VPS address, SSH user, port and path to your local key, not its contents. Leave the key field empty to use ssh-agent; load encrypted keys into the agent beforehand. Advanced settings accept the Guardian executable path on the VPS and an optional local known_hosts file. Access management requires Guardian 0.30.0+ on the VPS; Projects and Limits require 0.31.0+, Operations requires 0.32.0+, with the adjacent vps-guardian-access executable from the same installation. Older compatible installations can show monitoring only, with an upgrade warning. No server-side web service is installed.
To upgrade an existing pip installation on the VPS:
/opt/vps-guardian-mcp/.venv/bin/pip install --upgrade vps-guardian-mcp==0.32.1Upgrade any other Guardian environments used by your agents, then reconnect all agents once so they start the policy-aware server. Subsequent permission changes affect new calls in those sessions without a restart. In-flight operations are not cancelled, and cached client tool lists may still show disabled tools; calls to those tools are rejected.
Use Apply permissions to save a shared per-SSH-user policy on the VPS at ~/.local/share/vps-guardian-access/policy.json. Connection settings stay in local RAM, but the access policy persists across panel/server restarts. Reload policy fetches the current server values; concurrent edits are rejected rather than overwriting another operator's changes. No policy means the existing launch settings remain in force. Invalid or unsafe stored policies fail closed and pause normal MCP access; get_safety_status remains available for recovery. Unsafe file ownership, permissions or symlinks require manual repair by the operator.
Allow agent access: pause/resume normal MCP tool and resource entry points. The separate operator helper remains reachable over SSH so you can restore access.
Maximum safety mode: the effective mode is the stricter of the agent's launch mode and this cap. Setting
controlledcannot elevate a client launched asread-only. Confirmation tokens are the existing same-caller mechanism, not independent human approval.Allowed tools: allow all installed/future tools, or uncheck that option and choose an explicit allowlist. New tools are denied by default in an explicit allowlist.
get_safety_statuscannot be disabled. Read-only mode can still create bookkeeping records; use tool permissions when you also need to block those entry points.Project roots: inherit each process's
VPS_GUARDIAN_PROJECT_ROOTS, or replace it with up to 16 absolute VPS directory paths, one per line. An empty custom list blocks project workspaces. Filesystem roots and symlink roots are rejected. This controls project tools, not the separate fixed configuration-file directory whitelist.
The operator helper is not registered as an ordinary agent tool. Tools and the three read resources enforce access restrictions on the server. Policies apply to upgraded Guardian processes under the same SSH user, not to individually authenticated agent identities. An agent with independent root SSH access, the same account's shell, or permission to change Guardian's code can bypass this MCP boundary. For stronger separation use a restricted dedicated OS account; do not treat the panel as a sandbox or independent approval service.
Projects and live Limits
The Access, Projects and Limits sections require the current Python package on your computer. Projects and Limits additionally require Guardian 0.31.0+ on the VPS. A 0.30.0 server can still manage Access, but cannot enforce these new limits. Upgrade every server environment used by your agents and reconnect once; changing saved limits afterwards does not require another restart.
In Projects, click Find projects for a bounded metadata scan inside the current roots (at most 1,000 entries, 100 directories, 50 projects and a cooperative two-second deadline). Select a project or enter an absolute VPS path and click Inspect metadata. The panel shows top-level file/directory metadata, OS readability and policy decisions for common project tools, without reading source, invoking Git or executing project code. Links and sensitive/hidden names are excluded. Missing projects can be outside the roots, lack recognized markers or exceed the scan budget. When roots are inherited, the helper's launch defaults may differ from an agent's environment; use explicit managed roots for a shared boundary.
Use as the only project root only prepares a draft in Access. Review and click Apply permissions to replace the current roots with that single project. Browsing alone never grants or changes agent access. Policy allowance is not a guarantee that an agent's launch mode, roots or OS permissions permit an operation.
In Limits, choose Auto, Small VPS, Standard or Custom, then Apply limits. Values persist in the same private operator directory as limits.json, separately from policy.json. Revision checks reject concurrent edits. Auto and Standard currently request the same normal budgets; both retain automatic host guards. Small VPS requests smaller reads/searches, one-file patches and no Test Capsules. Custom accepts only the displayed integer ranges; it cannot disable guards or exceed hard ceilings. Reload live limits fetches a fresh server snapshot; it asks before discarding a local draft.
The table distinguishes Requested values from Effective on VPS values. Orange effective values have been reduced by host protection. The applied settings, not an unsaved draft, determine effective limits. Available memory below 512 MiB or one logical CPU reduces several budgets; below 384 MiB capsules are disabled; below 256 MiB additional critical-load restrictions apply. These are cooperative per-operation budgets, not a global CPU/RAM quota or OS sandbox. They do not cancel already running work. Every upgraded process under the same SSH user reads shared settings for subsequent operations. Invalid/unsafe limits use conservative Small VPS defaults and show an error; unsafe ownership/permissions or links may need manual repair.
Editable budgets cover project-file reads (up to 300,000 bytes when host guards permit), MCP tool/resource JSON, HTTP diagnostic bodies, directory entries, journal lines, code-search file/input budgets, project-patch file/staged-text budgets, configuration ChangeSet file/text budgets, Agent Job check concurrency (0 disables checks) and Test Capsules (0 disables, 1 permits). Existing stricter operation-specific bounds still apply. Capsule RAM/CPU/timeout/output bounds remain fixed, not editable. Resource response limits apply to each JSON document; MCP transport encoding and text/structured duplication add overhead. Large fields are omitted with response_truncated and omitted_fields, retaining small status/IDs/tokens where possible. A mutation has already returned: do not repeat it just to obtain omitted output. Request a narrower read or status instead.
Cached project patches recheck current roots and resource limits before preview, staging, testing and applying. A revoked or over-budget candidate must be replaced with an allowed, smaller patch. Configuration ChangeSets also recheck live limits before applying.
Shared Operations and reconnects
Operations requires Guardian 0.32.1+ on the VPS and on your computer. The English panel shows the latest 25 project patches/ChangeSets and up to 25 Agent Jobs, with search, state filters, file fingerprints and a bounded event timeline. Click a record to inspect metadata. Refresh is manual or every 30 seconds while the tab is open; it uses the separate short-lived operator helper, including while agent access is paused. This view does not execute, approve, cancel or retry changes. There is no additional always-on VPS worker.
Agents use list_operations(limit=10) and get_operation(operation_id="...") to retrieve the same metadata. next_before is an opaque cursor for older draft/history pages; pass it back as before. Job lists are separately bounded latest snapshots, not part of that cursor. Normal MCP tool permissions still apply; the human operator's helper is separate. History is per SSH user, not per individual agent or current project root: everyone granted these history tools under that account can see operation metadata.
Staged project patches and configuration ChangeSets now persist in private SQLite storage at ~/.local/share/vps-guardian-access/operations/. A draft survives MCP/panel reconnects for 24 hours, with the existing 24-patch/32-ChangeSet active caps. Active drafts are never silently evicted; storage-full errors require completing an allowed draft or waiting for expiry. Payloads are bounded to 9 MB each and 32 MB total; the SQLite database is capped at 48 MiB. Completed/expired source payloads are removed and their metadata retained for at most 30 days / 200 records. Retention is enforced on requests, not by a background cleanup daemon. Active candidate/original source can contain secrets: files use private permissions, not encryption, and must be protected with the SSH account and disk backups. Panel/history replies never include source, raw command output or confirmation tokens. Secret redaction is best-effort; avoid secrets in titles and paths.
After reconnecting, preview the same staged ID again to obtain a new session-local confirmation token, then apply with the existing tool. Current mode, roots, budgets and live-file fingerprints are rechecked; persistence does not grant access or approve a write. A claimed apply/check has a 120-second lease. If it outlives that lease without a stored outcome, history marks it uncertain, drops the candidate and never replays it. Inspect live files/services before preparing a replacement. This history is not a transactional filesystem journal or an automatic rollback guarantee.
New Agent Jobs use the same fixed private operation directory, independently of VPS_GUARDIAN_STATE_DIR; their existing leases, job TTL and 100-record cap remain. Their database is capped at 8 MiB. Upgrade note: previous releases kept jobs in the old state directory and drafts only in RAM. Old job databases remain untouched and are not automatically imported into shared storage; finish important old jobs before upgrading. Old in-memory drafts cannot be recovered after the old process exits. Upgrade all server environments used by agents and reconnect once.
SSH host-key verification is mandatory. Verify the fingerprint independently and connect once with ordinary SSH before using the panel. The dashboard does not accept unknown keys or use custom SSH config/aliases, ProxyCommand or jump hosts: enter a directly reachable IP/hostname. A changed host key must be investigated, not bypassed.
Connection settings and snapshots stay in memory; the panel never uploads or reads private-key contents. One MCP connection samples metrics every 30 seconds; manual refresh is limited to once per 10 seconds. Unavailable metrics are shown as unavailable, not zero. No failed systemd units is not a guarantee all applications are healthy. If the connection fails, retained metrics are marked stale.
Keep the printed local link private: its fragment is a per-launch operator access token. The token stays in tab memory and is removed from the address bar; after reloading, reopen the full printed link. Ctrl+C in the launch terminal stops the panel and its connection. --no-browser only prints the link; --port 8765 selects a loopback port. Do not reverse-proxy or expose the panel publicly. Loopback/token checks do not protect against malware running as your local user.
1. Install the server on the VPS
Run once on the VPS. This installs the published, pinned release:
sudo mkdir -p /opt/vps-guardian-mcp
sudo chown "$USER" /opt/vps-guardian-mcp
python3 -m venv /opt/vps-guardian-mcp/.venv
/opt/vps-guardian-mcp/.venv/bin/pip install --upgrade pip
/opt/vps-guardian-mcp/.venv/bin/pip install vps-guardian-mcp==0.32.1For development from source instead:
git clone --branch v0.31.0 https://github.com/murzirius/VPS-Guardian-MCP.git /opt/vps-guardian-mcp
cd /opt/vps-guardian-mcp
python3 -m venv .venv
.venv/bin/pip install -e .If the server runs as a non-root user, grant only the required read access. Docker and journal features gracefully report as unavailable when that access is absent.
sudo usermod -aG docker <MCP_USER>
sudo usermod -aG systemd-journal <MCP_USER>Log out and back in after changing groups.
2. Choose a safety mode
Mode | Use it when | Result |
| Inspecting or diagnosing | Default. Guarded server mutations are blocked; bookkeeping tools may still save records. |
| Assisted administration | Recommended. Each exact change needs a short-lived, single-use confirmation token. |
| A separately protected automation environment | Changes run immediately. Avoid on a general-purpose agent. |
Start with read-only; use controlled once the connection is verified.
3. Connect an AI client over SSH
First verify the VPS fingerprint independently and make a normal SSH connection once. That stores the host key in ~/.ssh/known_hosts (or %USERPROFILE%\.ssh\known_hosts on Windows). The launcher requires host-key verification by default.
Use this configuration for JSON-based MCP clients:
{
"mcpServers": {
"vps-guardian": {
"command": "npx",
"args": [
"-y",
"@murzirius/vps-guardian-mcp@0.32.1",
"--host", "<VPS_IP_OR_HOSTNAME>",
"--user", "root",
"--key", "~/.ssh/id_ed25519",
"--mode", "controlled",
"--tool-profile", "core"
]
}
}
}Every item in args is a separate argument. Do not join --host with its value or paste the entire command into one form field.
4. Codex / ChatGPT Desktop
Open Settings → MCP servers → Add server, choose STDIO, then enter:
Field | Value |
Name |
|
Command |
|
Environment variables | Leave empty |
Working directory | Leave empty/default |
Add these arguments as separate rows, in order:
-y
@murzirius/vps-guardian-mcp@0.32.1
--host
<VPS_IP_OR_HOSTNAME>
--user
root
--key
C:\Users\<WindowsUser>\.ssh\id_ed25519
--mode
controlled
--tool-profile
coreSave, restart the client, then use /mcp to confirm that vps-guardian is connected.
5. Common situations
Claude Code
claude mcp add vps-guardian -- npx -y @murzirius/vps-guardian-mcp@0.32.1 --host <VPS_IP_OR_HOSTNAME> --user root --key ~/.ssh/id_ed25519 --mode controlled --tool-profile coreA non-root SSH user — replace root after --user. Do not add passwordless sudo just for the MCP; grant the minimum group permissions needed.
A non-standard port — add separate arguments:
--port
2222A different server location — add:
--remote-path
/srv/vps-guardian/.venv/bin/vps-guardian-mcpHost key verification failed — do not disable verification. Check the VPS fingerprint through a trusted channel and correct known_hosts. Use --known-hosts <path> for a dedicated file. --accept-new-host-key is only for an intentional first-time bootstrap.
6. Verify and upgrade
Ask the agent: “Check CPU and RAM load on my server.” A correct setup returns structured VPS data rather than a shell command for you to run.
To upgrade the VPS server, install the matching version and restart the client connection:
/opt/vps-guardian-mcp/.venv/bin/pip install --upgrade vps-guardian-mcp==X.Y.ZThen replace @0.25.1 with @X.Y.Z in the client configuration. For source installations, fetch the tag, inspect local changes, check out the tag, and reinstall with .venv/bin/pip install -e ..
Related MCP server: Secure VPS Operations MCP Server
What it can do
VPS Guardian is built around a few workflows instead of a long, unstructured command list:
Observe: system pressure, processes, services, Docker, databases, ports, TLS, logs and updates.
Understand a workload: discover a site or Compose project, map its dependencies and health, then collect focused diagnostic evidence.
Coordinate agents: sessions, handoffs, leased work queues, durable Agent Jobs, runbooks, checkpoints, workload locks, maintenance windows and resumable server-event watches.
Change safely: preview impact, stage configuration changes, validate, back up, health-check and roll back when a deployment fails.
Recover deliberately: create baselines, compare drift, produce repair plans, verify isolated backups and require exact confirmation for changes.
Work with code: read a large file by line range, find Python symbols, stage a line edit, inspect a bounded Git diff and check a staged change in a temporary Docker capsule.
Understand an environment: inspect project venv metadata, compare direct dependencies and npm lock versions, and identify a running systemd service's launch path without executing project code.
Navigate Python code: map local imports, find symbol-use candidates, inspect likely change impact and gather short task context with related test candidates.
For smaller agent context, --tool-profile core exposes the everyday tools (including Agent Jobs); omit the flag or choose full for the complete catalogue. The launcher passes this profile to the server over SSH. Both profiles support compact JSON tool results, while new workload and log summaries return short answers by default. The profile takes effect when the MCP connection starts.
For a large project file, ask the agent to use get_project_symbols, then read_project_file_range around the relevant lines. A single range call returns at most 100 KB and includes a SHA-256 fingerprint. A subsequent stage_project_line_edit sends only changed lines and still uses the existing preview, confirmation, conflict check and backup flow. get_workload_brief, summarize_service_logs and get_server_event_delta provide compact operational context without background polling.
Agent Jobs: Create a job with 1-8 allowlisted checks, such as service_status and service_logs for target bot.service. Call advance_agent_job once per check. The job, bounded results, and progress survive MCP reconnects; get_agent_job(after_revision=...) returns only new results, and another authorized MCP client can continue by job ID. On Linux, advance_agent_job(background=true) starts just one detached read-only check that can finish after disconnection; it requires at least 512 MB available memory. After checking the exact service, an agent may propose one restart_service recovery. execute_agent_job_recovery uses the existing read-only/controlled/unrestricted safety mode; in controlled mode review its one-time confirmation and call again with the token. Then call verify_agent_job_recovery and record the conclusion. A possibly executed restart is never retried automatically after a disconnect. Jobs do not run an AI model, always-on worker, arbitrary shell commands, or automatic rollback on the VPS; a service restart cannot be undone. Existing reversible change tools retain their own backup and rollback rules.
Test Capsules: After begin_project_patch and stage_project_line_edit (or stage_project_file_change), call get_test_capsule_status, then test_project_patch(patch_id, check="auto"). In controlled mode, confirm this code-executing check with its own one-time token. auto syntax-checks staged Python or JavaScript files; python_unittest and npm_test explicitly run project tests. If it passes, call preview_project_patch to inspect the diff and obtain the separate apply confirmation, then promote_tested_project_patch with that token. A failed or edited candidate cannot be promoted through this tool. The existing apply_project_patch remains available for projects without Docker and does not claim a capsule test.
Capsules require a local Linux Docker daemon, an already-downloaded image (python:3.12-alpine or node:20-alpine by default) and at least 384 MiB available RAM. An operator may choose an already-local image with project dependencies via VPS_GUARDIAN_CAPSULE_PYTHON_IMAGE or VPS_GUARDIAN_CAPSULE_NODE_IMAGE. Guardian never pulls images or installs dependencies automatically. It copies at most 250 files / 8 MiB, omits common credential files and dependency directories, and allows one check at a time for 30 seconds. The container gets no network, host environment or live-project mount; CPU, RAM, processes and temporary storage are capped. Tests needing network, writable source files or missing dependencies will fail. Source files may still contain hard-coded secrets, so remove those before testing; Docker isolation reduces risk but is not a guarantee against malicious code or kernel vulnerabilities.
Environment Doctor: Ask the agent to call inspect_project_environment(project_path, service_name="bot.service"), then diagnose_project_dependencies for a focused problem report. It reads pyvenv.cfg, Python .dist-info/METADATA, static pyproject.toml/requirements.txt declarations, and direct npm dependencies against a v2/v3 package-lock.json. If both .venv and venv exist, supply an explicit authorized environment_path. include_dev=true includes npm development dependencies. Reuse after_fingerprint to receive only an unchanged reply when the relevant report has not changed. A dependency name is a distribution name, not necessarily its Python import name.
plan_environment_repair explains the next steps without installing packages or restarting anything. plan_capsule_environment prepares bounded direct dependency pins and runtime metadata for a reviewed local Capsule image; it does not build or download that image, verify its contents, or produce a complete transitive lockfile. Python versions come from venv metadata; running service evidence is limited to a Linux systemd MainPID. Stopped units, wrappers, containers, system Python without an authorized venv, inherited/legacy/editable packages, dynamic declarations, requirement directives, URLs, extras and unsupported npm locks may need separate review. Metadata is not proof that an import works, and Node runtime/engine compatibility is not probed. Reads are capped at 2 MB per report, 1,000 directory entries and 100 direct declarations per ecosystem; incomplete scans never report a clean result. No project interpreter, installer, npm script, package index or background worker is started.
Code Navigator: Start with get_project_import_map(project_path, relative_path="app/payments.py"), or use get_project_task_context(project_path, relative_path="app/payments.py", symbol_name="charge") to get a definition, short use-site fragments and related test candidates in one bounded answer. find_project_references distinguishes import-alias candidates from weaker name-only matches. assess_project_change follows reverse imports for up to three hops; test files are selected by import relationships and naming conventions, not measured coverage. For a class method use its exact qualified name, such as Gateway.refund; get_project_symbols helps choose it. All four tools accept after_fingerprint for compact unchanged replies; changing query arguments changes the fingerprint. No scan results are cached, so an unchanged reply saves output tokens but still requires a fresh scan.
No setup is needed beyond configured project roots. Navigator supports UTF-8 Python sources and bounded ASCII symbol names, including relative imports and common src/ layouts. Snippets are limited to one line; literals (including f-strings/template strings) and comments are masked. File, module and identifier names remain visible; use existing range reads for exact source bodies. This is not a runtime call graph, complete dependency analysis or proof that tests cover a change: alias shadowing, ambiguous modules, dynamic imports/reflection, instance types and non-Python/generated sources remain uncertain. Hidden, sensitive and dependency paths are excluded. Unreadable files, unsafe paths, syntax errors and exhausted budgets are reported as partial scans.
Scans are on demand, with no background indexer or disk cache: at most 150 Python files (60 on constrained hosts), 128 KB per file, 2 MB total (750 KB constrained), 1,000 directory entries, eight directory levels and a five-second cooperative scan deadline. AST nodes and extracted facts are capped; files with more than 200 import aliases are skipped. Below 96 MiB available memory, no scan starts. Only one scan per MCP process runs at a time. Separate MCP processes do not share this lock. Reference/import reports return at most 50 entries; task context defaults to 6,000 characters and can be capped between 2,000 and 8,000.
Examples of native MCP tools:
Request | Example tool | Result |
“Why is the API slow?” |
| Bounded health, logs, OOM and kernel evidence. |
“What will a restart affect?” |
| A read-only dependency and impact report. |
“Hand this incident to another agent.” |
| Secret-redacted context and outcome tracking. |
“Split this audit between agents.” |
| Prioritized work with dependencies and expiring ownership. |
“Check this service, then let another agent continue.” |
| Durable, bounded checks with delta results and gated recovery. |
“Deploy this Nginx change safely.” |
| Validated diff, backup, confirmation and rollback path. |
See the complete capability guide and full tool catalogue on the project site.
Security model
No arbitrary command-execution MCP tool.
Server-side allow-lists for files, paths, services and mutation types.
controlledmode uses parameter-bound, single-use confirmation tokens. This is not independent human approval: the caller receives the token and can repeat the operation. Client approval or a separately enforced operator policy is required for that guarantee. An agent with unrestricted SSH access can bypass MCP restrictions.Common secret patterns are redacted from file reads, sessions, audit data and diagnostic output. Redaction is best-effort, not a guarantee against every hard-coded credential or sensitive identifier.
Reads, logs, directory scans and stored state are bounded for small VPSs.
On Linux, hardened project/config reads refuse symlink components and special files. Atomic writes pin the parent directory, refuse unsafe backup paths, use private unique backups and preserve normal ownership/permissions without setuid/setgid bits. Existing private state directories must be owned by the service user with 0700, state/audit files with 0600; unsafe paths fail closed rather than being silently chmodded. Audit files stop accepting writes at 4 MiB each (primary/fallback), and reads inspect at most a 256 KiB tail; an operator must archive/reset full logs. The fallback audit filename is scoped to the effective UID. Python syntax checks use bounded source snapshots with an isolated interpreter, skip oversized files and never claim full success for incomplete scans. These protections do not make root execution or a shared SSH key an isolation boundary.
Details: security model · agent operating guide
Packages and releases
PyPI:
vps-guardian-mcpMCP Registry:
io.github.murzirius/vps-guardian-mcpGitHub Packages mirrors each npm release; npmjs is recommended for normal installation.
Development
python -m unittest discover -s tests -v
npm testPlease report security issues privately rather than publishing exploit details in a public issue.
License
MIT © 2026 murzirius.
Available Tools
31 toolsanalyze_disk_usageA
Analyze disk usage for a directory to discover space bottlenecks and large files.
Safely walks the filesystem without following symlinks and automatically skips virtual pseudo-filesystems (/proc, /sys, /dev, /run).
Args: target_path: Starting path to inspect (defaults to '/var'). max_depth: Depth of directory nesting to inspect (1 to 5, default 2). min_size_mb: Minimum size threshold in megabytes to include (default 50 MB). top_n: Maximum number of largest items to return (1 to 50, default 15).
Returns: JSON string with partition usage, largest directories, and largest files.
| Name | Required | Description | Default |
|---|---|---|---|
| top_n | No | ||
| max_depth | No | ||
| min_size_mb | No | ||
| target_path | No | /var |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility for behavioral disclosure. It does disclose meaningful behavior: it safely walks the filesystem, does not follow symlinks, and skips virtual pseudo-filesystems like /proc, /sys, /dev, and /run. It also states the return format as a JSON string. It does not mention permission requirements or exact error conditions, but the core behavioral traits are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and efficient. It front-loads the purpose, then provides a concise safety note, clearly labeled Args, and a clearly labeled Returns section. Every sentence provides necessary information and no filler is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only analysis tool with four optional parameters, the description covers all parameters, explains traversal behavior, notes exclusions, and specifies the return format. Even though an output schema exists, the description still appropriately summarizes what the returned JSON contains. Nothing critical is missing for correct tool invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must compensate, and it does. Every parameter is explicitly explained with its meaning, default, and applicable constraints: target_path with default '/var', max_depth with range 1-5, min_size_mb with threshold meaning, and top_n with range 1-50. This fully clarifies the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Analyze disk usage for a directory to discover space bottlenecks and large files.' It clearly distinguishes itself from sibling tools focused on logs, processes, security, or services, and immediately conveys what this tool does and why it exists.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: when disk-usage analysis or space-bottleneck discovery is needed. It does not explicitly name alternatives or exclusions, but none of the sibling tools clearly overlap with this disk-analysis function, so the lack of explicit when-not-to-use guidance is acceptable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_ssh_configA
Audit the SSH daemon configuration against security best practices.
Returns: JSON string with detected settings, security score (0-100), and remediation guidance.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must convey behavior. It states it 'audits' and returns results, implying a read-only operation, but it does not explicitly guarantee no modifications are made.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise, containing only two sentences. It quickly communicates the purpose and return value without any filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity and lack of parameters, the description is largely complete: it explains the action, the target (SSH daemon configuration), and the output format. It lacks explicit mention of scope (local system), but this is reasonably implied.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters, so the empty input schema leaves no room for missing documentation. The description accurately reflects that the tool operates without arguments, achieving the baseline for zero-parameter tools.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool audits the SSH daemon configuration against security best practices, specifying the action and resource. It is distinct from all sibling tools, which focus on other system aspects.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for security auditing but does not explicitly state when to prefer this tool over alternatives. However, since no sibling tool covers SSH config audits, the context is somewhat clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_dns_healthA
Audit system DNS resolution health, configured nameservers, and query responsiveness.
Args: domains: Optional custom list of domains to probe. Defaults to essential public services.
Returns: JSON string with configured nameservers, individual domain lookup latencies, and health verdict.
| Name | Required | Description | Default |
|---|---|---|---|
| domains | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It communicates that the tool performs DNS probes, can accept custom domains, defaults to public services, and returns a JSON verdict. 'Audit' implies a non-mutating diagnostic operation, which is useful transparency even though explicit safety or permission details are not stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured with Args and Returns sections. The main purpose is front-loaded, every sentence adds useful information, and there is no repetition of schema fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-optional-parameter diagnostic tool, the description covers purpose, input behavior, default behavior, and return content. It is complete enough for an agent to invoke it confidently. Minor gaps such as timeout behavior or possible failure modes are not described, but the output schema may cover return details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only defines 'domains' as an optional array/null with no description. The tool description compensates by explaining that domains is a custom probe list and defaults to essential public services. This adds real meaning beyond the schema, though it does not specify domain formatting constraints or example values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Audit') and clearly identifies the resource: system DNS resolution health, configured nameservers, and query responsiveness. This distinguishes it from the sibling health/network tools, which target different subsystems.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: when DNS resolution health, nameserver configuration, or query latency needs to be assessed. It does not explicitly name alternatives or provide when-not-to-use guidance, but the purpose is specific enough for an agent to route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_failed_loginsA
Inspect recent failed SSH login attempts to detect brute-force attackers.
Args: limit: Number of recent failed attempts to inspect (default: 20, max: 100).
Returns: JSON string with recent failed logins and top offending attacker IP addresses.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of explaining behavior. The verb 'Inspect' conveys a non-destructive read operation, and the description states what will be returned: a JSON string with failed logins and top offending IPs. It does not discuss authentication or data source, but nothing suggests hidden side effects or destructive behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured with clear Args and Returns sections. Every sentence adds useful information, and the main purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only tool with an output schema present, the description is sufficiently complete. It covers what the tool does, what the parameter means, and what kind of result to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, limit, is fully explained in the description: it is the number of recent failed attempts to inspect, with a default of 20 and a max of 100. This adds essential information not present in the schema, especially since schema description coverage is 0%.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Inspect'), a clear resource ('recent failed SSH login attempts'), and a goal ('detect brute-force attackers'). This separates it from sibling tools like get_fail2ban_status and audit_ssh_config.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: when checking recent failed SSH login attempts to detect brute-force activity. It gives clear context but does not explicitly mention alternatives or when-not-to-use conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_guardian_updatesA
Check if a newer version or commit of VPS-Guardian-MCP is available on GitHub.
Provides automated version verification and action guidance for self-updating. Cached for 5 minutes to minimize network and CPU overhead.
Args: force_refresh: Set to True to bypass cache and query GitHub API directly.
Returns: JSON string with current version, latest commit, update availability, and AI warning notice.
| Name | Required | Description | Default |
|---|---|---|---|
| force_refresh | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses caching behavior, network/CPU overhead, direct GitHub API usage, and the JSON return contents. It does not explicitly state that the operation is read-only or describe failure behavior, but the disclosed traits go well beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, followed by caching behavior, parameter details, and return format. Every sentence adds information and the structure is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter optional tool with no output schema, the description covers purpose, caching, the parameter, and the return fields. It is slightly less explicit about exact JSON key names or error handling, but it provides enough for an agent to invoke the tool and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema only provides a title and default for force_refresh, but the description adds meaningful semantics: 'Set to True to bypass cache and query GitHub API directly.' This explains the behavioral consequence of the parameter, which the schema alone does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Check if a newer version or commit of VPS-Guardian-MCP is available on GitHub.' This clearly distinguishes the tool from the unrelated server-operation siblings and states exactly what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context by explaining the 5-minute cache and explicitly telling the agent to set force_refresh to True to bypass cache and query GitHub API directly. It does not explicitly compare with alternatives, but no sibling tool offers this functionality, so the guidance is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_kernel_errorsA
Audit kernel logs for hardware failures, storage I/O errors, or application segfaults.
Args: limit: Maximum number of error entries to retrieve (1 to 50, default 20).
Returns: JSON string with categorized kernel errors, root causes, and critical issue counters.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description must carry behavioral context. It states the operation audits logs and describes the return as a JSON string with categorized errors, root causes, and critical counters. It does not explicitly mention side effects, permissions, or rate limits, though the read-only nature is reasonably inferred from 'audit'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact: one purpose sentence, one parameter line, and one return line. Every sentence adds information, and the key action and resource are front-loaded with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only diagnostic tool with an output schema, the description covers purpose, parameter semantics, and return shape. It lacks explicit alternative guidance and behavioral caveats, but nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has no property descriptions (0% coverage), so the description's Args section is the only source of parameter meaning. It defines limit as the maximum number of error entries and gives the range 1–50 and default 20, fully compensating for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action and resource: 'Audit kernel logs for hardware failures, storage I/O errors, or application segfaults'. The error categories give concrete purpose, but it does not explicitly distinguish itself from overlapping siblings such as check_oom_events or read_service_logs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as check_oom_events or get_system_health. The purpose sentence implies kernel-log triage, but no exclusions, contexts, or alternative routing are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_oom_eventsA
Inspect kernel logs for Linux Out-Of-Memory (OOM) Killer invocations.
Surfaces terminated processes, PIDs, and consumed RSS memory at time of termination.
Args: limit: Maximum number of recent OOM events to return (1 to 50, default 10).
Returns: JSON string with detected OOM incidents and diagnostic summary.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It uses the word 'Inspect' suggesting a read-only operation and describes the output, but it does not explicitly confirm no side effects, does not mention permissions required, and does not describe failure modes (e.g., missing logs, insufficient permissions). This is adequate but not thorough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured, with a clear purpose statement, a one-line parameter explanation, and a return-value note. No redundant or extraneous text is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple nature of the tool (one optional parameter, read-only intent), the description is nearly complete. It explains what the tool does and what it returns, but it omits details about behavior when no OOM events exist or how the JSON output is structured. These gaps are minor for this diagnostic tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter 'limit' has no schema description (0% coverage), but the description fully compensates by explaining its meaning, range (1 to 50), and default value (10). The parameter is effectively documented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Inspect kernel logs') and the specific subject ('Linux Out-Of-Memory (OOM) Killer invocations'). It also explains what is surfaced (terminated processes, PIDs, and RSS memory), leaving no ambiguity about the tool's purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly state when to use this tool versus alternatives. While the focus on OOM events implies a diagnostic use case, it lacks direct guidance such as 'use this when checking for memory pressure' or 'use check_kernel_errors for general kernel issues'. No when-not-to-use conditions are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_service_statusA
Check the operational status of a systemd service unit.
Args: service_name: Name of the system service (e.g. 'nginx', 'mysql', 'postgresql', 'ufw', 'docker').
Returns: JSON string with active state ('active', 'inactive', 'failed'), enabled state, and recent status logs.
| Name | Required | Description | Default |
|---|---|---|---|
| service_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It helps by stating the return payload: JSON with active state, enabled state, and recent status logs. Yet it does not explicitly disclose that the operation is read-only, what happens when the service does not exist, or whether elevated permissions are required to query status. This leaves some behavioral ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and efficiently structured with Args and Returns sections. The first sentence immediately states the tool's purpose, and every subsequent sentence adds necessary details about the parameter and output. There is no inflated or redundant wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one required string parameter, the description is largely complete: it names the service, gives examples, and describes the expected return content. It does not explain error cases or permission requirements, but the output schema already covers the return shape, and the tool's scope is narrow. A small gap remains in not explicitly distinguishing from failed-unit or service-log tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does. The Args section defines service_name as 'Name of the system service' and gives concrete examples ('nginx', 'mysql', 'postgresql', 'ufw', 'docker'). This adds real meaning beyond the schema's bare 'Service Name' title, though it could have specified whether the '.service' suffix is accepted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Check the operational status of a systemd service unit.' It clearly distinguishes this from sibling tools like list_systemd_timers, get_failed_systemd_units, and read_service_logs by focusing on the status of a single named service unit. The Returns section further clarifies what is inspected: active state, enabled state, and recent status logs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies the use case: you need the operational status of a specific systemd service. It provides a named service parameter with examples, which is enough context to select it over broader tools like get_system_health or get_failed_systemd_units. However, it does not explicitly name alternatives or provide when-not-to-use guidance, so it misses full 5-level routing clarity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_ssl_certificatesA
Audit SSL/TLS certificates configured on the host (Let's Encrypt / Certbot).
Returns: JSON string listing domains, expiration dates, days remaining, and warning flags.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the audit scope (host certificates, specifically Let's Encrypt / Certbot) and the exact return shape (JSON string with domains, expiration dates, days remaining, and warning flags). The verb 'Audit' implies a read-only operation, but it does not explicitly state that there are no side effects or mention any permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences: a clear purpose statement followed by a return-value summary. It is front-loaded and every sentence adds necessary information with no redundant text or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool, the description is complete: it states the exact subject (host SSL/TLS certificates, Let's Encrypt / Certbot) and enumerates the return fields. An output schema reportedly exists, so return format does not need explanation. There are no missing invocation details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so the baseline is 4. There is no parameter meaning to add, and the description correctly avoids inventing parameters. Nothing is lacking on this dimension.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Audit') with a clear resource ('SSL/TLS certificates configured on the host') and adds scope as Let's Encrypt / Certbot. It stands apart from all sibling tools, none of which target SSL/TLS certificate state, so an agent can confidently identify what this tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied by the tool's name and description: it is for auditing SSL/TLS certificates. However, there is no explicit statement of when to use this tool versus alternatives, nor any exclusion or comparison to sibling audit/health tools. That is implied usage, not explicit guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_system_updatesA
Audit available operating system package updates and pending security patches.
Checks reboot requirements (/var/run/reboot-required), total upgradable packages, and security CVE patches. Cached for 5 minutes to minimize CPU and disk usage.
Args: force_refresh: Set to True to bypass the 5-minute cache and query package managers directly.
Returns: JSON string with update counts, security status, reboot flag, and recommended recovery action.
| Name | Required | Description | Default |
|---|---|---|---|
| force_refresh | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses read-only audit behavior, checks the reboot-required file, queries package managers/CVE information, caches results for 5 minutes, and explains that force_refresh bypasses the cache. It does not explicitly state permission requirements or that it makes no system changes, but 'audit/checks/returns' strongly implies a non-mutating operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured. It front-loads the purpose, then adds concrete details about what is checked, caching, parameters, and return value. Every sentence earns its place with no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter audit tool with an output schema, the description is largely complete: purpose, checked sources, caching behavior, parameter semantics, and return content are all covered. Minor gaps include platform or root-permission requirements and potential failure modes, but these are not critical given the tool's low complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides 0% description coverage, and the description fully compensates for the only parameter. It explains that force_refresh=True bypasses the 5-minute cache and queries package managers directly, which gives the agent meaningful behavioral knowledge beyond the schema's default value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Audit available operating system package updates and pending security patches.' It is clear and distinct from most siblings, and enumerates concrete checks (reboot, upgradable packages, CVE patches). It does not explicitly contrast with the close sibling check_guardian_updates, so it stops short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: whenever OS package update or security patch status is needed. It provides useful operational context like the 5-minute cache and force_refresh behavior, but it never explicitly states when not to use it or names alternatives like check_guardian_updates. Usage guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_backupA
Create a compressed tar.gz archive of an authorized website or configuration directory.
Archives are saved into an isolated backup repository (/var/backups/vps-guardian/). Permitted source locations: /var/www/, /etc/nginx/, /etc/mysql/, /etc/postgresql/, /etc/docker/, /etc/caddy/
Args: backup_type: Identifier label for the archive (e.g. 'site', 'config', 'data'). source_path: Target directory to archive.
Returns: JSON string with archive file path, size, file count, and duration.
| Name | Required | Description | Default |
|---|---|---|---|
| backup_type | Yes | ||
| source_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden. It explains that a tar.gz archive is created, where it is stored, and what the return JSON contains. It does not mention overwrite behavior or permission prerequisites, but it is reasonably transparent for a non-destructive operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, front-loaded with the primary action, and uses a clear structure: action, destination, allowed source locations, Args, Returns. Every sentence contributes useful information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the essential aspects needed to call the tool: purpose, source restrictions, destination, parameters, and return format. Since an output schema exists, the return description is a helpful addition rather than a necessity. Minor gaps such as failure conditions and naming conventions prevent a perfect score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides zero description coverage, so the description must fully explain both parameters. It does: backup_type is defined as an identifier label with examples, and source_path is defined as the target directory. This is sufficient for an agent to populate both required arguments correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Create a compressed tar.gz archive of an authorized website or configuration directory.' It clearly distinguishes this from all sibling tools, which are mostly read-only audit and monitoring tools. The scope is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides concrete usage constraints by listing permitted source locations and stating that archives are saved to an isolated backup repository. It does not explicitly name alternative tools or state when not to use it, but the allowed paths and purpose provide clear contextual guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
execute_recoveryA
Execute an emergency recovery operation from a strictly whitelisted list.
Allowed actions:
'restart_service': Restarts a systemd service (requires target=service_name, e.g. target='nginx').
'clean_docker_cache': Deep prune of unused containers, networks, images, and volumes.
'clean_system_logs': Prunes journal logs older than 3 days and rotated archives in /var/log.
'kill_process': Terminates a runaway process by PID (requires target=PID, e.g. target='12345').
'restart_nginx': Restarts Nginx web server (legacy alias for restart_service target='nginx').
'vacuum_systemd_journal': Truncates journal logs to limit (target defaults to '200M').
'clean_package_cache': Cleans APT archive cache and removes obsolete packages.
'apply_security_updates': Non-interactively applies pending operating system security updates.
'update_guardian': Self-updates VPS-Guardian-MCP from GitHub and refreshes virtual environment.
Args: action_name: The exact recovery action to execute. target: Optional target parameter required by certain actions.
Returns: JSON string with operation outcome, freed resources, or security error.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | ||
| action_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full transparency burden and does it well: it lists exact side effects such as 'Deep prune of unused containers, networks, images, and volumes', 'Terminates a runaway process by PID', and 'Truncates journal logs to limit'. It also discloses the return format including 'security error'. It stops short of stating irreversibility or potential service-impact severity, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the primary purpose and then organized into a compact bullet list, an Args section, and a Returns section. Every sentence carries meaningful information; the length is justified by the number of actions described. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, the output schema presence, and the absence of annotations, the description covers all actions, their target semantics, and the return shape. It could additionally note prerequisites like root permissions or the fact that several actions are irreversible, but the current text is sufficiently complete for an agent to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must fully compensate. It does so by enumerating every allowed action_name value with concrete target requirements and examples such as target='nginx' and target='12345'. This is far more informative than the schema's bare strings and gives an agent everything needed to construct valid calls.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Execute an emergency recovery operation from a strictly whitelisted list.' It clearly differentiates itself from the read-only diagnostic sibling tools such as check_service_status and analyze_disk_usage by framing this as a recovery/action tool. The bullet-list of allowed actions makes the tool's purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'emergency recovery operation' establishes clear context for when this tool is appropriate, and the 'strictly whitelisted' phrasing imposes a strong constraint against arbitrary actions. It does not name sibling alternatives or explicitly say when not to use them, but the contrast with the read-only sibling names is evident. The per-action target requirements also give concrete operational guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_database_healthA
Discover running databases and verify responsiveness, latency, and socket states.
Detects Redis, PostgreSQL, MySQL/MariaDB, and SQLite databases in application directories.
Returns: JSON string with operational state, socket accessibility, and ping latency for each engine.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does a solid job: it reveals the discovery behavior, supported database types, scope ('in application directories'), and the exact JSON return content including operational state, socket accessibility, and ping latency. It does not explicitly state that it is read-only, but 'discover and verify' strongly implies non-destructive behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: first the core purpose, then the detection scope, then the return format. Every sentence adds useful information and there is no filler or redundancy that weakens the definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter health-check tool with an output schema, the description is largely complete: it defines scope, supported engines, and return contents. It leaves minor ambiguity about whether databases must be local or whether credentials are needed, but these are not critical for an initial discovery/health check call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is no parameter semantics burden on the description. The baseline for no-parameter tools is 4, and the description appropriately focuses on behavior and return value instead.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action with a clear resource: discovering running databases and verifying their responsiveness, latency, and socket states. It names the exact database engines covered, which further distinguishes it from generic system health tools among the siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for database health checks by listing supported engines and return details, but it does not explicitly say when to prefer it over siblings like get_system_health or check_service_status. There is no when-not-to-use guidance or mention of alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_docker_container_logsA
Safely read stdout/stderr logs from a specific Docker container.
Args: container_name: Container name or container short/full ID. lines_count: Number of recent log lines to retrieve (default: 50, max: 1000).
Returns: JSON string containing the container logs.
| Name | Required | Description | Default |
|---|---|---|---|
| lines_count | No | ||
| container_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations supplied, the description carries the safety burden; 'Safley read' and the limitation to 'recent log lines' with a maximum of 1000 convey a read-only, bounded operation. It also discloses the return type as a JSON string. It does not mention error cases or prerequisites, but the core behavioral profile is clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the one-sentence purpose, and organized into Args/Returns sections with no filler. Every sentence contributes useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read tool with an output schema and no annotations, the description covers purpose, parameter semantics, output format, and operational bounds. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description fully compensates: it explains that container_name accepts a name or short/full ID, and that lines_count controls recent-line volume with default and max. This adds meaning well beyond the bare schema properties.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('read'), a specific resource ('stdout/stderr logs from a specific Docker container'), and clearly distinguishes itself from siblings like read_service_logs and list_docker_containers. The scope is explicit and immediately actionable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes clear the tool applies to Docker containers, but it never explicitly says when to choose it over read_service_logs or list_docker_containers, nor does it mention exclusions. Usage guidance is only implied by the tool name and 'Docker container' wording.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_docker_statsA
Retrieve live resource utilization metrics for all running Docker containers.
Provides real-time CPU %, Memory %, Network I/O, and Block I/O (equivalent to docker stats).
Returns: JSON string listing resource metrics per running container.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It clearly indicates a read-only retrieval operation ('Retrieve', 'Provides', 'Returns') and discloses the return format (JSON string) and the coverage (all running containers, not stopped ones). It does not mention permissions or failure modes, but for a non-mutating stats tool these are minor omissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently written: one sentence for the action, one for the specific metrics and docker-stats equivalence, and one for the return format. Every sentence adds value without repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no parameters, no annotations, and an output schema is present, the description is fully sufficient. It states what the tool does, what metrics are included, and what the return looks like. No critical operational context is missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema is an empty object, so there are no parameter semantics to document. The baseline for zero-parameter tools is 4, and the description correctly avoids inventing parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Retrieve') and resource ('live resource utilization metrics for all running Docker containers'), and distinguishes itself from sibling tools like list_docker_containers and get_docker_container_logs by stating exactly what it returns: CPU, memory, network, and block I/O metrics. The 'equivalent to docker stats' reference further anchors its purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool: when live resource metrics for all running Docker containers are needed. It does not explicitly name alternatives or exclusion criteria, but the scope is unambiguous and no conflicting tool is suggested. This satisfies 'clear context, no exclusions'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_fail2ban_statusA
Check Fail2ban status, active protection jails, and currently banned IP addresses.
Returns: JSON string detailing active jails and banned IP addresses.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose that the tool returns a JSON string with active jails and banned IPs, which is useful. However, it does not mention whether elevated privileges are needed, how the tool behaves if Fail2ban is not installed, or whether this is strictly read-only.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with high signal-to-noise ratio. The brief redundancy between 'active protection jails...' and 'active jails and banned IP addresses' is minor and does not harm comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and an output schema present, the description is largely sufficient for a simple status-check tool. It names the key return content and the tool's focus, though it does not explicitly cover failure modes or privilege requirements.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description does not need to explain any parameter semantics. The baseline of 4 applies because there is nothing to document.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states a specific action ('Check') and resource ('Fail2ban status, active protection jails, and currently banned IP addresses'). This distinguishes it from sibling tools like check_failed_logins, which focus on login attempts rather than current ban state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives such as check_failed_logins or audit_ssh_config. While the purpose is clear, the description does not help an agent choose between overlapping security-related tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_failed_systemd_unitsA
Find all degraded or failed systemd services across the entire system.
Returns: JSON string with list of failed units ('systemctl --failed') and overall health indicator.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It discloses that the tool is read-only in nature by using 'Find' and explicitly references systemctl --failed, which is a non-mutating command. It also states the output shape (list of failed units and health indicator). It does not mention permissions or potential limitations, but for a simple diagnostic getter this is reasonably transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and efficiently structured. It leads with the action and scope, then separately provides the return contract. Every sentence adds value, and there is no redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless read-only diagnostic tool with an output schema, the description is complete. It identifies what is checked, the mechanism (systemctl --failed), and the return components. No additional context is necessary for an agent to invoke or interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is nothing ambiguous for the agent to resolve. The schema already covers everything; the description does not need to add parameter-level meaning. This matches the baseline for a parameterless tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Find') and resource ('degraded or failed systemd services') with a clear scope ('across the entire system'). This clearly differentiates from siblings like check_service_status, which targets a single service, and get_system_health, which is broader.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a system-wide health check use case, but it does not explicitly state when to prefer this over siblings such as check_service_status or get_system_health, nor does it mention exclusions or alternatives. The 'across the entire system' phrase provides some context, but the guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_open_portsA
Discover all listening network ports (TCP and UDP) and identify bound processes.
Returns: JSON string listing open ports, protocols (TCP/UDP), binding addresses (IPv4/IPv6), and process names/PIDs.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It states a read-oriented action and the return payload, but does not explicitly confirm it is read-only or note any privilege/permision requirements that could affect a successful call. Some context is added via the expected output, but explicit safety/ide-effect disclosure is missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: one action sentence followed by a clear 'Returns' list. It front-loads the core purpose and avoids unnecessary detail, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with an output schema, the description covers the essential behavior and return content. The main gap is the lack of an explicit read-only/privilege note, but the verb 'Discover' strongly implies a no-side-effect investigation, so this is a minor omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool accepts no parameters, so schema description coverage is trivially 100%. A baseline of 4 applies for zero-parameter tools, and the description appropriately focuses on output semantics rather than parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Discover') and clear resource ('ll listening network ports (TCP and UDP)') with a defined scope and output. This clearly distinguishes it from sibling tools like test_network_connectivity or get_ufw_status, which serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use guidance is provided. The usage is implied by the action described (inspect listening ports), but there is no mention of alternatives or conditions that would make this tool preferred over others in the sibling set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_system_healthA
Retrieve a complete system health snapshot of the Linux VPS.
Returns a JSON string containing:
CPU: overall percentage, per-core breakdown, core counts, 1/5/15m load averages.
RAM & Swap: total, used, available, percentage.
Disk: root partition usage, read/write I/O counters.
Network: sent/received bytes, packets, and error counts.
Uptime: boot timestamp and human-readable duration (e.g. '12d 4h 32m 10s').
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden for behavioral disclosure. It explicitly states that the tool returns a JSON string and details exactly which categories of data are included, including load averages, I/O counters, and uptime. It does not mention potential error conditions or latency, but for a zero-parameter read-only health snapshot this is reasonably transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and then uses a scannable bullet list to detail the return contents. Every line adds useful information, and the example uptime format helps clarify the expected shape. No filler or redundancy exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no parameters and a rich output schema, the description fully covers what an agent needs to know: what the tool does and what data it returns. The presence of an output schema means the agent can also inspect the exact return structure separately, and the description complements it well.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description need not explain parameter behavior. The baseline for a no-parameter tool is 4, and the description appropriately focuses on the output rather than input semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Retrieve a complete system health snapshot of the Linux VPS.' It enumerates CPU, RAM, Disk, Network, and Uptime, which clearly distinguishes this broad health tool from more targeted siblings like get_top_processes, analyze_disk_usage, or get_docker_stats.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The purpose is unambiguous enough that an agent can infer when to use it: whenever a broad system health overview is needed. It does not explicitly state when not to use it or name alternative tools, but the 'complete system health snapshot' framing provides clear context without excluding edge cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_top_processesA
Retrieve the top resource-consuming processes running on the VPS.
Args: sort_by: Metric to rank processes by ('cpu' or 'memory'). Default: 'cpu'. limit: Number of top processes to return (1 to 50, default: 10).
Returns: JSON string listing process PID, name, user, CPU %, RAM %, RSS memory, and command summary.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| sort_by | No | cpu |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It clearly describes the operation as a retrieval, documents the sort_by options and limit range, and specifies the return format and fields, making the tool's behavior predictable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, well-structured with Args and Returns sections, and every sentence adds relevant information. It front-loads the core purpose and then efficiently covers parameters and output.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read-only tool, the description is complete: it defines both parameters, their constraints, defaults, and the return shape. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema only provides types and defaults, with zero description coverage. The description compensates fully by explaining valid sort_by values ('cpu' or 'memory'), the limit range (1 to 50), defaults, and what data is returned.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Retrieve') and resource ('top resource-consuming processes running on the VPS'), which clearly distinguishes it from sibling monitoring tools. It immediately conveys the tool's scope and purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: when you need a ranked list of the most resource-consuming processes on the VPS. However, it does not explicitly name alternatives or state when not to use it, so it stops short of a perfect score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_ufw_statusA
Inspect the status and active filtering rules of the UFW firewall.
Returns: JSON string containing UFW active state, default incoming/outgoing policies, and all active firewall rules.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden of explaining behavior. It discloses that the tool inspects rather than modifies and explicitly states the return format (JSON string with active state, policies, and rules). This is adequate for a simple read-only status tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded, with the core purpose in the first sentence and return-value details in a clearly separated block. Every sentence adds value without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only status tool, the description is complete. It covers what the tool does and what output to expect, and the presence of an output schema reduces the need for further return-value explanation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema already fully covers parameter semantics. The description does not need to add parameter details; a baseline of 4 is appropriate for this no-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: inspecting UFW firewall status and active filtering rules. It identifies a specific resource (UFW firewall) and differentiates itself from sibling tools, none of which target UFW.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool: any time UFW firewall status, default policies, or active rules are needed. It does not explicitly discuss alternatives, but its unique scope among siblings makes the intended usage unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_cron_jobsA
Discover all scheduled cron jobs on the Linux system.
Audits /etc/crontab, /etc/cron.d/, /etc/cron.* periodic scripts, and user crontabs.
Returns: JSON string containing scheduled jobs with user, schedule expression, human-readable timing explanation, and command.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does well by disclosing exactly which cron sources are audited and the return structure. It does not mention potential permission requirements or failure behaviors, but the read-only nature is clear and the listed audit paths are strong behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: a one-line purpose, a concise list of audited locations, and a clearly formatted return summary. Every sentence adds value with no repetition or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless tool with an output schema, the description is complete enough: it identifies the system scope, enumerates audit sources, and describes returned fields. It could add a note on required privileges or the contrast with list_systemd_timers, but nothing essential is missing for successful invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and the schema coverage is 100%, so the description does not need to clarify parameter meanings. The baseline of 4 applies, and the description appropriately focuses on behavior and output rather than inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool discovers all scheduled cron jobs on the Linux system, with a specific verb and resource. It enumerates the exact audit locations (/etc/crontab, /etc/cron.d/, /etc/cron.*, user crontabs), making its scope concrete and easily distinguishable from systemd-timer-related siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: whenever cron jobs need to be enumerated. It does not explicitly name alternatives or exclusions, but the scope is unambiguous and the 'Discover all scheduled cron jobs' phrasing provides sufficient guidance for a zero-parameter tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_directoryA
Inspect file and directory structures within authorized administrative paths.
Permitted directories: /etc/nginx/, /etc/mysql/, /etc/postgresql/, /etc/docker/, /etc/caddy/, /var/www/
Args: dir_path: Path to the directory to inspect. max_depth: Exploration depth (1 to 3, default: 1).
Returns: JSON string with item list (names, types, sizes, modification dates).
| Name | Required | Description | Default |
|---|---|---|---|
| dir_path | Yes | ||
| max_depth | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It does disclose the permitted directory constraint and the return format, but does not explicitly state side-effect-free behavior or error handling. This partial transparency earns a 4.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured with a clear purpose, a list of permitted directories, parameter explanations, and return format. It is a bit verbose but front-loaded with the main action. No superfluous content; it earns a 4.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even though there is an output schema (not shown), the description provides sufficient detail about the return format (JSON string with item list including names, types, sizes, modification dates). It also outlines the permitted directory scope. Minor omissions like exact error behavior keep it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema lacks parameter descriptions, but the tool description adds meaningful context: it explains dir_path as 'path to the directory to inspect' and max_depth as 'exploration depth (1 to 3, default: 1)'. It also attaches the critical constraint of permitted directories to dir_path. This significantly enriches the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: to inspect file and directory structures within authorized administrative paths. It explicitly lists permitted directories, and the return format (JSON string with item list) confirms it lists directory contents. This is sufficiently specific to distinguish it from sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions 'authorized administrative paths' but does not explicitly state when to use this tool versus alternatives like view_file_content or list_virtual_hosts. An agent may infer usage from the name, but there is no explicit guidance on conditions or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_docker_containersA
List Docker containers with their status, image, port bindings, volumes, and health.
Args: all: Set to True to list all containers (running and stopped), False for running only.
Returns: JSON string with list of containers, port forwards, and mount mappings.
| Name | Required | Description | Default |
|---|---|---|---|
| all | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It clearly signals a read-only 'List' operation and discloses the return shape: a JSON string with containers, port forwards, and mount mappings. It does not mention failure modes or daemon availability, but for a list operation this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured with distinct Purpose, Args, and Returns sections. It leads with the action and resource, contains no filler, and every sentence contributes useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, read-only tool with an output schema, the description is nearly complete: it explains the parameter and the return payload. It only lacks explicit context about when to prefer this over related sibling commands, but this is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only the boolean name and default, with no property description. The description fully compensates by explaining the True and False behaviors: all containers vs running only. This adds meaningful semantic context beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action and resource: 'List Docker containers', and enumerates the contained fields: status, image, port bindings, volumes, and health. This clearly distinguishes it from sibling tools like get_docker_container_logs or get_docker_stats.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The tool is simple enough that usage is implied by its name, and the 'all' parameter is explained. However, there is no explicit guidance about when to choose this tool over siblings such as get_docker_stats or get_docker_container_logs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_systemd_timersA
Audit active and pending systemd timers via 'systemctl list-timers'.
Returns: JSON string with timer unit names, next execution time, countdown, and target services.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and it does well by naming the exact systemctl command and stating the return payload fields. It implies a read-only audit operation, though it never explicitly says it makes no system changes or whether elevated permissions are needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two compact sentences: the first states the action and command, the second enumerates the return values. There is no filler or redundant detail, making it easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-argument read-only auditing tool with an output schema, the description covers the command, scope, and return fields. It could add a note about permissions or how it relates to cron-job auditing, but those are not critical for invoking the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool accepts zero parameters, so there is no parameter semantics to clarify; the baseline of 4 applies. The description's mention of the command and output is sufficient for a no-argument tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'Audit active and pending systemd timers'. It also identifies the underlying command 'systemctl list-timers', making the tool's function unmistakable and naturally distinguishing it from sibling tools like list_cron_jobs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to prefer this tool over alternatives such as list_cron_jobs or get_failed_systemd_units. The description states what the tool does but not when it should be selected or avoided, leaving that decision entirely to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_virtual_hostsA
Inspect active Nginx virtual hosts, listening ports, SSL, and reverse proxy targets.
Returns: JSON string with parsed virtual hosts from /etc/nginx/sites-enabled/ and conf.d/.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the disclosure burden. It reveals the return format (JSON string), the source paths (/etc/nginx/sites-enabled/ and conf.d/), and the informational scope. It doesn't explicitly state that the operation is side-effect-free or describe error/failure behavior, so transparency is adequate but not complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences contain everything meaningful, with the primary action front-loaded and the return/source information compactly appended. There is no filler or redundant wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with an output schema, the description adequately covers the source scope and return type. It could note permission requirements or explicitly confirm read-only behavior, but given the simple inspection nature, the description is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool accepts zero parameters and the schema coverage is 100%, so there are no parameter details for the description to add. Baseline 4 is appropriate because no argument documentation is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Inspect) and resource (active Nginx virtual hosts), plus relevant aspects like listening ports, SSL, and reverse proxy targets. It doesn't explicitly differentiate itself from siblings such as check_ssl_certificates, get_open_ports, or test_nginx_config, so it falls short of a full 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage whenever Nginx virtual host configuration details are needed, and it names the source paths inspected. However, it doesn't state when to prefer this over sibling tools or mention any exclusions, leaving some selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_service_logsA
Safely fetch and optionally filter recent log lines for a service or Docker container.
Args: service_name: Target unit (e.g. 'nginx', 'systemd:cron', 'docker:my_container'). lines_count: Number of recent lines to retrieve (default: 50, maximum: 1000). grep_filter: Optional case-insensitive keyword to filter lines (e.g. 'ERROR', '403', 'denied').
Returns: JSON string containing the extracted log lines and matching statistics.
| Name | Required | Description | Default |
|---|---|---|---|
| grep_filter | No | ||
| lines_count | No | ||
| service_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burder. It conveys read-only intent with 'safely fetch' and reveals the return format as a JSON string with log lines and matching statistics. However, it does not disclose limits, error behavior, or access requirements beyond the line_count cap mentioned in the args.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The structure is clean: a one-sentence overview, an Args list with per-parameter details, and a Returns statement. Every sentence provides useful information and no filler. The key capability is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter read tool, the description is self-contained. It covers all parameters with defaults, examples, and constraints, and explains the output shape. Since an output schema exists, the explicit return description is a bonus. Nothing needed to invoke the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description fully compensates by documenting all three parameters with meaningful detail: service_name examples ('nginx', 'systemd:cron', 'docker:my_container'), lines_count default 50 and max 1000, and grep_filter as optional and case-insensitive. This far exceeds the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb and resource: 'fetch and optionally filter recent log lines for a service or Docker container.' This clearly communicates the tool's function. It does not explicitly differentiate it from the sibling tool get_docker_container_logs, which may overlap in the Docker case.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool (to retrieve recent logs, optionally filtered) but provides no explicit guidance about when to choose this over related siblings like get_docker_container_logs or check_service_status. No exclusion or alternative-routing context is included.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
test_network_connectivityA
Benchmark outbound network connectivity and latency using direct Python sockets.
Measures DNS resolution latency, TCP handshake time, and TLS handshake latency without shell ping.
Args: target_host: Destination hostname or IP address (e.g. 'api.github.com' or '8.8.8.8'). port: Destination port (1-65535, default 443). timeout_seconds: Network socket timeout (0.5 to 30.0 seconds, default 5.0).
Returns: JSON string with stage latency breakdown, resolved IP addresses, and TLS session details.
| Name | Required | Description | Default |
|---|---|---|---|
| port | No | ||
| target_host | Yes | ||
| timeout_seconds | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it delivers substantive behavior details: direct Python sockets, no shell ping, and exactly which latency stages are measured. It also discloses the return shape. It stops short of noting potential side effects like outbound connections to arbitrary hosts, but those are strongly implied by the TCP/TLS handshake wording.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well organized and front-loaded: a one-sentence summary, followed by bullet-style Args and a Returns section. Every sentence contributes information and nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter network diagnostic without annotations, the description covers the core purpose, parameter constraints, and return format. It does not provide explicit sibling guidance or safety notes, but the output schema exists and the description is otherwise sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has no descriptions for any property, so the description fully compensates. Each parameter gets a human-readable explanation with allowed ranges or defaults: target_host with examples, port with 1-65535 and default 443, timeout_seconds with 0.5-30.0 and default 5.0. This is exactly the semantic information an agent needs beyond the raw JSON schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Benchmark outbound network connectivity and latency using direct Python sockets,' naming a specific verb, resource, and method. It further distinguishes itself by listing the three measured stages (DNS, TCP, TLS) and explicitly noting it does not use shell ping, which separates it from diagnostic siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by explaining that it benchmarks network connectivity and latency, which gives an agent a clear sense of when it applies. However, it never explicitly states when to prefer this tool over related siblings such as check_dns_health or check_ssl_certificates, nor does it mention exclusions or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
test_nginx_configA
Test Nginx configuration for syntax errors ('nginx -t') without reloading.
Returns: JSON string indicating syntax validity, exit code, and syntax error messages.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description itself communicates that this tool only tests syntax and does not reload the server, which implies a safe read-only validation operation. It also discloses what the return value contains (validity, exit code, and error messages). This is sufficient behavioral transparency for this tool type.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no filler. The primary purpose is front-loaded and the return-value note is directly relevant to an agent calling the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with an output schema and a return format described inline, the description is complete. An agent knows exactly what action will be taken and what kind of response to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema description coverage is 100%, so there is no parameter meaning for the description to add. The baseline for a zero-parameter tool is 4, and nothing more is needed here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool tests Nginx configuration for syntax errors using `nginx -t` and explicitly says it does not reload. This gives a specific verb, resource, and scope. It is distinctive enough against the sibling tools, none of which appear to focus on nginx configuration..
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the tool's purpose and non-destructive behavior explicit, so an agent can decide to use it when validating Nginx configuration syntax. It does not name alternative tools or explicitly state when not to use it, but no sibling appears to compete directly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
view_file_contentA
Safely read the content of an authorized configuration or web file.
Permitted directories: /etc/nginx/, /etc/mysql/, /etc/postgresql/, /etc/docker/, /etc/caddy/, /var/www/ Strictly protected against path traversal attacks.
Args: file_path: Canonical path or relative path to the configuration file. max_bytes: Maximum bytes to return (default: 50,000, capped at 200,000).
Returns: JSON string with file content, size, and modification timestamp.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | ||
| max_bytes | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully carries the burden of behavioral disclosure. It states the operation is safe, read-only, protected against path traversal, limited to specific directories, capped at 200,000 bytes, and returns content, size, and modification timestamp.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with purpose, permitted directories, safety note, Args, and Returns. It is concise, front-loaded, and every sentence contributes to correct invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, scope, parameter semantics, and return shape, making it sufficient for a simple read tool. The main omission is explicit behavior for unauthorized or missing paths, but this is minor given the clear permitted-directory boundaries and output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain both parameters. It clearly defines file_path as a canonical or relative path and max_bytes with default 50,000 and cap 200,000, adding meaning beyond the raw input schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads the content of authorized configuration or web files and lists permitted directories. This makes it easy to distinguish from siblings like list_directory, read_service_logs, and write_file_content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context by enumerating the permitted directories and the file types the tool targets. It does not explicitly name alternative tools for exclusion, but the scope is specific enough for an agent to know when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
write_file_contentA
Atomically write or update a configuration file within authorized directories.
Creates an automatic timestamped backup (.bak.) before overwriting. Permitted directories: /etc/nginx/, /etc/mysql/, /etc/postgresql/, /etc/docker/, /etc/caddy/, /var/www/
Args: file_path: Path to the target configuration file. content: Text content to write. backup: Create a backup file before writing (default: True).
Returns: JSON string indicating write status and backup location.
| Name | Required | Description | Default |
|---|---|---|---|
| backup | No | ||
| content | Yes | ||
| file_path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and uses it well: it discloses atomic write behavior, creation of a timestamped backup before overwriting, optional backup via the backup parameter, authorized directory restrictions, and the JSON return format. The potentially destructive overwrite is revealed with mitigations rather than hidden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the core behavior, and then formatted into scope, parameters, and return sections. No filler sentences; each line adds needed information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter mutation tool, the description covers what the tool does, where it may act, what side effects occur, what the parameters mean, and what the caller receives. There is no destructive ambiguity or missing critical input needed to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description documents all three parameters with meaningful details: file_path as the target path, content as text to write, and backup with its default behavior. This more than compensates for the barren input schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Atomically write or update a configuration file within authorized directories.' This clearly distinguishes it from read-only siblings like view_file_content and list_directory, and the permitted-directory list narrows its scope further.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the tool is for writing or updating configuration files and explicitly enumerates permitted directories, which gives an agent clear conditions for use and exclusion (anything outside those directories). It does not name alternative tools or explicitly say when-not-to-use, but the scope is concrete enough to route selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
31 tool updates
v0.8.1- First observed
analyze_disk_usage - First observed
audit_ssh_config - First observed
check_dns_health - First observed
check_failed_logins - First observed
check_guardian_updates - First observed
check_kernel_errors - First observed
check_oom_events - First observed
check_service_status - First observed
check_ssl_certificates - First observed
check_system_updates - First observed
create_backup - First observed
execute_recovery - First observed
get_database_health - First observed
get_docker_container_logs - First observed
get_docker_stats - First observed
get_fail2ban_status - First observed
get_failed_systemd_units - First observed
get_open_ports - First observed
get_system_health - First observed
get_top_processes - First observed
get_ufw_status - First observed
list_cron_jobs - First observed
list_directory - First observed
list_docker_containers - First observed
list_systemd_timers - First observed
list_virtual_hosts - First observed
read_service_logs - First observed
test_network_connectivity - First observed
test_nginx_config - First observed
view_file_content - First observed
write_file_content
TDQS
Scored across 31 tools
Tool names and descriptions generally separate resources and actions well, but a few overlapping boundaries remain: read_service_logs can already fetch Docker logs, duplicating get_docker_container_logs, and list_virtual_hosts/get_open_ports both surface listening-port info. Overall an agent can usually pick the right tool with careful reading.
Every tool follows a snake_case verb_noun pattern with semantically appropriate verbs such as list, check, get, audit, test, execute, and create. Even with a large surface, the prefixes map predictably to action types and the resource nouns are clear.
31 tools is a heavy single-server surface and exceeds the 25+ threshold for comfortable agent selection. Several tools could be consolidated, such as Docker logs vs. service logs, or system health vs. separate status/disk tools.
The server has strong coverage for monitoring, auditing, Docker, config files, and whitelisted recovery actions, but there are notable dead ends: backups can be created but not restored, SSH config is audited but cannot be edited, and UFW/SSL have no modification or renewal path. These gaps will force agents to stop or go outside the MCP for common VPS remediation tasks.
Maintenance
Related MCP Connectors
Your org's AI agents, tasks, runs, search, and brain files as MCP tools and resources.
Operate Linux, macOS and Windows from your LLM. Every action runs through an auditable allowlist.
Hosted MCP with 91 agent tools: X, domains, SEO, Maps, Trends, Search, YouTube, TikTok, and more.
Provides capabilities that let LLM agents perform a range of infrastructure management tasks.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables AI assistants to monitor and manage Linux infrastructure including services, logs, processes, disk, memory, ports, cron, nginx, Docker, and system health checks via the Model Context Protocol.10MIT
- AlicenseAqualityCmaintenanceEnables secure, read-only inspection of a VPS over SSH through approved operations such as system health, disk usage, container logs, and service status, without giving the AI unrestricted shell access.81MIT
- AlicenseAqualityCmaintenanceProvides read-only access to host system metrics (CPU, memory, disk), Docker container health/logs, and sandboxed log file analysis via MCP tools, enabling AI agents to monitor enterprise infrastructure safely.3MIT
- FlicenseNot gradedqualityBmaintenanceEnables AI agents to safely explore and diagnose remote servers by providing a read-only sandbox with controlled access to files, logs, Docker, and databases. It exposes MCP tools that allow natural-language investigation and direct command execution without write permissions.3-