nginx-certbot-mcp
This server lets an AI agent manage production web infrastructure (nginx reverse proxies, DNS records, and Let's Encrypt certificates) through scoped, auditable MCP tools, without giving root shell access.
Inspect infrastructure: List nginx sites, view raw configs, check cert expiry, resolve DNS, probe upstream health, get nginx status, tail logs, list archives/backups, and run a full site diagnosis with suggested fixes.
Manage DNS (Route 53): Create/update or delete CNAME records (for pointing domains), and create/delete TXT records (e.g., for ACME DNS-01 challenges).
Manage nginx sites: Create new sites from a template, update an existing site's upstream, delete/site (with archive), restore from archive, rollback config changes, prune old archives, and reload nginx only if config test passes.
Manage certificates: Issue certificates via HTTP-01 or DNS-01 (including wildcards), renew certificates (dry-run by default), revoke certificates, delete certificate files, and benefit from a production rate-limit guard.
Safety controls: Read-only mode, tool allowlist, domain allowlist, per-tool confirmation gates, automatic config backups, audit logging, and staging-by-default for certificate issuance.
Provides tools for reading nginx configuration, managing reverse proxy server blocks, reloading nginx, and issuing/renewing SSL certificates via certbot.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@nginx-certbot-mcpcheck the SSL cert expiry for example.com"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
nginx-certbot-mcp
Let AI agents manage production web infrastructure without giving them root shell access.
nginx-certbot-mcp provisions Nginx reverse proxies, DNS records, and Let's Encrypt certificates through narrowly scoped, auditable MCP tools — never arbitrary shell commands. Privileged actions go through purpose-built wrappers and least-privilege sudo rules (see Why a wrapper script below).
Architecture

Related MCP server: npm-mcp
Tools
All 26 are implemented and exercised against real infrastructure — see Testing.
Tool | Description |
| List configured nginx server blocks with domain, upstream, and SSL status |
| Raw nginx config for one domain |
| List certbot-managed certs and days until expiry |
| Resolve a domain (CNAME, then A/AAAA) against public resolvers |
| TCP probe of an upstream |
| Whether nginx is running, plus its version |
| One-call health report for a domain: config, nginx, DNS, upstream, certificate, recent errors — with suggested next tools |
| Tail access/error logs, capped at 1000 lines |
| List configs archived by |
| List the automatic pre-change backups of site configs |
| Upsert a Route 53 CNAME |
| Delete a Route 53 CNAME — |
| Upsert a Route 53 TXT record, e.g. for ACME DNS-01 |
| Delete a Route 53 TXT record — |
| Create a websocket-capable nginx server block from the default template |
| Rewrite an existing site's |
| Disable, archive, and delete a server block — |
| Re-enable a site from its newest (or a chosen) archive — |
| Undo a config change by restoring the newest (or a chosen) backup — |
| Delete archives older than N days — |
|
|
| Issue via HTTP-01 ( |
| Issue |
|
|
| Revoke with Let's Encrypt, leaving the files in place — |
| Remove a cert's files from certbot's store — |
Setup
npm install
npm run build
npm run setup -- mcpuserRun the server as a dedicated non-root user (e.g. mcpuser) — never as
root. npm run setup -- <user> (scripts/setup.sh) grants that user
exactly the privileges below, nothing more, and is idempotent: safe to
re-run any time, including after changing the username or pulling an
update that adds a new allowed command.
Required permissions
npm run setup -- <user> installs two things:
/usr/local/bin/nginx-mcp-writesite— a narrow wrapper script that only accepts{write|enable|disable|remove|archive|restore|remove-archive|backup|restore-backup} <domain>orlog {access|error} <lines>, and only ever touches paths under/etc/nginx/sites-available/,/etc/nginx/sites-enabled/,/etc/nginx/sites-archived/,/etc/nginx/sites-backups/, and the two fixed nginx log files. It re-validates the domain (and, forrestore/remove-archive/restore-backup, the archive or backup filename) itself, independent of the Node-side validation./etc/sudoers.d/nginx-mcp— grants<user>passwordless sudo on exactlynginx -t,systemctl reload nginx,systemctl is-active --quiet nginx,certbot, and the wrapper above. Nothing broader. It also keepsAWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY/AWS_DEFAULT_REGIONthrough sudo (which strips the environment by default) socertbot --dns-route53can see them forissue_wildcard_cert.
issue_wildcard_cert also needs the certbot-dns-route53 plugin —
npm run install:deps installs it for you (see Testing).
Why a wrapper script instead of sudo on tee/ln/rm
An earlier version granted sudo on generic file tools (tee, ln, rm)
so create_site could write into /etc/nginx/. That works, but it's a
wider trust boundary than the task needs: those commands can touch any
root-owned file on the box, not just nginx configs. If the MCP server
process were ever compromised or triggered unexpectedly, the blast radius
would be the whole filesystem.
The wrapper narrows that to one purpose-built binary that can only act on nginx site configs, one archive at a time, or tail one of two fixed log files — nothing else. The trade-off is one more artifact to deploy and keep in sync with the server, in exchange for sudo that can only ever do what this project needs.
Network requirements for certificate issuance
issue_cert and issue_wildcard_cert prove domain ownership two
different ways, with different requirements on where the box sits on your
network:
issue_cert(HTTP-01) — Let's Encrypt makes an inbound HTTP request to the domain on port 80. If you're behind a home/office router doing NAT, that request lands on whichever one private IP your port-forwarding rule targets. The machine running nginx (and this MCP server) has to be that exact machine — not just any box on your network, and not the Docker sandbox (a different private IP on the Docker bridge network). Calling it from the wrong box fails every time, since the challenge request never arrives.issue_wildcard_cert(DNS-01) — validates via a TXT recordcertbot-dns-route53creates in Route 53. This is outbound-only (the box calls the AWS API; nothing calls back in), so it has no port-forwarding requirement and works identically from any network, including the Docker sandbox.
Either way, DNS still has to point at your public IP (create_domain_record
handles that) — DNS and port-forwarding are two separate requirements, and
issue_cert needs both.
Safety controls
Beyond the per-tool confirm gates and staging-by-default, the server has
operator-level controls that an agent can't turn off from inside a
conversation. Everything is configured through environment variables (see
Environment variables).
Audit log
Every mutating tool call is appended to a JSONL file — timestamp, tool, arguments, MCP client, whether it was a dry run, outcome, and a one-line result — including calls that were refused. Credentials in arguments are redacted and long values truncated.
{"ts":"2026-09-19T10:02:11.402Z","tool":"delete_site","mutating":true,"client":"claude-code/2.1.0","args":{"domain":"old.example.com","confirm":true},"dry_run":false,"outcome":"ok","message":"Removed \"old.example.com\" (disabled, archived ...","duration_ms":212}outcome is ok, failed (the tool ran and reported success:false,
including an unconfirmed dry run), error (it threw), or denied.
Default path is
~/.nginx-certbot-mcp/audit.jsonl(mode0600); override withAUDIT_LOG_PATH, or setAUDIT_LOG_PATH=offto disable.The server refuses to start if the log isn't writable — you find out at startup, not after the first change.
Read-only calls are skipped by default; set
AUDIT_LOG_READS=trueto include them.The file grows forever; point
logrotateat it if that matters.
Read-only mode and tool allowlist
Hand an agent visibility without write access, or expose only the tools a workflow needs. Tools that are switched off are never registered, so the agent can't see or call them.
MCP_MODE=readonly # only tools annotated read-only
MCP_ENABLED_TOOLS="check_*,list_sites" # only these (* is a wildcard)MCP_MODEisreadwrite(default) orreadonly. Read-only mode drops every tool that changes state — DNS, nginx config, certificates, reloads.MCP_ENABLED_TOOLSis a comma-separated list of tool names or*patterns. When both are set, a tool must pass both.An invalid
MCP_MODE, or an allowlist entry that matches no tool, is reported on stderr at startup (the former is fatal) rather than silently exposing the wrong set.
Domain allowlist
Confine the agent to the domains it's meant to manage, so a confused or manipulated agent can't edit, delete, or issue certificates for anything else on the box.
ALLOWED_DOMAINS="example.com,*.example.com"example.commatches exactly that name;*.example.commatches any subdomain at any depth, not the apex — list both if you want both. A bare TLD (*.com) is rejected at startup.Applies to every tool's
domainargument, read-only tools included, and to Route 53 record names (_acme-challenge.a.example.commatches*.example.com).issue_wildcard_certforexample.comalso covers*.example.com, so both must be allowed.Listings (
list_sites,check_cert_expiry,list_archived_sites) only show in-scope domains, andprune_archivesonly touches in-scope archives.renew_certandtail_site_logsnormally act on everything whendomainis omitted; with an allowlist they require one.Refused calls return a
Denied by policyerror and are written to the audit log with outcomedenied.
This limits which domains the tools act on; it doesn't change what the
sudo rules permit the server user to do — see
Required permissions.
Config backups and rollback
create_site, update_site, restore_site and rollback_site snapshot
a site's existing config before changing it, and each one restores that
snapshot itself if the new config fails nginx -t — so a broken change
never stays on disk. Backups also cover the case nginx -t can't catch: a
config that is valid but wrong (the wrong upstream, say).
list_site_backupsshows the snapshots, newest first;rollback_siterestores the newest one by default (the config as it was before the last change) or a chosenbackup_filename. It needsconfirm:true.Rolling back snapshots the current config first, so calling it again flips back — a rollback is never a one-way door.
The newest 10 backups per domain are kept in
/etc/nginx/sites-backups/; older ones are deleted automatically.If a snapshot can't be taken, the change is refused rather than made without a safety net. After upgrading, re-run
npm run setup -- <user>so the installed helper knows thebackupaction.Backups are separate from
delete_site's archives: those are "this site was deleted", backups are "what it looked like before the last change".reload_nginxstill only reloads a config that passesnginx -t; when it refuses, itshintpoints atrollback_site.
Production issuance guard
Let's Encrypt's production rate limits punish retry loops, and a lockout can
last a week. issue_cert and issue_wildcard_cert keep a local history of
production (staging:false) attempts and refuse a request that would
exceed a limit — before spending it:
Limit | Guard refuses when |
Duplicate certificates | 5 certs for the exact same set of names in 7 days |
Failed validations | 5 failed validations for a name in 1 hour |
Certificates per registered domain | 50 certs for one registered domain in 7 days |
A refusal says which limit was hit and when to retry; a Heads-up note in
rate_limit_note appears as a limit gets close. Staging requests are never
counted or refused.
It counts only what was issued through this server, so it's a guard rail, not a substitute for Let's Encrypt's own limits. "Registered domain" is approximated from the last two labels (three under
co.uk-style suffixes).Only failures that reached validation count towards the failed-validation limit; a missing sudo rule or a missing nginx block doesn't.
History lives in
issuance.jsonunderMCP_STATE_DIR(default~/.nginx-certbot-mcp/, mode0600) and is pruned after 7 days. SetRATE_LIMIT_GUARD=offto disable the guard.
Environment variables
Variable | Used by | Notes |
|
| Credentials for a Route-53-scoped IAM user — no other AWS permissions needed |
|
| Find with |
| Same as above, plus | Optional — Route 53 is global, but the AWS SDK/boto3 still need a signing region; defaults to |
| All mutating tools | Optional — audit log file; default |
| All tools | Optional — |
| All tools | Optional — comma-separated tool names / |
| All tools taking a | Optional — comma-separated |
|
| Optional — |
|
| Optional — where the guard keeps its issuance history; default |
| Read-only tools | Optional — |
Testing
Four layers, from "needs almost nothing" to "exercises everything":
1. Dependencies — npm run install:deps
Debian/Ubuntu only, idempotent. Installs nginx, certbot,
python3-certbot-nginx, and python3-certbot-dns-route53 via apt if
missing; for anything already installed, it only reports whether the
version is current, since silently upgrading a package that might be
serving traffic isn't this script's call to make. If certbot looks like
a snap install (common — certbot's own docs recommend it over apt's
often-outdated package), it also tries installing certbot-dns-route53 as
a snap plugin, since an apt-installed plugin can be invisible to snap
certbot's isolated Python environment. That's best-effort and additive,
never a replacement for the apt package, so issue_wildcard_cert has a
working path either way. Also reports your Node version against the
>=20 that @aws-sdk/client-route-53 will eventually require.
2. Route 53 round trip — npm run test:dns
The only requirement is a working ROUTE53_HOSTED_ZONE_ID (+ AWS
credentials) in .env. It discovers your zone's own domain from the
hosted zone, creates a disposable CNAME under a random subdomain, verifies
it (directly against Route 53, and best-effort via public DNS), then
deletes it — cleanup runs even if a check in between fails, so a bad run
can't leave an orphaned record.
3. Docker sandbox — real nginx + certbot, disposable
cp .env.example .env # fill in AWS credentials + hosted zone ID
docker compose up -d --build
docker compose exec sandbox npm run inspectThe container runs systemd as PID 1, so sudo systemctl reload nginx and
friends work exactly as they do in production (needs --privileged and a
cgroup mount, which docker-compose.yml already sets up). nginx/
certbot/the plugins are installed via install-deps.sh, and the
wrapper/sudoers via setup.sh, both at image build time. Tear down with
docker compose down — nothing persists; every rebuild is a fresh install.
4. Automated tool-by-tool suite
cp .env.test.example .env.test # AWS credentials + a domain you control
npm run test:tools # against the Docker sandbox
npm run test:tools:host # against this machine directlyDrives every tool over the real stdio JSON-RPC protocol and prints
✓/✗/– per tool. .env.test is separate from .env — read on the host and
injected directly into each MCP server process the runner spawns, so the
two files never need to match. Before touching anything it verifies your
AWS credentials work and that TEST_DOMAIN is the zone's apex or a
subdomain of it. Everything then runs under a random
mcp-test-<random>.<TEST_DOMAIN> subdomain, self-cleans after each phase,
and does a final best-effort cleanup regardless of pass/fail. Certificate
issuance is opt-in — asked interactively, or pass --certs for a
non-interactive run — since it hits real Let's Encrypt staging and adds a
minute or two.
The two targets differ in exactly one way, issue_cert:
npm run test:tools(default) runs in the Docker sandbox, which isn't reachable from the internet, soissue_cert(HTTP-01) is always skipped — the cert scenario only exercisesissue_wildcard_cert(DNS-01) and the renew/revoke/delete chain built on it.npm run test:tools:hostrunsnode dist/index.jsdirectly on this machine. If this is the box your router actually forwards 80/443 to, the cert scenario testsissue_certtoo (waiting up to 300s for the disposable CNAME to propagate first), and the renew/revoke/delete chain runs against that cert instead. It refuses to start unless passwordless sudo already works for the current user — i.e.npm run setup -- <user>was run for the user actually invoking it, not some other account. Everything this touches is real production state, not a sandbox.
Connecting a client
MCP Inspector — the fastest feedback loop for poking at a tool directly:
npm run inspectPrints a URL with a session token. Run it against a real box, or the
Docker sandbox (docker compose exec sandbox npm run inspect).
Claude Desktop / claude.ai — add to your MCP client config:
{
"mcpServers": {
"nginx-certbot": {
"command": "node",
"args": ["/absolute/path/to/nginx-certbot-mcp/dist/index.js"]
}
}
}Or against the running Docker sandbox:
{
"mcpServers": {
"nginx-certbot-sandbox": {
"command": "docker",
"args": ["exec", "-i", "nginx-certbot-mcp-sandbox", "node", "dist/index.js"]
}
}
}Diagnosing a site
When a site misbehaves, start with diagnose_site instead of calling the
individual checks one by one. It runs, in parallel, and reports each as
ok / warn / fail / skipped:
Check | Looks at |
| config exists, is enabled, |
| service running, and |
| the domain resolves (CNAME, then A/AAAA) |
| the |
| a certbot cert covers the domain (wildcards included), days left, and that the config actually uses it |
| the latest nginx error-log lines mentioning the domain or its upstream |
The result also has an overall healthy flag (no check failed — warnings
and skipped checks don't count) and next_steps: the specific tools that
would fix each problem, e.g. create_domain_record for a domain that
doesn't resolve or rollback_site after a config that no longer passes
nginx -t. A check that can't run (say, certbot isn't installed) is
reported as skipped with the reason; it doesn't sink the rest.
It is read-only, so it's available in MCP_MODE=readonly, and it respects
ALLOWED_DOMAINS.
Typical "add a new site" flow
create_domain_record— pointmysite.julcap.netatwww.julcap.net(wait for DNS propagation)
create_site— nginx serves the domain on port 80, reverse-proxied to the local service IP:portreload_nginxissue_cert— certbot validates via HTTP-01, updates nginx to redirect to 443
Contributing
Contributions are welcome. See CONTRIBUTING.md for guidelines.
License
nginx-certbot-mcp is source-available under the Elastic License 2.0. You may use, modify, and redistribute the software. However, you may not provide a substantial portion of its functionality to third parties as a hosted or managed service.
For commercial licensing or partnership enquiries, contact the maintainer.
TODO
Planned, not yet implemented:
provision_siteworkflow tool — DNS record, nginx site, reload and certificate in one call, rolling back the steps already taken if a later one fails. Today the agent has to sequence the add-a-site flow itself.Certificate expiry alerts — a
warn_daysthreshold oncheck_cert_expiryand an optional webhook (Slack / Discord) for certificates that are close to expiring.Per-site options in
create_site— custom headers, client body size, rate limiting, basic auth, IP allowlist, HTTP→HTTPS redirect, HSTS, and a choice of named templates instead of the single default one.More DNS providers — Cloudflare first, alongside Route 53 (certbot already has DNS plugins for it), so the DNS and DNS-01 tools aren't tied to AWS.
Available Tools
26 toolscheck_cert_expiryARead-only
List every certbot-managed certificate on the box with its expiry date and days remaining, via certbot certificates. Read-only. Covers all certs certbot knows about, not just domains with an active nginx site.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| certificates | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and the description redundantly states 'Read-only.' However, the description adds the behavioral detail that it executes 'certbot certificates' and covers all certs, which goes beyond the annotation. It does not describe potential edge cases (e.g., no certs) but the output schema likely covers the return shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no redundancy except the read-only note (which is harmless). The first sentence is action-oriented and front-loaded; the second adds necessary scope clarification. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only check tool, the description is complete: it states what it lists, how it does it, and its scope. The presence of an output schema means return value details are covered elsewhere. No additional information is needed for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description adds no parameter specifics—correctly so. With schema coverage at 100% (empty properties) and baseline for 0 params set at 4, the description needs no extra parameter info and appropriately remains silent on params.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'List' with a specific resource (certbot-managed certificates on the box) and the output (expiry date and days remaining). It also distinguishes its scope from similar tools by clarifying it covers all certs, not just those with an active nginx site, which differentiates it from sibling tools like list_sites or renew_cert.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context that this tool is for checking certificate expiry and explicitly notes it covers all certs, not only those with active nginx sites. While it doesn't name alternative tools, the scope clarification helps an agent decide when to use it over other cert-related tools. A higher score would require explicit when/not-when guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_dnsARead-only
Resolve a domain (CNAME first, then A/AAAA) against public resolvers (1.1.1.1, 8.8.8.8) rather than this box's own DNS, so the result matches what Let's Encrypt and the public internet see. Use before issue_cert / create_site to confirm a domain actually points where you expect. For confirming something is listening behind nginx, use check_upstream_health instead.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | Domain to resolve, e.g. mysite.julcap.net |
Output Schema
| Name | Required | Description |
|---|---|---|
| values | No | Resolved values (IPs, or the CNAME target); only present when resolves is true |
| resolves | Yes | |
| record_type | No | Only present when resolves is true |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint and openWorldHint, but the description adds behavioral nuance beyond that: it explains the resolution is against public resolvers to match Let's Encrypt's view, and mentions the query order. This is valuable context that isn't conveyed by annotations alone. No contradictions found.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise (two sentences) and front-loaded with the core function and rationale. It avoids redundancy and each clause adds value—explaining the method, the use case, and the sibling alternative without fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with one parameter, an output schema (which we don't see but is indicated), and annotations covering safety, the description is complete. It covers why the public resolver is used, when to call it, and how it differs from a nearby sibling, leaving no gap an agent would need filled.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is only one parameter, 'domain', and the input schema already provides a clear description with an example. The tool description doesn't add extra parameter-level details beyond what the schema covers, so it earns the baseline score for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: resolving a domain against public resolvers (1.1.1.1, 8.8.8.8) with a specific query order (CNAME first, then A/AAAA). It explicitly names the sibling tool check_upstream_health as the alternative for a different task, making its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use the tool ('Use before issue_cert / create_site to confirm a domain actually points where you expect') and when not to ('For confirming something is listening behind nginx, use check_upstream_health instead'). This leaves no ambiguity about the tool's intended context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_upstream_healthARead-only
TCP-probe an upstream host:port to confirm something is actually listening there, independent of nginx or DNS. Use to sanity-check an upstream before create_site, or to debug a 502 afterward. For confirming a public domain resolves, use check_dns instead.
| Name | Required | Description | Default |
|---|---|---|---|
| timeout_ms | No | Milliseconds to wait before reporting unreachable | |
| upstream_host | Yes | Hostname or IP that nginx would proxy_pass to | |
| upstream_port | Yes | TCP port to probe |
Output Schema
| Name | Required | Description |
|---|---|---|
| message | Yes | Human-readable outcome, e.g. success, timeout, or connection error detail |
| reachable | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds useful behavioral context beyond that: it is a TCP-level probe, independent of nginx or DNS, confirming that something is actually listening rather than relying on resolution.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler: the first states what the tool does, the second gives concrete usage contexts, and the third routes to the relevant alternative. Every sentence earns its place and the key information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a fully documented input schema, an output schema, and annotations covering read-only and open-world behavior, the description adds exactly what is missing: purpose, usage timing, and differentiation from check_dns. Nothing an agent needs to invoke it correctly is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameters are already well-documented in the schema. The description does not need to elaborate on individual parameters; it correctly focuses on the tool's purpose and usage context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'TCP-probe an upstream host:port to confirm something is actually listening there.' It also explicitly distinguishes itself from check_dns, making the tool's scope unmistakable even among many siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete use cases: sanity-check an upstream before create_site or debug a 502 afterward. It also names the alternative tool (check_dns) and the condition that selects it, leaving no ambiguity about when to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_domain_recordAIdempotent
Upsert a Route 53 CNAME record pointing domain at target. Safe to call repeatedly - it's an upsert, not create-only. Run this before create_site / issue_cert and allow a few minutes for DNS propagation; issue_cert re-checks resolution itself, so it's safe to retry issue_cert if it reports the domain isn't resolving yet. For an ACME DNS-01 TXT challenge record, use create_txt_record instead.
| Name | Required | Description | Default |
|---|---|---|---|
| ttl | No | DNS TTL in seconds | |
| domain | Yes | The domain to create/update, e.g. mysite.julcap.net | |
| target | Yes | CNAME target the domain should point to, e.g. www.julcap.net |
Output Schema
| Name | Required | Description |
|---|---|---|
| message | Yes | |
| success | Yes | |
| change_id | No | Route 53 change ID, useful for polling propagation status |
| change_status | No | Route 53 change status, e.g. PENDING or INSYNC |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Although idempotentHint already signals repeatability, the description adds valuable behavioral context: DNS propagation may take minutes, issue_cert re-checks resolution itself, and retrying issue_cert is safe if the domain is not yet resolving. This goes beyond the annotations and helps the agent anticipate timing and failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tight and front-loads the core purpose before providing workflow guidance. The sentence about repeatability overlaps with the idempotentHint annotation, which prevents a perfect score, but the remaining sentences all earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the rich schema, existing annotations, and presence of an output schema, the description fills the remaining practical gaps: where this tool fits in the provisioning flow, DNS propagation delays, retry safety, and which sibling tool to use instead. Nothing necessary for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema documents all three parameters with descriptions and examples, so parameter semantics are fully covered by structured data. The description restates domain→target but adds no new constraints, syntax, or edge-case guidance beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly identifies the operation as an upsert of a Route 53 CNAME record pointing `domain` at `target`, which is specific and unambiguous. It also distinguishes itself from create_txt_record and from create-only semantics, so an agent can immediately tell what this tool does and does not do.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit sequencing ('Run this before create_site / issue_cert'), timing guidance ('allow a few minutes for DNS propagation'), and a clear alternative for TXT records ('use create_txt_record instead'). This gives the agent concrete conditions for when to invoke this tool versus a sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_siteA
Create a new nginx server block from the websocket-capable default template. Validates and test-renders (nginx -t) before touching live config, and rolls back automatically if the test fails. Does NOT reload nginx or request a certificate - follow with reload_nginx to go live, then issue_cert to get SSL. To point an existing site at a different upstream later, use update_site instead of recreating it. If a config for the domain already exists it is replaced - after being backed up (see rollback_site).
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | Domain for the new server block, e.g. mysite.julcap.net | |
| upstream_host | Yes | Hostname or IP nginx should proxy_pass to | |
| upstream_port | Yes | TCP port on the upstream host |
Output Schema
| Name | Required | Description |
|---|---|---|
| success | Yes | |
| config_path | No | Present on success: absolute path of the written config |
| test_output | Yes | Output of `nginx -t` against the rendered config |
| backup_created | No | True if an existing config was replaced; it was backed up first (see rollback_site) |
| reload_required | Yes | True on success - nginx has not actually been reloaded yet |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses safety behaviors beyond annotations: runs `nginx -t` before touching live config, auto-rolls back on failure, and replaces existing configs only after backing them up via rollback_site. The annotations do not cover these details, so this is valuable added context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Each sentence earns its place: purpose, safety validation, follow-on steps, alternative tool, and replace/backup behavior. The most important action is front-loaded and the wording is dense without being wordy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema and detailed annotations, the description covers what the tool does, what it does not do, what happens on failure, how to handle existing configs, and the required follow-up steps. It gives an agent everything needed to invoke it correctly and sequence it with siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema covers all three parameters with full descriptions at 100% coverage, so the description does not need to repeat them. The description adds no extra parameter-level meaning, but the schema already carries that burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Create a new nginx server block' from a particular template. It also distinguishes itself from sibling tools like update_site and the follow-on reload_nginx/issue_cert, making its role unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit sequencing: do NOT reload or get a certificate; use reload_nginx then issue_cert. Also names update_site as the alternative for re-pointing an existing site, giving clear when-to-use vs when-not-to guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_txt_recordAIdempotent
Upsert a Route 53 TXT record - e.g. for an ACME DNS-01 challenge (_acme-challenge., as used by issue_wildcard_cert) or domain verification. Quotes the value automatically if the caller didn't. Clean up afterward with delete_txt_record. For a CNAME pointing a domain at an upstream, use create_domain_record instead.
| Name | Required | Description | Default |
|---|---|---|---|
| ttl | No | DNS TTL in seconds | |
| value | Yes | TXT record value; wrapped in double quotes automatically if not already | |
| domain | Yes | Record name, e.g. _acme-challenge.mysite.julcap.net |
Output Schema
| Name | Required | Description |
|---|---|---|
| message | Yes | |
| success | Yes | |
| change_id | No | Route 53 change ID, useful for polling propagation status |
| change_status | No | Route 53 change status, e.g. PENDING or INSYNC |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations (openWorldHint, idempotentHint, destructiveHint:false) already declare the safety profile; 'Upsert' is consistent with idempotentHint, so no contradiction. The description adds genuine behavioral value beyond annotations: automatic value quoting, the cleanup expectation, and the 'upsert' semantics. It doesn't mention auth needs or conflict behavior, but with annotations carrying the safety profile a 4 is fair.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences with zero filler. The core purpose is front-loaded, followed by use case, cleanup instruction, and sibling differentiation. Every sentence earns its place and the most actionable constraint (auto-quoting) is embedded early.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no elaboration. Given the moderate complexity (3 params, 2 required), the description covers purpose, use cases, cleanup, and alternative routing. In a rich sibling context of 23 tools, the description fully disambiguates create_txt_record from create_domain_record and delete_txt_record.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three params are already documented, setting the baseline at 3. The description adds meaningful value beyond the schema by explicitly stating the value auto-quoting behavior ('Quotes the value automatically if the caller didn't'), which maps directly to the value param, and reinforces the domain format with a concrete example.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Upsert a Route 53 TXT record') plus concrete use cases (ACME DNS-01 challenge, domain verification) with a real example name (_acme-challenge.<domain>). It also names the specific sibling it relates to (issue_wildcard_cert) and the sibling it is not (create_domain_record), so an agent can differentiate without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use context (DNS-01 challenge, domain verification), tells the agent to clean up afterward with delete_txt_record, and explicitly routes the CNAME case to create_domain_record instead. Alternatives and exclusions are both stated directly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_certADestructiveIdempotent
Delete a certificate's files from certbot's local store. Destructive - requires confirm:true. Does not revoke the certificate first - if it may be compromised, call revoke_cert before this. Any nginx config still referencing the deleted files will fail to reload afterward.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | certbot cert name | |
| confirm | No | Must be true to actually delete; false (default) is a dry run |
Output Schema
| Name | Required | Description |
|---|---|---|
| message | Yes | |
| success | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true, so the safety baseline is covered. The description adds genuinely valuable context beyond that: the confirm:true gate, the critical ordering constraint (revoke must happen before delete), and the concrete consequence that nginx config referencing deleted files will fail to reload. This transforms a bare destructive flag into an actionable risk profile. No contradiction with annotations - destructiveHint matches 'Destructive', idempotentHint is consistent with deleting already-absent files being a harmless repeat, and openWorldHint=false matches a purely local operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, zero waste. The core purpose is front-loaded first, safety warnings follow in order of criticality: confirmation gate, revocation ordering, then downstream impact. Even the slight redundancy of 'Destructive' against destructiveHint is justified for a tool where a misfire is harmful. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool, this is complete. The description covers purpose, confirmation requirement, the revoke-before-delete decision, and post-delete consequences. The output schema exists so return values need no description, and the strong annotations carry idempotency and destructiveness. An agent has everything needed to decide whether and how to call this tool safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both domain ('certbot cert name') and confirm ('Must be true to actually delete; false (default) is a dry run'). The description restates the confirm requirement but adds no new parameter-level meaning beyond the schema. Baseline 3 is correct when the schema carries the full parameter documentation burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Delete a certificate's files from certbot's local store.' It clearly differentiates from the sibling revoke_cert by explicitly stating deletion is not revocation, which is exactly the ambiguity this tool family needs resolved. An agent can distinguish delete_cert from revoke_cert without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit routing guidance: 'if it may be compromised, call revoke_cert before this' names the alternative and the condition that selects it. It also warns about the downstream failure mode (nginx reload failures) which implicitly tells the agent to confirm no config references the cert before deleting. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_domain_recordADestructiveIdempotent
Delete the Route 53 CNAME record for a domain. Destructive - requires confirm:true to actually act; without it, returns what would happen and changes nothing. Looks up the exact existing record first rather than guessing its TTL/value.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | The domain whose CNAME record should be deleted | |
| confirm | No | Must be true to actually delete; false (default) is a dry run |
Output Schema
| Name | Required | Description |
|---|---|---|
| message | Yes | |
| success | Yes | |
| change_id | No | Route 53 change ID, useful for polling propagation status |
| change_status | No | Route 53 change status, e.g. PENDING or INSYNC |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Explicitly discloses the destructive nature and the confirm:true safety gate, including the dry-run behavior. Adds the lookup-first behavior, which goes beyond the annotations and helps predict side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the main action, and each sentence adds a distinct piece of information without waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter destructive tool, the description plus annotations and output schema cover safety, dry-run, and lookup behavior. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description reinforces the confirm parameter's role and explains why no TTL/value parameters exist by noting the tool looks up the existing record. Schema coverage is 100%, so the baseline is 3, but this added context earns a slight upgrade.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: delete the Route 53 CNAME record for a domain. Differentiates from sibling tools like delete_txt_record and delete_cert by specifying the record type and service.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly implies when to use it: to remove a CNAME record from Route 53. Does not explicitly name alternatives or exclusions, but the domain/CNAME scope provides sufficient context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_siteADestructiveIdempotent
Disable, archive, and delete the nginx server block for a domain. Destructive - requires confirm:true to actually act; without it, returns what would happen. Does not touch any certbot certificate for the domain. Does NOT reload nginx - call reload_nginx afterward. The archived copy can be brought back with restore_site.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | Domain whose server block should be removed | |
| confirm | No | Must be true to actually act; false (default) is a dry run |
Output Schema
| Name | Required | Description |
|---|---|---|
| message | Yes | |
| success | Yes | |
| reload_required | Yes | True on success - nginx has not actually been reloaded yet |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true and idempotentHint=true, but the description goes well beyond them by disclosing the dry-run mode, the fact that certificates are left alone, that nginx is not reloaded, and that an archived copy can be restored. These are critical behavioral traits not inferable from the annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each carrying distinct information: what happens, the confirmation guard, the certificate non-effect, and the nginx reload/restore caveats. It is slightly dense but well-organized and front-loaded with the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with an output schema, full parameter schema coverage, and rich annotations, the description covers all necessary operational context: how to actually trigger deletion, what is not affected, what to do after, and how to undo. No important gap remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with both parameters fully described in the JSON schema itself. The description reinforces the confirm parameter's dry-run semantics but adds no new meaning beyond what the schema already documents, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('delete'), a specific resource ('nginx server block for a domain'), and ties in the related actions 'disable' and 'archive'. It clearly differentiates this from sibling tools like delete_cert and revoke_cert by scoping to the nginx server block only.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the confirm:true requirement and the dry-run behavior without it, explains that certbot certificates are untouched, says nginx is NOT reloaded and directs the agent to call reload_nginx afterward, and mentions restore_site as the reversal path. This fully routes the agent to the correct workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_txt_recordADestructiveIdempotent
Delete the Route 53 TXT record for a domain - e.g. to clean up an ACME DNS-01 challenge record left behind by create_txt_record or issue_wildcard_cert. Destructive - requires confirm:true to actually act; without it, returns what would happen and changes nothing. Looks up the exact existing record first rather than guessing its TTL/value. For a CNAME record, use delete_domain_record instead.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | Record name whose TXT record should be deleted, e.g. _acme-challenge.mysite.julcap.net | |
| confirm | No | Must be true to actually delete; false (default) is a dry run |
Output Schema
| Name | Required | Description |
|---|---|---|
| message | Yes | |
| success | Yes | |
| change_id | No | Route 53 change ID, useful for polling propagation status |
| change_status | No | Route 53 change status, e.g. PENDING or INSYNC |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Although destructiveHint is already true, the description adds meaningful detail: it requires confirm:true to act, otherwise it is a dry run that changes nothing. It also discloses that the tool looks up the exact existing record first rather than guessing TTL/value, which is important behavioral context beyond the annotations. No contradiction with annotations is present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four sentences with each one serving a distinct purpose: purpose/example, destructive behavior/confirm requirement, lookup behavior, and sibling alternative. It is front-loaded with the most important information and contains no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity, the description covers purpose, usage triggers, destructive behavior, confirmation requirement, and alternative tool routing. The output schema exists, so return value details are not needed in the description. An agent has enough information to decide when to use this tool and how to invoke it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both domain and confirm. The description reinforces the confirm parameter's dry-run behavior and adds the 'looks up the exact existing record' detail, but this is more about tool behavior than new parameter-level meaning. It meets the baseline but does not add substantial semantic value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Delete the Route 53 TXT record for a domain.' It also gives a concrete use case (cleaning up an ACME DNS-01 challenge record) and explicitly differentiates itself from delete_domain_record for CNAME records. An agent can immediately understand what this tool does and how it differs from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool: to clean up TXT records left by create_txt_record or issue_wildcard_cert. It also provides an explicit alternative: 'For a CNAME record, use delete_domain_record instead.' This gives clear context and exclusion criteria, leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
diagnose_siteARead-only
One-call health report for a domain: whether its nginx config exists and is enabled, whether nginx is running and passes nginx -t, whether the domain resolves, whether the upstream accepts TCP connections, whether a certificate covers it (and how long is left), and recent nginx error-log lines mentioning the site or its upstream. Read-only. Returns per-check status (ok/warn/fail/skipped), an overall healthy flag, and next_steps naming the tools that would fix each problem. Start here when a site is misbehaving instead of calling the individual check_* tools one by one.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | Domain to diagnose, e.g. mysite.julcap.net |
Output Schema
| Name | Required | Description |
|---|---|---|
| checks | Yes | |
| domain | Yes | |
| healthy | Yes | True when no check failed; warnings and skipped checks don't count |
| summary | Yes | |
| next_steps | Yes | Suggested tools/actions for each warning or failure; empty when all is well |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, and the description reinforces this with 'Read-only.' Beyond that, it discloses meaningful behavior: it runs nginx -t, checks certificate expiry, and returns per-check statuses plus next_steps. This adds color to what the tool actually does and what the agent should expect, going beyond the structured annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense, listing many checks in a single run-on sentence, but every clause carries information. It front-loads the core purpose ('One-call health report') and ends with clear usage guidance. Slight restructuring into bullets or shorter sentences would improve readability, but the content is efficient and earned.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a single fully documented parameter, an output schema, and sibling tools that this description explicitly routes around, nothing essential is missing. The description covers what is checked, what is returned, the read-only nature, and when to start with this tool. An agent has enough context to select and invoke it correctly without further research.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 100% coverage for the single parameter, including an example value in the description ('mysite.julcap.net'). The tool description adds no new parameter-level detail beyond confirming that the domain is the subject of the diagnostics. Since the schema fully documents the parameter, a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'One-call health report for a domain.' It enumerates the exact checks performed (nginx config, nginx -t, DNS, upstream TCP, certificate, error logs), which clearly distinguishes it from a generic or ambiguous tool. It also explicitly contrasts itself with the individual check_* siblings, making the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit selection guidance: 'Start here when a site is misbehaving instead of calling the individual check_* tools one by one.' This tells the agent when to use this tool and names the alternative approach it replaces. It also implies that the tool is the entry point for site health triage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_nginx_statusARead-only
Report whether the nginx service is active (via systemctl) and its version string. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| running | Yes | |
| version | Yes | `nginx -v` output, or an explanatory message if it couldn't be determined |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The read-only nature is disclosed both in the description and the readOnlyHint annotation, with no contradiction. The description adds value beyond the annotation by specifying the systemctl mechanism and the version-string output, which are not present in the annotation. No hidden side effects are omitted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no wasted words. The core behavior is front-loaded, and the 'Read-only' note reinforces safety without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter, read-only status tool with a defined output schema and annotations, the description is sufficient. It states the check performed, the mechanism, and the extra version output. There is no missing context that an agent would need to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema description coverage is 100% (empty schema). The description does not need to explain parameter meaning because there are none, so the zero-parameter baseline of 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('report'), a specific resource (nginx service), and the exact output (active status and version string). It clearly distinguishes itself from mutation siblings like reload_nginx by framing this as a status query.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when you need to know whether nginx is active or its version, but it does not explicitly mention alternatives or situations where this tool should not be used. No exclusion criteria or sibling routing is provided, so the guidance is only implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_site_configARead-only
Get the raw, unparsed nginx config file for one domain from sites-available. Throws if no config exists for that domain - call list_sites first if you're not sure it exists.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | Domain as it appears in sites-available, e.g. mysite.julcap.net |
Output Schema
| Name | Required | Description |
|---|---|---|
| domain | Yes | |
| raw_config | Yes | Full contents of the nginx config file, verbatim |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the read-only nature is covered. The description adds valuable behavior beyond annotations: it discloses the throw condition when the config is missing, and clarifies the output format as 'raw, unparsed' (implying no processing or validation). This goes beyond the annotations and provides actionable context for the agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with zero waste. The core action and resource are front-loaded in the first sentence, and the usage guidance (throw condition and alternative) appears in the second. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with a single parameter, and the description covers the error condition and provides a fallback path. An output schema exists (though not shown), and annotations handle the read-only safety. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides a clear description for the single 'domain' parameter ('Domain as it appears in sites-available, e.g. mysite.julcap.net') with 100% coverage. The tool description adds only the phrase 'for one domain,' which adds no new meaning beyond the schema. Baseline of 3 is appropriate since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Get the raw, unparsed nginx config file for one domain from sites-available.' It clearly distinguishes this from sibling tools like list_sites (which lists domains) and check_cert_expiry (which checks certificates). The scope is explicit — one domain, from sites-available — leaving no ambiguity about the tool's function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells the agent when to use it and when not to: 'Throws if no config exists for that domain - call list_sites first if you're not sure it exists.' This names the alternative (list_sites) and the condition that selects it, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
issue_certA
Request a certificate via certbot --nginx (HTTP-01 validation). Requires an nginx server block for domain to already exist (create_site) - certbot's nginx plugin edits that existing sites-available config in place, adding an SSL server block and an HTTP->HTTPS redirect; it does not create a new site from scratch, and it reloads nginx itself on success (no separate reload_nginx call needed). Pre-checks that the domain resolves and fails fast with guidance if not, avoiding a wasted attempt against Let's Encrypt's rate limits. Defaults to Let's Encrypt staging, which issues browser-untrusted certs but is exempt from rate limits - pass staging:false only when you're ready for a real, publicly CT-logged certificate: production Let's Encrypt enforces real per-domain issuance rate limits (a handful of certs per week), and a mis-issued cert isn't silently undone - call revoke_cert if you need to invalidate one. A local guard also refuses production requests that would exceed Let's Encrypt's duplicate-certificate, failed-validation or per-domain limits, reporting when to retry. For a *.domain wildcard, use issue_wildcard_cert instead - HTTP-01 can't validate wildcards.
| Name | Required | Description | Default |
|---|---|---|---|
| No | Contact email registered with the Let's Encrypt account, used for renewal-failure and expiry notices. Omitted registers with --register-unsafely-without-email, so Let's Encrypt cannot warn you if a future automated renewal fails. | ||
| domain | Yes | Domain to request a certificate for. Must already resolve (see check_dns) and already have an nginx server block from create_site - certbot edits that existing config rather than creating one. | |
| staging | No | True (default) uses Let's Encrypt's staging CA - browser-untrusted certs, but exempt from production rate limits; use for testing the flow. False requests a real, browser-trusted cert and counts against production rate limits. |
Output Schema
| Name | Required | Description |
|---|---|---|
| success | Yes | |
| dns_check | No | Present only when the DNS pre-check failed, before certbot was even invoked |
| certbot_output | Yes | Raw combined stdout/stderr from the certbot CLI invocation, on success or failure |
| rate_limit_note | No | Set when the local rate-limit guard refused a production request (with when to retry), or when a Let's Encrypt production limit is close |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial behavioral context beyond the openWorldHint/destructiveHint annotations: it edits existing nginx config in place, adds an SSL block and redirect, reloads nginx automatically, pre-checks DNS, fails fast on resolution errors, and enforces rate-limit guards. It also clarifies that production certs are not silently undone and recommends revoke_cert if needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: prerequisites, operational behavior, failure handling, staging trade-offs, production risks, and the wildcard alternative are all covered without fluff. It is front-loaded with the core action and requirements.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a high-complexity tool with rate limits, config mutation, and lifecycle implications. The description covers prerequisites, side effects, failure modes, retry guidance, staging/production trade-offs, and alternatives. Since an output schema exists, return-value details are not needed in the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though the schema already describes all three parameters at 100% coverage, the description enriches them further with real-world consequences: staging certs are browser-untrusted, production certs are CT-logged and rate-limited, omitted email registers without a contact for renewal warnings, and the domain must have an existing server block.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb and resource: requesting a certificate via certbot --nginx with HTTP-01 validation. It clearly distinguishes this from create_site (does not create a new site), reload_nginx (no separate call needed), and issue_wildcard_cert (for wildcard domains).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use this tool (standard domain cert via HTTP-01), when not to (wildcards should use issue_wildcard_cert), and the prerequisites (existing nginx server block from create_site, domain resolving via check_dns). It also explains staging vs. production usage with clear conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
issue_wildcard_certA
Request a wildcard certificate (domain and *.domain) via certbot --dns-route53 (DNS-01 validation, required since HTTP-01 can't prove ownership of a wildcard). Requires the certbot-dns-route53 plugin installed on the box and AWS credentials in the environment (see README) - fails fast with guidance if credentials are missing. Defaults to staging; a local guard refuses production requests that would exceed Let's Encrypt's rate limits, reporting when to retry. For a single non-wildcard domain, use issue_cert instead.
| Name | Required | Description | Default |
|---|---|---|---|
| No | Contact email for the Let's Encrypt account; omitted registers unsafely-without-email | ||
| domain | Yes | Base domain, e.g. julcap.net - issues it plus *.julcap.net | |
| staging | No | True (default) uses Let's Encrypt's staging CA: untrusted certs, but no rate-limit risk |
Output Schema
| Name | Required | Description |
|---|---|---|
| success | Yes | |
| certbot_output | Yes | |
| rate_limit_note | No | Set when the local rate-limit guard refused a production request (with when to retry), or when a Let's Encrypt production limit is close |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations by explaining that DNS-01 validation is required, that it fails fast with guidance if credentials are missing, that it defaults to staging, and that a local guard blocks production requests exceeding rate limits. This gives the agent a clear behavioral model without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: the core action, the required validation method, prerequisites, failure behavior, staging default, and the sibling alternative are all covered without redundancy. The description is front-loaded and efficiently structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists, return-value details are unnecessary. The description covers prerequisites, failure modes, default behavior, safety guards, and sibling routing, making it fully sufficient for an agent to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema descriptions are already informative. The description adds extra context by explaining the default staging behavior and the local guard's role in rate-limit protection, which helps an agent reason about setting `staging` to false. This is meaningful enrichment beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Request a wildcard certificate (`domain` and `*.domain`)' via certbot with DNS-01 validation. It clearly distinguishes this tool from the sibling `issue_cert` by specifying the wildcard scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool (for wildcard certificates) and when not to: 'For a single non-wildcard domain, use issue_cert instead.' It also gives required prerequisites such as the certbot-dns-route53 plugin and AWS credentials.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_archived_sitesARead-only
List archived nginx configs created by delete_site, newest first. Feed a filename from here into restore_site's archive_filename to restore a specific archive instead of the newest.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | No | Filter to one domain's archives |
Output Schema
| Name | Required | Description |
|---|---|---|
| archives | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=truecars, and the description adds useful behavioral context beyond that: these archives are created by delete_site, are sorted newest first, and their filenames are directly consumable by restore_site. It does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two efficient sentences front-load the core purpose and ordering, then immediately explain the downstream integration with restore_site. There is no filler or repetition of schema information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema, a single well-documented optional parameter, and read-only annotations, the description covers everything an agent needs: what to list, the order, and how the result is used by restore_site. Nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description covers the single optional domain parameter at 100%, so the baseline applies. The description does not add details about the domain filter, but it doesn't need to because the schema already documents it clearly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List'), names the exact resource ('archived nginx configs'), and clarifies the origin ('created by delete_site'). It also gives ordering ('newest first'), which distinguishes it from the sibling list_sites and makes the tool's role obvious.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly connects this tool to restore_site: filenames from this output feed into restore_site's archive_filename to restore a specific archive. It strongly implies when to use this tool over list_sites, though it does not explicitly state a negative condition like 'use list_sites for active sites.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_site_backupsARead-only
List the automatic backups of site configs, newest first. A backup is taken of the existing config just before create_site, update_site, restore_site or rollback_site changes it (the newest 10 per domain are kept). Feed a filename from here into rollback_site's backup_filename to return to a specific version instead of the newest. Not the same as list_archived_sites, which holds configs removed by delete_site.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | No | Filter to one domain's backups |
Output Schema
| Name | Required | Description |
|---|---|---|
| backups | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark readOnlyHint=true and openWorldHint=false, so the description's extra detail is pure added value: newest-first ordering, per-domain retention of 10 backups, and the creation trigger. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, all substantive, with the core action front-loaded. The lifecycle context, workflow hook, and sibling distinction each earn their place without padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with one optional filter and an output schema, this description covers scope, ordering, retention, usage, and exclusions. Nothing an agent needs to call or consume this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single domain parameter is already fully documented in the schema ('Filter to one domain's backups'), so schema_description_coverage is 100%. The description reinforces 'per domain' but does not need to add further semantics; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description opens with a specific verb and resource—'List the automatic backups of site configs, newest first'—and closes by distinguishing itself from list_archived_sites. An agent can immediately tell this tool from its siblings and understand the resource scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains exactly when backups exist (before create_site, update_site, restore_site, rollback_site), how to consume results (feed into rollback_site's backup_filename to restore a specific version), and explicitly says it is not list_archived_sites. This is explicit when/when-not and alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_sitesARead-only
List every nginx server block currently in sites-enabled, with domain, upstream, and whether SSL looks configured. Read-only - parses config files directly, does not shell out to nginx. Use get_site_config for one domain's full raw config.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| sites | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description adds useful behavioral detail beyond that: it parses config files directly and does not shell out to nginx. This clarifies how the read-only operation is performed and implies a lightweight, direct inspection.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core function and scope, then the read-only behavior, then the sibling alternative. Every sentence adds distinct value with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless read-only list tool with an output schema and strong annotations, the description fully covers purpose, scope, behavior, and the relevant sibling tool. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The description adds no parameter syntax, but that is unnecessary here; the empty schema and zero-parameter context fully describe invocation requirements.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('List'), a specific resource ('nginx server block'), and a precise scope ('sites-enabled'). It also lists the returned aspects (domain, upstream, SSL configuration), making the tool's function unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly provides a usage alternative: 'Use get_site_config for one domain's full raw config.' This tells an agent when to prefer this tool versus a sibling, leaving no ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prune_archivesADestructiveIdempotent
Delete archived site configs (from delete_site) older than a threshold. Destructive - requires confirm:true to actually act; without it, lists what would be deleted and changes nothing. Continues past individual failures and reports how many actually got removed.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | No | Must be true to actually delete; false (default) is a dry run | |
| older_than_days | No | Age threshold in days |
Output Schema
| Name | Required | Description |
|---|---|---|
| pruned | Yes | Archive filenames removed - or, when confirm is false, that would be removed |
| message | Yes | |
| success | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark destructiveHint=true and idempotentHint=true, but the description adds essential behavioral detail: the default is a dry run, deletion requires explicit confirmation, it continues past individual failures, and it reports the actual removal count. These are important safety and execution traits not visible from annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each carrying critical information: the core action, the safety mechanism, and the failure/reporting behavior. It is front-loaded with the main purpose and wastes no words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the destructive nature, the description covers the necessary safety context (dry run, confirm requirement), behavior on partial failure, and output reporting. Combined with full parameter schema coverage and an output schema, nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented. The description's 'older than a threshold' matches older_than_days and 'requires confirm:true' matches confirm, but it adds little beyond what the schema already states. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Delete'), a specific resource ('archived site configs (from delete_site)'), and a precise scope ('older than a threshold'). It clearly distinguishes this tool from siblings like list_archived_sites and delete_site by tying it to archived configs and threshold-based pruning.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly explains how to use the tool: confirm:true triggers deletion, while default is a dry-run that lists and changes nothing. It implies when to use it (for bulk pruning of old archived configs) but does not explicitly name alternatives or state when not to use it, leaving slight room for inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reload_nginxAIdempotent
Run nginx -t and reload the live service only if the config test passes - never reloads a broken config. Call this after create_site, delete_site, or restore_site to apply the change; those tools do not reload automatically.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| hint | No | Present on failure: how to recover, e.g. via rollback_site |
| success | Yes | |
| test_output | Yes | Combined stdout/stderr of `nginx -t` |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds valuable behavioral context beyond the annotations: it guarantees 'never reloads a broken config' and explains that it runs a config test first. This supplements the idempotentHint and destructiveHint annotations. However, it doesn't mention what happens if the reload fails (e.g., error handling or return status), which could be relevant for an agent, though the output schema likely covers that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The primary action and safety condition are front-loaded, followed by a direct usage instruction. Every word earns its place, making it highly efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with an output schema, the description is fully sufficient. It explains what the tool does, when to use it, and provides a safety guarantee. There are no missing pieces that an agent needs to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the input schema is empty. Per the rubric, 0 params implies a baseline score of 4. The description correctly omits parameter details since there are none, so nothing more is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Run nginx -t and reload the live service only if the config test passes') and resource ('live service'). It clearly distinguishes itself from sibling tools by specifying when to call it (after create_site, delete_site, or restore_site) and explicitly stating those tools do not reload automatically.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage context: 'Call this after create_site, delete_site, or restore_site to apply the change.' It also clarifies that those tools do not reload automatically, which tells the agent when this tool is necessary. This is clear, actionable guidance for when vs. when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
renew_certA
Run certbot renew, optionally scoped to one cert via --cert-name. Defaults to --dry-run (simulates against Let's Encrypt staging without touching the live cert or rate limit) - pass dry_run:false only when you mean to actually renew. Always disables certbot's random pre-renewal sleep so the call returns synchronously.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | No | Cert name to renew; omit to renew everything due | |
| dry_run | No | True (default) simulates the renewal without touching the live cert |
Output Schema
| Name | Required | Description |
|---|---|---|
| success | Yes | |
| certbot_output | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
It discloses non-obvious behaviors: the staging simulation for dry-run, no rate-limit impact, permanent disarming of certbot's pre-renewal sleep, and synchronous return. These details are valuable beyond the sparse annotations (only `openWorldHint: true`, `destructiveHint: false`) and do not contradict them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded: the first sentence states the core command and scoping, the second covers the default and its safety cue, and the third explains the synchronous return. Every sentence carries essential information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with a full output schema and straightforward behavior, this description covers the default state, the risky escape hatch, scoping, and execution behavior. Nothing needed for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is complete (100%), and the description adds the critical meaning of `dry_run: false` as a deliberate real renewal rather than simulation. It also clarifies that the `domain` parameter is actually a cert-name scope, which is not explicitly stated in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool runs `certbot renew`, scopes to a single cert name via `--cert-name`, and defaults to `--dry-run`. It is distinct from the sibling issue/revoke/delete certificate tools, so an agent can tell when to invoke it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear guidance on when to keep the default dry-run and when to pass `dry_run: false` ('only when you mean to actually renew'). It does not explicitly contrast with issue_cert/revoke_cert for choosing a renewal path, but the conditions for using the tool safely are well specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
restore_siteADestructive
Restore a domain's nginx config from its most recent archive (created by delete_site) and enable it. Destructive to any current config for that domain (which is backed up first, see rollback_site) - requires confirm:true. Test-renders before enabling and rolls back automatically if that fails. Does NOT reload nginx - call reload_nginx afterward.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | Domain to restore | |
| confirm | No | Must be true to actually act; false (default) is a dry run | |
| archive_filename | No | A filename from list_archived_sites; defaults to the newest archive for this domain |
Output Schema
| Name | Required | Description |
|---|---|---|
| message | Yes | |
| success | Yes | |
| reload_required | Yes | True on success - nginx has not actually been reloaded yet |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the destructiveHint annotation, the description discloses key behavior: it backs up the current config first, test-renders before enabling, auto-rolls back on failure, and does NOT reload nginx. It also makes the confirm requirement explicit. This is substantial behavioral context not available from annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly packed sentences: purpose first, then destructive warning/confirm requirement, then the critical non-reload follow-up. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with annotations and an output schema, the description covers the full workflow: source archive, backup behavior, confirmation, safety testing, rollback, and the required nginx reload afterward. No critical operational information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents domain, confirm, and archive_filename well. The description reinforces that confirm:true is required and that the default archive is the most recent, but it does not add new parameter-level meaning beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Restore a domain's nginx config from its most recent archive... and enable it.' It clearly identifies where the archive comes from (delete_site) and what the tool does. This differentiates it from related tools like rollback_site and reload_nginx.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: it is the restoration counterpart to delete_site, requires confirm:true, and must be followed by reload_nginx. It points to rollback_site for the backup aspect and list_archived_sites via archive_filename. It does not explicitly spell out when-not-to-use alternatives, but the context is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
revoke_certADestructive
Revoke a certificate with Let's Encrypt (e.g. after a key compromise) - the CA distrusts it immediately, browser-wide, for any site still serving it. Destructive and effectively irreversible - requires confirm:true. Leaves the cert files on disk; follow with delete_cert to remove them.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | certbot cert name | |
| confirm | No | Must be true to actually revoke; false (default) is a dry run |
Output Schema
| Name | Required | Description |
|---|---|---|
| message | Yes | |
| success | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Although destructiveHint=true already flags mutation, the description adds substantial behavioral specifics: immediate browser-wide distrust, effective irreversibility, the confirm:true requirement, and the fact that cert files are left on disk. This goes well beyond the annotation and prepares the agent for real consequences.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly packed sentences: purpose and trigger, consequence and safety requirement, and post-condition with follow-up. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema, 100% parameter coverage, and annotations. The description adds the missing real-world behavior, irreversibility, and next-step guidance, so an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and already explains both parameters: domain as a certbot cert name and confirm as a guard with false default. The description reinforces the confirm requirement but does not add meaning beyond the schema. A baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Revoke') and resource ('a certificate with Let's Encrypt'), then clarifies its immediate consequence. It clearly distinguishes itself from the sibling delete_cert by noting the files remain on disk, and from issue/renew by its destructive nature.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a concrete trigger ('after a key compromise'), warns that it is destructive and irreversible, and directs the user to follow with delete_cert for file removal. It does not explicitly enumerate when not to use it, but the destructive warning and follow-up relationship provide solid usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rollback_siteADestructive
Undo a config change: replace a site's current nginx config with a backup taken by create_site, update_site, restore_site or a previous rollback_site (see list_site_backups). Defaults to the newest backup, i.e. the config as it was before the last change. The current config is backed up first, so calling this again flips back. Requires confirm:true. Test-renders before keeping the result and leaves the current config in place if nginx -t fails. Does NOT reload nginx - call reload_nginx afterward.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | Domain whose config should be rolled back | |
| confirm | No | Must be true to actually act; false (default) is a dry run | |
| backup_filename | No | A filename from list_site_backups; defaults to the newest backup for this domain |
Output Schema
| Name | Required | Description |
|---|---|---|
| message | Yes | |
| success | Yes | |
| reload_required | Yes | True on success - nginx has not actually been reloaded yet |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the destructiveHint annotation, the description discloses that the current config is backed up first (making the operation reversible), that confirm:true is required for action, that nginx -t is tested before keeping the result with the current config left in place on failure, and that nginx is not reloaded. This is exemplary behavioral disclosure for a state-changing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every clause earns its place: purpose, source of backups, default selection, reversibility, confirmation requirement, safety behavior, and the reload follow-up. The most important guidance is front-loaded, and the length is justified for a destructive operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with destructiveHint, the description covers prerequisites (confirm), failure behavior (nginx -t leaves current config intact), reversibility (current config backed up first), data source (list_site_backups), and the required next step (reload_nginx). Together with the output schema and the input schema, an agent has everything needed to invoke it correctly and understand consequences.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents domain, confirm's dry-run behavior, and backup_filename's default. The description adds workflow context around the parameters (where backups come from, the flip-back behavior, the reload caveat) but does not materially change or extend the parameter semantics beyond what the schema states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Undo a config change' targeting the site's current nginx config. It also distinguishes the mechanism from related operations by identifying which tools produce the backups it reverts to, making it clear this is the revert-to-previous-change tool rather than a generic restore.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states when to use the tool ('undo a config change'), explains the default behavior (newest backup = state before last change), and provides a critical follow-up instruction ('Does NOT reload nginx - call reload_nginx afterward'). It does not explicitly say 'use restore_site instead when you want a specific older backup,' but the backup_filename parameter and reference to list_site_backups imply that workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tail_site_logsARead-only
Tail nginx's access or error log (capped at 1000 lines). domain is a best-effort substring filter, not a true per-vhost filter - see the result's note.
| Name | Required | Description | Default |
|---|---|---|---|
| lines | No | Most recent lines to return, capped at 1000 | |
| domain | No | Best-effort substring filter over each log line - see the result's `note` for its limits | |
| log_type | Yes | Which nginx log to read |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | No | Present when a domain filter was applied, or when the read failed |
| lines | Yes | |
| success | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the operation as read-only, and the description adds useful behavioral facts beyond that: the 1000-line cap and the best-effort substring filtering limitation with a pointer to the note in the result. It does not need to restate read safety.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences front-load the core purpose and cap, then deliver the key caveat about domain filtering. No filler or redundant clauses.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter, read-only log-tail tool with an output schema present, the description covers the essential behavioral constraints (line cap, domain caveat, result note). Additional details would be redundant with the schema and output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the prose largely mirrors the schema (e.g., 'capped at 1000' appears in both). The description adds no new parameter-level detail beyond what the schema already provides, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Tail') on a specific resource ('nginx's access or error log') with an important bound ('capped at 1000 lines'). This clearly distinguishes it from all siblings, none of which read logs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There are no sibling log-reading tools to contrast with, but the description gives clear context for invoking the tool and explicitly warns that the domain parameter is not a true per-vhost filter. It stops short of an explicit when-to-use/when-not-to-use statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_siteAIdempotent
Update an existing site's upstream by rewriting its proxy_pass directive(s) in place - everything else in the config, including any SSL server block issue_cert/certbot added, is left untouched. Fails if the domain has no existing config (use create_site instead) or has no proxy_pass directive to update. Test-renders before keeping the change and rolls back automatically if nginx -t fails. The previous config is backed up first, so a bad-but-valid change can be undone later with rollback_site. Does NOT reload nginx - call reload_nginx afterward.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | Domain of the existing site to update | |
| upstream_host | Yes | New hostname or IP nginx should proxy_pass to | |
| upstream_port | Yes | New TCP port on the upstream host |
Output Schema
| Name | Required | Description |
|---|---|---|
| success | Yes | |
| test_output | Yes | Output of `nginx -t` against the rewritten config, or an explanatory message if nothing was changed |
| backup_created | No | True once the previous config was backed up; undo with rollback_site |
| reload_required | Yes | True on success - nginx has not actually been reloaded yet |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations, disclosing that it rewrites in place, leaves SSL cert blocks untouched, fails on missing config or missing proxy_pass, test-renders with automatic rollback on nginx -t failure, backs up the prior config, and skips the reload. This is rich behavioral context for a mutation tool, and nothing contradicts the idempotentHint=true or destructiveHint=false annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each earning its place: purpose/scope, failure conditions with alternative, safety and backup behavior, and the critical no-reload caveat. Front-loaded with the primary action and progressively revealing caveats in logical order — appropriate length for a mutation tool with meaningful failure and recovery semantics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-required-param tool with a simple schema and an output schema present, the description covers everything an agent needs to call it correctly: purpose, prerequisites, failure modes, safety guarantees, recovery path, and required follow-up. No significant gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with clear descriptions for all three parameters, so the schema carries the semantic weight. The description marginally reinforces that upstream_host/upstream_port become the new proxy_pass target and that domain must reference an existing config, but it adds no substantial meaning beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb-resource pair: 'Update an existing site's upstream by rewriting its proxy_pass directive(s) in place'. This precisely distinguishes it from siblings like create_site (which creates new configs) and rollback_site (which restores backups). The scope is exact — it names what is changed and what is left untouched.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes to alternatives: 'use create_site instead' when no config exists, 'rollback_site' for undoing a bad-but-valid change, and 'call reload_nginx afterward' because it does not reload. The when-to-use, when-not-to-use, and follow-up steps are all stated without leaving anything to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.1.2- Changed
check_cert_expiry2 fields changed- added
Output schema / properties / certificates / items / properties / domainsAdded value: +{ + "description": "Every name on the certificate, e.g. example.com and *.example.com", + "items": { + "type": "string" + }, + "type": "array" +} - changed
Output schema / properties / certificates / items / requiredPrevious value: -[ - "domain", - "expires_at", - "days_remaining", - "auto_renew_enabled" -]New value: +[ + "domain", + "domains", + "expires_at", + "days_remaining", + "auto_renew_enabled" +]
- Changed
create_site1 field changed- added
Output schema / properties / backup_createdAdded value: +{ + "description": "True if an existing config was replaced; it was backed up first (see rollback_site)", + "type": "boolean" +}
- Added
diagnose_site - Changed
issue_cert1 field changed- added
Output schema / properties / rate_limit_noteAdded value: +{ + "description": "Set when the local rate-limit guard refused a production request (with when to retry), or when a Let's Encrypt production limit is close", + "type": "string" +}
- Changed
issue_wildcard_cert1 field changed- added
Output schema / properties / rate_limit_noteAdded value: +{ + "description": "Set when the local rate-limit guard refused a production request (with when to retry), or when a Let's Encrypt production limit is close", + "type": "string" +}
- Added
list_site_backups - Changed
reload_nginx1 field changed- added
Output schema / properties / hintAdded value: +{ + "description": "Present on failure: how to recover, e.g. via rollback_site", + "type": "string" +}
- Added
rollback_site - Changed
update_site1 field changed- added
Output schema / properties / backup_createdAdded value: +{ + "description": "True once the previous config was backed up; undo with rollback_site", + "type": "boolean" +}
24 tool updates
v0.1.1- Changed
check_cert_expiry1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "http://json-schema.org/draft-07/schema#", + "additionalProperties": false, + "properties": { + "certificates": { + "items": { + "additionalProperties": false, + "properties": { + "auto_renew_enabled": { + "type": "boolean" + }, + "days_remaining": { + "description": "Negative if the certificate has already expired", + "maximum": 9007199254740991, + "minimum": -9007199254740991, + "type": "integer" + }, + "domain": { + "type": "string" + }, + "expires_at": { + "description": "Expiry date/time as reported by certbot", + "type": "string" + } + }, + "required": [ + "domain", + "expires_at", + "days_remaining", + "auto_renew_enabled" + ], + "type": "object" + }, + "type": "array" + } + }, + "required": [ + "certificates" + ], + "type": "object" +}
- Added
check_dns - Added
check_upstream_health - Added
create_domain_record - Removed
create_server_block - Added
create_site - Added
create_txt_record - Added
delete_cert - Added
delete_domain_record - Added
delete_site - Added
delete_txt_record - Added
get_nginx_status - Changed
get_site_config3 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - changed
Input schema / properties / domain / descriptionPrevious value: -"The domain to look up, e.g. my.domain.com"New value: +"Domain as it appears in sites-available, e.g. mysite.julcap.net" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "http://json-schema.org/draft-07/schema#", + "additionalProperties": false, + "properties": { + "domain": { + "type": "string" + }, + "raw_config": { + "description": "Full contents of the nginx config file, verbatim", + "type": "string" + } + }, + "required": [ + "domain", + "raw_config" + ], + "type": "object" +}
- Changed
issue_cert5 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / domain / descriptionAdded value: +"Domain to request a certificate for. Must already resolve (see check_dns) and already have an nginx server block from create_site - certbot edits that existing config rather than creating one." - added
Input schema / properties / email / descriptionAdded value: +"Contact email registered with the Let's Encrypt account, used for renewal-failure and expiry notices. Omitted registers with --register-unsafely-without-email, so Let's Encrypt cannot warn you if a future automated renewal fails." - added
Input schema / properties / staging / descriptionAdded value: +"True (default) uses Let's Encrypt's staging CA - browser-untrusted certs, but exempt from production rate limits; use for testing the flow. False requests a real, browser-trusted cert and counts against production rate limits." - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "http://json-schema.org/draft-07/schema#", + "additionalProperties": false, + "properties": { + "certbot_output": { + "description": "Raw combined stdout/stderr from the certbot CLI invocation, on success or failure", + "type": "string" + }, + "dns_check": { + "additionalProperties": false, + "description": "Present only when the DNS pre-check failed, before certbot was even invoked", + "properties": { + "record_type": { + "description": "Only present when resolves is true", + "enum": [ + "A", + "AAAA", + "CNAME" + ], + "type": "string" + }, + "resolves": { + "type": "boolean" + }, + "values": { + "description": "Resolved values (IPs, or the CNAME target); only present when resolves is true", + "items": { + "type": "string" + }, + "type": "array" + } + }, + "required": [ + "resolves" + ], + "type": "object" + }, + "success": { + "type": "boolean" + } + }, + "required": [ + "success", + "certbot_output" + ], + "type": "object" +}
- Added
issue_wildcard_cert - Added
list_archived_sites - Changed
list_sites1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "http://json-schema.org/draft-07/schema#", + "additionalProperties": false, + "properties": { + "sites": { + "items": { + "additionalProperties": false, + "properties": { + "config_path": { + "description": "Absolute path to the enabled config file", + "type": "string" + }, + "domain": { + "description": "server_name parsed from the config, or the filename if not found", + "type": "string" + }, + "ssl_enabled": { + "description": "True if the config listens on 443 ssl or sets ssl_certificate", + "type": "boolean" + }, + "upstream": { + "description": "proxy_pass target, or null if none was found", + "type": [ + "string", + "null" + ] + } + }, + "required": [ + "domain", + "config_path", + "upstream", + "ssl_enabled" + ], + "type": "object" + }, + "type": "array" + } + }, + "required": [ + "sites" + ], + "type": "object" +}
- Added
prune_archives - Changed
reload_nginx1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "http://json-schema.org/draft-07/schema#", + "additionalProperties": false, + "properties": { + "success": { + "type": "boolean" + }, + "test_output": { + "description": "Combined stdout/stderr of `nginx -t`", + "type": "string" + } + }, + "required": [ + "success", + "test_output" + ], + "type": "object" +}
- Added
renew_cert - Added
restore_site - Added
revoke_cert - Added
tail_site_logs - Added
update_site
6 tool updates
v0.1.0- First observed
check_cert_expiry - First observed
create_server_block - First observed
get_site_config - First observed
issue_cert - First observed
list_sites - First observed
reload_nginx
TDQS
Scored across 26 tools
Most tools have clearly distinct resource-action targets, and descriptions frequently cross-reference the correct alternative when two tools could be confused. The main ambiguities are restore_site vs rollback_site (both restore saved configs, just from different sources) and create_site's replace-existing behavior overlapping with update_site.
Every tool follows a consistent snake_case verb_noun pattern: list_sites, create_site, delete_cert, issue_cert, check_dns, reload_nginx, etc. Complementary pairs like create_domain_record/create_txt_record and revoke_cert/delete_cert show a predictable and readable convention.
26 tools is above the comfortable range and even past the 16-25 heavy band, making the surface feel large. However, the count is largely justified by covering three distinct subdomains—nginx site lifecycle, certbot certificate lifecycle, and Route53 DNS records—so it is heavy but not chaotic.
The tool set covers the full nginx site lifecycle (create, read, update, delete, restore, rollback, backups, logs, reload), certificate lifecycle (issue, renew, revoke, delete, expiry), and DNS record management. Minor gaps exist: there is no tool to remove certificate references from an nginx config before delete_cert, and no general DNS record listing tool.
Maintenance
Related MCP Connectors
Provides capabilities that let LLM agents perform a range of infrastructure management tasks.
Read-only MCP access to a documented IT fleet: state, changes, posture. 15 tools.
Deploy and manage applications, databases, domains, and git repos
- emisarOAuthdev.emisar
Let AI operate servers without SSH. Choose actions, approve risky changes, and audit every step.
Related MCP Servers
- AlicenseBqualityCmaintenanceEnables management of Nginx Proxy Manager instances for configuring proxy hosts, requesting Let's Encrypt SSL certificates, and managing access lists. It allows users to control their web proxy infrastructure through natural language commands in MCP-compatible environments.503MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to manage Nginx Proxy Manager instances through natural language, covering 28 tools for proxy hosts, certificates, streams, and more.MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to manage Nginx Proxy Manager, including proxy hosts, SSL certificates, streams, and more.8 npmISC
- AlicenseBqualityCmaintenanceEnables AI agents to manage Nginx Proxy Manager (reverse proxies, streams, redirects, and Let's Encrypt certificates) through natural language commands.3011 npmMIT