Skip to main content
Glama

upkeep-mcp

An MCP server for the recurring checks behind ongoing website maintenance: domains, SSL certificates, uptime, technical SEO and accessibility — all from publicly available information.

Built for people who look after a portfolio of client sites on a retainer, not just a single domain. The goal is to answer one question quickly: what needs attention this week?

Status

Early development, built in public phase by phase. Everything that needs no browser is implemented and useful today.

Tool

Purpose

Status

domain_check

Registration expiry, registrar, nameserver agreement, DNS records, DNSSEC

Available

ssl_check

Certificate expiry, issuer, chain validity, SAN coverage, TLS version

Available

uptime_check

HTTP status, response time, redirect chain, HTTPS upgrade, security headers

Available

health

Server name, version, Node.js version, uptime

Available

seo_audit

Title, meta, headings, canonical, robots.txt, sitemap, broken links

Available

site_crawl

Duplicate titles, broken links and stray noindex across a whole site

Available

portfolio_report

All of the above across a portfolio, sorted by urgency

Available

accessibility_audit

WCAG violations via axe-core, in a real browser

Available

It also exposes the portfolio://sites resource (the site list, for a client to read without spending a tool call) and the quarterly_report prompt (turns a portfolio run into the report a client actually reads).

Published on npm, so it installs with one command — see Installation. A public instance is also running, for anyone who would rather point a client at a URL than run anything — see The hosted instance.

Related MCP server: websec-auditor

What it looks like

> Is example.com about to expire?

example.com expires 2027-08-13 (346 days).
Registrar: RESERVED-Internet Assigned Numbers Authority.
Nameservers: elliott.ns.cloudflare.com, hera.ns.cloudflare.com.
Resolves: apex yes, www yes. DNSSEC: delegation signed.
> Check the certificate on expired.badssl.com

expired.badssl.com:443 certificate expires 2015-04-12 (-4159 days).
Issued by COMODO RSA Domain Validation Secure Server CA.
Chain does not verify (CERT_HAS_EXPIRED). Negotiated TLSv1.2.
Host matched via *.badssl.com.
Revocation not established: http://ocsp.comodoca.com could not answer: the
responder is not authorised to answer for this certificate.

Needs attention:
- [critical] The certificate expired 4159 days ago.
- [critical] The certificate chain does not verify: CERT_HAS_EXPIRED.
> Check the certificate on revoked.grc.com

revoked.grc.com:443 certificate expires 2026-10-18 (42 days).
Issued by Certera RSA DV SSL CA 2.
Chain verifies. Negotiated TLSv1.2.
Host matched via revoked.grc.com.
Revoked on 2025-09-18.

Needs attention:
- [critical] The certificate was revoked on 2025-09-18; browsers that check
  revocation will refuse the site.

Note what the second one does: the chain verifies. Node performs no revocation lookup of its own, so that certificate completes a handshake and reports as trusted. Only asking its issuer finds the problem.

Every tool also returns structured data alongside the text, so results can be sorted, filtered and fed into a report. Full input and output for each tool is in examples/, and a whole portfolio session — the weekly triage, a drill-down, and what the comparison against the previous run will and will not claim — is in examples/conversation.md.

The tools

domain_check

Input: domain — a bare domain, a full URL, or an internationalised name. A subdomain is reduced to its registrable domain, since that is what a registration belongs to.

Returns the expiry date and days remaining, the registrar and its IANA ID, registry statuses, A/AAAA/NS/MX/TXT/CAA records, whether the apex and www resolve, whether the delegation is signed with DNSSEC, and what the domain's SPF and DMARC records say about who may send email as it.

Email authentication is read from the domain's own DNS — the SPF record at the apex, the DMARC record at _dmarc. A record that is absent is reported as information, because it is a standing improvement rather than something that broke this week. A record that is present and wrong is a warning, because it fails right now: two SPF records make receivers skip SPF entirely, and +all authorises the whole internet to send as the domain.

DKIM is deliberately not reported. A DKIM key lives at <selector>._domainkey, and a selector cannot be discovered — only guessed, one DNS query per guess. That is subdomain enumeration, which this project does not do, so a domain with no DKIM and one whose selector was not guessed are left indistinguishable rather than the second being reported as the first.

Each of the domain's own nameservers is then asked about the zone directly, over TCP port 53 with recursion off, which is the one question a recursive resolver cannot answer: it replies with whatever one server told it and does not say which. That finds a server left in the delegation after a migration — it answers REFUSED, or its own hostname stopped resolving, and every resolver query landing there is slow or fails — and it finds nameservers holding different versions of the zone, which is the "it works for me but not for my colleague" outage.

What is a fault and what is only unestablished are graded apart. Resolvers ask over UDP first and this server can only use TCP, so a nameserver that refuses TCP is not a broken one: sapo.pt's four all refuse it and the domain resolves perfectly. Different serials are info too — github.com runs two providers that do not transfer between them, so four of its nameservers report 1656468023 and four report 1, and nothing is wrong. A hostname that does not resolve, or an answer without authority for the zone, is broken for everybody and is a warning. Pass checkNameservers: false to skip the whole thing — portfolio_report does, because a portfolio would pay this per site and a deployment that cannot open port 53 would grade every site at once as unestablished. docs/adr/0020 records why it speaks DNS by hand and what it deliberately does not check.

ssl_check

Input: domain, optional port (443 by default).

Returns expiry and days remaining, issuer, whether the chain verifies and why not when it does not, which hostnames the certificate covers and via which SAN entry, the negotiated TLS version and cipher, and whether the certificate has been revoked. Expired, self-signed and untrusted certificates are inspected and reported rather than refused — those are the ones worth finding.

Revocation is checked over OCSP. The response a server staples to the handshake is preferred, because it costs no request at all; when there is none, the responder named in the certificate is asked directly. An answer is only believed once its signature verifies against the issuing authority, and only once its CertID is shown to be about the certificate that was actually served — a server serving a revoked certificate alongside a valid response for a different one is otherwise the easy way to fake a clean result.

Many healthy certificates cannot be checked at all: since 2025 the two largest issuers, Let's Encrypt and Google Trust Services, publish no OCSP responder and distribute revocation by CRL. That is reported as an unavailableReason and produces no finding, because it is a decision of the certificate authority and nothing the site owner can act on. A responder that was asked and would not answer is different, and gets one unknown.

uptime_check

Input: url — a full URL, or a bare domain, which is tried over HTTPS.

Returns the status code, response time, every hop of the redirect chain, whether plain HTTP is upgraded to HTTPS, the HSTS policy, and the security headers worth reporting on.

seo_audit

Input: url — the page to audit, plus optional checkLinks (true by default) and maxLinks (25 by default).

Returns title and meta description with their lengths, the heading structure, canonical, lang, viewport, Open Graph, hreflang alternates, the images with no alt attribute, the state of robots.txt and the sitemap, and which internal links are broken.

A gzipped sitemap is unpacked before it is read. Whether it is gzipped is decided by the first two bytes rather than by the file name or the content type, because plenty of files called .xml.gz are not, and plenty that are get labelled text/xml.

The sitemap is then checked against the rules of the protocol, because a file that answers 200 and parses is not the same as a file that works: a root element with no sitemaps.org namespace is dropped whole, an entry on another host is discarded, and one unescaped & makes the document ill-formed XML and costs every entry after it. Each broken rule is reported once with the number of entries that break it and one example, so a mistake repeated across fifty thousand URLs reads as one thing to fix. A rule that only costs a hint — a <lastmod> that is not a W3C Datetime, a <changefreq> outside its seven values — is graded info; one that costs the entry or the file is a warning.

robots.txt is read before anything else is requested and is obeyed — for the page itself and for every internal link. A page this crawler is not allowed to read is reported as such and is never fetched, and an unreadable robots.txt is treated as forbidding everything, as RFC 9309 requires.

> Audit the homepage of example.com

https://example.com/ answered 200. Title: "Example Domain".
1 h1, 0 images without alt, 0 internal links (0 checked, 0 broken).
Sitemap: the sitemap URL answered 404.

Needs attention:
- [warning] The page has no meta description, so search engines will write their own summary of it.
- [info] The page declares no canonical URL, which is how duplicate addresses for the same page get separated.
- [info] The page has no og:title or no og:image, so it will share poorly on social networks and in messaging apps.
- [info] There is no sitemap at https://example.com/sitemap.xml: the sitemap URL answered 404.
- [info] The site publishes no robots.txt. Nothing is blocked, but the sitemap cannot be declared there either.

site_crawl

Input: url — the page to start from — plus optional maxPages (25 by default, 100 at most) and maxDepth (3 by default).

Walks the site breadth-first from that page and reports what one page cannot tell you about itself: which pages share a title, and therefore compete with each other for the same search result; which share a meta description; which internal links are broken and, crucially, which page links to them; which pages still ask not to be indexed after a rebuild; and how much of the site was reachable at all.

It stays on the origin you start it on — https://example.com and https://www.example.com are different origins, and a crawl that wandered between them would report one site's pages as duplicates of the other's. A link that redirects off the site is counted and left alone: robots.txt was read for this origin and does not speak for anybody else.

Three budgets bound it: pages, depth, and a two-minute deadline. Whichever one ended the crawl is reported along with how many URLs were found and not visited, because a report that does not say it saw a quarter of the site is worse than no report. robots.txt is read first and obeyed for every URL before it is requested; an unreadable one refuses the crawl outright, per RFC 9309.

It is deliberately not part of portfolio_report. Twenty-five pages per site across a portfolio is a different order of cost, and docs/adr/0021 — which docs/adr/0010 predicted — records why it is a tool of its own rather than a depth parameter on seo_audit.

> Crawl example.com and tell me what needs fixing

Crawled 25 pages of https://example.com, 3 levels deep.
2 broken internal links, 1 title used more than once.
Stopped at the page budget; 11 URLs were not visited.

Needs attention:
- [warning] 2 internal links are broken: https://example.com/old-pricing (it answered 404) linked from https://example.com/, https://example.com/team/ana (it answered 404) linked from https://example.com/team.
- [warning] 1 title is used by more than one page, so those pages compete with each other in search results; the widest is "Services" on 4 pages.
- [info] 6 of 25 pages have no meta description, so search engines will write their own summary of them.
- [info] The crawl stopped at the page budget with 11 URLs found and not visited, so everything here describes the part of the site that was looked at.

portfolio_report

Input: sites inline, or file (defaults to sites.json), plus optional checks and tags.

Runs every check across the whole portfolio with bounded concurrency and returns one report ordered by what needs action first: what is down, what expires soonest, what regressed since the last run. A site that cannot be checked becomes a finding, never a failure of the whole report.

> What needs attention across my sites this week?

3 sites checked: 0 critical, 2 warning, 0 unknown, 1 fine.

Needs action:
- [warning] Example Ltd: Plain HTTP does not redirect to HTTPS.
- [warning] Example Ltd: No Strict-Transport-Security header is sent.
- [warning] Example Net: Plain HTTP does not redirect to HTTPS.
- [warning] Example Net: No Strict-Transport-Security header is sent.

No change is reported: this server has not run a report on this portfolio before,
and the portfolio names no history file, so nothing survived the last restart.
History for this portfolio is kept in memory only. To compare across restarts,
add a "history" path to the portfolio file.

Nothing to do: Example Foundation.

Comparing across restarts. By default the previous run lives in this server process and nowhere else, so a client that restarts daily gets a comparison that never spans more than a day — and a quarter-over-quarter report cannot be produced from it at all. Add one line to the portfolio file:

{ "version": 1, "history": "upkeep-history.json", "sites": [...] }

and the run is written there instead, beside the portfolio file itself. The next report compares against it however many restarts later, up to ninety days. The file names your clients and says which were broken, so it is created readable by your account alone and replaced on every run; without that line nothing is written at all. docs/adr/0018 records the reasoning.

Each site can set maxLinks — how many internal links the seo check may request, 0 for none. It is the setting that decides what a run costs: measured over twenty sites, the portfolio takes about eight seconds without seo and around forty with it, because link checking is one request per link paced at half a second per host.

The portfolio file format is documented in sites.example.json. Copy it to sites.json — which is gitignored, so a real client list never gets committed. file is resolved against the directory the client started the server in, so give the full path when that directory is not yours.

accessibility_audit

Input: url, plus optional standard (wcag2aa by default; also wcag2a, wcag21aa, wcag22aa, best-practice).

Opens the page in a headless browser and runs axe-core over it. Returns the rules that failed, how many elements failed each, CSS selectors for the first few, and how many rules axe could not decide on its own.

This is the only tool that needs a browser, and it is optional: nothing is downloaded when you install this project. Run npx playwright install chromium once if you want it. Without it the tool says so and names that command, and every other check carries on.

Automated rules find roughly a third of accessibility problems. A page with no violations passed the machine-checkable part, which is not the same as being usable — which is why the count of undecided rules is reported alongside.

Installation

Requires Node.js 22 or newer. Nothing else: installing downloads no browser, and every check works without one except accessibility_audit.

Claude Code

claude mcp add upkeep -- npx -y upkeep-mcp

Claude Desktop

Add the server to claude_desktop_config.json:

{
  "mcpServers": {
    "upkeep": {
      "command": "npx",
      "args": ["-y", "upkeep-mcp"]
    }
  }
}

Restart the client and ask it to run the health tool. It answers with the server version, the Node.js version and how long the process has been up.

A desktop client starts a server in a directory of its own choosing, usually /. That matters for one thing only: portfolio_report and the portfolio://sites resource look for sites.json there. Pass sites inline, or give portfolio_report the full path — file: "/Users/you/sites.json".

From source

For development, or to run a branch:

git clone https://github.com/tiagocalado86/upkeep-mcp.git
cd upkeep-mcp
npm install
npm run build
claude mcp add upkeep -- node /absolute/path/to/upkeep-mcp/dist/index.js

The hosted instance

For anyone who cannot or would rather not run a server, there is a public one:

https://upkeep-mcp-1080119881249.europe-west1.run.app/mcp

Point any MCP client that takes a remote server URL at it — in Claude, as a custom connector. Opening the host in a browser gives a plain page saying what it is.

It is a demo. No authentication, no availability promise, no support, and it may be switched off without notice. npx -y upkeep-mcp is the supported way to run this, and it is what you want if these checks matter to your work.

Every tool works there, accessibility_audit included — the image ships a browser, and every request that browser makes goes through the same rules as the rest of the server. Someone who cannot run a server themselves should not get a weaker tool than someone who can.

One thing genuinely differs, and it is a property of running in public rather than a compromise: it contacts only public addresses, and only three ports — 443, 80 when checking whether plain HTTP upgrades, and 53 for the nameservers a domain itself publishes. So it refuses to check anything on your own network, localhost included. Use the stdio server for those.

Running your own over HTTP

The transport is the same one the hosted instance uses:

npm run build
npm run start:http -- --port 8080

The HTTP entrypoint is not the stdio one with a socket attached: a stranger is not the person who started the process, so it applies the address and port rules above and admits traffic through a per-caller rate limit. docs/deploying.md covers running it on Google Cloud Run, and docs/adr/0012 explains the guard and what it does not close.

Security & privacy

This server never asks for, accepts or stores credentials. It reads only information that any person with a browser or a DNS resolver could read.

  • No API keys, tokens or passwords — for any service, ever.

  • No intrusive behaviour: no port scanning, no subdomain brute-forcing, no vulnerability probing. It inspects public configuration; it is not an offensive tool.

  • robots.txt is respected on any page crawl, with per-host rate limiting and an identifiable User-Agent carrying a contact URL.

  • No persistent sensitive state. Caching is in memory only, with a TTL. There is no database. The one file this server ever writes is a portfolio run snapshot, and only when a portfolio file names a history path for it.

  • The only third parties contacted are the ones that hold the answer: the registry's own RDAP server, IANA's RDAP bootstrap file, and cloudflare-dns.com for the one question node:dns cannot ask (whether a DNSSEC delegation is signed). SECURITY.md lists them and what each one learns.

Limitations

Stated plainly, because a tool that hides what it cannot do is worse than one that does less.

  • Certificate revocation lists are not downloaded. Revocation is checked over OCSP, and only over OCSP. A CRL is a file of every serial an authority has ever revoked — megabytes, fetched to answer one question about one certificate — so a certificate whose issuer publishes no responder is reported as unchecked, with the reason, rather than being judged from a file this tool declined to read. Since 2025 that covers Let's Encrypt and Google Trust Services, which is a large share of the web. See docs/adr/0017.

  • An OCSP answer is a point in time, not a subscription. Responders pre-sign about a week ahead, so good means "not revoked as of thisUpdate", and a certificate revoked an hour ago may still read as good until the authority publishes its next answer. producedAt, thisUpdate and nextUpdate are all reported so the age of the answer is visible rather than implied.

  • Some registries publish no expiry date. .de, .nl, .no, .au and .fi do not publish one over any protocol. The result names the registry and says so, instead of showing an indefinite "unknown". Registration data comes from RDAP only; there is no WHOIS fallback, and docs/adr/0004 explains why.

  • DNSSEC is not validated. The tool reports whether a delegation is signed and where it learned that. It never claims to have validated a chain.

  • The parent's delegation is not compared with the zone's own. "The nameservers at the registrar are not the nameservers in the zone" is the other classic delegation fault, and answering it means querying the parent zone's servers for a referral — a second hop and a different feature. What is checked is whether the zone's own nameservers agree with each other.

  • Response time includes connection setup. It is wall clock to the first response headers, covering DNS, TCP and TLS, so it is not a measure of server processing time.

  • seo_audit audits one page, not a site. It requests the page's internal links to find broken ones, but it does not crawl: there is no second level. Auditing a site means calling it for the pages that matter.

  • The sitemap is checked against the protocol's rules, not against its XSD. It establishes that the document exists, declares <urlset> or <sitemapindex>, how many <loc> entries it holds, and whether it arrived gzipped — a sitemap.xml.gz is unpacked before it is read, capped at 16 MiB so that a decompression bomb is refused rather than unpacked. It then checks the rules that decide whether a consumer keeps an entry: the namespace, <loc> present, absolute, escaped, within 2048 characters and on the sitemap's own host, <lastmod> a W3C Datetime, <changefreq> and <priority> within their ranges, and the 50,000-entry limit. There is no XML parser and no schema validator here, so four rules are deliberately left unchecked — the 50 MB size limit, an index listing another index, duplicate entries, and whether a <lastmod> is true. docs/adr/0019 lists them and says why. The rules are checked against what was read, so a sitemap past the read limit is judged on the entries before the cut and the report says it was truncated — nodejs.org's stray entries on another host sit past it, and are not reported.

  • A page nested thousands of levels deep is refused, not audited. HTML parsing costs roughly the square of the nesting depth, so a document built to be absurd would block the server for minutes. seo_audit measures the depth first and reports the refusal.

  • Accessibility is only checked as far as a machine can. Automated rules find roughly a third of real problems. The tool reports what axe could not decide rather than counting it as a pass, but no green result here is an accessibility statement.

  • Nothing here judges how a page ranks. seo_audit reports what is in the HTML. Rankings depend on things no public endpoint exposes.

  • "What changed since last time" lasts as long as the server process, unless you ask for otherwise. By default the previous run is held in memory and never written anywhere, so a restarted server — which for a desktop MCP client is a daily event — has nothing to compare against, and says so rather than implying nothing changed. Adding a history path to the portfolio file makes the comparison survive a restart, for up to ninety days; nothing is written without it. Only sites both runs measured the same way are compared, so a quick uptime-only pass never invents regressions in the run after it. docs/adr/0011 and docs/adr/0018 explain the trade.

  • Only the previous run is kept, never a series. Each run replaces the last. A trend over quarters is a different feature with a different storage question, and it has not been asked for.

  • Certificates and domains are judged on different clocks. A registration is a warning inside 30 days; a certificate only inside 14. ACME clients renew with 30 days left, so warning that early would fire on nearly every healthy site.

Documentation

License

MIT — see LICENSE.

Available Tools

8 tools
accessibility_auditWCAG violations on one pageA
Read-onlyIdempotent

Opens one page in a real browser and runs axe-core over it, reporting the WCAG rules it fails, how many elements fail each one, and where they are.

Use it to answer "does this page meet WCAG 2.2 AA?", "what would an accessibility audit flag?" or "which of these fixes matters most?". It is the check to run before a site goes live, and when a client asks about accessibility obligations.

Do not use it for metadata, headings or broken links, which seo_audit reports without a browser and far faster. It audits one page, not a site, and it renders that page: it is much slower than every other tool here.

It needs a browser. Nothing else in this server does, so if none is installed this tool says so and names the one command that fixes it, while every other check keeps working. A hosted instance runs it too: the published image ships a browser, and every request that browser makes goes through the same public-address and web-port rules as the rest of this server.

Automated rules find roughly a third of accessibility problems. A page with no violations is a page that passed the machine-checkable part, which is not the same as being usable, and the count of undecided rules is reported for exactly that reason. robots.txt is obeyed. Returns findings ordered by urgency, worst first.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe page to audit, e.g. "https://example.com/contact". A bare domain is accepted and read over HTTPS. One page is audited, not a site.
standardNoWhich rules to run. Defaults to "wcag2aa", the level most accessibility policies and procurement rules are written against. Each level includes the ones below it. "best-practice" runs axe's non-WCAG advice instead, which is useful but is not a legal standard anywhere.

Output Schema

ParametersJSON Schema
NameRequiredDescription
urlYesThe URL that was requested.
auditedYesWhether the page was audited at all. False when robots.txt forbids it.
finalUrlYesWhere the browser ended up.
findingsYesWhat needs attention, worst first.
severityYesHow much attention this needs. "critical" means act now; "warning" means act this month; "unknown" means the check could not establish the fact, which is not the same as it being fine.
standardYesWhich standard was asked for.
checkedAtYesWhen the audit ran, ISO 8601 in UTC.
pageTitleYesThe title as rendered, which markup alone may not show.
passCountYesHow many rules passed.
axeVersionYesWhich axe-core did the judging, so a result can be reproduced.
violationsYesEvery rule that failed, worst impact first.
violationCountYesHow many rules failed.
incompleteCountYesRules axe could not decide on its own. These are not passes: they are the ones needing a person, and a page with none is unusual rather than perfect.
affectedElementsYesHow many elements failed a rule, counting repeats.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive, but the description adds substantial context beyond them: a browser dependency unique in this server, a self-diagnosing error naming the fix command, hosted-image behavior and outbound public-address/web-port rules, robots.txt obedience, and worst-first ordering. It also discloses the fundamental limitation that automated rules catch roughly a third of issues and that a clean result is not usability.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening sentence is front-loaded with the mechanism and output, and the paragraphs are organized by purpose, exclusions, environment, and caveats. It runs long at five paragraphs, but nearly every sentence carries non-redundant selection or expectation-setting information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-shape explanation is unnecessary; the description still covers scope, cost, environment prerequisites, error behavior, network constraints, and interpretive limits on results. Nothing an agent needs to invoke this correctly or set expectations is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both `url` and `standard` are fully documented in the schema itself, including the wcag2aa default and the enum semantics. The description adds no syntax or format detail beyond restating the one-page scope, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (opens, runs axe-core) and resource (one page) with the exact output (WCAG rules failed, element counts, locations). It explicitly distinguishes itself from seo_audit by naming that sibling and the checks it does not cover, so an agent can separate the two without reading either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives the user questions it answers, the moment to run it ('before a site goes live', client accessibility questions), and an explicit exclusion routing other checks to seo_audit. It also warns that it audits one page rather than a site and is much slower than every other tool here, which is exactly the trade-off an agent needs to select correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

domain_checkDomain registration and DNSA
Read-onlyIdempotent

Reports when a domain registration expires, who the registrar is, and how the domain is configured in DNS — nameservers, address records, mail exchangers, TXT and CAA records, whether the delegation is signed with DNSSEC, and what its SPF and DMARC records say about who may send email as it.

Use it to answer "is this domain about to lapse?", "who do we renew this with?", "where does this domain point?" or "why does the apex not work when www does?". It is the right first call when a site has gone dark for no obvious reason.

Use it too for "why is this client's email going to spam?" or "can someone spoof this domain?" — SPF and DMARC are read from the domain's own DNS.

It also asks each of the domain's own nameservers, directly, whether they agree about the zone. That answers "why does this site work for some people and not others?" and "is this old nameserver still in the delegation?" — a question no recursive resolver can answer, because it replies with whatever one server told it and does not say which. Pass checkNameservers: false to skip it.

Do not use it to check whether a website responds — that is uptime_check — or to inspect an SSL certificate, which is ssl_check. It reads only what registries and DNS publish. It reports no DKIM: finding a DKIM key needs its selector, and a selector cannot be discovered without guessing at names, which this project will not do.

Registration data comes from RDAP. Some country registries (.de, .nl, .no, .au, .fi) publish no expiry date at all; the result says so explicitly rather than reporting a gap as if it were an unknown. An expiry inside 30 days is reported as a warning and inside seven days as critical — a manual renewal needs that much lead time. Returns findings ordered by how much attention they need, worst first.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesThe domain to check, e.g. "example.com" — no scheme, no trailing slash. A full URL such as "https://example.com/pricing" is also accepted and reduced to its hostname. Internationalised names ("café.pt") are accepted and converted automatically. Registration is a property of the registrable domain, so "www.shop.example.co.uk" is checked as "example.co.uk".
checkNameserversNoWhether to ask the domain's own nameservers whether they agree about it, which no recursive resolver can answer. Costs one DNS query over TCP to each nameserver the domain publishes, in parallel, and finds a server left in the delegation that no longer serves the zone and a zone edited on one server and never transferred to the others. Defaults to true. Set false to skip it — the rest of the check is unaffected. Example: false

Output Schema

ParametersJSON Schema
NameRequiredDescription
dnsYesDNS records for the registrable domain.
emailYesEmail authentication, read from the domain's own TXT records. DKIM is not reported: finding a DKIM key needs its selector, which cannot be discovered without guessing.
dnssecYesWhether the delegation is signed. This is not a validation of the DNSSEC chain.
domainYesThe hostname that was checked, in ASCII (punycode) form.
findingsYesWhat needs attention, worst first.
severityYesHow much attention this needs. "critical" means act now; "warning" means act this month; "unknown" means the check could not establish the fact, which is not the same as it being fine.
checkedAtYesWhen the check ran, ISO 8601 in UTC.
dnsResolvedYesWhether the DNS lookup answered at all. When false every field under "dns" is empty because nothing could be read, not because the domain has no records.
nameserversYesWhat the domain's own nameservers said when each was asked directly, over TCP port 53 with recursion off. This is the only way to see a nameserver left in the delegation after a migration, or a zone edited on one server and never transferred to the others; a recursive resolver answers with whatever one server told it and hides which.
registrationYesWhat the registry publishes about this registration.
unicodeDomainYesThe Unicode form when the domain is internationalised, otherwise null.
registrableDomainYesThe domain registration was checked against, e.g. "example.co.uk".

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), yet the description adds substantial operational context: RDAP as the data source, the fact that .de/.nl/.no/.au/.fi publish no expiry and say so rather than reporting a gap, warning at 30 days and critical at 7 days, and that findings are ordered worst-first. It also discloses the cost and mechanism of the nameserver check (one TCP query per nameserver, in parallel) and that it can be skipped without affecting the rest.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the capability, then usage, then exclusions, then caveats — a sensible order. It is on the long side (five paragraphs) with mild redundancy: 'Pass checkNameservers: false to skip it' appears both in the nameserver paragraph and again in the schema description, and the registrar-expiry warning thresholds could be tighter. Every paragraph still carries distinct information, so nothing is truly wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value formatting is correctly left out. Given that, the description covers everything an agent needs: what is queried, where the data comes from, known data gaps by TLD, threshold semantics for warnings, ordering of results, and the opt-out parameter. Nothing material to calling or interpreting the tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description earns extra credit by explaining why checkNameservers exists at all — 'a question no recursive resolver can answer, because it replies with whatever one server told it and does not say which' — and by describing the failure modes it detects (a stale server left in the delegation, a zone edited but never transferred). It does not restate the domain-normalisation rules that the schema already documents, which is appropriate restraint.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a precise verb+resource statement — reports registration expiry, registrar, and DNS configuration — and enumerates the concrete record types covered (NS, A, MX, TXT, CAA, DNSSEC, SPF, DMARC). It also explicitly distinguishes itself from the sibling tools uptime_check and ssl_check, so an agent can route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit trigger phrases ('is this domain about to lapse?', 'why is this client's email going to spam?') and an explicit negative boundary ('Do not use it to check whether a website responds — that is uptime_check — or to inspect an SSL certificate, which is ssl_check'). It even states a capability it deliberately omits (DKIM) and why, which prevents an agent from expecting it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

healthServer healthA
Read-onlyIdempotent

Confirms that the upkeep-mcp server is running and reports which version it is.

Use this to verify the connection after installing or reconfiguring the server, or when another upkeep-mcp tool behaves unexpectedly and you need to establish whether the server itself is reachable.

Do not use it to check whether a website is up — it says nothing about any domain or URL, only about this server process. Use uptime_check for websites.

Returns the server name and version, the Node.js version it runs under, how long the process has been alive, and the time of the check.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
nodeYesNode.js version the server runs under, e.g. "v20.11.0".
serverYesServer name, e.g. "upkeep-mcp".
statusYesAlways "ok". If the server can answer, it is healthy.
versionYesServer version, e.g. "0.1.0".
checkedAtYesWhen the report was produced, ISO 8601 with timezone.
uptimeSecondsYesWhole seconds since the server process started.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered; the description adds genuine extra context by bounding the tool's scope to the server process itself and enumerating what the check reflects. It stops short of anything richer (e.g. cost, latency, failure semantics), and the return-value sentence overlaps the output schema, so it is strong but not maximal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short paragraphs, each front-loaded with its point: purpose, when/when-not, then return shape. Every sentence carries information the agent needs, and the exclusions are placed before the alternative rather than buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter diagnostic tool with an output schema, annotations, and sibling tools that could be confused with it, the description supplies purpose, selection criteria, exclusions, and scope boundaries. Nothing an agent needs in order to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline for a no-parameter tool is 4. The description correctly implies no input is needed, but adds no parameter-level information beyond that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource — confirms the upkeep-mcp server is running and reports its version — and immediately scopes it against siblings by noting it says nothing about any domain or URL. An agent can distinguish it from uptime_check/domain_check without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states explicit triggering conditions (after installing or reconfiguring the server; when another upkeep-mcp tool behaves unexpectedly) and an explicit exclusion with the correct alternative named ("Do not use it to check whether a website is up ... Use uptime_check for websites"). This is textbook when/when-not/alternative guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

portfolio_reportWhole portfolio, ranked by urgencyA
Read-onlyIdempotent

Runs the maintenance checks across every site in a portfolio and returns one report ordered by what needs action first: what is down, what expires soonest, what regressed since the last run.

Use it to answer "what needs attention this week?", "is anything down?" or "what do I put in this quarter's report?" across a whole client list. It is the right call whenever the question is about more than one site — running domain_check, ssl_check and uptime_check once per site by hand is what this replaces.

The portfolio comes from the "sites" argument, or from a local JSON file ("sites.json" by default) whose format is documented in sites.example.json. Pass "checks" to override what each site asks for — ["uptime"] answers "is anything down right now?" in a fraction of the time — and "tags" to report on part of the portfolio.

Do not use it to answer a question about one site: domain_check, ssl_check, uptime_check and seo_audit answer those directly and in a fraction of the time. It runs checks and reports; it never changes a site, and it never writes to the portfolio file.

A site that cannot be checked is reported as a finding, never as a failure of the whole report. Only a portfolio that cannot be read at all is an error.

What changed since the previous run is compared against one recorded snapshot. By default that snapshot lives in memory and is lost when this server restarts; a portfolio file may add a "history" path, and then the snapshot is written there and comparisons survive a restart for up to 90 days. Nothing is written unless the portfolio asks for it. Either way the report says what it had to compare against, and why it had nothing when it had nothing, rather than implying nothing changed. Only sites that both runs measured the same way are compared, so a quick uptime-only pass never invents regressions in the run after it.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileNoPath to a portfolio JSON file, e.g. "/Users/you/sites.json". Give the full path: a relative one is resolved against the directory this server was started in, which for a desktop client is usually "/" and not the user's own. Defaults to "sites.json". Ignored when "sites" is given. The format is documented in sites.example.json; only .json files can be read.
tagsNoOnly report on sites carrying at least one of these tags, e.g. ["retainer"]. Matched case-insensitively. Omit to report on the whole portfolio.
sitesNoThe portfolio, passed inline. Use this for a one-off report over a handful of sites. When omitted, the portfolio is read from the local file instead.
checksNoRun exactly these checks for every site, ignoring what the file says. Use it for a quick pass — ["uptime"] answers "is anything down right now?" in a fraction of the time.

Output Schema

ParametersJSON Schema
NameRequiredDescription
fileYesThe file it was read from, when it was read from one.
notesYesAnything the report could not do, said plainly rather than left out.
sitesYesEvery site, most urgent first.
sourceYesWhere the portfolio came from.
changesYesWhat changed since the previous run in this server process.
summaryYesHow many sites landed at each severity.
severityYesHow much attention this needs. "critical" means act now; "warning" means act this month; "unknown" means the check could not establish the fact, which is not the same as it being fine.
siteCountYesHow many sites the report covers, after any tag filter.
generatedAtYesWhen the report was produced, ISO 8601 in UTC.
needsAttentionYesEverything worth acting on, worst first. This is the list the report exists to produce.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, yet the description adds substantial behavioral context: it never changes a site nor writes to the portfolio file, partial failures surface as findings rather than whole-report errors, and the snapshot mechanism (in-memory vs. a 'history' path surviving 90 days) is disclosed along with the caveat that only identically-measured sites are compared.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and routing, and every paragraph carries real information about behavior or arguments. It is on the long side and repeats 'in a fraction of the time' twice, but no sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a multi-parameter orchestration tool with an output schema (so return shape needs no restating), the description covers the remaining gaps an agent needs: alternate portfolio sources, check overriding, failure semantics, and snapshot/history persistence.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is 3, but the description adds genuine routing meaning: the 'sites' inline argument vs. the default sites.json file, 'checks' as a speed lever for a quick pass, and 'tags' for scoping to part of the portfolio.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with explicit scope: 'Runs the maintenance checks across every site in a portfolio and returns one report ordered by what needs action first.' It further distinguishes itself from the single-site siblings by naming them, so an agent can tell it apart without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use triggers ('what needs attention this week?', 'is anything down?'), names the fleet of alternatives (domain_check, ssl_check, uptime_check, seo_audit) and the condition that selects them, and adds an explicit when-not: 'Do not use it to answer a question about one site.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

seo_auditTechnical SEO of one pageA
Read-onlyIdempotent

Reads one page and reports the technical SEO facts a maintenance retainer is judged on: title and meta description, heading structure, canonical, Open Graph, hreflang, language, viewport, images with no alt text, the state of robots.txt and the sitemap, and which internal links are broken. The sitemap is checked against the rules of the sitemaps protocol, so a file that exists but that consumers drop is reported as broken rather than as present.

Use it to answer "why is this page not being indexed?", "does this page have the metadata it needs?" or "are there broken links on the homepage?". It is the check to run before a quarterly report, and after a site rebuild.

Do not use it to check whether a site is up, which is uptime_check, or to inspect a certificate, which is ssl_check. It audits one page: it does not crawl a site, and it judges only what is in the HTML, never how a page ranks.

robots.txt is read first and obeyed. A page this crawler is not allowed to read is reported as such and is not requested — and neither are internal links it forbids. An unreadable robots.txt is treated as forbidding everything, per RFC 9309.

Link checking is one HTTP request per link, paced politely, so an audit with links takes seconds rather than milliseconds; pass checkLinks: false when speed matters more. Returns findings ordered by urgency, worst first.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe page to audit, e.g. "https://example.com/pricing". A bare domain such as "example.com" is accepted and read over HTTPS. One page is audited, not a whole site, so pass the page you care about; the homepage is the usual choice.
maxLinksNoHow many internal links to check at most. Defaults to 25. Links beyond the limit are counted and reported as unchecked, never silently ignored.
checkLinksNoWhether to request each internal link to find broken ones. Defaults to true. Set it to false for a fast metadata-only audit: link checking is one request per link and is what makes this tool take seconds rather than milliseconds.

Output Schema

ParametersJSON Schema
NameRequiredDescription
urlYesThe URL that was requested.
pageYesWhat the document itself says.
linksYesInternal link health. Only internal links are requested; external ones are counted.
robotsYesWhat the host publishes about crawling, and what it means for this audit.
statusYesHTTP status of the page.
fetchedYesWhether the page was read at all. False when robots.txt forbids it.
sitemapYesThe sitemap declared in robots.txt, or /sitemap.xml when none is declared.
finalUrlYesWhere it ended up after redirects.
findingsYesWhat needs attention, worst first.
severityYesHow much attention this needs. "critical" means act now; "warning" means act this month; "unknown" means the check could not establish the fact, which is not the same as it being fine.
checkedAtYesWhen the audit ran, ISO 8601 in UTC.
truncatedYesWhether the document was longer than the read limit and was cut off.
contentTypeYesContent type, without parameters, e.g. "text/html".

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), and the description adds real behavioral context: robots.txt is read first and obeyed, an unreadable robots.txt is treated as forbidding everything per RFC 9309, forbidden links are not requested, sitemap validity is judged by protocol rules rather than mere existence, and results are ordered worst-first. It even explains the seconds-vs-milliseconds latency cost of link checking.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the resource and reported facts, then usage, then exclusions, then behavioral caveats. Dense and almost every sentence carries information, though the length sits at the upper end for a single-page audit tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return structure needn't be described, and the description still covers everything an agent needs to decide and call correctly: scope, exclusions, robots/RFC behavior, latency tradeoff, and result ordering.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds decision value beyond the schema: it explains why checkLinks costs time (one HTTP request per link, paced politely) and when to set it false, and it reinforces maxLinks' semantics (unchecked links are counted, never silently ignored).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (reads) and resource (one page) and enumerates exactly what is reported: title/meta, headings, canonical, OG, hreflang, viewport, alt-less images, robots.txt, sitemap, broken internal links. It is unmistakable against siblings like site_crawl (one page, not a whole site) and uptime_check.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete questions it answers, a lifecycle cue (before a quarterly report, after a rebuild), and explicit exclusions naming the alternatives uptime_check and ssl_check. It also states the negative scope: it does not crawl a site and does not judge rankings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

site_crawlTechnical SEO across a siteA
Read-onlyIdempotent

Walks a site from a starting page and reports what only a crawl can see: pages sharing a title or a meta description, internal links that are broken and which page links to them, pages asking not to be indexed, and how much of the site was reachable at all.

Use it after a site rebuild or a migration, and before a quarterly report — the findings here are the ones a client never notices and a search engine always does. Use it to answer "why are these pages competing with each other?", "are there broken links", anywhere on the site?" or "is anything still marked noindex from staging?".

Do not use it to audit one page in depth — that is seo_audit, which reports canonical, Open Graph, hreflang, images without alt text and the sitemap for a single URL. Do not use it to check whether a site is up (uptime_check) or to inspect a certificate (ssl_check). It judges only what is in the HTML, never how a page ranks.

It stays on one origin: https://example.com and https://www.example.com are different origins and only the one you start from is visited.

robots.txt is read first and obeyed for every URL before it is requested. An unreadable robots.txt is treated as forbidding everything, per RFC 9309, and the crawl reports that rather than proceeding.

It costs one request per page, paced at one every half second per host, so 25 pages is about fifteen seconds. Three budgets bound it — pages, depth, and a two-minute deadline — and whichever one stopped the crawl is reported along with how many URLs were left unvisited, because a report that does not say it saw a quarter of the site is worse than no report.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe page to start from, e.g. "https://example.com/" — a full URL with scheme. The crawl stays on this URL's origin, so start at the address the site actually serves: https://example.com and https://www.example.com are different origins and only one of them will be visited.
maxDepthNoHow many links deep from the starting page to follow. Defaults to 3. 0 fetches only the starting page. Three reaches everything a small business site has; past that the page budget is the real bound anyway. Example: 2
maxPagesNoHow many pages to fetch. Defaults to 25, at most 100. Every page is one request paced at one every half second, so this is the cost: 25 pages is about fifteen seconds, 100 about a minute. Pages found and not visited are counted and reported. Example: 50

Output Schema

ParametersJSON Schema
NameRequiredDescription
crawlYesWhat the crawl covered, and what it did not.
pagesYesEvery page fetched, in the order they were visited.
brokenYesInternal links that did not answer with a usable status, with the page each was found on — which is the half of a broken link report that makes it fixable.
originYesThe origin it stayed on, e.g. "https://example.com".
findingsYesWhat needs attention, worst first.
severityYesHow much attention this needs. "critical" means act now; "warning" means act this month; "unknown" means the check could not establish the fact, which is not the same as it being fine.
startUrlYesThe URL the crawl was asked to start from.
checkedAtYesWhen the crawl ran, ISO 8601 in UTC.
duplicatesYesWhat only a crawl can see. One page cannot tell you that its title is the same as forty others', and that is the commonest reason a site's pages compete with each other in search results.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover the read-only/idempotent safety profile, and the description goes far beyond that: origin confinement, robots.txt read first and obeyed per RFC 9309 with fail-closed behavior on unreadable files, one request per page paced at one per half second, three crawl budgets (pages, depth, two-minute deadline), and explicit reporting of which budget stopped the crawl and how many URLs were left unvisited.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then usage, then exclusions, then operational caveats — a sensible order with no redundancy. It runs long for a tool description and includes atmospheric lines ('the findings a client never notices and a search engine always does') that justify themselves as usage motivation but are not strictly load-bearing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation, and the description fills every remaining gap: scope limits (one origin), legal/ethical behavior (robots.txt), cost and time expectations, and what happens when the crawl is truncated. Nothing an agent needs to decide to call it, or to interpret a partial result, is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is already 100%, so the schema carries the definitions. The description nonetheless adds framing the schema does not: that pages, depth and a deadline are the three interacting budgets, and that the page count is the true cost driver because each page is one paced request. It does slightly restate the origin example already in the url field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('walks a site from a starting page') and enumerates exactly what it reports: duplicate titles/meta descriptions, broken internal links with their referrers, noindex pages, and reachability. It actively distinguishes itself from seo_audit, uptime_check and ssl_check by naming what each of those does instead.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit triggers ('after a site rebuild or a migration, and before a quarterly report') plus concrete example questions the tool answers. It also supplies explicit exclusions with the alternative named for each: single-page depth = seo_audit, is-the-site-up = uptime_check, certificate = ssl_check.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ssl_checkSSL certificateA
Read-onlyIdempotent

Inspects the TLS certificate a host actually serves: when it expires, who issued it, whether the chain verifies, which hostnames it covers, and which TLS version was negotiated.

Use it to answer "when does this certificate need renewing?", "why does the browser warn about this site?" or "does the certificate cover www as well as the bare domain?" — the last being one of the most common real-world misconfigurations, along with a missing intermediate certificate, which is reported as UNABLE_TO_VERIFY_LEAF_SIGNATURE.

Do not use it for domain registration expiry, which is a different date entirely — that is domain_check. It connects to the host but does not request a page; use uptime_check for that.

Certificates that are expired, self-signed or untrusted are inspected and reported rather than refused. Revocation is checked over OCSP, preferring the response a server staples to the handshake and otherwise asking the issuing authority directly; the answer is only believed once its signature verifies against that authority. Many healthy certificates cannot be checked at all, because since 2025 the two largest issuers publish no OCSP responder and distribute revocation by CRL instead — that is reported as an unavailable reason rather than as a problem with the site, and produces no finding. Certificate revocation lists are not downloaded.

A certificate is reported as a warning inside 14 days and as critical inside seven. That window is deliberately shorter than the one domain_check uses for registrations: ACME clients renew with 30 days left, so 28 days remaining is a healthy site in the middle of a normal renewal, not a problem. Returns findings ordered by urgency, worst first.

ParametersJSON Schema
NameRequiredDescriptionDefault
portNoTCP port to connect to. Defaults to 443. Use 8443, 993 and so on for other services.
domainYesThe host whose certificate should be inspected, e.g. "example.com" — no scheme, no trailing slash. A full URL is also accepted and reduced to its hostname. The certificate is read for exactly this host, so "www.example.com" and "example.com" are different checks.

Output Schema

ParametersJSON Schema
NameRequiredDescription
tlsYesWhat the handshake negotiated.
hostYesThe host that was contacted, in ASCII (punycode) form.
portYesThe port that was contacted.
chainYesThe certificate chain and whether it verified.
issuerYesIssuer common name, e.g. "R11" or "GTS CA 1P5".
subjectYesSubject common name.
coverageYesWhich hostnames this certificate is valid for.
findingsYesWhat needs attention, worst first.
issuedAtYesStart of the certificate validity, ISO 8601 UTC.
severityYesHow much attention this needs. "critical" means act now; "warning" means act this month; "unknown" means the check could not establish the fact, which is not the same as it being fine.
checkedAtYesWhen the check ran, ISO 8601 in UTC.
expiresAtYesWhen the certificate expires, ISO 8601 UTC.
revocationYesWhether the certificate has been revoked, and how confidently that was established.
serialNumberYesCertificate serial number, hexadecimal.
daysUntilExpiryYesWhole days until the certificate expires, negative if already expired. Floored, so a certificate expiring in a few hours reads as 0.
fingerprintSha256YesSHA-256 fingerprint of the certificate.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover the safety profile (readOnly/idempotent/non-destructive); the description goes well beyond them. It discloses that expired/self-signed/untrusted certs are inspected rather than refused, how OCSP revocation is verified, that CRL lists are not downloaded, how unverifiable certs are reported with no finding, and the exact 14-day/7-day severity windows.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose before routing guidance, and the sequencing (what it does → when to use → when not to use → behavior) is clean. It is longer than strictly needed for tool selection — the OCSP/CRL paragraph is dense — but every claim defends against a plausible misinterpretation, so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description fills the remaining gaps: severity thresholds, ordering (worst first), and the deliberate divergence from domain_check's renewal window. Nothing an agent needs to call or interpret this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already fully documented in the schema. The description adds no port guidance and only echoes the schema's note that www and bare domain are distinct checks. Baseline 3 is appropriate when the schema carries parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Inspects the TLS certificate a host actually serves') and enumerates exactly what it reports: expiry, issuer, chain verification, covered hostnames, negotiated TLS version. It explicitly distinguishes itself from sibling tools domain_check and uptime_check, so an agent can route correctly without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete use-case phrasings ('when does this certificate need renewing?'), names the two most likely alternative tools with the condition that selects them (domain_check for registration expiry, uptime_check for page requests), and explicitly states 'Do not use it for...'. Both when-to-use and when-not-to-use are covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

uptime_checkSite reachability and headersA
Read-onlyIdempotent

Requests a page and reports whether it answered, how long it took, the full redirect chain it went through, whether plain HTTP is upgraded to HTTPS, and which security headers came back.

Use it to answer "is this site up?", "why does this URL take four redirects to load?", "does http:// still work and should it?" or "does this site send HSTS?". It is the check to run when a client reports that a page is down or slow.

Do not use it to inspect a certificate — that is ssl_check — or for registration and DNS, which is domain_check. It fetches one page, not a whole site: use seo_audit for crawling.

Redirects are followed one hop at a time so the whole chain is visible, up to ten hops. Response time is wall clock to the first response headers and includes DNS, TCP and TLS setup, so it is not a measure of server processing time.

Status codes are graded rather than lumped together: 5xx, 404 and 410 are critical, 401 and 403 are a warning because they are normal for a staging site, and 429 is reported as unknown because a throttled check establishes nothing about real availability. Returns findings ordered by urgency, worst first.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe page to request, e.g. "https://example.com". A bare domain such as "example.com" is accepted and tried over HTTPS. Include the path when a specific page matters; the homepage is checked otherwise.

Output Schema

ParametersJSON Schema
NameRequiredDescription
urlYesThe URL that was requested first.
hstsYesThe HSTS policy, read from the HTTPS response only.
httpsYesWhether visitors arriving over plain HTTP are moved to HTTPS.
statusYesStatus code at the end of the redirect chain.
finalUrlYesWhere the redirect chain ended.
findingsYesWhat needs attention, worst first.
severityYesHow much attention this needs. "critical" means act now; "warning" means act this month; "unknown" means the check could not establish the fact, which is not the same as it being fine.
checkedAtYesWhen the check ran, ISO 8601 in UTC.
reachableYesWhether the server answered at all.
redirectsYesThe redirect chain, followed manually one hop at a time.
responseTimeMsYesWall-clock milliseconds to the first response headers. Includes DNS, TCP and TLS setup, so it is not server processing time.
securityHeadersYesSecurity-relevant response headers.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover the safety profile (readOnly, idempotent, non-destructive), so the description carries the real behavioral burden and does so richly: redirects followed one hop at a time up to ten hops, timing measured as wall clock to first response headers inclusive of DNS/TCP/TLS, and status-code grading with the rationale for 429 being 'unknown'. This is well beyond what structured fields disclose.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then usage, then exclusions, then behavioral detail. Despite its length, each paragraph maps to a distinct decision the caller must make, and nothing is redundant padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be restated. Combined with explicit routing to siblings, coverage of redirect/timing semantics, and status-code grading, an agent has everything needed to invoke this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single url parameter is fully documented in the schema (bare domain, path handling, homepage fallback). The description adds only the scoping nuance that it fetches one page rather than a whole site, so the baseline 3 applies with little added meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('requests a page') and enumerates exactly what it reports: reachability, timing, redirect chain, HTTP→HTTPS upgrade, and security headers. It explicitly distinguishes itself from siblings ssl_check, domain_check, and seo_audit, so an agent can route without opening another schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete triggering questions ('is this site up?', 'why does this URL take four redirects?', 'does this site send HSTS?') plus the operational context (run it when a client reports a page is down or slow). It also states explicit exclusions: not for certificates (ssl_check), not for DNS/registration (domain_check), not for whole-site crawling (seo_audit).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 8 tool updatesv0.6.1
    • First observedaccessibility_audit
    • First observeddomain_check
    • First observedhealth
    • First observedportfolio_report
    • First observedseo_audit
    • First observedsite_crawl
    • First observedssl_check
    • First observeduptime_check

TDQS

A4.6/5.0

Scored across 8 tools

Disambiguation5/5

Every tool targets a distinct maintenance concern (TLS, DNS/registration, uptime, one-page SEO, site crawl, accessibility, portfolio aggregation, server health) and the descriptions explicitly cross-reference each other with 'do not use for' guidance. There is no meaningful overlap between tools, so an agent can select correctly.

Naming Consistency4/5

Names follow a mostly consistent snake_case noun_action pattern (ssl_check, domain_check, uptime_check, seo_audit, site_crawl, accessibility_audit, portfolio_report). The single 'health' tool is a minor deviation, but all names remain readable and predictable.

Tool Count5/5

Eight tools is well-scoped for a website maintenance server: each tool handles a distinct class of check and the portfolio tool aggregates them without redundancy. No tool feels thin or gratuitous.

Completeness4/5

The surface covers core maintenance needs across certificates, domains, uptime, SEO, crawling, accessibility, and portfolio reporting with no obvious dead ends. Minor gaps exist for site-wide accessibility auditing and performance/Core Web Vitals, and DKIM is intentionally omitted, which agents can work around.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    Not graded
    maintenance
    Provides comprehensive website validation across performance, accessibility, SEO, and security dimensions using multiple testing services including WebPageTest, Google PageSpeed Insights, Axe DevTools, Mozilla Observatory, and SSL Labs. Enables automated website health assessments through browser automation and API integrations.
    12
    3 npm
    1
    -
  • F
    license
    Not graded
    quality
    B
    maintenance
    Passive website security and trust auditor that checks for security, SEO, AI surface, email, and other exposures, producing a score and remediation plan.
    -
  • A
    license
    A
    quality
    D
    maintenance
    Performs comprehensive website health audits including SSL, DNS, email authentication, performance, uptime, and broken link checks, all without requiring API keys. Returns a scored report with weighted metrics and actionable recommendations.
    7
    97 npm
    MIT