Skip to main content
Glama
cyanheads

devops-status-mcp-server

by cyanheads
README.md
<div align="center">
  <h1>@cyanheads/devops-status-mcp-server</h1>
  <p><b>Check vendor status pages, inspect SSL/TLS certificates, verify DNS propagation, and get incident-response playbooks via MCP. STDIO or Streamable HTTP.</b>
  <div>7 Tools • 1 Resource</div>
  </p>
</div>

<div align="center">

[![Version](https://img.shields.io/badge/Version-0.9.0-blue.svg?style=flat-square)](./CHANGELOG.md) [![License](https://img.shields.io/badge/License-Apache%202.0-orange.svg?style=flat-square)](./LICENSE) [![Docker](https://img.shields.io/badge/Docker-ghcr.io-2496ED?style=flat-square&logo=docker&logoColor=white)](https://github.com/users/cyanheads/packages/container/package/devops-status-mcp-server) [![MCP SDK](https://img.shields.io/badge/MCP%20SDK-^2.0.0-green.svg?style=flat-square)](https://modelcontextprotocol.io/) [![npm](https://img.shields.io/npm/v/@cyanheads/devops-status-mcp-server?style=flat-square&logo=npm&logoColor=white)](https://www.npmjs.com/package/@cyanheads/devops-status-mcp-server) [![TypeScript](https://img.shields.io/badge/TypeScript-^7.0.2-3178C6.svg?style=flat-square)](https://www.typescriptlang.org/) [![Bun](https://img.shields.io/badge/Bun-v1.4.0-blueviolet.svg?style=flat-square)](https://bun.sh/)

</div>

<div align="center">

[![Install in Claude Desktop](https://img.shields.io/badge/Install_in-Claude_Desktop-D97757?style=for-the-badge&logo=anthropic&logoColor=white)](https://github.com/cyanheads/devops-status-mcp-server/releases/latest/download/devops-status-mcp-server.mcpb) [![Install in Cursor](https://cursor.com/deeplink/mcp-install-dark.svg)](https://cursor.com/en/install-mcp?name=devops-status-mcp-server&config=eyJjb21tYW5kIjoibnB4IiwiYXJncyI6WyIteSIsIkBjeWFuaGVhZHMvZGV2b3BzLXN0YXR1cy1tY3Atc2VydmVyIl19) [![Install in VS Code](https://img.shields.io/badge/VS_Code-Install_Server-0098FF?style=for-the-badge&logo=visualstudiocode&logoColor=white)](https://vscode.dev/redirect?url=vscode:mcp/install?%7B%22name%22%3A%22devops-status-mcp-server%22%2C%22command%22%3A%22npx%22%2C%22args%22%3A%5B%22-y%22%2C%22%40cyanheads%2Fdevops-status-mcp-server%22%5D%7D)

[![Framework](https://img.shields.io/badge/Built%20on-@cyanheads/mcp--ts--core-67E8F9?style=flat-square)](https://www.npmjs.com/package/@cyanheads/mcp-ts-core)

</div>

<div align="center">

**Public Hosted Server:** [https://devops-status.caseyjhand.com/mcp](https://devops-status.caseyjhand.com/mcp)

</div>

---

## Overview

Vendor status pages, SSL/TLS certificates, and DNS propagation — normalized across Atlassian Statuspage, Status.io, Slack, AWS Health, Google Cloud Service Health, Azure status, and Firehydrant backends, plus direct TLS/DNS checks for any domain. List and check 52 built-in vendors, fetch incident timelines, watch a persisted stack, and get a tailored incident-response playbook, all without API keys. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.

### Tools

| Tool | Description |
|:-----|:------------|
| `devops_list_vendors` | List vendors in the built-in registry, optionally filtered by name or category. Returns slug, display name, category, and status page URL. |
| `devops_status_check` | Check the current health status for one or more vendors. Returns per-vendor indicator (`none` / `minor` / `major` / `critical` / `maintenance`), degraded components, and active incident summaries. |
| `devops_get_incidents` | Fetch incident history for a vendor — active, resolved, or scheduled maintenance. Returns the full incident timeline with per-update bodies and affected components. |
| `devops_watch_stack` | Check the health of a named vendor stack persisted in session state. Pass `vendors` once to save the list; subsequent calls reuse it. Returns an aggregate health rollup plus per-vendor detail. |
| `devops_check_certs` | Inspect SSL/TLS certificate health for one or more domains via a real TLS handshake. Reports expiry, chain depth, protocol version, cipher suite, and HSTS presence. Pure TypeScript — no external API. |
| `devops_check_dns` | Resolve DNS records and verify propagation for one or more domains across Google (8.8.8.8), Cloudflare (1.1.1.1), and Quad9 (9.9.9.9). Reports per-resolver latency and resolver discrepancies. Pure TypeScript — no external API. |
| `devops_suggest_action` | Instruction tool — returns a tailored incident-response playbook and pre-filled follow-up tool calls given a vendor name and optional incident context. No external calls; fully deterministic. |

### Resources

| Resource | Description |
|:-----|:------------|
| `devops-status://vendors/{name}` | Full registry entry for a vendor by slug — status page URL, category, and API type. |

All resource data is also reachable via tools. Tool-only agents are fully supported.

---

## Capability reference

### `devops_list_vendors` <sub>tool</sub>

- Accepts an optional free-text `query` (matches name and slug, case-insensitive) and an optional `category` filter — eight categories: `cloud`, `cdn-edge`, `dev-platform`, `data`, `comms`, `auth`, `monitoring`, `ai`
- Returns slug (what to pass to other tools), display name, category, and status page URL
- 52 built-in entries, most on Atlassian Statuspage; `aws`, `gcp`, `azure`, `gitlab`, `neon`, `slack`, and `redis-cloud` route through native-API adapters normalized to the same shape
- `azure` reads Microsoft's Azure status RSS feed, which carries no severity or lifecycle: each posted item is an open `minor` incident with its services and regions as affected components, and the feed is empty while nothing is posted
- Statuspage-compatible pages not in the registry are still reachable by passing a raw base URL to other tools

Built-in vendor registry:

| Category | Vendors |
|:---------|:--------|
| `cloud` | digitalocean, linode, aws, gcp, azure |
| `cdn-edge` | cloudflare, akamai |
| `dev-platform` | gitlab, github, npm, vercel, netlify, render, fly-io, circleci, travis-ci, snyk, atlassian, figma, launchdarkly |
| `data` | mongodb-atlas, planetscale, supabase, neon, redis-cloud, elastic, influxdb, upstash, cloudinary, segment |
| `comms` | slack, discord, twilio, sendgrid, mailgun, hubspot, brevo, courier, loops |
| `auth` | auth0, clerk, workos |
| `monitoring` | datadog, sentry, new-relic, grafana-cloud, honeycomb |
| `ai` | openai, anthropic, elevenlabs, pinecone, cohere |

---

### `devops_status_check` <sub>tool</sub>

- Accepts registered vendor slugs (e.g., `github`, `aws`) or raw Atlassian Statuspage base URLs, mixed freely — up to 20 per call
- `mode: "summary"` (default): indicator + degraded components + active incidents; `mode: "detailed"` adds the full component list (capped at `component_limit`, default 50, max 500) and scheduled maintenance windows
- `Promise.allSettled` fan-out — one failing vendor never blocks the rest; failures surface as a per-vendor `error` field
- Results served from a 60-second in-memory cache; `cached: true` on each result
- `summary` partitions the batch into `operational` / `degraded` / `down` / `maintenance` / `unavailable` counts
- `nextToolSuggestions` carries one pre-filled `devops_suggest_action` call per vendor with an active problem — indicator `minor`/`major`/`critical`, or an open incident of that impact — with the vendor slug (or normalized URL), indicator, affected components, and latest incident title filled in; a vendor that could not be checked never gets one, so the list is empty when no checked vendor has a problem

---

### `devops_get_incidents` <sub>tool</sub>

- `filter`: `all` (default, incidents + scheduled maintenances), `active` (investigating/identified/monitoring), `resolved` (fully resolved), or `scheduled` (maintenance windows only)
- Returns per-update bodies in chronological order, affected component names, duration in minutes for resolved incidents, and a direct shortlink to the incident page
- `limit` (1–50) with `offset` for paging; a truncated result discloses the total and the next `offset` to fetch
- Some vendor feeds cap their own history (`upstreamCeiling`). Atlassian Statuspage's API stops at the newest 50 incidents, which on a busy page is about a week
- `since` (`YYYY-MM-DD`, up to 24 months back, with `filter: "all"` or `"resolved"`) leaves out incidents that started before that date. On Statuspage vendors it also reads the status page's quarterly history archive back to that date, reaching past the 50-record ceiling
- Every incident carries `source`: `api` for the status API, or `history` for an archive record, which has the title, impact, start and end times, and final update message but no components or update timeline. When a record is in both, the API version is kept
- If the archive can't be read (the page doesn't publish one, it times out, or it comes back in an unexpected shape), the call still returns the status API result, with a `notice` saying how far history reached
- AWS keeps a resolved event listed for hours after it ends, so `filter: "resolved"` returns only those still listed; Azure's feed lists open items only, so `filter: "resolved"` is always empty for it; AWS, Azure, Google Cloud, and Slack publish no maintenance windows, so `filter: "scheduled"` is always empty for them

---

### `devops_watch_stack` <sub>tool</sub>

- On the first call, provide `vendors` to define the stack — it is saved to tenant-scoped session state under `stack_name`
- Subsequent calls can omit `vendors`; the saved list is reused automatically
- Multiple stacks coexist via distinct `stack_name` values (e.g., `"production"`, `"data-layer"`) — letters, digits, hyphens, and underscores, optionally separated by single dots or slashes, 1-64 characters
- Aggregate `health` rollup: `all_operational` / `maintenance` (a vendor in a scheduled window, nothing worse open) / `degraded` / `partial_outage` / `major_outage` / `unknown` (a vendor could not be reached) — never `all_operational` when any vendor errored or is in a window
- `nextToolSuggestions` pre-fills a `devops_suggest_action` call for each vendor with an active problem, same as `devops_status_check`
- Note: stack state is in-memory; it does not persist across server restarts

---

### `devops_check_certs` <sub>tool</sub>

- Accepts bare hostnames (no `https://` prefix) — up to 10 per call
- Reports: days to expiry (flagged `warning` at < 30 days, `critical` at < 7), certificate subject and SANs, issuer common name, chain depth, negotiated TLS version (flags 1.0 and 1.1 as insecure), cipher suite
- HSTS detection: sends a minimal HTTP/1.1 GET over the same TLS socket, reads the `Strict-Transport-Security` response header
- `status: "critical"` distinguishes a hostname mismatch (`hostname_verification_error`) from an untrusted chain (`authorization_error`) — both would be rejected by ordinary clients
- Per-domain failures are reported inline (`status: "error"`, reason in `error`) rather than throwing — useful partial results when checking multiple domains. `flags` is empty unless a handshake completed; a server that completes one without presenting a certificate still reports its TLS session
- Configurable port (default 443) and timeout per domain

---

### `devops_check_dns` <sub>tool</sub>

- Queries Google (8.8.8.8), Cloudflare (1.1.1.1), and Quad9 (9.9.9.9) in parallel per domain
- Supported record types: A, AAAA, CNAME, MX, TXT, NS (defaults to A, AAAA, MX, TXT)
- Reports per-resolver latency, propagation discrepancies, and human-readable flags
- Discrepancies are typed: `partial_resolution` (some resolvers answered, others didn't) signals a real problem; `value_variation` (all answered, different values) is normal for anycast/geo-steered domains
- Custom resolver list supported — pass public resolver IP literals to test resolver-specific behavior (private and loopback resolvers need `DEVOPS_STATUS_ALLOW_PRIVATE_TARGETS=true`); an empty `resolvers` or `record_types` array uses the defaults
- Up to 10 domains per call; per-domain timeouts configurable

---

### `devops_suggest_action` <sub>tool</sub>

- Category-tailored markdown playbook (cloud, CDN, dev-platform, data, comms, auth, monitoring, AI); falls back to generic guidance for unrecognized vendors
- Accepts a vendor slug or display name, resolved to the canonical slug so pre-filled follow-up arguments stay valid
- Optional `incident_summary` / `affected_components` prepend a targeted subsystem section (e.g. GitHub `Actions` → CI/CD steps, Cloudflare `DNS` → DNS/TTL guidance); optional `vendor_indicator` leads with severity-tailored urgency framing
- `nextToolSuggestions` pre-fills follow-up tool calls, including cert/DNS checks when `your_domain` is given — execute in sequence
- When `DEVOPS_STATUS_DISABLE_ACTIVE_PROBES=true`, guidance swaps the unregistered probe tools for equivalent manual commands (`dig`, `openssl s_client`)

---

### `devops-status://vendors/{name}` <sub>resource</sub>

- Returns the full registry entry for a vendor slug — status page URL, category, and API type (`statuspage`, `statusio`, `slack`, `aws`, `gcp`, `azure`, `firehydrant`)
- Cached publicly for 1 hour — the registry is compiled in and identical for every caller
- Same data is reachable via `devops_list_vendors` — tool-only agents are fully supported

---

## Features

Built on [`@cyanheads/mcp-ts-core`](https://github.com/cyanheads/mcp-ts-core): stdio and Streamable HTTP transports, pluggable auth (`none` / `jwt` / `oauth`), swappable storage (`in-memory`, `filesystem`, `Supabase`, `Cloudflare KV/R2/D1`), structured logging with optional OpenTelemetry tracing.

DevOps-status-specific:

- **No API keys required** — every status backend is a public API; TLS and DNS use Node.js stdlib (`node:tls`, `node:dns`)
- 52-vendor built-in registry covering cloud, CDN, dev-platform, data, comms, auth, monitoring, and AI categories; adapter layer normalizes Status.io, Slack, AWS Health, Google Cloud Service Health, Azure status, and Firehydrant backends into the Statuspage shapes; extendable via raw Statuspage URL passthrough
- 60-second in-memory cache on status reads shared across all tenants — prevents thundering-herd on batch calls
- `devops_watch_stack` persists named vendor lists in tenant-scoped state for repeat morning checks or pre-deploy sweeps
- `devops_suggest_action` dispatches category-specific playbooks deterministically — no LLM sampling dependency, works in all clients

Agent-friendly output:

- Batch tools (`devops_status_check`, `devops_watch_stack`, `devops_check_certs`, `devops_check_dns`) use `Promise.allSettled` — one failing target never blocks the rest; errors surface as inline `error` fields
- `cached: true` / `checked_at` on every status result — agents know when data was fetched
- Discriminated indicator and status enums (`none` / `minor` / `major` / `critical` / `maintenance`; `operational` / `degraded_performance` / `partial_outage` / `major_outage` / `under_maintenance`) — callers branch on data, not string parsing
- `nextToolSuggestions` in `devops_suggest_action` pre-fills tool arguments from incident context — agents can execute the playbook mechanically

---

## Getting started

### Public Hosted Instance

A public instance is available at `https://devops-status.caseyjhand.com/mcp` — no installation required. Point any MCP client at it via Streamable HTTP:

```json
{
  "mcpServers": {
    "devops-status-mcp-server": {
      "type": "streamable-http",
      "url": "https://devops-status.caseyjhand.com/mcp"
    }
  }
}
```

### Self-Hosted / Local

No API key required. Add the following to your MCP client configuration file:

```json
{
  "mcpServers": {
    "devops-status-mcp-server": {
      "type": "stdio",
      "command": "bunx",
      "args": ["@cyanheads/devops-status-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info"
      }
    }
  }
}
```

Or with npx (no Bun required):

```json
{
  "mcpServers": {
    "devops-status-mcp-server": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@cyanheads/devops-status-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info"
      }
    }
  }
}
```

Or with Docker:

```json
{
  "mcpServers": {
    "devops-status-mcp-server": {
      "type": "stdio",
      "command": "docker",
      "args": [
        "run", "-i", "--rm",
        "-e", "MCP_TRANSPORT_TYPE=stdio",
        "ghcr.io/cyanheads/devops-status-mcp-server:latest"
      ]
    }
  }
}
```

For Streamable HTTP, set the transport and start the server:

```sh
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcp
```

### Prerequisites

- [Bun v1.4.0](https://bun.sh/) or higher (or Node.js v24+).
- No API keys or external accounts required.

### Installation

1. **Clone the repository:**

```sh
git clone https://github.com/cyanheads/devops-status-mcp-server.git
```

2. **Navigate into the directory:**

```sh
cd devops-status-mcp-server
```

3. **Install dependencies:**

```sh
bun install
```

4. **Configure environment:**

```sh
cp .env.example .env
# edit .env if you want to override defaults
```

---

## Configuration

No API keys required. All environment variables are optional.

| Variable | Description | Default |
|:---------|:------------|:--------|
| `DEVOPS_STATUS_CACHE_TTL_MS` | In-memory cache TTL for vendor status reads (all backends) in milliseconds. | `60000` |
| `DEVOPS_STATUS_FETCH_TIMEOUT_MS` | Per-request timeout for vendor status API calls (all backends) in milliseconds. | `8000` |
| `DEVOPS_STATUS_CERT_TIMEOUT_MS` | Default `timeout_ms` for `devops_check_certs` (per-domain TLS handshake, milliseconds). A caller-passed `timeout_ms` overrides it. | `5000` |
| `DEVOPS_STATUS_DNS_TIMEOUT_MS` | Default `timeout_ms` for `devops_check_dns` (per domain+resolver query, milliseconds). A caller-passed `timeout_ms` overrides it. | `3000` |
| `DEVOPS_STATUS_ALLOW_PRIVATE_TARGETS` | When `true`, disables SSRF guards for user-supplied URLs and domains. For trusted local/intranet deployments only. | `false` |
| `DEVOPS_STATUS_DISABLE_ACTIVE_PROBES` | When `true`, omits the arbitrary-target probe tools (`devops_check_dns`, `devops_check_certs`) from the registered tool surface; the five vendor-registry/incident tools remain. For shared/public multi-tenant instances. | `false` |
| `MCP_TRANSPORT_TYPE` | Transport: `stdio` or `http`. | `stdio` |
| `MCP_HTTP_PORT` | Port for HTTP server. | `3010` |
| `MCP_SESSION_MODE` | HTTP session handling: `stateless`, `stateful`, or `auto`. The server declares `stateless` in source — it holds no per-session state — so setting this is only needed to override that. | `stateless` |
| `MCP_AUTH_MODE` | Auth mode: `none`, `jwt`, or `oauth`. | `none` |
| `MCP_LOG_LEVEL` | Log level (RFC 5424). | `info` |
| `LOGS_DIR` | Directory for log files (Node.js only). | `<project-root>/logs` |
| `OTEL_ENABLED` | Enable [OpenTelemetry instrumentation](https://github.com/cyanheads/mcp-ts-core/tree/main/docs/telemetry). | `false` |

See [`.env.example`](./.env.example) for the full list of optional overrides.

---

## Running the server

### Local development

- **Build and run:**

  ```sh
  bun run rebuild
  bun run start:stdio
  # or
  bun run start:http
  ```

- **Run checks and tests:**

  ```sh
  bun run devcheck   # Lint, format, typecheck, security
  bun run test       # Vitest test suite
  bun run lint:mcp   # Validate MCP definitions against spec
  ```

### Docker

```sh
docker build -t devops-status-mcp-server .
docker run --rm -p 3010:3010 devops-status-mcp-server
```

The Dockerfile defaults to HTTP transport, stateless session mode, and logs to `/var/log/devops-status-mcp-server`. OpenTelemetry peer dependencies are installed by default — build with `--build-arg OTEL_ENABLED=false` to omit them.

---

## Project structure

| Path | Purpose |
|:-----|:--------|
| `src/index.ts` | `createApp()` entry point — registers tools, resources, and inits services. |
| `src/config/` | Server-specific environment variable parsing and validation with Zod. |
| `src/mcp-server/tools/` | Tool definitions (`*.tool.ts`). |
| `src/mcp-server/resources/` | Resource definitions (`*.resource.ts`). |
| `src/services/cert/` | `node:tls` — TLS handshake, X.509 parsing, expiry and protocol flagging. |
| `src/services/dns/` | `node:dns` — multi-resolver DNS fan-out, propagation discrepancy detection. |
| `src/services/statuspage/` | Statuspage public API client with 60-second in-memory cache. |
| `src/services/status-adapters/` | Native-API adapters (Status.io, Slack, AWS Health, Google Cloud Service Health, Azure status, Firehydrant) + `api_type` dispatch, normalizing into the Statuspage shapes. |
| `src/services/vendor-registry/` | In-memory vendor registry loaded from `src/data/vendor-registry.ts`. |
| `src/data/` | Static vendor registry data file (`vendor-registry.ts`). |
| `tests/` | Vitest tests mirroring `src/`. |

---

## Development guide

See [`CLAUDE.md`](./CLAUDE.md) for development guidelines and architectural rules. The short version:

- Handlers throw, framework catches — no `try/catch` in tool logic
- Use `ctx.log` for request-scoped logging, `ctx.state` for tenant-scoped storage
- Register new tools and resources via the barrels in `src/mcp-server/*/definitions/index.ts`
- `devops_check_certs` and `devops_check_dns` use only Node.js stdlib — add no external deps for these paths

---

## Contributing

Issues are welcome. Run checks and tests before submitting:

```sh
bun run devcheck
bun run test
```

---

## License

Apache-2.0 — see [LICENSE](LICENSE) for details.