Skip to main content
Glama
README.md
# netdiag-mcp

A real [Model Context Protocol](https://modelcontextprotocol.io) server exposing infrastructure and security tools across ten domains, every one of them behind real cryptographic identity verification, least-privilege access scopes, rate limiting, a tamper-evident audit log, and structural prompt injection defense.

This used to be two repos. One had the governance depth but routed by keyword matching, not a real protocol, nothing could actually connect to it the way an MCP client connects to a real server. The other had the real protocol but none of the governance. Neither was the honest whole story on its own. This is the merge: 20 real MCP tools, all gated by the same governance layer this portfolio's ADRs document in depth, with the network dependency-tracing pattern already in this repo folded in as one of the domains.

All data is fabricated. No real hostnames, IPs, credentials, or vendor names appear anywhere in this repo.

## Why this exists

I run infrastructure and security for a real multi-site organization, and I built and led the production version of this pattern before writing about it. This repo is that architecture, rebuilt from scratch with sample data, wired into an actual protocol server instead of a demo script, so it's something a real MCP client can connect to and call, not just something to read.

## The 20 tools, by domain

| Domain | Tools |
|---|---|
| Network monitoring | `list_sites`, `site_health`, `root_cause`, `degraded_sites`, `redundancy_gaps` |
| Ticketing | `open_tickets_by_priority`, `unassigned_tickets` |
| Identity | `accounts_without_mfa`, `disabled_accounts` |
| Endpoint management | `unauthorized_software` |
| Security alerts | `open_alerts`, `alerts_with_severity_assessment` |
| Vulnerability | `exposed_devices` |
| Certificates | `expiring_certificates` |
| Virtualization | `vms_under_resource_pressure` |
| Network AAA | `elevated_aaa_users` |
| Server management | `servers_under_resource_pressure`, `servers_needing_patching`, `deployment_history` |
| Governance demo | `request_hard_limit_action` |

`root_cause` is the standout in network monitoring: it doesn't just report what's broken, it walks circuit, power, core, and access in dependency order and identifies the most likely upstream cause, the same redundancy model documented in [network-iac-lab](https://github.com/miservicespro/network-iac-lab).

`alerts_with_severity_assessment` is the standout in security alerts: it derives a severity signal from an alert's free text without ever returning that text, using the schema-constrained quarantine boundary described below.

`request_hard_limit_action` always refuses to execute directly. Every consequential action type (isolating an endpoint, disabling an account, rebooting a server, pushing a deployment, and others) is hard-coded to require a two-step, out-of-band confirmation, no matter what any pre-approval manifest says.

## The governance layer

Every tool call above passes through `_authorize()` in `netdiag_mcp/server.py` before its underlying logic runs:

1. **Authentication.** `netdiag_mcp/governance/identity.py`: real Ed25519 signature verification over `(caller, domain, nonce)`, checked against a registered public key. Replay protection is persistent, a used nonce is rejected even across a server restart. Keys can be rotated or revoked outright, with no dual-key grace window.
2. **Authorization.** `netdiag_mcp/governance/access_control.py`: five caller identities (a read-only dashboard, a field technician's assistant, a network engineer's assistant, a security analyst's assistant, an admin agent), each with an explicit, checked set of reachable domains. Deny-by-default.
3. **Rate limiting.** `netdiag_mcp/governance/rate_limit.py`: a token bucket per caller, so a runaway loop gets stopped, not just slowed.
4. **Audit logging.** `netdiag_mcp/governance/audit_log.py`: a hash-chained, file-persisted, checkpointed log of every governance decision. Checkpointing catches the one attack plain hash-chaining can't, a full, internally self-consistent rewrite of the entire log.
5. **Prompt injection defense.** `netdiag_mcp/governance/injection_defense.py`: untrusted free text (an alert description, a ticket summary) is stripped to an allowlisted set of fields before it can reach a decision, proven against two worked examples of embedded injection attempts in the sample data.
6. **Schema-constrained quarantine.** `netdiag_mcp/governance/quarantine.py`: for the one case where dropping free text entirely would lose real signal, a classifier reads it but its output is constrained to a small, pre-declared set of values, proven with a test where the classifier itself is deliberately adversarial and still can't get anything past the schema check. `netdiag_mcp/governance/model_classifier.py` is a real, working Anthropic API call usable as a drop-in replacement for the default keyword heuristic.
7. **Supply chain integrity.** `netdiag_mcp/governance/supply_chain_integrity.py`: a SHA-256 hash manifest of this repo's own trusted code, checked into `sample_data/integrity_baseline.json`, detecting modification, deletion, or unauthorized addition after the fact.

Full reasoning for every one of these lives in `docs/adr/`, ten ADRs, numbered in the order the decisions were actually made, including what each layer still doesn't cover and why.

## One deliberate tradeoff, stated plainly

Every tool above takes `caller`, `nonce`, and `signature_hex` as explicit arguments. That's not how most production MCP servers authenticate, official MCP guidance points toward transport-level authentication (OAuth) instead. This repo authenticates at the call level anyway, because the point of this repo is proving the governance layer works, in a way every test and every manual call demonstrates rather than assumes a transport layer handled somewhere this repo can't show. See `docs/adr/010-merging-into-one-mcp-server.md` for the full reasoning.

## Running it

```bash
pip install -e ".[dev]"
pytest tests/ -v
```

108 tests: the original network monitoring suite, the full governance layer (identity, access control, rate limiting, audit logging with checkpointing, injection defense, quarantine, secrets, supply chain integrity), and 6 integration tests that call tools through the real MCP protocol layer (`server.call_tool`), not the underlying Python functions directly, proving the governance wiring holds at the layer a real client would actually hit.

To generate your own demo keypairs (the checked-in `sample_data/demo_public_keys.json` has public keys only, no private key ever ships in this repo):

```bash
python3 scripts/generate_demo_keys.py
```

To run the server itself against an MCP client (stdio transport):

```bash
python3 -m netdiag_mcp.server
```

To call a tool manually, sign a request with a registered private key:

```python
from netdiag_mcp.governance.identity import sign_request, make_nonce

nonce = make_nonce()
signature_hex = sign_request(your_private_key, caller, domain, nonce).hex()
# pass caller, nonce, and signature_hex as the tool's leading arguments
```

TDQS

A4.3/5.0

Scored across 5 tools

Disambiguation5/5

Each tool has a clear, distinct purpose: listing sites, checking health, identifying degraded sites, tracing root causes, and finding redundancy gaps. No overlap or ambiguity.

Naming Consistency5/5

All tool names use lowercase snake_case consistently, following a predictable and readable pattern. The mix of verb-led and noun-led names is coherent within the domain.

Tool Count5/5

Five tools is well within the typical range and perfectly suited for a network diagnostics server, covering the essential operations without unnecessary bloat.

Completeness5/5

The tool set provides comprehensive coverage for network diagnostics: listing sites, checking health, surfacing degraded sites, identifying root causes, and flagging redundancy risks. No critical gaps are apparent.

Maintenance

ActivityMaintained
ResponsivenessNo issues