Skip to main content
Glama

netdiag-mcp

A real Model Context Protocol server exposing infrastructure and security tools across ten domains, every one of them behind real cryptographic identity verification, least-privilege access scopes, rate limiting, a tamper-evident audit log, and structural prompt injection defense.

This used to be two repos. One had the governance depth but routed by keyword matching, not a real protocol, nothing could actually connect to it the way an MCP client connects to a real server. The other had the real protocol but none of the governance. Neither was the honest whole story on its own. This is the merge: 20 real MCP tools, all gated by the same governance layer this portfolio's ADRs document in depth, with the network dependency-tracing pattern already in this repo folded in as one of the domains.

All data is fabricated. No real hostnames, IPs, credentials, or vendor names appear anywhere in this repo.

Why this exists

I run infrastructure and security for a real multi-site organization, and I built and led the production version of this pattern before writing about it. This repo is that architecture, rebuilt from scratch with sample data, wired into an actual protocol server instead of a demo script, so it's something a real MCP client can connect to and call, not just something to read.

Related MCP server: MCP IT Ops Server

The 20 tools, by domain

Domain

Tools

Network monitoring

list_sites, site_health, root_cause, degraded_sites, redundancy_gaps

Ticketing

open_tickets_by_priority, unassigned_tickets

Identity

accounts_without_mfa, disabled_accounts

Endpoint management

unauthorized_software

Security alerts

open_alerts, alerts_with_severity_assessment

Vulnerability

exposed_devices

Certificates

expiring_certificates

Virtualization

vms_under_resource_pressure

Network AAA

elevated_aaa_users

Server management

servers_under_resource_pressure, servers_needing_patching, deployment_history

Governance demo

request_hard_limit_action

root_cause is the standout in network monitoring: it doesn't just report what's broken, it walks circuit, power, core, and access in dependency order and identifies the most likely upstream cause, the same redundancy model documented in network-iac-lab.

alerts_with_severity_assessment is the standout in security alerts: it derives a severity signal from an alert's free text without ever returning that text, using the schema-constrained quarantine boundary described below.

request_hard_limit_action always refuses to execute directly. Every consequential action type (isolating an endpoint, disabling an account, rebooting a server, pushing a deployment, and others) is hard-coded to require a two-step, out-of-band confirmation, no matter what any pre-approval manifest says.

The governance layer

Every tool call above passes through _authorize() in netdiag_mcp/server.py before its underlying logic runs:

  1. Authentication. netdiag_mcp/governance/identity.py: real Ed25519 signature verification over (caller, domain, nonce), checked against a registered public key. Replay protection is persistent, a used nonce is rejected even across a server restart. Keys can be rotated or revoked outright, with no dual-key grace window.

  2. Authorization. netdiag_mcp/governance/access_control.py: five caller identities (a read-only dashboard, a field technician's assistant, a network engineer's assistant, a security analyst's assistant, an admin agent), each with an explicit, checked set of reachable domains. Deny-by-default.

  3. Rate limiting. netdiag_mcp/governance/rate_limit.py: a token bucket per caller, so a runaway loop gets stopped, not just slowed.

  4. Audit logging. netdiag_mcp/governance/audit_log.py: a hash-chained, file-persisted, checkpointed log of every governance decision. Checkpointing catches the one attack plain hash-chaining can't, a full, internally self-consistent rewrite of the entire log.

  5. Prompt injection defense. netdiag_mcp/governance/injection_defense.py: untrusted free text (an alert description, a ticket summary) is stripped to an allowlisted set of fields before it can reach a decision, proven against two worked examples of embedded injection attempts in the sample data.

  6. Schema-constrained quarantine. netdiag_mcp/governance/quarantine.py: for the one case where dropping free text entirely would lose real signal, a classifier reads it but its output is constrained to a small, pre-declared set of values, proven with a test where the classifier itself is deliberately adversarial and still can't get anything past the schema check. netdiag_mcp/governance/model_classifier.py is a real, working Anthropic API call usable as a drop-in replacement for the default keyword heuristic.

  7. Supply chain integrity. netdiag_mcp/governance/supply_chain_integrity.py: a SHA-256 hash manifest of this repo's own trusted code, checked into sample_data/integrity_baseline.json, detecting modification, deletion, or unauthorized addition after the fact.

Full reasoning for every one of these lives in docs/adr/, ten ADRs, numbered in the order the decisions were actually made, including what each layer still doesn't cover and why.

One deliberate tradeoff, stated plainly

Every tool above takes caller, nonce, and signature_hex as explicit arguments. That's not how most production MCP servers authenticate, official MCP guidance points toward transport-level authentication (OAuth) instead. This repo authenticates at the call level anyway, because the point of this repo is proving the governance layer works, in a way every test and every manual call demonstrates rather than assumes a transport layer handled somewhere this repo can't show. See docs/adr/010-merging-into-one-mcp-server.md for the full reasoning.

Running it

pip install -e ".[dev]"
pytest tests/ -v

108 tests: the original network monitoring suite, the full governance layer (identity, access control, rate limiting, audit logging with checkpointing, injection defense, quarantine, secrets, supply chain integrity), and 6 integration tests that call tools through the real MCP protocol layer (server.call_tool), not the underlying Python functions directly, proving the governance wiring holds at the layer a real client would actually hit.

To generate your own demo keypairs (the checked-in sample_data/demo_public_keys.json has public keys only, no private key ever ships in this repo):

python3 scripts/generate_demo_keys.py

To run the server itself against an MCP client (stdio transport):

python3 -m netdiag_mcp.server

To call a tool manually, sign a request with a registered private key:

from netdiag_mcp.governance.identity import sign_request, make_nonce

nonce = make_nonce()
signature_hex = sign_request(your_private_key, caller, domain, nonce).hex()
# pass caller, nonce, and signature_hex as the tool's leading arguments

Available Tools

5 tools
degraded_sitesA

List every site with at least one layer not reporting fully healthy.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden of disclosing behavioral traits. It does not mention side effects, read-only nature, authentication requirements, or potential errors. The phrase 'not reporting fully healthy' is also somewhat ambiguous regarding what 'reporting' means.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence with no unnecessary words. It front-loads the action and includes the key filtering condition without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is adequate for a simple list operation with no parameters and an output schema indicated. It could mention output fields or ordering, but the existence of an output schema reduces the need for the description to explain return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. No parameter descriptions are needed since there are no parameters to explain.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states a specific verb ('List') and resource ('sites') with a clear filtering condition ('at least one layer not reporting fully healthy'). It distinguishes itself from sibling tools like list_sites, which presumably lists all sites, by focusing on degraded ones.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when needing sites with unhealthy layers, but it does not explicitly state when to use this over alternatives like site_health or root_cause. Guidance is present only by inference from the wording.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_sitesA

List every site known to this diagnostics server.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. The verb 'list' clearly indicates a read-only operation, and there is no mention of side effects. While not explicitly stating 'read-only', the semantics are unambiguous enough for a safe interaction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence with no redundant words. It delivers the essential information efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of parameters and the simple listing action, the description is sufficient for an agent to invoke the tool correctly. No additional context is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has no parameters, so schema coverage is 100%. The description adds no parameter-specific information because there is nothing to describe. This meets the baseline for high coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (list) and the resource (sites) with specificity (every, known to this diagnostics server). It distinguishes itself from sibling tools like site_health or degraded_sites by being the comprehensive listing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it (as the general listing of all sites) but does not explicitly name alternatives or conditions for choosing this tool over the siblings. An agent can infer usage from the wording, but it is not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

redundancy_gapsA

List every (site, layer) pair currently lacking redundancy.

Flags layers that aren't redundant even if their current status is healthy, since a non-redundant healthy layer is one failure away from an outage.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of explaining behavior. It makes clear that the tool lists/flags redundancy gaps without indicating side effects or mutations, and it explains the rationale behind including healthy layers.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences, with the primary purpose front-loaded and the second sentence adding a valuable clarifying nuance. No unnecessary words or details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the low complexity, no parameters, and existing output schema, the description is complete enough for an agent to understand what the tool does and why it matters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has no parameters, so the baseline is 4. There is no parameter information needed beyond the empty schema, and the description focuses entirely on the tool's purpose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('List') and the specific resource ('every (site, layer) pair currently lacking redundancy'), which distinguishes it from sibling tools like list_sites and site_health.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool, especially the nuance that it flags non-redundant layers even when status is healthy. It does not explicitly mention alternatives or when not to use it, but the use case is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

root_causeA

Trace the most likely root-cause layer for a site's issues.

Walks circuit, power, core, and access in dependency order and reports the first layer that isn't healthy, since an upstream issue typically explains downstream symptoms rather than being a separate problem. site must exactly match a name returned by list_sites.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the walk order, the logic (first unhealthy layer), and the reporting behavior. It does not mention side effects (likely read-only) or error cases, but for a trace tool this is a reasonable level of transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each serving a purpose: purpose, algorithm/reasoning, and input constraint. Front-loaded with the primary action, no unnecessary fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, algorithm, and input constraint, but does not specify the exact return format (e.g., whether it returns the layer name, a report object) or the behavior when all layers are healthy. Given there is no output schema, this is a minor gap for an agent invoking the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does so by adding a crucial constraint: 'site must exactly match a name returned by list_sites.' This gives meaning to the single parameter beyond the bare schema, making the input unambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action: 'Trace the most likely root-cause layer for a site's issues.' Describes the algorithm (walks circuit, power, core, access in dependency order) and the output (first unhealthy layer). This clearly differentiates it from siblings like site_health or degraded_sites, which focus on status or listing rather than root-cause analysis.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides context for when it is appropriate ('an upstream issue typically explains downstream symptoms') and a prerequisite ('site must exactly match a name returned by list_sites'). However, it does not explicitly mention alternative tools or when NOT to use this one, leaving some ambiguity for the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

site_healthA

Get the layered health status (circuit, power, core, access) for a site.

site must exactly match a name returned by list_sites.

ParametersJSON Schema
NameRequiredDescriptionDefault
siteYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses that the operation is a read ('Get'), enumerates the status layers, and adds the exact-match input constraint. However, it doesn't describe error behavior for mismatched site names, the output/return format, or any side effects — though for a low-risk read operation the risk profile is modest. Adds useful context but has gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero waste. The primary purpose is front-loaded in the first sentence, and the operational constraint follows in the second. Every word earns its place with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with no output schema and no annotations, the description covers the essentials: what the tool does and what the parameter requires. It could add the return format of the health status values, but the core calling information is complete and an agent can invoke this tool correctly with what's provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: the second sentence clarifies that the 'site' parameter is a name that must exactly match one from list_sites. This adds meaningful semantic information beyond the bare schema, which only types it as a required string. The description fully compensates for the coverage gap for the sole parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get') and resource ('layered health status') and enumerates exactly which layers are covered (circuit, power, core, access). This distinguishes it from siblings like list_sites (which lists sites) and degraded_sites (which finds problem sites), though it doesn't name them explicitly. Clear purpose with minor room for explicit sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence gives a concrete, actionable constraint: 'site must exactly match a name returned by list_sites.' This tells the agent to call list_sites first and use an exact name value. It implies usage context well but stops short of stating when not to use this tool or naming alternatives as the preferred choice for other scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.1.0
    • First observeddegraded_sites
    • First observedlist_sites
    • First observedredundancy_gaps
    • First observedroot_cause
    • First observedsite_health

TDQS

A4.3/5.0

Scored across 5 tools

Disambiguation5/5

Each tool has a clear, distinct purpose: listing sites, checking health, identifying degraded sites, tracing root causes, and finding redundancy gaps. No overlap or ambiguity.

Naming Consistency5/5

All tool names use lowercase snake_case consistently, following a predictable and readable pattern. The mix of verb-led and noun-led names is coherent within the domain.

Tool Count5/5

Five tools is well within the typical range and perfectly suited for a network diagnostics server, covering the essential operations without unnecessary bloat.

Completeness5/5

The tool set provides comprehensive coverage for network diagnostics: listing sites, checking health, surfacing degraded sites, identifying root causes, and flagging redundancy risks. No critical gaps are apparent.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers