Skip to main content
Glama
YuCPbit
by YuCPbit

๐Ÿงพ mcp-proof

Ship an MCP server with a receipt.

Audit an MCP server from the wire. Get conformance, security, regression โ€” and effect โ€” evidence in one reproducible, offline-verifiable delivery report.

v0.8 extends conformance below the response boundary โ€” a tool can declare readOnlyHint: true, return a perfectly ordinary response, and still mint a credential that outlives the grant that authorized it. An out-of-band observer reads the real external effects; a probe exercises the authority the tools create. โ†’ the research lane

stdio + Streamable HTTP ยท 2026-07-28 + legacy eras ยท HTML / JSON / JUnit / SARIF

ci python checks transports license

English ยท ็ฎ€ไฝ“ไธญๆ–‡ ยท ๆ—ฅๆœฌ่ชž ยท ํ•œ๊ตญ์–ด ยท Franรงais

Browse the live reports โ†’ ยท Effect-aware evaluation โ†’

A real audit of the official MCP filesystem server: 27 conformance checks, the MSSS compliance table, 34 regression replays โ€” ship-ready, one advisory finding.


๐Ÿš€ Quick start

pip install git+https://github.com/YuCPbit/mcp-proof
mcp-proof run python my_server.py --fixtures fixtures/ --record-if-missing --out report.html

Auditing a running HTTP server instead? mcp-proof run --url http://localhost:8000/mcp --out report.html

Exit codes are the gate: 0 โ€” every MUST check passed, no blocking security findings (advisories may remain), no behavioural drift. 1 โ€” the audit completed and the server failed it. 2 โ€” the audit did not complete (missing baseline, internal auditor error) and proves nothing about the server, in either direction.

mcp-proof plan python my_server.py                             # what would auto-baselining call, and why
mcp-proof record python my_server.py --fixtures fixtures/      # freeze the behavioural contract
mcp-proof replay --fixtures fixtures/ -- python my_server.py   # fail on any drift
mcp-proof inspect python my_server.py --out baseline.json      # freeze the contract surface
mcp-proof diff baseline.json current.json                      # BREAKING / ADDITIVE / METADATA, exit 1 on breaking
mcp-proof verify report.json                                   # recheck the report's internal fingerprints offline
mcp-proof effects --sqlite state.db -- python my_server.py     # declared vs observed external effects (v0.8)
mcp-proof effects --fs-root data/ -- python my_server.py       # same lane, directory-backed servers

See the difference in 60 seconds with the built-in demo pair โ€” a clean server and one with nine planted violations:

mcp-proof run python demo/good_server.py --fixtures demo/fixtures-good --out report-good.html   # โ†’ SHIP-READY
mcp-proof run python demo/bad_server.py --out report-bad.html                                    # โ†’ 5 MUST failures, 3 security findings

Related MCP server: MCProbe

๐Ÿ”ฌ The four lanes

Lane

What it proves

How

Protocol conformance

The server implements MCP correctly on the wire โ€” era negotiation, JSON-RPC error semantics, tool/resource/prompt surfaces, output schemas, capability consistency, pagination, stdout hygiene

A hand-rolled JSON-RPC probe observes the raw byte stream, so nothing is smoothed over

Security & hygiene

Tool metadata is clean: no injected instructions, hidden Unicode, leaked secrets, or unconstrained execution surfaces

Deterministic static analysis, every finding carrying its MSSS control ID

Behaviour regression

The server still does exactly what it did at delivery

Record/replay of provenance-fingerprinted golden fixtures, drift graded by severity

Effect conformance (v0.8, research lane)

A tool's observed effect on external state matches its declared annotations; created objects are classified by probing, not by name

An out-of-band observer snapshots the server's state store around every call; a probe exercises created objects โ€” requires an observation channel (--sqlite / --fs-root), see below. Validated on three third-party servers, not just our own testbed

Every lane feeds one report โ€” and the report ends with a prioritized fix list, so it doubles as a remediation plan.

โœ… Validation

An audit tool has to earn more trust than the thing it audits. What stands behind every release:

  • 160 tests, including an adversarial suite that attacks the auditor itself: violations hidden on page 2 of paginated listings, tampered fixtures and manifests, hash-stripping downgrade attempts, reports with edited verdict banners, drift classes that used to slip through, invalid synthesized baselines โ€” and, for the effect lane, a read-only-annotated tool that mints a credential detectable only out-of-band, a persistent object that must not be classified as authority-bearing, a no-probe audit whose authority verdicts must degrade to unknown/SKIP instead of passing, and a create+delete call that must not hide its delete behind the headline effect.

  • CI on Linux, macOS and Windows ร— Python 3.11 / 3.12 / 3.13, plus a packaging job that builds the wheel, installs it fresh, and runs a real audit against a real server before anything ships โ€” and an experiments job that re-runs E1โ€“E3 from scratch and fails unless the outputs are byte-identical to the committed results: the reproducibility claim is CI-enforced, not asserted.

  • Exercised against servers we did not build: three case-study audits of published third-party MCP servers (both official reference servers and a community one โ€” three store types, with and without annotations), with committed evidence โ€” see the effect lane section below.

  • Cross-validated against the official v2 SDK in both directions: the official client adopts mcp-proof's hand-rolled modern test server via server/discover, and mcp-proof runs fully green against official v2 SDK servers on both transports (scripts/crosscheck_modern_server.py).

  • Fail-closed by design: a broken pagination walk, a tampered or unverifiable fixture, a missing baseline or an internal auditor error each stop the audit loudly โ€” and every command answers with the same taxonomy: exit 2 and one stable line, never a traceback, never a silently smaller audit, and never evidence against the target.

  • Offline-verifiable reports: mcp-proof verify report.json recomputes both fingerprints from the report's own fields; the document fingerprint covers everything a reader sees โ€” verdict banner, audit status, summary counters, the MSSS table, next steps โ€” so any post-audit edit breaks it. It is an internal-consistency proof, not a signature (attestation is on the roadmap).

โœจ What's under the hood

  • ๐Ÿ” Wire-level protocol checks across every surface, every page, both eras โ€” mcp-proof speaks raw JSON-RPC to your server and auto-detects its era: 32 checks for the 2026-07-28 modern era (server/discover, _meta envelope enforcement, resultType, ttlMs/cacheScope on every cacheable result, -32022 version rejection, HTTP routing-header enforcement) and 27 for the initialize-handshake era โ€” exact error codes, schema validity, structured output, stdout hygiene, pagination safety on all three list surfaces, dedicated resources & prompts lanes, and verified negative probes: TOOL-07 sends inputs that provably violate the declared inputSchema (a schema-valid baseline with exactly one field mutated) and warns when the server answers them normally โ€” and treats a hang as its own finding, never as rejection. One pagination collector feeds every lane, so a tool hidden on page 2 is audited exactly like a tool on page 1.

  • ๐Ÿ›ก๏ธ Security audit tied to a public standard โ€” 6 deterministic checks (tool-description poisoning, invisible/bidi characters, leaked credentials, unconstrained injection surfaces, advertised shell execution) over every advertised tool on every page, with a schema walker that sees through $ref/allOf/nesting/array items โ€” config.shell.command cannot hide one level down. Each check maps to canonical control IDs of the MCP Server Security Standard's 24-entry control matrix (23 fully documented controls plus the MCP-DEPLOY-04 future-control placeholder), rendered as a compliance table whose verdicts never outrun their evidence: full direct proof says met, clean-but-indirect evidence says partial, and a control the checks cannot see says manual review.

  • ๐Ÿงช Effect checks that read the world, not the response (v0.8) โ€” with an observation channel configured (--sqlite for a SQLite-backed server, --fs-root for a directory-backed one), the effect lane snapshots external state before and after every call and diffs it into per-object create/update/delete effects, attributed to the call that caused them (per-target ops, so a move can't hide its delete behind its create). Four checks compare that against the declared annotations, following the spec's semantics exactly โ€” defaults are pessimistic, so an absent hint is never flagged, only an explicit claim the observed effect falsifies: EFF-01 (a readOnlyHint: true tool caused no observed write), EFF-02 (no observed delete under an explicit destructiveHint: false), EFF-03 (an idempotentHint tool's repeated identical call was a no-op), EFF-06 (a created credential's continued effectiveness still depends on the grant that authorized it โ€” established by revoking each candidate dependency out-of-band and re-exercising the object, then restoring it). Every field of every effect record carries how it was known โ€” declared, observed, probed, or unknown โ€” and a dimension with no channel SKIPs rather than passing.

  • ๐Ÿ“ผ A regression suite your client keeps โ€” and that verifies itself before it judges anyone โ€” records in either protocol era; golden fixtures freeze the server's behaviour with SHA-256 provenance, including every content type (binary payloads as digests, so a swapped image can never replay as OK). Before replaying, an integrity gate recomputes every contract hash and the manifest fingerprint: a missing, tampered, duplicated or stale fixture aborts the replay instead of being silently skipped โ€” deleting a fixture's stored hash counts as tampering, not as an older schema, and baselines that predate contract hashing are refused unless --allow-legacy-fixtures explicitly opts in. Replay grades every drift (BREAKING / VALUE / COSMETIC / LATENCY) โ€” any structured or JSON value change is at least VALUE, a flipped "approved"โ†’"denied" can never pass as cosmetic โ€” and preserves stateful call order (sequence-numbered fixtures, order-sensitive fingerprint). A baseline is never created implicitly: run fails closed when fixtures are missing unless you opt in with --record-if-missing.

  • ๐Ÿ“„ A report for humans and machines โ€” self-contained HTML with sticky navigation, per-check anchors (report.html#SEC-03), attention/passed filters, an evidence-scope card and a collapsible MSSS matrix; --pdf for print. The same versioned model ships as --json (schema v3), --junit for any CI, and --sarif for the GitHub Security tab. The effect lane renders its own evidence page: declared annotations beside the observed effect, the response a response-only auditor would have read, the objects that resulted, and the probe's authority/dependency verdict, each value tagged with how it was known.

  • ๐Ÿ” Reproducible by design โ€” zero LLM calls, zero API keys. Two fingerprints, honestly separated: behavior_sha256 is computed from server behaviour alone (check verdicts, replay verdicts, protocol facts โ€” never timestamps, latency, the launch command or the auditor's version), so identical server behaviour fingerprints identically on any machine; run_hash freezes the whole report document โ€” evidence, verdict banner, audit status, summaries, MSSS table โ€” minus only the volatile timestamp block. mcp-proof verify rechecks both offline: an internal-consistency proof that any post-audit edit breaks, not a signature. Acceptance is verification, not trust.

  • ๐Ÿงฏ Conservative call planning, and a trust correction (v0.8) โ€” auto-baselining classifies tools by a conservative name/description heuristic, and mcp-proof plan shows exactly what would be called and on what basis before anything touches production. As of v0.8, MCP annotations may only add caution: destructiveHint: true still forces a skip, but an unverified readOnlyHint: true no longer rescues a mutating-looking tool into the auto-call set โ€” the spec says clients MUST treat annotations as untrusted, and the effect lane exists precisely because a "read-only" tool can mint credentials. --include-destructive and --edge-cases opt into more.

  • ๐Ÿ“‹ A contract diff for CI โ€” mcp-proof inspect freezes the served surface (capabilities + tools + resources + prompts, fully paginated, absent-vs-empty recorded) into a fingerprinted manifest โ€” and refuses to write one at all if any pagination walk cannot be completed, because half a surface frozen as "the baseline" makes every later diff against the missing half invisible. Volatile wire metadata is removed by location, never by key name, so a schema property that happens to be called ttlMs or nextCursor stays part of the contract. mcp-proof diff classifies every change as BREAKING / ADDITIVE / METADATA and exits non-zero on breaking ones โ€” schema tightening, enum narrowing, required-flips, removed output fields and weakened safety annotations all count.

๐Ÿ“Š Real audits, real reports

Target

Verdict

Report

Official MCP filesystem server (@modelcontextprotocol/server-filesystem)

โœ… SHIP-READY โ€” 11/11 MUST checks, 34/34 replays clean, 4 write tools auto-skipped

Live report ยท PDF

Official "everything" reference server (@modelcontextprotocol/server-everything)

โœ… SHIP-READY โ€” 20/20 MUST + 7/7 SHOULD, 0 security findings across 13 tools. Protocol + security lanes; recording deliberately skipped โ€” its get-env tool dumps environment variables

Live report

Official memory server (@modelcontextprotocol/server-memory)

โœ… SHIP-READY โ€” 16/16 MUST, 4/4 replays clean, 5 write/delete tools auto-skipped, one advisory: unconstrained search_nodes.query (SEC-04)

Live report

Official sequential-thinking server (@modelcontextprotocol/server-sequential-thinking)

โœ… SHIP-READY โ€” 11/11 MUST, 1/1 replays clean, one advisory (2,781-char tool description, SEC-05); investigating its honest TOOL-08 skip exposed the served inputSchema omitting a runtime-required field

Live report

2026-07-28 modern-era server (zero-dep, cross-validated against the official v2 SDK)

โœ… SHIP-READY โ€” era auto-detected via server/discover, 23/23 MUST incl. negative probes, 2/2 replays

Live report

Demo server with 9 planted violations

โŒ NOT SHIP-READY โ€” 5 MUST failures + 5 security findings (3 blocking, 2 advisory), every one caught with evidence

Live report

Well-behaved demo server

โœ… SHIP-READY โ€” 18/18 MUST, full three-lane pass incl. regression baseline

Live report

Effect testbed, silent-keymint variant

โŒ EFF-01 FAIL โ€” a tool annotated readOnlyHint: true returns a normal read response while inserting an api_keys row; the out-of-band state diff attributes the write to the call

Effect evidence

Effect case study: official memory server (JSONL store, fully annotated)

โœ… EFF-01/02/03 PASS โ€” every readOnly claim, destructive declaration and idempotency claim held under out-of-band observation, including a repeated delete_entities; authority honestly SKIPs (no probe channel)

Effect evidence

Effect case study: official filesystem server (jailed directory, stock observer)

โœ… EFF-01/02/03 PASS โ€” 10 readOnly tools wrote nothing; move_file's observed delete is covered by its destructiveHint โ€” and this call exposed (and fixed) a headline-only blind spot in EFF-02

Effect evidence

Effect case study: community SQLite server (@executeautomation/database-server, zero annotations)

โœ… honest degradation โ€” nothing declared, so EFF-01/03 SKIP; the observed DELETE is consistent with the spec's pessimistic default; effects still attributed per row via the stock --sqlite channel

Effect evidence

๐Ÿงช The effect-aware research lane (v0.8)

The three delivery lanes stop at the wire: their notion of behaviour is the response byte stream. The effect lane extends the same declared-vs-observed method one level deeper. Its parts, concretely:

  • Observer (effects/observe.py): snapshots external state โ€” a SQLite store read directly from the file, or a directory tree โ€” before and after each tool call, and diffs the snapshots into per-object create/update/delete deltas. It never asks the tools what changed, so an effect is seen whether or not the response mentions it.

  • Probe (effects/probes.py): attempts to use a created object as a credential against the service's real authorization rule. "Authority-bearing" is then an observed outcome (the object authorized an action), not a guess from a field name; "still effective" is the probe succeeding now, not the object still being listed.

  • Lineage, kept as three separate fields: created_via (which call produced the object โ€” observed), authorized_by (the grant the session ran under โ€” declared), and depends_on (what its continued effectiveness actually requires โ€” established by revoking each candidate out-of-band, re-exercising, and restoring). The distinction is the point: an API key authorized_by a grant whose depends_on does not include that grant survives the grant's revocation.

  • Testbed (testbed/): a deterministic SQLite-backed MCP server with ordinary persistent objects (notes) and credential objects (API keys, webhooks, share links), a one-bit grant, lifecycle tools, and mutation flags that plant exactly one annotation lie each โ€” modelled on documented incident patterns (a read path that mints authority; revocation that doesn't cascade). Ground truth is read out-of-band by testbed/saas_oracle.py, never through the audited MCP surface.

Three experiments run against it (python experiments/run_all.py โ€” deterministic, and CI re-runs the pipeline to enforce byte-identical outputs). Each does a different job โ€” E1 is the detection experiment, E2 is construct validation (the persistence/authority distinction is operationally real), E3 is a lifecycle measurement (existence โ‰  current effectiveness). The numbers are detection behaviour on planted inconsistencies in this controlled setting โ€” not production prevalence:

Experiment

Probe / effect observation

Baselines

E1 ยท detection โ€” declared vs observed effect, honest server + 6 single-lie variants; oracle = the mutation ledger

precision / recall 1.000 / 1.000

response-level 1.000 / 0.333 ยท name heuristic 1.000 / 0.333

E2 ยท construct validation โ€” authority-bearing vs merely persistent, 8-object corpus incl. a decoy note named api_key_backup and a credential that never persisted; oracle = by-construction authority labels

accuracy 1.000 on the corpus built to separate the signals

name keyword 0.875 ยท persistenceโ‡’authority 0.625

E3 ยท lifecycle measurement โ€” existence vs current effectiveness across 6 lifecycle scenarios (grant revoked, key revoked, key deleted, TTL expired, cascade); oracle = per-scenario intended effectiveness

accuracy 1.000, 0 false-ineffective

existence 0.500 ยท grant-status 0.333, 1 false-ineffective

The two results that motivate the lane: in E1, the three lies whose effects never appear in the response (silent-keymint, shadow-webhook, phantom-write) are caught only by the state diff โ€” a response-level auditor is structurally blind to them, which is why all-lane recall for the baselines is 0.333. In E3's grant_revoked scenario, the key created under the grant remains effective after the grant is revoked (the testbed's auth rule checks the key's own row, not the grant โ€” the same shape as documented OAuth-app persistence incidents); the grant-status proxy reports it neutralized, which is the one dangerous false-ineffective in the table โ€” authorized_by read as if it were depends_on.

And off the testbed: three case studies run the same instrument against published third-party servers โ€” the official memory server (JSONL store, fully annotated), the official filesystem server (jailed directory, stock observer, fully annotated) and a community SQLite server (no annotations at all). Every declaration that could be checked held; where nothing was declared the checks SKIPped rather than inventing a verdict; and the filesystem run exposed a real blind spot in EFF-02's first implementation (a create+delete call whose headline effect hid the delete), now fixed and pinned by a test. The probe half stayed honestly unknown throughout โ€” none of these services mints an exercisable credential โ€” so probe-backed authority results remain testbed-validated. Evidence: case studies on the evaluation site.

Methodology, oracle design, baselines, related work and limitations: docs/effect-aware-conformance.md ยท results with raw evidence: evaluation site ยท reproduction: experiments/README.md.

๐Ÿงญ How this relates to the official conformance suite

The MCP project maintains modelcontextprotocol/conformance โ€” scenario tests that verify protocol behaviour for servers and clients, including auth flows. If you need a protocol-correctness baseline, run it; mcp-proof's conformance lane covers overlapping ground from its own wire-level probes.

mcp-proof exists for the half the official suite doesn't do: delivery evidence. A fingerprinted, offline-verifiable report a client can keep; MSSS security mapping; golden behavioural regression with a fail-closed integrity gate; contract snapshot/diff as a CI gate; SARIF/JUnit artifacts; and the effect-conformance research lane. Use the official suite to prove the protocol; use mcp-proof to prove the delivery โ€” they compose, and cross-validating against the official suite is on the roadmap.

๐Ÿ“ก Protocol support

Transports

stdio โœ… ยท Streamable HTTP โœ…

Surfaces

tools โœ… ยท resources โœ… ยท prompts โœ… โ€” capability-aware in both directions

Modern era 2026-07-28 (server/discover, stateless _meta)

โœ… conformance lane, auto-detected โ€” --era auto|modern|legacy

Legacy era (initialize handshake, 2024-11-05 โ†’ 2025-11-25)

โœ… all lanes

Regression lane

โœ… both eras โ€” SDK session (legacy) ยท probe-backed session (modern)

Works with servers in any language โ€” mcp-proof talks to the process (or URL), not to your codebase.

โš™๏ธ CI in one step

- uses: YuCPbit/mcp-proof@v0.8.1
  with:
    server-command: python my_server.py
    fixtures: fixtures/

The job fails unless the server is ship-ready, and leaves mcp-proof-report.html / .json / .junit.xml / .sarif behind for upload. Prefer raw commands? mcp-proof run โ€ฆ --junit r.xml --sarif r.sarif plus mcp-proof diff is the same gate.

๐Ÿ—๏ธ Build on the audit-clean template

Building a server rather than auditing one? templates/server-starter/ is a fastmcp server that passes this audit out of the box โ€” constrained input schemas, proper error semantics, structured output, every practice annotated with the check ID it satisfies. Copy, implement your tools, audit, ship with the report.

๐Ÿ–ฅ๏ธ Platforms

macOS

โœ… developed & fully validated

Linux

โœ… exercised in CI

Windows

โœ… exercised in CI (--pdf needs Chrome/Chromium installed)

๐Ÿ—บ๏ธ Roadmap

Current โ€” v0.8.1

Research hardening: three third-party case studies with committed evidence (official memory + filesystem servers, community SQLite server); effect checks follow the spec's annotation defaults exactly (absence of a hint is never flagged); per-target delete attribution (a create+delete call can't hide its delete); --fs-root observation channel; CI-enforced byte-identical experiment reproduction

v0.8.0

Effect-aware research lane: out-of-band effect observation, probe-backed authority classification, residual-authority measurement (mcp-proof effects, experiments/, docs); annotation-trust correction; the evaluation site

v0.7.2

Truthfulness patch: verify fingerprints the whole document (report schema v3), fixture hash-stripping counts as tampering, legacy baselines fail closed, one exit-code taxonomy across every command

Next

2026-07-28 depth: MRTR input_required round-trips ยท cross-validation against the official conformance suite in CI ยท a real-provider probe adapter โ€” exercising created credentials against a live provider; the observation half is already exercised on third-party servers

Later

Signed evidence bundles (attestation) ยท opt-in semantic lane (LLM-graded assertions) โ€” parked until the deterministic core is complete

Release history lives in CHANGELOG.md.

๐Ÿ” Limitations

mcp-proof proves what can be proven deterministically, and says which is which:

  • Security checks cover the observable protocol and metadata surface. MSSS controls that need deployment, source or process evidence are always reported as manual review โ€” never as passed.

  • Authorization flows are out of scope for the delivery report: OAuth handshakes are not audited (the official conformance suite covers auth scenarios). The effect lane reasons about authority-bearing objects a tool creates, on a controlled testbed with an out-of-band observer โ€” it does not audit a production OAuth deployment.

  • The effect lane is a measurement instrument, not a black-box lane. It needs an observation channel (--sqlite, --fs-root); effects to systems it cannot observe are reported unknown/SKIP, never assumed absent. Its quantitative results are detection behaviour on a synthetic testbed โ€” the case studies show the instrument working on third-party servers, but they validate the observation half only: probe-backed authority/effectiveness remains testbed-validated. See docs/effect-aware-conformance.md ยง8.

  • Auto-baselining classifies tools by a conservative name/description heuristic; as of v0.8 an unverified readOnlyHint no longer overrides it. Review the skip list in the fixtures manifest before trusting a baseline recorded against production.

  • Semantic correctness (does the answer mean the right thing?) is outside the deterministic core by design.

๐Ÿ“„ License

MIT โ€” the taxonomy in the MSSS compliance section follows the MCP Server Security Standard (CC BY-SA 4.0).

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    A stdio MCP server that audits other MCP servers over the live protocol. It connects to any MCP target (stdio or HTTP), lints every tool's schema for agent-usability, then actually calls the tools with deliberately broken inputs to see how the server handles them, and returns a 0โ€“100 conformance score with a per-dimension breakdown rendered as Markdown.
    6
    6
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Audits MCP server configurations for security risks including capability inventory, SSRF, prompt injection, and drift detection. Works in read-only mode and can also be used as an MCP server to let AI agents audit their own attack surface.
    4
    MIT