mcpgov
README.md
# mcpgov
A production-grade MCP server over Postgres, where the hard 20% is the point:
authentication, tenant isolation, provably idempotent mutations, loop-aware
rate limiting, and a tamper-evident audit trail — each claim pinned by a test
that runs against a real database, and an end-to-end adversarial demo that
attacks the live server over real HTTP in CI.
## Why
Every company is wrapping internal systems in MCP servers so agents can use
them. Most wrappers are demos: the tools work, and nothing stops a retried
mutation from applying twice, a token from reading another tenant's rows, or an
agent loop from hammering the same failing call all night. This repo is the
missing 80-to-100 stretch, built as five controls that a tool author cannot
forget, because they live in the middleware chain and the database rather than
in tool bodies.
## The controls, and the evidence for each
**1. Identity is verified, not asserted.** OAuth-style short-lived Bearer
tokens (HS256, `exp`/`aud`/`iss`/`jti` all required), plus RFC 7523 jwt-bearer
federation: an IdP-issued assertion is exchanged for a local token after
signature-by-`kid`, audience, issuer, expiry-with-bounded-skew, and single-use
`jti` checks. Group-to-team mapping is a declarative allowlist; an unmapped
group grants nothing, including a group that happens to be named after a real
team. 19 tests, including forged signatures, `alg=none`, replayed assertions,
and cross-audience confusion.
**2. Tenant isolation is enforced by the database, not the queries.** Postgres
row-level security with `FORCE`, keyed on a GUC populated only from verified
token claims — there is no request field through which a client names a tenant.
The server runs as a role that is neither owner nor superuser (either would
bypass RLS silently; a test asserts this about the running role). Teams are
jsonb, not CSV, after measuring that CSV encoding let a principal entitled to
`gamma,delta` read the distinct team literally named `gamma,delta`. Cross-tenant
reads return `not_found` byte-identical to genuinely absent ids.
**3. Mutations are exactly-once under retry storms.** A claim-then-execute
ledger: the idempotency claim and the business write commit in one transaction,
duplicates replay the stored response marked `_replayed`, a key reused with
different arguments is refused, and only non-retryable failures are cached
(caching a transient one would convert a blip into a permanent failure that
looks healthy). Pinned by a 16-thread concurrent-duplicates test repeated 5
times, plus crash-recovery: an unclean death does not poison the key.
**4. Rate limiting distinguishes a runaway loop from a legitimate burst.** Two
layers, because "too fast" and "stuck" are different problems: a GCRA shaper
per `(principal, tool_class)` that delays, and a loop breaker that looks for
*repetition without progress* and opens a per-tool circuit. On the seeded
evaluation (200 trials/family, reproduced in CI byte-for-byte):
| workload | outcome |
|---|---|
| 5 runaway families (same-error, cycle, no-op write, infinite transient retry, idempotent replay) | broken in 100% of trials, median 8 wasted calls |
| 6 legitimate families (poll, paginate, fan-out, backoff retry, burst, small worklist) | 0 false breaks |
| declared long poll | bounded by its declared budget, by design |
| undeclared poll | known false positive, documented — without a declaration it is indistinguishable from a no-progress loop |
| token-bucket baseline | denies 41/80 of the runaway and 41/80 of the legitimate burst at identical arrival rates — its verdict is a function of rate, the label is a function of shape; it never breaks a loop, only slows it |
**5. The audit trail is tamper-evident, including truncation.** Every attempt —
allowed, denied, unauthenticated — is one row in an HMAC hash chain, with
arguments redacted to digests. The chain head lives in a singleton row read
under `FOR UPDATE` (the naive `ORDER BY seq DESC LIMIT 1` head forked under
concurrency: 16 concurrent appends produced 9 distinct predecessors, and
verification cried tamper on clean traffic). Verification requires a separate
auditor login the server never holds, because the writer being unable to read
the whole log — and the reader being unable to write — is what makes "the chain
verified" a statement about the data rather than the writer's self-report.
Deleting the tail is detected too, which hash-chaining alone cannot see.
## The demo: the live server, attacked over real HTTP
`demos/session.py` boots `mcpgov serve`, mints tokens with the CLI, and drives
13 acts through the official MCP client — no token (401 with the RFC 9728
challenge), forged signature, scope escalation, a retry storm of duplicated
mutations, key reuse with different arguments, cross-tenant probing, a runaway
loop broken on attempt 7 while other tools keep answering, tenant-scoped audit
reads, chain verification, and a superuser rewriting one row's `deny` to
`allow` — caught with the exact seq. Each act asserts its expected outcome; CI
fails if any deviates. Transcript: [`results/demo-session.json`](results/demo-session.json).
To point Claude Code or Cursor at it interactively: [`docs/clients.md`](docs/clients.md).
## Run it
Needs Python 3.12+, [uv](https://docs.astral.sh/uv/), and any Postgres 14+.
```sh
uv sync
# the full suite: 100 tests, most against the real database
MCPGOV_TEST_DSN='postgresql://postgres@127.0.0.1:5432/postgres' uv run pytest
# the adversarial demo (creates its own scratch database)
MCPGOV_DEMO_DSN='postgresql://postgres@127.0.0.1:5432/mcpgov_demo' \
uv run python demos/session.py
# the limiter evaluation (seeded; CI diffs the output against the committed file)
uv run python scripts/eval_limiter.py
```
## Design notes worth stealing
- **Controls in middleware, not tool bodies.** A check inside a tool is a check
a new tool can forget, silently. The guard wraps every inbound message, so it
also sees calls to tools that do not exist and calls that fail schema
validation — attempts worth auditing that a per-tool check never sees.
- **Middleware sees the wire format.** `call_next` returns a JSON-RPC dict
(`isError`, camelCase), not the typed `CallToolResult`. The attribute
spelling returned its default forever: every refusal was audited as `allow`,
and loop detection was structurally dead in the live server while the offline
harness scored it at full recall. The tests now drive a real client session.
- **`INSERT ... RETURNING` re-evaluates the SELECT policy** on the new row —
so the unscoped audit appender could not record tenant-tagged events at all.
- **Permissive RLS policies apply by role membership, not the active role.**
Granting the app role membership of the auditor role (the convenient wiring)
switched the full-read policy on for every app query and removed the tenant
boundary from audit reads without a single error.
- **Retryability is declared once, per error code**, and consulted by both the
idempotency cache and the loop breaker's thresholds — a permanent failure
misfiled as transient would let a runaway run 20 calls instead of 6.
## Limitations, honestly
- The IdP in the federation tests is a local fixture with published keys, not a
live Okta/Entra tenant; the validation logic is real, the network hop is not.
- `LocalTokenVerifier` is symmetric-key (HS256) — right for a single-server
deployment, not for a fleet where issuers and verifiers must not share a
secret.
- Loop detection keys on the verified `(principal, client_id, tool)`; a
principal that can provision many `client_id`s can fan out across buckets.
That is a provisioning-quota problem, stated rather than solved.
- The GitHub write path (`src/mcpgov/github.py`) is at-least-once made
effectively-once by reconciliation — exactly-once across two systems with no
shared transaction does not exist, and its list endpoint lags a successful
create by ~5-6s (measured), which is exactly why its reconciler polls past
the lag and refuses to create after a short-deadline miss. Its suite runs
against a recorded fake with an offline guard test.
## Layout
```
src/mcpgov/
auth.py tokens, RFC 7523 assertion exchange, group->team mapping
db.py pool, tenant-scoped connections, migration + grants
schema.sql tables, FORCE RLS policies, the audit chain head
idempotency.py the claim-then-execute ledger
limiter.py GCRA shaper + loop breaker (the two-layer argument)
audit.py HMAC hash chain: append, verify, truncation detection
server.py the five tools and the guard middleware
github.py writing to a second system that has no idempotency
cli.py migrate / seed / token / serve / verify-audit
tests/ 100 tests; Postgres-backed ones run against a real database
demos/session.py the 13-act adversarial session over real HTTP
scripts/ the seeded limiter evaluation
results/ committed evidence: eval numbers, demo transcript
```
This server cannot be deployed
Maintenance
ActivityStale
ResponsivenessNo issues