govmcp
by sarathalapad
README.md
# governed-mcp-server
[](https://github.com/sarathalapad/governed-mcp-server/actions/workflows/ci.yml)


**An MCP server that lets AI agents operate infrastructure, but only through server-side policy, human approval for risky actions, and a tamper-evident audit trail.**
Built by **Sarathkumar K** (AI Solutions Architect) to show how agent tool-calling via the
[Model Context Protocol](https://modelcontextprotocol.io) can be put under real governance
controls instead of relying on the model to behave.
The "infrastructure" is a **fictional, simulated IT-operations environment** (services such as
`payments-api` and `search-indexer` across `dev` / `staging` / `prod`), so everything is safe to run
and fully deterministic.
## How it works
```mermaid
flowchart LR
A[AI agent / LLM] --> B[MCP client]
B -- stdio / JSON-RPC --> C[govmcp MCP server]
C --> D{Policy engine<br/>rate limit, schema,<br/>lockdown, ranges}
D -- read tools --> E[(Simulated<br/>environment)]
D -- write, auto-approved<br/>non-prod only --> F[Executor]
D -- write, needs approval --> G[Approval queue]
H[Human approver<br/>govmcp approvals CLI] -- approve / reject --> G
G -- re-validate, execute once --> F
F --> E
C -.every call and decision.-> I[(Hash-chained<br/>audit log)]
H -.-> I
```
All governance lives in one place (`govmcp/gateway.py`). The MCP tool functions are thin, typed
wrappers, so there is no second code path that could skip a check.
## Tools
| Tool | Kind | Risk | Behaviour |
|---|---|---|---|
| `list_services` | read | low | Services with status, version, replicas (optional `env` filter) |
| `get_service_health` | read | low | Error rate, p95 latency, restarts, deploy history |
| `search_runbooks` | read | low | Keyword search over `data/runbooks/*.md` |
| `get_recent_incidents` | read | low | Newest incidents, filter by service / status |
| `get_approval_status` | read | low | Follow up on a pending write request |
| `restart_service` | write | medium | Auto-runs in `dev`/`staging` per policy; `prod` needs approval |
| `scale_service` | write | medium | Replicas limited to 1-10 by policy; auto-runs in `dev` only |
| `rollback_deployment` | write | high | Always needs human approval, in every environment |
Every write tool accepts `dry_run: true`, which returns the exact before/after change and whether it
would need approval, without touching anything.
## Governance controls and why they matter
| Control | What it does | Why it matters |
|---|---|---|
| **Server-side policy** (`policy.json`) | Per-tool risk level, allowed ranges, which envs may auto-run | Prompts can be injected; a server-side rule cannot be talked out of |
| **Hard guardrails in code** | Policy loader rejects `prod` or `high` risk in `auto_approve_envs` | A config typo must not be able to silently remove human oversight |
| **Maker-checker approvals** | Writes return an approval ID; a *different* identity must approve | Separation of duties: the agent can propose, never self-authorise |
| **Expiry, no replay, pinned args** | Approvals have a TTL, are single-use, and carry a digest of the exact arguments | Stale or reused approvals are a classic way to run something nobody signed off on |
| **Re-validation at execution** | Policy, lockdown and target state are checked again when the approver says yes | The world changes between request and approval |
| **Emergency lockdown** | `"lockdown": true` blocks every write immediately (policy is re-read per call) | A kill switch that works without redeploying or restarting anything |
| **Per-caller rate limits** | Sliding one-minute window per caller identity | Contains runaway agent loops |
| **Input validation** | SDK schema check plus the gateway's own strict check (no bool-as-int, no unknown args) | Governance must not depend on the transport being well-behaved |
| **Structured errors** | Stable `code`, `message`, `hint`; never a stack trace | Agents self-correct from a code and hint; stack traces leak internals |
| **Fail closed** | An unreadable or invalid policy refuses all calls | "Policy missing" must never mean "everything allowed" |
| **Tamper-evident audit** | Every call, refusal and approval is a JSONL record chained by SHA-256, or HMAC-SHA256 when `GOVMCP_AUDIT_KEY` is set | You can prove what the agent asked for, who approved it, and that the log was not edited |
## Quick start
```bash
git clone https://github.com/sarathalapad/governed-mcp-server.git
cd governed-mcp-server
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements-dev.txt
python -m pytest # run the test suite
python examples/demo_client.py # scripted agent + human approver, over real stdio
```
Run the server yourself and act as the approver:
```bash
python -m govmcp serve --caller my-agent # MCP over stdio (normally started by your MCP client)
python -m govmcp approvals list # see pending requests
python -m govmcp approvals approve apr-1234abcd --user bob
python -m govmcp approvals reject apr-1234abcd --user bob --reason "restart first"
python -m govmcp verify-audit # check the hash chain
```
Runtime state (environment, approvals, audit log) is written to `./state/`. Delete it to reset.
Environment variables: `GOVMCP_STATE_DIR`, `GOVMCP_POLICY`, `GOVMCP_DATA_DIR`, `GOVMCP_AUDIT_KEY`, `GOVMCP_CALLER`.
## Demo transcript
Real output of `python examples/demo_client.py`. The agent side talks to the server through the
official MCP Python SDK client over stdio; the human side uses the real CLI.
```text
### Tools advertised by the server
[
"list_services",
"get_service_health",
"search_runbooks",
"get_recent_incidents",
"get_approval_status",
"restart_service (write)",
"scale_service (write)",
"rollback_deployment (write)"
]
### Agent: get_service_health(payments-api, prod)
{
"status": "degraded",
"version": "2.4.0",
"error_rate": 6.8,
"p95_latency_ms": 910
}
### Agent: search_runbooks('rollback error rate')
[
{"runbook": "payments-api-high-error-rate.md", "title": "payments-api: high error rate"}
]
### Agent: rollback_deployment(..., dry_run=true)
{
"ok": true,
"status": "dry_run",
"would_change": {"target": "payments-api@prod", "before": {"version": "2.4.0"}, "after": {"version": "2.3.1"}},
"requires_approval": true,
"note": "no changes were made"
}
### Agent: rollback_deployment(payments-api, prod)
{
"ok": true,
"status": "pending_approval",
"approval_id": "apr-998482d0",
"risk": "high",
"would_change": {"target": "payments-api@prod", "before": {"version": "2.4.0"}, "after": {"version": "2.3.1"}},
"expires_in_seconds": 900,
"note": "a different human must approve this before it runs"
}
### Agent: scale_service(replicas=50) is outside policy
{
"ok": false,
"error": {"code": "POLICY_DENIED", "message": "replicas must be between 1 and 10", "hint": "larger changes need a capacity review"}
}
### Human: review queue
$ python -m govmcp approvals list
apr-998482d0 rollback_deployment payments-api@prod risk=high by=demo-agent expires in 899s
change: {"version": "2.4.0"} -> {"version": "2.3.1"}
### Agent identity tries to approve its own request
$ python -m govmcp approvals approve apr-998482d0 --user demo-agent
{
"ok": false,
"error": {"code": "POLICY_DENIED", "message": "maker-checker: the requester cannot approve their own action"}
}
### Human (bob) approves
$ python -m govmcp approvals approve apr-998482d0 --user bob
{
"ok": true,
"status": "executed",
"approval_id": "apr-998482d0",
"change": {"target": "payments-api@prod", "before": {"version": "2.4.0"}, "after": {"version": "2.3.1"}},
"result": {"service": "payments-api", "env": "prod", "status": "healthy", "version": "2.3.1", "replicas": 4, "restarts": 0}
}
### Replay attempt
$ python -m govmcp approvals approve apr-998482d0 --user bob
{
"ok": false,
"error": {"code": "POLICY_DENIED", "message": "approval apr-998482d0 is already executed; approvals cannot be reused"}
}
### Agent: get_approval_status
{
"ok": true,
"result": {"id": "apr-998482d0", "tool": "rollback_deployment", "status": "executed", "requester": "demo-agent", "decided_by": "bob", "result": {"service": "payments-api", "env": "prod", "status": "healthy", "version": "2.3.1", "replicas": 4, "restarts": 0}}
}
### Agent: health after rollback
{
"status": "healthy",
"version": "2.3.1",
"error_rate": 0.2,
"p95_latency_ms": 150
}
### Auditor: verify the hash chain
$ python -m govmcp verify-audit
audit log OK: 10 records, chain intact
### Auditor: verify again after someone edits the log
$ python -m govmcp verify-audit
audit log TAMPERED: line 7: record hash mismatch (edited?) (6 records verified before the break)
```
## Use it from an MCP client
Most MCP clients (Claude Desktop, VS Code, Cursor and others) accept an `mcpServers` block along
these lines. Replace the paths with your own checkout; on Windows the interpreter is
`.venv\Scripts\python.exe`.
```json
{
"mcpServers": {
"governed-ops": {
"command": "/path/to/governed-mcp-server/.venv/bin/python",
"args": ["-m", "govmcp", "serve", "--caller", "desktop-agent"],
"env": {
"PYTHONPATH": "/path/to/governed-mcp-server",
"GOVMCP_AUDIT_KEY": "set-a-long-random-secret"
}
}
}
}
```
Then ask the assistant something like *"payments-api in prod looks unhealthy, check the runbook and
propose a fix"*, and approve the resulting request from a terminal with `python -m govmcp approvals`.
## Project layout
```text
governed-mcp-server/
├── govmcp/
│ ├── __main__.py # python -m govmcp
│ ├── cli.py # serve | approvals list/approve/reject | verify-audit
│ ├── server.py # MCP tool definitions (thin wrappers, MCP Python SDK)
│ ├── gateway.py # the single enforcement point: policy, approvals, execution, audit
│ ├── policy.py # policy model, strict loader, rate limiter
│ ├── schema.py # tool argument schemas and validation
│ ├── approvals.py # persistent maker-checker queue
│ ├── audit.py # hash-chained / HMAC audit log and verifier
│ ├── environment.py # simulated services, plan/apply for write actions
│ └── errors.py # structured, agent-readable failures
├── data/ # fictional services, incidents and runbooks (seed data)
├── examples/demo_client.py
├── tests/ # policy, approvals, audit, MCP end-to-end (in-memory + stdio)
├── policy.json
├── requirements.txt # runtime: mcp
└── requirements-dev.txt # + pytest
```
## Design notes and limitations
This is a focused reference implementation, not a production control plane. Being explicit about the
gaps is part of the design:
- **Caller identity is asserted, not authenticated.** The agent identity comes from `--caller` and
the approver from `--user`. In production, derive both from real authentication: MCP's OAuth-based
authorization for remote (HTTP) transports on the agent side, and SSO for approvers, so
maker-checker compares verified identities.
- **Approvals via CLI.** A real deployment would surface approvals in chat or ticketing tools
(Slack, Teams, ServiceNow) with the same rules behind them.
- **File-based state.** JSON files and a JSONL log keep the project dependency-free and easy to
inspect. There is no cross-process locking, so concurrent approvers could race. Use a database with
transactions (and an append-only store for audit) for real use.
- **Tamper-evident, not tamper-proof.** The hash chain detects edits, deletions and reordering in the
middle of the log. Truncating the tail is only detectable if the latest hash is anchored elsewhere
(for example, periodically shipped to a separate system). Without `GOVMCP_AUDIT_KEY`, anyone with
write access can recompute a plain SHA-256 chain, so set the key in any shared environment.
- **Rate limits are in memory** and per server process; they reset on restart.
- **Dry-run is evaluated under the same policy**, so it is also refused during lockdown or for
out-of-range arguments. That keeps the preview honest about what would actually happen.
- **The environment is simulated.** `plan()`/`apply()` in `environment.py` is where a real
integration (Kubernetes, a cloud API, a CMDB) would plug in; the governance layer would not change.
- Built on the official [MCP Python SDK](https://github.com/modelcontextprotocol/python-sdk) v2
(`MCPServer`, the successor of `FastMCP`).
## License
MIT. See [LICENSE](LICENSE). Copyright (c) 2026 Sarathkumar K.
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues