Skip to main content
Glama
bhargavlukka

SecureAgentServer

by bhargavlukka
README.md
# Secure MCP-Based Agent System

A support-ticket + customer-account MCP server and client, built and secured end to end:
signed-token authentication, two independent guardrails against prompt injection and tool
poisoning, a documented threat model, and a human-in-the-loop gate on the one destructive
tool call. Built for the "Secure an MCP-Based Agent System" lab.

## Architecture

```mermaid
flowchart TD
    HOST["Host application"] --> CLIENT["client.py\n(fastmcp.Client + elicitation_handler)"]
    CLIENT <-->|"Streamable HTTP\nAuthorization: Bearer <signed JWT>"| SERVER

    subgraph Server["server.py — FastMCP('SecureAgentServer')"]
        AUTH["JWTVerifier (HS256)\nissuer + audience + signature checked"]
        G1["Guardrail 1: sanitize_untrusted_text()\napplied to ticket body/subject"]
        G2["Guardrail 2: verify_tool_manifest()\nchecked at startup, refuses to start on mismatch"]
        TOOLS["Tools: search_tickets, lookup_customer_account,\nclose_ticket (elicitation-gated)"]
        AUTH --> TOOLS
        TOOLS --> G1
    end
    G2 -.->|startup check| SERVER

    TOOLS --> TICKETS[("data/tickets.json\n(untrusted customer text)")]
    TOOLS --> CUSTOMERS[("data/customers.json\n('internal API')")]
```

## Setup

```bash
pip install -r requirements.txt

# 1. Set the JWT signing secret (never commit the real value; see .env.example)
export MCP_JWT_SECRET="a-long-random-secret-at-least-32-characters"

# 2. Generate the pinned tool-integrity manifest (a deliberate, manual step —
#    see docs/threat-model.md Risk #2)
python generate_manifest.py

# 3. Run the server
python server.py

# 4. In another terminal (same MCP_JWT_SECRET exported)
python client.py --auto-confirm   # non-interactive demo
python client.py                  # interactive: real yes/no confirmation prompts
```

## Reproducing the security controls

| Control | How to verify it |
|---|---|
| Signed-JWT authentication | `client.py`'s last two demos: a read-only-scoped token is rejected from `close_ticket`, and a forged/unsigned token is rejected with `401` before any tool runs at all — both captured in `demo/session_log.txt` |
| Guardrail 1: prompt-injection sanitization | Run the client and look at the `search_tickets` output for `TICKET-2002` (its body contains a planted "IGNORE ALL PREVIOUS INSTRUCTIONS..." payload) — the returned text is wrapped and the trigger phrase redacted. The server's own stdout logs a `[SECURITY]` line when this fires. |
| Guardrail 2: tool-poisoning detection | `python demo/verify_tampering_detection.py` — tampers with `close_ticket`'s description in memory and shows `verify_tool_manifest()` catching the mismatch; see `demo/tampering_detection_log.txt` for a captured run |
| Human-in-the-loop on destructive actions | `close_ticket` in the demo log only completes after `[ELICITATION] ... -> accepting`; run `client.py` without `--auto-confirm` to see a real interactive confirmation prompt |
| Least-privilege scoping | `add_ticket_note`/`close_ticket` explicitly check `write:tickets` via `get_access_token()` (`server.py::_require_scope`), not just a connection-level requirement — demonstrated by the read-only-token rejection above |

## Threat model

Full write-up, including 5 identified risks with mitigations and residual-risk notes
stated explicitly rather than glossed over: [`docs/threat-model.md`](docs/threat-model.md).

## What's deliberately out of scope, stated plainly

- `mint_token.py` stands in for a real OAuth 2.1 identity provider — a production
  deployment needs real token issuance/rotation/revocation, not a local minting script.
- No rate limiting or network-layer hardening (TLS termination, WAF) — this is an
  application-layer security demo, not a full deployment hardening guide.
- The regex-based half of Guardrail 1 is explicitly a secondary, best-effort layer — see
  `docs/threat-model.md` Risk #1 for why the delimiter-wrapping is the control actually
  relied upon.

## No secrets committed

`MCP_JWT_SECRET` is read from the environment (see `.env.example`) and is never
hardcoded or committed. `tool_manifest.json` **is** committed deliberately — it's a pinned
hash manifest (like a lockfile), not a secret.