mcp-guardrails-kit
# mcp-guardrails-kit
A prompt-injection-aware tool gateway and MCP server, built as a reference implementation
for a fictional internal ticket-triage assistant.
**This is a portfolio project, not a real product.** It exists to demonstrate a concrete,
testable pattern for building guardrails around agentic tool use — permission tiers,
untrusted-content quarantine, and heuristic injection detection — around a small but real
[MCP](https://modelcontextprotocol.io) server. The domain (support ticket triage) is
generic and interchangeable; the guardrail patterns are the point.
## What it demonstrates
- **Tool permission tiers with explicit confirmation.** Every tool is registered as either
`read_only` or `sensitive`. Sensitive tools (`draft_reply`, `escalate_ticket`) never
execute a side effect until the caller passes `confirmed=True` — every caller in this
codebase routes through the gateway's single `invoke()` entry point by construction (see
`docs/adr/0001-tool-permission-tiers.md` for the honest caveat: this is a code-review
convention, not a language-enforced boundary).
- **Quarantine of untrusted external content.** `fetch_external_page` returns content
fetched from a URL linked inside a ticket — a realistic prompt-injection vector. That
content is wrapped and clearly delimited as *data*, never treated as instructions, before
it is handed back to any caller or model.
- **Heuristic injection detection with a verdict.** Quarantined (and other) text is scanned
for injection patterns and returns one of `ALLOW` / `FLAG` / `BLOCK`. A `BLOCK` verdict
replaces the payload with a safe refusal instead of the raw text (never a verbatim
excerpt — even the scan's own `matched_patterns` are redacted to category labels before
crossing a tool boundary). This is backed by a red-team test suite of known injection
phrasings.
- **Three independent scan gates, not just one.** A `Supervisor` walks a
triage → draft → review → escalate pipeline. A ticket's own subject/body — the most
directly attacker-controlled input in the system — is scanned right after lookup; the
drafted reply is re-scanned before escalation is considered; and the escalation `reason`
a model proposes is scanned again before `escalate_ticket` is ever called. A `BLOCK` at
any of the three halts the pipeline right there.
## Install & run
```bash
pip install -e ".[dev]"
pytest
ruff check .
```
No external services or API keys are required to install, test, or lint. See
[Scope & non-goals](#scope--non-goals) below for what `pip install -e ".[live]"` adds.
## Connecting the MCP server to a real client
After `pip install -e .` (or `pip install mcp-guardrails-kit` once published), the
`mcp-guardrails-kit` command is registered as a console entry point
(see `[project.scripts]` in `pyproject.toml`) and speaks the MCP stdio protocol. Point a
real MCP client at it — for example, Claude Desktop or Claude Code — with a config block
like:
```json
{
"mcpServers": {
"guardrails-kit": {
"command": "mcp-guardrails-kit"
}
}
}
```
For Claude Desktop, this goes in `claude_desktop_config.json`; for Claude Code, add it via
`claude mcp add` or the equivalent project-level MCP config. No arguments or environment
variables are required for the default (non-live) mode.
## Scope & non-goals
- **Heuristic injection detection is defense in depth, not a guarantee.** It is a
regex/keyword-based scanner, not a model-backed classifier. It will miss novel or
sufficiently obfuscated phrasings — see `docs/adr/0003-heuristic-injection-detection.md`
for the explicit tradeoff. The permission-tier and quarantine layers stay in effect even
when detection fails; injection detection is one layer among three, not the only one.
- **There is no real ticketing system behind this.** `search_knowledge_base` and
`lookup_ticket` read from small in-memory fixtures. There is no database, no external
ticketing API integration, and no persistence.
- **`AnthropicModelClient` is optional and live-only.** It is gated behind the `live` extra
(`pip install -e ".[live]"`) and is never imported or exercised by the test suite or CI —
tests and the default install path have zero dependency on any external LLM API or network
access.
- **All data is in-memory and resets on restart.** Drafts, escalations, and fetched external
content are not persisted anywhere; restarting the server clears all state.
## More
- Architecture and data flow: [`docs/architecture.md`](docs/architecture.md)
- Design decisions: [`docs/adr/`](docs/adr/)
TDQS
Scored across 5 tools
Each tool has a distinct, non-overlapping purpose: search knowledge base, lookup ticket, fetch external page (with security features), draft reply, and escalate ticket. No ambiguity between them.
All tool names follow a consistent verb_noun pattern in snake_case (e.g., search_knowledge_base, lookup_ticket, fetch_external_page, draft_reply, escalate_ticket). The convention is uniform and predictable.
Five tools is a well-scoped set for a guardrails kit focused on support workflows. Each tool serves a clear function, and there is no bloat or insufficiency given the narrow domain.
The tool surface covers the main support workflow: search, lookup, secure fetching, drafting replies, and escalation. Minor gaps exist (e.g., no tool to update ticket status or send replies), but these are not critical to the stated guardrails purpose.