Skip to main content
Glama
tcharod
by tcharod
README.md
# mcp-guardrails-kit

A prompt-injection-aware tool gateway and MCP server, built as a reference implementation
for a fictional internal ticket-triage assistant.

**This is a portfolio project, not a real product.** It exists to demonstrate a concrete,
testable pattern for building guardrails around agentic tool use — permission tiers,
untrusted-content quarantine, and heuristic injection detection — around a small but real
[MCP](https://modelcontextprotocol.io) server. The domain (support ticket triage) is
generic and interchangeable; the guardrail patterns are the point.

## What it demonstrates

- **Tool permission tiers with explicit confirmation.** Every tool is registered as either
  `read_only` or `sensitive`. Sensitive tools (`draft_reply`, `escalate_ticket`) never
  execute a side effect until the caller passes `confirmed=True` — every caller in this
  codebase routes through the gateway's single `invoke()` entry point by construction (see
  `docs/adr/0001-tool-permission-tiers.md` for the honest caveat: this is a code-review
  convention, not a language-enforced boundary).
- **Quarantine of untrusted external content.** `fetch_external_page` returns content
  fetched from a URL linked inside a ticket — a realistic prompt-injection vector. That
  content is wrapped and clearly delimited as *data*, never treated as instructions, before
  it is handed back to any caller or model.
- **Heuristic injection detection with a verdict.** Quarantined (and other) text is scanned
  for injection patterns and returns one of `ALLOW` / `FLAG` / `BLOCK`. A `BLOCK` verdict
  replaces the payload with a safe refusal instead of the raw text (never a verbatim
  excerpt — even the scan's own `matched_patterns` are redacted to category labels before
  crossing a tool boundary). This is backed by a red-team test suite of known injection
  phrasings.
- **Three independent scan gates, not just one.** A `Supervisor` walks a
  triage → draft → review → escalate pipeline. A ticket's own subject/body — the most
  directly attacker-controlled input in the system — is scanned right after lookup; the
  drafted reply is re-scanned before escalation is considered; and the escalation `reason`
  a model proposes is scanned again before `escalate_ticket` is ever called. A `BLOCK` at
  any of the three halts the pipeline right there.

## Install & run

```bash
pip install -e ".[dev]"
pytest
ruff check .
```

No external services or API keys are required to install, test, or lint. See
[Scope & non-goals](#scope--non-goals) below for what `pip install -e ".[live]"` adds.

## Connecting the MCP server to a real client

After `pip install -e .` (or `pip install mcp-guardrails-kit` once published), the
`mcp-guardrails-kit` command is registered as a console entry point
(see `[project.scripts]` in `pyproject.toml`) and speaks the MCP stdio protocol. Point a
real MCP client at it — for example, Claude Desktop or Claude Code — with a config block
like:

```json
{
  "mcpServers": {
    "guardrails-kit": {
      "command": "mcp-guardrails-kit"
    }
  }
}
```

For Claude Desktop, this goes in `claude_desktop_config.json`; for Claude Code, add it via
`claude mcp add` or the equivalent project-level MCP config. No arguments or environment
variables are required for the default (non-live) mode.

## Scope & non-goals

- **Heuristic injection detection is defense in depth, not a guarantee.** It is a
  regex/keyword-based scanner, not a model-backed classifier. It will miss novel or
  sufficiently obfuscated phrasings — see `docs/adr/0003-heuristic-injection-detection.md`
  for the explicit tradeoff. The permission-tier and quarantine layers stay in effect even
  when detection fails; injection detection is one layer among three, not the only one.
- **There is no real ticketing system behind this.** `search_knowledge_base` and
  `lookup_ticket` read from small in-memory fixtures. There is no database, no external
  ticketing API integration, and no persistence.
- **`AnthropicModelClient` is optional and live-only.** It is gated behind the `live` extra
  (`pip install -e ".[live]"`) and is never imported or exercised by the test suite or CI —
  tests and the default install path have zero dependency on any external LLM API or network
  access.
- **All data is in-memory and resets on restart.** Drafts, escalations, and fetched external
  content are not persisted anywhere; restarting the server clears all state.

## More

- Architecture and data flow: [`docs/architecture.md`](docs/architecture.md)
- Design decisions: [`docs/adr/`](docs/adr/)

TDQS

A4.2/5.0

Scored across 5 tools

Disambiguation5/5

Each tool has a distinct, non-overlapping purpose: search knowledge base, lookup ticket, fetch external page (with security features), draft reply, and escalate ticket. No ambiguity between them.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern in snake_case (e.g., search_knowledge_base, lookup_ticket, fetch_external_page, draft_reply, escalate_ticket). The convention is uniform and predictable.

Tool Count5/5

Five tools is a well-scoped set for a guardrails kit focused on support workflows. Each tool serves a clear function, and there is no bloat or insufficiency given the narrow domain.

Completeness4/5

The tool surface covers the main support workflow: search, lookup, secure fetching, drafting replies, and escalation. Minor gaps exist (e.g., no tool to update ticket status or send replies), but these are not critical to the stated guardrails purpose.

Maintenance

ActivitySlowing
ResponsivenessNo issues