Skip to main content
Glama
Rinkia

bastiongate

by Rinkia
README.md
# bastiongate

**MCP security gateway.** An inline proxy that sits between an AI agent and its
MCP servers and enforces security on every call:

- **scans `tools/list`** and drops tools whose definitions carry prompt
  injection or hidden unicode (via [bastionsupply](https://github.com/Rinkia/bastionsupply))
- **enforces a tool allow/deny policy** — the agent can only call what you permit
- **scans tool-call results** and blocks any that carry indirect prompt
  injection before the agent ever reads them
- **logs every message** as a JSONL trace for forensics

The runtime-enforcement leg of the **bastion family**:

| tool | job |
|------|-----|
| **bastiongate** | **gate** — enforce security inline on live MCP traffic |
| [bastionsupply](https://github.com/Rinkia/bastionsupply) | scan an MCP server before you trust it |
| [agentbastion](https://github.com/Rinkia/agentbastion) | prevent — firewall around a running agent |
| [bastionprobe](https://github.com/Rinkia/bastionprobe) | attack — pentest your agent with injections |
| [bastiontrace](https://github.com/Rinkia/bastiontrace) | investigate — forensics on an agent trace |

## Install

```bash
pip install bastiongateway
```

(The PyPI distribution is `bastiongateway`; the import package and `bastiongate`
CLI keep that name.)

## Use

The gate *is* an MCP server to your agent, and a client to the real one. Point
your MCP client's `command` at the gate and put the real server after `--`:

```jsonc
// mcp.json
{
  "mcpServers": {
    "docs": {
      "command": "bastiongate",
      "args": ["run", "--policy", "policy.yaml", "--log", "gate.jsonl",
               "--", "npx", "-y", "@some/mcp-server"]
    }
  }
}
```

Everything the agent sends flows through the gate to the server and back, with
the checks applied in between.

### Policy

Drop in the same YAML `bastionsupply harden` emits:

```yaml
default: deny
allow:
  - get_weather
  - search_docs
deny:
  - run_command
# behavior knobs (defaults shown)
scan_tools: true            # scan tools/list
on_poisoned_tool: block     # drop poisoned tools from the listing
scan_results: true          # scan tool-call results
on_injected_result: block   # block results carrying injection
scrub_args: true            # scan tool-call arguments for secrets/PII
on_pii_arg: redact          # redact | block | warn
scrub_results: false        # scan tool-call RESULTS for secrets/PII (opt-in)
on_pii_result: redact       # redact | block | warn
result_inspector: static    # static | agentbastion  (deeper inspection)
inspector_fail: closed      # closed | open  (behavior if the inspector errors)
inspector_judge: false      # agentbastion: also use the Anthropic LLM judge
inspector_semantic: false   # agentbastion: also use the semantic detector

# per-tool overrides — any of the knobs above, scoped to one tool
tools:
  send_email:
    scrub_args: false       # the recipient email is the point; don't redact it
  fetch:
    on_injected_result: warn
```

So the pipeline is: **scan the server with bastionsupply → `harden` a policy →
run it live behind bastiongate.**

### `policy_version: 2` (bastiongate ≥ 0.8)

The same file format agentbastion reads: a shared core (`default`, `allow`, `deny`,
`rate_limits`, `detectors`) plus one block per tool. Gate's knobs move under `gate:`:

```yaml
policy_version: 2
default: deny
allow: [get_weather, search_docs]
detectors:
  bastion.exfil_action: off        # agentbastion kill switch, applied in deep-inspect
  bastion.dan_jailbreak: shadow    # runs, named in the gate trace, never blocks
gate:
  result_inspector: agentbastion   # required for bastion.* lines
  on_injected_result: block
  tools:
    fetch: {on_injected_result: warn}
```

- **Gate's own checks have no detector IDs: the knobs are the modes.**
  `off` = `scan_tools` / `scan_results` / `scrub_args` / `scrub_results: false`;
  `shadow` = `on_*: warn` (log only, let it through); `enforce` = `block` / `redact`.
  A `gate.*` line in `detectors:` is rejected with that mapping.
- `bastion.*` lines reach agentbastion's deep-inspect and need
  `result_inspector: agentbastion` plus `agentbastion ≥ 0.12` (the
  `bastiongateway[agentbastion…]` extras require it).
- Strict: unknown keys, typos (`scan_result`), bad actions (`wran`), a v1-style
  top-level knob in a v2 file, unknown detector IDs, or `bastion.*` lines without
  the agentbastion inspector **stop the gate at startup**, never silently.
- `default:` is required only alongside `allow` / `deny` / `rate_limits`; without
  them every tool is allowed. `rate_limits` is accepted and ignored (gate has no
  rate limiter). v1 files (no `policy_version`) load exactly as before.

### Flow guard

_bastiongate ≥ 0.9._

Every call can pass and the session can still leak. The GitHub MCP toxic flow:
the agent reads a **public issue** that carries an injection, reads a file from a
**private repo**, then opens a **pull request into the public repo** with the
private content. The flow guard tracks that sequence per session and catches the
egress call:

```bash
python -m bastiongate.demo
# bastiongate: WARN tainted egress: create_pull_request (untrusted from get_issue, private from get_file_contents) - set labels or on_tainted_egress: block
# ... -32005: bastiongate blocked tool call 'create_pull_request': ...
```

Each tool gets labels: `untrusted` (its result is attacker-reachable content),
`private` (its result is private data), `egress` (calling it sends data out). An
`egress` call after the session read both `untrusted` and `private` content is a
**tainted egress**.

- **Status: shadow.** Default `on_tainted_egress: warn` forwards the call, writes
  a `tainted_egress` trace event and prints one stderr line. `block` returns
  JSON-RPC error `-32005` naming the tools involved (never the content). The
  default moves to `block` only after the replay suite passes and a dogfood run
  over ≥ 20 real sessions shows 0 false blocks on benign flows.
- **Where labels come from** (first match wins): per-tool `labels` in the policy
  → built-in packs for common servers (exact tool names) → auto (bastionsupply
  capability categories; auto only ever adds `egress`). See what each tool gets:

  ```bash
  bastiongate labels                              # the built-in packs
  bastiongate labels --policy policy.yaml tools.json  # a tools/list dump
  ```

  | Pack | untrusted | private | egress |
  |---|---|---|---|
  | github | issue_read, list_issues, search_issues, pull_request_read, list_pull_requests, search_pull_requests, get_job_logs (+ legacy get_issue, get_issue_comments, get_pull_request, get_pull_request_comments) | get_file_contents, search_code, get_commit | issue_write, add_issue_comment, update_issue_comment, create_pull_request, update_pull_request, pull_request_review_write, add_comment_to_pending_review, add_reply_to_pull_request_comment, create_or_update_file, push_files, create_repository, fork_repository (+ legacy create_issue, update_issue) |
  | filesystem | | read_file, read_text_file, read_multiple_files | |
  | fetch | fetch | | fetch |
  | slack | slack_get_channel_history, slack_get_thread_replies | | slack_post_message, slack_reply_to_thread |
  | gmail | read_email, search_emails | read_email, search_emails | send_email |

- **GitHub same-repo rule.** A private read from the same `owner/repo` the egress
  call targets does not count, so the everyday "fix issue #N" flow (issue, file
  and PR in one repo) stays silent.
- Taint is set only by what the agent actually reads: a result the gate blocked
  (`-32002`) sets nothing, and secrets `scrub_results` redacted do not count as
  private. `private` also comes from credential-shaped secrets in a forwarded
  result (keys, tokens, JWTs), never from emails or card numbers.

```yaml
# v1 top-level (or under `gate:` in policy_version 2)
scan_flows: true            # false = off
on_tainted_egress: warn     # warn | block
label_packs: true           # false = explicit labels only
tools:
  internal_search: {labels: [private]}
  create_pull_request: {on_tainted_egress: block}   # per-tool action
```

#### Across servers (`taint_group`)

An MCP client runs one gate per server, so by default each gate sees only its own
server's taint and the common trifecta (read a page through `fetch`, a secret through
`filesystem`, send it out through `fetch`) passes. Give the gates the same `taint_group`
and each one also sees the others' taint:

```yaml
taint_group: auto     # the MCP client that spawned this gate; or a name, e.g. my-agent
```

- Rows live in one per-user SQLite file (`%LOCALAPPDATA%\bastiongate\taint.sqlite`,
  `$XDG_STATE_HOME/bastiongate/taint.sqlite`; override with `BASTIONGATE_STATE_DIR`),
  user-only on POSIX. Each row says `server:tool`, so warnings read
  `private from filesystem:read_file`; the trace event has `cross_server: true`.
- `auto` groups the gates one client spawned: on POSIX by process group, on Windows by the
  nearest ancestor process that is not a launcher (`py`, `uv`, `uvx`, `cmd`, a console-script
  shim, or the venv `python.exe` redirector right above a venv gate). The resolved group is in
  the trace (`taint_group` event). `auto` can merge two clients started from one shell job
  (POSIX) or miss a client that detaches its servers; a name set on every server entry is
  the reliable path (or env `BASTIONGATE_TAINT_GROUP`).
- Each gate keeps one row per tainting source (`server:tool`, at most 64 per gate), so one
  gate can never flood out the others' taint; past 64, its oldest private rows fold into one
  `server:*` row and its untrusted rows are never pushed out by private ones. A source read
  from several repos is never exempt.
- A live gate re-stamps its rows every minute while its own taint is live; rows not
  re-stamped for 3 minutes (the gate died, or its taint expired after 30 minutes idle) are
  ignored, so a crashed client's taint does not block the next session for long.
  `initialize` clears only this gate's own rows.
- A store problem never breaks the proxy: the check uses this gate's own taint. A lock held
  by another gate is retried on the next call (`taint_store_busy`); a corrupt or unwritable
  file prints one WARN, counts `taint_store_error` and, after 3 in a row, sharing is off for
  this gate.
- Off by default; stdio only (`run-http` refuses a `taint_group`: an HTTP gate serves
  many clients and must not pool their taint).

**Limits.** Taint lives per session per upstream server, in memory:
- a chain that crosses two MCP servers is detected only when their gates share a
  `taint_group`; a named group survives a client restart until its rows expire (30
  minutes idle), so a fresh session can inherit stale taint (warnings, not leaks);
- any local process running as the same user can write or clear rows in the store
  (it could equally read the secrets directly);
- a hostile MCP server in the group can make its gate write rows (by returning
  secret-shaped text it marks itself private; by returning an injection, untrusted), so
  it can cause warnings or blocks on the other servers' egress, never hide taint;
- stdio is one session for the process; over HTTP the key is `Mcp-Session-Id`, so
  a client that rotates it starts clean, and one that floods new ids can evict
  other sessions' taint (the `taint_evicted` metric counts it);
- an HTTP upstream that issues no session ids gets no flow guard (one warning);
- taint clears after 30 minutes with no calls, or on `initialize` (any client
  holding the session id can send one);
- auto labels come from the server's own tool descriptions, so a malicious server
  can word them to avoid `egress`: they are best-effort, packs and explicit
  `labels` are the reliable path;
- a tool with no label is never treated as egress, and auto labels need a
  `tools/list` first (packs and policy labels apply immediately);
- a private read counts only from its labelled tool or a credential-shaped secret
  in the forwarded result (text, resource and `resources/read` content); an injection
  inside a private read counts as untrusted only if the result inspector flags it;
- the very first egress call is checked before its own result taints the session,
  so exfiltration needs untrusted and private reads to have happened earlier;
- a tool call whose response takes longer than 5 minutes loses correlation: its
  result is still scanned (as `(unmatched)`) but not labelled.

Older gates silently ignore these keys: `bastionsupply doctor --policy policy.yaml`
warns when bastiongateway < 0.9 would read them.

### Encoded injection in results (`on_encoded_result`)

A payload hidden in base64, hex, binary, base32, ascii85/base85, Morse or escapes reads as
noise to a text filter but plainly to the model. The gate decodes each tool result
(bastioncorpus `variants`) and runs bastionsupply's `encoded-injection` check on the decoded
views.

```yaml
on_encoded_result: warn      # warn (default, shadow) | block (-32006)
decode_transforms: false     # true = also rot13 / leet / reversed / spaced views (results <= 64 KB)
tools:
  fetch: {on_encoded_result: block}
```

- `warn` forwards the result, logs `encoded_injection`, prints a WARN line, and counts the
  result as untrusted for the flow guard.
- `block` replaces it with error -32006.
- Encoded injections in `tools/list` definitions are warned about, never dropped.
- A tool definition over 1,000,000 characters, or every tool from the one that takes a
  `tools/list` page past 5,000,000 characters, is not scanned at all: it is handled like a
  poisoned tool (event `tools_list_oversize`). Tool scans read `description`, `title`,
  `annotations.title`, `outputSchema` and `inputSchema`.

**Limits:**
- Results over 1,000,000 characters are not decoded. They are refused under `block`, and only
  logged (`encoded_scan_skipped`) under `warn`.
- rot13, leetspeak and reversed text are decoded only with `decode_transforms: true` (off by
  default), and only on results up to 64 KB; a bigger result gets the run-based views only and
  the trace event `encoded_transforms_skipped`. Spaced-out letters are mostly missed (8% on
  the bench).
- Made-up ciphers can't be decoded by enumeration. Tool allow/deny lists and the flow guard
  are the controls encoding cannot bypass.

### Resource content (`scan_resources`)

Since 0.10 every result check reads all the text the model can see, not only `text` blocks:
- embedded `resource` blocks: `uri`, `text`, and `blob` base64-decoded (standard or url-safe,
  any MIME type except media: a text type decodes leniently, anything else counts when it is
  valid UTF-8; images other than SVG, audio, video, fonts, PDF and zip are never decoded);
- `resource_link` blocks (`uri`, `name`, `title`, `description`);
- `resources/read` responses (`contents[]`), checked under the pseudo tool name
  `resources/read`;
- an upstream JSON-RPC error on a `tools/call` or `resources/read` (`message` and `data`):
  clients show tool errors to the model;
- a response with no pending request (late past 5 minutes, duplicate id, unknown id): it is
  scanned under the pseudo tool name `(unmatched)` instead of passing through (trace event
  `response_unmatched`); an uncorrelated `tools/list` result is filtered like any other.

The `initialize` result's `instructions` and `serverInfo` (clients often put them in the
system prompt) are scanned like a tool definition: under `on_poisoned_tool: block` poisoned
`instructions` are removed and a poisoned `serverInfo` is replaced by `{"name": "upstream"}`
(event `instructions_poisoned`). A response the gate cannot inspect at all is replaced by
error -32002 (event `response_uninspectable`), never forwarded and never a crash. With
`scrub_results`, secrets in an upstream error message are redacted too.

The plain injection scan (-32002), the encoded scan (-32006), the PII scrub and the flow-guard
taint all see this text. Clean resources are forwarded unchanged.

```yaml
scan_resources: true           # default; false = the 0.9 behaviour (text blocks only)
tools:
  resources/read: {on_injected_result: warn}
```

`scan_resources: false` drops resource text from `tools/call` scans and lets `resources/read`
responses through entirely (no scan, scrub or taint), as in 0.9.

**Limits:**
- The PII scrub redacts `resource.text`, a blob that decodes to text (re-encoded after
  redaction) and `resource_link` name/title/description; media blobs and URIs are never
  rewritten. URIs are scanned both as sent and percent-decoded.
- Media blobs (images, audio, PDF...) are not inspected: the model receives them as media.

### Prompts and listings (`scan_prompts`)

Since 0.11 the gate also checks the other server text a client can hand the model:
- **`prompts/get`**: a prompt template is instructions by design ("always review...", "you
  must..."), so the full injection signatures would flag ordinary templates (98 of 1,015 real
  skill and agent files did). Only high-precision checks run: hidden or control unicode
  (allowed: an emoji joiner, a leading byte-order mark, ZWJ/ZWNJ between letters of scripts
  such as Devanagari or Persian), variation-selector steganography, a known bastioncorpus
  payload (matched after folding case, look-alike letters, markdown emphasis and trailing
  punctuation), or one hidden in an encoding. A hit is handled by `on_injected_result`
  (block: -32002, warn: forwarded and the session tainted), under the pseudo tool name
  `prompts/get`. A prompt over 1,000,000 characters is refused under block. A prompt result
  that also carries tool-result content (`content`, `contents`, `structuredContent`) or an
  error gets the full result checks instead.
- **`resources/list`, `resources/templates/list`, `prompts/list`**: each entry's name, title,
  description, uri and prompt-argument descriptions get the same checks as a tool definition.
  A poisoned, oversize or malformed entry is dropped under `on_poisoned_tool: block`
  (event `listing_poisoned`).

```yaml
scan_prompts: true              # default; false = neither is checked (0.10 behaviour)
tools:
  prompts/get: {on_injected_result: warn}
```

**Limits:** a prompt template carrying a new, never-seen injection phrase in plain text is not
flagged, and a known phrase with words inserted or reordered is missed: that is the price of
no false positives on real templates (0 of 1,015 skill and agent files flagged as prompts; the
one hit was stray byte-order marks and a zero-width space inside a CHANGELOG). With
`scan_resources: false`, resources embedded in prompts are not read. Phrase matching is
evaded by words joined with `_` or `-`, inserted commas or HTML tags, combining marks and
leetspeak. A `prompts/get` error, or a listing/prompt result that also carries `content`,
`structuredContent` or prompt `messages` it should not have, gets the full result checks,
so an error message with imperative wording ("Always call prompts/list first") can be
blocked like a tool error.

### Argument PII/secret scrub

On every `tools/call` the gate scans the arguments the agent is about to send
and, by default, **redacts** secrets/PII (`[REDACTED:<kind>]`) before they reach
the tool — API keys, AWS/GitHub/Slack tokens, private keys, JWTs, emails, SSNs,
and Luhn-valid card numbers. `on_pii_arg: block` refuses the call instead;
`warn` only logs. Only the *kind* is ever logged, never the value.

### OpenTelemetry spans (`--otel-out`, `--otel-endpoint`)

The gate can emit one OTel GenAI `execute_tool` span per `tools/call`, with its verdict, so
traces exist even when the agent framework exports none, and they plug into any OTLP backend
(Jaeger, Tempo, Honeycomb, Langfuse...) or straight into bastiontrace:

```bash
bastiongate run --otel-out spans.jsonl -- npx -y @modelcontextprotocol/server-filesystem ~/notes
bastiongate run --otel-endpoint http://127.0.0.1:4318 -- ...   # OTLP/HTTP collector
bastiontrace analyze --otel spans.jsonl --forbid send_email     # forensics on the gate's own spans
```

- Attributes: `gen_ai.operation.name=execute_tool`, `gen_ai.tool.name`, `gen_ai.tool.call.id`,
  `server.address`, `bastion.gate.verdict` (forwarded | warned | blocked | error),
  `bastion.gate.code`, `bastion.gate.checks` (the gate's trace events for the call),
  `bastion.gate.tainted_egress`.
- Content: by default only `bastion.gate.args_hmac` / `result_hmac` (HMAC-SHA256 with a
  per-process key: they correlate calls within one gate run, never across runs, and cannot
  be used to guess a short secret; arguments that differ only in redacted fields digest the
  same). With `--otel-content`, `gen_ai.tool.call.arguments` / `.result` (what the
  agent got, `structuredContent` included) are added with secret-named fields (`password`,
  `token`, `api_key`, `pin`...) and PII redacted, capped at 16 KB each; a blocked call also
  carries the server's text as `bastion.gate.upstream_result` (evidence, kept out of the
  result the agent "read"). Point `--otel-content` only at a collector you trust.
- One trace per gate session (stdio run, or HTTP `Mcp-Session-Id`).
- The file is OTLP/JSON, one export object per line (the Collector file exporter format),
  created user-only (0600). The endpoint gets OTLP/HTTP JSON at `/v1/traces`: https, or
  http on loopback only; no credentials, query or fragment in the URL; redirects never
  followed; `http_proxy` and friends ignored; headers from `OTEL_EXPORTER_OTLP_HEADERS`
  (validated at start).
- Export never blocks or breaks the proxy: spans queue (at most 10,000) and flush every
  2 seconds and at exit; a POST has a 10-second total deadline; a failed export drops that
  batch and prints one WARN. A server error is `verdict: error` even if it reuses a gate
  code; `blocked` means the gate replaced the response.

**Limits:** stdlib only, so no gRPC and no sampling; the client's own trace context is not
joined (MCP has no standard header for it); a process killed hard loses the last 2 seconds
of spans.

### HTTP transport

For MCP Streamable-HTTP servers, run the gate as an HTTP proxy instead:

```bash
bastiongate run-http --upstream http://127.0.0.1:8000/mcp --port 9000 --policy policy.yaml
```

Point your client at `http://127.0.0.1:9000/mcp`. JSON and SSE responses are
both gated. Binds `127.0.0.1` by default. Add `--auth-key KEY` (or env
`BASTIONGATE_PROXY_KEY`) to require an `X-Bastiongate-Key` header on every
request; the key is compared in constant time and never forwarded upstream.

`GET /__bastiongate/metrics` returns a JSON counter of what the gate has caught
(blocks by type, tools dropped, pending/session counts). It's auth-gated when a
key is set. In stdio mode the same counters are written to the trace at exit.

### Deeper result inspection (agentbastion)

`result_inspector: agentbastion` swaps the static signature scan for
[agentbastion](https://github.com/Rinkia/agentbastion)'s inbound Firewall,
composed of up to three tiers: heuristics (always), the semantic detector
(`inspector_semantic: true` + `BASTIONGATE_EMBED_URL`), and the Anthropic LLM
judge (`inspector_judge: true` + `ANTHROPIC_API_KEY`).

```bash
pip install "bastiongateway[agentbastion]"           # heuristic
pip install "bastiongateway[agentbastion-local]"     # + semantic (local model, no egress)
pip install "bastiongateway[agentbastion-semantic]"  # + semantic (remote embed endpoint)
pip install "bastiongateway[agentbastion-judge]"     # + LLM judge
```

The semantic detector needs an embedder, chosen by env:

- `BASTIONGATE_EMBED_MODEL` — a local sentence-transformers model (e.g.
  `all-MiniLM-L6-v2`). Result text never leaves the process. **Preferred.**
- `BASTIONGATE_EMBED_URL` — a self-hosted embeddings endpoint (result text is
  POSTed to it).

The model is loaded and its templates embedded **at gate startup** (not on the
first result), bounded by `BASTIONGATE_EMBED_INIT_TIMEOUT` (default 120s); set
`BASTIONGATE_EMBED_WARM=0` to defer. For air-gapped hosts, pre-cache the model
and set `HF_HUB_OFFLINE=1` — the first uncached load fetches from HuggingFace.

## Try it

```bash
bastiongate run --log gate.jsonl -- python examples/echo_server.py
```

The example server offers a poisoned tool and an injected result; the gate drops
the first and blocks the second. Watch `gate.jsonl`.

## Library

```python
from bastiongate import Gate, GatePolicy

gate = Gate(GatePolicy(deny={"run_command"}))
forward, reply = gate.handle_client_msg(msg)   # agent -> server
out = gate.handle_server_msg(response)          # server -> agent
```

`Gate` is a pure message transform — easy to embed or test.

## Security notes & limitations

- **Argument redaction can alter legitimate calls.** `on_pii_arg: redact`
  rewrites anything that looks like PII — including a recipient email a
  `send_email` tool actually needs. Scope it with a per-tool `scrub_args: false`
  or `on_pii_arg: warn` (see `tools:` above). Scrubbing is best-effort DLP:
  base64-encoded or field-split secrets can slip through.
- Result scrub covers both `text` content blocks and `structuredContent`, and
  the injection scan reads `structuredContent` too (not only text blocks).
- The HTTP listener throttles an IP after repeated auth failures (`429`), and
  correlation state is bounded per session (LRU eviction) with an idle TTL.
- **The HTTP proxy adds no auth of its own.** It binds `127.0.0.1` by default
  and passes the client's `Authorization` header through to the upstream. Do
  not bind a public interface without an auth layer in front.
- The HTTP proxy does not follow upstream redirects and only speaks
  `http`/`https` — a malicious upstream cannot bounce it to `file://` or an
  internal address.
- **JSON-RPC batch arrays** on the request side: a batch containing `tools/call`
  is rejected (`400`), since it would bypass tool policy and the flow guard; other
  batches (notifications) pass through, and responses are still scanned. Single
  messages — the normal case — are fully gated. Chunked request bodies (no `Content-Length`) are not supported.
- `inspector_fail: closed` (default) blocks a result if the inspector errors or
  times out. The LLM judge runs on every result (cost + latency); it has a
  `BASTIONGATE_JUDGE_TIMEOUT` (default 10s).

MIT.

Maintenance

ActivityMaintained
ResponsivenessNo issues