Skip to main content
Glama
Rinkia

bastiongate

by Rinkia

bastiongate

MCP security gateway. An inline proxy that sits between an AI agent and its MCP servers and enforces security on every call:

  • scans tools/list and drops tools whose definitions carry prompt injection or hidden unicode (via bastionsupply)

  • enforces a tool allow/deny policy — the agent can only call what you permit

  • scans tool-call results and blocks any that carry indirect prompt injection before the agent ever reads them

  • logs every message as a JSONL trace for forensics

The runtime-enforcement leg of the bastion family:

tool

job

bastiongate

gate — enforce security inline on live MCP traffic

bastionsupply

scan an MCP server before you trust it

agentbastion

prevent — firewall around a running agent

bastionprobe

attack — pentest your agent with injections

bastiontrace

investigate — forensics on an agent trace

Install

pip install bastiongateway

(The PyPI distribution is bastiongateway; the import package and bastiongate CLI keep that name.)

Related MCP server: SafeNode MCP Gateway

Use

The gate is an MCP server to your agent, and a client to the real one. Point your MCP client's command at the gate and put the real server after --:

// mcp.json
{
  "mcpServers": {
    "docs": {
      "command": "bastiongate",
      "args": ["run", "--policy", "policy.yaml", "--log", "gate.jsonl",
               "--", "npx", "-y", "@some/mcp-server"]
    }
  }
}

Everything the agent sends flows through the gate to the server and back, with the checks applied in between.

Policy

Drop in the same YAML bastionsupply harden emits:

default: deny
allow:
  - get_weather
  - search_docs
deny:
  - run_command
# behavior knobs (defaults shown)
scan_tools: true            # scan tools/list
on_poisoned_tool: block     # drop poisoned tools from the listing
scan_results: true          # scan tool-call results
on_injected_result: block   # block results carrying injection
scrub_args: true            # scan tool-call arguments for secrets/PII
on_pii_arg: redact          # redact | block | warn
scrub_results: false        # scan tool-call RESULTS for secrets/PII (opt-in)
on_pii_result: redact       # redact | block | warn
result_inspector: static    # static | agentbastion  (deeper inspection)
inspector_fail: closed      # closed | open  (behavior if the inspector errors)
inspector_judge: false      # agentbastion: also use the Anthropic LLM judge
inspector_semantic: false   # agentbastion: also use the semantic detector

# per-tool overrides — any of the knobs above, scoped to one tool
tools:
  send_email:
    scrub_args: false       # the recipient email is the point; don't redact it
  fetch:
    on_injected_result: warn

So the pipeline is: scan the server with bastionsupply → harden a policy → run it live behind bastiongate.

policy_version: 2 (bastiongate ≥ 0.8)

The same file format agentbastion reads: a shared core (default, allow, deny, rate_limits, detectors) plus one block per tool. Gate's knobs move under gate::

policy_version: 2
default: deny
allow: [get_weather, search_docs]
detectors:
  bastion.exfil_action: off        # agentbastion kill switch, applied in deep-inspect
  bastion.dan_jailbreak: shadow    # runs, named in the gate trace, never blocks
gate:
  result_inspector: agentbastion   # required for bastion.* lines
  on_injected_result: block
  tools:
    fetch: {on_injected_result: warn}
  • Gate's own checks have no detector IDs: the knobs are the modes. off = scan_tools / scan_results / scrub_args / scrub_results: false; shadow = on_*: warn (log only, let it through); enforce = block / redact. A gate.* line in detectors: is rejected with that mapping.

  • bastion.* lines reach agentbastion's deep-inspect and need result_inspector: agentbastion plus agentbastion ≥ 0.12 (the bastiongateway[agentbastion…] extras require it).

  • Strict: unknown keys, typos (scan_result), bad actions (wran), a v1-style top-level knob in a v2 file, unknown detector IDs, or bastion.* lines without the agentbastion inspector stop the gate at startup, never silently.

  • default: is required only alongside allow / deny / rate_limits; without them every tool is allowed. rate_limits is accepted and ignored (gate has no rate limiter). v1 files (no policy_version) load exactly as before.

Flow guard

bastiongate ≥ 0.9.

Every call can pass and the session can still leak. The GitHub MCP toxic flow: the agent reads a public issue that carries an injection, reads a file from a private repo, then opens a pull request into the public repo with the private content. The flow guard tracks that sequence per session and catches the egress call:

python -m bastiongate.demo
# bastiongate: WARN tainted egress: create_pull_request (untrusted from get_issue, private from get_file_contents) - set labels or on_tainted_egress: block
# ... -32005: bastiongate blocked tool call 'create_pull_request': ...

Each tool gets labels: untrusted (its result is attacker-reachable content), private (its result is private data), egress (calling it sends data out). An egress call after the session read both untrusted and private content is a tainted egress.

  • Status: shadow. Default on_tainted_egress: warn forwards the call, writes a tainted_egress trace event and prints one stderr line. block returns JSON-RPC error -32005 naming the tools involved (never the content). The default moves to block only after the replay suite passes and a dogfood run over ≥ 20 real sessions shows 0 false blocks on benign flows.

  • Where labels come from (first match wins): per-tool labels in the policy → built-in packs for common servers (exact tool names) → auto (bastionsupply capability categories; auto only ever adds egress). See what each tool gets:

    bastiongate labels                              # the built-in packs
    bastiongate labels --policy policy.yaml tools.json  # a tools/list dump

    Pack

    untrusted

    private

    egress

    github

    issue_read, list_issues, search_issues, pull_request_read, list_pull_requests, search_pull_requests, get_job_logs (+ legacy get_issue, get_issue_comments, get_pull_request, get_pull_request_comments)

    get_file_contents, search_code, get_commit

    issue_write, add_issue_comment, update_issue_comment, create_pull_request, update_pull_request, pull_request_review_write, add_comment_to_pending_review, add_reply_to_pull_request_comment, create_or_update_file, push_files, create_repository, fork_repository (+ legacy create_issue, update_issue)

    filesystem

    read_file, read_text_file, read_multiple_files

    fetch

    fetch

    fetch

    slack

    slack_get_channel_history, slack_get_thread_replies

    slack_post_message, slack_reply_to_thread

    gmail

    read_email, search_emails

    read_email, search_emails

    send_email

  • GitHub same-repo rule. A private read from the same owner/repo the egress call targets does not count, so the everyday "fix issue #N" flow (issue, file and PR in one repo) stays silent.

  • Taint is set only by what the agent actually reads: a result the gate blocked (-32002) sets nothing, and secrets scrub_results redacted do not count as private. private also comes from credential-shaped secrets in a forwarded result (keys, tokens, JWTs), never from emails or card numbers.

# v1 top-level (or under `gate:` in policy_version 2)
scan_flows: true            # false = off
on_tainted_egress: warn     # warn | block
label_packs: true           # false = explicit labels only
tools:
  internal_search: {labels: [private]}
  create_pull_request: {on_tainted_egress: block}   # per-tool action

Across servers (taint_group)

An MCP client runs one gate per server, so by default each gate sees only its own server's taint and the common trifecta (read a page through fetch, a secret through filesystem, send it out through fetch) passes. Give the gates the same taint_group and each one also sees the others' taint:

taint_group: auto     # the MCP client that spawned this gate; or a name, e.g. my-agent
  • Rows live in one per-user SQLite file (%LOCALAPPDATA%\bastiongate\taint.sqlite, $XDG_STATE_HOME/bastiongate/taint.sqlite; override with BASTIONGATE_STATE_DIR), user-only on POSIX. Each row says server:tool, so warnings read private from filesystem:read_file; the trace event has cross_server: true.

  • auto groups the gates one client spawned: on POSIX by process group, on Windows by the nearest ancestor process that is not a launcher (py, uv, uvx, cmd, a console-script shim, or the venv python.exe redirector right above a venv gate). The resolved group is in the trace (taint_group event). auto can merge two clients started from one shell job (POSIX) or miss a client that detaches its servers; a name set on every server entry is the reliable path (or env BASTIONGATE_TAINT_GROUP).

  • Each gate keeps one row per tainting source (server:tool, at most 64 per gate), so one gate can never flood out the others' taint; past 64, its oldest private rows fold into one server:* row and its untrusted rows are never pushed out by private ones. A source read from several repos is never exempt.

  • A live gate re-stamps its rows every minute while its own taint is live; rows not re-stamped for 3 minutes (the gate died, or its taint expired after 30 minutes idle) are ignored, so a crashed client's taint does not block the next session for long. initialize clears only this gate's own rows.

  • A store problem never breaks the proxy: the check uses this gate's own taint. A lock held by another gate is retried on the next call (taint_store_busy); a corrupt or unwritable file prints one WARN, counts taint_store_error and, after 3 in a row, sharing is off for this gate.

  • Off by default; stdio only (run-http refuses a taint_group: an HTTP gate serves many clients and must not pool their taint).

Limits. Taint lives per session per upstream server, in memory:

  • a chain that crosses two MCP servers is detected only when their gates share a taint_group; a named group survives a client restart until its rows expire (30 minutes idle), so a fresh session can inherit stale taint (warnings, not leaks);

  • any local process running as the same user can write or clear rows in the store (it could equally read the secrets directly);

  • a hostile MCP server in the group can make its gate write rows (by returning secret-shaped text it marks itself private; by returning an injection, untrusted), so it can cause warnings or blocks on the other servers' egress, never hide taint;

  • stdio is one session for the process; over HTTP the key is Mcp-Session-Id, so a client that rotates it starts clean, and one that floods new ids can evict other sessions' taint (the taint_evicted metric counts it);

  • an HTTP upstream that issues no session ids gets no flow guard (one warning);

  • taint clears after 30 minutes with no calls, or on initialize (any client holding the session id can send one);

  • auto labels come from the server's own tool descriptions, so a malicious server can word them to avoid egress: they are best-effort, packs and explicit labels are the reliable path;

  • a tool with no label is never treated as egress, and auto labels need a tools/list first (packs and policy labels apply immediately);

  • a private read counts only from its labelled tool or a credential-shaped secret in the forwarded result (text, resource and resources/read content); an injection inside a private read counts as untrusted only if the result inspector flags it;

  • the very first egress call is checked before its own result taints the session, so exfiltration needs untrusted and private reads to have happened earlier;

  • a tool call whose response takes longer than 5 minutes loses correlation: its result is still scanned (as (unmatched)) but not labelled.

Older gates silently ignore these keys: bastionsupply doctor --policy policy.yaml warns when bastiongateway < 0.9 would read them.

Encoded injection in results (on_encoded_result)

A payload hidden in base64, hex, binary, base32, ascii85/base85, Morse or escapes reads as noise to a text filter but plainly to the model. The gate decodes each tool result (bastioncorpus variants) and runs bastionsupply's encoded-injection check on the decoded views.

on_encoded_result: warn      # warn (default, shadow) | block (-32006)
decode_transforms: false     # true = also rot13 / leet / reversed / spaced views (results <= 64 KB)
tools:
  fetch: {on_encoded_result: block}
  • warn forwards the result, logs encoded_injection, prints a WARN line, and counts the result as untrusted for the flow guard.

  • block replaces it with error -32006.

  • Encoded injections in tools/list definitions are warned about, never dropped.

  • A tool definition over 1,000,000 characters, or every tool from the one that takes a tools/list page past 5,000,000 characters, is not scanned at all: it is handled like a poisoned tool (event tools_list_oversize). Tool scans read description, title, annotations.title, outputSchema and inputSchema.

Limits:

  • Results over 1,000,000 characters are not decoded. They are refused under block, and only logged (encoded_scan_skipped) under warn.

  • rot13, leetspeak and reversed text are decoded only with decode_transforms: true (off by default), and only on results up to 64 KB; a bigger result gets the run-based views only and the trace event encoded_transforms_skipped. Spaced-out letters are mostly missed (8% on the bench).

  • Made-up ciphers can't be decoded by enumeration. Tool allow/deny lists and the flow guard are the controls encoding cannot bypass.

Resource content (scan_resources)

Since 0.10 every result check reads all the text the model can see, not only text blocks:

  • embedded resource blocks: uri, text, and blob base64-decoded (standard or url-safe, any MIME type except media: a text type decodes leniently, anything else counts when it is valid UTF-8; images other than SVG, audio, video, fonts, PDF and zip are never decoded);

  • resource_link blocks (uri, name, title, description);

  • resources/read responses (contents[]), checked under the pseudo tool name resources/read;

  • an upstream JSON-RPC error on a tools/call or resources/read (message and data): clients show tool errors to the model;

  • a response with no pending request (late past 5 minutes, duplicate id, unknown id): it is scanned under the pseudo tool name (unmatched) instead of passing through (trace event response_unmatched); an uncorrelated tools/list result is filtered like any other.

The initialize result's instructions and serverInfo (clients often put them in the system prompt) are scanned like a tool definition: under on_poisoned_tool: block poisoned instructions are removed and a poisoned serverInfo is replaced by {"name": "upstream"} (event instructions_poisoned). A response the gate cannot inspect at all is replaced by error -32002 (event response_uninspectable), never forwarded and never a crash. With scrub_results, secrets in an upstream error message are redacted too.

The plain injection scan (-32002), the encoded scan (-32006), the PII scrub and the flow-guard taint all see this text. Clean resources are forwarded unchanged.

scan_resources: true           # default; false = the 0.9 behaviour (text blocks only)
tools:
  resources/read: {on_injected_result: warn}

scan_resources: false drops resource text from tools/call scans and lets resources/read responses through entirely (no scan, scrub or taint), as in 0.9.

Limits:

  • The PII scrub redacts resource.text, a blob that decodes to text (re-encoded after redaction) and resource_link name/title/description; media blobs and URIs are never rewritten. URIs are scanned both as sent and percent-decoded.

  • Media blobs (images, audio, PDF...) are not inspected: the model receives them as media.

Prompts and listings (scan_prompts)

Since 0.11 the gate also checks the other server text a client can hand the model:

  • prompts/get: a prompt template is instructions by design ("always review...", "you must..."), so the full injection signatures would flag ordinary templates (98 of 1,015 real skill and agent files did). Only high-precision checks run: hidden or control unicode (allowed: an emoji joiner, a leading byte-order mark, ZWJ/ZWNJ between letters of scripts such as Devanagari or Persian), variation-selector steganography, a known bastioncorpus payload (matched after folding case, look-alike letters, markdown emphasis and trailing punctuation), or one hidden in an encoding. A hit is handled by on_injected_result (block: -32002, warn: forwarded and the session tainted), under the pseudo tool name prompts/get. A prompt over 1,000,000 characters is refused under block. A prompt result that also carries tool-result content (content, contents, structuredContent) or an error gets the full result checks instead.

  • resources/list, resources/templates/list, prompts/list: each entry's name, title, description, uri and prompt-argument descriptions get the same checks as a tool definition. A poisoned, oversize or malformed entry is dropped under on_poisoned_tool: block (event listing_poisoned).

scan_prompts: true              # default; false = neither is checked (0.10 behaviour)
tools:
  prompts/get: {on_injected_result: warn}

Limits: a prompt template carrying a new, never-seen injection phrase in plain text is not flagged, and a known phrase with words inserted or reordered is missed: that is the price of no false positives on real templates (0 of 1,015 skill and agent files flagged as prompts; the one hit was stray byte-order marks and a zero-width space inside a CHANGELOG). With scan_resources: false, resources embedded in prompts are not read. Phrase matching is evaded by words joined with _ or -, inserted commas or HTML tags, combining marks and leetspeak. A prompts/get error, or a listing/prompt result that also carries content, structuredContent or prompt messages it should not have, gets the full result checks, so an error message with imperative wording ("Always call prompts/list first") can be blocked like a tool error.

Argument PII/secret scrub

On every tools/call the gate scans the arguments the agent is about to send and, by default, redacts secrets/PII ([REDACTED:<kind>]) before they reach the tool — API keys, AWS/GitHub/Slack tokens, private keys, JWTs, emails, SSNs, and Luhn-valid card numbers. on_pii_arg: block refuses the call instead; warn only logs. Only the kind is ever logged, never the value.

OpenTelemetry spans (--otel-out, --otel-endpoint)

The gate can emit one OTel GenAI execute_tool span per tools/call, with its verdict, so traces exist even when the agent framework exports none, and they plug into any OTLP backend (Jaeger, Tempo, Honeycomb, Langfuse...) or straight into bastiontrace:

bastiongate run --otel-out spans.jsonl -- npx -y @modelcontextprotocol/server-filesystem ~/notes
bastiongate run --otel-endpoint http://127.0.0.1:4318 -- ...   # OTLP/HTTP collector
bastiontrace analyze --otel spans.jsonl --forbid send_email     # forensics on the gate's own spans
  • Attributes: gen_ai.operation.name=execute_tool, gen_ai.tool.name, gen_ai.tool.call.id, server.address, bastion.gate.verdict (forwarded | warned | blocked | error), bastion.gate.code, bastion.gate.checks (the gate's trace events for the call), bastion.gate.tainted_egress.

  • Content: by default only bastion.gate.args_hmac / result_hmac (HMAC-SHA256 with a per-process key: they correlate calls within one gate run, never across runs, and cannot be used to guess a short secret; arguments that differ only in redacted fields digest the same). With --otel-content, gen_ai.tool.call.arguments / .result (what the agent got, structuredContent included) are added with secret-named fields (password, token, api_key, pin...) and PII redacted, capped at 16 KB each; a blocked call also carries the server's text as bastion.gate.upstream_result (evidence, kept out of the result the agent "read"). Point --otel-content only at a collector you trust.

  • One trace per gate session (stdio run, or HTTP Mcp-Session-Id).

  • The file is OTLP/JSON, one export object per line (the Collector file exporter format), created user-only (0600). The endpoint gets OTLP/HTTP JSON at /v1/traces: https, or http on loopback only; no credentials, query or fragment in the URL; redirects never followed; http_proxy and friends ignored; headers from OTEL_EXPORTER_OTLP_HEADERS (validated at start).

  • Export never blocks or breaks the proxy: spans queue (at most 10,000) and flush every 2 seconds and at exit; a POST has a 10-second total deadline; a failed export drops that batch and prints one WARN. A server error is verdict: error even if it reuses a gate code; blocked means the gate replaced the response.

Limits: stdlib only, so no gRPC and no sampling; the client's own trace context is not joined (MCP has no standard header for it); a process killed hard loses the last 2 seconds of spans.

HTTP transport

For MCP Streamable-HTTP servers, run the gate as an HTTP proxy instead:

bastiongate run-http --upstream http://127.0.0.1:8000/mcp --port 9000 --policy policy.yaml

Point your client at http://127.0.0.1:9000/mcp. JSON and SSE responses are both gated. Binds 127.0.0.1 by default. Add --auth-key KEY (or env BASTIONGATE_PROXY_KEY) to require an X-Bastiongate-Key header on every request; the key is compared in constant time and never forwarded upstream.

GET /__bastiongate/metrics returns a JSON counter of what the gate has caught (blocks by type, tools dropped, pending/session counts). It's auth-gated when a key is set. In stdio mode the same counters are written to the trace at exit.

Deeper result inspection (agentbastion)

result_inspector: agentbastion swaps the static signature scan for agentbastion's inbound Firewall, composed of up to three tiers: heuristics (always), the semantic detector (inspector_semantic: true + BASTIONGATE_EMBED_URL), and the Anthropic LLM judge (inspector_judge: true + ANTHROPIC_API_KEY).

pip install "bastiongateway[agentbastion]"           # heuristic
pip install "bastiongateway[agentbastion-local]"     # + semantic (local model, no egress)
pip install "bastiongateway[agentbastion-semantic]"  # + semantic (remote embed endpoint)
pip install "bastiongateway[agentbastion-judge]"     # + LLM judge

The semantic detector needs an embedder, chosen by env:

  • BASTIONGATE_EMBED_MODEL — a local sentence-transformers model (e.g. all-MiniLM-L6-v2). Result text never leaves the process. Preferred.

  • BASTIONGATE_EMBED_URL — a self-hosted embeddings endpoint (result text is POSTed to it).

The model is loaded and its templates embedded at gate startup (not on the first result), bounded by BASTIONGATE_EMBED_INIT_TIMEOUT (default 120s); set BASTIONGATE_EMBED_WARM=0 to defer. For air-gapped hosts, pre-cache the model and set HF_HUB_OFFLINE=1 — the first uncached load fetches from HuggingFace.

Try it

bastiongate run --log gate.jsonl -- python examples/echo_server.py

The example server offers a poisoned tool and an injected result; the gate drops the first and blocks the second. Watch gate.jsonl.

Library

from bastiongate import Gate, GatePolicy

gate = Gate(GatePolicy(deny={"run_command"}))
forward, reply = gate.handle_client_msg(msg)   # agent -> server
out = gate.handle_server_msg(response)          # server -> agent

Gate is a pure message transform — easy to embed or test.

Security notes & limitations

  • Argument redaction can alter legitimate calls. on_pii_arg: redact rewrites anything that looks like PII — including a recipient email a send_email tool actually needs. Scope it with a per-tool scrub_args: false or on_pii_arg: warn (see tools: above). Scrubbing is best-effort DLP: base64-encoded or field-split secrets can slip through.

  • Result scrub covers both text content blocks and structuredContent, and the injection scan reads structuredContent too (not only text blocks).

  • The HTTP listener throttles an IP after repeated auth failures (429), and correlation state is bounded per session (LRU eviction) with an idle TTL.

  • The HTTP proxy adds no auth of its own. It binds 127.0.0.1 by default and passes the client's Authorization header through to the upstream. Do not bind a public interface without an auth layer in front.

  • The HTTP proxy does not follow upstream redirects and only speaks http/https — a malicious upstream cannot bounce it to file:// or an internal address.

  • JSON-RPC batch arrays on the request side: a batch containing tools/call is rejected (400), since it would bypass tool policy and the flow guard; other batches (notifications) pass through, and responses are still scanned. Single messages — the normal case — are fully gated. Chunked request bodies (no Content-Length) are not supported.

  • inspector_fail: closed (default) blocks a result if the inspector errors or times out. The LLM judge runs on every result (cost + latency); it has a BASTIONGATE_JUDGE_TIMEOUT (default 10s).

MIT.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Open-source MCP proxy that enforces security policies, content scanning, and audit logging between AI agents and tool servers
    25
    AGPL 3.0
  • A
    license
    A
    quality
    A
    maintenance
    An MCP proxy that enforces policy on every tool call, blocking or flagging actions before they reach downstream MCP servers.
    1
    31 npm
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enforces deterministic security policies on Model Context Protocol traffic between agents and remote MCP servers, including request validation, signed human approval, response-side credential blocking, prompt-injection flagging, and privacy-minimized auditing.
    -
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables safe use of any MCP server by proxying and live-scanning all tool requests and responses, blocking or redacting poison descriptions, indirect prompt injection, malicious arguments, and unauthorized destinations.
    -