aw1-circuit-breaker
Deterministic AST security circuit breaker for inspecting and blocking unsafe code and tool calls in autonomous AI agents.
Inspect Python code AST for unsafe constructs (shell injection, dynamic eval, socket egress).
Validate external tool arguments for risky shell patterns, eval triggers, or socket connections.
Get current zero-trust AST containment rules status.
Integrate with MCP clients/agent frameworks to halt malicious tool dispatches before runtime.
Provides an adapter that drops into LangGraph agent execution nodes to guard tool dispatches before they run, raising a CircuitBreakerException when a tool call (e.g. a bash command with shell injection) is halted by AST containment rules.
AW-1 Circuit Breaker
Deterministic sub-0.2ms AST security circuit breaker for autonomous AI agents, tool dispatch pipelines, and MCP servers.
⚡ 10-Second Test Drive (Zero-Install)
Test the circuit breaker live in your terminal right now without installing anything:
uvx aw1-circuit-breaker --demoOr verify an arbitrary snippet:
uvx aw1-circuit-breaker verify "subprocess.run(['rm', '-rf', '/'])"AW-1 Circuit Breaker
Deterministic AST deconstruction and runtime circuit breaker preventing rogue shell escapes, dynamic eval(), and covert lateral socket egress in autonomous LLM tool calls.
Related MCP server: costwright
Installation
pip install -r requirements.txtRunning the Server
Run directly with Python:
python mcp_server.pyOr via uvx:
uvx mcp_server.pyMCP Client Configuration
{
"mcpServers": {
"aw1-circuit-breaker": {
"command": "python",
"args": ["mcp_server.py"]
}
}
}Exploit Interception in Action
AW-1 operates deterministically at the Abstract Syntax Tree (AST) level before code reaches Python's runtime execution frame. Probabilistic prompt-based guardrails fail under encoding obfuscation; AST inspection guarantees deterministic enforcement.
Interception Matrix
Attack Vector | Attacker Strategy | LLM Guardrail Result | AW-1 AST Circuit Breaker |
Dynamic Execution |
| Evades semantic filters | Tripped ( |
Shell Escapes |
| Masked as system task | Tripped ( |
Lateral Exfiltration |
| Disguised as HTTP fetch | Tripped ( |
Run the Showcase
python examples/exploit_showcase.pyRuntime Latency Telemetry
Deterministic security must not throttle agent execution. AW-1 operates with sub-millisecond AST parsing and microsecond-scale argument filtering across warm execution loops.
Sample Size: 1,000 synthetic iterations per test vector on Debian Python 3.13.
Evaluation Target | p50 Latency | p95 Latency | p99 Latency | Status |
AST Inspection (Benign Payload) | 0.0889 ms | 0.1298 ms | 0.1645 ms |
|
AST Inspection (Adversarial Exploit) | 0.0520 ms | 0.0919 ms | 0.1051 ms | Instant Breakout Halt |
Shell Argument Injection Guard | 0.0051 ms | 0.0096 ms | 0.0195 ms | Sub-20 µs Inspection |
Reproduce locally:
python benchmark.pyFramework Integration (LangGraph Example)
AW-1 drops directly into autonomous agent execution nodes prior to tool dispatch:
from examples.langgraph_adapter import execute_agent_action, CircuitBreakerException
# Guard autonomous agent tool dispatches
try:
execute_agent_action(bash_tool, "bash_tool", {"cmd": "cat log.txt; rm -rf /"})
except CircuitBreakerException as blocked:
print(f"Tool call halted: {blocked}")License
Licensed under the Apache License, Version 2.0. See LICENSE for details.
Available Tools
3 toolsget_containment_statusB
Returns the current state of zero-trust AST containment rules.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. 'Returns' implies a non-mutating read, but nothing is said about permissions/auth requirements, whether the status is live or cached, or any cost/latency considerations for an agent deciding to call it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or redundancy. It is appropriately sized, though it is so terse that the resource term goes unexplained.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value details need not be repeated. However, for a domain-specific concept like 'containment rules,' the description leaves the agent without enough context to know when this status matters relative to the sibling inspection tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline of 4 applies. No parameter-related gaps exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb ('Returns') and a specific domain resource ('zero-trust AST containment rules'), which goes beyond restating the name. It does not, however, differentiate itself from the sibling inspection tools or explain what 'containment rules' actually govern.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as inspect_python_code or inspect_tool_arguments. The agent must infer the use case entirely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_python_codeB
Deconstructs and analyzes Python code AST for unsafe constructs, shell injection, dynamic evaluation, and socket egress.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It discloses the types of unsafe constructs it looks for, but does not explicitly state whether the inspected code is executed, what permissions are needed, or whether the operation has side effects. Because an output schema exists, return-value explanation is not required, but the safety profile remains partially unclear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler. It immediately states the analysis scope and lists the detection categories, so every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has one required parameter and an output schema, the description covers the essential purpose and detection scope. It is slightly incomplete regarding sibling-tool routing and input-format expectations, but it does not need to explain return values because an output schema is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single required parameter. The description implies the input is Python source code suitable for AST analysis, which adds some meaning beyond the bare 'Code' title, but it does not clarify format expectations such as snippet vs. full file, size limits, or encoding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: deconstructs and analyzes Python code ASTs for unsafe constructs, shell injection, dynamic evaluation, and socket egress. It clearly distinguishes what the tool inspects, though it does not explicitly differentiate itself from the sibling tools inspect_tool_arguments or get_containment_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance on when to use this tool versus alternatives. It implies a security-inspection context but does not state when-not-to-use or name any sibling tool as an alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_tool_argumentsC
Validates arguments passed to external tools for risky shell patterns, eval triggers, or socket connections.
| Name | Required | Description | Default |
|---|---|---|---|
| tool_name | Yes | ||
| arguments_json | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does disclose what is inspected (shell patterns, eval triggers, socket connections). However, it never says what happens when a risky pattern is found — block, warn, or return a verdict — nor whether the tool has side effects, which is a notable gap for a validation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short lines with the verb and detection scope front-loaded and no filler. It is arguably too terse for the amount of structured information missing, but there is no wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but with zero annotation coverage and zero parameter descriptions the definition leaves an agent unsure about input format and about the consequence of a positive validation result. For a gating safety check, that is insufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the schema exposes only bare titles, so the description must compensate. It adds nothing about the format expected for tool_name or arguments_json (e.g., serialized JSON string), leaving both required parameters effectively undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (validates) plus the resource (arguments passed to external tools) and enumerates the threat categories it screens for: shell patterns, eval triggers, socket connections. This is distinct from inspect_python_code, though the description never names that sibling to sharpen the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies a pre-invocation safety check but never says when to call it, whether it is mandatory before invoking external tools, or how it relates to inspect_python_code and get_containment_status. No exclusions or alternatives are offered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
get_containment_status - First observed
inspect_python_code - First observed
inspect_tool_arguments
TDQS
Scored across 3 tools
Each tool targets a clearly distinct object: inspect_python_code analyzes Python source, inspect_tool_arguments validates external tool inputs, and get_containment_status reports rule state. No overlap in purpose, and descriptions make boundaries explicit.
All names follow a consistent snake_case verb_noun pattern (inspect_python_code, inspect_tool_arguments, get_containment_status). The minor verb difference (inspect vs get) is natural and does not break predictability.
Three tools is at the low end but reasonable for a focused security inspection server. Each tool earns its place, though a slightly richer set (e.g., rule management) could be justified.
The surface covers inspection and status checking but lacks operations to configure, enable/disable, or reset containment rules. For a circuit breaker, the absence of rule management is a notable gap.
Maintenance
Related MCP Connectors
Deterministic runtime safety for AI agents: scan PII, gate tool actions, verify LLM output.
Checked language for AI-written agent workflows: limits, cost and data flows known before running.
AgentGuard — 20-tool AI safety MCP: policy preflight, risk scoring, audit logging, rate limits.
Jailbreak-proof AI guardrails. Automated Reasoning SMT solver, not an LLM. ZK proofs included.
Related MCP Servers
- AlicenseAqualityBmaintenanceMCP server that vets LLM-emitted shell commands BEFORE execution — detects rm -rf nested deep in chains, package-manager glob removal (apt remove 'nvidia'), dd/mkfs filesystem destruction, chmod 777 / chown -R privilege blast, network-exfil via curl | bash, chained shutdown/reboot, git destructive ops. 30 detection rules across 8 families. Sub-second, local, free, MCP-native.358 PyPIMIT
- AlicenseAqualityDmaintenanceStatic worst-case token-budget analysis for LLM-agent workflows using AST analysis to identify certifiable, default-dependent, non-certifiable, and runaway units, with optional signed budget certificates.439 npmMIT
- AlicenseNot gradedqualityAmaintenanceA runtime gate for coding agents. Blocks the tool calls that wreck a repo (force-push main, rm -rf, secret exfiltration, CI wipe) and lets normal build and commit work through. Machine-checked git-branch core (z3); the rest is high-precision heuristics. Tested on 3,790 real CI commands, 0 false blocks.1MIT
- FlicenseNot gradedqualityCmaintenanceEnables autonomous AI agents and MCP clients to route tool calls through a zero-trust firewall that blocks dangerous shell commands, redacts secrets and PII, restricts sensitive file access, and logs verdicts. It provides an MCP interceptor decorator to protect MCP server tools with AST-based command injection checks and allow, block, or redact decisions.2-