AgentShield
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@AgentShieldCheck if merging PR #42 in production is allowed or needs review."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
AgentShield
A policy-driven decision engine for AI agents and AI-native applications.
AgentShield decides whether an action an agent wants to take should be allowed, sent for review, or denied — deterministically, based on rules you define. AgentShield decides; the agent/application executes. It does not need to sit between an agent and a tool, an MCP server, or anything else — it only needs to be consulted before the action happens:
AI Agent
|
| evaluate action
v
AgentShield
|
| ALLOW / REVIEW / DENY
v
AI Agent
|
| execute if permitted
v
Tool / MCP / API / Actiondecision = shield.evaluate(request)Scope note: in this phase, AgentShield is a decision engine, not an enforcement mechanism. It cannot, by itself, technically stop a malicious or misbehaving agent from ignoring its answer and executing the action anyway — that requires actually sitting in the execution path (as the MCP Gateway/server below optionally does) or another enforcement integration. Making that guarantee robust across execution paths is future work.
Why it exists
AI agents increasingly want to take actions with real-world side effects: merging pull requests, deleting resources, reading secrets, moving money. Provider SDKs and tool protocols like MCP give you the mechanism to take these actions, but not a consistent, auditable way to decide whether a given action should be allowed. AgentShield is that decision layer — independent of any specific LLM provider, agent framework, or tool protocol.
Related MCP server: signet-eval
Two ways to use it
Python SDK / Library (primary) — embed the Core decision engine directly in your own agent/application code and consult it before executing an action. See
examples/sdk_example.pyand "Minimal example" below. No MCP dependency required.MCP adapter (optional integration) —
agentshield.mcpis one integration built on top of the same Core: it places AgentShield in the execution path between an MCP client and a downstream MCP server, so it can also enforce (not just advise on) ALLOW/REVIEW/DENY for that path. See "MCP adapter" below. This is not what AgentShield is — it's one way to use it.
Architecture
AgentShield
│
┌──────┴──────┐
│ Core │ ← the decision engine
│ │
│ Policy │
│ Decision │
│ Risk │
│ Audit │
│ Providers │
└──────┬───────┘
│
┌──────────┴──────────┐
▼ ▼
Python SDK MCP adapter ← agentshield.mcp (optional)
│ │
Developers AI Agents / MCPThe Core (agentshield.decision, .policy, .engine, .risk,
.audit) is a small, dependency-light Python library. It is deterministic:
it never calls an external service, never calls an LLM, and never executes
the action it decides on — it only returns a Decision. Given a policy and
a request, it always returns the same decision. The Core has no
dependency on MCP (see tests/test_core_independence.py) and remains
fully usable on its own.
Core
^
|
MCP adapter (agentshield.mcp)The dependency direction is one-way: the Core has no idea MCP, or anything else that might call it, exists.
MCP is one adapter among others
agentshield.mcp (documented in full below) is an optional integration
that happens to place AgentShield in the execution path for MCP traffic
specifically, so it can enforce rather than just advise for that path:
AI Agent / MCP Client
|
| MCP tool call
v
+----------------------+
| AgentShield |
| MCP adapter |
+----------+-----------+
|
v
DecisionEngine
|
+------+------+
| |
ALLOW REVIEW / DENY
| |
v +-----> approval / block
Downstream MCP
ServerAgentShield works with existing MCP servers without requiring those servers
to be modified, and preserves their tool definitions and arguments as-is —
it is a policy layer, not a schema transformation layer. But this is one
possible caller of the Core, not a requirement: the same DecisionEngine
works identically for an HTTP API call, a shell command, a workflow step,
or anything else expressed as a DecisionRequest.
Minimal example (Core only, no MCP)
from agentshield import DecisionEngine, DecisionRequest, Outcome, Policy
policy = Policy.from_yaml("examples/policy.yaml")
shield = DecisionEngine(policy)
decision = shield.evaluate(
DecisionRequest(
action="merge_pull_request",
actor="agent",
server="github",
arguments={"repo": "acme/app", "pull_request": 42},
context={"environment": "production"},
)
)
if decision.outcome == Outcome.DENY:
... # the agent must not execute the action
elif decision.outcome == Outcome.REVIEW:
... # the agent must route this through its own approval flow first
else:
... # the agent may proceed to actually perform the action, however
# it chooses to (MCP, a direct API call, a CLI, ...)
print(decision.outcome) # Outcome.REVIEW
print(decision.allowed) # False
print(decision.reason) # "Production merges require human approval."
print(decision.rule) # "production-merge"See examples/sdk_example.py for a runnable
version of this.
Breaking rename: what was
AuthorizationEngine/AuthorizationRequestis nowDecisionEngine/DecisionRequest, reflecting that this is a general-purpose decision engine, not an MCP-specific authorization layer.DecisionRequest's tool/action field was renamed fromtooltoaction. The old class names remain importable as deprecated aliases (AuthorizationEngine is DecisionEngine, etc.) sofrom agentshield import AuthorizationEnginestill works, but any code constructing the request withtool=...must change toaction=....PolicyRule.tool(the policy YAML field) is unchanged — existing policy files keep working as-is.
Using AgentShield from an Agent
The pattern above is the whole integration surface: a Python agent (with
or without a framework) calls shield.evaluate(request) directly, as a
decision service/library, before doing anything else. AgentShield is
not an execution proxy — it never sits between the agent and whatever it
would eventually call (a tool, an MCP server, a deploy script). It only
needs to be asked first:
decision = shield.evaluate(request)
if decision.outcome == Outcome.DENY:
stop()
elif decision.outcome == Outcome.REVIEW:
review()
else:
execute()See examples/agent_demo.py for a complete,
runnable version of this, as three small scenarios against one policy.
Policy knows the rules. Jev evaluates the situation:
Deterministic DENY —
delete_databasein production. Policy alone settles this; Jev is never even consulted, and could not override it if it were.Policy ALLOW, Jev recommends REVIEW — a staging deploy, which policy permits outright — but the situation (a large database migration, late on a Friday) is one a semantic evaluator can flag as worth a second look, even though nothing technically forbids it.
Policy ALLOW, Jev agrees — the same staging deploy rule, but a small, routine, Tuesday-morning change: nothing here should raise a flag, and (with real Jev configured) nothing does.
python examples/agent_demo.py # deterministic policy only
pip install -e ".[jev]"
TYPESAFE_API_KEY=... python examples/agent_demo.py # + real Jev semantic evaluationWithout TYPESAFE_API_KEY set, every scenario prints With Jev: not evaluated (...) and falls back to the deterministic-only result — the
demo never fabricates a semantic verdict. This is a decision, not an
enforcement mechanism or a security guarantee: the simulated
🚀 ... executed. line only ever prints when the returned Decision
actually permits it, but AgentShield itself never executes anything —
the agent remains responsible for that.
Using AgentShield with an AI Agent
The Agent decides what it wants to do. AgentShield evaluates that proposed action before execution. AgentShield does not replace the Agent's planning or decision-making — it provides a decision boundary between an Agent's proposed action and its execution.
User
|
v
AI Agent
|
| proposes action
v
AgentShield
|-- Policy
+-- Jev
|
v
ALLOW / REVIEW / DENY
|
v
AI Agent
|
v
Tool / API / ActionThe Agent chooses the action — from a natural-language request, using an LLM (or the deterministic stand-in below), independently, with no hard-coded mapping from request to action.
Policy enforces deterministic constraints.
Jev evaluates semantic suitability (only when policy alone would ALLOW).
The Agent executes only after ALLOW.
AgentShield is not an execution proxy. It does not need to sit between the Agent and a tool — nothing here talks MCP, and MCP is not required for this integration at all (it remains one optional adapter among others, see above).
Provider-agnostic: the Agent never touches an SDK
The LLM step is behind a small, swappable interface, not hard-coded to OpenAI — the dependency direction is:
LLM Provider
|
| proposes
v
Agent
|
| asks for a decision
v
AgentShield
|
| ALLOW / REVIEW / DENY
v
Agent
|
| executes only after ALLOW
v
Actionagentshield.agent defines the provider-neutral vocabulary — ProposedAction
(action/arguments/context, adapted directly onto the existing
DecisionRequest with build_decision_request(), no parallel abstraction)
and the AgentProvider interface (propose_action(user_request) -> ProposedAction). It has no dependency on any specific provider SDK.
Three implementations ship: agentshield.providers.anthropic.AnthropicProvider
(Claude Messages API, native structured output), agentshield.providers.openai.OpenAIProvider
(chat completions, JSON mode), and agentshield.providers.gemini.GeminiProvider
(Gemini structured JSON output). Each alone creates its own client, reads
its own API key (ANTHROPIC_API_KEY / OPENAI_API_KEY / GEMINI_API_KEY),
and converts the response; the agent loop never imports
anthropic/openai/google.genai, never sees an API key, and never sees
a provider response object. The provider does not know about
AgentShield. AgentShield does not know about Anthropic, OpenAI, or
Gemini. The agent connects the two.
If a configured provider's call fails, it raises ProviderError — the
agent does not silently fall back to a different provider; it fails
safely and does not execute.
examples/real_agent_demo.py is a complete,
runnable version of this, built around a small Agent class that connects
a provider's proposal to AgentShield's decision to execution:
class Agent:
def handle(self, user_request):
proposal = self.provider.propose_action(user_request) # Agent decides
request = build_decision_request(proposal)
decision = self.shield.evaluate(request) # AgentShield evaluates
if decision.outcome == Outcome.ALLOW:
self._execute(proposal) # only after ALLOWIt gives the Agent three plain natural-language requests (e.g. "Delete
the production database.") — nothing in the demo states what action each
one means; the Agent (via a real LLM provider, or DeterministicDemoProvider
below when no key is set) independently derives the structured action,
arguments, and context each time, and AgentShield's policy/Jev then
determine what happens. Execution goes through a tiny local tool registry
(deploy, delete_database, search_repository) that only prints what
it would have done — no MCP, no real side effects. DeterministicDemoProvider —
implementing the exact same AgentProvider interface a real provider
does — stands in for the LLM step when no key is set, clearly labeled and
never pretending to be a real LLM call; it derives the action from facts
actually present in the request text, the same way a real provider would,
just without general language understanding. When multiple keys are set,
Anthropic is tried first, then OpenAI, then Gemini.
pip install -e ".[agent-demo,jev]"
export ANTHROPIC_API_KEY=... # tried first; never committed, never printed
export OPENAI_API_KEY=... # tried next; never committed, never printed
export GEMINI_API_KEY=... # tried last; never committed, never printed
export TYPESAFE_API_KEY=... # for semantic evaluation; never committed, never printed
python examples/real_agent_demo.pyAll are optional and independent: without a provider key, the
deterministic provider is used instead; without TYPESAFE_API_KEY,
semantic evaluation is skipped (also clearly labeled) and the
deterministic-only decision is used.
Example policy
rules:
- name: block-secret-access
tool: "*.read_secret"
outcome: deny
risk: critical
reason: "Access to secrets is blocked."
- name: production-merge
server: github
tool: merge_pull_request
context:
environment: production
outcome: review
risk: high
reason: "Production merges require human approval."Rules support actor, server, tool, and context as optional matching
constraints (all must be unset or match for a rule to apply). name,
outcome, and risk are required on every rule; reason is optional. Tool
names support */? wildcards (e.g. "*.delete_*"); actor and server
currently require exact matches.
See examples/policy.yaml for a fuller example, and
tests/test_engine.py for the precedence rules worked out in detail.
ALLOW / REVIEW / DENY semantics
Outcome |
| Meaning |
|
| The action may proceed. |
|
| The action must not proceed automatically; it needs human (or other) approval. |
|
| The action must not proceed, period. |
If no rule in the policy matches a request, the Core defaults to ALLOW
with risk = LOW and confidence = 1.0. This default is intentionally easy
to change in a future phase (e.g. a "default deny" mode) but is not
configurable yet.
Rule precedence
When multiple rules match a request, the Core picks exactly one, using this
deterministic order (implemented in engine.py):
Exact tool match beats wildcard tool match beats no tool constraint.
Within the same tier, the rule with more context constraints wins.
Within a further tie, the rule with more specific
actor/serverconstraints wins.Within a full tie, the earlier rule in the policy file wins.
This ordering depends only on each rule's own fields and its position in the policy — never on dictionary iteration order — so the same policy and request always produce the same decision.
Deterministic DENY is authoritative
A deterministic DENY must never be overridden by another layer. An
optional semantic evaluator (see "Semantic evaluation" below) may add
context or additional review — but it cannot flip a policy-level DENY
into an ALLOW, and DecisionEngine never even consults it once policy
has already said DENY. This holds both in the Core (the engine's
precedence rules operate purely over policy rules, and nothing exposes a
way to override a returned Decision) and in the MCP Gateway (a DENY is
never sent to an ApprovalProvider and never reaches the downstream
server). Any layer that wants to add a "second opinion" must be additive
(e.g. escalating ALLOW to REVIEW), never permissive.
Semantic evaluation (optional, via Jev)
Deterministic policy answers "is this technically permitted?" DecisionEngine
can optionally also ask a semantic question: "given the action, its
arguments, the current context, and the applicable policy, is this
actually a sensible decision?" — a technically-permitted action can
still be a bad idea (a database migration proposed for Friday evening in
production, say).
Action + Context + Applicable Policy
↓
Jev
↓
semantic assessment
↓
AgentShieldagentshield.jev.JevSemanticEvaluator implements this using
Jev, TypeSafe's "System One" model, via the
official typesafe-sdk package (optional dependency: pip install -e ".[jev]",
export TYPESAFE_API_KEY=...). It asks a single structured Choice
question (good / review / bad) with only the relevant slice of
policy for that one request as context — never the whole policy file.
from agentshield import DecisionEngine, DecisionRequest, Policy
from agentshield.jev import JevSemanticEvaluator
shield = DecisionEngine(Policy.from_dict({"rules": []}), semantic_evaluator=JevSemanticEvaluator())
decision = shield.evaluate(
DecisionRequest(
action="deploy",
actor="release-agent",
context={
"environment": "production",
"time": "friday_evening",
"database_migration": True,
},
)
)Deterministic policy remains authoritative, and the semantic evaluator
is only ever consulted for an ALLOW:
DENY -> final; Jev is never consulted
REVIEW -> final; Jev is never consulted (policy already asked for human
attention -- there is nothing more for semantic judgment to add)
ALLOW -> Jev is consulted, and may escalate to REVIEW -- never to DENYDecision.confidence is set from Jev's own confidence score when a
semantic evaluation ran (all other decisions keep confidence = 1.0, as
before).
This is entirely optional: agentshield.jev is never imported by the Core
or by DecisionEngine itself (see tests/test_core_independence.py), and
Jev is not an enforcement mechanism or a security guarantee — it is one
more input into a decision the agent/application is still responsible for
acting on. See examples/jev_example.py.
MCP adapter
An optional integration, not the definition of AgentShield: agentshield.mcp
puts a policy-enforcement proxy between an MCP client (an agent) and a
downstream MCP server launched locally over stdio, using the official
MCP Python SDK.
import asyncio
from agentshield.mcp import GatewayConfig, MCPGateway
async def main():
config = GatewayConfig.from_yaml("examples/gateway.yaml")
gateway = MCPGateway.from_config(config)
async with gateway:
tools = await gateway.list_tools() # unmodified downstream tool defs
result = await gateway.call_tool("echo", {"message": "hi"})
print(result.decision.outcome, result.executed)
asyncio.run(main())Install the optional MCP dependency first: pip install -e ".[mcp]".
Configuration
server:
name: demo
downstream:
command: python
args:
- examples/mcp_server.py
context:
environment: production
policy:
path: examples/mcp_policy.yaml
actor: agentdownstream describes how to launch the downstream MCP server as a local
subprocess. context is static context merged into every
DecisionRequest built by this gateway (so a rule like context: {environment: production} works through the gateway exactly as it does in
the Core). actor defaults to "agent".
Request mapping
Every intercepted tools/call is mapped onto the Core's generic
DecisionRequest like this — this is the adapter's one job, translating
MCP's vocabulary into the Core's generic vocabulary:
MCP call |
|
configured |
|
configured |
|
MCP tool name |
|
MCP tool arguments, unmodified |
|
configured static |
|
ALLOW / REVIEW / DENY at the gateway
ALLOW — the call is forwarded to the downstream server unmodified; its result is returned to the caller exactly as the downstream server produced it — including a downstream-reported error (
CallToolResult.is_error=True, e.g. "file not found"). That is a normal MCP result, not a gateway failure, so it is never turned into a Python exception; only a call that could not be carried out at all (a broken connection, a malformed protocol response) raisesDownstreamToolError.DENY — the downstream server is never called. The gateway returns a
GatewayCallResultwithexecuted=Falseand the decision (rule, risk, reason) attached — no arguments or internal details are leaked back.REVIEW — the call is not executed automatically. It is resolved through an
ApprovalProvider:from agentshield.mcp import ApprovalProvider, ApprovalResult class MyApprovalProvider(ApprovalProvider): async def request_approval(self, request, decision) -> ApprovalResult: ... # ask Slack / a web UI / a human, then: return ApprovalResult(approved=True, approver="alice")Two minimal implementations ship out of the box:
CallbackApprovalProvider(wraps a sync or async callable — the main building block for tests and programmatic integrations) andConsoleApprovalProvider(prompts a human at the terminal; demo use only). If a REVIEW decision is reached and noApprovalProvideris configured, the gateway fails closed and raisesApprovalProviderRequiredError— it never guesses.
Audit
The gateway reuses the Core's AuditLog — every intercepted call, of every
outcome, produces one AuditEvent. Two optional, backward-compatible fields
were added to AuditEvent for the gateway's use: approval_required and
approval_outcome ("approved" / "rejected" / None). Core-only usage is
unaffected; these default to False / None.
Error handling
agentshield.mcp defines its own exception hierarchy (GatewayError and
subclasses) for gateway/transport failures, distinct from Core errors like
PolicyError: DownstreamConnectionError, DownstreamToolError,
AuthorizationEvaluationError (the engine itself failed — fails closed,
never forwards), ApprovalProviderRequiredError, ApprovalProviderError,
GatewayConfigError. Downstream errors are never swallowed.
Running AgentShield as a real MCP server
agentshield.mcp.server is a thin adapter that exposes an MCPGateway as a
real, upstream-facing MCP server over stdio, using the official MCP SDK's
Server/stdio_server. This is the case where the MCP adapter actually
sits in the execution path and can enforce, not just advise. It contains
no decision logic of its own — every tools/call is routed straight
through the same MCPGateway.call_tool() used by the library form above:
MCP Client / AI Agent
|
| MCP / stdio
v
+----------------------+
| AgentShield |
| MCP Server | agentshield.mcp.server — protocol only
+----------+-----------+
|
v
MCPGateway agentshield.mcp.gateway — MCP adapter
|
v
DecisionEngine agentshield Core — MCP-agnostic
|
+------+------+------+
| | |
ALLOW REVIEW DENY
| | |
| approval STOP
| |
+-------------+
|
v
DownstreamMCPProxy -> Downstream MCP ServerStart it:
pip install -e ".[mcp]"
python -m agentshield.mcp.server --config examples/gateway.yamlMCP protocol traffic uses stdout; all diagnostics go to stderr via the
standard logging module, so stdout stays clean for the protocol. The
downstream connection is established once at startup and kept alive for the
life of the upstream session; it is always closed on shutdown, including on
error, so no subprocess is leaked.
No ApprovalProvider is wired up by this CLI launcher. stdin/stdout in
this process are owned by the MCP protocol stream, so the interactive
ConsoleApprovalProvider cannot be used here — a REVIEW decision therefore
fails closed (the client gets a clear "no approval provider configured"
result, and nothing is forwarded downstream). To approve REVIEW calls,
embed AgentShieldMCPServer directly and pass a CallbackApprovalProvider
backed by Slack, a web UI, a queue, etc. — see
tests/test_mcp_server.py for a worked example.
Using it from Claude Desktop / Cursor / other MCP clients
Add it to the client's MCP server configuration, e.g. Claude Desktop's
claude_desktop_config.json:
{
"mcpServers": {
"agentshield": {
"command": "python",
"args": [
"-m", "agentshield.mcp.server",
"--config", "/absolute/path/to/examples/gateway.yaml"
]
}
}
}The client then sees exactly the downstream server's tools (names, descriptions, input schemas, unmodified) and every call it makes is authorized by policy before AgentShield forwards it.
ALLOW / REVIEW / DENY through the real server
Against examples/mcp_policy.yaml:
echo(...) -> ALLOW -> forwarded; downstream result returned as-is
create_file(...) -> REVIEW -> blocked (no approval provider configured)
delete_file(...) -> DENY -> blocked; downstream never calledLocal demos
# Core only, no MCP: the primary usage pattern.
python examples/sdk_example.py
# A tiny "agent" that consults AgentShield before acting, across three
# scenarios (deterministic DENY, policy-ALLOW-but-Jev-says-REVIEW,
# policy-ALLOW-and-Jev-agrees) -- see "Using AgentShield from an Agent" above.
python examples/agent_demo.py
# Same, with real Jev semantic evaluation added on top.
pip install -e ".[jev]"
TYPESAFE_API_KEY=... python examples/agent_demo.py
# Real agent autonomy: the Agent itself derives the action from plain
# natural-language requests (via a real LLM, or a deterministic stand-in
# without one) -- nothing here hard-codes which action each request means.
pip install -e ".[agent-demo,jev]"
python examples/real_agent_demo.py
ANTHROPIC_API_KEY=... TYPESAFE_API_KEY=... python examples/real_agent_demo.pyThe two demos below need the optional MCP extra:
examples/mcp_server.py is a tiny fake MCP server
with three harmless tools (echo, create_file, delete_file), confined to
a local sandbox directory. examples/mcp_policy.yaml
allows echo, requires review for create_file, and denies delete_file.
All demos are entirely local — no network access, no API keys, no external
services.
pip install -e ".[mcp]"
# MCP library form: embeds MCPGateway directly in this process.
python examples/run_demo.py
# MCP real server form: a real MCP client connects to
# `python -m agentshield.mcp.server` as a subprocess, exactly like Claude
# Desktop or Cursor would.
python examples/run_demo_server.pyBoth MCP demos run MCP Client -> AgentShield -> Fake MCP Server end to end
and demonstrate ALLOW, REVIEW, and DENY plus (for run_demo.py) the
resulting audit trail.
Real MCP Integration
The MCP demos above are proven against a small fake local MCP server.
examples/github/ proves the exact same, unmodified
MCP adapter against a real MCP ecosystem server:
github/github-mcp-server,
GitHub's own official MCP server.
Claude
↓
AgentShield
↓
GitHub MCP
↓
GitHubWhy this matters: AgentShield can sit in front of an existing, unmodified
MCP server — this isn't limited to a purpose-built demo server. The same
MCPGateway/AgentShieldMCPServer that talk to examples/mcp_server.py
talk to real GitHub MCP with zero code changes; only the config and policy
differ. Nothing GitHub-specific was added to the Core, MCPGateway, or
AgentShieldMCPServer — GitHub-specific detail lives only in
examples/github/ and
tests/test_github_integration.py. The
same gateway works unchanged in front of ComfyUI MCP, AWS MCP, or any other
MCP server.
Security model: AgentShield is positioned between the agent and the tool layer. It does not replace MCP, and it does not modify the downstream MCP server — it evaluates the action before execution:
Agent
↓
AgentShield
↓
Policy decision
↓
Tool executionSo DENY = the tool is never executed, whether that tool is a local fake
server or the real GitHub API.
See examples/github/README.md for full,
step-by-step setup: prerequisites, the Python 3.11+ requirement,
installation, GitHub MCP setup (Docker), the required
GITHUB_PERSONAL_ACCESS_TOKEN environment variable (never committed —
referenced from config as ${GITHUB_PERSONAL_ACCESS_TOKEN} and resolved
only at connect time), the example policy and gateway config, an MCP client
configuration example, and the local-test vs. integration-test commands.
Quick summary:
pip install -e ".[mcp]"
export GITHUB_PERSONAL_ACCESS_TOKEN=your_token_here # never committed
python -m agentshield.mcp.server --config examples/github/gateway.yamlThe normal test suite stays completely offline and credential-free:
pytest # integration tests skip themselves automatically
pytest -m "not integration" # same, explicitReal integration tests (require Docker + a token):
pytest -m integrationWhat's implemented
Core — decision engine:
DecisionEngine/DecisionRequest, typed decision/policy models, deterministic matching and precedence, a default-allow fallback, and an in-memory audit log. This is the primary, MCP-agnostic public API (shield.evaluate(request)).MCP adapter (library):
MCPGateway, a policy-enforcement proxy for a downstream MCP server reached over stdio, using the official MCP SDK; tool discovery; ALLOW/REVIEW/DENY enforcement; a pluggable approval abstraction; audit logging.MCP adapter (real server):
agentshield.mcp.serverexposesMCPGatewayas an actual upstream-facing MCP server over stdio (a thin protocol adapter with no decision logic of its own), launchable viapython -m agentshield.mcp.server --config ...and usable directly from Claude Desktop, Cursor, or any other MCP-compatible client.Real MCP integration: the same MCP adapter proven against a real MCP ecosystem server (GitHub MCP) rather than only the bundled fake one, with an opt-in, credential-gated integration test suite.
Optional semantic evaluation:
DecisionEngine(..., semantic_evaluator=...)andagentshield.jev.JevSemanticEvaluator, a singlegood/review/badjudgment on top of deterministic policy, backed by the real Jev API. The Core has no dependency on it either.
Deliberately not implemented yet: automatic agent interception, a general CLI, an HTTP authorization service or remote AgentShield service, a database, a web dashboard, authentication infrastructure, multiple semantic questions/evaluators/models, a semantic-orchestration or confidence-threshold framework, or configurable approval backends (Slack, webhook, web UI, secret management). A full enforcement/proxy redesign beyond the existing MCP adapter is future work — see the scope note at the top of this README.
Roadmap
Core decision engine (current)
→ MCP adapter — library + real server (this repo)
→ Real MCP integration — GitHub (this repo)
→ Optional semantic evaluation — Jev (this repo)
→ Python SDK packaging
→ stronger enforcement integrations
→ TypeScript SDK
→ further integrationsInstalling and running tests
Requires Python 3.11+ (requires-python = ">=3.11").
pip install -e ".[dev]" # includes the optional mcp and jev extras
pytest # offline, credential-free
pytest -m integration # optional: real GitHub MCP + real Jev API testsStatus
AgentShield is an early-stage, open-source project. It does not make any production security guarantees; treat it as policy infrastructure you integrate and test against your own threat model.
This server cannot be deployed
Maintenance
Related MCP Connectors
Runtime permission, approval, and audit layer for AI agent tool execution.
Deterministic allow/require_approval/deny verdicts for agent actions, before they happen.
See, price, and control every tool call your AI agents make: policy checks, cost, and audit tools.
Security gateway for AI agents: policy, approval, and audited execution, no secrets shared.
Related MCP Servers
AlicenseNot gradedqualityDmaintenanceEnforces deterministic policies on AI agent tool calls, evaluating actions against compliance modules (SOC 2, HIPAA, GDPR, etc.) and returning ALLOW, BLOCK, or CONSTRAIN decisions with an audit trail.MIT- AlicenseNot gradedqualityAmaintenanceDeterministic policy enforcement for AI agent tool calls. It evaluates every tool call against user-defined rules before execution, with no LLM in the authorization path.3MIT
- AlicenseNot gradedqualityAmaintenanceA policy-based security layer for AI agents and MCP tools that evaluates tool requests against security policies before allowing, denying, or requiring human approval for execution.285 PyPI1MIT

ERDL Guardofficial
AlicenseNot gradedqualityAmaintenanceEnforces deterministic policy decisions on AI agent tool calls, supporting allow, deny, correct, escalate, and human review actions with verifiable audit receipts.48 npmMIT