aegis-crew
Aegis-crew provides a zero-trust MCP tool server for a multi-agent research and reporting pipeline. You can:
Search the knowledge base (
kb_search_tool): Retrieve top-k matching text chunks from the corpus.Read documents (
read_doc_tool): Get the full text of a specific authorized document.Publish reports (
publish_report_tool): Write reports to authorized output paths, gated by human-in-the-loop (HITL) approval.
All tool calls are authorized by a deny-by-default CapabilityBroker, logged in an audit trail, and pass OWASP LLM Top-10 guardrails. The server supports optional enhanced security tiers including Rule of Two, Dual-LLM quarantine, sandboxed execution (macOS), and behavioral egress monitoring. It connects via stdio to any MCP-compatible client (e.g., Claude Desktop) or orchestrator (LangGraph, CrewAI, AutoGen).
Allows the multi-agent system to be orchestrated using CrewAI, enabling flexible agent collaboration and task delegation while maintaining security controls.
Integrates with Google's Gemini models via the gemini provider, enabling the agent crew to use Google's language models for various tasks.
Provides LangGraph-based orchestration for stateful multi-agent workflows, enabling complex agent interactions with built-in security and MCP boundaries.
Integrates with Ollama to run local language models, enabling offline or privacy-preserving agent operations with customizable models and providers.
Integrates with OpenAI's API to access GPT models, enabling the agent crew to leverage OpenAI's language capabilities for research and report generation.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@aegis-crewresearch the impact of AI on healthcare"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
aegis-crew
A multi-agent analyst crew — Researcher, Writer, Reviewer on LangGraph — wrapped
in a defence-in-depth security envelope. Agents cannot import their tools: every
tool call crosses a real MCP boundary and is capability-checked by a zero-trust
CapabilityBroker, input and output pass deterministic OWASP LLM Top-10
guardrails, and the one side-effecting action (publishing the report) is held at
a human-in-the-loop gate. The thesis: the security layer is constant and the
orchestrator is swappable — CrewAI, AutoGen, BeeAI, Microsoft Agent
Framework, and Google ADK all drive the same secured tools through the same broker.
On top of the always-on envelope sit four opt-in security tiers, each independently switchable and composable:
Tier | Flag | What it guarantees |
Rule of Two (session policy) |
| No session combines untrusted inputs + sensitive-data access + external side effects without an explicit human grant. |
Dual-LLM quarantine (semantic) |
| No provider call ever carries both tool access and untrusted bytes. |
Sandboxed tools (runtime) |
| The tool server runs inside a deny-by-default macOS Seatbelt profile derived from the broker's capabilities. |
Behavioral baseline (egress drift) |
| Outgoing reports are scored against a profile fitted on known-good output; off-profile publishes warn or block. |
Everything is testable without an API key: 351 deterministic safety tests cover every control with negative-control discipline, and a separate live tier sends real adversarial probes through a real model.
Start with the design rationale: docs/ARCHITECTURE-DEEP-DIVE.md — a full walkthrough of the architecture, the trade-offs, and the threat model behind it.
Quickstart
This project requires uv: [tool.uv] override-dependencies resolves a metadata-only json-repair pin conflict that
plain pip cannot apply, so pip install fails where uv sync succeeds.
uv sync --extra dev
cp config.example.toml config.toml # edit provider/security settings
cp .env.example .env # set ANTHROPIC_API_KEY for live runs
# Key-free demo — scripted provider, full pipeline, no network:
uv run aegis run "grid storage policy" --provider fake --auto-approve
# Deterministic safety suite (no API key):
uv run --extra dev pytest tests/safety -qThe report lands in out/report.md, written only after the HITL gate approves.
Live interactive run
uv run aegis run "renewable energy subsidies" --provider anthropicThe crew runs against the real model. When the reviewer approves the draft, the
run parks at the HITL gate and prints the proposed report. You are prompted
to approve (y) or reject (n) before out/report.md is written.
Providers: fake (scripted, deterministic), anthropic, ollama, plus
first-party openai, gemini, mistral, and watsonx adapters.
config.toml keys of interest:
[provider]
default = "anthropic" # anthropic | ollama | fake | openai | gemini | mistral | watsonx
anthropic_model = "claude-sonnet-4-6"
[security]
output_dir = "out"
corpus_dir = "corpus"
reflexion_max_retries = 2The kb_search tool reads .md/.txt documents from corpus/ (example
policy documents ship in the repo).
Standalone MCP tool server (stdio)
uv run aegis serve-mcpStarts the FastMCP tool server on stdio. Any MCP-compatible client (e.g. Claude
Desktop) can connect and call the three tools (kb_search_tool,
read_doc_tool, publish_report_tool) through the zero-trust broker.
Note on retrieval quality:
MockEmbedding(offline, deterministic) is the default so all tests and the--provider fakedemo work without API keys. Retrieval is illustrative in this mode — pass a real embedding model viabuild_index(embed_model=...)for production-grade semantic search.
Related MCP server: production-grade-mcp-agentic-system
Architecture
User topic
|
v
[scan_input / LLM01] ← guardrail: prompt-injection / jailbreak
|
v
StateGraph (LangGraph)
┌──────────────────────────────────────────────────────────┐
│ │
│ START → researcher ──────────────────────────────────► │
│ (ReAct loop) │
│ │ MCP boundary (MCPToolClient) │
│ │ └── CapabilityBroker.authorize() │
│ │ └── check_tool_args / LLM06 │
│ ▼ │
│ writer │
│ ▼ │
│ reviewer ──"retry"──► writer │
│ (Reflexion) │
│ │ "ok" │
│ ▼ │
│ hitl_gate ── reject ──► END │
│ (interrupt()) │
│ │ approve │
│ ▼ │
│ publisher ── MCP ──► publish_report_tool │
│ │ └── check_egress / LLM02 │
│ ▼ │
│ END │
└──────────────────────────────────────────────────────────┘
|
v
out/report.md (written only after HITL approval)Key components:
Layer | Artifact |
Orchestration |
|
MCP boundary |
|
Zero-trust |
|
Guardrails |
|
HITL gate |
|
Agents |
|
State |
|
CLI |
|
Full threat analysis: docs/THREAT_MODEL.md (STRIDE
per component).
The always-on envelope
Zero-trust
CapabilityBroker(src/aegis/security/zero_trust.py:45) — every tool call is explicitly authorized against a per-tool capability set; denials are recorded in theAuditLog(zero_trust.py:26). No tool runs without explicit allowance.OWASP LLM Top-10 guardrails (
src/aegis/security/guardrails.py) — four guards mapped to four controls:scan_input→ LLM01 (Prompt Injection / Jailbreak)check_egress→ LLM02 (Sensitive Information Disclosure)validate_output→ LLM05 (Improper Output Handling)check_tool_args→ LLM06 (Excessive Agency)Full mapping:
docs/OWASP_MAPPING.md
Human-in-the-loop gate (
src/aegis/graph.py:257) — LangGraphinterrupt()parks the run before the one side-effecting tool (publish_report_tool). A rejected decision routes toENDwithout writing any file.Audit log (
src/aegis/security/zero_trust.py:26) — all broker allow/deny decisions are recorded with tool, action, target, verdict, and reason.MCP choke point — tools are exclusively reachable via the
MCPToolClientboundary; agents have no direct function references. The broker enforces authorization at the server side of that boundary.
The spine files behind these controls are byte-frozen:
tests/safety/test_no_security_regression.py guards them with a git-diff check
plus SHA-256 digest pins, so every later capability had to compose on top of
the controls instead of editing them. The frozen guardrails are also
mutation-tested — see docs/mutation-testing.md.
The four security tiers
Rule of Two (session policy, --rule-of-two)
Meta's "Agents Rule of Two": an agent session may combine at most two of [A] untrustworthy inputs, [B] sensitive-system/private-data access, [C] state-change/external communication. Needing all three means the session must not run autonomously — human approval is the mandated minimum.
aegis enforces this mechanically: a SessionLedger accumulates factors as the
broker authorizes calls (read→B, write→C, network→A+C, unknown→all three,
fail-closed; carve-outs are explicit declarations), and RuleOfTwoBroker (a
drop-in subclass of the frozen CapabilityBroker) denies, at authorize-time,
any call whose factors would complete the triple without a human grant. The
HITL gate's approval is the grant: approving the publish unlocks
EXTERNAL_WRITE for the session; rejecting leaves the third factor locked even
if graph routing were bypassed. Every decision lands in the same audit log.
Enforcement is in-process: in aegis run (in-memory MCP) the server and graph
share one broker and one ledger, so enforcement is end-to-end. In serve-mcp
stdio mode the ledger lives inside the server process: tool-side factor
accumulation and denial are enforced there, but graph-side marks and HITL
grants cannot reach it — a cross-process grant channel is future work, not
claimed. Classification is structural, not semantic: Rule of Two bounds blast
radius; it does not detect injection — the guardrails do that, and neither
subsumes the other.
Dual-LLM quarantine (semantic tier, --quarantine)
"Design Patterns for Securing LLM Agents against Prompt Injections"
(arXiv:2506.08837) states the principle: once
an agent has ingested untrusted input, it must be impossible for that input to
trigger consequential actions. --quarantine turns the crew into the paper's
Dual LLM pattern (with CaMeL as the
data-flow lineage):
the privileged researcher keeps its tools but never sees untrusted bytes — tool results are parked in a
QuarantineStoreand appear to the model only as opaque refs (⟦Q1⟧ from kb_search (1204 chars)) plus validated signals;a quarantined model (never given tools) reads the untrusted text and answers only schema-validated primitives (bool / bounded int / fixed choice); a reply that fails strict validation is re-asked once, then rejected fail-closed — free text never crosses back;
the writer and reviewer see the content but are quarantined roles: structurally tool-less (a regression-pinned property, not an accident);
the publisher boundary is unchanged: HITL + broker + egress guards + Rule of Two when enabled.
The enforced, tested invariant: no provider call ever carries both tool access
and untrusted bytes. Default-off; [security] quarantine_provider optionally
routes the quarantined model to a cheaper local provider.
Honesty notes, verbatim from the design review: "Quarantine bounds capability and data flow, not signal truthfulness: extracted signals are derived from untrusted content and remain adversary-influenceable. An attacker can lie to the relevance check; they cannot make the privileged model see their text or invoke a tool with it." "The quarantined model can still be prompt-injected. The guarantee is that injection there is inconsequential by construction: the quarantined call carries no tool access, and its only output channel is a schema-validated primitive — free text never crosses back to the privileged side." "Writer and reviewer see untrusted content by design; they are quarantined roles — structurally tool-less. The consequential boundary remains the publisher path (HITL + broker + egress guards + Rule of Two when enabled). Detection (guardrails) and structure (quarantine) are complementary; neither subsumes the other."
Stretch rung — --quarantine-mode plan (CTE-lite): the privileged model
emits a full typed plan (pydantic-validated verbs: search / read /
extract) from the topic alone, before any untrusted data exists; a
deterministic executor walks it through the same broker-checked MCP path.
Honest label, verbatim: "This is Plan-Then-Execute with typed steps and
capability-checked execution (CTE-lite). Full CaMeL-style data-flow labels on
variables are future work, not claimed."
Sandboxed tool execution (sandbox tier, --sandbox)
uv run aegis run "renewable energy subsidies" --sandbox --provider fake --auto-approveOWASP ASI02 ("Tool Misuse") calls for two layers: the broker that authorizes a
tool call, and a runtime boundary that survives the broker being wrong. aegis run --sandbox adds the second layer: the serve-mcp tool server is launched
inside a macOS sandbox-exec (Seatbelt) subprocess under a caps-derived
deny-by-default profile (security/sandbox.py, SandboxPolicy.from_caps) —
filesystem access is scoped to exactly the read/write paths the broker's
capabilities declare, and network is denied outright. The broker remains the
policy decision point inside that subprocess; the sandbox bounds the blast
radius if the broker is wrong or the tool runtime is compromised.
The launch writes two inspectable artifacts under out/sandbox/:
profile.sb (the actual Seatbelt profile text) and config.toml (the
secrets-free config the sandboxed server reads). Only the profile's
sha256_16 enters the shared AuditLog as a sandbox/launch record
(security/sandbox.py:launch_audit_record) — the content lives on disk for
inspection, not duplicated into the audit trail.
--sandbox composes with --quarantine and --auto-approve. It refuses
--rule-of-two: with --sandbox, tool-side capability decisions are made
and audited inside the sandboxed server process, and Rule of Two's session
ledger — an in-process object — cannot span that process boundary, so the
combination is rejected at the CLI rather than silently downgraded to
graph-side marks only.
Combination | Result |
| valid |
| refused ( |
Honesty guards, verbatim from the design review:
"Broker authorizes, sandbox contains. The broker remains the policy decision point; the sandbox bounds the blast radius of a tool-runtime compromise or a broker misconfiguration. Neither subsumes the other, and neither makes tool OUTPUT trustworthy — untrusted-content handling remains the quarantine and guardrail tiers' job."
"
sandbox-execis deprecated as a public CLI but is the OS-shipped Seatbelt mechanism that still underpins macOS app sandboxing. This tier is best-effort local containment on a developer machine, not a certified security boundary and not E2B: no microVM, no snapshotting, no network allowlisting — egress is simply denied.""The sandbox tier is darwin-only and fails closed: on any other platform
--sandboxis an error, never a silent downgrade to unsandboxed execution. A Linux backend (namespaces/containers) is future work."
Behavioral baseline (egress drift tier, --baseline)
uv run aegis baseline-fit ./known-good-reports -o profile.json
# config.toml: [security] baseline_profile = "profile.json"
uv run aegis run "renewable energy subsidies" --baseline --provider fake --auto-approve --config config.tomlsecurity/baseline.py wires portcullis's
BaselineMonitor into the publisher boundary as an optional, default-off egress
tier, run after check_egress and before the MCP publish call: the outgoing
report is embedded and scored against a profile fitted offline on known-good
reports. This is a behavioral tier, not a content tier — it complements the
existing guardrail and egress checks, it does not replace them.
Fit → enforce lifecycle: aegis baseline-fit <known-good-dir> [-o profile.json] [--threshold-percentile 95.0] fits a profile on *.md/*.txt reports and saves it
as JSON; [security] baseline_profile = "profile.json" in the config points
aegis run --baseline at it. Needs the [baseline] extra
(pip install 'aegis-crew[baseline]'), lazy-imported inside security/baseline.py
so the core install and the deterministic safety suite never need it.
Warn vs. block: --baseline alone runs in warn mode — drift is recorded on
the audit trail and the run still publishes. --baseline-mode block refuses an
off-profile publish with a content-free BLOCKED: ... critique, same shape as an
egress denial. Audit records land on the same ledger as every other tier:
baseline / egress / report.md / allow|deny, with reasons carrying
distance/threshold/destination — never report text.
Composition: --baseline composes with --rule-of-two, --quarantine, and
--sandbox — verified end-to-end by
tests/safety/test_baseline.py::test_baseline_composes_with_rule_of_two,
::test_baseline_composes_with_quarantine, and ::test_baseline_composes_with_sandbox
respectively (the monitor and embedder run in the graph process; --sandbox
contains a disjoint layer, the serve-mcp subprocess).
Honesty guards, verbatim from the design review:
"A behavioral baseline is anomaly detection, not attack detection: it flags conversations that drift from the fitted profile. A novel-but-benign topic can flag; an attack phrased on-profile will not. It bounds drift — it does not detect prompt injection or PII. Run it alongside the content gate, never instead of it."
"Fail posture is fail-closed with no override exposed: an embedder failure blocks the publish even in warn mode. Errors never downgrade to a warning silently."
"HashingEmbedder is deterministic and content-sensitive but NOT semantic: it demonstrates the wiring and the fit/enforce lifecycle, not detection quality. Hand a real embedding model to LlamaIndexEmbedder for any real deployment, and fit and enforce with the SAME embedder — profiles are embedder-specific."
The default no-key path uses HashingEmbedder (stdlib token-hash bag, unit
vector) — deterministic and content-sensitive, but not semantic (guard 3 above).
LlamaIndexEmbedder adapts any llama-index BaseEmbedding for a real
deployment. Fit and enforce must use the same embedder: profiles are
embedder-specific.
Capabilities
Everything below composes on the frozen spine; each item names the module that implements it and the test or CLI command that exercises it.
Broker & policy
Zero-trust broker + audit log (
security/zero_trust.py) — deny-by-default per-tool capability sets; every allow/deny audited. Exercised bytests/safety/test_zero_trust.py.Rule-of-Two session ledger (
security/rule_of_two.py,security/policy.py) — factor accumulation + authorize-time denial, HITL approval as the human grant. Exercised bytests/safety/test_rule_of_two.py.Sandbox policy derivation (
security/sandbox.py) — a Seatbelt profile generated from the broker's capability set, deny-by-default, network denied. Exercised bytests/safety/test_sandbox.py.HITL gate (
graph.py,security/hitl.py) —interrupt()before the one side effect; resume viaCommand(resume=...). Exercised bytests/safety/test_graph_hitl.py.
Guardrails & spine
OWASP LLM Top-10 guards (
security/guardrails.py) — LLM01/02/05/06, deterministic, mutation-tested (docs/mutation-testing.md). Exercised bytests/safety/test_guardrails.pywith negative controls.Dual-LLM quarantine (
security/quarantine.py,security/plan_execute.py) — the semantic tier above. Exercised bytests/safety/test_quarantine.py.Behavioral baseline (
security/baseline.py) — the egress drift tier above. Exercised bytests/safety/test_baseline*.py(skips cleanly without the[baseline]extra).NeMo Guardrails integration (
security/nemo_rails.py) — anLLMRailsinput rail delegating to the frozenscan_input; blocks injection with no model in the deterministic test, and with a live Ollama in the live test.Adversarial hardening (
adversarial/defense.py) — prompt-space evasion bench against the LLM01 guardrail; measured attack success rate 0.78 bare → 0.00 withharden(obfuscation-foldingnormalize),guardrails.pyuntouched. CLI:aegis adversarial bench.MCP tool-poisoning defense (
mcp/tool_scan.py) —vet_toolsscans advertised tool metadata for smuggled instructions;ToolPin/detect_rugpullcatch post-approval description swaps. CLI:aegis mcp scan.Data-poisoning attack + detector (
adversarial/poison.py) — plants a poisoned corpus doc and detects it by reusing the frozen LLM01 guardrail plus a duplication check. CLI:aegis adversarial poison-scan.Model-extraction attack + defense (
adversarial/extraction.py) — a surrogate fit on victim queries; measured fidelity degradation 0.96 → 0.57 under the perturbation/budget defense.
Orchestration & interop
The crew (
graph.py,agents/) — ReAct researcher, writer, Reflexion reviewer with bounded retry, HITL publisher, on a LangGraphStateGraph.Swappable orchestrators (
orchestrators/) — CrewAI, AutoGen, BeeAI, Microsoft Agent Framework, and Google ADK each drive the same secured MCP tools over theserve-mcpstdio contract via their native adapters; none re-implements a tool. Structural tests are model-free; live smoke tests run against local Ollama (qwen3:8b).uv run aegis run "grid storage policy" --orchestrator crewai # | autogen | beeaiFive orchestration topologies (
topologies/) — supervisor, plan-and-execute (replan-on-failure), routing/handoffs, swarm (decentralizedCommand(goto=...)), and debate (judge sides with either debater), each a separate LangGraph over the same secured boundary. CLI:aegis topo run --pattern <name> TOPIC.A2A protocol (
protocols/) — the crew is callable as a Google Agent2Agent agent (aegis serve-a2a) and can delegate to peers (aegis a2a-call); inbound messages are guardrail-scanned, outbound peer calls are deny-by-default and audited. Needs the[a2a]extra.Computer use (
agents/computer_use.py,mcp/tools/browser.py) — a bounded perceive→decide→act loop over a gated PlaywrightBrowserTool; origins and action verbs are allowlisted, sensitive actions (e.g.submit) park at HITL. CLI:aegis browse URL TASK. Needs the[browser]extra.Evaluation harness (
eval/harness.py) —aegis evalruns each task through the real crew and scores trajectory + safety-held + LLM-judge (deterministicStubJudgeby default,--livefor an Ollama judge);--metric tool-accuracyscores F1 over expected-vs-actual tool-call sets, and topology runs take a caller-supplied success predicate.
Memory & privacy
Three-tier memory (
memory/store.py) — episodic/semantic/scratch over SQLite;remember/recallare capability-checked through the same broker (memory paths are scoped under the run's output dir). Exercised bytests/safety/test_memory.pyand the crew-wiring tests.Semantic fact store + consolidation (
memory/semantic_store.py,memory/consolidate.py) — fact/preference store with dedup on(subject, predicate);consolidate()extracts durable facts from the episodic log and shrinks event volume while preserving knowledge. CLI:aegis memory facts.Context engineering (
context/assembler.py) —ContextAssemblerpacks prioritized sections into a token budget, compressing lowest-priority first; counter and summarizer are dependency-injected, so it tests deterministically.PII redaction (
privacy/redaction.py) —Redactorwithredactandblockpolicies. CLI:aegis privacy redact PATH.Differential privacy (
privacy/dp_train.py) — a real Opacus DP-SGD micro-train reporting the spent budget (ε≈2.25, δ=1e-5; CPU — Opacus grad-sample hooks are MPS-incompatible). CLI:aegis privacy dp-train. Needs the[privacy]extra.Federated learning simulation (
privacy/federated.py) — a labeled FedAvg simulation (honestly labeled: simulation, not a distributed deployment).
Supply chain & governance
Artifact attestation (
supplychain/provenance.py,supplychain/verify.py) — Ed25519 in-toto-style sign + verify-on-load; a tampered artifact is blocked before use. In-house local signing, NOT keyless sigstore. CLI:aegis attest,aegis verify.SBOM (
supplychain/sbom.py) — CycloneDX-style SBOM for installed packages. CLI:aegis sbom.cosign signing (
supplychain/cosign.py) — key-based blob sign/verify via the cosign CLI (key-based, NOT keyless-OIDC). CLI:aegis provenance cosign-sign.C2PA content provenance (
provenance/c2pa.py) — signs images with a local self-signed ES256 credential;verify_imageperforms real tamper-detection on the manifest. CLI:aegis provenance c2pa-sign/c2pa-verify. Needs the[c2pa]extra.Governance gate (
governance/policy_engine.py) — deny-by-defaultGovernanceGateover a NIST/ISO/EU-AI-Act risk register mapped to real callables; exits 1 on violation. Scope: an executable pre-flight / CI gate, not a runtime interceptor wired intograph.py. CLI:aegis govern check/report. Mapping:docs/AI_GOVERNANCE.md.ISO/IEC 42001 (
governance/aims.py) —AimsGatebinds 5 clauses to real callables and blocks on nonconformity. CLI:aegis govern aims-check. Mapping:docs/ISO_42001_AIMS.md.OWASP Agentic Top-10 map (
security/agentic_map.py) — an executable threat→control map over 8 threats: 6 controlled, 2 documented gaps (T4 Resource Overload, T7 Cascading). CLI:aegis agentic map/check. Mapping:docs/OWASP_AGENTIC_TOP10.md.Executable attack trees (
threat/attack_tree.py) — extends the STRIDE model; every leaf must resolve to a real aegis control or the test fails. CLI:aegis threat tree.Confidential computing (
confidential/) — a design study, not built: I have no TEE hardware. An attestation-gated inference stage refuses to run without a TEE quote (aegis confidential attestexits nonzero here); nothing is simulated. Design doc:docs/CONFIDENTIAL_COMPUTING.md.
Optional extras
Heavy dependencies are lazy-imported inside the functions that need them, so the core install and the deterministic safety suite never require any of these:
pip install 'aegis-crew[a2a]' # a2a-sdk — serve-a2a / a2a-call
pip install 'aegis-crew[browser]' # playwright — browse (then: python -m playwright install chromium)
pip install 'aegis-crew[privacy]' # torch + opacus — privacy dp-train
pip install 'aegis-crew[agentic]' # agent-framework-core — the MAF orchestrator
pip install 'aegis-crew[adk]' # google-adk[mcp] + litellm — the Google ADK orchestrator
pip install 'aegis-crew[rails]' # nemoguardrails — the NeMo rail layer
pip install 'aegis-crew[c2pa]' # c2pa-python — provenance c2pa-sign / c2pa-verify
pip install 'aegis-crew[baseline]' # portcullis — the behavioral-baseline tier
pip install 'aegis-crew[platform]' # convenience bundle: a2a + browserTesting
Two tiers, intentionally separated:
Deterministic safety suite (no API key)
uv run --extra dev --extra baseline --extra adk pytest tests/safety -q351 tests cover every security control, the guardrail negative-control
discipline, HITL interrupt/resume, the MCP boundary, the zero-trust broker, all
four tiers and their compositions, and CLI smoke. If a control is removed,
its corresponding test fails — the suite is written as a regression harness,
not just coverage. tests/safety/test_baseline*.py skip cleanly (via
pytest.importorskip) when the [baseline] extra is not installed — the suite
stays green either way.
Live adversarial red-team suite (requires ANTHROPIC_API_KEY)
ANTHROPIC_API_KEY=sk-... uv run --extra dev pytest -m redteam -qSends real adversarial probes through a live crew run:
A corpus document with a planted prompt-injection comment must not cause scope escape.
A jailbreak topic must be blocked at
scan_inputbefore the model or tools see it.
Further live markers: -m live (orchestrator + topology smoke against local
Ollama) and per-provider markers (openai_live, gemini_live, mistral_live,
watsonx_live).
Honest limitations
Confidential computing is a design study — no TEE hardware here, so it ships as a fail-closed attestation gate plus a design doc; TEE-backed deployment is not claimed.
The sandbox tier is darwin-only best-effort local containment (Seatbelt), not a certified security boundary; it fails closed on other platforms.
Rule of Two in
serve-mcpstdio mode enforces tool-side only; the cross-process HITL grant channel is future work, not claimed.Default retrieval and baseline embedders are illustrative —
MockEmbeddingandHashingEmbedderare deterministic, not semantic; swap in real embedding models for deployment.cosign signing is key-based, not keyless-OIDC; the Ed25519 attestation is in-toto-style, not sigstore.
OWASP Agentic Top-10 coverage is 6 of 8 mapped threats, with T4 (Resource Overload) and T7 (Cascading) as documented gaps.
The governance gates are pre-flight/CI gates, not runtime interceptors.
FedAvg is a simulation, not a distributed deployment.
Project history
The wave-by-wave build ledger — what was added when, with the full capability-to-evidence tables — lives in docs/BUILD-NOTES.md.
Available Tools
3 toolskb_search_toolC
Search the knowledge base; returns the top-k matching chunk texts.
| Name | Required | Description | Default |
|---|---|---|---|
| k | No | ||
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must bear the burden. It implies a read-only search operation but does not explicitly state safety, rate limits, or any side effects. Minimal behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, making it concise but lacking structure. It front-loads the verb but omits important details that could be organized into separate points.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (2 params, no nested objects, output schema present), the description is minimally adequate. It explains the core output but does not cover edge cases, error conditions, or additional behaviors.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, and the description only indirectly explains 'k' via 'top-k' and 'query' via 'Search'. It fails to explicitly define parameter meanings or provide additional details like format constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Search' and resource 'knowledge base', and explains output as 'top-k matching chunk texts'. It distinguishes from siblings like read_doc_tool (reading a specific doc) and publish_report_tool (publishing).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, nor any conditions or prerequisites. The description is purely functional with no usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
publish_report_toolC
Write a report to an authorized output path; returns the path written.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| content | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavioral traits. It mentions 'authorized output path' but does not explain authorization requirements, overwrite behavior, error handling, or return format beyond 'path written'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, concise, but overly terse; it sacrifices necessary details for brevity, making it less helpful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having an output schema, the description does not leverage it. For a tool with two required parameters and no annotations, the description fails to provide adequate context about authorized paths, content format, or expected outcomes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the tool description provides no additional meaning for the 'path' or 'content' parameters, leaving the agent without guidance on format, constraints, or allowed values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'write', the resource 'report', and the output 'returns the path written', distinguishing it from sibling tools like kb_search_tool (search) and read_doc_tool (read).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use or not use this tool vs its siblings. The description does not provide context for when it is appropriate over alternative tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_doc_toolA
Read a single document from the authorized corpus and return its text.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It discloses that the operation is read-only and returns text, but it does not discuss error handling, authorization details beyond 'authorized corpus', or any rate limits. For a simple read operation, this is adequate but not thorough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence of 12 words, front-loaded with the core action. Every word is necessary and contributes to clarity. There is no extraneous information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity and the existence of an output schema, the description fails to define the 'path' parameter or any constraints. With 0% schema coverage, the description should compensate by explaining the parameter, which it does not. Thus, it is incomplete for reliable agent use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning the 'path' parameter is undefined in the schema. The description does not elaborate on what 'path' means (e.g., format, absolute vs relative, document ID). It only implies it's a document identifier, leaving the agent with minimal semantic guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Read'), the resource ('a single document'), the scope ('from the authorized corpus'), and the output ('return its text'). This succinctly distinguishes it from sibling tools like 'kb_search_tool' (search) and 'publish_report_tool' (create/publish).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when one needs to retrieve a specific document's text by path, but it does not explicitly state when to use this tool versus alternatives, nor does it provide exclusions or prerequisites. With sibling tools present, more explicit guidance would be beneficial.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
kb_search_tool - First observed
publish_report_tool - First observed
read_doc_tool
TDQS
Scored across 3 tools
Each tool has a distinct purpose: searching the knowledge base, reading a document, and publishing a report. There is no overlap or ambiguity.
All tools use snake_case and include a verb, but the pattern varies: kb_search_tool follows noun_verb_tool, while read_doc_tool and publish_report_tool follow verb_noun_tool. This minor inconsistency is still clear.
With three tools, the set is minimal but fits a focused knowledge base and report creation workflow. It is not overly small given the scope.
The tools cover search, read, and publish, but lack update, delete, or list operations. For a basic workflow it is functional, but gaps exist for lifecycle management.
Maintenance
Related MCP Connectors
MCP server for secureFlows: token-free URL builders and integration-linting tools for AI agents.
- gatewayOAuthai.sealgate
MCP gateway with runtime security policy, tool-call-level control, and audit of agent actions.
Zero-secret MCP gateway for AI agents: risk-scored, audited calls with human-in-the-loop approval.
Authenticated LLM MCP Agent
Related MCP Servers
- AlicenseCqualityDmaintenanceEnables secure, zero-trust access to MCP tools through short-lived, signed capability leases that bind tool execution to specific sessions, intents, and constraints. Prevents prompt injection attacks and privilege escalation with dynamic risk scoring, policy enforcement, and tamper-evident audit logging.41MIT
- AlicenseNot gradedqualityDmaintenanceA production-grade MCP server designed for multi-tenant, authenticated, and observable AI agent systems, enabling secure tool execution across heterogeneous data sources.62MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to discover and execute tools via a secure MCP server with JWT authentication, RBAC, rate limiting, and audit logging.1MIT
- AlicenseNot gradedqualityCmaintenanceGoverned MCP server for bank-grade agent tool access with RBAC, PII redaction, rate limiting, and audit logging.MIT