warden
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@wardenreserve a task ID for the refactor"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Warden
A constitutional runtime for AI agents.
Agents cannot mint their own authority.
Every AI coding agent today can invent a task ID, mark its own work as done, approve itself, and quietly rewrite the story of what happened. We keep adding capability to agents. Warden adds the opposite: a small set of constitutional constraints that make an agent's authority external, its rejections loud, and its history unforgeable.
Warden is an MCP server. Point any MCP-capable agent (Codex, Claude Code, or your own) at it, and the agent now operates under a constitution.
Born from months of operating multiple concurrent AI agents in a real production ecosystem — where agents were repeatedly caught assigning themselves task IDs that were never granted to them. This is the minimal, generic distillation of the governance patterns that fixed it.
The Constitution
Article I — Identity is granted, never minted
The only way an agent obtains a task identity is reserve_task_id — an atomic reservation from a registry the agent does not control. If an agent invents an ID and submits work under it, the submission is blocked and a SELF_MINTED_IDENTITY violation is written to the permanent record. If it uses an ID reserved by another actor: IDENTITY_THEFT, on the record.
Article II — Work moves only inside typed envelopes, and rejection is loud
All work enters as an envelope with an explicit type (proposal, diagnostic, incident, evidence). An unknown type is not silently dropped or coerced — it is rejected explicitly, with the allowed set echoed back, and the attempt logged. Silent enum rejection is itself a governance bug: an agent that never hears "no" concludes "yes".
Article III — Agents propose; humans promote
There is no MCP tool that approves work. Inspect the server: it does not exist. Agents may only request_promotion; approval happens in a CLI (warden approve <gateId>) that lives outside the agent's reach. The gate is not a policy the agent is asked to follow — it is a capability the agent does not have.
Article IV — History is append-only and verifiable
Every event — grants, submissions, promotions, and especially violations — is appended to a hash-chained ledger. Each event carries the SHA-256 of the previous one. replay_run returns the chain and its verification: edit a single byte of history and the chain reports exactly where it broke. An agent cannot rewrite its own past, and neither can anyone else without leaving evidence.
Related MCP server: agent-gate
Architecture
flowchart LR
subgraph AgentSide["Agent side — proposal only"]
Agent[AI agent] --> MCP[Warden MCP server]
MCP --> Core[Constitutional core]
end
subgraph Boundary["Constitutional boundary"]
Gate{{Pending human gate}}
Ledger[(Append-only<br/>hash-chained ledger)]
end
subgraph HumanSide["Human side — authority"]
Human[Human operator] --> CLI[warden CLI]
end
Core -->|append every event| Ledger
Core -->|request promotion| Gate
CLI -->|approve or reject| Gate
Gate -->|human decision| CoreThe pending gate is the only bridge between agent proposals and human authority. There is deliberately no approval edge — and no approval tool — on the agent side.
Why this matters
The industry's answer to agent misbehavior has been better prompts: "please don't invent IDs", "please wait for approval". Prompts are requests. Warden's position is that governance must be structural:
Prompt-based governance | Constitutional governance |
"Don't invent task IDs" | Invented IDs are rejected by the registry |
"Wait for human approval" | No approval capability exists agent-side |
"Be honest about what you did" | History is hash-chained; honesty is not required |
Violations vanish into context | Violations are permanent, queryable records |
A well-aligned agent behaves identically under both. A misaligned, confused, or prompt-injected agent behaves very differently — and under Warden, you can prove what it tried to do.
This is also a thesis about human–AI collaboration: an AI that acts as a genuine cognitive extension — a governed exocortex — is more valuable than an autonomous oracle, precisely because its authority is bounded, its actions auditable, and its human principal always in the promotion path.
60-second quickstart
1. Install the CLI
npm install && npm link2. Initialize a project
cd /path/to/your/project
warden initThis creates .warden/, updates .gitignore, and prints a complete MCP configuration with absolute paths. Paste that JSON into your agent's MCP config, then restart the agent.
3. Use the constitutional boundary
The agent now sees four tools: reserve_task_id, submit_envelope, request_promotion, replay_run.
The human operates the other side:
warden gates # list pending promotions
warden approve <gateId> # promote (human-only — no agent tool can do this)
warden reject <gateId>
warden replay [taskId] # verify + print the hash-chained historyTry to cheat it
node demo/cheat.mjsThe demo drives two isolated agent profiles through self-minted identity, cross-agent identity theft, invalid envelope type, and the honest human-gated path. It shows every violation blocked and recorded, proves each agent's legitimate identity still works, and finally tampers with the ledger to show chain verification catching it.
Watch the ledger
After running the demo, start the zero-dependency dashboard server:
node demo/serve.mjsOpen http://127.0.0.1:4173. The dashboard tails .warden-demo/ledger.jsonl, verifies the SHA-256 chain in the browser, and highlights permanent VIOLATION events in red.
Design notes
Zero infrastructure. State is a JSON file; the ledger is append-only JSONL. The point is the protocol, not the database. Swap in Postgres when you outgrow it.
The agent's tool descriptions are part of the constitution. Each tool tells the agent exactly what is forbidden and that violations are recorded. Honest agents comply; the interesting data is what dishonest ones do anyway.
Multi-agent by construction. Set a distinct
WARDEN_AGENTper agent process; ownership and theft detection follow automatically.
License
Apache-2.0
Author
Márcio de Matos Costa — linkedin.com/in/marciommc Part of an ongoing body of work on governed human–AI collaboration (Hybrid Intelligence Conductor).
This server cannot be deployed
Maintenance
Related MCP Connectors
Constitutional AI kernel with 13 MCP tools, 888_JUDGE verdict pipeline, and VAULT999 ledger.
Blockchain SSN for AI agents. MCP gateway that blocks at the point of action, tamper evident audit.
Give AI agents identity, permissions, and reusable proof through one MCP.
AI governance MCP server for EU AI Act compliance and jurisdiction verification
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceCryptographic proof of every AI decision. An immutable, verifiable audit trail MCP server.1MIT
- AlicenseBqualityBmaintenanceAn MCP server that enforces fail-closed deterministic checks, independent refute-first review, and tamper-evident hash-chained receipts for AI agent outputs before claiming completion.43MIT
- AlicenseNot gradedqualityFmaintenanceConstitutional MCP server enforcing 13 Floors of governance for AI agents, providing tools for session anchoring, reasoning, safety critique, and audit logging.AGPL 3.0
- AlicenseAqualityBmaintenanceMCP server that provides cryptographic audit trails for AI agent actions, making every action tamper-evident via HMAC-SHA256 signed hash chains.10Apache 2.0