Mandare
OfficialMandare
Give your agents a budget they cannot talk their way out of.
Mandare is the accountability stack for AI agent fleets: signed agent identity (Passport), signed machine-readable authority (Mandate), a tamper-evident ledger of what agents actually did (Ledger), an offline kill switch — and external witnessing + public anchoring so ledger history cannot be truncated or rewritten without detection (with the witness on separate infrastructure, not even by the operator — the solo compose stack runs everything on one host and says so). Local-first: raw activity never leaves your machine.

Replay of the captured Demo 1 run
(pnpm demo, the no-docker path; the docker quickstart stops at call #24) — the
same script CI executes and asserts on every push. Regenerate:
node scripts/render-demo-gif.mjs.
Quickstart — 3 commands, no API keys needed
Needs Docker with Compose v2 (up --wait needs v2.1.1+). The images build
locally on first run, which takes a few minutes.
git clone https://github.com/mandarelabs/mandare && cd mandaredocker compose up -d --waitdocker compose run --rm demoThe demo releases a runaway agent loop against your gateway, priced
against a bundled mock provider: the enforcement is real, the money isn't.
The €20/day mandate kills it mid-run: 23 calls settle €19.17, call #24's reservation would
cross €20 and is refused 403 PER_DAY_EXCEEDED, the refusal is itself a ledger entry, and
mandare verify proves chain VALID, counters == replay(ledger), and the
witnessed head history covers the chain. Dashboard at
http://127.0.0.1:8788. Real providers: put keys in .env
(docs).
When you're done, docker compose down -v removes the containers and the
demo volumes (mandare-data, mandare-witness-state).
No docker (Node ≥ 22.13 and pnpm 10; if pnpm is missing, install.sh runs
corepack enable, a global change):
./install.shpnpm demo(pnpm demo runs a cheaper model and raises the gateway's default 60
calls/minute velocity limit so the budget is the only limit in play: 71 calls,
call #72 refused at the same €20 cap. The docker demo above runs the stack's
real defaults.)
pnpm demo runs without a witness, so mandare verify there can't see
entries dropped from the end of the ledger — refusals included — unless you
pass a saved --prev-head. The docker stack runs a witness, and
pnpm demo:witness shows it catching exactly that.
Or from npm (CLI + MCP server, no checkout; the MCP server reads the ledger
at MANDARE_LEDGER_DB, default ./mandare-ledger.db):
npm i -g @mandarelabs/cli && mandare helpnpx -y @mandarelabs/mcp-serverWhere state lands: the CLI's vault (./mandare-vault.db, used by passport issue and vault …) keeps its master key in the OS keychain by default;
passport issue writes the agent's credential and key under
~/.mandare/agents/ by default; and unless MANDARE_VAULT=1, the door's private key sits
next to the ledger (<ledger>.doorkey.pem, mode 0600).
Related MCP server: nobulex-mcp-server
Proofs, not data
Raw prompts and responses never leave your machine. What crosses a trust boundary is only ever a proof: salted tree heads to the witness, an integrity certificate to an auditor, a revocation bitstring to a verifier.
agent (any SDK, base URL → the door)
│ RFC 9421-signed request (passport) or scoped PoP token
▼
┌──────────────── gateway door ────────────────┐
│ kill-check → policy (mandate: caps/window/ │ ┌─ witness (external) ─┐
│ scope/approval) → INTENT entry (reserves │────▶│ salted head history, │
│ cost in the ledger tx) → provider → RESULT │ acks│ consistency-enforced,│
│ entry (settles true cost) │◀────│ public anchor (OTS) │
└───────────────┬──────────────────────────────┘ └──────────────────────┘
▼
append-only ledger (SQLite/Postgres)
hash-chained · door-signed · RFC 6962 tree
budget counters + revocation = PROJECTIONS (replay-checkable)
▼
mandare verify · certify (third-party checkable) · dashboard · killEvery door obeys three rules: fail-closed on spend, log-before-act, and agent input is hostile. Refusals are ledger entries — the system keeps its no's, and a witness keeps them from being quietly dropped.
The five demos are the acceptance tests (CI runs all of them)
# | Claim | Run |
1 | A runaway loop dies at €20, with proof |
|
2 | A stolen token is dead paper; kill bites mid-task |
|
3 | One signed mandate replaces 40 prompts; humans approve async |
|
4 | The card declines AT THE NETWORK; one cap governs both rails |
|
5 | Truncation and rewrites can't hide from an independent witness (verified with the door key held out-of-band) |
|
Each demo also exists as a self-contained, narrated scenario in
examples/ — the story, the real captured output, and the code
to read next.
Integrations
Surface | Where | What |
TypeScript SDK |
| A signed fetch for your existing Anthropic/OpenAI SDK (token PoP + passport RFC 9421) |
Python client |
| Zero-dependency token-mode client (stdlib only) |
MCP server |
| The door as MCP tools: verify, budgets, issuance, kill — stdio, env-configured |
OpenClaw skill |
| Native AgentSkills skill (also works in Claude Code): budget awareness, honest refusals, proofs, kill |
Self-host |
| gateway + witness + dashboard, no secrets needed for dry-run |
Dashboard |
| Local-first fleet view over the ledger; zero telemetry |
Docs |
| Quickstart, concepts, threat model, reference |
Security & provenance
Adversarially reviewed before launch — by AI, not yet by an external auditor: four parallel AI-assisted review passes (crypto/integrity · spend/enforcement · packaging/supply-chain · docs-vs-claims) were prompted to break the system. 15 findings — 3 HIGH — all fixed with regression tests or documented as accepted residuals, none silent. Full report:
docs/SECURITY-REVIEW-S8.md, including what was probed and held, the honest residuals, and the target list for the external audit. A second AI-assisted audit pass (2026-09) found further spend, witnessing and packaging defects; their fixes and red-team cases are logged inTASKS.md(S10-fix 2A–2D). The external audit is still pending.Fail-closed by construction: no mandate → no spend; ledger down → no action; witness dead → high-value actions refuse (the kill switch never depends on anything remote).
Red-team suites run in CI (rule R5): edit/delete/truncate/rollback/ replay/forge on SQLite AND Postgres, token theft + replay, signature coverage attacks, webhook forgery, budget races, witness split-view — and they may never be weakened to make a change pass.
Supply chain: pnpm 10 with install scripts off, 3-day dependency cooldown, frozen lockfiles, hand-rolled security primitives pinned to official test vectors where they exist (RFC 6962 CT vectors, did:key/base58) and otherwise tested against the published wire scheme with adversarial round-trip suites (Stripe signatures, OpenTimestamps). From launch: npm Trusted Publishing (OIDC provenance), cosign-signed images, signed skill envelopes. Honest reproducibility bar in
REPRODUCING.md.Verify without trusting us: the verifier, passport, and witness protocol are Apache-2.0 and embeddable;
mandare certifyproduces integrity certificates a third party checks with no ledger access.Vulnerabilities: see
SECURITY.md(private reporting, safe harbor, 90-day disclosure).
License
AGPL-3.0-only, except the embeddable packages listed in
LICENSING.md (spec, policy-engine, verifier, passport,
witness-protocol, sdk, sdk-py — Apache-2.0). The split is permanent; we do
not relicense.
Available Tools
8 toolsmandare_budget_statusBudget statusARead-only
Per-mandate spend from the ledger: settled and reserved amounts (integer micro-units of the ledger currency), intent counts, refusal count, and whether the live budget counters equal a fresh replay of the ledger. Use before starting expensive work. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes the safety profile, so the trailing "Read-only" adds nothing. However, the description discloses real behavioral substance beyond annotations: the reconciliation check telling the agent whether live budget counters match a fresh ledger replay, which signals the tool doubles as a consistency diagnostic.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The field inventory, usage trigger, and safety note are packed into two short sentences with the scope front-loaded. The trailing "Read-only" is the one redundant clause since readOnlyHint already conveys it, costing a bit of tightness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema present, the description carries the burden of describing return values, and it does so by enumerating the reported quantities and their unit convention. It is complete enough to call correctly, though it does not specify types or formats for the individual counts.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, which is the baseline-4 case. The description does not need to document inputs, and it usefully clarifies the output unit convention (integer micro-units of the ledger currency) even though that belongs to the return value rather than a parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (per-mandate spend from the ledger) and enumerates exactly what is reported: settled/reserved amounts, intent counts, refusal count, and a counter-vs-replay consistency check. This is clearly distinct from write-oriented siblings like mandare_issue_mandate or mandare_kill without needing to open any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use before starting expensive work" gives a concrete trigger for invoking the tool, which is genuine guidance rather than a restatement. It stops short of naming alternatives or stating when not to call it (e.g., how it relates to mandare_gateway_health), so it is clear context without full when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mandare_certifyExport an integrity certificateARead-only
Build the selective-disclosure integrity certificate (chain valid · witnessed · anchored) a third party can verify WITHOUT ledger access. Optionally disclose specific entries by seq; undisclosed entries stay salted hashes. Requires a configured witness.
| Name | Required | Description | Default |
|---|---|---|---|
| disclose_seqs | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes this is a non-mutating operation, and the description is consistent with that, so there is no contradiction. It adds real context beyond the annotation: the prerequisite of a configured witness, and the specific disclosure semantics (only seqs you name are revealed; everything else remains salted hashes). It does not describe the certificate's return format or size limits, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: the artifact and its purpose first, disclosure behavior second, prerequisite last. No filler, no restatement of the tool name, and every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, read-only tool with no output schema, the description covers purpose, the gating prerequisite, and the disclosure/hashing behavior sufficiently to call it correctly. The main remaining gap is the shape of the returned certificate, but without an output schema that omission is a minor shortfall rather than a blocking one.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With one parameter and 0% schema description coverage, the description must carry the burden, and it partly does: disclose_seqs reveals specific entries by sequence number, is optional, and non-disclosed entries remain salted hashes. It adds meaningful semantics beyond the raw array-of-integers schema, though it omits the maxItems=100 cap and the positive-integer constraint.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb and resource: build/export a selective-disclosure integrity certificate. It defines the artifact's properties (chain valid, witnessed, anchored) and its key use case (third-party verification without ledger access), which makes the tool's role legible. It stops short of naming how it differs from the sibling mandare_verify, so sibling differentiation is only implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: an agent can infer this is the tool to reach for when an external party needs verifiable proof without ledger access. The precondition 'Requires a configured witness' is a genuinely useful gating rule, but there is no explicit statement of when to prefer this over mandare_verify or other siblings, nor any when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mandare_gateway_healthGateway healthARead-only
Read the running gateway door /healthz (halted state, card rail, witness gating). Requires MANDARE_GATEWAY_URL in the MCP server environment. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint, so the description's disclosure that MANDARE_GATEWAY_URL must be set in the server environment is genuinely useful setup context, and the hint at what /healthz reports (halted state, card rail, witness gating) adds substance. 'Read-only' merely restates the annotation, so it earns no credit there.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the verb and endpoint, with the prerequisite stated before the read-only note. Slight density of undefined domain jargon is the only cost.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does well to enumerate the health signals returned and to flag the required environment variable. Missing only a note on when a caller should prefer this over the other mandare status/verify tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing to document; the baseline for a parameterless tool applies. The description appropriately does not invent parameter guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read') and resource ('gateway door /healthz'), which is clearly distinct from sibling mutation tools like mandare_kill and the mandare_issue_* family. The parenthetical domain terms ('halted state, card rail, witness gating') are jargon that a reader may not fully parse, but the core purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No statement of when to call this versus mandare_verify or mandare_budget_status, and no exclusions or preconditions beyond the environment variable. The health-check framing implies usage but nothing is spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mandare_issue_mandateIssue a mandateA
Issue an owner-signed mandate (SD-JWT VC) for an agent: spend caps in WHOLE currency units (per transaction / day / task / total), optional human-approval threshold, validity window. The mandate file path is returned; point the gateway at it (MANDARE_MANDATE_PATH).
| Name | Required | Description | Default |
|---|---|---|---|
| total | No | ||
| per_tx | No | Max per single call/transaction, whole currency units | |
| per_day | No | ||
| purpose | No | ||
| currency | No | ||
| per_task | No | ||
| agent_did | Yes | The agent's did:key (from mandare_issue_passport) | |
| valid_hours | No | ||
| approval_above | No | Spends above this wait for an async human approval |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full burden. It usefully discloses that the mandate is owner-signed, that a file path is returned for gateway configuration, and that approval_above causes spends to wait for async human approval. It omits whether it overwrites an existing mandate, what owner signing authority is required, and whether issuance is idempotent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences that are front-loaded with the core action, then the cap model, then the operational follow-up step. No filler; every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 9 parameters, 33% schema coverage, no annotations, and no output schema, the description does meaningful work (units, tiers, approval, validity) but still leaves several parameters unexplained and does not cover auth/prerequisite or replacement behavior. Adequate but with clear gaps for a tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, so the description must compensate, and it does: it maps the cap tiers (per transaction / day / task / total) and stresses 'WHOLE currency units', which is critical numeric semantics not present in the schema. It also introduces the approval threshold and validity window. It leaves 'purpose' and 'currency' semantics unaddressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Issue an owner-signed mandate (SD-JWT VC) for an agent') and enumerates the content of the mandate: spend caps, approval threshold, validity window. This clearly distinguishes it from sibling issuance tools like mandare_issue_passport and mandare_issue_token, though it never explicitly names those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides a downstream usage step ('point the gateway at it (MANDARE_MANDATE_PATH)'), which is genuinely helpful operational context, but gives no guidance on when to use this versus mandare_issue_token or mandare_issue_passport, nor any prerequisites such as needing an agent passport first.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mandare_issue_passportIssue an agent passportA
Issue an Agent Delegation Credential (SD-JWT VC) for a NEW agent: owner + local attestation authority keys from the vault, fresh agent did:key, revocation slot registered on the ledger. The credential and the agent private key are written to the operator-configured home directory; the result references them by path and never contains key material.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_name | Yes | Label for the new agent (also the artifact file prefix) | |
| valid_days | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the side effects (credential + private key written to an operator-configured home directory, revocation slot registered on the ledger) and the security posture (result references artifacts by path and never contains key material). It omits required permissions and re-issuance/overwrite behavior, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two information-dense sentences that front-load the verb and resource, then enumerate the mechanics. Nothing is wasted, though the mechanics list is slightly packed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no output schema and no annotations, the description covers creation side effects, storage location, and the absence of key material in the result, which is what an agent needs to call it safely. Gaps are the undocumented valid_days semantics and error/idempotency behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%: agent_name is documented in the schema, but valid_days has no schema description and the description says nothing about it (validity window, defaults, bounds). The description adds context about artifacts, but it does not compensate for the undocumented parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Issue) and resource (Agent Delegation Credential / SD-JWT VC) and scopes it to a NEW agent, with concrete mechanics (vault keys, did:key, revocation slot). It is clear what the tool produces, but it never names or contrasts with the sibling issuance tools (issue_mandate, issue_token), so sibling differentiation is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'for a NEW agent' phrasing implies the appropriate context (provisioning, not re-issuing), which is useful. However, there is no explicit when-to-use/when-not guidance and no mention of alternatives among the three issue_* siblings, so the agent must infer routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mandare_issue_tokenIssue a scoped access tokenA
Mint a short-lived proof-of-possession token an agent presents to the gateway (TTL ≤ 30 minutes). The grant INCLUDING ITS ONE-TIME SECRET is written 0600 to the operator-configured home directory; the result references it by path only — hand the FILE to the agent process (e.g. @mandarelabs/sdk tokenCredentialsFromIssueJson), never paste its contents into chat.
| Name | Required | Description | Default |
|---|---|---|---|
| actor_did | Yes | The agent DID the token is scoped to | |
| mandate_id | Yes | ||
| ttl_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: TTL cap (≤30 minutes), a disk side effect (grant written 0600 to the operator-configured home dir), one-time secret semantics, and a return value that is a path reference rather than the secret. It omits caller privileges/auth requirements and revocation or idempotency behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the action and lifetime, then the file-handling caveat. The parenthetical SDK example is load-bearing, though the sentence grows long and slightly overloaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with no annotations and no output schema, the description covers the essentials an agent needs: lifetime limit, where the artifact lands, that the result is a path, and how not to leak the secret. Failure modes and the relationship to the mandate lifecycle are left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 33% (only actor_did is documented in-schema), so the description should compensate. It adds real value by tying the token to a grant and restating the TTL ceiling that matches the schema's maximum of 1800, but mandate_id is never explained beyond the word 'grant' and no format guidance is offered for either DID or mandate identifier.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a precise verb+resource ('Mint a short-lived proof-of-possession token') plus its scope and audience ('an agent presents to the gateway'). It is clearly distinguishable from siblings like mandare_issue_passport and mandare_issue_mandate without needing the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the operational context (presented to the gateway) and the correct consumption path (hand the FILE to the agent process via tokenCredentialsFromIssueJson). It does not name alternatives or state when to prefer this over issue_mandate/issue_passport, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mandare_killKill switchA
The LOCAL, offline, fail-closed kill: revoke ONE agent, ONE mandate (the agent survives, the permission slip dies), or --all (halts the whole door). Writes the revocation to the ledger; the gateway refuses the subject on its very next request. Bound virtual cards are revoked too. Kills only CLOSE doors — reversing one requires the operator (mandare reinstate).
| Name | Required | Description | Default |
|---|---|---|---|
| all | No | Revoke the whole door (every subject it governs) | |
| reason | No | ||
| agent_did | No | ||
| mandate_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so well. It discloses that the operation is LOCAL and offline, fail-closed, persists the revocation to the ledger, takes effect on the subject's very next request at the gateway, cascades to bound virtual cards, and is irreversible without an operator. That is exactly the side-effect and reversibility information an agent needs before firing a destructive call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the most important qualifier (LOCAL, offline, fail-closed) and the mode selection, then layers on effects and the reversal path. Every clause is load-bearing — ledger write, next-request enforcement, card cascade, operator-gated reversal — with no filler or restated boilerplate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-annotation, four-parameter mutation tool with no output schema, the description covers effects, scope, persistence, and reversibility thoroughly. Gaps remain: no mention of the 'reason' parameter's role, no guidance on what happens with conflicting or missing subject arguments (required count is 0), and no error/idempotency behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 25%, so the description must compensate and largely does: it explains the semantics of the three subject selectors (agent_did, mandate_id, all) and what each actually revokes, including that agent revocation kills the mandate rather than the agent identity. The 'reason' parameter remains undocumented in both schema and description, and the implicit mutual exclusivity of agent_did/mandate_id/all is only implied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (revoke/kill) against specific resources (one agent, one mandate, or the entire door) and clarifies the distinction between killing an agent versus its mandate. It is immediately distinguishable from siblings like mandare_issue_mandate or mandare_verify, which provision or check credentials rather than revoke them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear conditions for each mode: an agent, a mandate (with the key note that the agent survives but the permission slip dies), or --all for the whole door. It also names the reversal path (mandare reinstate) and the operator requirement, telling the agent when this is NOT the right recovery tool. It stops short of explicitly contrasting against siblings such as mandare_gateway_health or mandare_budget_status for non-revocation diagnostics.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mandare_verifyVerify the Mandare ledgerARead-only
Verify the local ledger: hash chain, door signatures, RFC 6962 tree head, spend trail and budget-counter replay, approval trail, and revocation state. With check_witness=true also checks the chain against the externally witnessed head history (catches truncation and rewrites). Returns the machine-readable verification report. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| check_witness | No | Also verify against the configured witness (detects truncation/rewrites) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description aligns ('Read-only'), so no contradiction. Beyond the annotation it adds real behavioral value: it names the specific verification checks performed, states that the witness mode detects truncation and rewrites, and notes the output is a machine-readable report. It does not cover failure semantics (what a negative report looks like) or any auth requirements, keeping it out of 5 territory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and the scope of checks, followed by the conditional behavior and return note. The enumerated check list is dense but each item is meaningful. Slightly heavy packing into a single long sentence keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so adequately ('Returns the machine-readable verification report'). For a zero-required-parameter verification tool with full schema coverage and a readOnly annotation, this is close to complete; only failure/error reporting specifics are absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and there is only one optional boolean, so the schema already documents check_witness fully. The description's phrasing ('catches truncation and rewrites') largely restates the schema's own parenthetical, adding little new syntax or semantics. Baseline 3 is appropriate when the schema carries the parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ('Verify the local ledger') and enumerates exactly what is verified: hash chain, door signatures, RFC 6962 tree head, spend trail, budget-counter replay, approval trail, revocation state. This scope is precise enough to distinguish it from siblings like mandare_certify or mandare_gateway_health, which are attesting/health tools rather than ledger verifiers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states the condition that changes behavior: set check_witness=true to verify against the externally witnessed head history and catch truncation/rewrites. That is actionable when-to-use guidance for the single parameter. It stops short of naming alternate tools (e.g., when to prefer mandare_certify) or stating when verification is unnecessary, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
mandare_budget_status - First observed
mandare_certify - First observed
mandare_gateway_health - First observed
mandare_issue_mandate - First observed
mandare_issue_passport - First observed
mandare_issue_token - First observed
mandare_kill - First observed
mandare_verify
TDQS
Scored across 8 tools
The three issue_* tools target clearly distinct artifacts (agent identity passport, spend mandate, session token), and the read-only tools split cleanly by target (ledger integrity, budget, gateway health). Minor potential overlap between mandare_verify and mandare_certify, since both concern ledger integrity, but their descriptions distinguish local verification from third-party disclosure certificates.
All tools share the mandare_ prefix and are largely verb-led (verify, issue_*, kill, certify), which reads predictably. Two names (budget_status, gateway_health) are noun phrases rather than verbs, a small deviation from the otherwise consistent verb pattern.
Eight tools is well-scoped for a delegated-authority gateway: identity issuance, mandate issuance, token minting, verification, budget inspection, revocation, certification, and health. Each tool earns its place with no obvious redundancy.
The issue/verify/budget/kill/certify lifecycle is mostly covered, but the kill description explicitly notes that reversal requires an operator-side 'reinstate' that has no MCP counterpart, and there is no tool to list or enumerate existing agents, mandates, or tokens. These are notable gaps an agent cannot work around from the tool surface alone.
Maintenance
Related MCP Connectors
Command your AI agents: verifiable passports, credential injection, full audit, revoke in 60s.
AI-agent trust infrastructure for discovery, authority, execution, verification, and receipts.
Bitcoin-anchored, tamper-evident audit log for AI agents — record, disclose and verify actions.
Issue Agent Passports and verify agent authority before value moves. Signed verification records.
Related MCP Servers
- AlicenseAqualityCmaintenanceCryptographic accountability for AI agents. Ed25519-signed receipts for every MCP tool call. Constraints, chains, AI judgment, invoicing, and local dashboard included.2413 npm1MIT
- AlicenseAqualityAmaintenanceProof-of-behavior enforcement for AI agents. Declare behavioral constraints, enforce at runtime, produce SHA-256 hash-chained audit trails. Supports covenants (permit/forbid/require), real-time verification, and cross-agent trust handshakes.440MIT

evermint-mcpofficial
AlicenseNot gradedqualityDmaintenanceTamper-evident receipts for AI agent actions. The notary layer for agent-to-agent transactions.39 npm1MIT
emilia-mcp-serverofficial
AlicenseAqualityAmaintenanceThe accountability layer for AI agents — a named human's signed yes before an agent does anything irreversible (payment, record change, deploy), then an offline-verifiable Trust Receipt. Apache-2.0, formally verified.3612Apache 2.0