Skip to main content
Glama
attestedintelligence

AGA-mcp-server

AGA - Attested Governance Artifacts

Verifiable decision records for AI agents: each governed tool-call decision becomes a signed, hash-chained receipt, exported in evidence bundles anyone can verify offline against the published format.

npm PyPI License: MIT npm provenance

Status: published to npm; this release carries SLSA build provenance (check it: npm audit signatures). The server tools and the aga-proxy emit the canonical SEP evidence bundle, verifiable offline by the published @attested-intelligence/aga-verify and the reference verifier aga-receipt-spec/verify/verify-sep.mjs. Since 3.2.0 the verifier is algorithm-agile and ships a post-quantum profile: v1 Ed25519-SHA256-JCS (the default the gateway emits) and v2 ML-DSA-65+Ed25519-SHA256-JCS (a NIST FIPS-204 ML-DSA-65 + RFC-8032 Ed25519 composite, both must verify), selected per-bundle by the algorithm field with a VERIFIED / FAILED / UNSUPPORTED_PROFILE trichotomy. Pre-3.0 releases (a legacy continuity-chain bundle that does not verify under the SEP verifier) are deprecated; use ^3.0.0. Claim scope and residual attack surface are documented honestly in THREAT_BOUNDARY.md. 3.5.0 (2026-08-29) changes one behavior: an artifact's TTL now fails closed — on expiry the portal terminates and a further measurement is refused, where earlier releases degraded and kept measuring. If you depend on the old post-expiry behavior, pin 3.3.3. 3.6.0 (2026-09-18) changes no existing behavior: it makes aga-proxy honour AGA_GATEWAY_KEY / AGA_GATEWAY_KEY_FILE, which it had silently ignored. See CHANGELOG.md.

# This package IS the AGA MCP server (TypeScript, runs over stdio). Use it from any MCP client:
npx -y @attested-intelligence/aga-mcp-server

A Python companion SDK (aga-governance) is documented in the Python SDK section below.

Verify this yourself (don't take our word)

You do not have to take any of this on faith. The repo ships the reference verifier, the canonical vectors, and sample bundles, so you can check one offline right now, with no network and no callback to us:

git clone https://github.com/attestedintelligence/aga-mcp-server
cd aga-mcp-server
# A canonical SEP bundle verifies; a one-byte-tampered copy is rejected.
node aga-receipt-spec/verify/verify-sep.mjs fixtures/valid_minimal.json   # OVERALL: VERIFIED (integrity only; no key pinned)
node aga-receipt-spec/verify/verify-sep.mjs fixtures/tampered.json        # OVERALL: FAILED

The published @attested-intelligence/aga-verify CLI renders the identical verdict, and npm run conformance:cross-stack (first: npm run build && npm --prefix independent-verifier run build) proves six v1 verifier configurations — spanning three independent toolchains (JavaScript, Go, and Python, including a pure-stdlib, no-third-party-crypto path) — agree on the 54 object-level cases, and the five file-parsing verifiers agree on the 7 raw-byte/file-parse cases (61 total). The in-server engine is library-only, receiving parsed objects rather than raw file bytes, so it does not run the file-parse cases; six configurations do not agree on all 61 and this no longer claims they do. npm run conformance:cross-stack-v2 proves two genuinely independent-language oracles (@noble/JS and CIRCL/Go) agree on the v2 composite corpus. For a full trust-free reproduction (build the package yourself, reproduce the published tarball byte-for-byte, re-run every gate), see the REVIEWER_GUIDE.md (a command-by-command self-service path), REPRODUCIBILITY.md, and the step-by-step SKEPTICAL_AUDITOR.md. This release carries SLSA build provenance, checkable with npm audit signatures.

Related MCP server: Proofpane

What This Does

This is built for teams shipping agentic-AI products into financial services and insurance, at the moment a customer's vendor-risk, model-risk, or internal-audit review asks what your agent did and how anyone would know.

Tool calls routed through the AGA gateway are evaluated against the operator's policy, and each decision (PERMITTED or DENIED) is recorded as a signed, hash-linked governance receipt, except the calls known issue 7 below describes as refused without one. aga-proxy also signs the SHA-256 of its policy's canonical JSON into every receipt; see KNOWN_LIMITATIONS.md for what that field binds. Receipts are collected into evidence bundles that anyone holding the published format and the public key can verify offline, with no callback to us.

Record. Prove. Verify.

Scope: a verified bundle proves the integrity of the receipts present: each is authentic, correctly ordered, Merkle-included, and (when a key is pinned) provenance-bound. It does not prove non-omission (that every action the agent took was logged); completeness is bounded by the tamper-evidence of the interception point, which is outside the bundle. See KNOWN_LIMITATIONS.md for the full honest boundary, and THREAT_BOUNDARY.md for the per-field detail.

Use with Claude Desktop

Add to your Claude Desktop MCP config (claude_desktop_config.json):

{
  "mcpServers": {
    "aga": {
      "command": "npx",
      "args": ["-y", "@attested-intelligence/aga-mcp-server"]
    }
  }
}

Claude can then seal artifacts, measure integrity, generate evidence bundles, and verify them offline through natural language.

Persist the signing key (do this first)

By default the gateway signs with an ephemeral key that rotates on every restart. That is fine for a first look, but evidence-bundle provenance cannot be pinned across restarts (and the server warns about it on stderr). Set one stable 64-hex Ed25519 seed so provenance stays pinnable:

Since 3.6.0 this applies to both binaries. aga-proxy reads the same two variables through the same resolver and prints the active public key at startup so you can pin it out of band; --ephemeral makes a throwaway key a stated choice. In 3.5.0 and earlier aga-proxy ignored both variables silently — a key you set had no effect and no warning was printed, so evidence from such a proxy is integrity-verifiable but not provenance-pinnable across restarts. See DEPLOYMENT.md §2.

# generate a seed once (32 random bytes, hex)
node -e "console.log(require('node:crypto').randomBytes(32).toString('hex'))"

Provide it via AGA_GATEWAY_KEY, or AGA_GATEWAY_KEY_FILE (a path to the seed). In Claude Desktop, add an env block:

{
  "mcpServers": {
    "aga": {
      "command": "npx",
      "args": ["-y", "@attested-intelligence/aga-mcp-server"],
      "env": { "AGA_GATEWAY_KEY": "<your-64-hex-seed>" }
    }
  }
}

Keep the seed secret and out of version control; see DEPLOYMENT.md for key handling.

MCP Tools (15)

Category

Tools

Identity

get_server_info, get_portal_state

Lifecycle

init_chain, attest_subject, revoke_artifact

Measurement & decision

measure_integrity, measure_behavior, verify_chain

Evidence

generate_evidence_bundle, verify_bundle_offline

Privacy

request_claim, list_claims

Delegation

delegate_to_subagent

Audit

get_receipts, get_chain_events

measure_behavior is detective-only by default: it observes tool-usage patterns and records a signed, provable drift finding, but does not block. Enforcement (drift → quarantine) is opt-in via enforce=true and off by default. Hard governance decisions (PERMITTED/DENIED) are made by the portal/PEP, not the behavioral monitor.

Quick Start: verify a bundle offline

A bundle this package emits (via the generate_evidence_bundle MCP tool) is a canonical SEP bundle. Verify it offline, with no network and no callback to us:

# Published verifier CLI — ships on npm, nothing to clone. Pin the gateway key (from get_server_info) to prove provenance.
npx -y @attested-intelligence/aga-verify evidence-bundle.json --pubkey <gateway-public-key>

# Or, from a clone of this repo, the zero-dep reference verifier (Node 18+) renders the identical verdict:
node aga-receipt-spec/verify/verify-sep.mjs evidence-bundle.json --pubkey <gateway-public-key>

The published @attested-intelligence/aga-verify CLI is the shipped path (the older forgeable 1.0.0 is deprecated); the reference verify-sep.mjs renders the identical verdict from a repo clone. Without --pubkey you get an integrity-only result (issuerVerified=false); pin the key to also prove who issued it. See THREAT_BOUNDARY.md §3.7. A hosted browser verifier is linked under Links.

The reference §6 algorithm is implemented in three languages: JavaScript (aga-receipt-spec/verify/verify-sep.mjs), Go (verify.go, stdlib crypto/ed25519), and Python (verify.py, pure-stdlib RFC-8032 Ed25519). A cross-stack harness (npm run conformance:cross-stack; first: npm run build && npm --prefix independent-verifier run build) proves all three, plus the in-server engine and aga-verify, render identical verdicts on the canonical vectors (valid, adversarial, and every small-order forgery). The v2 composite profile (ML-DSA-65+Ed25519-SHA256-JCS) is held to the same bar by a second harness (npm run conformance:cross-stack-v2): a @noble/JavaScript engine and a CIRCL/Go oracle, two genuinely independent toolchains, render identical verdicts on the pinned v2 corpus, and the reference v1 verifier (verify-sep.mjs/verify.py/verify.go) returns UNSUPPORTED_PROFILE (exit 3) on a v2 bundle, signalling "profile not implemented" rather than a misleading "invalid". (The published aga-verify CLI does not implement this profile trichotomy: on a v2 bundle it returns FAILED (exit 1). Use exit 3 as the unsupported-profile signal only with the reference verifiers.)

Check-name mapping across implementations

The JS reference verifier and the Python SDK (aga-governance) decompose the same seven-check verification differently. Overall verdicts and exit codes agree on all 61 conformance-corpus cases as the cross-stack harness feeds them (object-level cases re-serialized, so float spellings arrive as integers; measured on aga-governance 0.3.2 on 2026-09-25) and on the 10 cells re-proven 2026-07-01 (pristine and tampered bundles with unpinned, correct and wrong keys). On the literal file bytes of the corpus's float-spelled leaf_index case (0.0), aga-governance 0.3.2 reports FAILED where the JS reference, aga-verify, Go and Python reference verifiers report VERIFIED; see https://attestedintelligence.com/spec. The sub-check that reports a given tamper can differ:

JS reference check

Python result field

What it covers

structural

algorithm_valid + parts of bundle_consistent

algorithm id, key well-formedness, receipt/proof counts

receipt_signatures

receipt_signatures_valid

Ed25519 over canonical receipt bytes

chain_and_ordering

chain_integrity_valid

prev-leaf linkage, canonical non-decreasing timestamps (ids are not ordering fields and are not checked)

merkle_and_bijection

merkle_proofs_valid

leaf recompute, single-root walk, index bijection

signed_checkpoint

checkpoint_valid

gateway-signed root + count + chain-head binding

envelope_consistency

envelope_consistent

envelope gateway_id, generated_at, merkle_root vs signed content (bundle_id, schema_version, the envelope policy_reference and offline_capable are unsigned and unchecked)

gateway_key_match (with --pubkey)

gateway_key_match / provenance

pinned issuer key

Known decomposition difference: the JS reference recomputes every Merkle leaf from full receipt content, so a receipt-signature tamper also fails merkle_and_bijection; the Python verifier surfaces the same tamper in receipt_signatures_valid, chain_integrity_valid, and bundle_consistent while its merkle_proofs_valid can remain true. Neither is looser: the bundle fails in both stacks, exit 1. A --pubkey KEY that is not 64 lowercase hex characters is a usage error (exit 2) in the JS reference, aga-verify and the Python SDK, and a 64-hex pin that is not a valid curve point is honored, fails to match, and fails the bundle (exit 1). Written as --pubkey=KEY, the pin is ignored by the JS reference, aga-verify, verify.go and v2/verify-v2.go, and by verify.py when it follows the bundle path (integrity only, exit 0), and read by the Python SDK. Other verifiers differ as well. The in-server engine (the package's ./verify export, which verify_bundle_offline calls) treats a pin that is not a well-formed key for the bundle's profile (for a v1 bundle, a small-order point or a non-canonical encoding) as no pin, and returns VERIFIED with pinned: false. v2/verify-v2.go does the same and prints integrity only; no key pinned (exit 0). A 64-hex value that is not a curve point counts as well-formed, so both take it as a pin and the bundle fails. The Go and Python reference verifiers in aga-receipt-spec/verify/ treat a pin that is not 64 lowercase hex the same way and print integrity only; no key pinned (exit 0). Read pinned before taking a VERIFIED as provenance; in CI, pass the key after a space and check that the output says provenance verified. A --pubkey given with no value is also treated as no pin (exit 0, integrity only) by aga-verify, verify-sep.mjs, verify.py, verify.go and v2/verify-v2.go, and is a usage error in the Python SDK. Other differences concern the bundle rather than the pin, and https://attestedintelligence.com/security lists the ones measured, including which algorithm labels each verifier leaves unchecked; outside the conformance corpus the verifiers differ in both directions. Two examples, where the failing side fails closed: a proof leaf_index spelled as an integral float (1.0) reports FAILED in aga-governance 0.3.2 and VERIFIED in the others, and object keys outside the Basic Multilingual Plane (possible only in a non-string field value, which no shipped producer emits) sort differently in verify.py, verify.go, v2/verify-v2.go and aga-governance than in the JavaScript verifiers, so such a bundle reports VERIFIED in JavaScript and FAILED in Go and Python. These wait for the next reviewed release.

How It Works

AI Agent                  AGA Gateway                    Verifier
   |                          |                              |
   |-- tools/call ----------->|                              |
   |                    [Evaluate Policy]                    |
   |                    [Sign Receipt]                       |
   |                    [Chain to Previous]                  |
   |<-- PERMITTED/DENIED -----|                              |
   |                          |                              |
   |                    [Export Bundle]                       |
   |                          |--------- evidence.json ----->|
   |                          |                  [Verify Signatures]
   |                          |                  [Verify Chain + Order]
   |                          |                  [Verify Merkle Tree]
   |                          |                  [Verify Signed Checkpoint]
   |                          |                  [PASS / FAIL]

MCP Governance Proxy

Run AGA as a proxy in front of an MCP server that it starts as a stdio child process (the hardened default), or one it reaches with a plain JSON-RPC POST (--upstream-url; no Streamable HTTP session or SSE handling). The proxy's agent port speaks newline-delimited JSON-RPC 2.0 over raw TCP, not stdio or Streamable HTTP. A stdio MCP client needs a relay you provide (a few lines that pipe stdin to the port and the port to stdout); none ships. A scripted client can speak that framing directly. Every tools/call request with a non-empty string tool name and arguments the proxy can canonicalize is evaluated against the policy and produces a signed receipt, except the calls that known issue 7 below describes as refused without one. Other methods that are not benign are forwarded with a signed passthrough receipt and are not policy-evaluated, and benign protocol methods (initialize, initialized, ping, tools/list, prompts/list, resources/list, resources/templates/list, logging/setLevel, completion/complete and notifications/*) produce no receipt (THREAT_BOUNDARY.md section 3 item 2). Read the known issues below before you expose the port.

# Start the proxy (the `aga-proxy` bin) in front of an upstream MCP server.
# stdio upstream = the hardened default (the upstream is a child process, not network-reachable).
npx -p @attested-intelligence/aga-mcp-server aga-proxy start \
  --upstream "npx -y @modelcontextprotocol/server-filesystem /tmp/test" --profile permissive

permissive records each tools/call it evaluates (known issue 7 describes the exceptions) and denies nothing on policy grounds. standard and restrictive allow only generic example tool names, so they deny every tool this example server exposes; to permit some of your server's tools and deny the rest, pass a --policy file that names them.

Exporting the evidence bundle from a running proxy

The proxy records receipts in its own process and keeps the SEP ledger in memory. To make that live ledger reachable from a separate shell, aga-proxy start opens a loopback-only control channel — an HTTP listener bound to 127.0.0.1 (never a routable interface), on its own port (default 18801, override with --control-port), distinct from the agent-facing proxy port (18800). It exposes only read routes (/export, /status, /receipts); nothing on it mutates policy or state. It does not check a request's Host or Origin header, so a web page in a browser on the same host can read its responses through DNS rebinding unless the browser blocks it (known issue 12). The proxy writes the chosen control port to ~/.aga-proxy/control.json alongside proxy.pid.

A separate aga-proxy export invocation reads that file and fetches the same signed bundle the running proxy would emit:

# Terminal A — start the proxy in front of an upstream MCP server
npx -p @attested-intelligence/aga-mcp-server aga-proxy start \
  --upstream "npx -y @modelcontextprotocol/server-filesystem /tmp/test" --profile permissive

# (First, drive at least one tools/call through the proxy from your MCP client — an empty
#  ledger has no receipts to checkpoint, and the export reports there is nothing to export.)
# Terminal B — export the live ledger from a different shell, then verify it offline
npx -p @attested-intelligence/aga-mcp-server aga-proxy export -o evidence.json
npx -y @attested-intelligence/aga-verify evidence.json --pubkey <gateway-public-key>

Export and verify before you stop the proxy: aga-proxy stop ends the process without exporting, and the in-memory chain goes with it (known issue 9 covers export time and bounding the chain).

If no proxy is running, aga-proxy export prints no running proxy found; start it first, or export from within the session and exits non-zero — it never emits an empty or placeholder bundle. Within the MCP server session you can also call the generate_evidence_bundle tool and save the returned JSON.

In-memory ledger: the exported bundle is the durable cryptographic record, but the live in-process chain does not survive a proxy restart. This flow makes the live ledger reachable from another process; it does not add cross-restart persistence, which needs the persistent (SQLite) backend and remains roadmap (see KNOWN_LIMITATIONS.md).

The proxy intercepts tools/call requests, evaluates them against the loaded policy (a JSON file or a built-in profile; the SHA-256 of its canonical JSON is signed into every receipt), and generates a signed SEP receipt for every decision (except the calls that known issue 7 below describes as refused without one). Permitted calls are forwarded to the downstream server; denied calls return an MCP error and never reach it. Every decision is hash-linked and checkpoint-bound into a tamper-evident bundle. (Methods other than tools/call aren't policy-evaluated, but non-benign ones are recorded as signed passthrough receipts for auditability, and a library caller can pass a method denylist (denyMethods) to reject them; the aga-proxy CLI has no flag for it; see THREAT_BOUNDARY.md §3.2.)

Three built-in policy profiles:

  • permissive - audit_only: denies nothing on policy grounds and records each tools/call it evaluates (default); the fail-closed refusals below and known issue 7 still apply

  • standard - an allowlist of ten generic example tool names (filesystem_read, shell_execute, web_search and others) with rate limits, and substring denials on two of them; every other tool is denied, so a real server's tools need a --policy file

  • restrictive - an allowlist of three generic example tool names with lower rate limits and a path prefix on one; every other tool is denied

Because the default (permissive) is audit-only, starting with an audit_only policy prints a loud stderr banner stating that every call is permitted and recorded and no call is denied in that mode. No call is denied on policy grounds, but the proxy still refuses, fail-closed, a tools/call with no tool name or with arguments it cannot canonicalize (nested past 100 levels, for example), and signs a DENIED receipt for each; a name of 0, false, null or an empty string counts as no name. Policy denial needs --profile standard or restrictive, or a --policy file in allowlist or denylist mode. In denylist mode a policy denies each tool it lists as an object whose allowed value is missing or falsy (false, 0, null or an empty string); any other allowed value allows the tool, including the string "false", and so does a listed entry that is itself false, 0, null or an empty string rather than an object. In 3.6.0 through 3.6.2 it also denies an unlisted tool named like a built-in object property, such as constructor. It applies the rate limits of the listed tools it allows; path and pattern rules apply in allowlist mode only and check only top-level string arguments (known issue 10). A --policy file in audit_only mode permits every call; one with any other mode, or none, denies every tools/call, and in 3.6.0 through 3.6.2 one in allowlist or denylist mode whose constraints member is missing or null refuses, with no receipt and no response, every tools/call that has a tool name and arguments the proxy can canonicalize (known issue 7). An unrecognized --profile value is a hard error (exit 2 listing the valid names), never a silent fallback to permissive.

Verification (canonical SEP 3.0; normative §6 algorithm in aga-receipt-spec/verify/verify-sep.mjs)

  1. Structural floor - Bundle declares Ed25519-SHA256-JCS, public key well-formed (all small-order encodings + non-canonical y ≥ p rejected), receipts.length > 0, proof count = receipt count

  2. Receipt Signatures - Ed25519 over JCS-profile canonical JSON, sorted-key (signature field excluded)

  3. Chain + ordering - Each receipt's previous_receipt_hash = leaf of the preceding receipt; non-decreasing timestamps

  4. Merkle Proofs - Recompute every leaf from receipt content, walk siblings/directions to one root, leaf indices form the complete 0..N-1 bijection

  5. Signed checkpoint - Verify the gateway-signed checkpoint binding merkle_root, leaf_count, and chain head (this makes the no-prefix construction truncation-safe)

  6. Provenance (when a key is pinned) - public_key == expected key; otherwise integrity-only is reported

Cryptographic Primitives

Primitive

Purpose

Ed25519

Receipt signatures

SHA-256

Hash chaining, Merkle trees, leaf computation

JCS-profile (sorted-key canonical JSON)

Deterministic signing (canon is byte-compatible with the reference verifier)

Merkle Trees

Binding all receipts to a single verifiable root

Live Gateway

A demo gateway is deployed on Cloudflare Workers (a separate deployment that may track its own version; treat it as a convenience mirror, and always verify what it returns offline against a pinned key, not as the canonical artifact):

# Check status
curl https://aga-mcp-gateway.attested-intelligence.workers.dev/health

# Export evidence bundle
curl https://aga-mcp-gateway.attested-intelligence.workers.dev/bundle -o evidence-bundle.json

Python SDK

Status, rechecked against PyPI on 2026-09-23. aga-governance 0.3.1 fixed the depth-bomb crash: on a deeply nested receipts payload the verifier returns a FAILED verdict instead of raising, and every later release carries the fix. 0.3.0 was yanked for that crash; 0.2.6 raises on the same input and is not yet yanked, so any version specifier that excludes 0.3.1 and later (~=0.2.0 or <0.3.1, for example) still installs it. Install 0.3.1 or later before you verify untrusted bundles with the Python SDK. The JavaScript reference verifier and the @attested-intelligence/aga-verify CLI are unaffected.

pip install "aga-governance>=0.3.1"
from aga import AgentSession

with AgentSession(gateway_id="my-gateway") as session:
    session.record_tool_call(
        tool_name="search_web",
        decision="PERMITTED",
        reason="tool in allowlist",
        request_id="req-1",
    )
    bundle = session.export_bundle()
    result = session.verify()
    assert result["overall_valid"]

Test Suite

Automated tests across TypeScript and Python, plus a conformance corpus:

  • TypeScript MCP server: 428 automated tests (vitest), including provable-denial and behavioral-monitor regressions

  • SEP conformance corpus: npm run test:conformance (valid → VERIFIED, negatives → FAILED)

  • Python companion SDK: the separately-published aga-governance PyPI package (install + smoke-checked here; its full pytest suite runs from the source tree). The smoke check imports the package and prints its version. It does not exercise the verifier.

npm test                              # TypeScript tests (vitest)
npm run test:conformance              # SEP conformance corpus
pip install aga-governance && python -c "import aga; print(aga.__version__)"   # Python SDK smoke check

Benchmarks

Receipt-format determinism is reproducible here: npm test runs the cross-language vectors, and npm run conformance:cross-stack (first: npm run build && npm --prefix independent-verifier run build) shows the six v1 verifier configurations (across three independent toolchains: JS, Go, Python) agree on the 54 object-level cases of the canonical 61-case corpus — the remaining 7 are raw-byte/file-parse cases run by the five file-parsing verifiers, since the in-server engine never receives raw bytes — while npm run conformance:cross-stack-v2 shows the two independent-language v2 oracles agree on the composite corpus.

Project Structure

src/
  sep/                 # Canonical SEP evidence engine: single source of truth (canon, merkle, receipt, checkpoint, bundle, verify)
  core/                # Governance primitives (portal, artifact, attestation, disclosure, delegation, behavioral) + internal continuity-chain profile
  crypto/              # Internal continuity-chain crypto: Ed25519 (node:crypto), SHA-256/blake2b, salt
  proxy/               # MCP governance proxy (transparent interception + policy evaluation; emits SEP bundles)
  middleware/          # Governance PEP wrapper (records a signed PERMITTED/DENIED receipt per governed call)
independent-verifier/  # @attested-intelligence/aga-verify: standalone SEP verifier, zero AGA imports
scenarios/             # Demo scenarios (SCADA, autonomous vehicle, AI agent) that emit SEP bundles
tests/                 # TypeScript test suite (428 automated tests)

Known issues in 3.6.0 to 3.6.2 and the published verifiers

3.6.1 and 3.6.2 change only the documentation and the version number; the runtime is 3.6.0's. Items 1 to 4 were reproduced on 2026-09-23 on @attested-intelligence/aga-mcp-server 3.6.0 installed from npm, and concern the aga-proxy gateway. Item 5, added 2026-09-25, concerns the verifiers and was reproduced on 2026-09-25 on the current releases. Item 6, also added 2026-09-25, concerns aga-proxy with an HTTP upstream and was reproduced on 2026-09-25 on 3.6.2. Item 7, also added 2026-09-25, concerns tools/call messages that aga-proxy refuses without a receipt and was reproduced on 2026-09-25 on 3.6.2 (its oversized-message case was added and reproduced on 2026-09-26). Items 8 to 12, added 2026-09-26, concern non-ASCII text, the cost of exporting evidence, what policy constraints check, oversized tool results and the control channel; each was reproduced on 2026-09-26 on 3.6.2, as were the memory and Windows port cases added to item 1 that day. The same list is kept at https://attestedintelligence.com/security.

  1. The agent port listens on every network interface, with no authentication. Anyone who can reach the host on that port can send governed calls through the proxy. The proxy also sets no limit on the number of connections, so although each connection's unfinished message is capped (item 7), the memory they hold together is not: on 3.6.2, ten connections that each sent 7.5 MiB without ending a message raised the proxy's working set from 68 MiB to 392 MiB while they stayed open, and a governed call on another connection was still answered. On Windows, a process running under the same user account as the proxy can bind 127.0.0.1 on the same port while the proxy listens, and a local client that connects to 127.0.0.1 then reaches that process instead of the proxy, with no policy check and no receipt (measured on 3.6.2 with both processes under one account). Block inbound traffic to the port in the host firewall, or admit only the agent with network policy; on Windows, run nothing untrusted under the proxy's account, and check while the proxy runs that its process is the only listener on the port. The control port is bound to loopback; see item 12.

  2. Two clients that reuse a JSON-RPC id through one proxy can receive each other's tool results. Workaround, measured on 3.6.0: give each client its own id range, or run one proxy per client. With disjoint ids, every result reached the client that asked for it.

  3. When the gateway key is supplied through AGA_GATEWAY_KEY (with or without --ephemeral) or AGA_GATEWAY_KEY_FILE, the stdio upstream inherits that variable (the seed, or the file's path), so the upstream sits inside the key's trust domain. Workaround, measured on 3.6.0: run without either variable. The upstream then sees neither, but the proxy signs with a per-process key that cannot be pinned across restarts.

  4. The --upstream-url mode forwards raw JSON-RPC over HTTP POST with only a content-type header. It does not implement the MCP Streamable HTTP transport (the Accept header, the request metadata headers and event-stream handling), so a spec-conformant HTTP MCP server rejects its requests. Workaround: bridge to the server over stdio.

  5. A bundle file can repeat a field name anywhere: in a receipt, in the checkpoint or in the envelope. For example, a forged "decision": "PERMITTED" can be placed ahead of the signed "decision": "DENIED", or a forged checkpoint leaf_count ahead of the signed one. The published verifiers (aga-verify 2.2.2, aga-governance 0.3.2, the verifier in this package, and the reference verifiers in aga-receipt-spec/verify/) and the site's /verify page read the last occurrence, and when it holds the genuine value they report VERIFIED, with provenance when the key is pinned. They do not reject the file, so a tool or a person reading the first occurrence can see a value that was never signed. Measured on the public sample bundle, pinned to the sample key: with the repeated name inserted in a receipt, in the checkpoint or in the envelope, aga-verify 2.2.2, aga-governance 0.3.2, the verifier in this package and /verify report VERIFIED, and a real change of the checkpoint value fails. Workaround: treat the verifier's parsed output as the record's content, and reject or flag files with repeated field names before displaying them. A strict rejection of repeated field names is planned for the reviewed release.

  6. A message can repeat the "method" member when aga-proxy has an HTTP upstream (--upstream-url). With tools/call first and another method last, aga-proxy reads the last copy, so it never checks the tool call against the policy, and it forwards every method other than tools/call to the HTTP upstream as the exact bytes it received. An upstream whose JSON parser keeps the first copy of a repeated name then runs the tool call, even one the policy denies. The bundle holds no receipt for that call: nothing at all when the last method is one the proxy passes through without a receipt (such as ping, initialize, a list method or a notification), and otherwise only a passthrough receipt that names the last method. An upstream that keeps the last copy handles the message as the method the proxy read. The stdio upstream, the default, is not affected: the proxy re-serializes each message before writing it, so the upstream receives one method. Measured on 3.6.2 from npm on 2026-09-25, with the restrictive profile and with the default permissive profile. Workaround: keep the stdio default, or have the HTTP upstream reject any message that repeats a member name. A strict rejection of repeated member names in the proxy is planned for the reviewed release.

  7. aga-proxy does not record every tools/call it refuses. It signs a tools/call's receipt before it forwards the call, so with a stdio upstream a call it cannot record never reaches the tool (for an HTTP upstream, see item 6), but such a call leaves no receipt. These cases were measured on 3.6.2 from npm on 2026-09-25, each with no receipt and no response to the client: a tools/call whose tool name is a non-zero number, an array or an object (a name of 0, false, null or an empty string counts as no name and gets a DENIED receipt and an error); one whose tool name or string id holds an unpaired surrogate (the JSON escape \ud800, for example); under a --policy file in allowlist or denylist mode whose constraints member is missing or null, every tools/call with a tool name and arguments the proxy can canonicalize; and, under an allowlist file, a call it would otherwise permit that carries a string path when that tool's path_prefix is neither a string nor false, 0 or null. The proxy starts with such a policy file, and it reports each refusal listed above only on its own stderr. A message sent as a JSON-RPC batch array or without "jsonrpc": "2.0" is refused differently: the client gets an error, and there is no receipt. A message of 8,388,608 characters or more, not counting its newline (UTF-16 code units, about 8.4 million), also gets an error and no receipt, and the proxy then closes the connection, dropping any reply still due on it. The limit counts input not yet split into messages, so a message just under the limit can be refused the same way when the read that completes it also carries enough of the next message to pass the limit; whether that happens depends on where the reads fall. On 3.6.2, when a 100-character message, a message 58 characters under the limit and a 100,000-character message were sent in one write, the first was answered, the second got the error and the connection closed; sent without the 100,000-character message, or without the 100-character one, every message was forwarded. These are the cases measured, not a proof that no other input does the same. Workaround: give every policy file a constraints object whose path_prefix values are strings, and have the client time out a call that gets no reply. A DENIED receipt and an error for a malformed tool name, and a check of the policy file at startup, are planned for the reviewed release.

  8. aga-proxy can alter non-ASCII text whose bytes are split between two reads. It decodes each chunk it reads from the agent's connection, and from a stdio upstream's output, on its own, so a character split between two chunks becomes one or more replacement characters (U+FFFD). A tool call's arguments can then reach the upstream altered, and the receipt's arguments hash is the hash of the altered arguments; a large non-ASCII result from a stdio upstream can reach the agent altered. An HTTP upstream's result is decoded whole. Whether a split happens depends on how the bytes arrive, so any message with non-ASCII text can be affected, and large ones more often. Measured on 3.6.2 from npm on 2026-09-26: a forced split inside "é" reached the upstream as two replacement characters, and a result of 200,000 "€" reached the client with 15 replacement characters in it. Workaround: send JSON whose non-ASCII characters are written as \uXXXX escapes, so every byte the proxy reads from the agent is ASCII, and have a stdio upstream do the same; a forced split of the escaped message then arrived intact. A fix is planned for the reviewed release.

  9. Exporting the evidence bundle takes time that grows with the square of the number of receipts, and aga-proxy handles nothing else while it runs: every governed call waits until the export ends. A call forwarded to a stdio upstream that has not answered when an export starts gets a timeout error if the export ends more than 30 seconds after the call was forwarded, although the upstream ran it and its receipt says PERMITTED, so an agent that retries can run the tool twice. Measured on 3.6.2 from npm on 2026-09-26 through the control channel's GET /export: 2.8 seconds at 1,000 receipts and 17.4 seconds at 2,500, with a tools/call sent during the export waiting as long; at 4,000 receipts an export took 43.8 seconds, and a call forwarded just before it, which the upstream answered in 2 seconds, got the timeout error. At 1,000 to 4,000 receipts a compact bundle took about 1.7 to 1.9 KB per receipt, rising with the count. An audit the same day measured 88 to 113 seconds at 5,000 receipts. Verification time grows close to linearly. Workaround: each export covers every receipt since the proxy started and does not shorten the chain, so bound the chain by restarting the proxy on a schedule: pause the agents, export and verify, then restart. A restart begins a new chain that is not linked to the last one and resets the rate-limit counts; with a per-process key (--ephemeral, or neither AGA_GATEWAY_KEY nor AGA_GATEWAY_KEY_FILE set) it also begins a new signing key, printed at startup, that cannot be pinned across restarts (item 3). Export outside busy periods, and export and verify before any stop, because the live chain is kept in memory and a stop loses receipts not yet exported. A fix that leaves the bundle's bytes unchanged is planned for the reviewed release.

  10. Policy constraints check less than their names suggest. A path_prefix is checked only when the value under the checked key (path, or the keys a rule lists in path_keys) is a string, so the same path sent inside an array or an object is not checked. denied_patterns match case-sensitively and only in top-level string arguments, so an uppercased command, or one inside an array or a nested object, is not matched. A constraint key the proxy does not recognise, such as a misspelling, is ignored without a warning, and a value of the wrong JSON type is not rejected: allowed: "false", a string, allows the tool; in denylist mode a tool listed as false, null or 0 instead of an object is allowed; and a non-empty path_keys string instead of an array makes the check read each character as a key name, so the intended key goes unchecked. Rate limits count per tool name across the whole proxy, shared by every client, and in allowlist mode the limit is checked before the path and pattern rules, so a call those rules deny still uses up a slot. Measured on 3.6.2 from npm on 2026-09-26 with allowlist and denylist policy files: a path_prefix of /home denied "/etc/passwd" and forwarded ["/etc/passwd"]; a denied pattern of rm -rf denied rm -rf / and forwarded RM -RF / and the same command inside an array; a rule spelled denied_pattern denied nothing; allowed: "false" forwarded the call in both modes; a denylist entry of false forwarded the call; path_keys: "path" forwarded /etc/passwd past a /home prefix; and, under a limit of 2 a minute, two calls denied by a /home prefix left a third call, to a path under /home, denied for the rate limit, while after one such denial that call was forwarded. Workaround: treat path and pattern rules as a convenience rather than a boundary, restrict paths in the upstream server itself, and check a policy file's keys against the constraint names in dist/proxy/types.d.ts. Checks that fail closed on these inputs, and a check of the policy file at startup, are planned for the reviewed release.

  11. A stdio upstream's response whose JSON line is 8,388,608 characters or more (UTF-16 code units, counting JSON escaping but not its newline) is dropped. The call already has a PERMITTED receipt, and the agent gets a timeout error after 30 seconds, so an agent that retries can run the tool twice; the proxy reports the drop only on its own stderr. The limit counts upstream output not yet split into lines, so a response just under it can be dropped when the read that completes it also carries enough of the next response to pass the limit, and that next response, possibly to another client's call, is dropped with it; whether that happens depends on where the reads fall. Measured on 3.6.2 from npm on 2026-09-26: a 9,000,000-character result was dropped and the agent got the timeout after 30.0 seconds, while a 1,000,000-character result came back in 31 milliseconds. Through a running proxy, a response line of 8,388,607 characters was returned and one of 8,388,608 was dropped; and when the upstream answered three calls in one write (100 characters and 58 under the limit for one client, then 100 for the other), the first was answered and the other two, one of them the other client's, timed out, while the same answers without the leading 100-character answer were both returned. These are the cases measured. An HTTP upstream's result is read whole and is not bounded this way. Workaround: keep tool results well under the bound, for example by reading large files in parts, and where results are large run one proxy per client. An error returned at once is planned for the reviewed release.

  12. The control channel does not check a request's Host or Origin header. It listens on 127.0.0.1 (port 18801 by default) so that a separate aga-proxy export can fetch the live bundle (routes /export, /status and /receipts). Measured on 3.6.2 from npm on 2026-09-26: GET /receipts and GET /export sent with the Host and Origin of another site returned 200, and both carried a denied call's argument path in its denial reason. A web page open in a browser on the same host can therefore read the live receipts and the evidence bundle through DNS rebinding, unless the browser blocks a public site's requests to the loopback address; any local user on the host can read them as well. A page can also start an export with a plain GET /export without rebinding, unless the browser blocks it, and each export holds up governed calls while it runs (item 9). The CLI has no option meant to turn the channel off; a --control-port of 70000, or any number above 65535, leaves it unstarted while governance runs, but then no command can export the running proxy's receipts. Workaround: do not browse the web on the host while the proxy runs, or run the proxy on a host where no one does, and on a host shared with other users treat the live receipts as readable by all of them. A Host and Origin check is planned for the reviewed release.

No fixed version is named until one is published.

Security

See SECURITY.md for vulnerability reporting.

Contributing

See CONTRIBUTING.md for development setup and guidelines.

License

MIT. The aga-receipt-spec/ directory carries its own Apache-2.0 license (see aga-receipt-spec/LICENSE).


Attested Intelligence Holdings LLC

Available Tools

15 tools
attest_subjectC

Attest subject, generate sealed Policy Artifact. Auto-loads into portal.

ParametersJSON Schema
NameRequiredDescriptionDefault
evidence_itemsNo
subject_contentYesContent/bytes of the subject
subject_metadataYes
behavioral_baselineNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description only indicates that it generates and auto-loads an artifact. It does not disclose whether the operation is destructive, idempotent, requires authorization, or has side effects beyond loading into the portal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with two short sentences. It front-loads the core action and provides an important side effect (auto-loading) without unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 parameters, nested objects, and no output schema, the description is too sparse. It does not explain return values, the format of the sealed artifact, or how parameters like evidence_items and behavioral_baseline contribute to the attestation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description does not mention any parameters, despite schema description coverage being only 25%. It fails to explain the purpose or usage of subject_content, subject_metadata, evidence_items, or behavioral_baseline, leaving a significant gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool attests a subject and generates a sealed Policy Artifact, with auto-loading into a portal. However, it does not distinguish from sibling tools like generate_evidence_bundle or request_claim, which could have overlapping purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. There is no mention of prerequisites, use cases, or scenarios where other tools would be more appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delegate_to_subagentB

Derive a constrained policy artifact for a sub-agent. Scope can only diminish, never expand. (NCCoE constrained delegation)

ParametersJSON Schema
NameRequiredDescriptionDefault
measurement_typesYesSubset of parent measurement types
delegation_purposeYesPurpose of the delegation
enforcement_triggersYesSubset of parent enforcement triggers
requested_ttl_secondsYesRequested TTL (will be clamped to parent remaining)

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden. It discloses the key behavioral trait of non-expanding scope, but lacks details on permissions, side effects, or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, concise and to the point, with no wasted words. However, it could benefit from slightly more structure or context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 4 required parameters, no output schema, and no annotations, the description is too minimal. It lacks details on return values, prerequisites, or how this fits with siblings, leaving an agent with insufficient information for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and each parameter is described in the schema. The description adds no additional meaning beyond the general purpose of constrained delegation, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool derives a constrained policy artifact for a sub-agent, with specific scope restrictions. It uses domain-specific terminology (NCCoE constrained delegation) but is not contrasted with siblings like request_claim or attest_subject.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The constraints 'Scope can only diminish, never expand' provide some usage context, but there is no explicit guidance on when to use this tool versus alternatives, nor when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_evidence_bundleA

Export the canonical SEP evidence bundle: signed PERMITTED/DENIED tool-call receipts + Merkle proofs + a mandatory signed checkpoint, for offline third-party verification (verify_bundle_offline, aga-verify, or aga-receipt-spec/verify/verify-sep.mjs). Pin gateway_public_key to prove provenance.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description bears full burden. It discloses the output components (receipts, proofs, checkpoint) and purpose (offline verification). However, it does not mention potential side effects, authentication, or whether it is read-only (likely safe).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with two sentences that front-load the main action. It contains no fluff, though it could be slightly more structured (e.g., bullet list) for clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no parameters and no output schema, the description provides a comprehensive overview of what the bundle contains and how it is used. Minor gaps: no mention of prerequisites (like needing a chain) or the bundle's size/format.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, and schema coverage is 100%. The description adds value by explaining what the tool outputs and its purpose, which is more than the schema provides. Baseline for 0 parameters is 4, and the description meets it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Export' and the resource 'canonical SEP evidence bundle', detailing its content (signed receipts, Merkle proofs, signed checkpoint). It distinguishes from the sibling 'verify_bundle_offline' by specifying this is for generation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the use case: 'for offline third-party verification'. Mentions verification methods (verify_bundle_offline, etc.) and a key prerequisite (pin gateway_public_key). Does not include when-not-to-use or alternative tools beyond verify_bundle_offline.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_chain_eventsC

Get continuity chain events.

ParametersJSON Schema
NameRequiredDescriptionDefault
end_seqNo
start_seqNo

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the burden of behavioral disclosure. It does not mention side effects, authorization requirements, rate limits, or any constraints beyond the name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At only four words, the description is underspecified rather than concisely informative. It fails to provide necessary details that would make it useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, no annotations, and no parameter descriptions, the description leaves the agent with minimal understanding of tool behavior, inputs, or outputs. It is insufficient for the given complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has 0% parameter description coverage, and the description does not explain the meaning, format, or use of 'start_seq' and 'end_seq'. The agent remains uninformed about these parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states it retrieves 'continuity chain events', identifying a specific resource and action. However, it lacks clarity on what these events represent, and does not distinguish it from sibling tools like 'get_portal_state' or 'get_server_info'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. The description provides no context about scenarios or prerequisites for invoking it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_portal_stateA

Get current portal state, loaded artifact info, and enforcement status.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, description must disclose behavioral traits. It only names what is retrieved, but does not mention side effects, permissions, or whether it is read-only. This is insufficient for a no-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, no wasted words. Every part contributes to understanding the tool's output.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so description should explain return structure. It names major components but lacks detail on format or fields. Given low complexity (no params, no annotations), it is minimally adequate but not thorough.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters in schema, so baseline is 4. Description adds meaning by specifying what the tool returns, going beyond the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states what is retrieved: 'current portal state, loaded artifact info, and enforcement status'. Verb and resource are specific and distinct from siblings like get_server_info or get_chain_events.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives (e.g., get_server_info). No context about prerequisites or typical use cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_receiptsC

Get all signed receipts, optionally filtered by artifact.

ParametersJSON Schema
NameRequiredDescriptionDefault
artifact_hashNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full burden. It states 'Get' implying read-only, but does not explicitly confirm no side effects, idempotency, or permission requirements. It does not disclose behavior for missing artifact_hash or large result sets.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, concise but lacks structure. It could be improved by separating purpose and usage details. It earns its place but does not excel in providing organized information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description is incomplete. It omits return format, pagination, error handling, and behavior for missing parameters. For a tool with one optional param, it is borderline but still leaves significant questions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description adds minimal meaning: 'optionally filtered by artifact' links the parameter artifact_hash to filtering but does not explain the format or valid values. This is insufficient for a parameter with no schema description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves all signed receipts with optional filtering by artifact. The verb 'Get' and resource 'signed receipts' are specific. Sibling tools like list_claims or get_chain_events are distinct, so no confusion arises.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description does not mention prerequisites, context, or exclusions. For example, it does not clarify if this tool is preferred over list_claims for certain cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_server_infoA

Get AGA server info, public keys, and portal state.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so the description carries full burden. It does not disclose whether the operation is safe, requires authentication, or has any side effects. For a read-only tool, minimal behavioral disclosure is given.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that is front-loaded and concise with no waste. Every word adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although the description mentions the types of information retrieved (server info, public keys, portal state), it does not describe the output structure or format. With no output schema, the description should provide more detail on what the response contains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has no parameters, and the description does not need to add parameter meaning. Schema coverage is 100% trivially. Baseline for 0 parameters is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves AGA server info, public keys, and portal state. It uses a specific verb 'Get' and distinguishes from sibling tools like get_portal_state, which is more specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. Given siblings like get_portal_state and other info tools, the description should include context on when to prefer this over others.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

init_chainC

Initialize continuity chain with genesis event.

ParametersJSON Schema
NameRequiredDescriptionDefault
specification_hashNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. It only says 'initialize' and 'genesis event', but does not mention side effects, idempotency, error conditions, or required permissions. Falls short for a mutating operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely concise at 5 words, but at the expense of critical information. It is front-loaded but does not earn its place as it omits parameter details and usage guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has one parameter, no output schema, and no annotations, the description is incomplete. It lacks explanation of the parameter's meaning, return value, and overall behavior, making it insufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter 'specification_hash' has no description in the schema and is not mentioned in the description. The agent has no way to understand its purpose or format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Initialize' and the resource 'continuity chain' with 'genesis event', which differentiates it from sibling tools like verify_chain. However, the term 'continuity chain' is not explained, assuming domain knowledge.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives like get_chain_events or verify_chain. No prerequisites, exclusions, or contextual conditions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_claimsA

List available claims with sensitivity levels.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It only states that the tool lists claims with sensitivity levels, but does not disclose side effects (likely none), authorization needs, or data freshness. For a tool with no parameters, the behavior is simple, but more transparency could help.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that conveys the essential information without any fluff. It is front-loaded and efficient, earning its place with precise wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no parameters or output schema, the description is sufficient for a simple list operation. It tells the agent what the tool provides (claims with sensitivity levels). However, it could mention if the list is complete or filtered in any way, but overall it is adequately complete for its simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0 parameters and 100% schema description coverage. The description adds no parameter information, but since none exist, this is appropriate. The baseline is 4 for no parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: listing available claims with sensitivity levels. It uses a specific verb ('list') and resource ('claims'), and adds detail about sensitivity levels, distinguishing it from siblings like 'request_claim' or 'verify_bundle_offline' which have different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. For example, it does not explain how 'list_claims' differs from 'get_receipts' or other listing operations. No context about prerequisites or ideal scenarios is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

measure_behaviorA

Measure behavioral patterns (unauthorized tools, rate violations, forbidden sequences). DETECTIVE-ONLY by default: it records and PROVES drift but does not block. Pass enforce=true to also trip the portal into phantom quarantine on drift (opt-in; off by default). (NIST-2025-0035)

ParametersJSON Schema
NameRequiredDescriptionDefault
enforceNo

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the detective-only default, that it records and proves drift but does not block, and that enforce=true triggers phantom quarantine. It could elaborate on what 'phantom quarantine' entails, but it's sufficiently transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with purpose and followed by details on default behavior and enforce option. Every sentence adds value with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one parameter and no output schema, the description covers purpose, default behavior, and the enforce option. It does not describe the output or what triggers measurement, but it is fairly complete given the context. A brief mention of output would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides only the parameter name and type with 0% coverage. The description adds valuable meaning: it explains that enforce (boolean, default false) opts in to phantom quarantine on drift. This fully compensates for the schema's lack of description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool measures behavioral patterns (unauthorized tools, rate violations, forbidden sequences) and distinguishes itself from sibling tools by highlighting 'DETECTIVE-ONLY' default behavior and an optional enforce parameter for quarantine.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use the tool (for detecting behavioral drift) and how to opt into enforcement via enforce=true. It does not explicitly state when not to use it or compare to siblings like measure_integrity, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

measure_integrityB

Measure subject state, compare to sealed reference. Generates signed receipt for every measurement.

ParametersJSON Schema
NameRequiredDescriptionDefault
subject_contentYesCurrent content of the subject
subject_metadataYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden for behavioral disclosure. It mentions receipt generation but does not disclose side effects, permission requirements, or what happens on mismatch. The comparison to a sealed reference is implied but not detailed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no redundancy, front-loaded with the core action. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having two parameters (one nested) and no output schema, the description is too brief. It omits the source of the sealed reference, return format, and any prerequisites, leaving significant gaps for an agent to successfully invoke the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only 50% of parameters have schema descriptions; the description adds no parameter-level details. Subject_content is described in schema; subject_metadata lacks description in both schema and tool description. The description does not clarify their roles beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: measuring subject state and comparing to a sealed reference, and generating signed receipts. It uses specific verbs and resources, distinguishing it from sibling tools like 'attest_subject' and 'measure_behavior'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, no prerequisites, and no exclusions. It only states the function without contextual usage advice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

request_claimC

Request disclosure of a claim. Auto-substitutes if denied.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoREVEAL_MIN
claim_idYes
requester_idNoanonymous

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden for behavioral disclosure. It mentions 'Auto-substitutes if denied', which is a key behavior, but fails to explain what 'substitutes' means, whether it's destructive, or any other side effects. No mention of permissions or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise (two sentences), but it lacks structure and essential information. While it is not verbose, the conciseness comes at the cost of completeness. Every sentence should earn its place, but here they are insufficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 3 parameters, no output schema, and no annotations, the description is severely incomplete. It does not explain the return value, the meaning of disclosure modes, or the auto-substitution behavior in detail. The tool is for a potentially complex operation, but the description leaves many gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, and the description adds no parameter explanations. It doesn't mention claim_id, mode, or requester_id. The mode enum (PROOF_ONLY, REVEAL_MIN, REVEAL_FULL) is not described. The description fails entirely to add meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Request disclosure of a claim.' The verb 'request disclosure' and resource 'claim' are specific. It also mentions auto-substitution on denial, adding context. However, it doesn't differentiate from siblings like 'list_claims', but the purpose is distinct enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. There is no mention of prerequisites, scenarios, or exclusions. The tool is for requesting claim disclosure, but no context on when to choose this over other claim-related tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

revoke_artifactA

Revoke an active policy artifact mid-session. Portal terminates on next measurement. (NCCoE Phase 3b)

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonYesReason for revocation
sealed_hashYesSealed hash of artifact to revoke

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description partially carries the burden. It discloses the termination effect on next measurement, but does not detail permissions, side effects on other artifacts, or confirmation mechanics. It adds value but leaves gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero waste. Every word contributes to understanding purpose and key behavior. Efficient and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is minimal but covers the primary action and a notable behavioral effect. However, given no output schema and potential complexity around revocation, more detail (e.g., what happens to the session, how to confirm success) would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with clear descriptions for both parameters ('sealed_hash', 'reason'). The tool description does not add any additional meaning beyond what the schema already provides, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly specifies the action 'Revoke' on the resource 'active policy artifact', adds context 'mid-session', and notes a specific effect 'Portal terminates on next measurement'. This distinguishes it from sibling tools like 'attest_subject' or 'delegate_to_subagent'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context ('mid-session') but no explicit guidance on when to use this tool versus alternatives (e.g., when not to use, or which sibling to choose instead). It lacks exclusion criteria or comparative hints.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_bundle_offlineA

Verify a canonical SEP evidence bundle offline (full §6 algorithm: structural floor, receipt signatures, chain+ordering, leaf-recompute + Merkle bijection, signed checkpoint). Pass pinned_public_key to also prove provenance.

ParametersJSON Schema
NameRequiredDescriptionDefault
bundleNo
pinned_public_keyNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, but the description details the verification algorithm: structural floor, receipt signatures, chain+ordering, leaf-recompute + Merkle bijection, signed checkpoint. This fully discloses the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences: first describes purpose and algorithm, second adds optional parameter. No unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the algorithm well and parameter use, but lacks output specification (return value) and prerequisites. For a complex verification tool, this is a notable gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has zero description coverage. The description explains pinned_public_key for provenance but does not elaborate on the bundle parameter's structure or constraints, leaving a gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool verifies a canonical SEP evidence bundle offline, listing specific algorithm steps from §6. It distinguishes from sibling tools like verify_chain by focusing on bundle verification.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies use when offline bundle verification is needed and mentions optional provenance via pinned_public_key. It does not explicitly state when not to use, but context with siblings is implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_chainC

Verify continuity chain integrity.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.6/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, and the description does not disclose behavioral traits. It does not state whether the tool performs a read or mutation, what it returns, or any side effects, leaving the agent with no insight into the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise (one sentence). However, it sacrifices informativeness for brevity; it could provide more detail without being verbose. It is not tautological but minimally sufficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity of the tool (no parameters, no output schema), the description still fails to explain what 'verify' entails, what the output indicates, or how it relates to sibling tools like verify_bundle_offline. It is incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so schema coverage is trivially 100%. The description adds no parameter information, but baseline for zero parameters is 4. No further meaning is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Verify' and resource 'continuity chain integrity', making the general purpose clear. However, it does not differentiate from sibling tools like verify_bundle_offline or get_chain_events, which limits clarity in context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description lacks any context about prerequisites, intended use cases, or when to avoid it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 33 tool updatesv2.0.1
    • Removedaga_create_artifact
    • Removedaga_delegate_to_subagent
    • Removedaga_demonstrate_lifecycle
    • Removedaga_disclose_claim
    • Removedaga_export_bundle
    • Removedaga_generate_receipt
    • Removedaga_get_chain
    • Removedaga_get_portal_state
    • Removedaga_init_chain
    • Removedaga_measure_behavior
    • Removedaga_measure_subject
    • Removedaga_quarantine_status
    • Removedaga_revoke_artifact
    • Removedaga_rotate_keys
    • Removedaga_server_info
    • Removedaga_set_verification_tier
    • Removedaga_start_monitoring
    • Removedaga_trigger_measurement
    • Removedaga_verify_artifact
    • Removedaga_verify_bundle
    • Addedattest_subject
    • Addeddelegate_to_subagent
    • Addedgenerate_evidence_bundle
    • Addedget_chain_events
    • Addedget_portal_state
    • Addedget_receipts
    • Addedlist_claims
    • Addedmeasure_behavior
    • Addedmeasure_integrity
    • Addedrequest_claim
    • Addedrevoke_artifact
    • Addedverify_bundle_offline
    • Addedverify_chain
  2. 22 tool updatesv2.0.0
    • First observedaga_create_artifact
    • First observedaga_delegate_to_subagent
    • First observedaga_demonstrate_lifecycle
    • First observedaga_disclose_claim
    • First observedaga_export_bundle
    • First observedaga_generate_receipt
    • First observedaga_get_chain
    • First observedaga_get_portal_state
    • First observedaga_init_chain
    • First observedaga_measure_behavior
    • First observedaga_measure_subject
    • First observedaga_quarantine_status
    • First observedaga_revoke_artifact
    • First observedaga_rotate_keys
    • First observedaga_server_info
    • First observedaga_set_verification_tier
    • First observedaga_start_monitoring
    • First observedaga_trigger_measurement
    • First observedaga_verify_artifact
    • First observedaga_verify_bundle
    • First observedget_server_info
    • First observedinit_chain

TDQS

B3.4/5.0

Scored across 15 tools

Disambiguation5/5

Each tool has a unique purpose with no overlapping functionality. For example, measure_behavior and measure_integrity measure different aspects, and verify_chain vs. get_chain_events are distinct operations. The descriptions clearly differentiate them.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern in snake_case (e.g., attest_subject, delegate_to_subagent, verify_chain). There is no mixing of conventions or irregular naming.

Tool Count5/5

With 15 tools, the server is well-scoped for its attestation and policy domain. Each tool earns its place, covering setup, measurement, verification, and revocation without being excessively numerous.

Completeness5/5

The tool surface appears complete for the server's purpose, covering artifact creation, delegation, evidence export, chain initialization, state querying, claims handling, behavioral and integrity measurement, revocation, and offline verification. No obvious gaps in the core workflow.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Local zero-trust permission gateway for AI agents. Enforces policy-based tool authorization, human approvals, scoped permissions, and cryptographically verifiable audit logs.
    4
    94 PyPI
    5
    Apache 2.0
  • A
    license
    B
    quality
    B
    maintenance
    A governance proxy for AI tools — every MCP/agent tool call is policy-gated, secret-redacted, and written to a hash-chained, offline-verifiable audit trail.
    13
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Self-hosted MCP gateway that applies deterministic, compiled policy to tool discovery, invocation, and outbound data flow, with no model in the enforcement path. Every decision emits a hash-chained receipt sealed with Ed25519 and verifiable using public keys only.
    Apache 2.0
  • A
    license
    Not graded
    quality
    B
    maintenance
    Governed MCP gateway that lets AI agents call tools with policy enforcement, prompt-injection screening, a kill-switch, and tamper-evident signed audit logs.
    Apache 2.0