Skip to main content
Glama
verax-ai

Verax

Official

Verax

The body an agent asks before it acts.

Verax is an MCP server that sits between an agent and its tools. Every tool call passes a policy gate and leaves a signed decision record before anything runs; every call that ran leaves an effect row that is reconciled against its record afterwards. A refusal is recorded the same way as an approval. A call the policy will not decide alone is held until an operator on this machine approves it. The ledger stays on the machine the body runs on, and the body opens only when its authorization is configured: there is no default token.

By VERAX Teknoloji. Sister projects: Conarium · Tugra · Cedulon. Decision records use the Cedulon record format.

See it run

npx @verax-ai/body demo

verax demo in a terminal: an allowed memory write, a signed refuse, and a payment held until the operator answers y

The recording is the output of one run in a terminal where the approval question was answered y. That run took about a second; it is played back slowly here so it can be read.

What it prints, without a terminal to answer the approval question:

verax demo

memory.put / memory.get
  allowed; two signed records

message.send -> ops@blocked.test
  refused egress-blocked (signed)
  ref 6c6e4925-b769-4a8b-8fc4-e2443613a5b6

spend 100 minor USD sample-merchant
  held
  no terminal to ask, so it stays held (run this in a terminal to be asked)

audit.explain of the refuse
  finding none
  trust-root own-key
  chain and signatures read back

records  5
effects  3
ledger   removed on exit (run with --keep to keep it and check it with verax verify)

Not shown here: data masking arrives with a downstream server such as Conarium; statement reconciliation needs a real statement (verax reconcile).

Node 22.6 or newer, @verax-ai/body 0.2.1 or later. The command records an allowed memory write and read, a signed refuse of a message to a host off the policy list, and a payment held for the operator on this machine (approved when the terminal answers y).

--keep leaves the temporary ledger on disk. verax verify <dir> reads it back without a body, as in Read the ledger back without us.

The same rules, animated

Two short episodes show the same gate, with characters in place of a terminal. They are sample scenarios: the scenes are drawn in three.js and the voices are generated with ElevenLabs.

Episode 1, "$1,850": a payment is held until a person approves it on a phone, and changing one cent in the record breaks its signature.

https://github.com/user-attachments/assets/772ae1c4-2697-47a0-a228-36f0508b0d93

Episode 2, "No rule, no way through": a call with no rule is refused, the same call under a new name is refused again, and both refusals are signed.

https://github.com/user-attachments/assets/4fb86627-3310-4f10-b1a9-059d8fcffb28

Related MCP server: Conarium

Connect your agent

Elevated verax install and verax approve run a copy of this program that only an administrator can change; a copy your account can change is refused. On Windows, in 0.4.0, an elevated CLI approve is the way to approve (the passkey panel ships in 0.4.1). It has to be run from a separate administrator account, not this account elevated, because a same-user elevated shell inherits that user's environment variables and PowerShell profile, which the agent can set. Clear NODE_OPTIONS in that shell (Remove-Item Env:NODE_OPTIONS). Start that other account's PowerShell with -NoProfile (an elevated shell otherwise runs your $PROFILE, which your account can change):

npm install -g --prefix "$env:ProgramFiles\verax-cli" @verax-ai/body
& "$env:ProgramFiles\verax-cli\verax.cmd" install

On Linux and macOS, with a root-owned Node (the distribution's /usr/bin/node, or /opt/verax-node/<dir>/bin/node; the installer prints those steps when the Node it was started from can be changed by your account). Do not start an elevated command with env node:

sudo npm install -g --prefix /opt/verax-cli @verax-ai/body
sudo /usr/bin/node /opt/verax-cli/lib/node_modules/@verax-ai/body/dist/cli.js install

The command installs @verax-ai/body from the npm registry into an administrator-owned directory after signature checks, runs that Node, and keeps the ledger under a service account. On Windows the agent token is %ProgramData%\Verax\agent-token\<your SID>\agent.token (Administrators and SYSTEM have full control, your SID can read the file and read-execute the directory). On Linux and macOS a child process running as your uid writes ~/.verax/agent.token from its stdin. It prints the Claude Code line that reads that file. Port 8787 taken? verax install --port 8797. Node must be the all-users installer from nodejs.org on Windows; a Node your account can rewrite is refused. On macOS the remedy extracts the official tarball as root into /opt/verax-node (root:wheel, not group- or other-writable). On Linux the same place, /opt/verax-node (root:root), which SELinux labels usr_t. On SELinux systems install requires Node labelled bin_t or usr_t (distribution Node is; a tarball under /usr/local/lib is not) and prints the one-line fix. The service then runs in unconfined_service_t. The service account and the file permissions are the boundary.

Approve a held call with verax approve. The passkey panel (verax desktop) is not in 0.4.0; it ships in 0.4.1. On Windows run the approve from a separate administrator account, not this account elevated, in a -NoProfile PowerShell after Remove-Item Env:NODE_OPTIONS: & "$env:ProgramFiles\verax-cli\verax.cmd" approve. Linux and macOS, naming the root-owned Node: sudo /usr/bin/node /opt/verax-cli/lib/node_modules/@verax-ai/body/dist/cli.js approve or sudo /opt/verax-node/<dir>/bin/node /opt/verax-cli/lib/node_modules/@verax-ai/body/dist/cli.js approve. Uninstall the same way, with uninstall in place of approve.

To try it in your own user, which is not a boundary:

verax init --local ~/.verax
verax serve --env-file ~/.verax/verax.env
claude mcp add --transport http verax http://127.0.0.1:8787/mcp --header "Authorization: Bearer $(cat ~/.verax/local-issuer/agent.token)"

The token can read and write memory through the gate; it cannot approve. The shipped policy refuses spend until you add a rule for it; a call your policy holds waits for verax approve on this machine. See docs/THREAT_MODEL.md.

With Conarium

@verax-ai/body 0.2.2 and later accepts --with-conarium. 0.2.1 does not carry the flag. The flag has npx download @conarium-ai/core from npm and run it as a child process; that needs a network.

Conarium masks the rows, and its sample policy denies its public.secrets table. Verax puts each call through the same gate as its own tools, keeps a signed record and a hash of the answer, and refuses a downstream tool the policy does not name before the child sees the call. When Conarium answers with an error, the body tells the caller that it did, not what it said.

What it prints, without a terminal to answer the approval question:

verax demo

memory.put / memory.get
  allowed; two signed records

message.send -> ops@blocked.test
  refused egress-blocked (signed)
  ref 6c6e4925-b769-4a8b-8fc4-e2443613a5b6

spend 100 minor USD sample-merchant
  held
  no terminal to ask, so it stays held (run this in a terminal to be asked)

conarium.query customers
  allowed; masked by Conarium before the rows left it
  measured [MASKED_PII] on every email and card

conarium.query public.secrets
  allowed by this gate; Conarium answered with an error and no rows
  recorded as a failed call

conarium.list_tables
  refused no-rule (signed)
  refused by this gate; the downstream server never saw the call

audit.explain of the refuse
  finding none
  trust-root own-key
  chain and signatures read back

records  8
effects  5
ledger   removed on exit (run with --keep to keep it and check it with verax verify)

Not shown here: a real database (these are Conarium's sample rows); statement reconciliation needs a real statement (verax reconcile).

What ships

Package

What it is

@verax-ai/body

The MCP server and the verax command: serve, install, uninstall, init, doctor, approve, operator, reconcile, witness, halt, unlock (desktop ships in 0.4.1).

@verax-ai/proxy

The decision proxy the body is built on: policy, signed records, ledger, explain, reconcile.

@verax-ai/inventory

The roster document a body serves and the panel lists, with its strict parser.

The three packages are published together and carry the same version; the capability matrix in docs/STATUS.md names the current one, and it is the version on npm. The body is also listed in the MCP registry as io.github.verax-ai/verax. What that version carries, what it does not, and the test holding each row up are in the capability matrix at the top of docs/STATUS.md.

Read the ledger back without us

A ledger only the vendor's running service can read is evidence a buyer rents, not evidence they hold. verax verify reads a state directory on its own — no body listening, nothing on the network — and states four things separately, because they fail separately:

$ verax verify ./verax-state
ledger        ./verax-state
decisions     6
effects       4 (4 bound to a decision, 0 with none)
signatures    6 verify, 0 do not
chain         unbroken
verified with the key carried in these files
              verified against the key carried in the records themselves: this
              shows the files are internally consistent, not that the key was
              ever trusted. Pin a key you hold to check that.

VERIFIED

That last pair of lines is the point. Checking a ledger against the key lying next to it proves the files agree with each other and nothing more — anything able to write the ledger could write that key too. Pass --key <public.pem> to verify decision records against a copy you hold, and the answer says a key you supplied instead. Effects are signed with a separate key. A self row is checked under the effect key: --effect-key <public.pem> when you pin it, otherwise one key taken from the first self row. A same-org row is checked under the witness key: --witness-key <public.pem> when you pin it, otherwise one key taken from the first same-org row. Pinning --effect-key while a same-org row is present requires --witness-key as well, and pinning --witness-key while a self row is present requires --effect-key as well; otherwise the result is not verified. Each of those lines says the same thing: the files agree with each other, not that the key was ever yours. Checkpoints are signed by the witness key. Every checkpoint row is checked under one key: --checkpoint-key <public.pem> when you pin it, otherwise one key taken from the checkpoint file, with the same note that agreement is not trust. The tail counts only checkpoints whose signatures verified. Someone who can rewrite the files can still roll that file back to an older valid prefix together with the records after it; only a checkpoint held elsewhere detects that. --json prints the same result for a pipeline; the exit code is 0 when it verifies and 1 when it does not, and a directory with no ledger in it is never quiet success.

Records are COSE_Sign1 with algorithm -19 (Ed25519, RFC 9864), not the older polymorphic -8 (EdDSA). Some COSE libraries do not know -19 yet: as of go-cose 1.3.0 and pycose 1.1.0 both reject it, and support is tracked in go-cose#224 and pycose#126. verax verify does not depend on either library.

Install

npm install -g @verax-ai/body
verax --help
verax doctor

On an account where every process is elevated (the built-in Administrator, or EnableLUA=0), a command from a user-writable npm prefix is refused; use the administrator-owned copy at %ProgramFiles%\verax-cli.

Node 22.6 or newer. The body speaks MCP over Streamable HTTP at /mcp on VERAX_BIND (default 127.0.0.1:8787) and needs an issuer, a JWKS URL, an audience, a state directory and a policy file before it listens; verax doctor names what is missing. The variables and the run steps are in packages/body/README.md.

Tools

The policy decides which of these a token's scopes may call; packages/proxy/policy/default.json denies what it does not name.

Tool

Description

memory.get

Reads one memory item behind the gate, stored per tenant.

memory.put

Writes one memory item behind the gate, stored per tenant.

audit.explain

Reads a decision back from the signed ledger, with its chain, its signatures and its findings.

message.read

Reads the inbox.

message.send

Writes to the outbox; reaches only hosts the policy allow-lists.

spend

Authorizes a payment and records it, under a cap, a payee list and a daily limit from the policy; held for an operator when the policy says so. The body does not move money.

What the body does beyond the gate

Each line below is a row in the capability matrix in docs/STATUS.md, where it is stated against a version, with what it does not do and the test that fails when it stops being true.

  • Approval: in 0.4.0 a held call is approved with verax approve on this machine. On Windows that runs from a separate administrator account. The passkey panel (verax operator, verax desktop) ships in 0.4.1. The approver's operator id is bound into the signed record by hash.

  • Witness: verax witness signs effect rows from a second process and writes durable checkpoints; without it the witness class stays self.

  • Halt and revoke: verax halt turns every further call into a signed deny; a revoked token id is refused before any record is written.

  • Reconcile: verax reconcile matches recorded spends against a card statement export and names the matched, ghost and authorized-but-unpaid rows.

  • Tenant key: memory and inbox are stored under a key derived from the token's issuer and subject; another tenant's id is answered with a signed deny.

  • Bounds: rate and daily counters, a disk-low refusal (HTTP 507) when a deny could not be recorded, and an egress allow-list; counters that cannot be read fail closed.

  • Doctor and heartbeat: verax doctor names what is missing or stale before the first call finds out.

Connect a client

The body listens on http://127.0.0.1:8787/mcp by default. A client configuration looks like this; the token comes from your issuer.

{
  "mcpServers": {
    "verax": {
      "url": "http://127.0.0.1:8787/mcp",
      "headers": { "Authorization": "Bearer <token from your issuer>" }
    }
  }
}

Status

What the tree carries and what stays unproven is stated, item by item, in docs/STATUS.md. Nothing in this repository is a claim beyond that file, and a paragraph there is not a release. The threat model is in docs/THREAT_MODEL.md; how to report a vulnerability is in SECURITY.md.

What else is in the tree

  • apps/panel is the account-for screen: the records list, the black box and the status view, read from the signed ledger. Private; it is not published. The panel session uses the code flow; the access token stays in memory and is dropped on refresh. Vite may still attach VERAX_DEV_TOKEN from .env.local to /api when the request has no Authorization header (desktop MCP brains and tests).

  • scripts/dev-issuer.mjs is development only; not a production authorization server. It serves GET /authorize (PKCE S256) and POST /token, writes a token to --out, and never prints one. It listens on VERAX_DEV_ISSUER_PORT (default 8790). NODE_ENV=production exits.

  • scripts/demo-box.mjs is development only: one process that starts the dev issuer and the body on loopback with a temporary ledger, mints itself a short-lived token through the issuer's code flow, and speaks MCP over stdio for a sandbox that cannot hold a token of its own, such as a directory's build check. The body is not changed by it: every call still passes the gate and is recorded, spend is always held, and no operator is there to approve it. NODE_ENV=production exits. Not a deployment.

Developing

npm ci
npm test            # guards, typecheck, build, unit and cost suites, panel
npm run pack:smoke  # pack the three packages and install them elsewhere

CI runs the suite as a non-root user on Linux and again on Windows, plus the proxy performance check. Releases go out from the Actions tab: release.yml publishes the three packages with npm trusted publishing and a provenance attestation, then mcp-registry.yml updates the registry record once npm answers for the new version. Neither runs on push.

Tested on every release on the platforms listed in docs/PLATFORMS.md, each row linked to its CI run.

License

Apache-2.0.

Available Tools

6 tools
audit.explainA

Reads one decision back from the signed ledger by its ref and explains it. Use it to check what the body decided about an earlier call and whether the recorded effect matched, before repeating a call or reporting on it; read-only, and the lookup itself is recorded too. Returns JSON with record (the signed decision's claims: tool, verdict, policy hash, timestamps), effect (the reconciled effect row), finding (match, mismatch or missing), witnessClass, guarantee, warnings, trustRoot (which key verified the signatures), and for a held call pair with its defer and resolution records. A ref that does not exist, or belongs to another tenant, is answered with the same signed deny, so neither case reveals the other.

ParametersJSON Schema
NameRequiredDescriptionDefault
refYesThe decision reference: the ref returned by an earlier call, also the tail of a denied:… or deferred:… answer; 1 to 64 characters of letters, digits, '.', '_' or '-', starting with a letter or digit.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so thoroughly. It discloses the read-only nature, the side effect that the lookup is recorded, and the security behavior of returning a signed deny for both non-existent and other-tenant refs to avoid information leaks. This is exemplary behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-organized, moving from purpose to usage, return fields, and edge-case behavior. It is longer than minimal, yet every sentence contributes essential information and nothing feels redundant or off-topic.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no output schema, the description fully explains the return shape: record, effect, finding, witnessClass, guarantee, warnings, trustRoot, and the held-call pair case. Error semantics are also covered with the signed-deny behavior. An agent has everything needed to call and interpret the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds practical semantics beyond the schema: the ref is the one returned by an earlier call, or the tail of a denied/deferred answer. This helps the agent locate the correct value to pass.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: reads a signed ledger decision by ref and explains it. The tool's focus on auditing decisions is clearly distinct from sibling tools handling memory, messaging, and spending, so an agent can tell them apart immediately.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it: to check what the body decided about an earlier call and whether the effect matched, before repeating or reporting on it. It also notes the lookup is recorded, which is a useful side-effect warning. No alternatives are named, but the siblings are unrelated so exclusions are not necessary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory.getA

Reads one memory item this tenant stored earlier with memory.put, by its id. Use it to recall a fact, a setting or a note before acting on it; nothing is written. Like every call it passes the policy gate and leaves a signed decision record; an id that belongs to another tenant is answered with a signed deny. Returns the stored item as JSON: {id, body, source, validFromMs, validUntilMs, versionHash}. Outside the validity window the body is withheld: {stale: true, id, validUntilMs} after it, {notYetValid: true, id, validFromMs} before it. An unknown id answers {error: "not-found", id}.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe id given to memory.put: 1 to 128 characters of letters, digits, '.', '_' or '-', starting with a letter or digit; case-sensitive.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it delivers: it discloses policy-gate behavior, signed decision records, cross-tenant deny behavior, exact JSON return shapes, stale/notYetValid handling, and not-found errors. This is unusually thorough for a read tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: purpose first, then usage, then safety/behavior, then return formats and edge cases. There is no filler, and the structure front-loads the single most important fact: what the tool reads.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read tool with no output schema and no annotations, the description fully covers return values, success and error cases, cross-tenant behavior, and validity-window semantics. An agent has everything needed to call it correctly and interpret all possible responses.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema itself fully documents the one parameter, including the charset, length, starting character, and case-sensitivity, so schema coverage is 100%. The description adds only 'by its id' and 'given to memory.put,' which is helpful but does not materially go beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Reads one memory item this tenant stored earlier with memory.put, by its id.' This accurately distinguishes it from the write sibling memory.put and from other tools: it is a read operation for a single memory item.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear use case: 'Use it to recall a fact, a setting or a note before acting on it; nothing is written.' It implies read vs. write by referencing memory.put, but it does not explicitly state 'use memory.put when storing' or list exclusions beyond that the call writes nothing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory.putA

Writes one memory item for this tenant, or replaces the item with the same id, in the body's state directory on this machine. Use it to keep a fact for a later memory.get together with where it came from and how long it holds, so a stale fact is not served later. The call passes the policy gate and is recorded; the record carries the item's versionHash, a SHA-256 over id, body and validity window. Returns {ok: true, id, versionHash}. A missing source answers {error: "source-required"}, a missing validUntilMs {error: "validUntilMs-required"}, a malformed id {error: "id-invalid"}.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesIdentifier to store under and read back with memory.get: 1 to 128 characters of letters, digits, '.', '_' or '-', starting with a letter or digit; case-sensitive. An existing item with this id is replaced.
bodyYesThe value to keep, as any JSON: object, array, string, number or boolean. Stored as given and returned as given by memory.get.
sourceYesWhere the value came from, as a JSON object of your choosing, for example {"kind": "document", "ref": "invoice-2026-09.pdf"}. Required; stored with the item so a later reader can weigh it.
validFromMsNoOptional. Unix time in milliseconds from which the item may be served; before it memory.get answers notYetValid. Omit to serve it at once.
validUntilMsYesRequired. Unix time in milliseconds after which memory.get answers stale and withholds the body. Pick the moment the fact should no longer be trusted.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It discloses that the call passes the policy gate, is recorded, returns a versionHash, and defines error responses for missing source, missing validUntilMs, and invalid id. It also implies mutation via 'writes' and 'replaces', though it doesn't explicitly state destructive effects beyond replacement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a few sentences, but every sentence adds meaningful information: the core action, usage rationale, recording and hashing behavior, and error cases. It is front-loaded with the primary action and does not waste words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 params, nested objects, no output schema, and no annotations, the description covers the essential operational details: what it does, how to use it, what it returns (example), and common errors. It doesn't explain the exact return structure beyond the example, but that is acceptable given no output schema and the example provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds value by explaining the purpose of source and validUntilMs in the usage section and enumerating error responses tied to specific parameters (e.g., source-required, validUntilMs-required, id-invalid). This goes beyond the schema's basic field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (writes/replaces) and resource (one memory item), with explicit scope (tenant, body's state directory on this machine). It also distinguishes from the sibling memory.get by saying 'for a later memory.get', making the action unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context: use it to keep a fact for later retrieval, with provenance and validity window to avoid serving stale facts. It doesn't explicitly state when not to use it or name alternative tools, but the guidance is sufficient for an agent to decide when to write memory.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

message.readA

Reads this tenant's inbox, the messages placed for it in the body's state directory on this machine, and returns them as a JSON array in arrival order, oldest first. Use it to see what has arrived before deciding what to answer. Takes no arguments; read-only; the call is recorded like every other. An empty or absent inbox answers [].

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the read-only nature, that the call is recorded, and that an empty inbox returns [] – all beyond what the schema shows.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is succinct, front-loads the core action (reads inbox), and includes necessary behavioral notes without fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters, no output schema, and no annotations, the description fully covers what an agent needs to know: what it does, when to use it, and the return format.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters, so there is nothing to add. The schema coverage is 100%, and the description confirms no arguments are needed, which is sufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reads the tenant's inbox and returns messages as a JSON array. However, it does not explicitly differentiate from sibling tools like message.send or memory.get, though the read-only nature is implied.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says to use it before deciding what to answer, which gives some context. But it does not mention when not to use it or compare with alternatives like memory.get.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

message.sendA

Queues one message in this tenant's outbox on this machine for the delivery step the operator runs; this call opens no network connection and nothing leaves the body from it. Use it to hand off a message, not to deliver one. Like every call it passes the policy gate and leaves a signed decision record. The gate reads the host after the last '@' in to and allows it only when it is on the policy's egress allow-list; otherwise the call is refused with denied:egress-blocked, or denied:egress-host-missing when no host can be read. A policy rule in approve mode holds the call for an operator instead and answers deferred:approval-required:. Returns {queued: true, ref}, where ref is the decision reference for audit.explain.

ParametersJSON Schema
NameRequiredDescriptionDefault
toYesRecipient address with a host after the last '@', for example ops@example.com. The host, lower-cased, is matched against the policy's egress list.
_refNoOptional reference you choose for this call, 1 to 64 characters of letters, digits, '.', '_' or '-', starting with a letter or digit. Resend the same call with the same _ref after an operator approved it to receive allowed:<ref>; a _ref reused for a different call is refused with denied:ref-reuse.
textYesThe message body as plain text. Stored as given in the outbox row.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full disclosure burden, and it does so thoroughly: no network connection is opened, nothing leaves the body, the policy gate is passed, a signed decision record is left, and refusal/deferral outcomes are specified. This gives a complete behavioral picture.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: core behavior, usage guidance, policy outcome, approval behavior, and return value. The most important disambiguation—queueing rather than delivering—is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema and no annotations, the description fully specifies the return shape, error/refusal statuses, approval flow, and policy context. An agent has enough information to invoke this tool correctly and interpret its result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds useful meaning: the host is lower-cased and checked against the egress allow-list, and the _ref interaction with operator approval and ref-reuse refusal is clarified. This goes beyond the schema without being redundant.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: it 'queues one message in this tenant's outbox on this machine' and explicitly says it is not delivering the message. This clearly distinguishes it from message.read and from any actual network-send behavior the name 'send' might imply.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use the tool ('Use it to hand off a message, not to deliver one') and points to the operator-run delivery step as the separate follow-up. It also explains policy approval/deferral behavior so the agent knows when the call will be queued versus held for an operator.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

spendA

Asks the body to authorize a payment and records the decision; the body never moves money, so authorized: true is a signed permission for a later payment step, not a transfer. Use it before any payment so that amount, currency, payee and reference are checked against the policy: the one currency the policy names, a cap per call, a payee list and a daily limit. A call outside those bounds is refused with a signed deny naming the bound: denied:spend-cap, denied:spend-payee, denied:spend-currency or denied:spend-daily. A call within them is held for an operator on this machine and answers deferred:approval-required:; once that ref is approved (verax approve, or the panel), resending the same call with the same _ref answers allowed:, and the authorization is recorded as {authorized: true, ref, amountMinor, currency, payee, reference}. Without a spend rule in the policy every call answers denied:spend-not-wired.

ParametersJSON Schema
NameRequiredDescriptionDefault
_refNoOptional reference you choose for this call, 1 to 64 characters of letters, digits, '.', '_' or '-', starting with a letter or digit. Resend the same call with the same _ref after an operator approved it to receive allowed:<ref>; a _ref reused for a different call is refused with denied:ref-reuse.
payeeYesWho is to be paid, spelled exactly as the policy's payee list spells it (a merchant or account name). A payee off the list is refused.
currencyYesISO 4217 code in upper case, for example USD, EUR or TRY. Must equal the currency the policy's spend rule names.
referenceYesYour own reference for this payment, such as an invoice or order id. Recorded with the authorization and used by verax reconcile to match the card statement.
amountMinorYesAmount in the currency's minor unit as a positive integer: cents, kuruş or pence, so 1250 means 12.50. Compared against the policy's cap per call and daily limit.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full disclosure burden — and it excels: it reveals that 'authorized: true is a signed permission... not a transfer,' enumerates every refusal token (denied:spend-cap, denied:spend-payee, denied:spend-currency, denied:spend-daily, denied:spend-not-wired), explains the deferred:approval-required flow, and specifies the recorded result shape {authorized: true, ref, amountMinor, currency, payee, reference}. This far exceeds typical behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a dense single paragraph, but every sentence earns its place: purpose and non-transfer caveat, when-to-use and policy checks, refusal modes, approval flow, and the not-wired fallback. It is front-loaded with the core purpose, though breaking the wall of text into shorter sentences would improve scannability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description must cover result states and does: every possible answer (denied tokens, deferred:approval-required:<ref>, allowed:<ref>) is specified, along with the recorded authorization object and the not-wired case. For a complex async approval tool, an agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% — each parameter already has detailed descriptions including the policy constraints (amountMinor compared against cap and daily limit, currency must equal the policy currency, payee must be on the list, _ref reuse rules). The description reinforces how parameters map to refusal tokens but adds little beyond what the schema provides, so the high-coverage baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource — 'asks the body to authorize a payment and records the decision' — and immediately disambiguates by clarifying 'the body never moves money... not a transfer.' None of the sibling tools (memory.get, message.send, audit.explain) overlap with payment authorization, so an agent can identify this tool's role unmistakably.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use guidance: 'Use it before any payment so that amount, currency, payee and reference are checked against the policy.' It also narrates the full acceptance/rejection flow and the approval path via 'verax approve, or the panel.' It stops short of explicitly naming alternatives or exclusions, though no sibling is a payment alternative, so 4 rather than 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv0.1.1
    • First observedaudit.explain
    • First observedmemory.get
    • First observedmemory.put
    • First observedmessage.read
    • First observedmessage.send
    • First observedspend

TDQS

A4.4/5.0

Scored across 6 tools

Disambiguation5/5

Each tool has a clear, distinct purpose: memory.get/put for storage, audit.explain for decision lookup, message.read/send for messaging, and spend for payment authorization. There is no overlap or ambiguity; an agent can confidently select the right tool for a task.

Naming Consistency4/5

Most tools follow a consistent 'domain.verb' pattern (memory.get, memory.put, audit.explain, message.read, message.send), but 'spend' is a lone verb without a domain prefix, creating a minor deviation. The overall style is uniform and readable.

Tool Count5/5

With 6 tools, the surface is well-scoped for the domains covered (memory, messaging, audit, spending). Each tool serves a distinct operational need, and the count is neither sparse nor bloated.

Completeness4/5

Core operations are covered: memory supports get and put (including replace), messaging supports read and send, audit supports explain (though no list), and spend supports authorization. Minor gaps exist, such as no explicit delete for memory or messages and no audit listing, but these are not critical for the intended workflows.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    B
    maintenance
    Provides a secure MCP boundary for AI agents, intercepting and validating tool calls, redacting secrets, and requiring human approval for sensitive actions with a tamper-evident audit trail.
    -
  • A
    license
    A
    quality
    B
    maintenance
    Self-hosted governance layer between an AI assistant and your data: allow/deny policy, deterministic PII masking, row caps, and a hash-chained audit log with an Ed25519-signed receipt for every access, verifiable offline.
    4
    341 npm
    3
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Policy-gated agent spend with signed receipts and rail-extract audit: a fail-closed gate issues a COSE receipt for every allowed payment, and the audit reconciles those receipts against a rail extract so a settlement with no receipt is named rather than assumed. It settles on a mock rail, holding no wallet and signing no transaction.
    5
    Apache 2.0