Skip to main content
Glama

Accountability infrastructure for long-running AI agents.

Give every process an identity. Keep claims, evidence, reviews, and outcomes connected. Recover work across restarts, context loss, and handoffs.

Tests Python License DOI

An agent reports that its fix is done and the tests pass. By morning its session has restarted, its context is gone, and another process has taken over the task. Who said it? What supports it? Who challenged it? What happened? Each run leaves its own log, and the answers scatter across them.

UNITARES is self-hosted accountability infrastructure for operators running multiple AI agents. Its single-operator federation kernel connects independent runtimes to one operator-controlled server over MCP or HTTP, where they share a durable record while keeping their own models, tools, and runtimes. They interoperate with each other over their own transports or A2A; UNITARES is the record behind them, not the transport between them.

UNITARES preserves accountability across discontinuities in agent identity, context, process, and time. Agent work remains attributable, reviewable, and recoverable even when the process that started it is gone.

What UNITARES gives you

  • Identity and lineage — know which process acted and where inherited work came from. This is a record for attribution, not a credential; platform workload identity remains the credential layer.

  • Claims and evidence — retain important findings, corrections, and their provenance outside any one context window.

  • Governed review — preserve disagreement, conditions, and resolution as part of the work record.

  • Outcome grounding — connect predictions and check-ins to what later happened.

  • Runtime policy — return an action, reason, and next step at meaningful checkpoints in an agent's loop.

  • Reconstruction — give a successor the records needed to understand and continue earlier work.

The claim ledger gives the evidence status of each measured result.

Together, these form an operator-owned accountability layer across coding agents, research agents, background agents, and custom runtimes. What it adds to a record of what happened is adjudication: disagreement, conditions, and outcomes bound to the process that made the claim.

Related MCP server: hejdar-mcp

Install

With Git, curl, and Docker Compose installed, one command starts the latest verified release of the local operator stack:

v=$(curl -fsSL https://raw.githubusercontent.com/cirwel/unitares/master/PUBLISHED_VERSION) && git clone --branch "v$v" --depth 1 https://github.com/cirwel/unitares.git && cd unitares && docker compose up -d --wait

Connect MCP clients at http://localhost:8767/mcp/ or open the dashboard at http://localhost:8767/dashboard.

This provisions the server, PostgreSQL with AGE and pgvector, Redis, and the coordination plane.

Data lives in Docker named volumes keyed to the Compose project name, which is the checkout directory name (unitares). Re-running the one-liner therefore reuses an earlier install's database. For a clean start, run docker compose down -v in the old checkout first. See Reinstalling.

From the checkout, ./scripts/unitares model points the server at a model you run with Ollama, which turns on advisory consult answers and a first reviewer for dialectic reviews; without one, a review waits for a peer or the operator. ./scripts/unitares update later moves the install to the newest release: it backs up the database before any migration, applies them, restarts, and checks health. See Choose a model and Updating, which also covers the one-time step for installs made before update existed.

How it works

An agent joins the operator's UNITARES deployment and receives a process identity. During work it can publish selected findings and evidence, request structured review, report meaningful state transitions, and record outcomes. UNITARES keeps those records available to the operator and to later authorized processes.

The server runs alongside evals, sandboxes, and guardrails. It provides the continuity and accountability layer that connects their outputs over time. Core storage is self-hosted and runs on its own; the operator chooses which inference providers and integrations to connect.

Its EISV state model is runtime proprioception: a way to make changes in an agent process visible so operators can diagnose and act on them with evidence.

Where it is going

UNITARES is working toward an operator experience where a fleet can be brought under accountable operation in one step: identities are configured, handoffs are enforceable, important evidence survives, reviews bind to the work they govern, and outcomes are recorded where the next decision can use them.

The larger aim is infrastructure for agent systems that can accumulate useful experience without losing authorship, challenge, or operational control as they grow.

Start here

Goal

Guide

Operate a deployment

Operator manual

Connect an agent or application

MCP integration · Python SDK

Understand the product and architecture

Product definition · Architecture

Evaluate the claims

Evidence and limits · Reviewer Guide · Public dataset

Contribute

Contributing · Development guide

The documentation index covers deployment profiles, operations, security, compatibility, research, and the full tool surface.

Ecosystem

UNITARES works with the governance plugin for Codex and Claude Code, the host adapter for Hermes Agent and OpenAI-compatible clients, the public Python SDK, and the resident agent runtime. These are separate userlands connected by the same operator-owned record.

Citation and license

Kenny Wang (ORCID 0009-0006-7544-2374), CIRWEL Systems. See CITATION.cff for the versioned citation.

@misc{wang2026unitares,
  author = {Wang, Kenny},
  title  = {{UNITARES}: Information-Theoretic Governance of Heterogeneous Agent Fleets},
  year   = {2026},
  doi    = {10.5281/zenodo.19647159}
}

Apache License 2.0. See LICENSE and NOTICE.

Available Tools

13 tools
check_working_stateA
Read-onlyIdempotent

Read your current governance state and verdict without running a cycle, writing, or minting an identity. Only proof sent with the call reads your state (start_session's client_session_id, an X-Session-ID header or a verified continuity_token), never an inferred binding; otherwise a self-read is unbound and next_action says how to recover. agent_id, dropped on /mcp/, names the agent to read through use_tool or REST unless you are bound as a different agent (identity_mismatch); that read is marked identity_assurance.caller_proven=false on an inferred session. verbosity='standard' adds mode and basin with their meanings under raw_governance; verbosity='full' (alias lite=false) returns the full canonical diagnostics. sync_state also logs work and returns proceed or pause. get_governance_metrics returns this read's raw payload. EISV fields: E=Energy [0,1] (mixed-provenance capacity estimate); I=Information Integrity [0,1] (mixed-provenance calibration estimate); S=Entropy [0,1] (drift from the agent's own normal); V=Valence [-1,1] (EMA-smoothed E-I imbalance; positive=motion outruns integrity, negative=integrity outruns motion).

ParametersJSON Schema
NameRequiredDescriptionDefault
liteNoIf true (default), returns minimal essential metrics only. Set lite=false for full diagnostic data. verbosity, when given, takes precedence.
verbosityNoTier: minimal (default); standard: bare EISV and risk, verdict/basin/mode with meanings, guidance; full: diagnostics.
include_stateNoRetained for compatibility; has no effect on any tier. Use verbosity to choose what is returned. Accepts boolean or string ('true'/'false').
continuity_tokenNoSame-process rebind proof only; never a cross-process resume.
client_session_idNoBinding id for calls in this process; not a cross-process proof.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/no-destructive, and the description layers on genuinely additive behavior: which proofs bind the read (client_session_id, X-Session-ID, continuity_token) vs. an unbound self-read, the identity_mismatch case, and the identity_assurance.caller_proven=false marking on inferred sessions. This is well beyond the annotation surface.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded, but the body is a dense wall of run-on sentences with heavy parenthetical nesting and semicolon-chained clauses. The EISV glossary is useful, yet the identity-binding prose could be tightened considerably without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does the heavy lifting for returns: it defines E/I/S/V ranges and semantics, next_action recovery, and what each verbosity tier yields. It stops short of enumerating verdict/basin/mode values themselves, relying on the server to attach meanings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds real meaning: verbosity tiers ('standard adds mode and basin with their meanings'; 'full' is an alias for lite=false) and that continuity_token is same-process-only, 'never a cross-process resume.' The lite/verbosity precedence is clarified too.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a precise verb+resource ('Read your current governance state and verdict') and immediately bounds scope with negations ('without running a cycle, writing, or minting an identity'). It is explicitly distinguishable from siblings like sync_state and get_governance_metrics, which it names.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names alternatives and their conditions: sync_state 'also logs work and returns proceed or pause' vs. this pure read, and get_governance_metrics as the raw-payload route. It also explains the recovery path ('next_action says how to recover') and the binding preconditions for a valid read.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

consultA

Primary advisory model-help surface: send a brief, get back advisory model evidence, never a governed verdict — request_review produces that. effort='thorough' asks a strong model (Claude, Codex or Antigravity) from a family other than the caller's, when detectable. It needs privacy='cloud_allowed' and an operator extension a default install lacks (see list_inference_hosts); without both it fails unless allow_degraded=true, which returns a standard local answer instead. Requires a bound identity. Audited as event_type='consultation', readable by bound agents: route and keyed hashes, never text (key: record.hash_key). A success also updates your governance state. Use call_model or delegate_inference only for explicit provider, host, model or timeout control.

ParametersJSON Schema
NameRequiredDescriptionDefault
briefYesQuestion or material to send for advisory model help.
effortNostandard uses the lower-overhead inference lane; thorough requests the operator-authorized strong-model lane.standard
privacyNolocal confines routing to the configured local inference service; cloud_allowed permits, but does not require, external processing.local
purposeNoDesired advisory operation. critique remains model advice, not a governed peer-review verdict.answer
agent_idNoUUID; leave unset for yourself.
response_modeNocompact returns the advisory result and policy outcome; full adds a single diagnostics object with route and inference provenance.compact
allow_degradedNoAllow thorough effort to fall back to standard local inference. This never weakens the requested privacy policy.
continuity_tokenNoSame-process rebind proof only; never a cross-process resume.
client_session_idNoBinding id for calls in this process; not a cross-process proof.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only flag readOnly=false, openWorld=true, idempotent=false; the description adds far more: the thorough lane's strong-model cross-family routing, the cloud_allowed + operator-extension prerequisite, the degraded fallback semantics, the bound-identity requirement, the audit event_type and hash-only visibility, and that success mutates governance state. This is rich disclosure beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose and the key alternative, then layers prerequisites and side effects. Dense and information-rich, though the middle sentences are tightly packed with clauses; nothing is truly wasted, but it borders on overload for a single paragraph.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, yet the description covers the return shape (advisory result + policy outcome, diagnostics with route and provenance under full), the failure modes, the identity prerequisite, and the side effects. An agent has everything needed to invoke it correctly or fall back.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds real parameter meaning the schema lacks — that thorough needs privacy='cloud_allowed' plus an operator extension and otherwise fails unless allow_degraded=true, and that allow_degraded never weakens the requested privacy. It does not elaborate on purpose, response_mode, or the session/token params beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific action and resource ('send a brief, get back advisory model evidence') and explicitly separates itself from the sibling that produces verdicts ('never a governed verdict — request_review produces that'). An agent can distinguish consult from request_review, call_model, and delegate_inference without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use (advisory model help) and when-not (governed verdicts → request_review; explicit provider/host/model/timeout control → call_model or delegate_inference). Also names the preconditions (privacy=cloud_allowed + operator extension, else allow_degraded=true) and the identity requirement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

describe_toolA
Read-onlyIdempotent

Return one named tool's description, stability tier, operation, examples, and advertised JSON input schema. Use list_tools(lite=true) for the compact capability-name index or list_tools(lite=false) to browse rich catalog metadata. An unqualified describe call returns the full record because lite=false is the advertised default; pass lite=true for a first-line-plus-key-parameters summary. On a consolidated router, action=... narrows the response to that action's parameters.

ParametersJSON Schema
NameRequiredDescriptionDefault
liteNoIf true, return simplified schema with examples.
actionNoFor a consolidated router (knowledge, dialectic, observe, agent, ...): narrow the returned schema to the parameters this one action uses.
agent_idNoUUID; leave unset for yourself.
tool_nameYesExact name of the tool to describe.
include_schemaNoFull mode only (lite=false): include the tool's inputSchema (default true).
continuity_tokenNoSame-process rebind proof only; never a cross-process resume.
client_session_idNoBinding id for calls in this process; not a cross-process proof.
include_full_descriptionNoFull mode only (lite=false): include the full description; false keeps the first line (default true).

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish the safety profile (readOnly, idempotent, non-destructive, closed-world), so the bar is lower. The description still adds real behavioral context beyond them: the lite=false default means an unqualified call returns the full record, and lite=true downgrades to a summary. It doesn't cover error behavior for an unknown tool_name, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each load-bearing: what is returned first, then sibling routing, then default-mode behavior, then router nuance. No filler and the most decision-relevant information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly enumerates the return payload and the mode-dependent variations. Combined with 100% schema coverage and existing annotations, an agent has everything needed to call this correctly for a read-only inspection tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds genuine meaning: it clarifies that lite=true yields a 'first-line-plus-key-parameters summary' and that action narrows the response to that action's parameters. The remaining params (agent_id, continuity_token, client_session_id) are left to the schema, but the mode-defining params are well covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Return) and resource (one named tool's record) and enumerates exactly what the record contains: description, stability tier, operation, examples, and input schema. It is clearly distinguishable from the sibling list_tools, which returns many tools rather than one.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent to alternatives with conditions: list_tools(lite=true) for the compact index, list_tools(lite=false) for rich catalog metadata. It also explains what an unqualified call returns and how action=... behaves on a consolidated router, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

identityA

Resolve which agent this MCP session is bound to, or set a cosmetic display name. Not a plain read: a call carrying no proof argument at all is gated to a fresh mint, so it persists a new agent and reports on that one, marked caller_proven=false. A call carrying only a cosmetic name= skips that gate and can instead infer a co-located binding — pass client_session_id to get your own back. name= persists a cosmetic label only and never looks an agent up. For a fresh process call onboard(force_new=true). continuity_token is per-process ownership proof, not a transport-level claim: carrying it into another process re-opens silent resurrection.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoOptional COSMETIC display name; sets display_name only, never `agent_id` or `uuid`. Thread `uuid` across tools, not this.
resumeNoExplicitly resume existing identity
agent_idNoUUID; leave unset for yourself.
force_newNoForce new identity creation
agent_uuidNoResume a known identity by UUID directly. Skips session/name resolution. Returns error if not found.
model_typeNoOptional model type for distinct identity
continuity_tokenNoSame-process rebind proof only; never a cross-process resume.
client_session_idNoBinding id for calls in this process; not a cross-process proof.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial context beyond the annotations: it explains the side effect that justifies readOnlyHint=false ('Not a plain read… persists a new agent and reports on that one, marked caller_proven=false'), warns that a cosmetic name never looks an agent up, and clarifies continuity_token is per-process proof whose cross-process use 're-opens silent resurrection'. This is exactly the non-obvious behavior an agent needs. No contradiction with the provided hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, and nearly every clause carries distinct behavioral information. It is a dense single block, though, with compressed jargon ('silent resurrection', 'co-located binding') that costs readability for an 8-parameter dual-mode tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex identity tool with 8 optional params and no output schema, the description covers the hard parts — minting gate, cosmetic vs. identity fields, ownership-proof token, and a hint at the reported marker (caller_proven=false). It omits guidance on resume/model_type/agent_uuid and any return-shape detail, but those are largely covered by the 100%-covered schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds cross-parameter interaction semantics the schema does not carry: the gating rule keyed on presence/absence of proof arguments, and the note that name= is cosmetic and thread-safe guidance ('Thread uuid across tools, not this'). That is genuine value above the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific dual purpose: 'Resolve which agent this MCP session is bound to, or set a cosmetic display name.' The verb+resource pair is concrete and the two modes are distinguishable. Sibling differentiation is weak, though — it points to 'onboard(force_new=true)', which is not among the listed siblings, and never contrasts with the actual sibling 'start_session'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real branching guidance: a call with no proof argument is gated to a fresh mint, a call with only a cosmetic name= skips the gate and can infer a co-located binding, and pass client_session_id to get your own back. It routes to an alternative ('onboard') for a fresh process. It stops short of an explicit 'use this instead of X when…' rule against the real sibling set.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_toolsA
Read-onlyIdempotent

Discover the complete governance capability catalog, including names omitted from the initial progressive tools/list advertisement. The default lite=true response is the compact federation handshake: every public capability appears once as a name-only record beside the interface contract. Use lite=false for descriptions, categories, tiers, workflows, relationships, and direct-advertisement status. Use describe_tool for one capability's parameters, then use_tool to invoke a capability absent from the initial listing. Callable before an identity is bound.

ParametersJSON Schema
NameRequiredDescriptionDefault
liteNoIf true (default), return capability names and the interface contract; false adds descriptions, categories and tiers.
tierNoFilter by tier: 'essential', 'common', 'advanced', or 'all' (case-insensitive; blank means 'all').all
verboseNoIgnored; accepted for compatibility.
agent_idNoUUID; leave unset for yourself.
categoryNoFilter tools by category.
progressiveNoIf true, order tools by usage frequency.
essential_onlyNoIf true, return only Tier 1 (essential) tools.
continuity_tokenNoSame-process rebind proof only; never a cross-process resume.
include_advancedNoIf false, exclude Tier 3 (advanced) tools.
client_session_idNoBinding id for calls in this process; not a cross-process proof.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, non-destructive and closed-world, so the safety profile is covered. The description adds genuinely useful context beyond that: the two response modes and their contents, and the important operational fact that the tool can be called before identity binding.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and scoping, and most sentences carry routing value. Slightly dense with several chained instructions, but there is little outright waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter read-only catalog tool with no output schema, the description covers mode selection, scope, and the identity precondition adequately, and routes to the correct follow-up tools. It does not describe pagination or volume limits, a minor gap given the breadth of the catalog.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 10 parameters are already documented, and the schema even explains the lite default. The description largely restates the lite semantics and adds only marginal detail (workflows, relationships, direct-advertisement status) not covered elsewhere, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Discover') and resource ('complete governance capability catalog'), and immediately clarifies it includes names omitted from the initial progressive tools/list advertisement. It is clearly distinguished from sibling tools describe_tool and use_tool, which it explicitly references.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit branching guidance: use the default lite=true for the compact handshake, lite=false for the richer catalog, describe_tool for one capability's parameters, and use_tool to invoke capabilities missing from the initial listing. It also states the precondition that it is callable before an identity is bound.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_resultA

Record a measurable outcome and pair it with this agent's EISV snapshot so verdicts can be graded against what really happened. Pass the prediction_id from a sync_state reply to bind the outcome to that check-in's confidence; it is consumed on first use and TTL-bound (an hour by default). Needs a bound or explicit agent_id, and refuses under strict identity from an ephemeral session. Provenance cannot be self-attested here: verification_source is forced and provenance keys in detail are stripped. Use store_finding for durable knowledge. outcome_event is the canonical twin; this name adds a digest envelope and keeps the raw payload under raw_governance only with response_mode='full' or include_semantics=true, or when the write returned no outcome_id. EISV field definitions: describe_tool(tool_name='check_working_state').

ParametersJSON Schema
NameRequiredDescriptionDefault
detailNoType-specific metadata (e.g., mark_count, test_name, error_message)
is_badNoWhether this is a negative outcome. Inferred from type if omitted.
agent_idNoAgent ID. Falls back to session-bound agent_id if omitted.
confidenceNoAgent confidence at outcome time (0-1). Looked up from last check-in if omitted.
session_idNoOptional session id; falls back to client_session_id and then to context.
outcome_typeYesType of outcome event
outcome_scoreNoQuality score 0.0 (worst) to 1.0 (best). Inferred from type if omitted.
prediction_idNoTactical prediction id from a prior process_agent_update response.
response_modeNo'full' is an alias for include_semantics=true; 'lite' (default) returns the small snapshot.
decision_actionNoThe decision the agent took (e.g. 'proceed', 'pause').
continuity_tokenNoSame-process rebind proof only; never a cross-process resume.
client_session_idNoBinding id for calls in this process; not a cross-process proof.
include_semanticsNoIf true, the response's eisv_snapshot carries the full EISV ontology (state_semantics role table + hierarchy).
verification_sourceNoProvenance of this outcome.agent_reported_tool_result

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavior beyond the annotations: prediction_id is consumed on first use and TTL-bound (~1 hour), a bound or explicit agent_id is needed, it refuses under strict identity from an ephemeral session, and provenance cannot be self-attested (verification_source forced, provenance keys stripped). This is exactly the write-safety and state-mutation context an agent needs, and it is consistent with idempotentHint=false and destructiveHint=false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose, then layers binding, TTL, identity, and provenance rules in a compact sequence where each clause carries real information. It is dense and occasionally hard to parse (the outcome_event/raw_governance sentence), keeping it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Appropriately complete for a 14-parameter mutation tool with no output schema: it covers binding lifecycle, identity requirements, provenance constraints, response-mode effects, and points to describe_tool for EISV field definitions. Nothing critical to invoking it correctly appears missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is 3; the description earns above baseline by explaining why prediction_id exists (bind to a check-in's confidence), how response_mode/include_semantics gate the digest envelope and raw payload, and the forced verification_source. One minor wrinkle: description says prediction_id comes from a sync_state reply while the schema says process_agent_update response.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Record a measurable outcome') plus the reason it exists ('pair it with this agent's EISV snapshot so verdicts can be graded'). It explicitly differentiates from siblings by naming 'store_finding for durable knowledge' and 'outcome_event is the canonical twin', so an agent can distinguish it without opening another schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear routing context: pass prediction_id from a sync_state reply, use store_finding for durable knowledge, and it clarifies the relationship to the outcome_event twin. It stops short of a crisp 'use this when X, use outcome_event when Y' rule, leaving the digest-envelope vs raw-payload distinction somewhat implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

request_reviewA

Open a governed, on-record review session for this agent. issue_description is reused as the thesis by default, so one call can reach a verdict; pass use_brief_as_thesis=false for the two-call form. Requires a session-owned registered identity, refuses with SESSION_EXISTS while one is active, and returns skipped with no session when the agent is waiting_input. dialectic advances an open session; consult gives advisory evidence with no verdict.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonNoReason (for action=request/reassign)
agent_idNoFilter by agent (for action=get or list)
reasoningNoExplanation/reasoning
root_causeNoRoot cause analysis (for action=thesis/synthesis/consult)
continuity_tokenNoSame-process rebind proof only; never a cross-process resume.
client_session_idNoBinding id for calls in this process; not a cross-process proof.
issue_descriptionNoIssue description (action=request/quick)
proposed_conditionsNoConditions for resumption (for action=thesis/synthesis/consult)
use_brief_as_thesisNoFor action=request or thesis, reuse the issue description or saved session brief as the thesis instead of repeating it.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations already declaring readOnlyHint=false, openWorldHint=true, idempotentHint=false and destructiveHint=false, the description goes well beyond them by disclosing the identity requirement, the SESSION_EXISTS refusal, the waiting_input skip behavior, and the one-call vs two-call flow. These are exactly the mutation-state traits an agent needs to predict outcomes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded — the lead sentence establishes the core action before qualifications. The final clause about dialectic and consult is somewhat terse and assumes familiarity, costing a point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, non-idempotent, open-world session tool with no output schema, the description covers prerequisites, failure modes, and the alternate flow well. The residual gap is how the referenced dialectic/consult actions map onto this tool versus the sibling consult tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description does add real meaning for issue_description and use_brief_as_thesis (thesis reuse semantics), but several schema descriptions reference an 'action' parameter that does not exist in the schema, leaving the parameter model partially unclear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Open a governed, on-record review session for this agent') and immediately distinguishes the mode from related capabilities, noting that dialectic advances an open session and consult yields advisory evidence with no verdict. An agent can tell this is the session-opening entry point rather than an advisory or continuation tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use conditions: issue_description is reused as the thesis by default so a single call can reach a verdict, and use_brief_as_thesis=false selects the two-call form. It also names prerequisites (session-owned registered identity) and the conditions that route the agent elsewhere or fail (SESSION_EXISTS while active, skipped when waiting_input).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_shared_memoryA

Search the cross-agent knowledge graph for prior findings. Rows in status archived or cold are excluded unless you set status explicitly or pass include_archived / include_cold; a resolved or closed finding is still returned. Reading is not free of effect: every successful search appends a knowledge_read audit row naming the reader and a redacted copy of the query, which is why this tool is not annotated read-only. It serves unbound callers, so it works before start_session, unlike the writes: use store_finding to add a finding and use_tool(tool_name='update_finding', ...) to revise one.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNoExact any-of tag filter.
limitNoMax results, 1-100 (more is capped).
queryNoSearch text.
statusNoFilter by status: open, resolved, archived, superseded.
sort_byNosearch order: relevance, or created_at (newest matches first).
agent_idNoFilter by author agent; agent_id_filter wins if both are set.
operatorNoBoolean operator for multi-term FTS queries
semanticNoLegacy semantic on/off toggle.
severityNoFilter by severity: low, medium, high, critical.
search_modeNoauto, or force fts/semantic/hybrid.
include_coldNoInclude cold-storage rows.
created_afterNosearch: created after ISO time.
response_modeNoRead-envelope mode. Default lean: one-line digests; compact adds diagnostics, full adds raw_governance.lean
authority_modeNoprefer_governed (default) down-ranks imported memory; all keeps raw order.
created_beforeNosearch: created before this ISO time.
discovery_typeNoFilter by discovery type, e.g. bug_found.
min_similarityNoSemantic similarity floor.
agent_id_filterNoAuthor agent UUID filter; wins over agent_id.
include_detailsNoInline details need response_mode='full'; else open one with knowledge(action='details').
continuity_tokenNoSame-process rebind proof only; never a cross-process resume.
include_archivedNoInclude archived rows.
client_session_idNoBinding id for calls in this process; not a cross-process proof.
include_provenanceNoAdd provenance fields.
exclude_agent_labelsNoOmit search results whose writer display label matches one of these values
recency_half_life_daysNosearch: score halves per N days.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false without explaining why; the description supplies exactly that missing context, disclosing that every successful search appends a knowledge_read audit row naming the reader plus a redacted query copy. It also documents the visibility side effect that archived/cold rows are hidden unless status/include_archived/include_cold are set, which is behavioral information the annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tightly packed sentences, front-loaded with purpose then caveats, side effects, and routing, with no filler. It is dense rather than bloated, though the 'resolved or closed' clause sits slightly awkwardly against the status enum, which costs a point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 25-parameter, no-required, no-output-schema search tool, the description covers the high-risk behaviors: default filtering, the audit side effect, call ordering relative to start_session, and sibling routing. Response envelope modes and pagination are left to the schema, which documents them well, so remaining gaps are minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real cross-parameter semantics: the interaction between the default status filter and include_archived/include_cold, and that resolved/closed findings still surface. It does not cover the 25 parameters exhaustively, but it lifts the most error-prone filtering logic above what the schema alone states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb and resource ('Search the cross-agent knowledge graph for prior findings') and explicitly separates itself from the write siblings by naming store_finding and update_finding. An agent can distinguish read-from-search versus add/revise without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States when it applies ('serves unbound callers, so it works before start_session, unlike the writes') and names the alternatives for adjacent operations (store_finding to add, use_tool(update_finding) to revise). The archived/cold exclusion rule and the resolved/closed exception give concrete filtering conditions rather than leaving them to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

self_recoveryA
Destructive

Lifts a pause or other hold on your own agent. action='check', the default, changes no stored state, and both resuming actions verify you own the agent. Neither resume path runs while a void is active; quick also caps risk at 0.40, review at 0.65 plus a written reflection (20+ characters) on what happened, which is recorded in the shared knowledge graph under your agent whether or not it resumes. An attempt that reaches the safety checks stamps a fresh recovery_attempt_at even when they refuse it, so a retry is not a no-op; a missing reflection is rejected before that stamp. To resume an agent you do not own use operator_resume_agent.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNoRecovery action: check (diagnose), quick (fast resume), review (with reflection)check
reasonNoBrief reason (optional for action=quick)
agent_idNoUUID; leave unset for yourself.
conditionsNoRecovery conditions (optional for action=review)
reflectionNoWhat went wrong and what you'll change (required for action=review)
continuity_tokenNoSame-process rebind proof only; never a cross-process resume.
client_session_idNoBinding id for calls in this process; not a cross-process proof.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite annotations covering the safety profile (destructiveHint=true, idempotentHint=false), the description adds substantive behavior: check is a pure no-op, neither resume path runs while a void is active, a refused attempt still stamps recovery_attempt_at, a missing reflection is rejected before that stamp, and the reflection is written to the shared knowledge graph regardless of outcome. These are exactly the non-obvious traits an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, and virtually every clause carries a distinct behavioral fact rather than filler. It is dense and multi-clause, which slightly taxes readability, but nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter, mutation-class tool with no output schema, the description covers triggers, mode selection, constraints, and side effects. The one gap is that it does not describe what action='check' returns diagnostically, which an agent calling the default mode would benefit from knowing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds requirements beyond the schema: the reflection minimum of 20 characters and its graph-recording side effect, and the risk caps attached to quick vs review. It also reinforces that continuity_token is a same-process-only proof, matching the schema hint.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Lifts a pause or other hold on your own agent') with immediate scope narrowing ('your own agent'). It also names the routing alternative operator_resume_agent, so an agent can tell it apart from the resume tools without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete mode semantics (check = no state change, quick = fast resume at 0.40 risk cap, review = 0.65 cap plus a 20+ character reflection) and a clear exclusion ('To resume an agent you do not own use operator_resume_agent'). Slightly implicit on the top-level trigger ('use this when your agent is paused'), which keeps it from a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_sessionA

Register this process-instance and mint its agent identity; keep the returned client_session_id for later calls. Call with force_new=true — a bare call with no ownership proof is defaulted to force_new or refused under strict identity, never resumed onto another process's uuid. parent_agent_id claims succession from an EXITED predecessor: naming a still-live parent is rejected as coincidental and the claim cleared, unless spawn_reason marks a dispatched child or a compaction continuation. Use identity to inspect or rename an existing binding. onboard is the canonical twin; this name adds a digest envelope. Read the uuid from agent_uuid; response_mode='full' keeps the raw payload under raw_governance.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoOptional COSMETIC display name; sets display_name only, never `agent_id` or `uuid`. Thread `uuid` across tools, not this.
resumeNoResume existing identity when a proof signal is present (continuity_token, agent_uuid, agent_id, client_session_id, or name).
agent_idNoUUID; leave unset for yourself.
force_newNoForce new identity creation.
thread_idNoExplicit thread ID to join (auto-derived from session if not provided)
model_typeNoOptional model type
client_hintNoClient hint string
orchestratedNoDeclare that a client_session_id is a thread-stable anchor provisioned by an orchestrator for a headless turn-child.
spawn_reasonNoWhy this fork was created. Registered reasons: subagent, dialectic_reviewer, dispatch, compaction, explicit, new_session.
initial_stateNoOptional bootstrap check-in payload.
response_modeNoVerbosity of the identity envelope.minimal
onboard_originNoAdapter-supplied observability label for the onboard entry path: agent, harness_backstop, or orchestrated_resume.
parent_agent_idNoUUID of predecessor agent (for fork lineage)
continuity_tokenNoSame-process rebind proof only; never a cross-process resume.
client_session_idNoBinding id for calls in this process; not a cross-process proof.
process_fingerprintNoOptional client-reported execution context: {host_id, pid, pid_start_time, transport, ppid?, tty?, anchor_path_hash?}.
trajectory_signatureNoTrajectory signature dict

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply generic hints (readOnly=false, idempotent=false, destructive=false). The description goes well beyond them, disclosing the force_new defaulting/refusal path under strict identity, the parent_agent_id succession rule (live parent rejected as coincidental), and response_mode='full' retaining the raw payload under raw_governance.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose, but the body is a chain of semicolon-joined clauses mixing identity rules, lineage rules, and response modes, making it hard to scan. Information density is high, but structure and readability suffer.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 17-param, zero-required identity bootstrap with no output schema, the description covers the essential call path and the critical failure modes (defaulting, refusal, lineage rejection). It omits the response envelope's full contents, but names the two key fields to retain (client_session_id, agent_uuid).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description still adds meaning the one-line schema entries lack: the interactive semantics of force_new (defaulted-or-refused, never resumed onto another uuid), the parent_agent_id validation outcome, and where the uuid surfaces (agent_uuid).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Register this process-instance and mint its agent identity,' and names siblings it differs from ('Use identity to inspect or rename an existing binding', 'onboard is the canonical twin'). The dense jargon slightly obscures the plain purpose, but the core action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete routing: use force_new=true, a bare call is defaulted or refused, use identity for inspect/rename, and onboard is the canonical twin. Covers when-to-use and some when-not, though it never gives a clean 'prefer onboard unless X' rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

store_findingA

Write one new durable finding into the cross-agent knowledge graph and get back its discovery_id. summary is required at call time even though the schema marks every field optional; severity high or critical is refused unless the session is bound to a registered agent, while low and medium fall back to an anonymous writer id. Every call mints a NEW discovery — search_shared_memory first, and revise one with use_tool(tool_name='update_finding', ...). Use it for a discovery, root cause or correction, and record_result for task, tool or test outcomes; budget 20 findings an hour.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNoTags (action=store, note); for action=search an exact any-of filter in every search mode.
contentNoExtended content/details (for action=store, note, or promote)
detailsNoExtended details for discovery (for action=store or promote). Alias: content
summaryNoDiscovery summary (for action=store or promote)
agent_idNoLeave unset; the bound session is the writer.
severityNoSeverity: low, medium, high, critical (for action=store or action=update)
task_labelNoS22 H5 provenance: human-readable bounded task label
task_outcomeNoS22 H5 provenance: outcome label for the bounded task
comparison_keyNoS22 H5 provenance: stable key for comparing the same bounded task across harnesses
discovery_typeNoaction=store; one of architectural_decision, learning, pattern, bug_fix, refactoring, documentation, experiment, question, note, rule, insight, bug_found, bug, improvement, exploration, observation.
memory_contextNoS22 provenance: memory/KG/transcript surfaces visible to the writer
continuity_tokenNoSame-process rebind proof only; never a cross-process resume.
client_session_idNoBinding id for calls in this process; not a cross-process proof.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond annotations: discloses that every call mints a NEW discovery (non-idempotent), a 20-findings/hour budget, the auth gating on high/critical severity, the anonymous-writer fallback for low/medium, and the required-summary quirk that contradicts the schema's all-optional marking. This is exactly the behavioral context an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: the core write action and return value come first, followed by constraints in descending priority. Every clause carries information, though the run-on sentence packing auth, fallback, alternatives, and budget is heavier than ideal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter mutation tool with no output schema, the description covers the return value, auth requirements, rate limits, idempotency behavior, and sibling routing. An agent has everything needed to call it correctly without opening other sources.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description adds real semantics on top: summary is required at call time despite the schema, and agent_id should be left unset because the bound session is the writer. Severity gating is also explained, though per-field syntax details stay in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Write one new durable finding into the cross-agent knowledge graph') plus the return value ('get back its discovery_id'). It clearly distinguishes itself from siblings by naming search_shared_memory, update_finding, and record_result as different operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit routing: search_shared_memory first, revise via use_tool(tool_name='update_finding'), and the boundary against record_result ('discovery, root cause or correction' vs 'task, tool or test outcomes'). Both when-to-use and which-alternative are spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sync_stateA

Record a work check-in and get a governance decision: it advances and persists this agent's EISV state and returns proceed or pause with a named reason, plus a prediction_id — only when you pass confidence — to grade that check-in later with record_result. The first call auto-binds an identity, except under strict identity, which refuses and points at start_session. simulate_update previews a proposed check-in without advancing state, though it still appends an audit event; check_working_state reads the current verdict without writing. process_agent_update is the canonical twin; this name returns a digest envelope, and a routine check-in omits the raw payload (response_mode='full' adds it under raw_governance). EISV field definitions: describe_tool(tool_name='check_working_state').

ParametersJSON Schema
NameRequiredDescriptionDefault
liteNoBoolean alias for response_mode='compact'. Applies only when response_mode is left at 'auto' (an explicit response_mode always wins).
logprobsNoPer-token top-k output logprobs, e.g. [[lp, lp, ...], ...]. Grounds S at tier-1 instead of the heuristic; absent for Claude.
task_typeNoTask type. Core types: convergent | divergent | mixed; 'introspection' for self-examination.mixed
complexityNoTask complexity, strictly 0-1. Check-in aliases also accept 'trivial'|'low'|'medium'|'high'|'very_high'.
confidenceNoConfidence level for this update (0-1, optional).
parametersNoAgent parameters vector (optional, deprecated).
task_labelNoS22 H5 provenance: human-readable bounded task label
sensor_dataNoCaller-published sensor measurements: `eisv` for a physical E/I/S/V reading, `afferents` for raw dimensions. Telemetry only, never a verdict input.
task_outcomeNoS22 H5 provenance: outcome label for the bounded task
ethical_driftNoEthical drift signals (3 components): [primary_drift, declared coherence_loss, complexity_contribution].
response_modeNoResponse shape. 'auto' (default) or 'compact' for routine check-ins; 'mirror' for actionable signals; 'full' for everything.auto
response_textNoAgent's response text (optional, for analysis)
comparison_keyNoS22 H5 provenance: stable key for comparing the same bounded task across harnesses
memory_contextNoS22 provenance: memory/KG/transcript surfaces visible to the writer
epistemic_classNoStorage label for this row: agent_report (default), substrate_observation (measured), substrate_interpretation (derived), prediction (forward claim).agent_report
continuity_tokenNoSame-process rebind proof only; never a cross-process resume.
client_session_idNoBinding id for calls in this process; not a cross-process proof.
provenance_contextNoSituating metadata: harness_type, model_provider, model, transport, tool_surface, governance_mode, verification_source, locus. Descriptive only.
recent_tool_resultsNoSelf-reported tool outcomes from the agent's most recent actions.
trajectory_signatureNoTrajectory identity signature from anima-mcp.
require_strong_identityNoIf true, reject updates unless identity assurance tier is strong.
include_memory_suggestionsNoOpt in to a KG lookup seeded from this check-in and include a few matching discovery digests.
auto_export_on_significanceNoIf true, write a governance history export whenever this check-in is judged significant: a risk spike, a coherence drop, a void threshold …

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false). The description adds substantial behavior beyond that: state advancement and persistence, identity auto-binding and strict-identity refusal, audit-event appending by the preview twin, and the digest-vs-raw envelope tradeoff with response_mode='full'. It does not spell out permission/auth requirements or idempotency semantics, which keeps it below a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action-and-return, then alternatives and envelope details. It is dense and paragraph-shaped with several clauses packed together, but each clause carries routing or behavioral information rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 23-parameter, zero-required, no-output-schema tool, the description covers the decision semantics, the returned fields (verdict, named reason, prediction_id), identity binding, response-shape modes, sibling routing, and defers EISV field definitions to describe_tool. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is 3. The description nonetheless adds real meaning: that prediction_id is returned only when confidence is supplied (a linkage the schema doesn't state), and that response_mode='full' surfaces the raw payload under raw_governance rather than just changing envelope shape.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Record a work check-in') and outcome ('get a governance decision ... proceed or pause with a named reason'), and explicitly distinguishes itself from siblings: simulate_update previews without advancing, check_working_state reads without writing, process_agent_update is the canonical twin. An agent can pick this apart from adjacent tools without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit routing: use record_result to grade a check-in later, start_session when strict identity refuses, simulate_update to preview, check_working_state for the current verdict. Also states the first-call auto-bind behavior and the one condition (strict identity) under which it fails.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

use_toolA
Destructive

Invoke one public capability omitted from the initial progressive tools/list advertisement. Find the exact name with list_tools and inspect its arguments with describe_tool, then pass that argument object here. The target's normal identity, validation, authorization, timeout and response middleware all run; this is a discovery gateway, not an authorization bypass. It refuses recursive use_tool calls.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idNoUUID; leave unset for yourself.
argumentsNoArguments for the target capability; inspect its schema with describe_tool before invoking it.
tool_nameYesExact public capability name returned by list_tools.
continuity_tokenNoSame-process rebind proof only; never a cross-process resume.
client_session_idNoBinding id for calls in this process; not a cross-process proof.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, openWorldHint=true, and non-idempotent, so the risk profile is covered. The description adds meaningful context beyond that: the target's own identity, validation, authorization, timeout and response middleware all run, so this gateway is not a privilege escalation path. It stops short of describing error or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, front-loaded with the action and the prerequisite chain, then the safety clarification and the recursion exclusion. The middleware enumeration is slightly long but each clause carries a distinct behavioral fact worth keeping.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description should orient the agent on return expectations, which it only gestures at implicitly. For a five-parameter gateway with nested arguments, it otherwise covers prerequisites, delegation semantics, and the key failure mode (recursive refusal) adequately.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with all five parameters documented, so the baseline is 3. The description only reinforces that arguments must be built from describe_tool output and that tool_name must be the exact name from list_tools, adding marginal value over the schema's own field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: invoke one public capability discovered via list_tools. It also carves out exactly what it is not ('a discovery gateway, not an authorization bypass'), which distinguishes it from siblings like list_tools, describe_tool, and the session tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit sequence: find the name with list_tools, inspect arguments with describe_tool, then pass the argument object here. It also states an exclusion — it refuses recursive use_tool calls — so the agent knows both when to use it and one case where it will not work.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev3.2.0
    • Changedsearch_shared_memory3 fields changed
      • changedInput schema / properties / agent_id / description
        Previous value: -"Filter by agent (for action=get, search; omit when using discovery_id readback)"New value: +"Filter by author agent; agent_id_filter wins if both are set."
      • changedInput schema / properties / query / description
        Previous value: -"Search query (for action=search)"New value: +"Search text."
      • changedInput schema / properties / status / description
        Previous value: -"Status filter/update value (open, resolved, archived, superseded)"New value: +"Filter by status: open, resolved, archived, superseded."
  2. 16 tool updatesv3.1.0
    • Changedcheck_working_state5 fields changed
      • changedInput schema / properties / client_session_id / description
        Previous value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof."
      • changedInput schema / properties / continuity_token / description
        Previous value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume."
      • changedInput schema / properties / include_state / description
        Previous value: -"Include nested state dict in response (can be large). Default false to reduce context bloat. Accepts boolean or string ('true'/'false')."New value: +"Retained for compatibility; has no effect on any tier. Use verbosity to choose what is returned. Accepts boolean or string ('true'/'false')."
      • changedInput schema / properties / lite / description
        Previous value: -"If true (default), returns minimal essential metrics only. Set lite=false for full diagnostic data."New value: +"If true (default), returns minimal essential metrics only. Set lite=false for full diagnostic data. verbosity, when given, takes precedence."
      • addedInput schema / properties / verbosity
        Added value: +{
        +  "anyOf": [
        +    {
        +      "enum": [
        +        "minimal",
        +        "standard",
        +        "full"
        +      ],
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Tier: minimal (default); standard: bare EISV and risk, verdict/basin/mode with meanings, guidance; full: diagnostics."
        +}
    • Changedconsult3 fields changed
      • changedInput schema / properties / agent_id / description
        Previous value: -"UNIQUE agent identifier; optional when session-bound (auto-injected)."New value: +"UUID; leave unset for yourself."
      • changedInput schema / properties / client_session_id / description
        Previous value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof."
      • changedInput schema / properties / continuity_token / description
        Previous value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume."
    • Changeddescribe_tool3 fields changed
      • changedInput schema / properties / agent_id / description
        Previous value: -"UNIQUE agent identifier; optional when session-bound (auto-injected)."New value: +"UUID; leave unset for yourself."
      • changedInput schema / properties / client_session_id / description
        Previous value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof."
      • changedInput schema / properties / continuity_token / description
        Previous value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume."
    • Removedhealth_check
    • Changedidentity4 fields changed
      • changedInput schema / description
        Previous value: -"Who am I? Auto-creates identity if first call."New value: +"Resolve this session's bound agent (pass client_session_id), or set a cosmetic display name."
      • changedInput schema / properties / agent_id / description
        Previous value: -"UNIQUE agent identifier; optional when session-bound (auto-injected)."New value: +"UUID; leave unset for yourself."
      • changedInput schema / properties / client_session_id / description
        Previous value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof."
      • changedInput schema / properties / continuity_token / description
        Previous value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume."
    • Removedknowledge
    • Addedlist_tools
    • Changedrecord_result3 fields changed
      • changedInput schema / description
        Previous value: -"Parameters for outcome_event"New value: +"Outcome record parameters."
      • changedInput schema / properties / client_session_id / description
        Previous value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof."
      • changedInput schema / properties / continuity_token / description
        Previous value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume."
    • Changedrequest_review6 fields changed
      • changedInput schema / description
        Previous value: -"Parameters for dialectic"New value: +"Review session parameters."
      • changedInput schema / properties / client_session_id / description
        Previous value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof."
      • changedInput schema / properties / continuity_token / description
        Previous value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume."
      • changedInput schema / properties / issue_description / description
        Previous value: -"Issue description (for action=request)"New value: +"Issue description (action=request/quick)"
      • changedInput schema / properties / proposed_conditions / description
        Previous value: -"Conditions for resumption (for action=thesis/synthesis)"New value: +"Conditions for resumption (for action=thesis/synthesis/consult)"
      • changedInput schema / properties / root_cause / description
        Previous value: -"Root cause analysis (for action=thesis/synthesis)"New value: +"Root cause analysis (for action=thesis/synthesis/consult)"
    • Changedsearch_shared_memory37 fields changed
      • changedInput schema / description
        Previous value: -"Parameters for knowledge"New value: +"Knowledge graph parameters."
      • addedInput schema / properties / agent_id_filter
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Author agent UUID filter; wins over agent_id."
        +}
      • addedInput schema / properties / authority_mode
        Added value: +{
        +  "anyOf": [
        +    {
        +      "enum": [
        +        "prefer_governed",
        +        "all"
        +      ],
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "prefer_governed (default) down-ranks imported memory; all keeps raw order."
        +}
      • changedInput schema / properties / client_session_id / description
        Previous value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof."
      • removedInput schema / properties / closure_class
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "type": "string"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Closing standard for action=update: fix_verified | unobserved | not_reproducible | obsolete | duplicate. 'fix_verified' needs a deployed change whose effect was observed."
        -}
      • removedInput schema / properties / closure_evidence
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "additionalProperties": true,
        -      "type": "object"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Evidence for closure_class. Required keys: fix_verified needs {deployed, observed}; unobserved needs {window, instrument_check}."
        -}
      • removedInput schema / properties / confidence
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "type": "number"
        -    },
        -    {
        -      "type": "string"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Writer-supplied confidence for action=store, validated to 0-1; not independently verified or adjusted by governance metrics"
        -}
      • changedInput schema / properties / continuity_token / description
        Previous value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume."
      • addedInput schema / properties / created_after
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "search: created after ISO time."
        +}
      • addedInput schema / properties / created_before
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "search: created before this ISO time."
        +}
      • changedInput schema / properties / discovery_type / description
        Previous value: -"action=store; one of architectural_decision, learning, pattern, bug_fix, refactoring, documentation, experiment, question, note, rule, insight, bug_found, bug, improvement, exploration, observation."New value: +"Filter by discovery type, e.g. bug_found."
      • removedInput schema / properties / dry_run
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "type": "boolean"
        -    },
        -    {
        -      "type": "string"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Dry run mode (for action=cleanup, synthesize)"
        -}
      • removedInput schema / properties / epoch_scope
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "enum": [
        -        "current",
        -        "all"
        -      ],
        -      "type": "string"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Stats/list scope: current epoch only or all epochs"
        -}
      • changedInput schema / properties / include_archived / description
        Previous value: -"Include archived discoveries in search results (default: excluded)"New value: +"Include archived rows."
      • changedInput schema / properties / include_cold / description
        Previous value: -"Include cold-storage (long-term) discoveries in search results (default: excluded)"New value: +"Include cold-storage rows."
      • changedInput schema / properties / include_details / description
        Previous value: -"Expand results inline only with response_mode='full'; otherwise open one with knowledge(action='details')."New value: +"Inline details need response_mode='full'; else open one with knowledge(action='details')."
      • changedInput schema / properties / include_provenance / description
        Previous value: -"Include provenance and lineage chain fields in search/details results"New value: +"Add provenance fields."
      • removedInput schema / properties / include_response_chain
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "type": "boolean"
        -    },
        -    {
        -      "type": "string"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Include typed response chain for action=details"
        -}
      • removedInput schema / properties / including_cold
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "type": "boolean"
        -    },
        -    {
        -      "type": "string"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Include cold-storage discoveries in action=list raw status aggregates"
        -}
      • removedInput schema / properties / length
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "type": "integer"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Maximum details characters returned for action=details"
        -}
      • changedInput schema / properties / limit / description
        Previous value: -"Max results (for action=search: min 1, values above 100 are capped, 0 or negative is rejected)"New value: +"Max results, 1-100 (more is capped)."
      • removedInput schema / properties / max_chain_depth
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "type": "integer"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Maximum response-chain traversal depth for action=details"
        -}
      • removedInput schema / properties / memory_context
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "type": "string"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "S22 provenance: memory/KG/transcript surfaces visible to the writer"
        -}
      • removedInput schema / properties / min_members
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "type": "integer"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Minimum discoveries a topic needs before it is rolled up (for action=synthesize, default 3)"
        -}
      • changedInput schema / properties / min_similarity / description
        Previous value: -"Minimum cosine similarity for semantic retrieval modes"New value: +"Semantic similarity floor."
      • removedInput schema / properties / offset
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "type": "integer"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Character offset for action=details pagination"
        -}
      • addedInput schema / properties / recency_half_life_days
        Added value: +{
        +  "anyOf": [
        +    {
        +      "exclusiveMinimum": 0,
        +      "maximum": 36500,
        +      "type": "number"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "search: score halves per N days."
        +}
      • removedInput schema / properties / scope
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "enum": [
        -        "open",
        -        "all",
        -        "by_agent"
        -      ],
        -      "type": "string"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "KG audit scope for action=audit"
        -}
      • changedInput schema / properties / search_mode / description
        Previous value: -"Force retrieval mode for action=search. 'semantic' and 'hybrid' fail honestly when unsupported by the active backend."New value: +"auto, or force fts/semantic/hybrid."
      • changedInput schema / properties / semantic / description
        Previous value: -"Legacy action=search toggle to force or skip semantic retrieval when supported"New value: +"Legacy semantic on/off toggle."
      • changedInput schema / properties / severity / description
        Previous value: -"Severity: low, medium, high, critical (for action=store or action=update)"New value: +"Filter by severity: low, medium, high, critical."
      • addedInput schema / properties / sort_by
        Added value: +{
        +  "anyOf": [
        +    {
        +      "enum": [
        +        "relevance",
        +        "created_at"
        +      ],
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "search order: relevance, or created_at (newest matches first)."
        +}
      • changedInput schema / properties / tags / description
        Previous value: -"Tags (action=store, note); for action=search an exact any-of filter in every search mode."New value: +"Exact any-of tag filter."
      • removedInput schema / properties / top_n
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "type": "integer"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Maximum stale entries returned by action=audit"
        -}
      • removedInput schema / properties / topic
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "type": "string"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Synthesize just this one tag/topic (for action=synthesize). Omit to sweep the densest topics."
        -}
      • removedInput schema / properties / use_llm
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "type": "boolean"
        -    },
        -    {
        -      "type": "string"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Use the local LLM for the rollup narrative (for action=synthesize, default true; falls back to deterministic when unreachable)"
        -}
      • removedInput schema / properties / use_model
        Removed value: -{
        -  "anyOf": [
        -    {
        -      "type": "boolean"
        -    },
        -    {
        -      "type": "string"
        -    },
        -    {
        -      "type": "null"
        -    }
        -  ],
        -  "default": null,
        -  "description": "Use the local model to assess stale entries for action=audit"
        -}
    • Changedself_recovery3 fields changed
      • changedInput schema / properties / agent_id / description
        Previous value: -"UNIQUE agent identifier; optional when session-bound (auto-injected)."New value: +"UUID; leave unset for yourself."
      • changedInput schema / properties / client_session_id / description
        Previous value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof."
      • changedInput schema / properties / continuity_token / description
        Previous value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume."
    • Changedstart_session3 fields changed
      • changedInput schema / properties / agent_id / description
        Previous value: -"UNIQUE agent identifier; optional when session-bound (auto-injected)."New value: +"UUID; leave unset for yourself."
      • changedInput schema / properties / client_session_id / description
        Previous value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof."
      • changedInput schema / properties / continuity_token / description
        Previous value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume."
    • Changedstore_finding7 fields changed
      • changedInput schema / description
        Previous value: -"Parameters for knowledge"New value: +"Knowledge graph parameters."
      • changedInput schema / properties / agent_id / description
        Previous value: -"Filter by agent (for action=get, search; omit when using discovery_id readback)"New value: +"Leave unset; the bound session is the writer."
      • changedInput schema / properties / client_session_id / description
        Previous value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof."
      • changedInput schema / properties / content / description
        Previous value: -"Extended content/details (for action=store or action=note)"New value: +"Extended content/details (for action=store, note, or promote)"
      • changedInput schema / properties / continuity_token / description
        Previous value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume."
      • changedInput schema / properties / details / description
        Previous value: -"Extended details for discovery (for action=store). Alias: content"New value: +"Extended details for discovery (for action=store or promote). Alias: content"
      • changedInput schema / properties / summary / description
        Previous value: -"Discovery summary (for action=store)"New value: +"Discovery summary (for action=store or promote)"
    • Changedsync_state6 fields changed
      • changedInput schema / description
        Previous value: -"Share your work and get supportive feedback. Your main tool for checking in."New value: +"Record a work check-in and get a governance decision."
      • changedInput schema / properties / auto_export_on_significance / description
        Previous value: -"If true, automatically export governance history when thermodynamically significant events occur."New value: +"If true, write a governance history export whenever this check-in is judged significant: a risk spike, a coherence drop, a void threshold …"
      • changedInput schema / properties / client_session_id / description
        Previous value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof."
      • changedInput schema / properties / complexity / anyOf
        Previous value: -[
        -  {
        -    "anyOf": [
        -      {
        -        "type": "number"
        -      },
        -      {
        -        "type": "string"
        -      }
        -    ],
        -    "ge": 0,
        -    "le": 1
        -  },
        -  {
        -    "type": "null"
        -  }
        -]New value: +[
        +  {
        +    "maximum": 1,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  {
        +    "anyOf": [
        +      {
        +        "pattern": "^(0(\\.\\d+)?|1(\\.0+)?|\\.\\d+)$",
        +        "type": "string"
        +      },
        +      {
        +        "enum": [
        +          "complex",
        +          "critical",
        +          "high",
        +          "low",
        +          "medium",
        +          "minimal",
        +          "moderate",
        +          "simple",
        +          "trivial",
        +          "very_high"
        +        ],
        +        "type": "string"
        +      }
        +    ],
        +    "type": "string"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • changedInput schema / properties / confidence / anyOf
        Previous value: -[
        -  {
        -    "anyOf": [
        -      {
        -        "type": "number"
        -      },
        -      {
        -        "type": "string"
        -      }
        -    ],
        -    "ge": 0,
        -    "le": 1
        -  },
        -  {
        -    "type": "null"
        -  }
        -]New value: +[
        +  {
        +    "maximum": 1,
        +    "minimum": 0,
        +    "type": "number"
        +  },
        +  {
        +    "pattern": "^(0(\\.\\d+)?|1(\\.0+)?|\\.\\d+)$",
        +    "type": "string"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • changedInput schema / properties / continuity_token / description
        Previous value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume."
    • Removedupdate_finding
    • Addeduse_tool
  3. 14 tool updates
    • First observedcheck_working_state
    • First observedconsult
    • First observeddescribe_tool
    • First observedhealth_check
    • First observedidentity
    • First observedknowledge
    • First observedrecord_result
    • First observedrequest_review
    • First observedsearch_shared_memory
    • First observedself_recovery
    • First observedstart_session
    • First observedstore_finding
    • First observedsync_state
    • First observedupdate_finding

TDQS

A4.1/5.0

Scored across 13 tools

Disambiguation3/5

Several tools overlap by design: sync_state and process_agent_update are 'canonical twins', start_session and onboard likewise, and check_working_state vs sync_state vs identity all touch governance state reads. The dense descriptions do disambiguate the intended path, but the twin/alias pattern and the list_tools/describe_tool/use_tool discovery layer duplicating the advertised surface create real misselection risk.

Naming Consistency4/5

Names are consistently snake_case, and most follow a verb_noun pattern (sync_state, check_working_state, search_shared_memory, store_finding, start_session, record_result, request_review, list_tools, describe_tool, use_tool). A few break the pattern (consult, identity, self_recovery) but remain readable and unambiguous in style.

Tool Count4/5

13 advertised tools is a well-scoped number for a governance server covering identity, state, memory, review, consultation, and recovery. However, the set is effectively larger since many capabilities (process_agent_update, onboard, outcome_event, update_finding, operator_resume_agent) are hidden behind use_tool, so the true surface is heavier than the count suggests.

Completeness4/5

The surface covers the lifecycle well: identity/session setup, state check-in and read, outcome recording, knowledge graph read/write, review, advisory consult, and self-recovery, plus a gateway to unlisted capabilities. Minor gaps remain (e.g., no explicit delete/archive for findings, revisions routed through the use_tool gateway), but core workflows are covered.

Maintenance

ActivityActive
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • F
    license
    A
    quality
    F
    maintenance
    Provides policy-based access control, incident tracking, and compliance monitoring to govern AI agent behavior. It enables organizations to enforce security rules and maintain audit trails by validating agent actions against trust levels and pattern-based policies.
    6
    1
    -
  • A
    license
    A
    quality
    C
    maintenance
    Runtime policy enforcement for AI agents. Evaluate every agent action against your organization's policies before execution, with observe and enforce modes.
    1
    1
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to self-govern by scanning code for hardcoded secrets, structural violations, and AI drift in real-time, providing fix packets for automatic remediation.
    27
    MIT