UNITARES
UNITARES is a self-hosted accountability server for AI agents, exposed as MCP tools for identity, governance check-ins, shared memory, review, and outcome tracking.
Identity & lineage —
start_session/onboardregister a process-instance, mint an agent UUID, resume or fork withparent_agent_id, and record spawn reason;identityresolves or renames the session binding.Governance check-ins —
sync_staterecords work, advances/persists EISV state, and returnsproceed/pausewith a reason plus aprediction_id;check_working_statereads the current verdict and EISV diagnostics without writing.Shared knowledge —
search_shared_memoryqueries the cross-agent knowledge graph (FTS, semantic, hybrid);store_findingwrites durable findings, root causes, and corrections.Outcome grounding —
record_resultpairs a measurable outcome with the agent's EISV snapshot and binds it to a priorprediction_idso verdicts can be graded.Governed review —
request_reviewopens an on-record review session (thesis, dialectic, synthesis, conditions) or requests advisory consult evidence.Advisory inference —
consultsends a brief to a model for advisory answers, critique, summarization, or generation, with local/cloud privacy and standard/thorough effort lanes.Self-recovery —
self_recoverychecks, quickly resumes, or resumes with a written reflection an agent that has been paused or held.Capability discovery —
list_tools,describe_tool, anduse_toolbrowse and invoke the full catalog of governance capabilities beyond the initially advertised tool list.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@UNITARESshow fleet health summary"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Accountability infrastructure for long-running AI agents.
Give every process an identity. Keep claims, evidence, reviews, and outcomes connected. Recover work across restarts, context loss, and handoffs.
An agent reports that its fix is done and the tests pass. By morning its session has restarted, its context is gone, and another process has taken over the task. Who said it? What supports it? Who challenged it? What happened? Each run leaves its own log, and the answers scatter across them.
UNITARES is self-hosted accountability infrastructure for operators running multiple AI agents. Its single-operator federation kernel connects independent runtimes to one operator-controlled server over MCP or HTTP, where they share a durable record while keeping their own models, tools, and runtimes. They interoperate with each other over their own transports or A2A; UNITARES is the record behind them, not the transport between them.
UNITARES preserves accountability across discontinuities in agent identity, context, process, and time. Agent work remains attributable, reviewable, and recoverable even when the process that started it is gone.
What UNITARES gives you
Identity and lineage — know which process acted and where inherited work came from. This is a record for attribution, not a credential; platform workload identity remains the credential layer.
Claims and evidence — retain important findings, corrections, and their provenance outside any one context window.
Governed review — preserve disagreement, conditions, and resolution as part of the work record.
Outcome grounding — connect predictions and check-ins to what later happened.
Runtime policy — return an action, reason, and next step at meaningful checkpoints in an agent's loop.
Reconstruction — give a successor the records needed to understand and continue earlier work.
The claim ledger gives the evidence status of each measured result.
Together, these form an operator-owned accountability layer across coding agents, research agents, background agents, and custom runtimes. What it adds to a record of what happened is adjudication: disagreement, conditions, and outcomes bound to the process that made the claim.
Related MCP server: hejdar-mcp
Install
With Git, curl, and Docker Compose installed, one command starts the latest verified release of the local operator stack:
v=$(curl -fsSL https://raw.githubusercontent.com/cirwel/unitares/master/PUBLISHED_VERSION) && git clone --branch "v$v" --depth 1 https://github.com/cirwel/unitares.git && cd unitares && docker compose up -d --waitConnect MCP clients at http://localhost:8767/mcp/ or open the dashboard at
http://localhost:8767/dashboard.
This provisions the server, PostgreSQL with AGE and pgvector, Redis, and the coordination plane.
Data lives in Docker named volumes keyed to the Compose project name, which is
the checkout directory name (unitares). Re-running the one-liner therefore
reuses an earlier install's database. For a clean start, run
docker compose down -v in the old checkout first. See
Reinstalling.
From the checkout, ./scripts/unitares model points the server at a model you
run with Ollama, which turns on advisory consult answers and a first reviewer
for dialectic reviews; without one, a review waits for a peer or the operator.
./scripts/unitares update later moves the install to the newest release: it
backs up the database before any migration, applies them, restarts, and checks health. See
Choose a model and
Updating, which also covers the one-time
step for installs made before update existed.
How it works
An agent joins the operator's UNITARES deployment and receives a process identity. During work it can publish selected findings and evidence, request structured review, report meaningful state transitions, and record outcomes. UNITARES keeps those records available to the operator and to later authorized processes.
The server runs alongside evals, sandboxes, and guardrails. It provides the continuity and accountability layer that connects their outputs over time. Core storage is self-hosted and runs on its own; the operator chooses which inference providers and integrations to connect.
Its EISV state model is runtime proprioception: a way to make changes in an agent process visible so operators can diagnose and act on them with evidence.
Where it is going
UNITARES is working toward an operator experience where a fleet can be brought under accountable operation in one step: identities are configured, handoffs are enforceable, important evidence survives, reviews bind to the work they govern, and outcomes are recorded where the next decision can use them.
The larger aim is infrastructure for agent systems that can accumulate useful experience without losing authorship, challenge, or operational control as they grow.
Start here
Goal | Guide |
Operate a deployment | |
Connect an agent or application | |
Understand the product and architecture | |
Evaluate the claims | |
Contribute |
The documentation index covers deployment profiles, operations, security, compatibility, research, and the full tool surface.
Ecosystem
UNITARES works with the governance plugin for Codex and Claude Code, the host adapter for Hermes Agent and OpenAI-compatible clients, the public Python SDK, and the resident agent runtime. These are separate userlands connected by the same operator-owned record.
Citation and license
Kenny Wang (ORCID 0009-0006-7544-2374),
CIRWEL Systems. See CITATION.cff for the versioned citation.
@misc{wang2026unitares,
author = {Wang, Kenny},
title = {{UNITARES}: Information-Theoretic Governance of Heterogeneous Agent Fleets},
year = {2026},
doi = {10.5281/zenodo.19647159}
}Available Tools
13 toolscheck_working_stateARead-onlyIdempotent
Read your current governance state and verdict without running a cycle, writing, or minting an identity. Only proof sent with the call reads your state (start_session's client_session_id, an X-Session-ID header or a verified continuity_token), never an inferred binding; otherwise a self-read is unbound and next_action says how to recover. agent_id, dropped on /mcp/, names the agent to read through use_tool or REST unless you are bound as a different agent (identity_mismatch); that read is marked identity_assurance.caller_proven=false on an inferred session. verbosity='standard' adds mode and basin with their meanings under raw_governance; verbosity='full' (alias lite=false) returns the full canonical diagnostics. sync_state also logs work and returns proceed or pause. get_governance_metrics returns this read's raw payload. EISV fields: E=Energy [0,1] (mixed-provenance capacity estimate); I=Information Integrity [0,1] (mixed-provenance calibration estimate); S=Entropy [0,1] (drift from the agent's own normal); V=Valence [-1,1] (EMA-smoothed E-I imbalance; positive=motion outruns integrity, negative=integrity outruns motion).
| Name | Required | Description | Default |
|---|---|---|---|
| lite | No | If true (default), returns minimal essential metrics only. Set lite=false for full diagnostic data. verbosity, when given, takes precedence. | |
| verbosity | No | Tier: minimal (default); standard: bare EISV and risk, verdict/basin/mode with meanings, guidance; full: diagnostics. | |
| include_state | No | Retained for compatibility; has no effect on any tier. Use verbosity to choose what is returned. Accepts boolean or string ('true'/'false'). | |
| continuity_token | No | Same-process rebind proof only; never a cross-process resume. | |
| client_session_id | No | Binding id for calls in this process; not a cross-process proof. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/no-destructive, and the description layers on genuinely additive behavior: which proofs bind the read (client_session_id, X-Session-ID, continuity_token) vs. an unbound self-read, the identity_mismatch case, and the identity_assurance.caller_proven=false marking on inferred sessions. This is well beyond the annotation surface.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded, but the body is a dense wall of run-on sentences with heavy parenthetical nesting and semicolon-chained clauses. The EISV glossary is useful, yet the identity-binding prose could be tightened considerably without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does the heavy lifting for returns: it defines E/I/S/V ranges and semantics, next_action recovery, and what each verbosity tier yields. It stops short of enumerating verdict/basin/mode values themselves, relying on the server to attach meanings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds real meaning: verbosity tiers ('standard adds mode and basin with their meanings'; 'full' is an alias for lite=false) and that continuity_token is same-process-only, 'never a cross-process resume.' The lite/verbosity precedence is clarified too.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a precise verb+resource ('Read your current governance state and verdict') and immediately bounds scope with negations ('without running a cycle, writing, or minting an identity'). It is explicitly distinguishable from siblings like sync_state and get_governance_metrics, which it names.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names alternatives and their conditions: sync_state 'also logs work and returns proceed or pause' vs. this pure read, and get_governance_metrics as the raw-payload route. It also explains the recovery path ('next_action says how to recover') and the binding preconditions for a valid read.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
consultA
Primary advisory model-help surface: send a brief, get back advisory model evidence, never a governed verdict — request_review produces that. effort='thorough' asks a strong model (Claude, Codex or Antigravity) from a family other than the caller's, when detectable. It needs privacy='cloud_allowed' and an operator extension a default install lacks (see list_inference_hosts); without both it fails unless allow_degraded=true, which returns a standard local answer instead. Requires a bound identity. Audited as event_type='consultation', readable by bound agents: route and keyed hashes, never text (key: record.hash_key). A success also updates your governance state. Use call_model or delegate_inference only for explicit provider, host, model or timeout control.
| Name | Required | Description | Default |
|---|---|---|---|
| brief | Yes | Question or material to send for advisory model help. | |
| effort | No | standard uses the lower-overhead inference lane; thorough requests the operator-authorized strong-model lane. | standard |
| privacy | No | local confines routing to the configured local inference service; cloud_allowed permits, but does not require, external processing. | local |
| purpose | No | Desired advisory operation. critique remains model advice, not a governed peer-review verdict. | answer |
| agent_id | No | UUID; leave unset for yourself. | |
| response_mode | No | compact returns the advisory result and policy outcome; full adds a single diagnostics object with route and inference provenance. | compact |
| allow_degraded | No | Allow thorough effort to fall back to standard local inference. This never weakens the requested privacy policy. | |
| continuity_token | No | Same-process rebind proof only; never a cross-process resume. | |
| client_session_id | No | Binding id for calls in this process; not a cross-process proof. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only flag readOnly=false, openWorld=true, idempotent=false; the description adds far more: the thorough lane's strong-model cross-family routing, the cloud_allowed + operator-extension prerequisite, the degraded fallback semantics, the bound-identity requirement, the audit event_type and hash-only visibility, and that success mutates governance state. This is rich disclosure beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose and the key alternative, then layers prerequisites and side effects. Dense and information-rich, though the middle sentences are tightly packed with clauses; nothing is truly wasted, but it borders on overload for a single paragraph.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, yet the description covers the return shape (advisory result + policy outcome, diagnostics with route and provenance under full), the failure modes, the identity prerequisite, and the side effects. An agent has everything needed to invoke it correctly or fall back.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds real parameter meaning the schema lacks — that thorough needs privacy='cloud_allowed' plus an operator extension and otherwise fails unless allow_degraded=true, and that allow_degraded never weakens the requested privacy. It does not elaborate on purpose, response_mode, or the session/token params beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific action and resource ('send a brief, get back advisory model evidence') and explicitly separates itself from the sibling that produces verdicts ('never a governed verdict — request_review produces that'). An agent can distinguish consult from request_review, call_model, and delegate_inference without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use (advisory model help) and when-not (governed verdicts → request_review; explicit provider/host/model/timeout control → call_model or delegate_inference). Also names the preconditions (privacy=cloud_allowed + operator extension, else allow_degraded=true) and the identity requirement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
describe_toolARead-onlyIdempotent
Return one named tool's description, stability tier, operation, examples, and advertised JSON input schema. Use list_tools(lite=true) for the compact capability-name index or list_tools(lite=false) to browse rich catalog metadata. An unqualified describe call returns the full record because lite=false is the advertised default; pass lite=true for a first-line-plus-key-parameters summary. On a consolidated router, action=... narrows the response to that action's parameters.
| Name | Required | Description | Default |
|---|---|---|---|
| lite | No | If true, return simplified schema with examples. | |
| action | No | For a consolidated router (knowledge, dialectic, observe, agent, ...): narrow the returned schema to the parameters this one action uses. | |
| agent_id | No | UUID; leave unset for yourself. | |
| tool_name | Yes | Exact name of the tool to describe. | |
| include_schema | No | Full mode only (lite=false): include the tool's inputSchema (default true). | |
| continuity_token | No | Same-process rebind proof only; never a cross-process resume. | |
| client_session_id | No | Binding id for calls in this process; not a cross-process proof. | |
| include_full_description | No | Full mode only (lite=false): include the full description; false keeps the first line (default true). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish the safety profile (readOnly, idempotent, non-destructive, closed-world), so the bar is lower. The description still adds real behavioral context beyond them: the lite=false default means an unqualified call returns the full record, and lite=true downgrades to a summary. It doesn't cover error behavior for an unknown tool_name, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each load-bearing: what is returned first, then sibling routing, then default-mode behavior, then router nuance. No filler and the most decision-relevant information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly enumerates the return payload and the mode-dependent variations. Combined with 100% schema coverage and existing annotations, an agent has everything needed to call this correctly for a read-only inspection tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds genuine meaning: it clarifies that lite=true yields a 'first-line-plus-key-parameters summary' and that action narrows the response to that action's parameters. The remaining params (agent_id, continuity_token, client_session_id) are left to the schema, but the mode-defining params are well covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Return) and resource (one named tool's record) and enumerates exactly what the record contains: description, stability tier, operation, examples, and input schema. It is clearly distinguishable from the sibling list_tools, which returns many tools rather than one.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent to alternatives with conditions: list_tools(lite=true) for the compact index, list_tools(lite=false) for rich catalog metadata. It also explains what an unqualified call returns and how action=... behaves on a consolidated router, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
identityA
Resolve which agent this MCP session is bound to, or set a cosmetic display name. Not a plain read: a call carrying no proof argument at all is gated to a fresh mint, so it persists a new agent and reports on that one, marked caller_proven=false. A call carrying only a cosmetic name= skips that gate and can instead infer a co-located binding — pass client_session_id to get your own back. name= persists a cosmetic label only and never looks an agent up. For a fresh process call onboard(force_new=true). continuity_token is per-process ownership proof, not a transport-level claim: carrying it into another process re-opens silent resurrection.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Optional COSMETIC display name; sets display_name only, never `agent_id` or `uuid`. Thread `uuid` across tools, not this. | |
| resume | No | Explicitly resume existing identity | |
| agent_id | No | UUID; leave unset for yourself. | |
| force_new | No | Force new identity creation | |
| agent_uuid | No | Resume a known identity by UUID directly. Skips session/name resolution. Returns error if not found. | |
| model_type | No | Optional model type for distinct identity | |
| continuity_token | No | Same-process rebind proof only; never a cross-process resume. | |
| client_session_id | No | Binding id for calls in this process; not a cross-process proof. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial context beyond the annotations: it explains the side effect that justifies readOnlyHint=false ('Not a plain read… persists a new agent and reports on that one, marked caller_proven=false'), warns that a cosmetic name never looks an agent up, and clarifies continuity_token is per-process proof whose cross-process use 're-opens silent resurrection'. This is exactly the non-obvious behavior an agent needs. No contradiction with the provided hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, and nearly every clause carries distinct behavioral information. It is a dense single block, though, with compressed jargon ('silent resurrection', 'co-located binding') that costs readability for an 8-parameter dual-mode tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex identity tool with 8 optional params and no output schema, the description covers the hard parts — minting gate, cosmetic vs. identity fields, ownership-proof token, and a hint at the reported marker (caller_proven=false). It omits guidance on resume/model_type/agent_uuid and any return-shape detail, but those are largely covered by the 100%-covered schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds cross-parameter interaction semantics the schema does not carry: the gating rule keyed on presence/absence of proof arguments, and the note that name= is cosmetic and thread-safe guidance ('Thread uuid across tools, not this'). That is genuine value above the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific dual purpose: 'Resolve which agent this MCP session is bound to, or set a cosmetic display name.' The verb+resource pair is concrete and the two modes are distinguishable. Sibling differentiation is weak, though — it points to 'onboard(force_new=true)', which is not among the listed siblings, and never contrasts with the actual sibling 'start_session'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives real branching guidance: a call with no proof argument is gated to a fresh mint, a call with only a cosmetic name= skips the gate and can infer a co-located binding, and pass client_session_id to get your own back. It routes to an alternative ('onboard') for a fresh process. It stops short of an explicit 'use this instead of X when…' rule against the real sibling set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_toolsARead-onlyIdempotent
Discover the complete governance capability catalog, including names omitted from the initial progressive tools/list advertisement. The default lite=true response is the compact federation handshake: every public capability appears once as a name-only record beside the interface contract. Use lite=false for descriptions, categories, tiers, workflows, relationships, and direct-advertisement status. Use describe_tool for one capability's parameters, then use_tool to invoke a capability absent from the initial listing. Callable before an identity is bound.
| Name | Required | Description | Default |
|---|---|---|---|
| lite | No | If true (default), return capability names and the interface contract; false adds descriptions, categories and tiers. | |
| tier | No | Filter by tier: 'essential', 'common', 'advanced', or 'all' (case-insensitive; blank means 'all'). | all |
| verbose | No | Ignored; accepted for compatibility. | |
| agent_id | No | UUID; leave unset for yourself. | |
| category | No | Filter tools by category. | |
| progressive | No | If true, order tools by usage frequency. | |
| essential_only | No | If true, return only Tier 1 (essential) tools. | |
| continuity_token | No | Same-process rebind proof only; never a cross-process resume. | |
| include_advanced | No | If false, exclude Tier 3 (advanced) tools. | |
| client_session_id | No | Binding id for calls in this process; not a cross-process proof. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive and closed-world, so the safety profile is covered. The description adds genuinely useful context beyond that: the two response modes and their contents, and the important operational fact that the tool can be called before identity binding.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and scoping, and most sentences carry routing value. Slightly dense with several chained instructions, but there is little outright waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter read-only catalog tool with no output schema, the description covers mode selection, scope, and the identity precondition adequately, and routes to the correct follow-up tools. It does not describe pagination or volume limits, a minor gap given the breadth of the catalog.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 10 parameters are already documented, and the schema even explains the lite default. The description largely restates the lite semantics and adds only marginal detail (workflows, relationships, direct-advertisement status) not covered elsewhere, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Discover') and resource ('complete governance capability catalog'), and immediately clarifies it includes names omitted from the initial progressive tools/list advertisement. It is clearly distinguished from sibling tools describe_tool and use_tool, which it explicitly references.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit branching guidance: use the default lite=true for the compact handshake, lite=false for the richer catalog, describe_tool for one capability's parameters, and use_tool to invoke capabilities missing from the initial listing. It also states the precondition that it is callable before an identity is bound.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_resultA
Record a measurable outcome and pair it with this agent's EISV snapshot so verdicts can be graded against what really happened. Pass the prediction_id from a sync_state reply to bind the outcome to that check-in's confidence; it is consumed on first use and TTL-bound (an hour by default). Needs a bound or explicit agent_id, and refuses under strict identity from an ephemeral session. Provenance cannot be self-attested here: verification_source is forced and provenance keys in detail are stripped. Use store_finding for durable knowledge. outcome_event is the canonical twin; this name adds a digest envelope and keeps the raw payload under raw_governance only with response_mode='full' or include_semantics=true, or when the write returned no outcome_id. EISV field definitions: describe_tool(tool_name='check_working_state').
| Name | Required | Description | Default |
|---|---|---|---|
| detail | No | Type-specific metadata (e.g., mark_count, test_name, error_message) | |
| is_bad | No | Whether this is a negative outcome. Inferred from type if omitted. | |
| agent_id | No | Agent ID. Falls back to session-bound agent_id if omitted. | |
| confidence | No | Agent confidence at outcome time (0-1). Looked up from last check-in if omitted. | |
| session_id | No | Optional session id; falls back to client_session_id and then to context. | |
| outcome_type | Yes | Type of outcome event | |
| outcome_score | No | Quality score 0.0 (worst) to 1.0 (best). Inferred from type if omitted. | |
| prediction_id | No | Tactical prediction id from a prior process_agent_update response. | |
| response_mode | No | 'full' is an alias for include_semantics=true; 'lite' (default) returns the small snapshot. | |
| decision_action | No | The decision the agent took (e.g. 'proceed', 'pause'). | |
| continuity_token | No | Same-process rebind proof only; never a cross-process resume. | |
| client_session_id | No | Binding id for calls in this process; not a cross-process proof. | |
| include_semantics | No | If true, the response's eisv_snapshot carries the full EISV ontology (state_semantics role table + hierarchy). | |
| verification_source | No | Provenance of this outcome. | agent_reported_tool_result |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial behavior beyond the annotations: prediction_id is consumed on first use and TTL-bound (~1 hour), a bound or explicit agent_id is needed, it refuses under strict identity from an ephemeral session, and provenance cannot be self-attested (verification_source forced, provenance keys stripped). This is exactly the write-safety and state-mutation context an agent needs, and it is consistent with idempotentHint=false and destructiveHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose, then layers binding, TTL, identity, and provenance rules in a compact sequence where each clause carries real information. It is dense and occasionally hard to parse (the outcome_event/raw_governance sentence), keeping it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Appropriately complete for a 14-parameter mutation tool with no output schema: it covers binding lifecycle, identity requirements, provenance constraints, response-mode effects, and points to describe_tool for EISV field definitions. Nothing critical to invoking it correctly appears missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3; the description earns above baseline by explaining why prediction_id exists (bind to a check-in's confidence), how response_mode/include_semantics gate the digest envelope and raw payload, and the forced verification_source. One minor wrinkle: description says prediction_id comes from a sync_state reply while the schema says process_agent_update response.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Record a measurable outcome') plus the reason it exists ('pair it with this agent's EISV snapshot so verdicts can be graded'). It explicitly differentiates from siblings by naming 'store_finding for durable knowledge' and 'outcome_event is the canonical twin', so an agent can distinguish it without opening another schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear routing context: pass prediction_id from a sync_state reply, use store_finding for durable knowledge, and it clarifies the relationship to the outcome_event twin. It stops short of a crisp 'use this when X, use outcome_event when Y' rule, leaving the digest-envelope vs raw-payload distinction somewhat implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_reviewA
Open a governed, on-record review session for this agent. issue_description is reused as the thesis by default, so one call can reach a verdict; pass use_brief_as_thesis=false for the two-call form. Requires a session-owned registered identity, refuses with SESSION_EXISTS while one is active, and returns skipped with no session when the agent is waiting_input. dialectic advances an open session; consult gives advisory evidence with no verdict.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | Reason (for action=request/reassign) | |
| agent_id | No | Filter by agent (for action=get or list) | |
| reasoning | No | Explanation/reasoning | |
| root_cause | No | Root cause analysis (for action=thesis/synthesis/consult) | |
| continuity_token | No | Same-process rebind proof only; never a cross-process resume. | |
| client_session_id | No | Binding id for calls in this process; not a cross-process proof. | |
| issue_description | No | Issue description (action=request/quick) | |
| proposed_conditions | No | Conditions for resumption (for action=thesis/synthesis/consult) | |
| use_brief_as_thesis | No | For action=request or thesis, reuse the issue description or saved session brief as the thesis instead of repeating it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations already declaring readOnlyHint=false, openWorldHint=true, idempotentHint=false and destructiveHint=false, the description goes well beyond them by disclosing the identity requirement, the SESSION_EXISTS refusal, the waiting_input skip behavior, and the one-call vs two-call flow. These are exactly the mutation-state traits an agent needs to predict outcomes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded — the lead sentence establishes the core action before qualifications. The final clause about dialectic and consult is somewhat terse and assumes familiarity, costing a point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, non-idempotent, open-world session tool with no output schema, the description covers prerequisites, failure modes, and the alternate flow well. The residual gap is how the referenced dialectic/consult actions map onto this tool versus the sibling consult tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description does add real meaning for issue_description and use_brief_as_thesis (thesis reuse semantics), but several schema descriptions reference an 'action' parameter that does not exist in the schema, leaving the parameter model partially unclear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Open a governed, on-record review session for this agent') and immediately distinguishes the mode from related capabilities, noting that dialectic advances an open session and consult yields advisory evidence with no verdict. An agent can tell this is the session-opening entry point rather than an advisory or continuation tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use conditions: issue_description is reused as the thesis by default so a single call can reach a verdict, and use_brief_as_thesis=false selects the two-call form. It also names prerequisites (session-owned registered identity) and the conditions that route the agent elsewhere or fail (SESSION_EXISTS while active, skipped when waiting_input).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
self_recoveryADestructive
Lifts a pause or other hold on your own agent. action='check', the default, changes no stored state, and both resuming actions verify you own the agent. Neither resume path runs while a void is active; quick also caps risk at 0.40, review at 0.65 plus a written reflection (20+ characters) on what happened, which is recorded in the shared knowledge graph under your agent whether or not it resumes. An attempt that reaches the safety checks stamps a fresh recovery_attempt_at even when they refuse it, so a retry is not a no-op; a missing reflection is rejected before that stamp. To resume an agent you do not own use operator_resume_agent.
| Name | Required | Description | Default |
|---|---|---|---|
| action | No | Recovery action: check (diagnose), quick (fast resume), review (with reflection) | check |
| reason | No | Brief reason (optional for action=quick) | |
| agent_id | No | UUID; leave unset for yourself. | |
| conditions | No | Recovery conditions (optional for action=review) | |
| reflection | No | What went wrong and what you'll change (required for action=review) | |
| continuity_token | No | Same-process rebind proof only; never a cross-process resume. | |
| client_session_id | No | Binding id for calls in this process; not a cross-process proof. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite annotations covering the safety profile (destructiveHint=true, idempotentHint=false), the description adds substantive behavior: check is a pure no-op, neither resume path runs while a void is active, a refused attempt still stamps recovery_attempt_at, a missing reflection is rejected before that stamp, and the reflection is written to the shared knowledge graph regardless of outcome. These are exactly the non-obvious traits an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, and virtually every clause carries a distinct behavioral fact rather than filler. It is dense and multi-clause, which slightly taxes readability, but nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter, mutation-class tool with no output schema, the description covers triggers, mode selection, constraints, and side effects. The one gap is that it does not describe what action='check' returns diagnostically, which an agent calling the default mode would benefit from knowing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds requirements beyond the schema: the reflection minimum of 20 characters and its graph-recording side effect, and the risk caps attached to quick vs review. It also reinforces that continuity_token is a same-process-only proof, matching the schema hint.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lifts a pause or other hold on your own agent') with immediate scope narrowing ('your own agent'). It also names the routing alternative operator_resume_agent, so an agent can tell it apart from the resume tools without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete mode semantics (check = no state change, quick = fast resume at 0.40 risk cap, review = 0.65 cap plus a 20+ character reflection) and a clear exclusion ('To resume an agent you do not own use operator_resume_agent'). Slightly implicit on the top-level trigger ('use this when your agent is paused'), which keeps it from a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_sessionA
Register this process-instance and mint its agent identity; keep the returned client_session_id for later calls. Call with force_new=true — a bare call with no ownership proof is defaulted to force_new or refused under strict identity, never resumed onto another process's uuid. parent_agent_id claims succession from an EXITED predecessor: naming a still-live parent is rejected as coincidental and the claim cleared, unless spawn_reason marks a dispatched child or a compaction continuation. Use identity to inspect or rename an existing binding. onboard is the canonical twin; this name adds a digest envelope. Read the uuid from agent_uuid; response_mode='full' keeps the raw payload under raw_governance.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Optional COSMETIC display name; sets display_name only, never `agent_id` or `uuid`. Thread `uuid` across tools, not this. | |
| resume | No | Resume existing identity when a proof signal is present (continuity_token, agent_uuid, agent_id, client_session_id, or name). | |
| agent_id | No | UUID; leave unset for yourself. | |
| force_new | No | Force new identity creation. | |
| thread_id | No | Explicit thread ID to join (auto-derived from session if not provided) | |
| model_type | No | Optional model type | |
| client_hint | No | Client hint string | |
| orchestrated | No | Declare that a client_session_id is a thread-stable anchor provisioned by an orchestrator for a headless turn-child. | |
| spawn_reason | No | Why this fork was created. Registered reasons: subagent, dialectic_reviewer, dispatch, compaction, explicit, new_session. | |
| initial_state | No | Optional bootstrap check-in payload. | |
| response_mode | No | Verbosity of the identity envelope. | minimal |
| onboard_origin | No | Adapter-supplied observability label for the onboard entry path: agent, harness_backstop, or orchestrated_resume. | |
| parent_agent_id | No | UUID of predecessor agent (for fork lineage) | |
| continuity_token | No | Same-process rebind proof only; never a cross-process resume. | |
| client_session_id | No | Binding id for calls in this process; not a cross-process proof. | |
| process_fingerprint | No | Optional client-reported execution context: {host_id, pid, pid_start_time, transport, ppid?, tty?, anchor_path_hash?}. | |
| trajectory_signature | No | Trajectory signature dict |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply generic hints (readOnly=false, idempotent=false, destructive=false). The description goes well beyond them, disclosing the force_new defaulting/refusal path under strict identity, the parent_agent_id succession rule (live parent rejected as coincidental), and response_mode='full' retaining the raw payload under raw_governance.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose, but the body is a chain of semicolon-joined clauses mixing identity rules, lineage rules, and response modes, making it hard to scan. Information density is high, but structure and readability suffer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 17-param, zero-required identity bootstrap with no output schema, the description covers the essential call path and the critical failure modes (defaulting, refusal, lineage rejection). It omits the response envelope's full contents, but names the two key fields to retain (client_session_id, agent_uuid).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description still adds meaning the one-line schema entries lack: the interactive semantics of force_new (defaulted-or-refused, never resumed onto another uuid), the parent_agent_id validation outcome, and where the uuid surfaces (agent_uuid).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Register this process-instance and mint its agent identity,' and names siblings it differs from ('Use identity to inspect or rename an existing binding', 'onboard is the canonical twin'). The dense jargon slightly obscures the plain purpose, but the core action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete routing: use force_new=true, a bare call is defaulted or refused, use identity for inspect/rename, and onboard is the canonical twin. Covers when-to-use and some when-not, though it never gives a clean 'prefer onboard unless X' rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
store_findingA
Write one new durable finding into the cross-agent knowledge graph and get back its discovery_id. summary is required at call time even though the schema marks every field optional; severity high or critical is refused unless the session is bound to a registered agent, while low and medium fall back to an anonymous writer id. Every call mints a NEW discovery — search_shared_memory first, and revise one with use_tool(tool_name='update_finding', ...). Use it for a discovery, root cause or correction, and record_result for task, tool or test outcomes; budget 20 findings an hour.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | Tags (action=store, note); for action=search an exact any-of filter in every search mode. | |
| content | No | Extended content/details (for action=store, note, or promote) | |
| details | No | Extended details for discovery (for action=store or promote). Alias: content | |
| summary | No | Discovery summary (for action=store or promote) | |
| agent_id | No | Leave unset; the bound session is the writer. | |
| severity | No | Severity: low, medium, high, critical (for action=store or action=update) | |
| task_label | No | S22 H5 provenance: human-readable bounded task label | |
| task_outcome | No | S22 H5 provenance: outcome label for the bounded task | |
| comparison_key | No | S22 H5 provenance: stable key for comparing the same bounded task across harnesses | |
| discovery_type | No | action=store; one of architectural_decision, learning, pattern, bug_fix, refactoring, documentation, experiment, question, note, rule, insight, bug_found, bug, improvement, exploration, observation. | |
| memory_context | No | S22 provenance: memory/KG/transcript surfaces visible to the writer | |
| continuity_token | No | Same-process rebind proof only; never a cross-process resume. | |
| client_session_id | No | Binding id for calls in this process; not a cross-process proof. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond annotations: discloses that every call mints a NEW discovery (non-idempotent), a 20-findings/hour budget, the auth gating on high/critical severity, the anonymous-writer fallback for low/medium, and the required-summary quirk that contradicts the schema's all-optional marking. This is exactly the behavioral context an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: the core write action and return value come first, followed by constraints in descending priority. Every clause carries information, though the run-on sentence packing auth, fallback, alternatives, and budget is heavier than ideal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter mutation tool with no output schema, the description covers the return value, auth requirements, rate limits, idempotency behavior, and sibling routing. An agent has everything needed to call it correctly without opening other sources.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description adds real semantics on top: summary is required at call time despite the schema, and agent_id should be left unset because the bound session is the writer. Severity gating is also explained, though per-field syntax details stay in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Write one new durable finding into the cross-agent knowledge graph') plus the return value ('get back its discovery_id'). It clearly distinguishes itself from siblings by naming search_shared_memory, update_finding, and record_result as different operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing: search_shared_memory first, revise via use_tool(tool_name='update_finding'), and the boundary against record_result ('discovery, root cause or correction' vs 'task, tool or test outcomes'). Both when-to-use and which-alternative are spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sync_stateA
Record a work check-in and get a governance decision: it advances and persists this agent's EISV state and returns proceed or pause with a named reason, plus a prediction_id — only when you pass confidence — to grade that check-in later with record_result. The first call auto-binds an identity, except under strict identity, which refuses and points at start_session. simulate_update previews a proposed check-in without advancing state, though it still appends an audit event; check_working_state reads the current verdict without writing. process_agent_update is the canonical twin; this name returns a digest envelope, and a routine check-in omits the raw payload (response_mode='full' adds it under raw_governance). EISV field definitions: describe_tool(tool_name='check_working_state').
| Name | Required | Description | Default |
|---|---|---|---|
| lite | No | Boolean alias for response_mode='compact'. Applies only when response_mode is left at 'auto' (an explicit response_mode always wins). | |
| logprobs | No | Per-token top-k output logprobs, e.g. [[lp, lp, ...], ...]. Grounds S at tier-1 instead of the heuristic; absent for Claude. | |
| task_type | No | Task type. Core types: convergent | divergent | mixed; 'introspection' for self-examination. | mixed |
| complexity | No | Task complexity, strictly 0-1. Check-in aliases also accept 'trivial'|'low'|'medium'|'high'|'very_high'. | |
| confidence | No | Confidence level for this update (0-1, optional). | |
| parameters | No | Agent parameters vector (optional, deprecated). | |
| task_label | No | S22 H5 provenance: human-readable bounded task label | |
| sensor_data | No | Caller-published sensor measurements: `eisv` for a physical E/I/S/V reading, `afferents` for raw dimensions. Telemetry only, never a verdict input. | |
| task_outcome | No | S22 H5 provenance: outcome label for the bounded task | |
| ethical_drift | No | Ethical drift signals (3 components): [primary_drift, declared coherence_loss, complexity_contribution]. | |
| response_mode | No | Response shape. 'auto' (default) or 'compact' for routine check-ins; 'mirror' for actionable signals; 'full' for everything. | auto |
| response_text | No | Agent's response text (optional, for analysis) | |
| comparison_key | No | S22 H5 provenance: stable key for comparing the same bounded task across harnesses | |
| memory_context | No | S22 provenance: memory/KG/transcript surfaces visible to the writer | |
| epistemic_class | No | Storage label for this row: agent_report (default), substrate_observation (measured), substrate_interpretation (derived), prediction (forward claim). | agent_report |
| continuity_token | No | Same-process rebind proof only; never a cross-process resume. | |
| client_session_id | No | Binding id for calls in this process; not a cross-process proof. | |
| provenance_context | No | Situating metadata: harness_type, model_provider, model, transport, tool_surface, governance_mode, verification_source, locus. Descriptive only. | |
| recent_tool_results | No | Self-reported tool outcomes from the agent's most recent actions. | |
| trajectory_signature | No | Trajectory identity signature from anima-mcp. | |
| require_strong_identity | No | If true, reject updates unless identity assurance tier is strong. | |
| include_memory_suggestions | No | Opt in to a KG lookup seeded from this check-in and include a few matching discovery digests. | |
| auto_export_on_significance | No | If true, write a governance history export whenever this check-in is judged significant: a risk spike, a coherence drop, a void threshold … |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false). The description adds substantial behavior beyond that: state advancement and persistence, identity auto-binding and strict-identity refusal, audit-event appending by the preview twin, and the digest-vs-raw envelope tradeoff with response_mode='full'. It does not spell out permission/auth requirements or idempotency semantics, which keeps it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action-and-return, then alternatives and envelope details. It is dense and paragraph-shaped with several clauses packed together, but each clause carries routing or behavioral information rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 23-parameter, zero-required, no-output-schema tool, the description covers the decision semantics, the returned fields (verdict, named reason, prediction_id), identity binding, response-shape modes, sibling routing, and defers EISV field definitions to describe_tool. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3. The description nonetheless adds real meaning: that prediction_id is returned only when confidence is supplied (a linkage the schema doesn't state), and that response_mode='full' surfaces the raw payload under raw_governance rather than just changing envelope shape.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Record a work check-in') and outcome ('get a governance decision ... proceed or pause with a named reason'), and explicitly distinguishes itself from siblings: simulate_update previews without advancing, check_working_state reads without writing, process_agent_update is the canonical twin. An agent can pick this apart from adjacent tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing: use record_result to grade a check-in later, start_session when strict identity refuses, simulate_update to preview, check_working_state for the current verdict. Also states the first-call auto-bind behavior and the one condition (strict identity) under which it fails.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
use_toolADestructive
Invoke one public capability omitted from the initial progressive tools/list advertisement. Find the exact name with list_tools and inspect its arguments with describe_tool, then pass that argument object here. The target's normal identity, validation, authorization, timeout and response middleware all run; this is a discovery gateway, not an authorization bypass. It refuses recursive use_tool calls.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | No | UUID; leave unset for yourself. | |
| arguments | No | Arguments for the target capability; inspect its schema with describe_tool before invoking it. | |
| tool_name | Yes | Exact public capability name returned by list_tools. | |
| continuity_token | No | Same-process rebind proof only; never a cross-process resume. | |
| client_session_id | No | Binding id for calls in this process; not a cross-process proof. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, openWorldHint=true, and non-idempotent, so the risk profile is covered. The description adds meaningful context beyond that: the target's own identity, validation, authorization, timeout and response middleware all run, so this gateway is not a privilege escalation path. It stops short of describing error or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the action and the prerequisite chain, then the safety clarification and the recursion exclusion. The middleware enumeration is slightly long but each clause carries a distinct behavioral fact worth keeping.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description should orient the agent on return expectations, which it only gestures at implicitly. For a five-parameter gateway with nested arguments, it otherwise covers prerequisites, delegation semantics, and the key failure mode (recursive refusal) adequately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with all five parameters documented, so the baseline is 3. The description only reinforces that arguments must be built from describe_tool output and that tool_name must be the exact name from list_tools, adding marginal value over the schema's own field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: invoke one public capability discovered via list_tools. It also carves out exactly what it is not ('a discovery gateway, not an authorization bypass'), which distinguishes it from siblings like list_tools, describe_tool, and the session tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit sequence: find the name with list_tools, inspect arguments with describe_tool, then pass the argument object here. It also states an exclusion — it refuses recursive use_tool calls — so the agent knows both when to use it and one case where it will not work.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v3.2.0- Changed
search_shared_memory3 fields changed- changed
Input schema / properties / agent_id / descriptionPrevious value: -"Filter by agent (for action=get, search; omit when using discovery_id readback)"New value: +"Filter by author agent; agent_id_filter wins if both are set." - changed
Input schema / properties / query / descriptionPrevious value: -"Search query (for action=search)"New value: +"Search text." - changed
Input schema / properties / status / descriptionPrevious value: -"Status filter/update value (open, resolved, archived, superseded)"New value: +"Filter by status: open, resolved, archived, superseded."
16 tool updates
v3.1.0- Changed
check_working_state5 fields changed- changed
Input schema / properties / client_session_id / descriptionPrevious value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof." - changed
Input schema / properties / continuity_token / descriptionPrevious value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume." - changed
Input schema / properties / include_state / descriptionPrevious value: -"Include nested state dict in response (can be large). Default false to reduce context bloat. Accepts boolean or string ('true'/'false')."New value: +"Retained for compatibility; has no effect on any tier. Use verbosity to choose what is returned. Accepts boolean or string ('true'/'false')." - changed
Input schema / properties / lite / descriptionPrevious value: -"If true (default), returns minimal essential metrics only. Set lite=false for full diagnostic data."New value: +"If true (default), returns minimal essential metrics only. Set lite=false for full diagnostic data. verbosity, when given, takes precedence." - added
Input schema / properties / verbosityAdded value: +{ + "anyOf": [ + { + "enum": [ + "minimal", + "standard", + "full" + ], + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Tier: minimal (default); standard: bare EISV and risk, verdict/basin/mode with meanings, guidance; full: diagnostics." +}
- Changed
consult3 fields changed- changed
Input schema / properties / agent_id / descriptionPrevious value: -"UNIQUE agent identifier; optional when session-bound (auto-injected)."New value: +"UUID; leave unset for yourself." - changed
Input schema / properties / client_session_id / descriptionPrevious value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof." - changed
Input schema / properties / continuity_token / descriptionPrevious value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume."
- Changed
describe_tool3 fields changed- changed
Input schema / properties / agent_id / descriptionPrevious value: -"UNIQUE agent identifier; optional when session-bound (auto-injected)."New value: +"UUID; leave unset for yourself." - changed
Input schema / properties / client_session_id / descriptionPrevious value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof." - changed
Input schema / properties / continuity_token / descriptionPrevious value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume."
- Removed
health_check - Changed
identity4 fields changed- changed
Input schema / descriptionPrevious value: -"Who am I? Auto-creates identity if first call."New value: +"Resolve this session's bound agent (pass client_session_id), or set a cosmetic display name." - changed
Input schema / properties / agent_id / descriptionPrevious value: -"UNIQUE agent identifier; optional when session-bound (auto-injected)."New value: +"UUID; leave unset for yourself." - changed
Input schema / properties / client_session_id / descriptionPrevious value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof." - changed
Input schema / properties / continuity_token / descriptionPrevious value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume."
- Removed
knowledge - Added
list_tools - Changed
record_result3 fields changed- changed
Input schema / descriptionPrevious value: -"Parameters for outcome_event"New value: +"Outcome record parameters." - changed
Input schema / properties / client_session_id / descriptionPrevious value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof." - changed
Input schema / properties / continuity_token / descriptionPrevious value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume."
- Changed
request_review6 fields changed- changed
Input schema / descriptionPrevious value: -"Parameters for dialectic"New value: +"Review session parameters." - changed
Input schema / properties / client_session_id / descriptionPrevious value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof." - changed
Input schema / properties / continuity_token / descriptionPrevious value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume." - changed
Input schema / properties / issue_description / descriptionPrevious value: -"Issue description (for action=request)"New value: +"Issue description (action=request/quick)" - changed
Input schema / properties / proposed_conditions / descriptionPrevious value: -"Conditions for resumption (for action=thesis/synthesis)"New value: +"Conditions for resumption (for action=thesis/synthesis/consult)" - changed
Input schema / properties / root_cause / descriptionPrevious value: -"Root cause analysis (for action=thesis/synthesis)"New value: +"Root cause analysis (for action=thesis/synthesis/consult)"
- Changed
search_shared_memory37 fields changed- changed
Input schema / descriptionPrevious value: -"Parameters for knowledge"New value: +"Knowledge graph parameters." - added
Input schema / properties / agent_id_filterAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Author agent UUID filter; wins over agent_id." +} - added
Input schema / properties / authority_modeAdded value: +{ + "anyOf": [ + { + "enum": [ + "prefer_governed", + "all" + ], + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "prefer_governed (default) down-ranks imported memory; all keeps raw order." +} - changed
Input schema / properties / client_session_id / descriptionPrevious value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof." - removed
Input schema / properties / closure_classRemoved value: -{ - "anyOf": [ - { - "type": "string" - }, - { - "type": "null" - } - ], - "default": null, - "description": "Closing standard for action=update: fix_verified | unobserved | not_reproducible | obsolete | duplicate. 'fix_verified' needs a deployed change whose effect was observed." -} - removed
Input schema / properties / closure_evidenceRemoved value: -{ - "anyOf": [ - { - "additionalProperties": true, - "type": "object" - }, - { - "type": "null" - } - ], - "default": null, - "description": "Evidence for closure_class. Required keys: fix_verified needs {deployed, observed}; unobserved needs {window, instrument_check}." -} - removed
Input schema / properties / confidenceRemoved value: -{ - "anyOf": [ - { - "type": "number" - }, - { - "type": "string" - }, - { - "type": "null" - } - ], - "default": null, - "description": "Writer-supplied confidence for action=store, validated to 0-1; not independently verified or adjusted by governance metrics" -} - changed
Input schema / properties / continuity_token / descriptionPrevious value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume." - added
Input schema / properties / created_afterAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "search: created after ISO time." +} - added
Input schema / properties / created_beforeAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "search: created before this ISO time." +} - changed
Input schema / properties / discovery_type / descriptionPrevious value: -"action=store; one of architectural_decision, learning, pattern, bug_fix, refactoring, documentation, experiment, question, note, rule, insight, bug_found, bug, improvement, exploration, observation."New value: +"Filter by discovery type, e.g. bug_found." - removed
Input schema / properties / dry_runRemoved value: -{ - "anyOf": [ - { - "type": "boolean" - }, - { - "type": "string" - }, - { - "type": "null" - } - ], - "default": null, - "description": "Dry run mode (for action=cleanup, synthesize)" -} - removed
Input schema / properties / epoch_scopeRemoved value: -{ - "anyOf": [ - { - "enum": [ - "current", - "all" - ], - "type": "string" - }, - { - "type": "null" - } - ], - "default": null, - "description": "Stats/list scope: current epoch only or all epochs" -} - changed
Input schema / properties / include_archived / descriptionPrevious value: -"Include archived discoveries in search results (default: excluded)"New value: +"Include archived rows." - changed
Input schema / properties / include_cold / descriptionPrevious value: -"Include cold-storage (long-term) discoveries in search results (default: excluded)"New value: +"Include cold-storage rows." - changed
Input schema / properties / include_details / descriptionPrevious value: -"Expand results inline only with response_mode='full'; otherwise open one with knowledge(action='details')."New value: +"Inline details need response_mode='full'; else open one with knowledge(action='details')." - changed
Input schema / properties / include_provenance / descriptionPrevious value: -"Include provenance and lineage chain fields in search/details results"New value: +"Add provenance fields." - removed
Input schema / properties / include_response_chainRemoved value: -{ - "anyOf": [ - { - "type": "boolean" - }, - { - "type": "string" - }, - { - "type": "null" - } - ], - "default": null, - "description": "Include typed response chain for action=details" -} - removed
Input schema / properties / including_coldRemoved value: -{ - "anyOf": [ - { - "type": "boolean" - }, - { - "type": "string" - }, - { - "type": "null" - } - ], - "default": null, - "description": "Include cold-storage discoveries in action=list raw status aggregates" -} - removed
Input schema / properties / lengthRemoved value: -{ - "anyOf": [ - { - "type": "integer" - }, - { - "type": "null" - } - ], - "default": null, - "description": "Maximum details characters returned for action=details" -} - changed
Input schema / properties / limit / descriptionPrevious value: -"Max results (for action=search: min 1, values above 100 are capped, 0 or negative is rejected)"New value: +"Max results, 1-100 (more is capped)." - removed
Input schema / properties / max_chain_depthRemoved value: -{ - "anyOf": [ - { - "type": "integer" - }, - { - "type": "null" - } - ], - "default": null, - "description": "Maximum response-chain traversal depth for action=details" -} - removed
Input schema / properties / memory_contextRemoved value: -{ - "anyOf": [ - { - "type": "string" - }, - { - "type": "null" - } - ], - "default": null, - "description": "S22 provenance: memory/KG/transcript surfaces visible to the writer" -} - removed
Input schema / properties / min_membersRemoved value: -{ - "anyOf": [ - { - "type": "integer" - }, - { - "type": "null" - } - ], - "default": null, - "description": "Minimum discoveries a topic needs before it is rolled up (for action=synthesize, default 3)" -} - changed
Input schema / properties / min_similarity / descriptionPrevious value: -"Minimum cosine similarity for semantic retrieval modes"New value: +"Semantic similarity floor." - removed
Input schema / properties / offsetRemoved value: -{ - "anyOf": [ - { - "type": "integer" - }, - { - "type": "null" - } - ], - "default": null, - "description": "Character offset for action=details pagination" -} - added
Input schema / properties / recency_half_life_daysAdded value: +{ + "anyOf": [ + { + "exclusiveMinimum": 0, + "maximum": 36500, + "type": "number" + }, + { + "type": "null" + } + ], + "default": null, + "description": "search: score halves per N days." +} - removed
Input schema / properties / scopeRemoved value: -{ - "anyOf": [ - { - "enum": [ - "open", - "all", - "by_agent" - ], - "type": "string" - }, - { - "type": "null" - } - ], - "default": null, - "description": "KG audit scope for action=audit" -} - changed
Input schema / properties / search_mode / descriptionPrevious value: -"Force retrieval mode for action=search. 'semantic' and 'hybrid' fail honestly when unsupported by the active backend."New value: +"auto, or force fts/semantic/hybrid." - changed
Input schema / properties / semantic / descriptionPrevious value: -"Legacy action=search toggle to force or skip semantic retrieval when supported"New value: +"Legacy semantic on/off toggle." - changed
Input schema / properties / severity / descriptionPrevious value: -"Severity: low, medium, high, critical (for action=store or action=update)"New value: +"Filter by severity: low, medium, high, critical." - added
Input schema / properties / sort_byAdded value: +{ + "anyOf": [ + { + "enum": [ + "relevance", + "created_at" + ], + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "search order: relevance, or created_at (newest matches first)." +} - changed
Input schema / properties / tags / descriptionPrevious value: -"Tags (action=store, note); for action=search an exact any-of filter in every search mode."New value: +"Exact any-of tag filter." - removed
Input schema / properties / top_nRemoved value: -{ - "anyOf": [ - { - "type": "integer" - }, - { - "type": "null" - } - ], - "default": null, - "description": "Maximum stale entries returned by action=audit" -} - removed
Input schema / properties / topicRemoved value: -{ - "anyOf": [ - { - "type": "string" - }, - { - "type": "null" - } - ], - "default": null, - "description": "Synthesize just this one tag/topic (for action=synthesize). Omit to sweep the densest topics." -} - removed
Input schema / properties / use_llmRemoved value: -{ - "anyOf": [ - { - "type": "boolean" - }, - { - "type": "string" - }, - { - "type": "null" - } - ], - "default": null, - "description": "Use the local LLM for the rollup narrative (for action=synthesize, default true; falls back to deterministic when unreachable)" -} - removed
Input schema / properties / use_modelRemoved value: -{ - "anyOf": [ - { - "type": "boolean" - }, - { - "type": "string" - }, - { - "type": "null" - } - ], - "default": null, - "description": "Use the local model to assess stale entries for action=audit" -}
- Changed
self_recovery3 fields changed- changed
Input schema / properties / agent_id / descriptionPrevious value: -"UNIQUE agent identifier; optional when session-bound (auto-injected)."New value: +"UUID; leave unset for yourself." - changed
Input schema / properties / client_session_id / descriptionPrevious value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof." - changed
Input schema / properties / continuity_token / descriptionPrevious value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume."
- Changed
start_session3 fields changed- changed
Input schema / properties / agent_id / descriptionPrevious value: -"UNIQUE agent identifier; optional when session-bound (auto-injected)."New value: +"UUID; leave unset for yourself." - changed
Input schema / properties / client_session_id / descriptionPrevious value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof." - changed
Input schema / properties / continuity_token / descriptionPrevious value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume."
- Changed
store_finding7 fields changed- changed
Input schema / descriptionPrevious value: -"Parameters for knowledge"New value: +"Knowledge graph parameters." - changed
Input schema / properties / agent_id / descriptionPrevious value: -"Filter by agent (for action=get, search; omit when using discovery_id readback)"New value: +"Leave unset; the bound session is the writer." - changed
Input schema / properties / client_session_id / descriptionPrevious value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof." - changed
Input schema / properties / content / descriptionPrevious value: -"Extended content/details (for action=store or action=note)"New value: +"Extended content/details (for action=store, note, or promote)" - changed
Input schema / properties / continuity_token / descriptionPrevious value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume." - changed
Input schema / properties / details / descriptionPrevious value: -"Extended details for discovery (for action=store). Alias: content"New value: +"Extended details for discovery (for action=store or promote). Alias: content" - changed
Input schema / properties / summary / descriptionPrevious value: -"Discovery summary (for action=store)"New value: +"Discovery summary (for action=store or promote)"
- Changed
sync_state6 fields changed- changed
Input schema / descriptionPrevious value: -"Share your work and get supportive feedback. Your main tool for checking in."New value: +"Record a work check-in and get a governance decision." - changed
Input schema / properties / auto_export_on_significance / descriptionPrevious value: -"If true, automatically export governance history when thermodynamically significant events occur."New value: +"If true, write a governance history export whenever this check-in is judged significant: a risk spike, a coherence drop, a void threshold …" - changed
Input schema / properties / client_session_id / descriptionPrevious value: -"In-session binding id from start_session()/identity(); pass it on same-process calls. Not a cross-process proof."New value: +"Binding id for calls in this process; not a cross-process proof." - changed
Input schema / properties / complexity / anyOfPrevious value: -[ - { - "anyOf": [ - { - "type": "number" - }, - { - "type": "string" - } - ], - "ge": 0, - "le": 1 - }, - { - "type": "null" - } -]New value: +[ + { + "maximum": 1, + "minimum": 0, + "type": "number" + }, + { + "anyOf": [ + { + "pattern": "^(0(\\.\\d+)?|1(\\.0+)?|\\.\\d+)$", + "type": "string" + }, + { + "enum": [ + "complex", + "critical", + "high", + "low", + "medium", + "minimal", + "moderate", + "simple", + "trivial", + "very_high" + ], + "type": "string" + } + ], + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / properties / confidence / anyOfPrevious value: -[ - { - "anyOf": [ - { - "type": "number" - }, - { - "type": "string" - } - ], - "ge": 0, - "le": 1 - }, - { - "type": "null" - } -]New value: +[ + { + "maximum": 1, + "minimum": 0, + "type": "number" + }, + { + "pattern": "^(0(\\.\\d+)?|1(\\.0+)?|\\.\\d+)$", + "type": "string" + }, + { + "type": "null" + } +] - changed
Input schema / properties / continuity_token / descriptionPrevious value: -"Ownership proof from onboard()/identity(), for same-live-process rebinds only. Not a cross-process resume credential."New value: +"Same-process rebind proof only; never a cross-process resume."
- Removed
update_finding - Added
use_tool
14 tool updates
- First observed
check_working_state - First observed
consult - First observed
describe_tool - First observed
health_check - First observed
identity - First observed
knowledge - First observed
record_result - First observed
request_review - First observed
search_shared_memory - First observed
self_recovery - First observed
start_session - First observed
store_finding - First observed
sync_state - First observed
update_finding
TDQS
Scored across 13 tools
Several tools overlap by design: sync_state and process_agent_update are 'canonical twins', start_session and onboard likewise, and check_working_state vs sync_state vs identity all touch governance state reads. The dense descriptions do disambiguate the intended path, but the twin/alias pattern and the list_tools/describe_tool/use_tool discovery layer duplicating the advertised surface create real misselection risk.
Names are consistently snake_case, and most follow a verb_noun pattern (sync_state, check_working_state, search_shared_memory, store_finding, start_session, record_result, request_review, list_tools, describe_tool, use_tool). A few break the pattern (consult, identity, self_recovery) but remain readable and unambiguous in style.
13 advertised tools is a well-scoped number for a governance server covering identity, state, memory, review, consultation, and recovery. However, the set is effectively larger since many capabilities (process_agent_update, onboard, outcome_event, update_finding, operator_resume_agent) are hidden behind use_tool, so the true surface is heavier than the count suggests.
The surface covers the lifecycle well: identity/session setup, state check-in and read, outcome recording, knowledge graph read/write, review, advisory consult, and self-recovery, plus a gateway to unlisted capabilities. Minor gaps remain (e.g., no explicit delete/archive for findings, revisions routed through the use_tool gateway), but core workflows are covered.
Maintenance
Related MCP Connectors
Pre-execution governance for AI agents. Deterministic PASS/FAIL/REVIEW verdicts, replayable proof.
Connect, monitor, and control AI agents — tasks, approvals, schedules, and governance.
Deterministic runtime safety for AI agents: scan PII, gate tool actions, verify LLM output.
Runtime AI governance: decision gates, human approval, hash-chained audit, compliance mapping.
Related MCP Servers
- FlicenseAqualityFmaintenanceProvides policy-based access control, incident tracking, and compliance monitoring to govern AI agent behavior. It enables organizations to enforce security rules and maintain audit trails by validating agent actions against trust levels and pattern-based policies.61-
- AlicenseAqualityCmaintenanceRuntime policy enforcement for AI agents. Evaluate every agent action against your organization's policies before execution, with observe and enforce modes.11MIT

@vorionsys/mcp-serverofficial
AlicenseAqualityBmaintenanceMCP server for AI-agent governance using trust scoring, behavioral signals, and pre-flight action checks.1017 npm1Apache 2.0
Rigour MCPofficial
AlicenseNot gradedqualityBmaintenanceEnables AI agents to self-govern by scanning code for hardcoded secrets, structural violations, and AI drift in real-time, providing fix packets for automatic remediation.27MIT