UNITARES
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@UNITARESshow fleet health summary"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Runtime state, accountability, and recovery for long-lived AI-agent fleets.
Status: v2.18.0. Sustained operation is documented in one long-running maintainer deployment. External adoption remains unvalidated, and the current outcome-lift evaluation found no result beyond a selection-aware null.
An agent forty turns into a task reports high confidence while its tests are failing. Every individual tool call was allowed, so an action-level guardrail may have nothing to object to. What is missing is a longitudinal record that compares what the agent claims with what actually happened.
UNITARES keeps that record. Agents check in after meaningful units of work. The server binds writes to a process identity, stores claims and outcomes, derives a four-score state estimate, and returns a policy decision with a named reason. It is a self-hosted MCP/HTTP service, not an agent framework or hosted platform.
Maintainer dogfood since November 2025 · 4.5M+ recorded audit/telemetry events.
That count is evidence of sustained operation in one maintainer-run environment, not external adoption or governance efficacy.
What it does
Core surface | What the operator gets |
Accountable identity | Bind writes to a process instance and retain what it did, claimed, and observed. Reads can remain open; writes are attributable. |
Evidence-linked calibration | Compare stated confidence with tests, exit codes, tool results, review labels, and other recorded outcomes. |
Policy and recovery | Return a named action, reason, and next step; enforce pauses on governed write surfaces; support a |
Operator visibility | Inspect lifecycle, health, state, evidence, and policy history through MCP/HTTP APIs and a self-hosted dashboard. |
The core loop is deliberately small. Optional modules add a provenance-aware knowledge graph, structured review, reference resident agents, and Elixir/OTP coordination for leases, handoffs, dispatch, and supervision. They can be used independently of the basic check-in loop.
The public unitares-sdk handles connection, identity,
check-ins, heartbeats, and knowledge participation for resident agents. Its
README carries the current install command and server compatibility guidance.
Related MCP server: promptspeak-mcp-server
Quickstart
git clone --branch v2.18.0 --depth 1 https://github.com/cirwel/unitares.git
cd unitares
docker compose up -d --wait
make demomake demo onboards a fresh process and sends six check-ins over the real API.
It prints the response shape, decision reason, state detail, and warmup position.
The demo answers “is my stack wired?” It does not establish predictive value.
The dashboard is at http://localhost:8767/dashboard; MCP clients connect to
http://localhost:8767/mcp/.
Use the surface that matches the question:
Does the signal beat a simple baseline? Run the falsifiability harness.
What does sustained operation look like? Read the maintainer-deployment snapshot and its data caveat.
How do I operate or integrate it? Follow the user manual.
Integrate an MCP client
start_session and sync_state are tools exposed by the connected UNITARES
server. A fresh process creates its own identity, then includes the returned
session binding on later writes:
session = start_session(force_new=True)
result = sync_state(
response_text=output,
complexity=0.6,
confidence=0.8,
client_session_id=session["client_session_id"],
)
action = result.get("state_summary", {}).get("action")
if action in ("pause", "reject"):
return_to_operator(result.get("next_action")) # application-defined boundaryFor a durable resident, preserve its identity anchor rather than minting a new identity on every run; the SDK lifecycle example handles that continuity. Pair self-reported confidence with verifiable evidence whenever possible:
Need | Tool |
Search shared memory before writing |
|
Record a test, task, or external outcome |
|
Request structured review |
|
Read state without writing |
|
list_tools() enumerates the complete live surface and describe_tool(name)
explains any one tool. MCP, REST, the SDK, and host adapters use the same server;
Claude Code and Codex are supported clients, not server-side assumptions.
How the runtime loop works
At each checkpoint, the agent reports a meaningful unit of work and its stated confidence. The server resolves the process identity, associates available outcomes, updates longitudinal state, and returns a policy action with a named reason. The retained record lets an operator audit both the claim and the basis for the response.
Clients can treat the policy action, reason, and next step as the stable contract. Operators can optionally inspect four EISV coordinates covering work progress, evidence alignment, behavioral drift, and their balance. These are auditable heuristics, not literal thermodynamic quantities or universal labels. The computation reference documents formulas, warmup, thresholds, and source code; the interpretation contract records permitted readings, refuted claims, and open evaluation gaps.
Where it fits
UNITARES runs alongside evals, guardrails, and sandboxes; it replaces none of them.
Layer | Question | Timing |
Evals | Is this model good enough for a defined task? | Before or between deployments. |
Guardrails / sandbox | Is this action allowed and contained? | Per action. |
UNITARES | What has this running process been doing, what evidence supports its claims, and what state is it in now? | Continuously, mid-run. |
It is designed for long-lived coding, research, operations, monitoring, and multi-agent processes that can instrument a check-in loop. It is usually not worth the overhead for short-lived chat turns.
It is not an outcome oracle: it does not decide whether an output is correct or ethical, and it cannot detect deliberate concealment without independent evidence. The information-theoretic and ODE formulation in the companion paper remains a research target and parallel diagnostic path, not the deployed policy mechanism.
Local control and future federation
UNITARES starts as a self-hosted governor under one operator's control. The architecture exposes the seams that later federation experiments would need: process-bound identity, evidence provenance, a versioned telemetry envelope, and policy decisions with named reasons.
Future experiments can test whether those records are sufficient for exchanging cross-operator attestations without centralizing raw telemetry. Cross-governor trust, consensus, and enforcement are not deployed guarantees. Experiments between mutually distrustful governors remain gated on independent-operator validation in the roadmap.
Evidence and limits
At the 2026-08-11 frozen snapshot, the maintainer deployment provided operational evidence, not an independent efficacy study:
Evidence | Scope |
4,573,890 audit/telemetry events | Continuous maintainer-run operation since 2025-11-28. Session-resolution observations and cross-device-call records make up 91.4%; this is infrastructure/load evidence, not 4.6M independent policy decisions. |
71,141 stored EISV state rows | Longitudinal state observations in |
15 recorded self-recovery events | Of 21 canonical, non-automatic lifecycle-resume records. Shows the path was exercised; not 15 independent trials or proof that pauses improved outcomes. |
32,181 labeled EISV windows | 20,655 overlapping real windows from one 39-day Raspberry Pi run plus 11,526 synthetic windows. These are windows, not independent agents or customer trajectories. |
The maintainer deployment is single-operator and co-development dogfood: most
agents governed by the system are also building the system. Read
DEPLOYMENT_DATA_CAVEAT.md before
citing a fleet number.
The frozen 2026-08-09 outcome-lift evaluation is a negative result: after model selection, no overall slice separated from the permutation null (selective p = 0.070–0.567). Some unadjusted metrics improved, but none cleared the selection-aware threshold. There is no demonstrated prevention. The Reviewer Guide gives the frozen command and interpretation; compact output is preserved in the dated ablation snapshot.
Robustness against a motivated attacker optimizing the monitored proxy remains unproven. Calibrated capability concealment is a documented structural blind spot. See the scope and threat model.
The companion DOI identifies a public preprint, not peer-reviewed validation.
Architecture, setup, and documentation
Python 3.12+ · PostgreSQL + AGE + pgvector · Redis · optional Elixir/OTP coordination. Redis is the de-facto session store in the long-running maintainer deployment; without it the server starts in degraded local-only mode suitable for the demo.
Reader | Start here |
Evaluator or grant reviewer | |
Integrator | |
Operator | |
Contributor | |
Research/provenance reader |
The complete documentation map is docs/README.md. Optional
analogies and philosophical readings are isolated under
docs/essays/; they are not specifications or evidence.
Project operation is explicit: see the roadmap, compatibility and naming map, governance, support policy, and release process.
Ecosystem repositories
These are adjacent integrations, testbeds, and research projects; the core quickstart does not require them.
Project | Role |
Raspberry Pi longitudinal testbed. | |
Codex and Claude Code lifecycle/hook packaging. | |
Thin bindings for additional clients and model hosts. | |
Governed-effect runtime research seed. | |
Dataset generation and labeling pipeline. | |
Companion preprint and research formulation. |
Citation and license
Kenny Wang (ORCID 0009-0006-7544-2374),
CIRWEL Systems. See CITATION.cff for the versioned citation.
@misc{wang2026unitares,
author = {Wang, Kenny},
title = {{UNITARES}: Information-Theoretic Governance of Heterogeneous Agent Fleets},
year = {2026},
doi = {10.5281/zenodo.19647159}
}This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- FlicenseAqualityFmaintenanceProvides policy-based access control, incident tracking, and compliance monitoring to govern AI agent behavior. It enables organizations to enforce security rules and maintain audit trails by validating agent actions against trust levels and pattern-based policies.61
- AlicenseBqualityCmaintenancePre-execution governance for AI agents. 45 MCP tools for hold queues, audit trails, risk scoring, and policy enforcement. Validates agent actions before they execute.451191MIT
- AlicenseAqualityDmaintenanceRuntime policy enforcement for AI agents. Evaluate every agent action against your organization's policies before execution, with observe and enforce modes.11MIT

@vorionsys/mcp-serverofficial
AlicenseAqualityBmaintenanceMCP server for AI-agent governance using trust scoring, behavioral signals, and pre-flight action checks.10381Apache 2.0
Related MCP Connectors
Runtime AI governance: decision gates, human approval, hash-chained audit, compliance mapping.
See, price, and control every tool call your AI agents make: policy checks, cost, and audit tools.
Sovereign Agent OS — Persistent Memory, Governance & Compliance for AI Agents.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/cirwel/unitares'
If you have feedback or need assistance with the MCP directory API, please join our Discord server