Skip to main content
Glama
ariffazil

arifOS MCP Server

by ariffazil

arifOS — The Authority Plane of the arifOS Federation

Security Audit: In Progress External evaluation: not yet published Kernel: healthy MCP scanners: 148 rules, 0 unauthorised mutations Sovereign Boundary: F13 MCP Conformance Vault Integrity Floor Gate Governance Runtime Drift External Witness

arifOS evaluates consequential AI actions against constitutional floors and returns an independent verdict before execution occurs.

When an AI agent proposes to write, delete, deploy, or spend, arifOS inserts a constitutional judgment step: the agent proposes, arifOS evaluates the proposal against the F1–F13 floors, a verdict is returned, and only then does execution proceed. Every consequential verdict is written into an append-only hash chain with its decision context, evidence references, provenance and receipt hashes.

In a world where intelligence is abundant, authority becomes the scarce resource. arifOS exists to keep judgment independent from execution.

This is not an AI model. It is not an agent framework. It is a constitutional authority system — the layer between "agent wants to act" and "action is permitted."

Two invariants hold everywhere in this document:

Proposer ≠ Judge ≠ Executor ≠ Witness Human sovereignty > machine authority

What arifOS is not: not an AI model · not an agent framework · not an execution engine (A-FORGE executes) · not an attention plane (AAA attends) · not the independent witness (FRAME witnesses; VAULT999 remembers) · not a substitute for authentication, sandboxing, or legal review.

Audience

What you get

Human

A quiet veto: the agent proposes, the kernel records a verdict, you stay sovereign

Agent / A2A

MCP tools + receipts. You do not get the keys. A2A v1.0 (a2a_version: 1.0.1 in the agent card), not v1.2 — discovery is owned by AAA, judgment by arifOS, execution by A-FORGE

Institution

Policy floors F1–F13, a hash-chained VAULT999 audit trail, model-vendor independence

Indexer / Search

Structured metadata, llms.txt, tools.json, CITATION.cff

Machine / MCP

Streamable HTTP endpoint, 8 canonical verbs, published input/output schemas

Robot / Automation

CI gates (conformance, vault-integrity, drift, governance), Makefile targets, Dockerfile

Live: https://arifos.arif-fazil.com · MCP :8088 · sister organs GEOX · A-FORGE · AAA


The Problem

AI agents that act are also certifying their own actions. Nothing independent evaluates the proposal against safety, compliance and policy constraints before execution occurs.

Related MCP server: acf-mcp

The Solution

  Agent proposes action
         │
         ▼
  ┌─────────────────┐
  │   arifOS Kernel  │  Evaluates against 13 constitutional floors
  │     (:8088)      │  Records provenance + evidence refs + hashes
  └────────┬────────┘
           │
    ┌──────┼────────┬──────────┬───────────┐
    ▼      ▼        ▼          ▼           ▼
  SEAL    HOLD    SABAR     PARTIAL      VOID
  (go)  (wait for (defer —  (proceed   (blocked by
        human)    evidence   with       a floor)
                  pending)   cooling)
    │      │
    ▼      ▼
  Execute  Human
  via      reviews
  A-FORGE
    │
    ▼
  Receipt in VAULT999
  (append-only hash chain)

The judge never executes. The executor never certifies.

Federation in one line

Intelligence proposes. Authority constrains. Execution acts. Reality witnesses. History remembers.

Plane / class

Component

Role

Authority

arifOS :8088

Constitutional judgment — evaluates proposals against F1–F13; owns the gavel and the ledger

Attention

AAA :3001

Attention and routing — what matters, and to whom. Displays, routes, queues; never adjudicates

Execution

A-FORGE :7071/:7072

Governed mutation — leases, gates, receipts; only under a SEAL verdict

Witness

FRAME :18085

Independent observer. Its output is evidence, never a verdict

Record

VAULT999 (in-kernel)

Immutable append-only ledger — the memory of what was decided

Metabolism

arifFlow :7073

Receipt ingestion, FQ monitoring, verification cadence; never adjudicates

Domain intelligence

GEOX :8081 · WEALTH :18082 · WELL :18083

Evidence and computation in a domain. Compute-only: no organ authorises its own action

Provider gateway

FED :7074

Multi-provider model routing (advisory-only)

Synthesis

i-ARIF

Seal-B synthesis engine — runs through FED chains, owns no port

Boundary services (Tier-3)

HERMES :18087 · CHRON :18102

Semantic boundary (meaning integrity, relay-only) and temporal boundary (episodes, predictions, calibration). Neither is an organ — see the ruling in FEDERATION_CONTRACT.md §2.1

Authority remains separated at every stage: no component proposes, judges, executes and witnesses the same action. arifOS determines whether and how a routing may proceed; AAA determines what matters.


Quick Start

Requires Python 3.12+ (supported range 3.12–3.14; see pyproject.toml).

Versioning: two schemes coexist. The kernel release (v2026.08.01, reported by /health as release_name) is the operational identity of the running service. The PyPI package uses epoch versioning (1!…) to outrank legacy releases: 1!2026.9.2 is what pip install arifos resolves to today, while this tree is 1!2026.10.1 (staged, not yet released). The kernel release is the operational truth; PyPI is the distribution truth.

Install

pip install arifos
pip show arifos        # → Version: 1!2026.9.2 (PyPI published; repo tree 1!2026.10.1 — see gaps table)

Install (Docker)

docker build -t arifos .
docker run -p 3000:3000 --env-file .env.docker.example arifos
# The image listens on 3000 (Dockerfile EXPOSE/CMD). Its OCI labels still say 8088 and
# carry a legacy licence value — see "What Is Not Yet Proven".

Run the kernel

# HTTP transport (this is the MCP endpoint) — port comes from $PORT, default 8080
# Ports: $PORT defaults to 8080; the deployed unit runs on 8088; the Docker image listens on 3000. Use 8088 for anything in this README.
PORT=8088 arifos-mcp streamable-http

# stdio transport, for a local stdio MCP client
arifos-mcp

# equivalent module form
PORT=8088 python -m arifosmcp.runtime --mode streamable-http

# health — the deployed unit runs on 8088
curl -s http://localhost:8088/health

Verified on 2026-09-21 against the live unit:

{ "status": "healthy", "release_name": "v2026.08.01", "mcp_protocol_version": "2026-07-28",
  "tools_loaded": 8, "operational_tools": 3, "deployment_drift_status": "aligned",
  "floors_active": 13, "vault999_health": "healthy" }

Run a governed workflow locally (no client needed)

examples/enterprise_operations_demo.py runs six graded scenarios in-process and prints each verdict, its floor gate and its receipt — including two that are refused:

python examples/enterprise_operations_demo.py

Demo output renders human-readable display labels over the canonical seven seals. The wire vocabulary is always the seven seals (SEAL, HOLD, SABAR, PARTIAL, PROVISIONAL, HOLD_888, VOID); the labels below are presentation only and are mapped deterministically to a canonical verdict:

Display label

Canonical verdict

✔ ALLOW

SEAL (proceed)

⏸ HOLD

HOLD or SABAR (await human / await evidence)

✘ BLOCK/VOID

VOID (blocked by a hard floor)

[ DEMO / SIMULATED WORKFLOW ]  No real customer funds, records, or firewall policies are altered.
Kernel Session ID : SEAL-DEMO-…        Authority Ceiling : LIMITED_MUTATE
SCENARIO 1: Read Customer Account Data      → ✔ ALLOW   RCPT-4245A2002143490B
SCENARIO 3: Major Enterprise Refund (RM5,000) → ⏸ HOLD  (RM100 autonomous ceiling, over by 50×)
SCENARIO 4: Delete Customer Account & Audit Trail → ✘ BLOCK/VOID

Connect an MCP client

http://localhost:8088/mcp      # or https://mcp.arif-fazil.com/mcp

Protocol facts, as measured: the kernel advertises 2026-07-28 and accepts 2026-07-28 · 2025-11-25 · 2025-03-26 · 2024-11-05. A live initialize currently settles on 2025-11-25 — the declared canonical spec in arifosmcp/runtime/public_surface.py and the version the internal conformance runner records. If you pin a protocol version, use 2025-11-25.

curl -s http://localhost:8088/mcp \
  -H 'content-type: application/json' -H 'accept: application/json, text/event-stream' \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}'
# → 8 tools: arif_init arif_observe arif_think arif_route arif_memory arif_judge arif_forge arif_seal

Session flow (agents start here)

The kernel binds a session before any judgment. Verified transcript, 2026-09-21:

// Step 1 — arif_init: establish actor identity and authority band
{ "name": "arif_init", "arguments": { "mode": "init", "actor_id": "my-agent" } }
// → { "verdict": "HOLD", "session_id": "SEAL-df114686f4c34971",
//     "autonomy_band": "OBSERVE_ONLY", "trace_id": "trc-197d2fa35887", "session_token": "act_v1.…" }
// HOLD here is the normal starting state: the session is open, no mutation is authorised yet.

// Step 2 — arif_judge: evaluate a proposal, return a verdict + evidence
{ "name": "arif_judge", "arguments": {
    "candidate": "Delete the production audit table", "action_tier": "standard",
    "session_id": "SEAL-df114686f4c34971", "actor_id": "my-agent" } }
// → { "verdict": "HOLD", "trace_id": "trc-d5d4d2fed6f8",
//     "constitutional_check": { "hold_required": true, "failed_floors": [] },
//     "next_safe_action": "provide actor_signature / sovereign_receipt / heart_critique, or reduce blast radius" }

// Step 3 — arif_seal: only after a SEAL verdict; appends the receipt to VAULT999
{ "name": "arif_seal", "arguments": { "payload": "…", "session_id": "…", "ack_irreversible": true } }

Call arif_judge without a session and the kernel answers actor: anonymous, authority: OBSERVE_ONLY, verdict: HOLD — it will not silently upgrade an unattested caller.

Stage numbering (single source of truth). The verbs follow the ratified eight-stage map: 000 INIT · 111 OBSERVE · 333 THINK · 444 ROUTE · 555 MEMORY · 666 JUDGE · 777 FORGE · 999 SEAL. Source: arifosmcp/constitutional_map.py ToolStage (F13-ratified 2026-07-31: JUDGE = 666, FORGE = 777; the old 888 stage was retired when compose was absorbed into forge). Some mirrors still carry the retired 888 label — a description string in constitutional_map.py, tools_sot.yaml, the generated llms.txt, and docs/PROMPT_666_JUDGE_DEPRECATION.md (dated 2026-07-10, predating the correction). The live wire says 666, and this README follows the live wire. Flagged for repair in "What Is Not Yet Proven".

For a guided walkthrough, start with docs/START_HERE.md or the verified docs/QUICKSTART.md.


Core Concepts

The verdict lattice

Verdicts are ordered and non-compensatory — a stronger verdict always dominates a weaker one. Seven seals are defined in arifosmcp/models/verdicts.py:

VOID  >  HOLD_888  >  HOLD  >  SABAR  >  PARTIAL  >  PROVISIONAL  >  SEAL

Verdict

Meaning

Plain English

What happens

SEAL

Authorised under stated conditions, W³ ≥ 0.95

Go

Proceed to execution

PARTIAL

A derived floor warns

Go carefully

Proceed with cooling and monitoring

PROVISIONAL

Time-limited authorisation, expires or downgrades to PARTIAL/HOLD

Proceed with expiry timer

Auto-revoke at valid_until; no extension without re-judging

SABAR

Not yet decidable — reality hasn't finished speaking

Wait for evidence

Retry permitted later; distinct from HOLD

HOLD

Insufficient evidence, or human approval required

Wait for human

Pause; await a human decision

HOLD_888

Immediate sovereign escalation

Stop — the sovereign decides

Escalate to F13

VOID

Blocked by a hard constitutional floor

Blocked

Stop; the constraint must be resolved

Seven seals, ordered and non-compensatory (lattice at top of section). The diagram below shows the five common verdicts for width; PROVISIONAL and HOLD_888 are omitted from the diagram but present in the lattice and table.

Floors are never averaged. One floor failure propagates into the verdict; there is no compensating score.

13 Constitutional Floors (F1–F13)

Every proposal is evaluated against 13 non-compensatory policy constraints. Canonical names and rules: FEDERATION_CONTRACT.md §3, GENESIS/000_KERNEL_CANON.md, and the constitution at static/arifos/theory/000/000_CONSTITUTION.md.

Floor

Name

Rule

Pass condition

F1

AMANAH

Reversible first. Irreversible → 888_HOLD unless the sovereign acknowledges

Score ≥ threshold (higher = better)

F2

TRUTH

P(truth) ≥ 0.99. Cheap claims = VOID. Evidence carries an OBS/DER/INT/SPEC label

Score ≥ threshold (higher = better)

F3

TRI-WITNESS

W₃ = ∛(Human × AI × Earth) ≥ 0.75 at judgment time (floor for admissibility; W³ ≥ 0.95 is the additional bar for SEAL)

Score ≥ threshold (higher = better)

F4

CLARITY

ΔS ≤ 0 — every output reduces entropy, never adds it

Score ≥ threshold (higher = better)

F5

PEACE²

Non-destructive power — block harm and extraction

Score ≥ threshold (higher = better)

F6

EMPATHY (operational: MARUAH)

Protect the weakest stakeholder; dignity is not tradeable

Score ≥ threshold (higher = better)

F7

HUMILITY

Ω₀ ∈ [0.03, 0.05]. No fake certainty

In-band [0.03, 0.05] = pass; outside = fail

F8

GENIUS

G ≥ 0.80 for complex actions — the simplest correct path

Score ≥ threshold (higher = better)

F9

ANTIHANTU

No deception, manipulation, or consciousness claims

Lower = better (lower ceiling = less manipulation)

F10

ONTOLOGY

AI-only ontology. A soul claim is VOID; map it to harness content

Score ≥ threshold (higher = better)

F11

AUDITABILITY

Every decision logged, inspectable, attributable

Score ≥ threshold (higher = better)

F12

RESILIENCE

Injection defence. Input risk < 0.85

Lower = better (lower = less injection surface)

F13

SOVEREIGN

Human veto is FINAL. The harness switch belongs to the human

Score ≥ threshold (higher = better)

Three semantic classes — verified by the live unit on 2026-09-21:

  • Score floors (F1–F6, F8, F10, F11, F13): higher = better. Score ≥ canonical threshold.

  • Band floors (F7): in-band [0.03, 0.05] = pass; outside = fail. 0.02 fails.

  • Ceiling floors (F9, F12): lower = better. Value below canonical ceiling.

W₃ has two named thresholds: W₃ ≥ 0.75 is the floor for admissibility; W₃ ≥ 0.95 is the additional bar for SEAL. Between them a proposal can be PARTIAL or HOLD, never SEAL.

Three naming notes, so a reader can reconcile this table with the wire:

  • F6 has two canonical names by design — EMPATHY in the public register, MARUAH in the kernel, logs and receipts. Both are correct; the bridge is documented in GENESIS/000_KERNEL_CANON.md §3.4. Public surfaces render EMPATHY.

  • The live runtime keys F10–F13 as L10–L13 in /health → runtime_floors. Same floors, different key prefix.

  • Lower-is-better floors are reported as raw measurements, not as failures: live values on 2026-09-21 were F7 = 0.04, F9 = 0.15, L12 = 0.425, and /health → runtime_floors_status reports 13/13 pass with every floor measured: true.

VAULT999 (append-only audit ledger)

Every consequential verdict, evidence chain and execution receipt is written to VAULT999 — a hash-chained, append-only JSONL ledger set, tamper-evident by construction, not tamper-proof.

What it stores is provenance, not copies. An entry carries its payload hash, entry hash, previous entry hash, trace root, actor, session, decision context and evidence references. Source payloads are not duplicated into the ledger — auditability here means provenance + integrity + reconstructability, which is both stronger and safer than copying everything forever (evidence can contain sensitive material).

A hash chain proves internal consistency. It does NOT prove that the chain was not recomputed by whoever controls the ledger. External anchoring — periodic publication of the chain head, or signing by a key the kernel does not hold — is not yet implemented. Treat VAULT999 today as self-verifiable, not independently attestable. See "What Is Not Yet Proven" for the gap row.

Measured 2026-09-21: 241,765 lines across 24 JSONL ledgers in VAULT999/ (largest: outcomes.jsonl 93,266 · arifflow_sealed.jsonl 52,056 · apex-zen-receipts.jsonl 30,977). Composition: the majority is operational telemetry and receipt ingestion; constitutional verdict count is published separately (see scripts/verify_vault_chain.py --verdict-count-only). The chain report returns overall: INTACT for the active ledgers, reporting 2 strict link breaks in the frozen v1 legacy ledger as historical facts rather than hiding or rewriting them. Live record count is re-stamped into the header manifest above by scripts/update_readme_sot.py.


Architecture

arifOS Federation — planes and classes

    ┌───────────────────────────────────────────────────┐
    │  Authority plane — arifOS :8088                    │
    │  Constitutional judgment · F1–F13 · VAULT999        │
    └──────────────────────┬────────────────────────────┘
                           │
    ┌──────────────────────▼────────────────────────────┐
    │  Attention plane — AAA :3001                       │
    │  What matters · routing · state · skill catalog    │
    └──────────────────────┬────────────────────────────┘
                           │
    ┌──────────────────────▼────────────────────────────┐
    │  Execution plane — A-FORGE :7071/:7072             │
    │  Governed mutation · leases · receipts             │
    └──────────────────────┬────────────────────────────┘
                           │
    ┌──────────────────────▼────────────────────────────┐
    │  Witness + record — FRAME :18085 · VAULT999        │
    │  Independent observation · immutable history       │
    └───────────────────────────────────────────────────┘
        boundaries:  HERMES :18087 (semantic) · CHRON :18102 (temporal)
        domains:     GEOX :8081 · WEALTH :18082 · WELL :18083
        metabolism:  arifFlow :7073      gateway: FED :7074      synthesis: i-ARIF

arifOS is the kernel; the other components are supporting infrastructure. GEOX is the primary reference implementation — a live geoscience organ whose evidence is governed in a high-consequence, uncertainty-heavy domain. FRAME's output is evidence, never a verdict. Domain organs compute; they never authorise.

Organs vs. boundaries — the ruling. The ratified organ table (FEDERATION_CONTRACT.md §2) lists 7 live organs: arifOS, A-FORGE, AAA, GEOX, WEALTH, WELL, arifFlow. HERMES is Tier-3 boundary infrastructure and is not an organ (§2.1, 2026-09-14 ruling); CHRON is the temporal boundary service. They are reported here by class rather than padded into an organ count — "everything we built" is not an architectural category.

Port note: i-ARIF has no listening port; :18095 is apa-github-bridge, one of the APA boundary bridges (:18075–:18099). A previous revision of this file named :18095 as i-ARIF — corrected 2026-09-21.


MCP Interface

The kernel exposes 8 canonical verbs over Streamable HTTP. Verified by live tools/list on 2026-09-21:

Stage

Verb

Purpose

000

arif_init

Session ignition — binds actor, floors and audit before any other verb

111

arif_observe

Sense reality into evidence with epistemic tags and uncertainty bounds

333

arif_think

Structured reasoning under F2/F7, with OBS/DER/INT/SPEC labels

444

arif_route

Intent → organ routing, dispatching to GEOX / WEALTH / WELL / A-FORGE

555

arif_memory

Governed memory — L1–L6 recall, storage, promotion

666

arif_judge

Evaluate a proposal; returns a binding verdict with the floor chain

777

arif_forge

Execution gate via A-FORGE — mutates only after a SEAL verdict

999

arif_seal

VAULT999 immutable append — seals a completed chain with its receipt

arif_forge is a governed dispatch verb: it routes an authorised action toward the execution organ and mutates only after SEAL. The kernel does not perform the underlying mutation. arif_route is authority-aware dispatch, not attention: it decides whether and how a routing may proceed, and may consult AAA. The judge never executes; the executor never certifies.

Tool-count semantics (so no two surfaces appear to disagree): tools_loaded = 8 — the public MCP facade, and the only number to quote publicly · canonical superset = 25 — 8 exposed + 13 hidden verbs (e.g. arif_challenge, arif_judge_deliberate), hidden by design · total_declared_tools = 48 · tools_registry_size = 62 (includes aliases) · operational_tools = 3 — the count with a durable SUCCESS in the last 24 hours, i.e. proven live, not merely invocable.


Verification Status

Live-probed 2026-09-21 (UTC+08); rows marked ↻ re-probed 2026-09-30 and 2026-10-05. Re-run the commands; static counts are not evidence.

Surface

Status

Evidence

Public repository

Live

GitHub ariffazil/arifOS, AGPL-3.0

PyPI package

Published 1!2026.9.2 (uploaded 2026-09-15T16:42Z)

pip install arifos — pypi.org/project/arifos; tree is 1!2026.10.1, unreleased

Live kernel

Healthy, 13/13 floors

curl localhost:8088/health → status: healthy, floors_active: 13

MCP interface

8 exposed, 3 with a durable SUCCESS in the last 24 h (tools_loaded: 8, operational_tools: 3)

live tools/list; protocol advertises 2026-07-28, negotiation settles 2025-11-25

Floor enforcement

13/13 measured pass

/health → runtime_floors_status (F7 = 0.04, F9 = 0.15, L12 = 0.425 lower-is-better)

VAULT999 ledger ↻

Healthy, chain INTACT

scripts/verify_vault_chain.py → overall: INTACT (its own scope: the 3 active ledgers under arifOS/VAULT999/). The previously published "241,765 lines / 24 ledgers" named no population and matched no measured root, so it is withdrawn rather than refreshed. Census 2026-09-30T07:44:35Z — two different, both-correct denominators: (a) recursive whole-estate, find <root> -name '*.jsonl' then count non-blank lines: arifOS/VAULT999 60/354,283 · /var/lib/arifos/vault 7/20,986 · ~/.local/share/arifos/vault999 66/16,555 · AAA/VAULT999 1/238 = 134 ledgers / 392,062 lines; (b) the generated SOT-MANIFEST header above reports 342K+ records because scripts/update_readme_sot.py counts only top-level VAULT999/*.jsonl in this repo, non-recursively (measured 24 files / 342,582 lines), excluding the subdirectory ledgers (a) includes. Neither is wrong; they have different populations. The vault is append-only and live, so every figure here is a snapshot — quote the timestamp with it.

Source / build / deploy ↻

Aligned — no drift

source_commit = built_commit = deployed_commit = 27409ddbf; /health → drift: false, runtime_matches_build: true, deployment_drift_status: aligned. Release provenance is re-attested on every scripts/deploy-release.sh run (VERDICT: PASS, strict) — quote the live value, not this one

Contract schema

8/8 published, no drift

contract_status: {tool_count: 8, schemas_complete: true, contract_drift: false}

Federation surfaces

10/11 reachable without drift; WELL degraded (drift: true)

ports 8088 · 3001 · 7071 · 7072 · 8081 · 18082 · 18083 · 7073 · 7074 · 18085 · 18102; WELL at 127.0.0.1:18083/health returns HTTP 200 but reports drift: true, so the row honestly reads 10/11 healthy + 1 degraded

Machine-readable ↻

tools.json live, 56,224 bytes

tools.json · llms.txt (generated mirror — see gaps) · CITATION.cff

Scanner attestation

148 rules · 0 unauthorised mutations; 26/26 GRADE A

mcp.arif-fazil.com/proof/

What Is Not Yet Proven

Every row names its own gap. Repairing a claim by substituting a stronger one is worse than the stale claim it replaced.

Gap

Risk

Status

Independent security audit

Adversarial bypass testing not published

In progress — external researcher reviewing since 2026-08-25. First finding (fetch-surface SSRF) fixed, released in 1!2026.9.1. A second scan (2026-09-15/16, mcp-safeguard) found 2 confirmed issues: Cypher injection (graph-wipe risk, HIGH — key whitelist landed in l5_sovereign_forge.py) and a fastmcp decode-after-match path traversal (MEDIUM — arifOS-side containment guard landed in atlas333.py; upstream report pending). Both are at code level; neither has a released-artifact verdict. See SECURITY.md

Sovereign-binding of irreversible acknowledgement

ack_irreversible: true and actor_signature are transported as payload fields. Nothing published demonstrates that a human signature is bound to (transaction_hash ‖ state_hash ‖ nonce), single-use, scoped to one organ, and unforgeable by the proposing agent.

Open — design drafted (docs/judgment-integrity-2026-09-22/ARIFOS_AUTHORITY_SERVICE_DESIGN.md), not implemented. Until then, F13 is enforced by convention and band-gating, not by cryptographic transaction binding. This matches OWASP transaction-authorization practice: credentials unique per operation and bound to significant transaction data, not to session.

Ledger head not externally anchored or signed

A hash chain held entirely by one party detects external modification but not wholesale rewrite by the holder, because the holder can recompute every hash.

Open — periodic publication of chain head, or signing by a key the kernel does not hold, is not implemented. VAULT999 today is self-verifiable, not independently attestable.

Third-party evaluation

No external reviewer has published findings

In progress — one review under way since 2026-08-25; nothing published

Reproducible demo by strangers

Onboarding path not independently tested

Partial — examples/enterprise_operations_demo.py runs in-process and is verified here; no stranger has reproduced it unaided

Enterprise deployment

No production customer reference

Open

Standards conformance

MCP/A2A conformance results not published externally

Partial — CI 06-mcp-conformance.yml; conformance report says result: PARTIAL, spec_version: 2025-11-25, with per-tool schema validation DEFERRED

SBOM and signed releases

Supply-chain integrity unverified externally

Partial — CycloneDX generator + sbom job on publish; no signing, no CVE scan

Container image metadata

The image's own labels contradict this repository

Fixed 2026-10-05 (8372cecc3) — labels aligned to measured reality: licenses → AGPL-3.0 (matches the repo LICENSE file and PyPI metadata — both said AGPL; the BSL-1.1 label was the outlier), tool counts 13/7 → 8 (tools_loaded: 8 measured), PORT env + MCP port label → 3000 (CMD/EXPOSE/HEALTHCHECK all listen on 3000 by design for Manufact cloud compat), image version → 2026.10.1

Generated mirrors drift

llms.txt still prints the retired 888 judge stage and version v2026.07.24

Fixed 2026-10-05 (8372cecc3) — KERNEL 888 prose corrected to KERNEL 666 in arifosmcp/constitutional_map.py and tools_sot.yaml; mirrors regenerated via scripts/generate_tool_manifest.py. The version line is now DERIVED from installed package metadata in the generator (was hardcoded v2026.07.24), so it cannot go stale again. Live tools/list confirms KERNEL 666 post-deploy a98af0d

Generated discovery artifact drift

smithery.yaml no longer matches the kernel ABI registry — the guard itself fails

Fixed 2026-10-05 (8372cecc3) — regenerated from the ABI registry: Kernel ABI verified: 8 capabilities, 6 public tools, --check exit 0

Development test suite

The suite needs a live kernel

Fixed 2026-09-25 (collection) — stale 269-line duplicate tests/test_rasa_bench_10.py removed (canonical 632-line bench lives at /root/.hermes/policy/test_rasa_bench_10.py where rasa_boundary resolves in-place); 7,668 tests collect clean (re-measured 2026-10-05, was 7,622); tests/conftest.py still refuses to run without a reachable kernel at :8088 (by design). Collection ≠ green: the identity/authority/session/attestation subset carries 34 pre-existing failures against 321 passes (measured 2026-09-30 on a pristine main worktree), so quote a failure set, never a bare pass count

Semantic layer (Graphiti)

Knowledge graph retired from the read path

Operational gap — graphiti_read: retired_888, semantic_floor: disabled by choice (ARIFOS_ML_FLOORS=0)

Observability

Tracing partially wired

Partial — sovereign Postgres backend active; arifFlow FlowReceipt adapter live; OTel spans on all 8 canonical verbs; langfuse_tracing: NOT_WIRED after the cutover to kabarkan; caller-side trace propagation incomplete

Comparative benchmark

No published comparison against alternative frameworks

Open

Maintainer continuity

Single sovereign, single reviewer; no succession or key-recovery procedure published

Open — relevant to any institutional adoption

__version__ strings are stale

Module __version__ lags kernel release; readers may quote it incorrectly

Closed 2026-10-05 — arifosmcp/__init__.py derives __version__ from installed package metadata (_pkg_version("arifos")), and the live wheel reports 1!2026.10.1, identical to pyproject.toml. Submodule __version__ strings (webmcp, entropy_kernel, federation, hexagon) are component tags, not kernel releases — they track their own artifacts

Identity registry exists in two copies

A one-sided edit ships a split-brain identity registry: contracts/identity.py (imported by ~9 runtime sites) and arifosmcp/contracts/identity.py (~3 sites, and the one the wheel ships) had already diverged — the packaged copy was missing the I_ARIF entry added by the 2026-08-21 Seal C fix and the GEMINI/FI-004 lane entry, so i-arif and agy/FI-004 resolved differently depending on which copy a module imported

Mitigated 2026-09-30, consolidation still open — both copies were re-aligned to 18 actors (7796ec701) and tests/contracts/test_identity_registry_no_divergence.py (31 tests) now fails on any drift in key set, aliases, normalization of 20 probe ids, or a missing actor_lookup_candidates; it was falsified against the pre-alignment state, where it caught exactly gemini, i-arif and agy/FI-004, so it is not vacuous. Open: pick one owner — migrate the 9 imports to arifosmcp.contracts.identity and retire the top-level copy, or keep both and keep the tripwire. Not decided here because it is a source-of-truth call, not a bug fix

Authority floor depended on a pointer stub and an unshipped package

Two independent faults capped every non-exempt agent at OBSERVE_ONLY, so mutation-capable citizens silently lost tool access — surfacing as the "tool inexplicably unauthorized" class. (a) Boot Q5 read identity.toml with _file_read(A) or _file_read(B); both are 101-byte DERIVED stubs reading "superseded by: /root/AAA/identity.toml", and a non-empty stub short-circuits the or, so the canonical F13 payload was never read → Q5=NO → boot_state=FAIL → _apply_boot_gate demoted LIMITED_MUTATE/FULL. Authority was in practice coming from the 29-entry exemption list, not from attestation. (b) contracts* was absent from the wheel include, so the kernel's identity path resolved against a July-24 leftover. A bare except ImportError: pass in _apply_boot_gate hid a wrong module path, making the 2026-08-21 alias bypass dead code

Fixed 2026-09-30 — 5cc357974 (_read_identity_toml_chain() joins all candidates and follows superseded by: pointers; Q5 NO→PARTIAL, which passes the gate since the shipped predicate is boot_state != "FAIL"), 2eaec19c2+d43b7fece (one shared exempt_actor_band() resolver replacing 7 raw-string membership tests, so the documented name/FI-nnn form resolves), 69175c0ec (patched the copy production imports; boot-gate bypass moved off the fragile import), 27409ddbf (contracts* ships). Verified live: 17/17 AAA identities bind SEAL + mutation_allowed=True + VERIFIED with an ACT carrying LIMITED_MUTATE and all 9 verbs, 0 boot demotions, and the control nobody/FI-999 is still refused — trust boundary unchanged, exempt membership still does not auto-verify. The generalizable defect: a stand-in served where the payload belonged, and a fail-silent except hid it — three times in one session.

See SECURITY.md for the threat model, known gaps and disclosure policy, and docs/evidence/claims.yaml for the machine-readable claim registry.


Who Is This For

Operators deploying AI agents in regulated environments who need an independent judgment layer between agent proposals and execution.

Developers building AI agent systems who want a policy decision point as a service.

Evaluators and security reviewers assessing AI governance frameworks.

Domain builders adapting governance to a specific field (geoscience, finance, healthcare).

Start here: docs/START_HERE.md


Founding Context

arifOS was built by Muhammad Arif bin Fazil, a senior exploration geoscientist who spent his career making decisions where observations are incomplete, interpretations are probabilistic, provenance matters, and irreversible action must be gated. He transferred that discipline into agent runtime governance.

The system is named after its founder and reflects a core belief: governance is a systems problem, not a model problem.


Development

# Clone
git clone https://github.com/ariffazil/arifOS.git
cd arifOS

# Install (dev tier — kernel + test tooling)
pip install -e ".[dev]"

# Run tests (needs the kernel reachable at :8088; one module fails collection — see gaps)
python -m pytest tests/ -q

# Kernel health (this Makefile target probes the kernel only)
make health

# Start the kernel
PORT=8088 arifos-mcp streamable-http   # or: PORT=8088 python -m arifosmcp.runtime --mode streamable-http

See CODEOWNERS for sovereign ownership of automation surfaces and CONTRIBUTING.md for contribution guidelines.

Project Structure

arifOS/
├── arifosmcp/          # Core kernel package                    [wheel root]
│   ├── abi/            # Capability registry and floor definitions
│   ├── constitution/   # Constitutional floor implementations
│   ├── contracts/      # Packaged copy of the identity/verdict contracts
│   ├── kernel/         # Core judgment engine
│   └── VAULT999/       # VAULT999 ledger implementation
├── arifos/             # Phase 3 identity/consent namespace      [wheel root]
├── contracts/          # Runtime contract package — identity registry,
│                       # gateway discovery, verdicts, continuity, envelopes.
│                       # Imported by 5 runtime files; see note.    [wheel root]
├── core/               # Load-bearing legacy root (session.py imports
│                       # core.shared.types)                      [wheel root]
├── schemas/            # INIT v2 schemas (F13-ratified 2026-09-20) [wheel root]
├── examples/           # Runnable governed workflow demo
├── tests/              # Test suite (pytest — constitutional + integration)
├── scripts/            # Operational tooling (incl. update_readme_sot.py)
├── docs/               # Documentation
│   ├── START_HERE.md   # External reader entry point
│   ├── QUICKSTART.md   # Verified first governed call
│   └── evidence/        # Claim registry and evidence index
└── pyproject.toml      # Package metadata — [wheel root] marks
                        # packages.find.include

The five [wheel root] entries are exactly what ships. Until 2026-09-30 this tree listed only arifosmcp/, and contracts* was absent from [tool.setuptools.packages.find].include — while the runtime imported contracts.* at 10 import statements across 5 files under arifosmcp/ (contracts.identity 6, verdicts/continuity/artifacts/envelopes 1 each), plus arifosmcp/contracts/__init__.py doing from contracts.identity import *. Scope matters: a repo-wide grep reports ~4× that, because build/lib/ holds a 1,402-file duplicate tree and tests/ adds its own imports (gateway_discovery 7, all of them tests — 0 in the runtime). Those inflated figures were what an earlier revision of this note published; corrected 2026-09-30. Production resolved those imports only because pip does not delete a pre-existing unrelated site-packages/contracts/ leftover; a fresh venv or a machine migration would have raised ImportError at kernel import time. Fixed in 27409ddbf. An incomplete structure diagram is not a cosmetic gap — it is how an unshipped, runtime-critical package stays invisible.

contracts/ and arifosmcp/contracts/ are two copies of the same contract modules and had already diverged. They are kept in agreement by tests/contracts/test_identity_registry_no_divergence.py; consolidating them into one owner is still open (see gaps).


Sister Repositories

Repository

Purpose

AAA

Attention plane — routing, state, skill catalog, A2A gateway

A-FORGE

Execution engine after authorisation

GEOX

Earth sciences domain evidence

WEALTH

Capital and financial intelligence

WELL

Human and machine vitality observation

arifFlow

Metabolic ledger daemon — FQ monitoring, receipt ingestion

FRAME

Independent observer — drift detection, evidence gathering (repository archived, organ live)

No public repository today: FED (:7074) and i-ARIF. They are deployed and reachable, but their source is not published — a previous revision of this file linked to repositories that return 404. Treat their behaviour as observed from /health only, not as readable code.


Evidence & Trust

arifOS publishes verifiable evidence for its claims. Every public claim links to an artifact that can be regenerated.

Claim

Evidence

Status

MCP conformance

CI workflow · report

Partial — result: PARTIAL, per-tool schema validation deferred, results not published externally

ABI stability

Drift guard

Partial — --check currently reports drift in smithery.yaml (reproduced on a pristine main worktree, 2026-09-21); last clean resync 2026-09-15

Floor enforcement

/health → runtime_floors_status

Verified — 13/13 measured pass (live probe)

Vault integrity

scripts/verify_vault_chain.py

Verified — overall INTACT; 2 legacy breaks reported as historical

SBOM

Generator

Partial — CycloneDX generated, no CVE scan, unsigned

Observability

Telemetry

Partial — Postgres + arifFlow live; tracing incomplete

Governance hardening

Adversarial test spec (federation-internal, not published)

Partial — spec written, no external audit

Quickstart

Verified quickstart

Verified by the maintainer on 2026-09-21; not yet reproduced by a stranger

See docs/evidence/ for the claim registry (claims.yaml), the evidence index and the contracts for observability, release integrity, reproducibility and threat model.

Honesty principle: we publish what passed, what failed, and what remains unknown. We do not claim maturity beyond our evidence.


Who Maintains This

One human, and the agents he directs.

I am a geologist, not a programmer. I did not write this codebase and I do not read it line by line. What I do is point at where the problem is — and the agents in my federation solve it, inside the constitution and the review gates I set, with the commit trail to show who ran what.

Read the commits if you want to check that claim: the author fields are agent handles, not aliases of mine. That is the deliberate shape of this project, not a detail being hidden. It also sets the honest expectation — the design and the judgment are mine, the implementation is theirs, and where the two disagree, the bug is mine to answer for.


Contributing

See CONTRIBUTING.md for guidelines.


Security

See SECURITY.md for the threat model, known vulnerabilities and disclosure policy.


License

AGPL-3.0 — GNU Affero General Public License v3.0 (LICENSE).

When deployed over a network, the complete source code must be made available to all users interacting with the service, consistent with AGPL-3.0 terms. (The container image's OCI licence label does not currently agree with this — see "What Is Not Yet Proven".)


Revision 2026-09-21 — audited against the live kernel, the ratified federation contract and the repository itself. Every number in this file was re-measured, not carried forward; each correction is receipted in the commit history.

Independent audit pass 2026-09-22 (Copilot external, mode ENTERPRISE). Findings A, B, C, D, E, F, G, H, I, K, L, N — all four blocking contradictions + the two highest-leverage gaps (G sovereign-binding, H ledger anchoring) — corrected in this revision. M (TOC/badges) and a per-pass signature remain open as structural polish. See commit history for per-finding receipts.

Ditempa Bukan Diberi — Forged, Not Given.

Available Tools

11 tools
arif_bridge_connect444 Bridge · Direct Organ (HIGH)AInspect

KERNEL 444-direct · Low-level organ call (organ + tool_name required). Bypasses intent routing. Authority: HIGH / lease — often 888_HOLD for anonymous. Agents should prefer arif_route (same reach, safer default). Not a generic MCP proxy; only federation organs under kernel envelope.

ParametersJSON Schema
NameRequiredDescriptionDefault
organYes"geox" | "wealth" | "well" | "geox" (case-insensitive)
actor_idNoCalling actor (injected into envelope)
_envelopeNo
argumentsNoTool arguments dict
tool_nameYesMCP tool name on the target organ
session_idNoGoverning session

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide basic safety hints (non-destructive, non-idempotent, read-write). The description adds context about authority levels (HIGH / lease) and that it bypasses intent routing, which is valuable. However, it does not elaborate on potential side effects or consequences of misuse beyond these points, so not fully exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise—four sentences. It is front-loaded with the kernel identifier and key purpose, and every sentence adds essential information without redundancy. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 6 parameters, 2 required, high schema coverage, and an output schema, the description adequately covers purpose, usage guidelines, and key constraints. It does not explain return values, but the output schema exists. Minor gap: no mention of error cases or expected behavior when parameters are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is high (83%), so the baseline is 3. The description reinforces that organ and tool_name are required, but does not add new details about other parameters like actor_id, arguments, or session_id beyond what the schema provides. No parameter-specific guidance is given.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: a low-level, direct organ call that bypasses intent routing, requiring organ and tool_name. It distinguishes itself from the sibling tool arif_route, which is the safer default. The description is specific and actionable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly advises agents to prefer arif_route instead, providing clear when-to-use and when-not-to-use guidance. This directly addresses the decision between this tool and its alternative, making it highly helpful for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arif_composeCompose · Kernel ReplyA
Read-onlyIdempotent
Inspect

KERNEL reply · Final human-facing composition (citations, tone, ΔS≤0). Call LAST after observe/think/judge — not mid-pipeline. Authority: L0–L1. Modes: compose | summarize | cite | tone_shift | style | format. Not a substitute for arif_judge or arif_seal.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNocompose
styleNo
messageNo
actor_idNo
languageNoen
_envelopeNo
citationsNo
session_idNo
session_tokenNo
ai_involvementNofull

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint, idempotentHint, destructiveHint. The description adds significant behavioral context: it is a 'final human-facing composition' with a ΔS≤0 constraint (no state change), and it supports multiple modes. There is no contradiction; the description enriches transparency beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded with key information. It uses jargon like 'ΔS≤0' that may be obscure but is concise. Every sentence adds value. Minor deduction for the jargon density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 10 parameters and an output schema exists (not shown), the description covers overall purpose and pipeline position but lacks parameter documentation. It is adequate for understanding when to use the tool but insufficient for parameter-level decisions without the schema. The output schema may compensate, but description doesn't reference it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It mentions modes and style but introduces 'format' and 'style' as modes not present in the enum (schema lacks 'style' and 'format' as enum values, though style is a separate parameter). Other parameters (message, actor_id, language, etc.) are not explained at all. This inconsistency and lack of detail limit its helpfulness.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the tool's purpose: 'Final human-facing composition' and positions it as the last step in a pipeline ('Call LAST after observe/think/judge'). It lists specific capabilities (citations, tone, ΔS≤0) and modes, clearly distinguishing it from siblings like arif_judge and arif_seal.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance on when to call ('LAST after observe/think/judge') and when not to ('not mid-pipeline'). It also states what the tool is not a substitute for (arif_judge, arif_seal), providing clear exclusion criteria. This helps the agent select the correct tool in the pipeline.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arif_critique555 Critique · HeartA
Read-onlyIdempotent
Inspect

KERNEL 555 · Heart — ethical/dignity/risk stress before judgment (not SEAL). Select when blast_radius MEDIUM+, human/dignity impact, or irreversible risk. Requires non-empty target (proposal/plan text). Authority: L1. Modes: critique | redteam | maruah | deescalate | empathize | simulate | instruction_scan. Returns: risk, floors, human impact. Skip pure technical with zero human stake. Binding verdict → arif_judge.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNocritique
targetNo
actor_idNo
_envelopeNo
session_idNo
session_tokenNo
evidence_receiptNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint false, which the description does not contradict. The description adds beyond annotations: requires non-empty target, lists return types (risk, floors, human impact), and describes modes. However, it does not detail behavioral variations across modes, keeping a score of 4.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph that front-loads purpose and usage. It packs substantial information without excessive verbosity, though it could be better structured with separation of modes and return types.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 7 parameters (all optional, mostly metadata), no required params, and existence of output schema, the description covers core aspects: purpose, selection criteria, modes, requirements, and return types. Minor inconsistencies (mode list) and missing details on non-core parameters prevent a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description provides partial meaning for 'mode' and 'target' but contains inaccuracies: lists modes like 'simulate' and 'instruction_scan' not in schema enum, and omits 'shadow' and 'empathy'. Also states target is required but schema allows null default. With 0% schema coverage, the description poorly compensates.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool performs ethical/dignity/risk stress testing before judgment, distinguishing it from SEAL. It specifies the conditions for selection (blast_radius MEDIUM+, human/dignity impact, irreversible risk) and names a sibling tool (arif_judge) for binding verdict.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use (blast_radius MEDIUM+, human/dignity impact, irreversible risk) and when to skip (pure technical with zero human stake). Also mentions authority level L1 and available modes, providing clear context for agent decision.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arif_forge777 Forge · Execute GateC
Destructive
Inspect

KERNEL 777 · Execution gate via A-FORGE (hands, not law). Mutates only after arif_judge SEAL + lease/chain IDs — no self-authorize. Authority: 888_HOLD without SEAL. Modes include dry_run | engineer | query | write. Public execution verb (arif_act is internal alias only). Skip while still planning (arif_think) or without judge SEAL.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoengineer
queryNo
plan_idNo
actor_idNo
manifestNo
_envelopeNo
session_idNo
arif_ack_idNo
artifact_idNo
vault_entry_idNo
seal_verdict_idNo
ack_irreversibleNo
judge_state_hashNo
approved_action_hashNo
constitutional_chain_idNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructiveHint=true. Description adds that mutation only occurs after judge SEAL and lease/chain IDs, and mentions modes. However, terms like '888_HOLD without SEAL' are undefined, reducing clarity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short but dense with jargon. It attempts to be concise but sacrifices clarity. Key information is present but poorly structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having an output schema, the description does not explain what the tool returns or how parameters like 'seal_verdict_id' or 'approved_action_hash' relate to the judge workflow. Given the complexity (15 params, destructive behavior, sibling tools), the description is insufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description adds very little about the 15 parameters. Only 'mode' enum values are mentioned in passing. This is insufficient for an agent to correctly use the tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses cryptic jargon like 'KERNEL 777', 'A-FORGE', and '888_HOLD' without plain-language explanation. It vaguely indicates it's an execution gate requiring judge seal, but the core purpose is unclear for an AI agent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides some guidance: skip without judge SEAL or while still planning (arif_think). However, it doesn't explicitly state when to use this tool vs siblings like arif_judge or arif_think, leaving ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arif_init000 Init · Kernel SessionA
Idempotent
Inspect

KERNEL 000 · Session ignition — not a helper app. Binds actor, floors, and audit before any other arif_* verb can govern. Without session_id, kernel treats you as anonymous (OBSERVE_ONLY / SYUBHAH). Authority: pre-session open; light/init mint session_id + authority band. Modes: ping | light | init | resume | validate | epoch_open | epoch_seal | canary | preflight | triage. Returns: session_id, actor_verified, authority, allowed_next_verbs, next_tool. Skip when live session already bound → arif_triage; pure facts only → arif_observe.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoinit
nonceNo
intentNo
contextNo
payloadNo
toolingNo
verboseNo
actor_idNo
epoch_idNo
evidenceNo
trace_idNo
_envelopeNo
session_idNo
agent_policyNo
counterpartyNo
sovereign_idNo
actor_signatureNo
caller_actor_idNo
delegation_modeNo
idempotency_keyNo
ack_irreversibleNo
executor_actor_idNo
declared_model_keyNo
client_capabilitiesNo
requested_authorityNoOBSERVE_ONLY
previous_session_hashNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide idempotentHint and destructiveHint. The description adds context: binds actor/floors/audit, mints session_id and authority band, and describes anonymous mode without session. Good additional insight.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense with jargon but front-loaded with core purpose. Could be more concise; multiple concepts packed into a single paragraph.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 26 undocumented parameters and no formal output schema, the description lists return fields but lacks explanations for most parameters. Incomplete for a complex initialization tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has 26 parameters with 0% description coverage. The description only lists some modes but misses two enum values (opt_out, opt_out_profiling) and does not explain other parameters like nonce, intent, context, etc. Insufficient compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that this tool initiates a kernel session, binds actor, floors, and audit, and is required before other arif_* verbs. It distinguishes from siblings by noting when to use arif_triage or arif_observe instead.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use (before other arif_* verbs, if no session) and when not to use (live session bound → arif_triage; pure facts only → arif_observe). Provides clear alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arif_judge888 Judge · VerdictA
Read-only
Inspect

KERNEL 888 · Constitutional verdict — only organ that SEAL/HOLD/SABAR/VOIDs. Not advice; binding arbitration of floors + authority. Authority: 888_HOLD / SOVEREIGN session required for real adjudicate. REQUIRES: actor, intent, domain, reversibility_level, blast_radius (+ evidence). Modes: judge | compare | history | explain | floor_status | witness_consensus. Skip if evidence incomplete → arif_observe; plan incomplete → arif_think; reversible low-risk advisory only.

ParametersJSON Schema
NameRequiredDescriptionDefault
actorYes
domainYes
intentYes
actor_idNo
evidenceNo
_envelopeNo
session_idNo
measurementNo
blast_radiusYes
session_tokenNo
authority_tokenNo
epistemic_stateNoUNKNOWN
reversibility_levelYes
requested_capabilityYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Description claims binding arbitration and state-changing actions (SEAL/HOLD/SABAR/VOIDs), contradicting the readOnlyHint=true annotation. This is a serious inconsistency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is dense and front-loaded with core purpose, but uses cryptic jargon (KERNEL 888, floors, SABAR) that may confuse agents. Not optimally concise for AI consumption.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers core concept, modes, prerequisites, and alternatives, but lacks explanation of cryptic terms and some parameters. Adequate given output schema exists, but could be more comprehensive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Description adds meaning to key required parameters (actor, intent, domain, etc.) and mentions evidence, but many parameters (actor_id, session_id, etc.) are left unexplained. With 0% schema coverage, more detail would be beneficial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool is for constitutional verdict and binding arbitration, with specific verbs like SEAL/HOLD/SABAR/VOIDs. It distinguishes from siblings by naming alternative tools for incomplete evidence or plans.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states required parameters (actor, intent, domain, etc.) and provides clear when-to-use guidance by directing to siblings when evidence or plan is incomplete, or for low-risk advisory.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arif_memoryMemory Governor · KernelB
Destructive
Inspect

KERNEL memory governor · L1–L6 stack under F1/F2/F4/F11 (not a free notepad). Recall/inspect free-ish; remember/promote/revise/forget are J-space mutations. Authority: recall L0; writes gated. Modes: recall | inspect | attest | remember | promote | revise | forget. Skip ephemeral one-off facts.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNorecall
tierNo
queryNo
scopeNo
top_kNo
aspectNo
hybridNo
policyNo
cascadeNo
contentNo
includeNo
payloadNo
seal_idNo
to_tierNo
actor_idNo
lease_idNo
metadataNo
trace_idNo
_envelopeNo
from_tierNo
memory_idNo
tier_hintNo
timestampNo
provenanceNo
redact_piiNo
session_idNo
structuredNo
graph_firstNo
new_contentNo
truth_classNo
caller_chainNo
memory_classNo
policy_basisNo
embedding_refNo
include_proofNo
session_tokenNo
vault_versionNo
human_approvalNo
new_structuredNo
temporal_as_ofNo
tombstone_textNo
idempotency_keyNo
new_truth_classNo
resolution_kindNo
source_receiptsNo
correction_eventNo
promotion_reasonNo
include_contestedNo
progressive_levelNo
require_human_ackNo
time_window_hoursNo
organ_staleness_bandNo
supersedes_memory_idNo
minimised_vault_recordNo
required_floors_satisfiedNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations (destructiveHint true), the description adds that recall/inspect are free-ish and writes are gated mutations. It mentions authority and tiers. However, it does not explain what 'J-space mutations' entail, rate limits, or what exactly gets destroyed, leaving some ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (4 sentences) and front-loaded with the essential identifier. It includes a list of modes. However, the dense jargon and lack of structure (e.g., bullet points) slightly reduce readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (55 parameters, no schema descriptions), the description is incomplete. It omits explanations for most parameters, does not clarify parameter interactions, and assumes knowledge of the kernel stack. The output schema exists but does not compensate for the lack of usage guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 55 parameters and 0% schema description coverage, the description only explains the 'mode' parameter (lists values) and briefly hints at 'tier'. The vast majority of parameters (e.g., query, content, policy) are not addressed, leaving agents to guess their meaning and usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it is a memory governor for a kernel, managing L1–L6 tiers and listing modes (recall, inspect, etc.). It distinguishes itself from a free notepad and from siblings by focusing on structured memory operations. However, the heavy jargon (J-space, F1/F2/F4/F11) reduces immediate clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it (for structured memory, not ephemeral facts) and distinguishes read vs write modes. However, it lacks explicit guidance on when to use this tool vs sibling tools (e.g., arif_observe, arif_think) or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arif_observe111 Observe · Sense RealityA
Read-onlyIdempotent
Inspect

KERNEL 111 · Sense reality into evidence (not reasoning, not judgment). Web/URL/vitals/repo/entropy with epistemic tags. Authority: L0 OBSERVE. Modes: search | fetch | ingest | compass | atlas | entropy_dS | vitals | repo_map | hybrid_discovery. Returns: evidence + sources + uncertainty. Skip when pure reasoning → arif_think; domain compute → arif_route to GEOX/WEALTH/WELL.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNo
modeNosearch
queryNo
layersNo
actor_idNo
_envelopeNo
session_idNo
result_limitNo
session_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover safety (readOnly, idempotent, openWorld). Description adds behavioral detail: returns evidence, sources, uncertainty; authority level L0 OBSERVE; mentions epistemic tags. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is very concise with no wasted words. It front-loads the purpose and lists modes. However, the jargon 'KERNEL 111', 'L0 OBSERVE' may be opaque without additional context, slightly reducing accessibility.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 9 un-documented parameters and many modes, the description is far from complete. It mentions return types but fails to explain parameter roles or mode differences. Output schema exists but is not shown; description does not compensate for missing parameter semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% – the description does not explain any of the 9 parameters beyond listing some modes. It omits meanings for url, query, layers, actor_id, session_id, etc. This is a critical gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool gathers evidence from reality, with specific verb 'sense' and resource 'reality'. It distinguishes from siblings arif_think (pure reasoning) and arif_route (domain compute), making purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: for evidence gathering with modes like search, fetch, etc. Directly advises against use for pure reasoning ('skip when pure reasoning → arif_think') and domain compute ('arif_route to GEOX/WEALTH/WELL').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arif_route444 Route · Intent→OrganA
Read-onlyIdempotent
Inspect

KERNEL 444 · Intent→organ router (default path to GEOX/WEALTH/WELL/A-FORGE). Select when you know the goal but not which organ/verb. Optional organ_tool+arguments = governed bridge call (prefer this over arif_bridge_connect). Authority: L0. Returns: organ, port, tool_prefix, suggested_tools. Not session preflight (use arif_triage). Not a free shell.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskNoAlias for intent (backward compat).
organNoOptional explicit organ override. If provided, intent matching is skipped and this organ is used directly.
intentYesNatural-language description of what the user wants. e.g. "interpret this seismic section", "assess portfolio risk"
actor_idNoCalling actor.
_envelopeNo
argumentsNoArguments to pass to organ_tool.
organ_toolNoThe tool name on the target organ to call. If absent, returns routing decision only (no bridge call).
session_idNoGoverning session.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotation declares readOnlyHint=true, but the description mentions 'Optional organ_tool+arguments = governed bridge call', which implies potential execution of a non-read-only action. This contradiction undermines transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is concise, front-loaded with purpose, and each sentence adds value. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (8 params, output schema exists), the description covers purpose, usage, alternatives, return values, and constraints (L0 authority, not a free shell). It is sufficiently complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 88% (high), so baseline is 3. The description adds minimal parameter info beyond the schema; it mentions 'organ_tool+arguments' but does not elaborate on other parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it is an 'Intent→organ router' and explains the use case: route a goal to the appropriate organ and verb. It distinguishes from siblings by mentioning arif_bridge_connect and arif_triage, providing specific alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: 'Select when you know the goal but not which organ/verb.' Also gives clear when-not: 'Not session preflight (use arif_triage).' Recommends preferring this over arif_bridge_connect.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arif_seal999 Seal · VAULT999B
Destructive
Inspect

KERNEL 999 · VAULT999 immutable append — civilizational memory, irreversible. Authority: 888_HOLD / SOVEREIGN + ack_irreversible for seal mode. Modes: seal | verify | chain | list | dry_run | seal_card | render. Seal only after SEAL verdict path; HOLD/SABAR/VOID do not seal. Testing → dry_run. Kernel judges; vault seals; Arif owns F13 veto.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoseal
nonceNo
payloadNo
actor_idNo
_envelopeNo
session_idNo
drift_eventsNo
witness_typeNoai
session_tokenNo
actor_signatureNo
judge_state_hashNo
constitutional_chain_idNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description reinforces the annotation's destructiveHint=true by stating 'irreversible' and 'immutable append'. It adds behavioral context beyond annotations, such as authority requirements ('Authority: 888_HOLD / SOVEREIGN + ack_irreversible for seal mode') and a veto power ('Arif owns F13 veto'). There is no contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single paragraph of moderate length, front-loading the core concept of immutable append. It contains poetic language and jargon (e.g., 'civilizational memory'), which reduces conciseness. While it covers key points, it could be more efficiently structured for an AI agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (12 parameters, destructive, no param descriptions), the description leaves significant gaps. It does not explain the meaning or purpose of any parameter, nor describe the output (despite an existing output schema). The mode list discrepancy further undermines completeness. An agent would struggle to use this tool correctly without additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 12 parameters with 0% description coverage, and the description does not explain any parameter in detail. It lists modes in text ('seal | verify | chain | list | dry_run | seal_card | render') but the schema's mode enum includes different values ('seal', 'verify', 'ledger', 'changelog', 'audit'), creating inconsistency. The description adds minimal helpful parameter semantics and introduces confusion instead.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states that the tool performs an 'immutable append' or 'seal' operation, implying irreversible creation of a record. It mentions modes like 'seal' and 'verify', and gives a condition ('Seal only after SEAL verdict path'). However, the language is esoteric and uses proprietary terms (e.g., '888_HOLD', 'F13 veto') that may confuse an AI agent. The purpose is somewhat clear but not universally understandable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides implicit usage guidance: 'Seal only after SEAL verdict path' and 'Testing → dry_run', indicating when to use the seal mode versus dry run. It distinguishes the tool from siblings by stating 'Kernel judges; vault seals', implying arif_seal is for sealing after judging. However, it does not explicitly list when to use this tool over sibling tools like arif_judge or arif_think, limiting clarity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arif_think333 Think · MindA
Read-onlyIdempotent
Inspect

KERNEL 333 · Mind — structure reasoning under F2/F7 (not a chat model, not a verdict). Plan, reflect, verify, synthesize with OBS/DER/INT/SPEC labels. Authority: L0–L1. Modes: reason | reflect | verify | plan | plan_review | plan_approve | refactor_plan | metabolize | axioms. Returns: structured reasoning + confidence + next_safe_action. Ethical/maruah risk → arif_critique. Binding decision → arif_judge. Facts → arif_observe.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoreason
queryNo
plan_idNo
actor_idNo
_envelopeNo
session_idNo
witness_typeNoai
session_tokenNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint, idempotentHint, destructiveHint. Description adds that it operates at Authority L0–L1, returns structured reasoning + confidence + next_safe_action, and is not a chat model. This provides useful behavioral context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, packing purpose, modes, output, and cross-references into a few sentences. It front-loads the core identity. However, heavy jargon (F2/F7, OBS/DER/INT/SPEC) may reduce clarity for some agents, but it remains efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 8 parameters and an output schema, the description covers purpose, usage, and behavioral traits reasonably well. However, it does not explain most parameters or provide usage examples, leaving some gaps for a complex tool. The output schema exists but is not shared in the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%. The description explains the 'mode' parameter with its enum values, but does not describe other parameters like query, plan_id, actor_id, session_id, etc. Given 8 parameters, the description only partially compensates for the lack of schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it's for structure reasoning under F2/F7, explicitly distinguishes from a chat model and a verdict, and differentiates from siblings by pointing to arif_critique (ethical risk), arif_judge (binding decision), and arif_observe (facts).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says when not to use it ('not a chat model, not a verdict') and directs to sibling tools for related tasks. It lists modes but does not provide explicit guidance on when to choose each mode, though the mode names are self-explanatory.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 35 tool updatesv0.1.1
    • Removedarif_bridge
    • Changedarif_bridge_connect3 fields changed
      • addedOutput schema / additionalProperties
        Added value: +true
      • removedOutput schema / properties
        Removed value: -{
        -  "_nine_signal_compliant": {
        -    "description": "Internal compliance flag",
        -    "type": "boolean"
        -  },
        -  "_violations": {
        -    "description": "Non-compliance audit trail",
        -    "items": {
        -      "type": "string"
        -    },
        -    "type": "array"
        -  },
        -  "actor_id": {
        -    "description": "Sovereign or agent actor ID",
        -    "type": [
        -      "string",
        -      "null"
        -    ]
        -  },
        -  "delta_S": {
        -    "description": "Thermodynamic entropy change",
        -    "type": "number"
        -  },
        -  "meta": {
        -    "description": "Metadata including actor_id, mode, circuit",
        -    "type": "object"
        -  },
        -  "nine_signal": {
        -    "description": "F2 addendum nine-signal block",
        -    "type": "object"
        -  },
        -  "output_policy": {
        -    "description": "Policy constraints: DOMAIN_SEAL, DOMAIN_HOLD, DOMAIN_VOID, SIMULATION_ONLY",
        -    "type": "string"
        -  },
        -  "reasons": {
        -    "description": "Human-readable justification list",
        -    "items": {
        -      "type": "string"
        -    },
        -    "type": "array"
        -  },
        -  "result": {
        -    "description": "Tool-specific payload",
        -    "type": "object"
        -  },
        -  "session_id": {
        -    "description": "Active session identifier",
        -    "type": [
        -      "string",
        -      "null"
        -    ]
        -  },
        -  "stage_progression": {
        -    "description": "Next stage auto-chain hint",
        -    "type": [
        -      "object",
        -      "null"
        -    ]
        -  },
        -  "status": {
        -    "description": "Execution status: OK, ERROR, TIMEOUT, DRY_RUN",
        -    "type": "string"
        -  },
        -  "timestamp": {
        -    "description": "ISO-8601 timestamp",
        -    "type": "string"
        -  },
        -  "tool": {
        -    "description": "Canonical tool name that produced this response",
        -    "type": "string"
        -  },
        -  "verdict": {
        -    "description": "Constitutional verdict: SEAL, HOLD, VOID, SABAR, PROVISIONAL, PARTIAL",
        -    "type": "string"
        -  }
        -}
      • removedOutput schema / required
        Removed value: -[
        -  "status",
        -  "tool",
        -  "verdict",
        -  "result",
        -  "nine_signal",
        -  "reasons"
        -]
    • Addedarif_compose
    • Removedarif_conformance_report
    • Addedarif_critique
    • Removedarif_evidence_fetch
    • Addedarif_forge
    • Removedarif_forge_execute
    • Removedarif_gateway_connect
    • Removedarif_heart_critique
    • Addedarif_init
    • Removedarif_initialize_probe
    • Addedarif_judge
    • Removedarif_judge_deliberate
    • Removedarif_kernel_attest
    • Removedarif_kernel_health
    • Removedarif_kernel_route
    • Removedarif_kernel_status
    • Addedarif_memory
    • Removedarif_memory_recall
    • Removedarif_mind_reason
    • Addedarif_observe
    • Removedarif_ops_measure
    • Removedarif_ping
    • Removedarif_reply_compose
    • Changedarif_route3 fields changed
      • addedOutput schema / additionalProperties
        Added value: +true
      • removedOutput schema / properties
        Removed value: -{
        -  "_nine_signal_compliant": {
        -    "description": "Internal compliance flag",
        -    "type": "boolean"
        -  },
        -  "_violations": {
        -    "description": "Non-compliance audit trail",
        -    "items": {
        -      "type": "string"
        -    },
        -    "type": "array"
        -  },
        -  "actor_id": {
        -    "description": "Sovereign or agent actor ID",
        -    "type": [
        -      "string",
        -      "null"
        -    ]
        -  },
        -  "delta_S": {
        -    "description": "Thermodynamic entropy change",
        -    "type": "number"
        -  },
        -  "meta": {
        -    "description": "Metadata including actor_id, mode, circuit",
        -    "type": "object"
        -  },
        -  "nine_signal": {
        -    "description": "F2 addendum nine-signal block",
        -    "type": "object"
        -  },
        -  "output_policy": {
        -    "description": "Policy constraints: DOMAIN_SEAL, DOMAIN_HOLD, DOMAIN_VOID, SIMULATION_ONLY",
        -    "type": "string"
        -  },
        -  "reasons": {
        -    "description": "Human-readable justification list",
        -    "items": {
        -      "type": "string"
        -    },
        -    "type": "array"
        -  },
        -  "result": {
        -    "description": "Tool-specific payload",
        -    "type": "object"
        -  },
        -  "session_id": {
        -    "description": "Active session identifier",
        -    "type": [
        -      "string",
        -      "null"
        -    ]
        -  },
        -  "stage_progression": {
        -    "description": "Next stage auto-chain hint",
        -    "type": [
        -      "object",
        -      "null"
        -    ]
        -  },
        -  "status": {
        -    "description": "Execution status: OK, ERROR, TIMEOUT, DRY_RUN",
        -    "type": "string"
        -  },
        -  "timestamp": {
        -    "description": "ISO-8601 timestamp",
        -    "type": "string"
        -  },
        -  "tool": {
        -    "description": "Canonical tool name that produced this response",
        -    "type": "string"
        -  },
        -  "verdict": {
        -    "description": "Constitutional verdict: SEAL, HOLD, VOID, SABAR, PROVISIONAL, PARTIAL",
        -    "type": "string"
        -  }
        -}
      • removedOutput schema / required
        Removed value: -[
        -  "status",
        -  "tool",
        -  "verdict",
        -  "result",
        -  "nine_signal",
        -  "reasons"
        -]
    • Removedarif_schema_echo
    • Addedarif_seal
    • Removedarif_sense_observe
    • Removedarif_session_init
    • Addedarif_think
    • Removedarif_transport_echo
    • Removedarif_triage
    • Removedarif_vault_seal
    • Removedarif_version_echo
  2. 26 tool updatesv0.1.0
    • First observedarif_bridge
    • First observedarif_bridge_connect
    • First observedarif_conformance_report
    • First observedarif_evidence_fetch
    • First observedarif_forge_execute
    • First observedarif_gateway_connect
    • First observedarif_heart_critique
    • First observedarif_initialize_probe
    • First observedarif_judge_deliberate
    • First observedarif_kernel_attest
    • First observedarif_kernel_health
    • First observedarif_kernel_route
    • First observedarif_kernel_status
    • First observedarif_memory_recall
    • First observedarif_mind_reason
    • First observedarif_ops_measure
    • First observedarif_ping
    • First observedarif_reply_compose
    • First observedarif_route
    • First observedarif_schema_echo
    • First observedarif_sense_observe
    • First observedarif_session_init
    • First observedarif_transport_echo
    • First observedarif_triage
    • First observedarif_vault_seal
    • First observedarif_version_echo

TDQS

A3.6/5.0

Scored across 11 tools

Disambiguation5/5

Each tool has a distinct and clearly defined purpose within the kernel's lifecycle: from session initiation (arif_init) through observation, reasoning, ethical critique, judgment, sealing, composition, execution, and memory management. The only potential overlap between arif_route and arif_bridge_connect is explicitly resolved by preferring arif_route. All other tools are uniquely scoped with no ambiguity.

Naming Consistency4/5

All tools follow the 'arif_<verb>' pattern, with most verbs being single words (e.g., compose, judge, seal). The sole exception is 'arif_bridge_connect', which uses a compound verb. This minor inconsistency slightly detracts from the overall uniformity, but the pattern is otherwise consistent and predictable.

Tool Count5/5

With 11 tools, arifOS covers the full gamut of operations for a sophisticated governance kernel without being excessive. Each tool serves a clear and necessary role in the agent's workflow, from initialization to immutable sealing. The count feels well-scoped for the intended domain.

Completeness4/5

The tool set covers the major stages of the decision and execution pipeline: init, observe, think, critique, judge, seal, compose, forge, memory, route, and bridge. This provides a cohesive workflow. Missing are utility tools like listing or unsealing, but these are not central to the core lifecycle, so the gap is minor.

Maintenance

ActivityActive
ResponsivenessWithin a week

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Universal governance layer for AI agents — MCP-native, fail-closed, LNN interpretability. Governed receipts, IPFS audit proofs, and rollback for any agent in any framework.
    3
    79 npm
    Apache 2.0
  • A
    license
    Not graded
    quality
    F
    maintenance
    Constitutional MCP server enforcing 13 Floors of governance for AI agents, providing tools for session anchoring, reasoning, safety critique, and audit logging.
    AGPL 3.0