arifOS — Constitutional AI Kernel
Server Details
Constitutional AI kernel with 13 MCP tools, 888_JUDGE verdict pipeline, and VAULT999 ledger.
- Status
- Healthy
- Uptime
- 89.8% over 55 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- ariffazil/APEX
- GitHub Stars
- 0
TDQS
Scored across 8 tools
Each tool targets a clearly distinct kernel operation (session ignition, evidence gathering, reasoning, arbitration, execution, memory, routing, permanent sealing), and descriptions explicitly state boundaries (e.g., judge arbitrates but does not gather or mutate; forge is the only general mutation verb). Mode names overlap slightly across tools, but tool-level purposes are unambiguous.
All tools use a consistent arif_ prefix and snake_case, which is predictable. However, seven are verbs (forge, init, judge, observe, route, seal, think) while one is a noun (memory), a minor deviation from a uniform verb pattern.
Eight tools is well-scoped for a constitutional AI kernel. Each tool corresponds to a core lifecycle function (init, observe, think, judge, forge, memory, seal, route) and no tool feels redundant or missing.
The surface covers the full kernel lifecycle: session management, evidence collection, reasoning, arbitration, execution, memory tiers, routing, and immutable sealing. Intentional omissions (e.g., no unseal) are documented, and no obvious gaps exist for the stated purpose.
Available Tools
8 toolsarif_forge777 Forge · Execute GateADestructiveInspect
KERNEL 777 · Governed execution: applies a mutation through A-FORGE only when carrying a SEAL verdict (seal_verdict_id) from arif_judge and a live session. The manifest/query defines exactly what changes; mode=dry_run previews the plan without applying it. This is the only verb that executes general mutations — kernel memory writes go to arif_memory and permanent records to arif_seal.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | 'engineer' (default) — plan and apply the work; 'dry_run' — preview what would happen, nothing is applied; 'query' — read-only questions about the workspace; 'write' — write files; 'generate' — generate code or content; 'commit' — commit prepared work; 'recall' — retrieve past forge artifacts. | engineer |
| query | No | The task in plain language when no manifest is supplied, e.g. 'restart the gateway service'. | |
| plan_id | No | Plan id from arif_think's plan mode, when executing an approved plan. | |
| actor_id | No | Your agent identity, e.g. 'kimi-code/FI-008' — recorded in the audit log. | |
| manifest | No | The work order (usually JSON or markdown) describing exactly what to build or change — the more specific, the tighter the gate. | |
| _envelope | No | Internal transport envelope handle — server-side; omit. | |
| session_id | No | Session id returned by arif_init; scopes this execution to your governed session. | |
| arif_ack_id | No | Acknowledgment id from a prior step, for multi-step execution chains. | |
| artifact_id | No | Id of an existing artifact (from an earlier forge or judge response) that this call operates on. | |
| session_token | No | Session Continuity Token (SCT) returned by arif_init — proves the session is yours. | |
| vault_entry_id | No | Id of the permanent-ledger entry linked to this execution, when one already exists. | |
| seal_verdict_id | No | The SEAL verdict id arif_judge returned — without it, nothing is executed. | |
| ack_irreversible | No | Set true to confirm you accept this execution cannot be undone, when the judged action was irreversible. | |
| judge_state_hash | No | Hash string from the arif_judge SEAL response — proves the verdict has not been tampered with. | |
| approved_action_hash | No | Hash of the exact action that was judged — execution is refused if what you submit differs from it. | |
| constitutional_chain_id | No | Id of the evidence chain (observe→think→judge) that led to the SEAL — copy it from the judge response. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the description doesn't need to repeat that. It adds valuable context about the gating requirement (SEAL verdict and live session) and the dry_run safety mode, which goes beyond what annotations state. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single focused paragraph that front-loads the core purpose, then adds the dry_run note and sibling differentiation. It's concise without being under-specified, though a bulleted list could improve scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 16 parameters and no output schema, the description provides the essential context: it's the general mutation executor with a governance gate, and it routes to siblings for specialized writes. It doesn't explain every parameter, but the schema covers those. The description is sufficient to understand when and how to invoke it correctly, given the rich schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter in detail. The description adds some high-level context (the manifest/query defines changes, dry_run previews) but does not provide per-parameter semantics beyond what's in the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('applies a mutation through A-FORGE') and the exact resource (general mutations), and explicitly distinguishes it from siblings by naming arif_memory and arif_seal as the destinations for kernel memory writes and permanent records. It also names the required gate (SEAL verdict) making it unmistakable what this tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says this is the only verb that executes general mutations, and provides routing guidance: kernel memory writes go to arif_memory, permanent records to arif_seal. It also mentions dry_run mode for previewing without applying. This is clear when-to-use and when-not-to-use guidance with named alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arif_init000 Init · Kernel SessionAInspect
KERNEL 000 · Ignite a governed kernel session: binds actor identity, constitutional floors F1–F13, and the audit chain. Returns the session_id + session_token that every other arif_* verb requires. Use mode=preflight to inspect an existing session without re-igniting, mode=resume to continue one. Modes: init, preflight, resume, validate, canary, triage, epoch_open, epoch_seal, light, opt_out.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | What to do: 'init' (default) starts a new governed session; 'preflight' checks an existing session's state without creating one; 'resume' re-attaches to the session in session_id; 'validate' re-checks credentials; 'canary' is a transport probe; 'triage' reads session state; 'epoch_open'/'epoch_seal' bracket a long working window; 'light' is a minimal session; 'opt_out' records a privacy opt-out. | init |
| nonce | No | One-time random string (e.g. a UUID) making this request unique; protects against replay of the same call. | |
| intent | No | Plain-language statement of what this session is for, e.g. 'repair the failing MCP tests'. Recorded for audit. | |
| context | No | JSON object of background facts to bind into the session, e.g. {"repo": "arifOS", "task": "fix-tests"}. | |
| payload | No | Extra mode-specific data as a JSON object; the mode's response tells you which fields it expects. | |
| tooling | No | JSON list of the tools you can use, e.g. ["Bash", "Read"] — lets the kernel scope what it will permit you. | |
| verbose | No | Legacy on/off verbosity ('true'/'false'); prefer verbosity. | |
| actor_id | No | Your agent identity, e.g. 'kimi-code/FI-008' or 'claude/sonnet'. Recorded in the audit log so this action is attributed to you. | |
| epoch_id | No | Optional id grouping related sessions into one long-running epoch, e.g. '2026-H2-ops'. | |
| evidence | No | JSON list of facts you have already verified, e.g. [{"fact": "pytest 299 passed", "source": "CI run"}] — carried into the session record. | |
| trace_id | No | Correlation id of your choosing (e.g. 'req-8f3a') — lets you find this call later across kernel logs and organ systems. | |
| _envelope | No | Reserved for the transport layer — never fill this in. | |
| verbosity | No | How detailed responses should be: 'minimal' (default), 'standard', or 'full'. | minimal |
| session_id | No | Session id returned by arif_init (looks like 'sess-…'). Pass it on every call after init so the action is attributed to your governed session. | |
| agent_policy | No | JSON constraints on your own behavior, e.g. {"autonomy": "reversible_only", "forbidden": ["git push"]}. | |
| auth_context | No | JSON carrying external authentication material, e.g. {"token_class": "bearer", "issuer": "github"} — used for gating. | |
| counterparty | No | JSON describing the other party in a two-agent exchange, e.g. {"agent_id": "hermes/1", "role": "verifier"}. | |
| sovereign_id | No | Identifier of the human principal you act for, when acting under their explicit delegation. | |
| session_token | No | Session Continuity Token (SCT) — the credential string arif_init returned alongside session_id. Pass it back on follow-up calls to prove the session is yours; without it the kernel treats you as unauthenticated. | |
| actor_signature | No | Cryptographic signature over the request, if your agent holds a key — proves the call genuinely came from actor_id. Omit if you have no key. | |
| caller_actor_id | No | If you are calling on behalf of another agent, that agent's id — builds a delegation chain for the audit log. | |
| delegation_mode | No | Contract governing delegated calls, e.g. 'read_only' or 'governed'. | |
| idempotency_key | No | Client-chosen key (e.g. 'job-42-attempt-1'). Retrying with the same key will not repeat the effect — safe retries on flaky networks. | |
| ack_irreversible | No | Set true to acknowledge that session records are permanent audit artifacts and cannot be deleted afterwards. | |
| executor_actor_id | No | Id of the agent that will actually carry out work under this session's permissions, if different from actor_id. | |
| declared_model_key | No | The model you run on, e.g. 'zai-coding-plan/glm-5.3'. Informational only — the kernel records but never trusts it. | |
| client_capabilities | No | JSON declaring what your client supports, e.g. {"transports": ["http"], "protocol": "2025-11-25"}. | |
| requested_authority | No | The highest class of action you may take: 'OBSERVE_ONLY' (read-only, default) or a higher governed class granted by your policy. | OBSERVE_ONLY |
| previous_session_hash | No | Hash string returned when your previous session closed; supplying it chains this session to that one, proving continuity. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, so the description is consistent with a write operation. It adds context about binding actor identity, constitutional floors, and the audit chain, and notes the return of a session token. However, it does not explicitly disclose that session records are permanent and irreversible, which is only mentioned in the ack_irreversible parameter schema, not the main description. No contradiction with annotations, but the description could be more explicit about the durability of created sessions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences: it states the core purpose, the return value, and mode guidance, then lists all modes. It is front-loaded with the most important information, contains no fluff, and is well-structured for an agent to quickly grasp the tool's role and when to use different modes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (29 parameters, no output schema), the description covers the essential context: it explains the purpose, the return credentials, and the existence of modes for preflight/resume. It does not explicitly mention the ack_irreversible requirement or the exact return format, but the parameter schema and the mention of session_id/session_token cover the critical aspects. It is complete enough for an agent to call it correctly, though slightly more detail on side effects would improve it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers all 29 parameters with detailed descriptions, including the mode enum, so the description need not add parameter-level detail. The description's mention of preflight and resume modes aligns with the schema, but adds no new semantics beyond what the schema already provides. Baseline of 3 is appropriate for 100% schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Ignite a governed kernel session' and specifies what it binds (actor identity, constitutional floors, audit chain) and what it returns (session_id + session_token). It also implicitly differentiates itself from siblings by noting every other arif_* verb requires these credentials, making it the entry point.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit mode guidance ('Use mode=preflight to inspect an existing session without re-igniting, mode=resume to continue one') and states that every other arif_* verb requires the returned session credentials, which tells an agent this is the first call to make. It doesn't explicitly list when not to use it or name alternative tools, but the entry-point role is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arif_judge888 Judge · VerdictRead-onlyInspect
KERNEL 666 · Binding constitutional arbitration: evaluates a candidate action or claim and returns SEAL / HOLD / SABAR / VOID with the full reason chain. Weighs action class, blast radius, reversibility, entropy pathway, and cooling state; a SEAL verdict here is what arif_forge requires before it will execute anything. Judge arbitrates — it does not gather (evidence comes via arif_observe, reasoning via arif_think) and it does not mutate (execution is arif_forge, permanence is arif_seal). Modes: judge, intercept, validate, hold, escalate.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | 'judge' (default) — render a verdict on candidate; 'intercept' — pre-flight gate before an action runs; 'validate' — re-check a prior verdict; 'hold' — place an action into a cooling period; 'escalate' — refer the decision to the human owner. | judge |
| nonce | No | One-time random string (e.g. a UUID) making this request unique; protects against replay of the same call. | |
| domain | No | Subject area for claim evaluation, e.g. 'geoscience' or 'finance'. | |
| key_id | No | Which of your registered keys produced actor_signature, e.g. 'key-1'. | |
| actor_B | No | Your calibrated confidence for a prediction, 0.0–1.0 (Brier-style score component). | |
| actor_id | No | Your agent identity, e.g. 'kimi-code/FI-008' or 'claude/sonnet'. Recorded in the audit log so this action is attributed to you. | |
| evidence | No | The facts the verdict should rest on — a JSON list like [{"fact": "tests pass", "source": "CI"}], or an object. | |
| _envelope | No | Reserved for the transport layer — never fill this in. | |
| actor_Phi | No | JSON map of supporting signals about you, e.g. {"consistency": 0.9, "track_record": 0.7}. | |
| candidate | No | The action or claim being judged, stated plainly, e.g. 'delete table users in prod'. | |
| claim_text | No | The exact claim sentence under judgment, e.g. 'uptime exceeded 99% in August'. | |
| session_id | No | Session id returned by arif_init (looks like 'sess-…'). Pass it on every call after init so the action is attributed to your governed session. | |
| action_tier | No | How risky the action is: 'standard', 'high', or 'critical' — higher tiers demand stronger evidence. | standard |
| claim_class | No | Epistemic strength of the claim: 'OBS' (directly observed), 'DER' (derived from observations), 'INT' (interpreted), 'SPEC' (speculative), 'UNKNOWN' (cannot witness — honest null; maps to TruthClass.UNK, weight 0.30, cannot authorize mutation). RULE absence-null: 'not in my knowledge' is NOT evidence a claim is false — label UNKNOWN, never FALSE. | |
| claim_tense | No | ||
| niat_params | No | JSON intent-calibration settings, e.g. {"sincerity": 0.8, "stated_goal": "verify the claim"}. | |
| action_class | No | Kind of action under judgment: 'OBSERVE' (read-only), 'DRAFT' (compose text), 'MUTATE' (change state), 'IRREVERSIBLE' (permanent). | |
| blast_radius | No | Who or what the action can affect: 'self', 'session', 'organ', or 'federation'. | |
| seal_purpose | No | One sentence saying why this verdict/record must exist, e.g. 'closing deployment D-17'. Stored permanently alongside the record. | |
| session_token | No | Session Continuity Token (SCT) — the credential string arif_init returned alongside session_id. Pass it back on follow-up calls to prove the session is yours; without it the kernel treats you as unauthenticated. | |
| context_source | No | Where the candidate came from: 'session', 'file', or 'memory'. | |
| heart_critique | No | JSON result of an ethics/dignity check, e.g. {"coercion": false, "dignity": 0.9} — feeds the verdict. | |
| vault_entry_id | No | Id of the permanent-ledger entry this verdict should attach to (from an earlier arif_seal response). | |
| actor_signature | No | Cryptographic signature over the request, if your agent holds a key — proves the call genuinely came from actor_id. Omit if you have no key. | |
| authority_token | No | ||
| entropy_pathway | No | How the action changes system order: 'reduces', 'neutral', or 'increases' complexity. | |
| entropy_receipt | No | JSON receipt from an entropy computation, bound to this judgment. | |
| authority_effect | No | What permission a SEAL verdict would grant, e.g. 'execute forge plan P-9'. | |
| cooldown_entry_id | No | Id of an existing cooling-period record to consult before this action may proceed. | |
| sovereign_receipt | No | Reference to the human owner's explicit approval, for the rare case where they have already decided directly. | |
| reversibility_level | No | How hard the action would be to undo: 'reversible', 'hard', or 'irreversible'. | |
| requested_capability | No | The specific capability being requested, e.g. 'forge.execute' — checked against the capability registry. | |
| constitutional_chain_id | No | Id of the evidence chain this call belongs to (observe→think→judge→seal). Copy it from the earlier step's response to link the steps together. |
arif_memoryMemory Governor · KernelADestructiveInspect
KERNEL 555 · Governed memory of the kernel itself: six tiers (L1–L6) with per-mode gating. recall, inspect, and audit are read paths; remember, revise, promote, and forget mutate tiers — promote and forget additionally require human_approval=true. Use for cross-session lessons, canon, and memory audit; external evidence belongs to arif_observe and reasoning artifacts to arif_think. Modes: recall, inspect, attest, remember, promote, revise, forget, audit, metabolize.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | 'recall' (default) — semantic search of stored memories; 'inspect' — read one memory in full; 'attest' — vouch for a memory's accuracy; 'remember' — store new text; 'promote' — move a memory up a tier; 'revise' — replace a memory's text; 'forget' — remove a memory (gated); 'audit' — integrity scan; 'metabolize' — compact and consolidate tiers. | recall |
| tier | No | Which memory tier to target, 'L1'–'L6' (L1 = hot working memory, L6 = sealed canon). | |
| query | No | What to search for, in plain language (mode=recall/audit), e.g. 'past deploy rollback steps'. | |
| content | No | The text to store (mode=remember) — write it as a self-contained lesson or fact. | |
| payload | No | Extra mode-specific fields as JSON — e.g. {"truth_class": "DERIVED", "provenance": "CI log"} for remember. | |
| to_tier | No | Destination tier when promoting, e.g. 'L4'. | |
| actor_id | No | Your agent identity, e.g. 'kimi-code/FI-008' — recorded in the audit log. | |
| lease_id | No | Id of a short-lived permission grant (lease) authorizing this write, when one was issued to you. | |
| trace_id | No | Correlation id of your choosing (e.g. 'req-8f3a') to find this call later in the logs. | |
| memory_id | No | UUID of the memory entry to inspect/revise/forget — it was returned when the memory was created or last listed. | |
| session_id | No | Session id returned by arif_init; attributes this call to your governed session. | |
| new_content | No | Replacement text for the memory (mode=revise) — must refer to the same memory_id. | |
| session_token | No | Session Continuity Token (SCT) returned by arif_init — proves the session is yours. | |
| human_approval | No | Set true ONLY when the human owner explicitly approved this promote/forget — the gate refuses without it. | |
| idempotency_key | No | Client-chosen key (e.g. 'memo-42'); retries with the same key will not store duplicates. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark destructiveHint=true RF, but the description adds valuable specifics: it categorizes read paths ('recall, inspect, and audit') vs. mutating modes ('remember, revise, promote, and forget') and highlights the human_approval=true gate for promote/forget. This goes beyond annotations to explain which operations are destructive and what guardrails apply. It does not contradict annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core concept ('Governed memory of the kernel itself'), then the next sentence covers reading vs. mutation, then usage guidance, then a mode list (which is somewhat redundant with the schema but serves as a quick reference). It earns its length with no fluff, though the 'KERNEL 555' preamble adds a little flavor without substance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 15 optional parameters and no output schema, the description gives the essential context: purpose, modes, tier system, gating, and sibling boundaries. It explains the tool's role in the environment and how to use it, while the schema handles parameter details. Some specifics like return-value behavior are absent but not required without an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for all 15 parameters, so the schema already documents each parameter. The description does not add new parameter semantics beyond noting that promote/forget require human_approval=true, which is also in the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('governed memory') and resource ('kernel itself'), defines six tiers (L1–L6), and explicitly differentiates from siblings: 'external evidence belongs to arif_observe and reasoning artifacts to arif_think.' An agent can immediately tell this is the memory-management tool, not the observe or think tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use for cross-session lessons, canon, and memory audit' gives positive use cases, while 'external evidence belongs to arif_observe and reasoning artifacts to arif_think' provides clear exclusions and routes to alternatives. This is explicit when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arif_observe111 Observe · Sense RealityARead-onlyIdempotentInspect
KERNEL 111 · Collect evidence — facts and sources with epistemic tags (OBS) and uncertainty bounds, never conclusions. mode=search queries the open web/literature; mode=fetch retrieves a URL and records its provenance; mode=vitals reads kernel machine telemetry. Reason over what you gathered with arif_think; delegate domain analysis to an organ with arif_route. Modes: search, fetch, hybrid_discovery, ingest, compass, atlas, entropy_dS, vitals.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | Exact URL to retrieve when mode=fetch, e.g. 'https://example.com/report.pdf'. | |
| mode | No | 'search' (default) — web/literature query; 'fetch' — retrieve one URL with provenance; 'hybrid_discovery' — combine sources; 'ingest' — absorb a document; 'compass'/'atlas' — guided navigation; 'entropy_dS' — measure system change; 'vitals' — kernel machine telemetry. | search |
| query | No | What to look for, in plain language, e.g. 'TDQS scoring rubric MCP'. | |
| layers | No | Restrict where to look, e.g. ["web"] or ["canon", "memory"]; omit to search everything available. | |
| actor_id | No | Your agent identity, e.g. 'kimi-code/FI-008' or 'claude/sonnet'. Recorded in the audit log so this action is attributed to you. | |
| _envelope | No | Reserved for the transport layer — never fill this in. | |
| session_id | No | Session id returned by arif_init (looks like 'sess-…'). Pass it on every call after init so the action is attributed to your governed session. | |
| result_limit | No | Maximum number of results to return in search modes; default 10. | |
| session_token | No | Session Continuity Token (SCT) — the credential string arif_init returned alongside session_id. Pass it back on follow-up calls to prove the session is yours; without it the kernel treats you as unauthenticated. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare a safe read-only, idempotent, open-world profile, and the description is consistent with them. Beyond that, it adds a distinctive behavioral contract — 'never conclusions' — telling the agent the tool deliberately withholds conclusions and only emits evidence with epistemic tags and uncertainty bounds, plus provenance recording for fetch. This shapes how an agent must interpret results, which exceeds the annotation coverage; a 5 would require output-format or error-behavior disclosure it doesn't provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well front-loaded: the first sentence delivers the core purpose, and the mode explanations are compressed into one efficient sentence. Punished for the decorative 'KERNEL 111 ·' prefix and the trailing 'Modes: search, fetch, hybrid_discovery, ingest, compass, atlas, entropy_dS, vitals.' enumeration, which largely duplicates the preceding sentence and the schema enum.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For the complexity — 9 parameters, 8 modes, no output schema — the description is only partially complete. It covers the epistemic contract and three modes, while five modes are unexplained in the description (the schema's mode parameter partially compensates by glossing each). With no output schema, the agent must rely on the abstract 'facts and sources with epistemic tags (OBS) and uncertainty bounds' for the return contract; error behavior, result structure, and cross-mode interplay are absent. The rich schema props it up to minimum viable but no higher.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3 per rubric. The description adds value above that baseline by explaining the semantics of three modes in prose (search, fetch, vitals) and framing result output as evidence-with-tags, enriching the mode enum. The other five enum modes (hybrid_discovery, ingest, compass, atlas, entropy_dS) get their semantics only from the schema, so it does not reach 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource — 'Collect evidence — facts and sources with epistemic tags (OBS) and uncertainty bounds, never conclusions' — and differentiates from siblings by explicitly routing downstream work to arif_think and arif_route. It also names three concrete modes (search, fetch, vitals) with behaviors. However, the multi-mode sprawl of 8 modes (compass, atlas, entropy_dS left unexplained) and the poetic 'Observe · Sense Reality' title blur the central purpose slightly, preventing a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit division-of-labor guidance: 'Reason over what you gathered with arif_think; delegate domain analysis to an organ with arif_route,' and ties mode selection to task ('mode=search queries the open web/literature; mode=fetch retrieves a URL...; mode=vitals reads kernel machine telemetry'). Lacks when-not-to-use guidance for the remaining siblings (arif_memory, arif_judge, arif_seal) and for 5 of the 8 modes, so it falls short of fully explicit exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arif_route444 Route · Intent→OrganARead-onlyIdempotentInspect
KERNEL 444 · Intent→organ router: classifies a natural-language intent and dispatches it to the specialist organ (GEOX geoscience, WEALTH capital, WELL vitality, A-FORGE execution). Returns the routing decision only — no organ call — unless organ_tool names the target tool and arguments carries its inputs. Prefer this over guessing organs yourself; use arif_think for reasoning you keep in-kernel.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | route | |
| task | No | Alias for intent (backward compat). | |
| organ | No | Optional explicit organ override. If provided, intent matching is skipped and this organ is used directly. | |
| intent | No | Natural-language description of what the user wants. e.g. "interpret this seismic section", "assess portfolio risk" | |
| actor_id | No | Your agent identity, e.g. 'kimi-code/FI-008' or 'claude/sonnet'. Recorded in the audit log so this action is attributed to you. | |
| _envelope | No | Reserved for the transport layer — never fill this in. | |
| arguments | No | Arguments to pass to organ_tool. | |
| mission_id | No | Explicit human-cockpit mission binding (investigate|interpret| decide|build|monitor|remember). When set, skips keyword classification and binds the six-mission plan. Preferred over free-text when the agent already knows the mission. | |
| organ_tool | No | The tool name on the target organ to call. If absent, returns routing decision only (no bridge call). | |
| session_id | No | Session id returned by arif_init (looks like 'sess-…'). Pass it on every call after init so the action is attributed to your governed session. | |
| session_token | No | Session Continuity Token (SCT) — the credential string arif_init returned alongside session_id. Pass it back on follow-up calls to prove the session is yours; without it the kernel treats you as unauthenticated. | |
| contract_c_kwargs | No | Extra keyword arguments passed through to the organ tool call, as a JSON object. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description usefully discloses that the default is a routing decision with no downstream call, and that organ_tool + arguments triggers an actual bridge call. However, that bridge call can invoke an arbitrary organ tool whose effects are unknown, which sits uneasily against the readOnlyHint=true annotation — the definition is the only place this write-capable passthrough is surfaced, and it is not reconciled with the read-only annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the routing purpose, then the default behavior, then the alternative routing to arif_think. No filler or restated name/title content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter router with no output schema and high schema coverage, the description covers the essential decision logic (route vs. bridge) and the sibling alternative. Auth/session prerequisites are left entirely to the schema field descriptions, which is acceptable given their coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 92%, so the baseline is 3, but the description still adds semantic value by explaining the organ_tool/arguments interplay and the two-mode dispatch that the schema documents only per-field. It adds the mode-switching meaning rather than repeating field docs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (classifies) and resource (natural-language intent) and names the dispatch targets (GEOX, WEALTH, WELL, A-FORGE). It also explicitly distinguishes itself from a sibling ('use arif_think for reasoning you keep in-kernel'), so an agent can tell it apart without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear directive ('prefer this over guessing organs yourself') and names the alternative sibling (arif_think) with the condition that selects it. It also explains the two operating modes (route-only vs. bridge via organ_tool), though it stops short of enumerating when-not-to-use scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arif_seal999 Seal · VAULT999ADestructiveInspect
KERNEL 999 · Append an entry to VAULT999, the immutable ledger — accepted entries can never be edited or removed; there is no unseal. Use for permanent records of verified outcomes, lessons, and session closure once a verdict exists. ack_irreversible=true is the explicit acknowledgment of permanence, and judge_state_hash binds the entry to the verdict that authorized it. Reversible changes belong in arif_forge under a SEAL. Modes: seal, verify, ledger, changelog, audit, session_close.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | 'seal' (default) — append a permanent entry; 'verify' — check one entry; 'ledger' — read the ledger head; 'changelog' — recent appends; 'audit' — integrity check; 'session_close' — close out a session into the ledger. | seal |
| nonce | No | One-time random string (e.g. a UUID) making this request unique; protects against replay of the same call. | |
| payload | No | The content to store forever, as a string (usually JSON-serialized) — the outcome, lesson, or record itself. | |
| actor_id | No | Your agent identity, e.g. 'kimi-code/FI-008' or 'claude/sonnet'. Recorded in the audit log so this action is attributed to you. | |
| _envelope | No | Reserved for the transport layer — never fill this in. | |
| session_id | No | Session id returned by arif_init (looks like 'sess-…'). Pass it on every call after init so the action is attributed to your governed session. | |
| drift_events | No | JSON list of deviations observed, e.g. [{"what": "schema drift", "where": "tools/list"}] — stored with the record. | |
| seal_purpose | No | One sentence saying why this verdict/record must exist, e.g. 'closing deployment D-17'. Stored permanently alongside the record. | |
| witness_type | No | Who witnessed the sealed fact: 'ai', 'human', or 'external' system. | ai |
| session_token | No | Session Continuity Token (SCT) — the credential string arif_init returned alongside session_id. Pass it back on follow-up calls to prove the session is yours; without it the kernel treats you as unauthenticated. | |
| constitutional | No | JSON block of governance metadata (floors consulted, chain references) — normally built by the kernel, not by callers. | |
| actor_signature | No | Cryptographic signature over the request, if your agent holds a key — proves the call genuinely came from actor_id. Omit if you have no key. | |
| ack_irreversible | No | Set true to confirm you understand sealed entries are PERMANENT — they can never be edited or removed. | |
| judge_state_hash | No | Hash string from the arif_judge SEAL response — ties this entry to the verdict that authorized it. | |
| constitutional_chain_id | No | Id of the evidence chain this call belongs to (observe→think→judge→seal). Copy it from the earlier step's response to link the steps together. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses the critical irreversible behavior: 'accepted entries can never be edited or removed; there is no unseal.' It explains the significance of ack_irreversible and judge_state_hash, going beyond the generic destructiveHint annotation. No contradiction with annotations; it enriches them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense yet efficiently structured. It front-loads the core purpose, then gives usage, key parameter explanations, and alternatives, all in about five sentences. Every sentence earns its place, and the format is scannable for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 15 parameters and no output schema, the description covers the essential context: core purpose, irreversibility, usage timing, alternatives, and modes. It references session attribution indirectly through parameter hints but the schema covers that. The description is complete enough for an agent to decide when and how to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides 100% description coverage, so the baseline is 3. The description adds value by explaining ack_irreversible as 'explicit acknowledgment of permanence' and judge_state_hash as 'binds the entry to the verdict that authorized it', giving semantic weight to these parameters. It also lists modes but the schema already details them, so it's a moderate addition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Append an entry to VAULT999, the immutable ledger'. It clearly states the tool's core action and domain, and distinguishes it from siblings by noting 'Reversible changes belong in arif_forge under a SEAL'. The purpose is explicit and actionable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use: 'Use for permanent records of verified outcomes, lessons, and session closure once a verdict exists.' It also names the alternative tool for reversible changes and provides a clear exclusion. The mention of modes further clarifies usage scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
arif_think333 Think · MindARead-onlyIdempotentInspect
KERNEL 333 · Structured reasoning pass: decomposes a query and returns reasoning steps labeled OBS (observed), DER (derived), INT (interpretation), SPEC (specification) under truth floors. Produces reasoning records only — no verdicts (those come from arif_judge) and no state changes. The plan-family modes draft/review/approve execution plans; simulate and wonder explore counterfactuals. Modes: reason, reflect, verify, axioms, plan, plan_review, plan_approve, refactor_plan, metabolize, simulate, wonder, atlas.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | 'reason' (default) — decompose a question; 'reflect' — self-review of prior reasoning; 'verify' — check a derivation; 'axioms' — surface hidden assumptions; 'plan'/'plan_review'/'plan_approve'/'refactor_plan' — execution-plan lifecycle; 'metabolize' — consolidate past reasoning; 'simulate' — what-if; 'wonder' — open exploration; 'atlas' — map the problem space. | reason |
| query | No | The question, claim, or problem to reason about, in plain language. | |
| plan_id | No | Plan reference from an earlier plan-mode response — needed for the review/approve/refactor steps. | |
| actor_id | No | Your agent identity, e.g. 'kimi-code/FI-008' or 'claude/sonnet'. Recorded in the audit log so this action is attributed to you. | |
| _envelope | No | Reserved for the transport layer — never fill this in. | |
| session_id | No | Session id returned by arif_init (looks like 'sess-…'). Pass it on every call after init so the action is attributed to your governed session. | |
| witness_type | No | Who vouches for the reasoning record: 'ai' (default), 'human', or 'external' system. | ai |
| session_token | No | Session Continuity Token (SCT) — the credential string arif_init returned alongside session_id. Pass it back on follow-up calls to prove the session is yours; without it the kernel treats you as unauthenticated. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the description adds value by specifying 'Produces reasoning records only — no verdicts and no state changes,' which aligns with the annotations. It also mentions 'under truth floors' as a behavioral constraint. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph that front-loads the core purpose and then lists modes. It is efficient and avoids redundancy, though a bulleted mode list might improve scannability. Every sentence contributes to understanding the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (8 parameters, 12 modes) and the rich parameter descriptions in the schema, the description provides a solid high-level overview and distinguishes the tool from arif_judge. It does not explain the exact output format or the meaning of 'truth floors,' but the schema covers parameter details and no output schema is expected. It is adequate for an agent to select and invoke the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already explains each parameter. The description adds marginal value by grouping modes (e.g., plan-family, simulate/wonder) but largely restates information already present in the schema's mode enum. It does not compensate for any coverage gap because there is none.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('decomposes a query'), a clear resource ('structured reasoning pass'), and distinguishes its output (reasoning records labeled OBS/DER/INT/SPEC) from verdicts (which arif_judge provides). It also enumerates the modes, making the tool's purpose unambiguous and differentiating it from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly notes that verdicts come from arif_judge, providing an alternative, and groups plan-family modes and exploratory modes. However, it does not give explicit 'when to use this vs. that' conditions for all siblings, leaving some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
- Changed
arif_judge1 field changed- added
Input schema / properties / claim_tenseAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null +}
Publisher details
- Operator
- ARIF FAZIL · Publisher source
- Operator website
- https://arif-fazil.com
- Vendor relationship
- First-party
- Documentation
- Unknown
- Trust center
- Unknown
- Restrictions
- Unknown
Related MCP Connectors
Constitutional AI governance: 11 mega-tools, 13 floors, VAULT999 ledger. Human-in-loop by design.
AI Reasoning Cache & Consensus Layer with 11 MCP tools via Streamable HTTP.
Shared long-term memory vault for AI agents with 20 MCP tools.
Free OpenAI-compatible inference with signed provenance receipts and 3 focused MCP tools.
Related MCP Servers
- AlicenseAqualityAmaintenanceProvides a constitutional governance framework for AI agents, offering 13 MCP tools for session initiation, reasoning, evidence fetching, judgment, and execution, all governed by hard invariant laws and a hierarchical set of constitutional floors.11133 PyPI53AGPL 3.0
- AlicenseNot gradedqualityFmaintenanceConstitutional MCP server enforcing 13 Floors of governance for AI agents, providing tools for session anchoring, reasoning, safety critique, and audit logging.AGPL 3.0
- AlicenseNot gradedqualityCmaintenanceA constitutional MCP server that enforces external authority, loud rejections, and an unforgeable hash-chained ledger, preventing AI agents from minting their own identity, approving themselves, or rewriting history.2 npmApache 2.0
- AlicenseNot gradedqualityBmaintenanceRuntime constitutional verification for AI answers — claim extraction with reasoning chains, Epistemic Confidence Score (ECS), 7-angle Glassbox Court red team, constitution compilation, Trust Card assembly, and deterministic SHA-256 audit logs.11Apache 2.0
Glama MCP Gateway
Add one secure layer between your agents and this server.