Skip to main content
Glama

arifOS — Constitutional AI Kernel

Server Details

Constitutional AI kernel with 13 MCP tools, 888_JUDGE verdict pipeline, and VAULT999 ledger.

Ownership verified
Status
Healthy
Uptime
89.8% over 55 days
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
ariffazil/APEX
GitHub Stars
0

TDQS

A4.3/5.0

Scored across 8 tools

Disambiguation5/5

Each tool targets a clearly distinct kernel operation (session ignition, evidence gathering, reasoning, arbitration, execution, memory, routing, permanent sealing), and descriptions explicitly state boundaries (e.g., judge arbitrates but does not gather or mutate; forge is the only general mutation verb). Mode names overlap slightly across tools, but tool-level purposes are unambiguous.

Naming Consistency4/5

All tools use a consistent arif_ prefix and snake_case, which is predictable. However, seven are verbs (forge, init, judge, observe, route, seal, think) while one is a noun (memory), a minor deviation from a uniform verb pattern.

Tool Count5/5

Eight tools is well-scoped for a constitutional AI kernel. Each tool corresponds to a core lifecycle function (init, observe, think, judge, forge, memory, seal, route) and no tool feels redundant or missing.

Completeness5/5

The surface covers the full kernel lifecycle: session management, evidence collection, reasoning, arbitration, execution, memory tiers, routing, and immutable sealing. Intentional omissions (e.g., no unseal) are documented, and no obvious gaps exist for the stated purpose.

Available Tools

8 tools
arif_forge777 Forge · Execute GateA
Destructive
Inspect

KERNEL 777 · Governed execution: applies a mutation through A-FORGE only when carrying a SEAL verdict (seal_verdict_id) from arif_judge and a live session. The manifest/query defines exactly what changes; mode=dry_run previews the plan without applying it. This is the only verb that executes general mutations — kernel memory writes go to arif_memory and permanent records to arif_seal.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo'engineer' (default) — plan and apply the work; 'dry_run' — preview what would happen, nothing is applied; 'query' — read-only questions about the workspace; 'write' — write files; 'generate' — generate code or content; 'commit' — commit prepared work; 'recall' — retrieve past forge artifacts.engineer
queryNoThe task in plain language when no manifest is supplied, e.g. 'restart the gateway service'.
plan_idNoPlan id from arif_think's plan mode, when executing an approved plan.
actor_idNoYour agent identity, e.g. 'kimi-code/FI-008' — recorded in the audit log.
manifestNoThe work order (usually JSON or markdown) describing exactly what to build or change — the more specific, the tighter the gate.
_envelopeNoInternal transport envelope handle — server-side; omit.
session_idNoSession id returned by arif_init; scopes this execution to your governed session.
arif_ack_idNoAcknowledgment id from a prior step, for multi-step execution chains.
artifact_idNoId of an existing artifact (from an earlier forge or judge response) that this call operates on.
session_tokenNoSession Continuity Token (SCT) returned by arif_init — proves the session is yours.
vault_entry_idNoId of the permanent-ledger entry linked to this execution, when one already exists.
seal_verdict_idNoThe SEAL verdict id arif_judge returned — without it, nothing is executed.
ack_irreversibleNoSet true to confirm you accept this execution cannot be undone, when the judged action was irreversible.
judge_state_hashNoHash string from the arif_judge SEAL response — proves the verdict has not been tampered with.
approved_action_hashNoHash of the exact action that was judged — execution is refused if what you submit differs from it.
constitutional_chain_idNoId of the evidence chain (observe→think→judge) that led to the SEAL — copy it from the judge response.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, so the description doesn't need to repeat that. It adds valuable context about the gating requirement (SEAL verdict and live session) and the dry_run safety mode, which goes beyond what annotations state. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single focused paragraph that front-loads the core purpose, then adds the dry_run note and sibling differentiation. It's concise without being under-specified, though a bulleted list could improve scannability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 16 parameters and no output schema, the description provides the essential context: it's the general mutation executor with a governance gate, and it routes to siblings for specialized writes. It doesn't explain every parameter, but the schema covers those. The description is sufficient to understand when and how to invoke it correctly, given the rich schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every parameter in detail. The description adds some high-level context (the manifest/query defines changes, dry_run previews) but does not provide per-parameter semantics beyond what's in the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('applies a mutation through A-FORGE') and the exact resource (general mutations), and explicitly distinguishes it from siblings by naming arif_memory and arif_seal as the destinations for kernel memory writes and permanent records. It also names the required gate (SEAL verdict) making it unmistakable what this tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says this is the only verb that executes general mutations, and provides routing guidance: kernel memory writes go to arif_memory, permanent records to arif_seal. It also mentions dry_run mode for previewing without applying. This is clear when-to-use and when-not-to-use guidance with named alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arif_init000 Init · Kernel SessionAInspect

KERNEL 000 · Ignite a governed kernel session: binds actor identity, constitutional floors F1–F13, and the audit chain. Returns the session_id + session_token that every other arif_* verb requires. Use mode=preflight to inspect an existing session without re-igniting, mode=resume to continue one. Modes: init, preflight, resume, validate, canary, triage, epoch_open, epoch_seal, light, opt_out.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoWhat to do: 'init' (default) starts a new governed session; 'preflight' checks an existing session's state without creating one; 'resume' re-attaches to the session in session_id; 'validate' re-checks credentials; 'canary' is a transport probe; 'triage' reads session state; 'epoch_open'/'epoch_seal' bracket a long working window; 'light' is a minimal session; 'opt_out' records a privacy opt-out.init
nonceNoOne-time random string (e.g. a UUID) making this request unique; protects against replay of the same call.
intentNoPlain-language statement of what this session is for, e.g. 'repair the failing MCP tests'. Recorded for audit.
contextNoJSON object of background facts to bind into the session, e.g. {"repo": "arifOS", "task": "fix-tests"}.
payloadNoExtra mode-specific data as a JSON object; the mode's response tells you which fields it expects.
toolingNoJSON list of the tools you can use, e.g. ["Bash", "Read"] — lets the kernel scope what it will permit you.
verboseNoLegacy on/off verbosity ('true'/'false'); prefer verbosity.
actor_idNoYour agent identity, e.g. 'kimi-code/FI-008' or 'claude/sonnet'. Recorded in the audit log so this action is attributed to you.
epoch_idNoOptional id grouping related sessions into one long-running epoch, e.g. '2026-H2-ops'.
evidenceNoJSON list of facts you have already verified, e.g. [{"fact": "pytest 299 passed", "source": "CI run"}] — carried into the session record.
trace_idNoCorrelation id of your choosing (e.g. 'req-8f3a') — lets you find this call later across kernel logs and organ systems.
_envelopeNoReserved for the transport layer — never fill this in.
verbosityNoHow detailed responses should be: 'minimal' (default), 'standard', or 'full'.minimal
session_idNoSession id returned by arif_init (looks like 'sess-…'). Pass it on every call after init so the action is attributed to your governed session.
agent_policyNoJSON constraints on your own behavior, e.g. {"autonomy": "reversible_only", "forbidden": ["git push"]}.
auth_contextNoJSON carrying external authentication material, e.g. {"token_class": "bearer", "issuer": "github"} — used for gating.
counterpartyNoJSON describing the other party in a two-agent exchange, e.g. {"agent_id": "hermes/1", "role": "verifier"}.
sovereign_idNoIdentifier of the human principal you act for, when acting under their explicit delegation.
session_tokenNoSession Continuity Token (SCT) — the credential string arif_init returned alongside session_id. Pass it back on follow-up calls to prove the session is yours; without it the kernel treats you as unauthenticated.
actor_signatureNoCryptographic signature over the request, if your agent holds a key — proves the call genuinely came from actor_id. Omit if you have no key.
caller_actor_idNoIf you are calling on behalf of another agent, that agent's id — builds a delegation chain for the audit log.
delegation_modeNoContract governing delegated calls, e.g. 'read_only' or 'governed'.
idempotency_keyNoClient-chosen key (e.g. 'job-42-attempt-1'). Retrying with the same key will not repeat the effect — safe retries on flaky networks.
ack_irreversibleNoSet true to acknowledge that session records are permanent audit artifacts and cannot be deleted afterwards.
executor_actor_idNoId of the agent that will actually carry out work under this session's permissions, if different from actor_id.
declared_model_keyNoThe model you run on, e.g. 'zai-coding-plan/glm-5.3'. Informational only — the kernel records but never trusts it.
client_capabilitiesNoJSON declaring what your client supports, e.g. {"transports": ["http"], "protocol": "2025-11-25"}.
requested_authorityNoThe highest class of action you may take: 'OBSERVE_ONLY' (read-only, default) or a higher governed class granted by your policy.OBSERVE_ONLY
previous_session_hashNoHash string returned when your previous session closed; supplying it chains this session to that one, proving continuity.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, so the description is consistent with a write operation. It adds context about binding actor identity, constitutional floors, and the audit chain, and notes the return of a session token. However, it does not explicitly disclose that session records are permanent and irreversible, which is only mentioned in the ack_irreversible parameter schema, not the main description. No contradiction with annotations, but the description could be more explicit about the durability of created sessions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences: it states the core purpose, the return value, and mode guidance, then lists all modes. It is front-loaded with the most important information, contains no fluff, and is well-structured for an agent to quickly grasp the tool's role and when to use different modes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (29 parameters, no output schema), the description covers the essential context: it explains the purpose, the return credentials, and the existence of modes for preflight/resume. It does not explicitly mention the ack_irreversible requirement or the exact return format, but the parameter schema and the mention of session_id/session_token cover the critical aspects. It is complete enough for an agent to call it correctly, though slightly more detail on side effects would improve it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers all 29 parameters with detailed descriptions, including the mode enum, so the description need not add parameter-level detail. The description's mention of preflight and resume modes aligns with the schema, but adds no new semantics beyond what the schema already provides. Baseline of 3 is appropriate for 100% schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Ignite a governed kernel session' and specifies what it binds (actor identity, constitutional floors, audit chain) and what it returns (session_id + session_token). It also implicitly differentiates itself from siblings by noting every other arif_* verb requires these credentials, making it the entry point.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit mode guidance ('Use mode=preflight to inspect an existing session without re-igniting, mode=resume to continue one') and states that every other arif_* verb requires the returned session credentials, which tells an agent this is the first call to make. It doesn't explicitly list when not to use it or name alternative tools, but the entry-point role is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arif_judge888 Judge · Verdict
Read-only
Inspect

KERNEL 666 · Binding constitutional arbitration: evaluates a candidate action or claim and returns SEAL / HOLD / SABAR / VOID with the full reason chain. Weighs action class, blast radius, reversibility, entropy pathway, and cooling state; a SEAL verdict here is what arif_forge requires before it will execute anything. Judge arbitrates — it does not gather (evidence comes via arif_observe, reasoning via arif_think) and it does not mutate (execution is arif_forge, permanence is arif_seal). Modes: judge, intercept, validate, hold, escalate.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo'judge' (default) — render a verdict on candidate; 'intercept' — pre-flight gate before an action runs; 'validate' — re-check a prior verdict; 'hold' — place an action into a cooling period; 'escalate' — refer the decision to the human owner.judge
nonceNoOne-time random string (e.g. a UUID) making this request unique; protects against replay of the same call.
domainNoSubject area for claim evaluation, e.g. 'geoscience' or 'finance'.
key_idNoWhich of your registered keys produced actor_signature, e.g. 'key-1'.
actor_BNoYour calibrated confidence for a prediction, 0.0–1.0 (Brier-style score component).
actor_idNoYour agent identity, e.g. 'kimi-code/FI-008' or 'claude/sonnet'. Recorded in the audit log so this action is attributed to you.
evidenceNoThe facts the verdict should rest on — a JSON list like [{"fact": "tests pass", "source": "CI"}], or an object.
_envelopeNoReserved for the transport layer — never fill this in.
actor_PhiNoJSON map of supporting signals about you, e.g. {"consistency": 0.9, "track_record": 0.7}.
candidateNoThe action or claim being judged, stated plainly, e.g. 'delete table users in prod'.
claim_textNoThe exact claim sentence under judgment, e.g. 'uptime exceeded 99% in August'.
session_idNoSession id returned by arif_init (looks like 'sess-…'). Pass it on every call after init so the action is attributed to your governed session.
action_tierNoHow risky the action is: 'standard', 'high', or 'critical' — higher tiers demand stronger evidence.standard
claim_classNoEpistemic strength of the claim: 'OBS' (directly observed), 'DER' (derived from observations), 'INT' (interpreted), 'SPEC' (speculative), 'UNKNOWN' (cannot witness — honest null; maps to TruthClass.UNK, weight 0.30, cannot authorize mutation). RULE absence-null: 'not in my knowledge' is NOT evidence a claim is false — label UNKNOWN, never FALSE.
claim_tenseNo
niat_paramsNoJSON intent-calibration settings, e.g. {"sincerity": 0.8, "stated_goal": "verify the claim"}.
action_classNoKind of action under judgment: 'OBSERVE' (read-only), 'DRAFT' (compose text), 'MUTATE' (change state), 'IRREVERSIBLE' (permanent).
blast_radiusNoWho or what the action can affect: 'self', 'session', 'organ', or 'federation'.
seal_purposeNoOne sentence saying why this verdict/record must exist, e.g. 'closing deployment D-17'. Stored permanently alongside the record.
session_tokenNoSession Continuity Token (SCT) — the credential string arif_init returned alongside session_id. Pass it back on follow-up calls to prove the session is yours; without it the kernel treats you as unauthenticated.
context_sourceNoWhere the candidate came from: 'session', 'file', or 'memory'.
heart_critiqueNoJSON result of an ethics/dignity check, e.g. {"coercion": false, "dignity": 0.9} — feeds the verdict.
vault_entry_idNoId of the permanent-ledger entry this verdict should attach to (from an earlier arif_seal response).
actor_signatureNoCryptographic signature over the request, if your agent holds a key — proves the call genuinely came from actor_id. Omit if you have no key.
authority_tokenNo
entropy_pathwayNoHow the action changes system order: 'reduces', 'neutral', or 'increases' complexity.
entropy_receiptNoJSON receipt from an entropy computation, bound to this judgment.
authority_effectNoWhat permission a SEAL verdict would grant, e.g. 'execute forge plan P-9'.
cooldown_entry_idNoId of an existing cooling-period record to consult before this action may proceed.
sovereign_receiptNoReference to the human owner's explicit approval, for the rare case where they have already decided directly.
reversibility_levelNoHow hard the action would be to undo: 'reversible', 'hard', or 'irreversible'.
requested_capabilityNoThe specific capability being requested, e.g. 'forge.execute' — checked against the capability registry.
constitutional_chain_idNoId of the evidence chain this call belongs to (observe→think→judge→seal). Copy it from the earlier step's response to link the steps together.
arif_memoryMemory Governor · KernelA
Destructive
Inspect

KERNEL 555 · Governed memory of the kernel itself: six tiers (L1–L6) with per-mode gating. recall, inspect, and audit are read paths; remember, revise, promote, and forget mutate tiers — promote and forget additionally require human_approval=true. Use for cross-session lessons, canon, and memory audit; external evidence belongs to arif_observe and reasoning artifacts to arif_think. Modes: recall, inspect, attest, remember, promote, revise, forget, audit, metabolize.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo'recall' (default) — semantic search of stored memories; 'inspect' — read one memory in full; 'attest' — vouch for a memory's accuracy; 'remember' — store new text; 'promote' — move a memory up a tier; 'revise' — replace a memory's text; 'forget' — remove a memory (gated); 'audit' — integrity scan; 'metabolize' — compact and consolidate tiers.recall
tierNoWhich memory tier to target, 'L1'–'L6' (L1 = hot working memory, L6 = sealed canon).
queryNoWhat to search for, in plain language (mode=recall/audit), e.g. 'past deploy rollback steps'.
contentNoThe text to store (mode=remember) — write it as a self-contained lesson or fact.
payloadNoExtra mode-specific fields as JSON — e.g. {"truth_class": "DERIVED", "provenance": "CI log"} for remember.
to_tierNoDestination tier when promoting, e.g. 'L4'.
actor_idNoYour agent identity, e.g. 'kimi-code/FI-008' — recorded in the audit log.
lease_idNoId of a short-lived permission grant (lease) authorizing this write, when one was issued to you.
trace_idNoCorrelation id of your choosing (e.g. 'req-8f3a') to find this call later in the logs.
memory_idNoUUID of the memory entry to inspect/revise/forget — it was returned when the memory was created or last listed.
session_idNoSession id returned by arif_init; attributes this call to your governed session.
new_contentNoReplacement text for the memory (mode=revise) — must refer to the same memory_id.
session_tokenNoSession Continuity Token (SCT) returned by arif_init — proves the session is yours.
human_approvalNoSet true ONLY when the human owner explicitly approved this promote/forget — the gate refuses without it.
idempotency_keyNoClient-chosen key (e.g. 'memo-42'); retries with the same key will not store duplicates.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark destructiveHint=true RF, but the description adds valuable specifics: it categorizes read paths ('recall, inspect, and audit') vs. mutating modes ('remember, revise, promote, and forget') and highlights the human_approval=true gate for promote/forget. This goes beyond annotations to explain which operations are destructive and what guardrails apply. It does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core concept ('Governed memory of the kernel itself'), then the next sentence covers reading vs. mutation, then usage guidance, then a mode list (which is somewhat redundant with the schema but serves as a quick reference). It earns its length with no fluff, though the 'KERNEL 555' preamble adds a little flavor without substance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with 15 optional parameters and no output schema, the description gives the essential context: purpose, modes, tier system, gating, and sibling boundaries. It explains the tool's role in the environment and how to use it, while the schema handles parameter details. Some specifics like return-value behavior are absent but not required without an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for all 15 parameters, so the schema already documents each parameter. The description does not add new parameter semantics beyond noting that promote/forget require human_approval=true, which is also in the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('governed memory') and resource ('kernel itself'), defines six tiers (L1–L6), and explicitly differentiates from siblings: 'external evidence belongs to arif_observe and reasoning artifacts to arif_think.' An agent can immediately tell this is the memory-management tool, not the observe or think tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use for cross-session lessons, canon, and memory audit' gives positive use cases, while 'external evidence belongs to arif_observe and reasoning artifacts to arif_think' provides clear exclusions and routes to alternatives. This is explicit when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arif_observe111 Observe · Sense RealityA
Read-onlyIdempotent
Inspect

KERNEL 111 · Collect evidence — facts and sources with epistemic tags (OBS) and uncertainty bounds, never conclusions. mode=search queries the open web/literature; mode=fetch retrieves a URL and records its provenance; mode=vitals reads kernel machine telemetry. Reason over what you gathered with arif_think; delegate domain analysis to an organ with arif_route. Modes: search, fetch, hybrid_discovery, ingest, compass, atlas, entropy_dS, vitals.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoExact URL to retrieve when mode=fetch, e.g. 'https://example.com/report.pdf'.
modeNo'search' (default) — web/literature query; 'fetch' — retrieve one URL with provenance; 'hybrid_discovery' — combine sources; 'ingest' — absorb a document; 'compass'/'atlas' — guided navigation; 'entropy_dS' — measure system change; 'vitals' — kernel machine telemetry.search
queryNoWhat to look for, in plain language, e.g. 'TDQS scoring rubric MCP'.
layersNoRestrict where to look, e.g. ["web"] or ["canon", "memory"]; omit to search everything available.
actor_idNoYour agent identity, e.g. 'kimi-code/FI-008' or 'claude/sonnet'. Recorded in the audit log so this action is attributed to you.
_envelopeNoReserved for the transport layer — never fill this in.
session_idNoSession id returned by arif_init (looks like 'sess-…'). Pass it on every call after init so the action is attributed to your governed session.
result_limitNoMaximum number of results to return in search modes; default 10.
session_tokenNoSession Continuity Token (SCT) — the credential string arif_init returned alongside session_id. Pass it back on follow-up calls to prove the session is yours; without it the kernel treats you as unauthenticated.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare a safe read-only, idempotent, open-world profile, and the description is consistent with them. Beyond that, it adds a distinctive behavioral contract — 'never conclusions' — telling the agent the tool deliberately withholds conclusions and only emits evidence with epistemic tags and uncertainty bounds, plus provenance recording for fetch. This shapes how an agent must interpret results, which exceeds the annotation coverage; a 5 would require output-format or error-behavior disclosure it doesn't provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well front-loaded: the first sentence delivers the core purpose, and the mode explanations are compressed into one efficient sentence. Punished for the decorative 'KERNEL 111 ·' prefix and the trailing 'Modes: search, fetch, hybrid_discovery, ingest, compass, atlas, entropy_dS, vitals.' enumeration, which largely duplicates the preceding sentence and the schema enum.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For the complexity — 9 parameters, 8 modes, no output schema — the description is only partially complete. It covers the epistemic contract and three modes, while five modes are unexplained in the description (the schema's mode parameter partially compensates by glossing each). With no output schema, the agent must rely on the abstract 'facts and sources with epistemic tags (OBS) and uncertainty bounds' for the return contract; error behavior, result structure, and cross-mode interplay are absent. The rich schema props it up to minimum viable but no higher.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3 per rubric. The description adds value above that baseline by explaining the semantics of three modes in prose (search, fetch, vitals) and framing result output as evidence-with-tags, enriching the mode enum. The other five enum modes (hybrid_discovery, ingest, compass, atlas, entropy_dS) get their semantics only from the schema, so it does not reach 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource — 'Collect evidence — facts and sources with epistemic tags (OBS) and uncertainty bounds, never conclusions' — and differentiates from siblings by explicitly routing downstream work to arif_think and arif_route. It also names three concrete modes (search, fetch, vitals) with behaviors. However, the multi-mode sprawl of 8 modes (compass, atlas, entropy_dS left unexplained) and the poetic 'Observe · Sense Reality' title blur the central purpose slightly, preventing a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit division-of-labor guidance: 'Reason over what you gathered with arif_think; delegate domain analysis to an organ with arif_route,' and ties mode selection to task ('mode=search queries the open web/literature; mode=fetch retrieves a URL...; mode=vitals reads kernel machine telemetry'). Lacks when-not-to-use guidance for the remaining siblings (arif_memory, arif_judge, arif_seal) and for 5 of the 8 modes, so it falls short of fully explicit exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arif_route444 Route · Intent→OrganA
Read-onlyIdempotent
Inspect

KERNEL 444 · Intent→organ router: classifies a natural-language intent and dispatches it to the specialist organ (GEOX geoscience, WEALTH capital, WELL vitality, A-FORGE execution). Returns the routing decision only — no organ call — unless organ_tool names the target tool and arguments carries its inputs. Prefer this over guessing organs yourself; use arif_think for reasoning you keep in-kernel.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoroute
taskNoAlias for intent (backward compat).
organNoOptional explicit organ override. If provided, intent matching is skipped and this organ is used directly.
intentNoNatural-language description of what the user wants. e.g. "interpret this seismic section", "assess portfolio risk"
actor_idNoYour agent identity, e.g. 'kimi-code/FI-008' or 'claude/sonnet'. Recorded in the audit log so this action is attributed to you.
_envelopeNoReserved for the transport layer — never fill this in.
argumentsNoArguments to pass to organ_tool.
mission_idNoExplicit human-cockpit mission binding (investigate|interpret| decide|build|monitor|remember). When set, skips keyword classification and binds the six-mission plan. Preferred over free-text when the agent already knows the mission.
organ_toolNoThe tool name on the target organ to call. If absent, returns routing decision only (no bridge call).
session_idNoSession id returned by arif_init (looks like 'sess-…'). Pass it on every call after init so the action is attributed to your governed session.
session_tokenNoSession Continuity Token (SCT) — the credential string arif_init returned alongside session_id. Pass it back on follow-up calls to prove the session is yours; without it the kernel treats you as unauthenticated.
contract_c_kwargsNoExtra keyword arguments passed through to the organ tool call, as a JSON object.

TDQS

A4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description usefully discloses that the default is a routing decision with no downstream call, and that organ_tool + arguments triggers an actual bridge call. However, that bridge call can invoke an arbitrary organ tool whose effects are unknown, which sits uneasily against the readOnlyHint=true annotation — the definition is the only place this write-capable passthrough is surfaced, and it is not reconciled with the read-only annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the routing purpose, then the default behavior, then the alternative routing to arif_think. No filler or restated name/title content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter router with no output schema and high schema coverage, the description covers the essential decision logic (route vs. bridge) and the sibling alternative. Auth/session prerequisites are left entirely to the schema field descriptions, which is acceptable given their coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 92%, so the baseline is 3, but the description still adds semantic value by explaining the organ_tool/arguments interplay and the two-mode dispatch that the schema documents only per-field. It adds the mode-switching meaning rather than repeating field docs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (classifies) and resource (natural-language intent) and names the dispatch targets (GEOX, WEALTH, WELL, A-FORGE). It also explicitly distinguishes itself from a sibling ('use arif_think for reasoning you keep in-kernel'), so an agent can tell it apart without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear directive ('prefer this over guessing organs yourself') and names the alternative sibling (arif_think) with the condition that selects it. It also explains the two operating modes (route-only vs. bridge via organ_tool), though it stops short of enumerating when-not-to-use scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arif_seal999 Seal · VAULT999A
Destructive
Inspect

KERNEL 999 · Append an entry to VAULT999, the immutable ledger — accepted entries can never be edited or removed; there is no unseal. Use for permanent records of verified outcomes, lessons, and session closure once a verdict exists. ack_irreversible=true is the explicit acknowledgment of permanence, and judge_state_hash binds the entry to the verdict that authorized it. Reversible changes belong in arif_forge under a SEAL. Modes: seal, verify, ledger, changelog, audit, session_close.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo'seal' (default) — append a permanent entry; 'verify' — check one entry; 'ledger' — read the ledger head; 'changelog' — recent appends; 'audit' — integrity check; 'session_close' — close out a session into the ledger.seal
nonceNoOne-time random string (e.g. a UUID) making this request unique; protects against replay of the same call.
payloadNoThe content to store forever, as a string (usually JSON-serialized) — the outcome, lesson, or record itself.
actor_idNoYour agent identity, e.g. 'kimi-code/FI-008' or 'claude/sonnet'. Recorded in the audit log so this action is attributed to you.
_envelopeNoReserved for the transport layer — never fill this in.
session_idNoSession id returned by arif_init (looks like 'sess-…'). Pass it on every call after init so the action is attributed to your governed session.
drift_eventsNoJSON list of deviations observed, e.g. [{"what": "schema drift", "where": "tools/list"}] — stored with the record.
seal_purposeNoOne sentence saying why this verdict/record must exist, e.g. 'closing deployment D-17'. Stored permanently alongside the record.
witness_typeNoWho witnessed the sealed fact: 'ai', 'human', or 'external' system.ai
session_tokenNoSession Continuity Token (SCT) — the credential string arif_init returned alongside session_id. Pass it back on follow-up calls to prove the session is yours; without it the kernel treats you as unauthenticated.
constitutionalNoJSON block of governance metadata (floors consulted, chain references) — normally built by the kernel, not by callers.
actor_signatureNoCryptographic signature over the request, if your agent holds a key — proves the call genuinely came from actor_id. Omit if you have no key.
ack_irreversibleNoSet true to confirm you understand sealed entries are PERMANENT — they can never be edited or removed.
judge_state_hashNoHash string from the arif_judge SEAL response — ties this entry to the verdict that authorized it.
constitutional_chain_idNoId of the evidence chain this call belongs to (observe→think→judge→seal). Copy it from the earlier step's response to link the steps together.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses the critical irreversible behavior: 'accepted entries can never be edited or removed; there is no unseal.' It explains the significance of ack_irreversible and judge_state_hash, going beyond the generic destructiveHint annotation. No contradiction with annotations; it enriches them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense yet efficiently structured. It front-loads the core purpose, then gives usage, key parameter explanations, and alternatives, all in about five sentences. Every sentence earns its place, and the format is scannable for an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 15 parameters and no output schema, the description covers the essential context: core purpose, irreversibility, usage timing, alternatives, and modes. It references session attribution indirectly through parameter hints but the schema covers that. The description is complete enough for an agent to decide when and how to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides 100% description coverage, so the baseline is 3. The description adds value by explaining ack_irreversible as 'explicit acknowledgment of permanence' and judge_state_hash as 'binds the entry to the verdict that authorized it', giving semantic weight to these parameters. It also lists modes but the schema already details them, so it's a moderate addition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Append an entry to VAULT999, the immutable ledger'. It clearly states the tool's core action and domain, and distinguishes it from siblings by noting 'Reversible changes belong in arif_forge under a SEAL'. The purpose is explicit and actionable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use: 'Use for permanent records of verified outcomes, lessons, and session closure once a verdict exists.' It also names the alternative tool for reversible changes and provides a clear exclusion. The mention of modes further clarifies usage scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arif_think333 Think · MindA
Read-onlyIdempotent
Inspect

KERNEL 333 · Structured reasoning pass: decomposes a query and returns reasoning steps labeled OBS (observed), DER (derived), INT (interpretation), SPEC (specification) under truth floors. Produces reasoning records only — no verdicts (those come from arif_judge) and no state changes. The plan-family modes draft/review/approve execution plans; simulate and wonder explore counterfactuals. Modes: reason, reflect, verify, axioms, plan, plan_review, plan_approve, refactor_plan, metabolize, simulate, wonder, atlas.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo'reason' (default) — decompose a question; 'reflect' — self-review of prior reasoning; 'verify' — check a derivation; 'axioms' — surface hidden assumptions; 'plan'/'plan_review'/'plan_approve'/'refactor_plan' — execution-plan lifecycle; 'metabolize' — consolidate past reasoning; 'simulate' — what-if; 'wonder' — open exploration; 'atlas' — map the problem space.reason
queryNoThe question, claim, or problem to reason about, in plain language.
plan_idNoPlan reference from an earlier plan-mode response — needed for the review/approve/refactor steps.
actor_idNoYour agent identity, e.g. 'kimi-code/FI-008' or 'claude/sonnet'. Recorded in the audit log so this action is attributed to you.
_envelopeNoReserved for the transport layer — never fill this in.
session_idNoSession id returned by arif_init (looks like 'sess-…'). Pass it on every call after init so the action is attributed to your governed session.
witness_typeNoWho vouches for the reasoning record: 'ai' (default), 'human', or 'external' system.ai
session_tokenNoSession Continuity Token (SCT) — the credential string arif_init returned alongside session_id. Pass it back on follow-up calls to prove the session is yours; without it the kernel treats you as unauthenticated.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the description adds value by specifying 'Produces reasoning records only — no verdicts and no state changes,' which aligns with the annotations. It also mentions 'under truth floors' as a behavioral constraint. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph that front-loads the core purpose and then lists modes. It is efficient and avoids redundancy, though a bulleted mode list might improve scannability. Every sentence contributes to understanding the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (8 parameters, 12 modes) and the rich parameter descriptions in the schema, the description provides a solid high-level overview and distinguishes the tool from arif_judge. It does not explain the exact output format or the meaning of 'truth floors,' but the schema covers parameter details and no output schema is expected. It is adequate for an agent to select and invoke the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already explains each parameter. The description adds marginal value by grouping modes (e.g., plan-family, simulate/wonder) but largely restates information already present in the schema's mode enum. It does not compensate for any coverage gap because there is none.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('decomposes a query'), a clear resource ('structured reasoning pass'), and distinguishes its output (reasoning records labeled OBS/DER/INT/SPEC) from verdicts (which arif_judge provides). It also enumerates the modes, making the tool's purpose unambiguous and differentiating it from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly notes that verdicts come from arif_judge, providing an alternative, and groups plan-family modes and exploratory modes. However, it does not give explicit 'when to use this vs. that' conditions for all siblings, leaving some inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool update
    • Changedarif_judge1 field changed
      • addedInput schema / properties / claim_tense
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null
        +}

Publisher details

Operator
ARIF FAZIL · Publisher source
Operator website
https://arif-fazil.com
Vendor relationship
First-party
Documentation
Unknown
Trust center
Unknown
Restrictions
Unknown

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Provides a constitutional governance framework for AI agents, offering 13 MCP tools for session initiation, reasoning, evidence fetching, judgment, and execution, all governed by hard invariant laws and a hierarchical set of constitutional floors.
    11
    133 PyPI
    53
    AGPL 3.0
  • A
    license
    Not graded
    quality
    F
    maintenance
    Constitutional MCP server enforcing 13 Floors of governance for AI agents, providing tools for session anchoring, reasoning, safety critique, and audit logging.
    AGPL 3.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    A constitutional MCP server that enforces external authority, loud rejections, and an unforgeable hash-chained ledger, preventing AI agents from minting their own identity, approving themselves, or rewriting history.
    2 npm
    Apache 2.0
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.