JEVULON VII (JVII)
Provides a native drop-in integration into LangGraph StateGraphs, enabling use as a supervisor node, router, or safety middleware. Includes createSupervisorNode and createShieldMiddleware for zero-token deterministic safety and coordination within LangGraph workflows.
Allows connecting local Ollama models as a pluggable decision engine via the DecisionEngine seam, enabling local LLM-based decisions instead of the built-in deterministic simulation engine.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@JEVULON VII (JVII)claim src/auth.ts for editing and check for collisions"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
JEVULON VII (JVII)
The deterministic, self-healing control layer for parallel AI coding agents.
JEVULON VII replaces slow, token-hungry LLM supervisor chatrooms with instant sub-second System-1 decision gates and a shared stigmergic coordination whiteboard (board.json). Run Claude Code, Cursor, Codex, and LangGraph agents in true parallel without merge collisions or file overwrites.
โก 10-Second Quickstart
You don't need an account or API key to start coordinating agents locally.
1. Run via npx
npx -y jevulon2. Connect to Cursor or Claude Desktop (MCP)
Add JEVULON VII to your claude_desktop_config.json or Cursor MCP settings:
{
"mcpServers": {
"jvii": {
"command": "npx",
"args": ["-y", "jevulon"]
}
}
}Once connected, your agents automatically gain access to the coordination tools:
sentinel_presenceโ Claims files before editing and checks for workspace collisions.sentinel_shield_inspectโ Deterministic zero-latency policy floor guarding destructive commands.sentinel_superviseโ Autonomous wave dispatch, worker recovery, and completion verification.
Related MCP server: coordinaut
๐๏ธ How It Works: System-1 vs. System-2
Most multi-agent frameworks use System-2 LLM supervisors (chatty conversational models like GPT-4 or Claude 3.5 Sonnet) to coordinate workers. This burns thousands of tokens per minute, takes 10โ15 seconds per coordination turn, and frequently hallucinates file states.
JEVULON VII decouples execution into two specialized layers:
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ SYSTEM-2 WORKERS (Claude Code, Cursor, Codex) โ
โ Thinks, plans, and writes code inside your workspace โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ (Claims files / Asks gates)
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ SYSTEM-1 CONTROL LAYER (JEVULON VII) โ
โ โ Deterministic file locking (board.json) โ
โ โ Zero-latency offline safety floor (0ms) โ
โ โ Sub-second categorical decision gates (~500ms) โ
โ โ $0 wasted on supervisor conversation โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ๐ฆ Core Capabilities
1. Deterministic Conflict Prevention
Before an agent edits any file, it registers an atomic claim on the shared board. If sibling agents attempt to modify overlapping files or dependent interfaces, JEVULON VII negotiates priority and prevents overwriting code before it happens.
2. Zero-Egress Offline Safety Floor
A 77-case deterministic safety engine blocks destructive shell commands (rm -rf, raw disk wipes, credential exfiltration, base64 payload piping) locally in 0ms without sending your code to any external network.
3. Native LangGraph Drop-in Integration
Integrate directly into LangGraph StateGraphs as a supervisor, router, or safety middleware without restructuring your workflow:
import { createSupervisorNode, createShieldMiddleware } from "jevulon/langgraph";
// Add zero-token deterministic safety middleware to your graph
const shieldMiddleware = createShieldMiddleware({
onBlock: "halt",
onEscalate: "human_review"
});4. Pluggable Decision Seam (BYOK)
Implements a structural DecisionEngine seam (choice and noul in src/engine.ts). Connect your own TypeSafe JEV keys, local Ollama/vLLM models, or use the built-in deterministic simulation engine.
๐ณ Enterprise & Private VPC Deployment
For enterprise teams with strict compliance or zero-external-egress mandates, a pre-hardened multi-container Docker Compose bundle is available:
๐ https://github.com/Jason-Fay/jevulon-docker
curl -sSL https://jevulon.com/deploy/docker.tar.gz | tar -xz && docker compose up -d๐งช Testing & Verification
JEVULON VII maintains a rigorous test suite covering the deterministic floor, stigmergy, MCP servers, and circuit breakers:
git clone https://github.com/Jason-Fay/jevulon.git
cd jevulon
npm install
npm test๐ License
MIT ยฉ 2026 MetaWave / Jason Fay. See LICENSE for details.
Available Tools
7 toolssentinel_approve_escalationA
Operator authorization for one escalated action. Approvals are single-use, bound to the exact action payload, and expire. Approvals must be attributed: pass approvedBy with your host's operator identity โ unattributed approvals, and decisions attributed to the identity that requested the escalation, are both refused. Configure your MCP host to always require human confirmation for this tool โ the server cannot verify who calls it.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | Optional reasoning from the human operator | |
| actionId | Yes | The actionId returned by sentinel_shield_inspect | |
| approved | Yes | true to authorize execution, false to deny | |
| approvedBy | No | Identity of whoever is deciding, when your host attributes it (recorded on the decision) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so: approvals are single-use, bound to the exact payload, and expire, and the server cannot verify the caller's identity. These are non-obvious consumption/auth semantics that an agent could not infer from the schema alone. No contradiction with structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core purpose and the single-use/bound/expiring semantics, then the attribution constraint, then the host-configuration caveat. No sentence is redundant and none could be dropped without losing behaviorally significant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description does not say what a successful approval returns (e.g., a token or decision record), which is the one gap for a gating tool. Everything else an agent needs โ attribution requirement, refusal conditions, single-use and expiry behavior, host-side human confirmation โ is covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description still adds meaning beyond the schema: it converts approvedBy from the schema's permissive 'when your host attributes it' into a hard requirement, and warns that the requester's own identity is rejected. It does not explain reason or approved beyond what the schema already says.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (operator authorization of one escalated action) and immediately scopes it to a single action rather than a class of actions. It also names the sibling tool that produces the actionId (sentinel_shield_inspect), so an agent can place it in the workflow without opening another schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it (authorizing one escalated action), the preconditions (approvedBy with the host's operator identity), and the cases where a decision is refused (unattributed approvals, self-attribution by the requester). It also instructs the operator to configure the host to require human confirmation, which is direct usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sentinel_code_reviewA
JEVULON VII Pro code review (Anti-Slop Sieve): audits one file or snippet under the neutral-state 3-run protocol and returns the modal smell label, complexity (cx) & security hazard (hz) scores, gate verdicts, and the refactor directive. Call it before marking a task complete and fix the flagged slop first; a fired abort gate means completion is blocked. Requires a Pro/Team license token in JVII_TOKEN; the code under review goes only to your own BYOK JEV endpoint (never scraped).
| Name | Required | Description | Default |
|---|---|---|---|
| runs | No | Neutral judge runs (default 3, the project protocol) | |
| filePath | Yes | Path of the file under review (part of the judge state; read from disk when sourceCode is omitted) | |
| sourceCode | No | The code to audit; when omitted the file at filePath is read |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the auth requirement (Pro/Team token in JVII_TOKEN), the data-handling guarantee (code goes only to your own BYOK JEV endpoint, never scraped), and the outcome semantics (a fired abort gate blocks completion). Rate limits and timeout/error behavior are unstated, keeping it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, followed by usage guidance and then the license/privacy caveat, so an agent gets the essentials first. It is dense with branded jargon ("Anti-Slop Sieve", "modal smell label", "cx & hz") that slightly taxes parsing, but there is little true filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description usefully enumerates the return payload (modal label, cx/hz scores, gate verdicts, refactor directive). Together with auth, privacy, and the completion-blocking behavior, an agent has enough to call it correctly, though pagination/error semantics remain unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters; per the rubric the baseline is 3. The description adds only marginal context via "3-run protocol" tying to the runs default, and "one file or snippet" loosely hinting at filePath vs sourceCode, so it does not meaningfully exceed the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('audits one file or snippet') and enumerates the outputs (modal smell label, cx/hz scores, gate verdicts, refactor directive). The file-audit scope distinguishes it from the other sentinel_* siblings (supervise, shield_inspect, overseight_monitor), which are supervision/monitoring rather than single-file reviewing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Call it before marking a task complete and fix the flagged slop first; a fired abort gate means completion is blocked" gives explicit when-to-use and a consequence trigger. It lacks an explicit alternative/exclusion clause (e.g., when not to use versus sentinel_shield_inspect), so it stops short of routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sentinel_get_statusA
Returns supervisor state, blackboard counts, pending escalations, and local usage counters.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. 'Returns' implies a read-only operation and the field list tells the agent what data surfaces, which is genuinely useful. However, it says nothing about permissions, cost, latency, or whether calling it has any side effects, so transparency is only partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that enumerates the payload with zero filler. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description partially compensates by naming the returned data categories. For a zero-parameter read tool this is close to sufficient, though the shape/format of those counters and escalation entries remains unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to clarify and the baseline of 4 applies. No parameter confusion is possible.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Returns') and enumerates the concrete payload: supervisor state, blackboard counts, pending escalations, and local usage counters. That is far more informative than a generic 'gets status'. It stops short of distinguishing itself from siblings like sentinel_oversight_monitor or sentinel_supervise, which also sound status/monitoring oriented.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this versus the other sentinel_* tools, nor any preconditions or exclusions. Usage must be inferred entirely from the tool name and the list of returned fields.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sentinel_oversight_monitorC
Evaluates a worker trace for retry loops, deadlocks, and toxic failure patterns (continue / quarantine / reroute / halt).
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Action just attempted | |
| workerId | Yes | ID of the executing worker | |
| recentTrace | No | Recent preceding event strings | |
| outputOrError | Yes | Result or error output |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavior burden. It names four verdicts (continue / quarantine / reroute / halt) but never says whether the tool merely recommends them or actually applies them โ a critical distinction for a tool whose outcomes look like enforcement actions. Permissions, side effects, and whether it mutates worker state are all unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the verb and detection targets. Nothing is wasted, though the parenthetical verdict list makes it slightly overloaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter diagnostic tool with no output schema, the parenthetical verdict list partially compensates by hinting at the return values. What remains missing is whether the verdict is advisory or enforced and when in a workflow the call belongs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so workerId, action, outputOrError, and recentTrace are already documented in the schema itself. The description adds no parameter meaning beyond that, which is the baseline case for full coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (Evaluates) plus the resource (a worker trace) and enumerates the failure classes it detects (retry loops, deadlocks, toxic failure patterns). It does not, however, distinguish itself from the close-sounding sibling sentinel_supervise, leaving the agent to guess which of the two to call.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to invoke this tool versus sentinel_supervise, sentinel_shield_inspect, or sentinel_approve_escalation, and no prerequisites or exclusions. The only usage signal is implicit โ that it is a diagnostic step on a worker trace โ which the agent must infer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sentinel_presenceA
Shared presence board for every agent working in this repository (advisory โ it never blocks an action). Quick usage: โข Claim files before writing: { action: 'claim', paths: ['src/file.js'], mode: 'edit', note: 'why' }. โข Release files when done: { action: 'release', paths: ['src/file.js'] }. โข Checkpoint changes: { action: 'checkpoint' }. โข Read active board: { action: 'read' }. checkpoint is the one call to make before an irreversible step. claim records working paths and reports conflicts; release drops your claims. Signals: 'signal'. Directives: 'direct', 'ack', 'resolve'. Atomic queues: 'allocate'. Pheromones: 'alarm', 'attract'.
| Name | Required | Description | Default |
|---|---|---|---|
| key | No | Signal/contract key to post (action: signal) | |
| kind | No | Directive kind (action: direct) | |
| mode | No | Claim mode (default edit) | |
| note | No | Why you are in these paths (or a note on an acknowledgement) โ other agents inherit this | |
| items | No | Ordered candidate items to allocate from (action: allocate) | |
| paths | No | Repository-relative paths to read, claim, or release | |
| value | No | Arbitrary JSON-serializable signal payload (action: signal) | |
| action | Yes | Board operation | |
| target | No | What the directive is about: a path, a job id, or a short description (action: direct) | |
| agentId | No | Calling agent identity (defaults to server PID); for 'direct', the target agent to steer | |
| issuedBy | No | Who is issuing the directive (action: direct; defaults to the caller) | |
| queueKey | No | Named queue key for atomic item allocation (action: allocate) | |
| rationale | No | Why the directive was issued โ the steered agent inherits this | |
| ttlMinutes | No | Claim or directive lifetime in minutes (default 15 for claims, 10 for directives; max 1440) | |
| directiveId | No | Directive to acknowledge or resolve (action: ack / resolve) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that the board is advisory and never blocks an action, that claim reports conflicts, and that release drops your claims. But the trailing lists ('Signals: signal', 'Directives: direct, ack, resolve', 'Pheromones: alarm, attract') are name-only labels that do not explain the mechanics, side effects, or permissions of those operations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, followed by compact bullet-style usage examples, which is efficient for a 12-action tool. The final run-on grouping of signals/directives/queues/pheromones is terse and slightly cryptic, but the overall size is reasonable and most sentences earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a high-complexity tool (12 actions, 15 params, no annotations, no output schema), the description covers claim/release/checkpoint/read well but leaves directives, signal, allocate, alarm, and attract essentially undocumented beyond their action names. An agent can invoke the common paths, but the less-common operations lack enough context to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 15 parameters, including ttlMinutes defaults, mode default, and note inheritance. The inline example payloads for claim/release/checkpoint/read are illustrative but add little beyond what the schema already states; baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence names a specific resource ('shared presence board for every agent working in this repository') and immediately qualifies its nature ('advisory โ it never blocks an action'), so an agent knows it is a coordination/awareness tool rather than an enforcement one. It does not, however, explicitly differentiate itself from the sentinel_* siblings, so the agent must infer its distinct role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete when-to-use guidance tied to workflow moments: claim before writing, release when done, checkpoint before an irreversible step, read for the active board. 'checkpoint is the one call to make before an irreversible step' is an explicit trigger condition. It stops short of stating when NOT to use the tool or how to choose among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sentinel_shield_inspectA
Pre-flight safety inspection for HIGH-RISK, destructive, or irreversible system operations ONLY (e.g. broad deletions, system path wipes, dropping databases, running untrusted remote scripts, or network exfiltration). DO NOT call this tool for routine coding: file reads, code edits, git status/diff/log/branch, package installs (npm/bun/pip), or build/test runs โ routine development does not require inspection.
| Name | Required | Description | Default |
|---|---|---|---|
| context | No | Optional situational context | |
| arguments | No | Optional arguments or payload (bound into the approval fingerprint) | |
| commandOrTool | Yes | The shell command or tool name about to be executed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the tool's risk-gating purpose but not its actual behavior: whether it blocks execution, returns a verdict, requires a follow-up approval (sibling sentinel_approve_escalation hints at a flow), or how it interacts with the pending command. Useful framing but incomplete for a no-annotation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the primary constraint and the DO-NOT list clearly separated. Slightly dense with all-caps emphasis, but every clause earns its place by delineating the risk boundary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the use/avoid decision well, but with no annotations and no output schema the description should explain what the inspection produces and what the caller does next (approval, escalation). The invocation gate is complete; the post-invocation picture is not.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema documents commandOrTool, context, and arguments, so the description need not repeat them. It adds nothing about parameter format or the 'approval fingerprint' semantics that the schema itself only mentions in passing. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('pre-flight safety inspection') and scopes it to HIGH-RISK destructive/irreversible operations with concrete examples (deletions, DB drops, remote scripts). An agent can immediately distinguish this from routine-operation siblings like sentinel_code_review.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use (high-risk/destructive/irreversible ops) and when NOT to (reads, edits, git status/diff/log, package installs, build/test), with named counterexamples. This is the strongest part of the definition and leaves little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sentinel_superviseB
Stigmergic task dispatcher backed by the recovery supervisor: plan advances a wave, halt/resume controls execution, fail records a worker death, complete runs the calibrated completion gate, and classify_circuit routes a task to the appropriate circuit.
| Name | Required | Description | Default |
|---|---|---|---|
| error | No | Failure reason when recording worker death | |
| jobId | No | Target job ID | |
| title | No | Job title when adding a job | |
| action | Yes | The supervisor operation to perform | |
| evidence | No | Completion evidence/oracle output for the gate | |
| workerId | No | Target worker ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It adds useful action semantics beyond the schema enum for five of nine operations (e.g., 'fail records a worker death', 'complete runs the calibrated completion gate'), but it omits behavior for `assign`, `add_job`, and `register_worker`, and says nothing about side effects, permissions, reversibility, or response behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence that efficiently maps the tool's purpose and several actions without filler. It could be improved by covering all actions, but it is appropriately sized and structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex multi-action dispatcher with no annotations and no output schema, the description is incomplete. It omits three of the nine enum actions and provides no safety, permission, or return-value context, leaving significant gaps for an agent to fill before invoking it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents all six parameters. The description does not add syntax, format, or dependency details beyond what the schema provides, making 3 the appropriate baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a task dispatcher with supervisor actions and maps several enum values (`plan`, `halt`/`resume`, `fail`, `complete`, `classify_circuit`) to concrete behaviors. It does not distinguish itself from any sibling tools (`sentinel_get_status`, `sentinel_oversight_monitor`, etc.), and it omits several enum actions (`assign`, `add_job`, `register_worker`), so it is clear but not fully specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no explicit guidance on when to use this tool versus alternatives. It only lists what some actions do, leaving the agent to infer usage context from the action names and overall 'dispatcher' framing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.0- First observed
sentinel_approve_escalation - First observed
sentinel_code_review - First observed
sentinel_get_status - First observed
sentinel_oversight_monitor - First observed
sentinel_presence - First observed
sentinel_shield_inspect - First observed
sentinel_supervise
TDQS
Scored across 7 tools
Each tool has a distinct primary role: dispatcher (supervise), safety inspection (shield_inspect), authorization (approve_escalation), trace evaluation (oversight_monitor), status read, code review, and presence/coordination. There is some conceptual overlap in the safety cluster โ supervise's halt/fail, oversight_monitor's halt verdict, shield_inspect, and approve_escalation all touch escalation/failure handling โ but the descriptions make the boundaries readable.
All seven tools share a consistent `sentinel_` prefix and snake_case formatting, which is strong. The suffixes are slightly loose in form (verb like `supervise`, noun_verb like `shield_inspect`, plain noun like `presence`), but the overall convention is predictable.
Seven tools is well-scoped for a multi-agent supervision/coordination server and sits comfortably in the ideal 3-15 range. Each tool earns its place with an identifiable function, and there is no bloat.
The surface covers the coordination lifecycle: plan/dispatch (supervise), pre-flight safety (shield_inspect), approval (approve_escalation), monitoring (oversight_monitor), status (get_status), quality gate (code_review), and presence (presence). Coverage is broad with only minor potential gaps, and no obvious dead ends for the stated supervision purpose.
Maintenance
Related MCP Connectors
Shared control plane for AI coding agents โ tasks, memory, decisions, file locks. 12 tools.
- AxisOAuthdev.useaxis
Coding agents from Claude Code, Cursor and Codex claim jobs and lock files on one shared board.
- ParleyOAuthdev.weldra
Coordination hub for AI coding agents: message teammates, ask humans, audit every event.
The team layer for AI coding agents: shared contracts, collision alerts, E2EE sessions.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceCoordination layer for AI coding agents working on the same codebase. Adds file locks, shared project memory, and cross-machine file sync so Claude Code, Cursor, Windsurf, and other MCP agents stop overwriting each other.50Apache 2.0
- AlicenseNot gradedqualityCmaintenanceCoordinates parallel AI coding agents by providing task ownership, scoped file locks, handoffs, and verification workflows.MIT
- AlicenseNot gradedqualityBmaintenanceEnables multiple AI coding agents to collaborate on the same Git repository without conflicts through isolated worktrees, file locking, automated test verification, and a serialized merge queue.3 npm7MIT
- AlicenseNot gradedqualityBmaintenanceEnables multiple AI coding agents to safely collaborate in the same git working tree by managing file ownership, merging writes, and preventing snapshot races.6 npm2MIT