agent-reasoning-mcp
Manage autonomous agent reasoning with BDI goal planning, utility-based decisions, beliefs/intentions, and fast System 1 primitives.
Create, update, decompose, inspect, and abandon hierarchical goals and task DAGs.
Evaluate situations, rank candidate actions, and compute multi-attribute expected utility.
Replan after blockers or environment changes and preserve completed subgoals.
Assess quantitative risk for actions, plans, comparisons, and spatial trajectories.
Query learned heuristics, tactics, anti-patterns, and similar past situations.
Configure, inspect, activate, and nudge utility weight profiles.
Retrieve and explain decision traces, utility breakdowns, and candidate rankings.
Manage belief state with confidence scores, decay rates, TTL expiration, and spatial reconciliation.
Queue, dispatch, track, cancel, and resolve runtime behavior intentions.
Gate intention dispatch with safety checks and signed HMAC dispatch tokens.
Maintain reasoning database integrity via stats, Merkle audit, doctor, snapshots, diffs, and restore.
Run low-latency System 1 tools: classify, ask_noul, ask_choice, ask_score, and gate_intention.
Use CLI commands such as init, doctor, inspect, and run; supports air-gapped deterministic evaluation.
@putervision/agent-reasoning-mcp
Strategic BDI Reasoning, Multi-Attribute Expected Utility Theory & Decision Intelligence for Autonomous AI Agents
@putervision/agent-reasoning-mcp is a formal Model Context Protocol (MCP) server that provides strategic belief-desire-intention (BDI) reasoning, hierarchical goal decomposition, multi-attribute expected utility calculation ((E[U] = \sum w_i u_i)), exponential belief decay, quantitative risk evaluation, and reactive replanning across multi-modal memory bridges.
π Official Documentation: putervision.com β’ Interactive Web Docs
β‘ 15-Second Quick Start
# 1. Initialize reasoning database & seed default utility profiles
npx @putervision/agent-reasoning-mcp init
# 2. Run health diagnostics and Merkle audit checks
npx @putervision/agent-reasoning-mcp doctor
# 3. Inspect active goals, intentions, and belief states
npx @putervision/agent-reasoning-mcp inspectRelated MCP server: Summon-MCP
π οΈ 15 Core MCP Tools
BDI Strategic Deliberation (10 Tools)
Tool | Actions | Purpose |
|
| Manage goal hierarchy, task DAGs, and success criteria |
|
| Score and rank candidate actions from environment snapshots |
|
| Adaptively reconstruct subgoals upon obstacles and abort stale intentions |
|
| Quantitative threat and risk calculation across candidate actions |
|
| Search learned heuristics, tactical knowledge, and past decision patterns |
|
| Configure utility weights (aggression, caution, greed, efficiency, exploration) |
|
| Explainable chain-of-thought rationale and latency telemetry |
|
| Structured belief state with exponential confidence decay ($C = C_0 e^{-\lambda t}$) |
|
| Wire contract directives queue for runtime execution engines |
|
| Reasoning database statistics, SHA-256 Merkle audit, and snapshot rollback |
System 1 Fast Decision Layer (5 Tools)
Inspired by the typed System 1 pattern pioneered by TypeSafe's Jev (evaluating typed
Choice,Score, andNoulprimitives over compact state without token generation), implemented locally via in-memory LRU caches and deterministic heuristics (<2ms) without external API calls.
Tool | Purpose | Latency Target | L1 Cache (p50) | Throughput |
| Low-latency categorical labeling over multi-modal StatePacks |
|
| ~90,000 ops/s |
| Typed probabilistic hypothesis and Boolean verification ($p \in [0.0, 1.0]$) |
|
| ~127,000 ops/s |
| Discrete $1$-of-$N$ choice selection ($N \le 16$) with probability simplex |
|
| ~64,000 ops/s |
| Bounded numeric scalar scoring and calibrated utility rating |
|
| ~129,000 ops/s |
| Pre-dispatch blast-radius audit gate issuing signed HMAC dispatch tokens |
|
| ~12,000 ops/s |
See docs/benchmarks.md for full benchmark reproduction commands, latency percentiles (p50/p95/p99), and multi-tier caching architecture details.
ποΈ PuterVision Pentad Multi-Modal Ecosystem
agent-reasoning-mcp coordinates the closed-loop PuterVision Super-Loop:
π§
agent-reasoning-mcp: Decides what to do (BDI Strategic Reasoning, Utility Theory, Replanning)β‘
behavior-mcp: Executes how to act at ~60Hz in browser runtimesπ
state-memory-mcp: Durable workflow memory, tasks, blockers, decisionsποΈ
vision-memory-mcp: Perceptual caching, visual grounding, video timelinesπ
world-model-mcp: 3D/2D spatial layout, entity permanence, collision simulation
π Deep Documentation Guides
π Formal API Reference: Full parameter tables, type definitions, and tool schemas for all 15 tools.
π Performance Benchmarks: Empirical throughput and microsecond latency metrics across all 5 System 1 tools.
π‘ Core Architecture & Concepts: BDI model, utility formulation, and belief decay dynamics.
π₯οΈ CLI Usage Guide: Complete CLI command reference (
init,doctor,inspect,run).πΎ Database Schema: SQLite table structures, indexes, and Merkle audit ledger.
βοΈ Configuration Reference:
.agent-reasoning-mcp.jsonparameters and environment variables.
π Client Configuration & Environment
Add to .cursor/mcp.json or .vscode/mcp.json:
{
"mcpServers": {
"agent-reasoning-mcp": {
"command": "agent-reasoning-mcp",
"args": ["run"],
"env": {
"PENTAD_HMAC_SECRET": "your-secure-shared-secret-here",
"DISPATCH_TOKEN_TTL_MS": "30000"
}
}
}
}Key Environment Variables
PENTAD_HMAC_SECRET: 256-bit shared key for cryptographic intention dispatch token signing.DISPATCH_TOKEN_TTL_MS: Dispatch token expiration window (default: 30,000ms).SKIP_MODEL_LOAD: Set to1(orOFFLINE=1) to force air-gapped L1/L2 deterministic evaluation.
π§ͺ Testing & Benchmarks
# Run full unit and integration test suites
npm test
# Run System 1 fast decision layer benchmark suite (throughput & latency percentiles)
npm run benchmark
# Run air-gapped verification
OFFLINE=1 SKIP_MODEL_LOAD=1 npm testπ License
MIT Β© PuterVision
Available Tools
15 toolsask_choiceARead-onlyIdempotent
Select 1 option from a discrete set of alternatives (N <= 16) with probability distribution and utility margin. Use ask_choice instead of classify when choosing the best action or alternative under active utility profiles rather than categorizing an entity.
Returns selected option ID, probability distribution, and utility margin.
| Name | Required | Description | Default |
|---|---|---|---|
| options | Yes | Mutually exclusive options (strictly capped at 16) | |
| project | Yes | Target project slug | |
| question | Yes | Decision prompt | |
| state_pack | No | Optional explicit StatePack | |
| utility_profile | No | Named utility profile override |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, closed-world behavior, so the safety profile is covered. The description adds real behavioral context beyond them: selection is probabilistic, produces a probability distribution and a utility margin, and depends on active utility profiles. It does not explain how state_pack or utility_profile resolution affects determinism, which is the one remaining ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, correctly front-loaded with the core action and the sibling-differentiation rule. The return values are stated twice ('with probability distribution and utility margin' and 'Returns selected option ID, probability distribution, and utility margin'), which is redundant but not damaging.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates return values, which is necessary. However, for a tool with nested objects and a 5-parameter signature involving an opaque 'StatePack' and 'utility_profile override', the description never explains what those are or how they influence the choice, leaving conceptual gaps for a non-obvious domain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already carries parameter meaning and the baseline is 3. The description restates the 16-option cap (already expressed as maxItems and in the schema text) and adds nothing further about 'project', 'state_pack', or 'utility_profile'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Select 1 option from a discrete set of alternatives') and adds the scope constraint 'N <= 16' plus what it produces. It explicitly distinguishes itself from the sibling 'classify' by contrasting decision-under-utility with entity categorization.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit selection rule: use ask_choice rather than classify 'when choosing the best action or alternative under active utility profiles rather than categorizing an entity.' That covers the most confusable sibling, but other decision-oriented siblings (ask_score, ask_noul, evaluate_situation) are not addressed, so the routing guidance is clear yet incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ask_noulARead-onlyIdempotent
Evaluate whether a specific proposition is true given the current state pack with calibrated probability and abstain safeguards. Use ask_noul instead of ask_score when evaluating binary truth/falsehood rather than scoring an entity on a continuous scale.
Returns boolean answer, confidence probability, and abstain flag.
| Name | Required | Description | Default |
|---|---|---|---|
| prior | No | Optional Bayesian prior [0.0 - 1.0] | |
| project | Yes | Target project slug | |
| statement | Yes | Proposition to test | |
| state_pack | No | Optional explicit StatePack |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), so the bar is lower. The description adds genuine behavioral context beyond them: the result is calibrated rather than a raw score, it carries abstain safeguards, and it includes a confidence probability. It does not say when abstention triggers or how the prior influences the result, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and scope, followed by the sibling-routing rule and the return shape. Every sentence carries distinct information with no repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming the return fields (boolean answer, confidence, abstain flag), and it covers the key sibling alternative. Minor gaps remain around how abstention is decided and how the optional explicit state_pack relates to the 'current' one, but nothing blocking for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents project, statement, prior and state_pack. The description adds only the notion of a 'current state pack' and nothing about the prior's role or the expected statement phrasing, so it sits at the baseline for a fully documented schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: evaluating whether a proposition is true against the current state pack, with calibrated probability and abstain safeguards. It also differentiates itself from the sibling ask_score by naming the binary-vs-continuous distinction, so an agent can select correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative (ask_score) and the exact condition that selects this tool over it: binary truth/falsehood rather than scoring an entity on a continuous scale. That is a direct when-to-use routing rule, not an implied one.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ask_scoreARead-onlyIdempotent
Evaluate an entity, plan, or action on a bounded continuous scale against weighted criteria. Use ask_score instead of ask_noul when evaluating continuous numeric quality or fitness rather than binary truth.
Returns normalized score within scale bounds, criterion breakdown, and evaluation confidence.
| Name | Required | Description | Default |
|---|---|---|---|
| scale | No | [min, max] range (defaults to [0.0, 1.0]) | |
| metric | Yes | Metric name | |
| target | Yes | Subject to evaluate | |
| project | Yes | Target project slug | |
| criteria | No | Evaluation criteria | |
| state_pack | No | Optional explicit StatePack |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive, so the safety profile is covered. The description adds valuable context beyond them by disclosing the return shape (normalized score within scale bounds, criterion breakdown, evaluation confidence), which matters since no output schema exists. It does not mention auth or rate-limit behavior, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two well-formed sentences, with the core purpose front-loaded and the sibling comparison second. No filler, though the return-value sentence could be marginally tighter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only evaluation tool with a rich 6-parameter schema and no output schema, the description covers purpose, routing, and return contents. Minor gaps remain around the optional state_pack and criteria weighting semantics, but nothing needed to call it correctly is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters are self-documented; the baseline is 3. The description reinforces 'weighted criteria' and 'scale bounds' conceptually but adds no syntax, format, or default details beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (evaluate) and resource (entity, plan, or action) with the exact evaluation mode: a bounded continuous scale against weighted criteria. It explicitly names the sibling it differs from (ask_noul), so an agent can distinguish them without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit routing rule: use ask_score instead of ask_noul when the question is continuous numeric quality/fitness rather than binary truth. This is a when-to-use plus an alternative, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
assess_riskARead-onlyIdempotent
Compute quantitative risk and threat assessment for candidate actions, plans, or 3D spatial rollouts against active utility weights (actions: action, plan, compare, spatial_rollout). Use assess_risk instead of evaluate_situation when estimating failure probability and threat exposure rather than ranking overall utility.
Returns risk score (0.0-1.0), threat breakdown, and comparative risk ratings.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Risk assessment mode: action, plan, compare, spatial_rollout | |
| project | No | Target project slug | |
| obstacles | No | Obstacles with position/bounding_box and affordance_mask | |
| parameters | No | Action parameters | |
| trajectory | No | Array of [x, y, z] waypoints for spatial rollout risk assessment | |
| candidate_action | No | Action name to evaluate | |
| candidate_actions | No | Multiple actions to compare risk scores | |
| situation_context | No | Current environment telemetry & vitals | |
| clearance_threshold | No | Minimum clearance threshold in meters (default: 1.0) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so safety is covered. The description adds real behavioral context: the assessment is computed 'against active utility weights' (a state dependency) and it returns a 0.0-1.0 score with threat breakdown and comparative ratings. It does not describe cost, latency, or what happens when no weights are set, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly packed sentences with the purpose and routing rule front-loaded. The parenthetical enumeration of modes is mildly redundant with the schema enum, but it aids scanning and costs little.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter analyzer with nested objects and no output schema, the description covers purpose, routing, state dependency, and the shape of the return value (score range, threat breakdown, comparative ratings). Only edge behaviors, such as missing utility weights or invalid trajectories, are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all nine parameters including the mode enum, trajectory waypoints, and clearance_threshold default. The description's parenthetical mode list and mention of '3D spatial rollouts' restate the schema rather than adding format or constraint detail. Baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Compute quantitative risk and threat assessment') and scopes the accepted modes. It explicitly names the sibling it is not, so an agent can separate assess_risk from evaluate_situation without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit routing rule: 'Use assess_risk instead of evaluate_situation when estimating failure probability and threat exposure rather than ranking overall utility.' This names the alternative and the condition that selects it, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
classifyARead-onlyIdempotent
Assign semantic categorical labels to an entity, visual state, task, or state snapshot using deterministic System One calculus. Use classify instead of ask_choice when assigning predefined taxonomy labels rather than selecting among runtime decision alternatives.
Returns top class label, probability distribution, and classification margin.
| Name | Required | Description | Default |
|---|---|---|---|
| classes | Yes | Candidate classes (capped at 16) | |
| project | Yes | Target project slug | |
| target_id | No | Identifier of the target | |
| state_pack | No | Optional explicit StatePack | |
| target_type | Yes | Type of target to classify |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive, closed-world behavior. The description adds that the operation is deterministic (System One calculus) and specifies return values (top label, probability distribution, margin), which is valuable since no output schema exists. It does not cover error behavior or rate limits, but the annotation set carries the safety profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with purpose, followed by routing guidance and return values. Every sentence contributes new information without repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema, the description usefully states what is returned (top class label, probability distribution, classification margin). Annotations cover safety and idempotency, and the schema covers parameters fully. Minor gaps around state_pack semantics and error handling remain, but the core needs are met.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema itself documents all five parameters. The description mentions the target types broadly but adds no syntax, constraints, or meaning beyond what the schema already provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (assign semantic categorical labels) and resource scope (entity, visual_state, task, snapshot) and names the sibling tool it replaces (ask_choice). An agent can distinguish it from ask_choice and ask_score without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit when-to-use rule: assign predefined taxonomy labels rather than selecting among runtime decision alternatives, and names the alternative tool ask_choice. This is clear routing guidance with an implicit when-not condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evaluate_situationA
Ingest multi-modal situation snapshot, compute expected utilities against active weights, and output prioritized action recommendations (actions: snapshot, quick). Use evaluate_situation instead of assess_risk when ranking candidate actions across multi-attribute utility dimensions rather than calculating isolated threat probabilities.
Returns ranked candidate actions, expected utility scores, and top recommendation.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Snapshot evaluation mode or quick text context: snapshot, quick | |
| project | No | Target project slug | |
| snapshot | No | Normalized SituationSnapshot with world, vision, state, and vitals | |
| trace_id | No | Distributed trace ID | |
| session_id | No | Linked state-memory session ID | |
| quick_context | No | Text summary of current situation for quick evaluation | |
| lookahead_depth | No | Bounded heuristic lookahead plies (e.g. 2-3 plies, discount gamma=0.85) | |
| utility_profile | No | Named utility profile to score against (defaults to active) | |
| candidate_actions | No | Candidate actions to score and rank |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, idempotentHint=false, non-destructive and closed-world, so the safety profile is already supplied. The description adds useful return-shape context (ranked candidate actions, expected utility scores, top recommendation), but says nothing about what state is persisted or whether trace/session records are written, which matters for a non-read-only tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is front-loaded with the core verb sequence, then the routing rule, then the return values, with no filler sentences. The utility/decision-theory terminology is dense but each sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with nested objects, an enum mode switch and no output schema, the description does cover the return values and the key routing decision. It leaves the role of trace_id, session_id and project unaddressed, but those are documented in the schema, so the gap is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across 9 parameters, so the schema already explains action, snapshot, quick_context, lookahead_depth, utility_profile and candidate_actions. The description adds only the two enum mode names and the notion of scoring against active weights, which is marginal beyond the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific pipeline of verbs and resources: ingest a multi-modal snapshot, compute expected utilities against active weights, and output prioritized action recommendations, with the modes (snapshot, quick) called out. It also explicitly distinguishes itself from the sibling assess_risk, so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states an explicit routing rule: use evaluate_situation rather than assess_risk when ranking candidate actions across multi-attribute utility dimensions instead of computing isolated threat probabilities. That is a genuine when-to-use/when-not rule, though it covers only one sibling out of many and gives no prerequisites or preconditions for the modes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gate_intentionARead-onlyIdempotent
Evaluate an intention before dispatching to behavior-mcp and issue a cryptographic HMAC dispatch token if approved. Use gate_intention instead of manage_intentions when verifying precondition safety and issuing execution authorization rather than tracking intention state.
Returns gate verdict (approved/rejected), risk evaluation, and HMAC dispatch token.
| Name | Required | Description | Default |
|---|---|---|---|
| project | Yes | Target project slug | |
| state_pack | No | Optional explicit StatePack | |
| intention_id | No | Target intention ID | |
| context_goal_id | No | Active goal being pursued | |
| proposed_action | Yes | Proposed behavior action to execute |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint=false, openWorldHint=false), and the description adds real context beyond them: this is a precondition gate whose approval produces a cryptographic HMAC token, and it returns a verdict plus risk evaluation. It stops short of saying what a rejection looks like to the caller or whether the token has a lifetime/one-time-use constraint, so it is not fully transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, roughly fifty words, with the core purpose front-loaded in the first clause and no filler. The return-value sentence earns its place because no output schema exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so adequately (verdict, risk evaluation, HMAC token). Given a five-parameter nested schema and a gating role in a behavior-dispatch pipeline, it could say more about rejection handling and token use, but nothing essential to calling it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters (project, state_pack, intention_id, context_goal_id, proposed_action and its nested fields) are already documented structurally. The description references the intention and proposed action conceptually but adds no syntax, format, or interaction detail beyond the schema, which is the baseline-3 case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Evaluate an intention before dispatching to behavior-mcp') plus the concrete output ('issue a cryptographic HMAC dispatch token if approved'). It also explicitly names the sibling it is not, manage_intentions, so an agent can separate the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit routing rule: 'Use gate_intention instead of manage_intentions when verifying precondition safety and issuing execution authorization rather than tracking intention state.' This names the alternative and the condition that selects it, which is exactly the when/when-not guidance an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_decision_traceARead-onlyIdempotent
Retrieve explainable step-by-step chain-of-thought rationale, candidate utilities, and risk assessment for past decisions (actions: latest, get, list, explain). Use get_decision_trace instead of query_knowledge when performing deep forensic analysis of a specific historical decision.
Returns step-by-step reasoning trace, utility breakdown, candidate rankings, and explanation text.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max traces to list | |
| action | Yes | Trace retrieval operation: latest, get, list, explain | |
| goal_id | No | Filter traces by linked goal ID | |
| project | No | Target project slug | |
| trace_id | No | Trace ID for get/explain action |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and closed-world, so the safety profile is covered. The description adds value by enumerating the return structure (trace, utility breakdown, candidate rankings, explanation text), which matters since no output schema exists. It does not discuss filtering semantics or limits, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, purpose front-loaded and the sibling call-out second. The parenthetical action list duplicates the enum and the return enumeration slightly overlaps, but nothing is buried or wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description helpfully summarizes return values, and annotations carry the safety profile. What is missing is guidance on how action values interact with the other params (which params apply to get vs list), leaving a small gap for a 5-param tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are already documented, including the action enum. The description only restates the action values parenthetically and adds no syntax or interaction detail (e.g., trace_id required for get/explain), so baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Retrieve) and resource (decision trace) plus the payload it contains (chain-of-thought rationale, candidate utilities, risk assessment). It explicitly contrasts itself with query_knowledge, so an agent can distinguish it from siblings without schema inspection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternative (query_knowledge) and the precise condition that selects this tool: deep forensic analysis of a specific historical decision. The action set (latest, get, list, explain) further signals usage modes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_beliefsADestructive
Maintain structured belief state with TTL expiration sweeps, exponential confidence decay, and category filtering (actions: update, query, expire, reconcile, reconcile_spatial). Use manage_beliefs instead of query_knowledge when managing dynamic agent epistemic state rather than static heuristic patterns.
Returns belief record, query matches, expired belief count, or reconciliation report.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Belief operation: update, query, expire, reconcile, reconcile_spatial | |
| object | No | Belief value / state payload | |
| source | No | Belief provenance source | |
| project | No | Target project slug | |
| subject | No | Belief subject (e.g. "north_gate", "enemy_patrol") | |
| category | No | Belief category | |
| belief_id | No | Belief ID for specific lookup | |
| predicate | No | Predicate relationship (e.g. "is_locked", "status") | |
| confidence | No | Confidence score (0.0 to 1.0) | |
| decay_rate | No | Exponential decay rate lambda per hour | |
| expires_at | No | ISO-8601 expiration timestamp | |
| novel_entities | No | Novel entities for reconcile_spatial action | |
| matched_entities | No | Matched entities for reconcile_spatial action | |
| missing_entities | No | Missing entities for reconcile_spatial action | |
| client_request_id | No | Idempotency key | |
| unexpected_entities | No | Unexpected entities for reconcile_spatial action |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false, and the description adds real behavioral context beyond them: TTL expiration sweeps, exponential confidence decay, and category filtering, plus the shape of the return value. It still doesn't say what the destructive expire path actually removes or how update overwrites prior state, so it falls short of full disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the capability and the sibling comparison, then closes with the return shapes. Two focused sentences with little waste; the return sentence earns its place because there is no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 16-parameter, five-action mutation tool with no output schema, the description compensates by listing return values (belief record, query matches, expired count, reconciliation report) and the action set. The remaining gap is that it never indicates which parameters are required by which action, which matters at this parameter count.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 16 parameters are already documented in the schema and the baseline is 3. The description adds the action vocabulary but no per-action parameter mapping or format detail (e.g., which fields reconcile_spatial requires) beyond what the schema states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Maintain") and resource ("structured belief state") and enumerates the five supported actions, so the scope is unambiguous. It also explicitly distinguishes itself from the sibling query_knowledge, letting an agent route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative tool (query_knowledge) and the condition that selects this one instead (dynamic agent epistemic state vs static heuristic patterns). That is a direct when-to-use/when-not-to-use rule rather than implied guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_intentionsA
Queue, dispatch, track, cancel, or resolve behavior directives for behavior-mcp runtime execution (actions: create, dispatch, get, list, cancel, resolve). Use manage_intentions instead of set_goal when dispatching immediate execution instructions to runtime behaviors rather than managing abstract objectives.
Returns intention record, dispatch status, wire contract payload, or intention list.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Intention operation: create, dispatch, get, list, cancel, resolve | |
| result | No | Outcome payload for resolve action | |
| status | No | Status filter or update | |
| goal_id | No | Linked goal ID | |
| project | No | Target project slug | |
| priority | No | Execution priority (0.0 to 1.0) | |
| trace_id | No | Linked decision trace ID | |
| parameters | No | Runtime behavior parameters | |
| deadline_at | No | ISO-8601 completion deadline | |
| intention_id | No | Intention ID for dispatch/get/cancel/resolve | |
| behavior_name | No | Target behavior tree name (e.g. "combat_kite", "gather_loop") | |
| abort_conditions | No | Auto-abort trigger conditions | |
| client_request_id | No | Idempotency key |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), so the agent knows this is a non-destructive, non-idempotent mutation surface. The description adds the multi-action nature and the return categories, but does not disclose side effects of cancel/resolve, permission needs, or why the operation is non-idempotent despite having a client_request_id idempotency key.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action verbs and resource, followed by the routing rule and the return-value summary. Tight and well-ordered, with only slight redundancy between the action list in prose and the enum in the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter multi-action tool with nested objects and no output schema, the description covers purpose, routing, and return categories, and the fully-described schema fills in the per-action parameter requirements. Minor gap: it does not map which parameters apply to which action (e.g., intention_id for dispatch/get/cancel/resolve).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented in the schema; the description adds no per-parameter meaning beyond what is structured. Baseline 3 is appropriate when the schema carries the semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States specific verbs (queue, dispatch, track, cancel, resolve) applied to a clear resource (behavior directives / intentions) and enumerates the six actions. It explicitly distinguishes itself from the sibling set_goal, so an agent can tell the two apart without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternative tool (set_goal) and the exact condition that selects this one instead: dispatching immediate execution instructions to runtime behaviors vs. managing abstract objectives. This is explicit when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_reasoning_dbADestructive
Database maintenance, diagnostics, SHA-256 Merkle audit verification, and snapshot management (actions: stats, audit, doctor, snapshot, diff, restore). Use manage_reasoning_db instead of manage_beliefs when performing SQLite storage integrity verification or database snapshot restore.
Returns database diagnostics, Merkle audit tree, snapshot metadata, or diff reports.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Snapshot name | |
| action | Yes | Database maintenance operation: stats, audit, doctor, snapshot, diff, restore | |
| project | No | Target project slug | |
| description | No | Description for snapshot |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false and idempotentHint=false, so the safety profile is covered by structured data. The description adds useful return-shape context ('Returns database diagnostics, Merkle audit tree, snapshot metadata, or diff reports'), but never warns that 'restore' overwrites existing database state or that actions carry differing risk β information an agent would want beyond the blunt destructive flag.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the capability summary followed by routing and return information; no filler. The parenthetical action list is slightly redundant with the schema enum but keeps the definition self-contained.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by stating what categories of results are returned, and annotations carry the safety profile. It is close to complete for a six-action multiplexed tool, though per-action behavior (especially restore semantics) remains unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the enum is fully documented in the schema, so the baseline is 3. The description only restates the action list and adds no syntax, format, or dependency details (e.g., that 'name'/'description' are only relevant to snapshot) beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource ('Database maintenance, diagnostics, SHA-256 Merkle audit verification, and snapshot management') and enumerates the six concrete actions, so an agent knows exactly what domain this covers. It also explicitly contrasts itself with the sibling manage_beliefs, making it distinguishable without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit routing rule: 'Use manage_reasoning_db instead of manage_beliefs when performing SQLite storage integrity verification or database snapshot restore.' That is a clear when-to-use-this-vs-alternative condition. It stops short of covering the other actions (stats, diff, doctor) or any when-not/exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_knowledgeARead-onlyIdempotent
Search learned heuristic patterns, tactics, and past decision traces by context similarity (actions: search, patterns, similar_situations). Use query_knowledge instead of get_decision_trace when retrieving generalized patterns across sessions rather than inspecting a single execution trace.
Returns matching heuristics, anti-patterns, tactics, and similarity scores.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max patterns to return | |
| query | No | Semantic search query string | |
| action | Yes | Knowledge query mode: search, patterns, similar_situations | |
| project | No | Target project slug | |
| context_tags | No | Filter by context tags | |
| pattern_type | No | Filter by pattern category |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive, and closed-world behavior. The description adds useful return context by naming the kinds of matches returned (heuristics, anti-patterns, tactics, similarity scores) and the cross-session scope, though it does not discuss pagination or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose and routing guidance, then states the return contents. Every sentence adds useful information without repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter read-only search tool with rich annotations, the description supplies purpose, usage routing, supported action modes, and the shape of returned results. With no output schema, the return summary is especially valuable and sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters are already documented in the schema, including the action and pattern_type enums. The description repeats the action modes but adds no syntax, format, or interaction details beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Search) and resources (learned heuristic patterns, tactics, past decision traces) and names the supported actions. It also explicitly distinguishes the tool from get_decision_trace by contrasting generalized cross-session patterns with a single execution trace.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit routing guidance: use query_knowledge instead of get_decision_trace when retrieving generalized patterns across sessions, rather than inspecting one execution trace. This covers when to use it, when to prefer an alternative, and the alternative itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
replanA
Regenerate sub-task DAG and adjust intentions upon unexpected obstacles or state changes (actions: blocker, event, full). Use replan instead of set_goal when recovering from execution blockers or environment shifts rather than creating new goals.
Returns replanned goal DAG, invalidated intentions, and newly synthesized subgoals.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Replanning trigger type: blocker, event, full | |
| goal_id | Yes | ID of goal to replan | |
| project | No | Target project slug | |
| trigger_event | No | Event description triggering replanning | |
| preserve_completed | No | Whether to preserve already completed subgoals | |
| blocker_description | No | Description of the obstacle or blocker encountered |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and idempotentHint=false, so the mutation profile is partly covered. The description adds real context beyond them by enumerating the trigger actions and describing what is produced/side-effected (replanned DAG, invalidated intentions, synthesized subgoals). It stops short of explaining permissions or whether invalidation is reversible, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences with clear routing guidance and no filler. Repeating the action list is mildly redundant with the schema enum, which keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter mutation tool with no output schema, the description usefully names the return payload and the alternative tool, and annotations cover safety. It is nearly complete, lacking only error/precondition behavior for when replanning fails.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters including the action enum are already documented in the schema. The description only restates the action values ('blocker, event, full') without adding semantics such as how 'full' differs from 'event'. Baseline 3 applies when the schema carries the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource β 'Regenerate sub-task DAG and adjust intentions' β and explicitly scopes the trigger to 'unexpected obstacles or state changes'. It also names and contrasts against the sibling set_goal, so an agent can distinguish it without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit when-to-use rule plus the alternative: 'Use replan instead of set_goal when recovering from execution blockers or environment shifts rather than creating new goals.' The condition that selects this tool over the sibling is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_goalA
Register, update, decompose, inspect, or abandon hierarchical goals and task DAGs (actions: create, update, decompose, get, list, abandon). Use set_goal instead of manage_intentions when formulating multi-step objectives and sub-goal DAGs rather than queueing concrete execution directives.
Returns goal record, sub-goal hierarchy, completion progress, or goal lists.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Goal ID (required for update, get, abandon) | |
| limit | No | Max items to return for list action | |
| title | No | Goal title or objective summary | |
| action | Yes | The goal management operation to perform: create, update, decompose, get, list, abandon | |
| status | No | Goal status | |
| project | No | Target project slug | |
| priority | No | Goal priority (0.0 to 1.0) | |
| progress | No | Completion progress (0.0 to 1.0) | |
| subgoals | No | Array of sub-goals for decompose action | |
| parent_id | No | Parent goal ID for hierarchical sub-goals | |
| deadline_at | No | ISO-8601 deadline timestamp | |
| description | No | Detailed goal description | |
| utility_weights | No | Goal-specific utility weight overrides | |
| success_criteria | No | List of verifiable conditions | |
| client_request_id | No | Idempotency key to prevent duplicate creation |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and idempotentHint=false, so the mutation profile is covered. The description adds an enumeration of return shapes (goal record, sub-goal hierarchy, progress, lists), which is useful since there is no output schema, but it says nothing about permissions, whether 'abandon' is reversible, or side effects of decompose. It adds some context but not rich behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the action set and the sibling routing rule, then closes with a brief returns sentence. It is appropriately sized for a multi-action tool, with no filler, though the opening verb list is slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter, nested-object, no-output-schema tool, the description covers the action space, sibling routing, and return shapes, which is most of what an agent needs. It stops short of per-action parameter guidance (e.g. which fields apply to decompose vs update), leaving a modest gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every one of the 15 parameters is already documented, including which actions require `id` and the action enum. The description only restates the action names and does not enrich parameter meaning beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (hierarchical goals and task DAGs) and enumerates the exact verb set (create, update, decompose, get, list, abandon), so an agent knows precisely what the tool does. It also distinguishes itself from the sibling manage_intentions, making it identifiable without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit routing rule: use set_goal rather than manage_intentions 'when formulating multi-step objectives and sub-goal DAGs rather than queueing concrete execution directives.' The alternative and the condition that selects it are both stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_utility_weightsA
Configure, inspect, activate, or incrementally nudge multi-attribute utility weight profiles (actions: configure, get, list, activate, nudge). Use set_utility_weights instead of evaluate_situation when defining decision preferences (aggression, caution, greed, exploration) rather than evaluating actions.
Returns configured utility profile, active weight map, or profile directory.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Profile name (e.g. "aggressive", "cautious", "explorer") | |
| delta | No | Key-value map of weight deltas for nudge action | |
| action | Yes | Profile operation: configure, get, list, activate, nudge | |
| project | No | Target project slug | |
| weights | No | Key-value map of weight values (0.0 to 1.0) | |
| is_active | No | Whether to set as currently active profile | |
| description | No | Profile description |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-read-only, non-destructive, non-idempotent mutation, so the safety profile is covered. The description adds the return surface (configured profile, active weight map, or profile directory), which is genuinely useful given no output schema exists, but says nothing about side effects of activate/nudge or auth requirements. Adequate, not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and action list, then the routing rule, then returns β a sensible ordering with no filler sentences. The action list is restated from the enum, a minor redundancy that keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter multi-action tool with nested objects and no output schema, the definition covers purpose, action routing, alternative selection, and the shape of returns. It is close to complete; the only gap is per-action behavioral detail for the mutating actions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter including the action enum is already documented in the schema; baseline 3 applies. The description adds no syntax, range, or interaction detail (e.g. how delta composes with weights) beyond what the schema states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (multi-attribute utility weight profiles) with an explicit verb set and enumerates the five supported actions, so the agent knows exactly what operations this tool performs. It also names the sibling it must not be confused with (evaluate_situation) and the distinction between them, which is more than most definitions provide.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly gives the selection rule: use this instead of evaluate_situation when defining decision preferences rather than evaluating actions. That is a clear when-to-use-vs-alternative statement. It does not, however, give per-action guidance (e.g. when to nudge vs configure), which keeps it short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.4.0- Changed
assess_risk5 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Risk assessment mode: action, plan, compare"New value: +"Risk assessment mode: action, plan, compare, spatial_rollout" - changed
Input schema / properties / action / enumPrevious value: -[ - "action", - "plan", - "compare" -]New value: +[ + "action", + "plan", + "compare", + "spatial_rollout" +] - added
Input schema / properties / clearance_thresholdAdded value: +{ + "description": "Minimum clearance threshold in meters (default: 1.0)", + "type": "number" +} - added
Input schema / properties / obstaclesAdded value: +{ + "description": "Obstacles with position/bounding_box and affordance_mask", + "items": {}, + "type": "array" +} - added
Input schema / properties / trajectoryAdded value: +{ + "description": "Array of [x, y, z] waypoints for spatial rollout risk assessment", + "items": { + "items": { + "type": "number" + }, + "type": "array" + }, + "type": "array" +}
- Changed
manage_beliefs6 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Belief operation: update, query, expire, reconcile"New value: +"Belief operation: update, query, expire, reconcile, reconcile_spatial" - changed
Input schema / properties / action / enumPrevious value: -[ - "update", - "query", - "expire", - "reconcile" -]New value: +[ + "update", + "query", + "expire", + "reconcile", + "reconcile_spatial" +] - added
Input schema / properties / matched_entitiesAdded value: +{ + "description": "Matched entities for reconcile_spatial action", + "items": {}, + "type": "array" +} - added
Input schema / properties / missing_entitiesAdded value: +{ + "description": "Missing entities for reconcile_spatial action", + "items": {}, + "type": "array" +} - added
Input schema / properties / novel_entitiesAdded value: +{ + "description": "Novel entities for reconcile_spatial action", + "items": {}, + "type": "array" +} - added
Input schema / properties / unexpected_entitiesAdded value: +{ + "description": "Unexpected entities for reconcile_spatial action", + "items": {}, + "type": "array" +}
- Changed
set_utility_weights3 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Profile operation: configure, get, list, activate"New value: +"Profile operation: configure, get, list, activate, nudge" - changed
Input schema / properties / action / enumPrevious value: -[ - "configure", - "get", - "list", - "activate" -]New value: +[ + "configure", + "get", + "list", + "activate", + "nudge" +] - added
Input schema / properties / deltaAdded value: +{ + "description": "Key-value map of weight deltas for nudge action", + "type": "object" +}
15 tool updates
v0.3.1- Added
ask_choice - Added
ask_noul - Added
ask_score - Changed
assess_risk2 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Risk assessment mode"New value: +"Risk assessment mode: action, plan, compare" - added
Input schema / properties / action / enumAdded value: +[ + "action", + "plan", + "compare" +]
- Added
classify - Changed
evaluate_situation2 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Snapshot evaluation mode or quick text context"New value: +"Snapshot evaluation mode or quick text context: snapshot, quick" - added
Input schema / properties / action / enumAdded value: +[ + "snapshot", + "quick" +]
- Added
gate_intention - Changed
get_decision_trace2 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Trace retrieval operation"New value: +"Trace retrieval operation: latest, get, list, explain" - added
Input schema / properties / action / enumAdded value: +[ + "latest", + "get", + "list", + "explain" +]
- Changed
manage_beliefs2 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Belief operation"New value: +"Belief operation: update, query, expire, reconcile" - added
Input schema / properties / action / enumAdded value: +[ + "update", + "query", + "expire", + "reconcile" +]
- Changed
manage_intentions2 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Intention operation"New value: +"Intention operation: create, dispatch, get, list, cancel, resolve" - added
Input schema / properties / action / enumAdded value: +[ + "create", + "dispatch", + "get", + "list", + "cancel", + "resolve" +]
- Changed
manage_reasoning_db2 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Database maintenance operation: stats, audit, doctor (health diagnostics), snapshot, diff, restore"New value: +"Database maintenance operation: stats, audit, doctor, snapshot, diff, restore" - added
Input schema / properties / action / enumAdded value: +[ + "stats", + "audit", + "doctor", + "snapshot", + "diff", + "restore" +]
- Changed
query_knowledge2 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Knowledge query mode"New value: +"Knowledge query mode: search, patterns, similar_situations" - added
Input schema / properties / action / enumAdded value: +[ + "search", + "patterns", + "similar_situations" +]
- Changed
replan2 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Replanning trigger type"New value: +"Replanning trigger type: blocker, event, full" - added
Input schema / properties / action / enumAdded value: +[ + "blocker", + "event", + "full" +]
- Changed
set_goal2 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"The goal management operation to perform"New value: +"The goal management operation to perform: create, update, decompose, get, list, abandon" - added
Input schema / properties / action / enumAdded value: +[ + "create", + "update", + "decompose", + "get", + "list", + "abandon" +]
- Changed
set_utility_weights2 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Profile operation"New value: +"Profile operation: configure, get, list, activate" - added
Input schema / properties / action / enumAdded value: +[ + "configure", + "get", + "list", + "activate" +]
10 tool updates
v0.2.1- Changed
assess_risk16 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / action / descriptionAdded value: +"Risk assessment mode" - removed
Input schema / properties / action / enumRemoved value: -[ - "action", - "plan", - "compare" -] - added
Input schema / properties / candidate_action / descriptionAdded value: +"Action name to evaluate" - added
Input schema / properties / candidate_actions / descriptionAdded value: +"Multiple actions to compare risk scores" - removed
Input schema / properties / candidate_actions / items / additionalPropertiesRemoved value: -true - removed
Input schema / properties / candidate_actions / items / properties / parameters / additionalPropertiesRemoved value: -true - removed
Input schema / properties / candidate_actions / items / properties / parameters / propertiesRemoved value: -{} - removed
Input schema / properties / parameters / additionalPropertiesRemoved value: -true - added
Input schema / properties / parameters / descriptionAdded value: +"Action parameters" - removed
Input schema / properties / parameters / propertiesRemoved value: -{} - added
Input schema / properties / project / descriptionAdded value: +"Target project slug" - removed
Input schema / properties / situation_context / additionalPropertiesRemoved value: -true - added
Input schema / properties / situation_context / descriptionAdded value: +"Current environment telemetry & vitals" - removed
Input schema / properties / situation_context / propertiesRemoved value: -{}
- Changed
evaluate_situation17 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / action / descriptionAdded value: +"Snapshot evaluation mode or quick text context" - removed
Input schema / properties / action / enumRemoved value: -[ - "snapshot", - "quick" -] - added
Input schema / properties / candidate_actions / descriptionAdded value: +"Candidate actions to score and rank" - removed
Input schema / properties / candidate_actions / items / additionalPropertiesRemoved value: -true - removed
Input schema / properties / candidate_actions / items / properties / parameters / additionalPropertiesRemoved value: -true - removed
Input schema / properties / candidate_actions / items / properties / parameters / propertiesRemoved value: -{} - added
Input schema / properties / lookahead_depthAdded value: +{ + "description": "Bounded heuristic lookahead plies (e.g. 2-3 plies, discount gamma=0.85)", + "type": "number" +} - added
Input schema / properties / project / descriptionAdded value: +"Target project slug" - added
Input schema / properties / quick_context / descriptionAdded value: +"Text summary of current situation for quick evaluation" - added
Input schema / properties / session_id / descriptionAdded value: +"Linked state-memory session ID" - removed
Input schema / properties / snapshot / additionalPropertiesRemoved value: -true - added
Input schema / properties / snapshot / descriptionAdded value: +"Normalized SituationSnapshot with world, vision, state, and vitals" - removed
Input schema / properties / snapshot / propertiesRemoved value: -{} - added
Input schema / properties / trace_id / descriptionAdded value: +"Distributed trace ID" - added
Input schema / properties / utility_profile / descriptionAdded value: +"Named utility profile to score against (defaults to active)"
- Changed
get_decision_trace8 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / action / descriptionAdded value: +"Trace retrieval operation" - removed
Input schema / properties / action / enumRemoved value: -[ - "latest", - "get", - "list", - "explain" -] - added
Input schema / properties / goal_id / descriptionAdded value: +"Filter traces by linked goal ID" - added
Input schema / properties / limit / descriptionAdded value: +"Max traces to list" - added
Input schema / properties / project / descriptionAdded value: +"Target project slug" - added
Input schema / properties / trace_id / descriptionAdded value: +"Trace ID for get/explain action"
- Changed
manage_beliefs15 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / action / descriptionAdded value: +"Belief operation" - removed
Input schema / properties / action / enumRemoved value: -[ - "update", - "query", - "expire", - "reconcile" -] - added
Input schema / properties / belief_id / descriptionAdded value: +"Belief ID for specific lookup" - added
Input schema / properties / category / descriptionAdded value: +"Belief category" - added
Input schema / properties / client_request_id / descriptionAdded value: +"Idempotency key" - added
Input schema / properties / confidence / descriptionAdded value: +"Confidence score (0.0 to 1.0)" - added
Input schema / properties / decay_rate / descriptionAdded value: +"Exponential decay rate lambda per hour" - added
Input schema / properties / expires_at / descriptionAdded value: +"ISO-8601 expiration timestamp" - added
Input schema / properties / object / descriptionAdded value: +"Belief value / state payload" - added
Input schema / properties / predicate / descriptionAdded value: +"Predicate relationship (e.g. \"is_locked\", \"status\")" - added
Input schema / properties / project / descriptionAdded value: +"Target project slug" - added
Input schema / properties / source / descriptionAdded value: +"Belief provenance source" - added
Input schema / properties / subject / descriptionAdded value: +"Belief subject (e.g. \"north_gate\", \"enemy_patrol\")"
- Changed
manage_intentions22 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / abort_conditions / descriptionAdded value: +"Auto-abort trigger conditions" - removed
Input schema / properties / abort_conditions / items / additionalPropertiesRemoved value: -true - removed
Input schema / properties / abort_conditions / items / propertiesRemoved value: -{} - added
Input schema / properties / action / descriptionAdded value: +"Intention operation" - removed
Input schema / properties / action / enumRemoved value: -[ - "create", - "dispatch", - "get", - "list", - "cancel", - "resolve" -] - added
Input schema / properties / behavior_name / descriptionAdded value: +"Target behavior tree name (e.g. \"combat_kite\", \"gather_loop\")" - added
Input schema / properties / client_request_id / descriptionAdded value: +"Idempotency key" - added
Input schema / properties / deadline_at / descriptionAdded value: +"ISO-8601 completion deadline" - added
Input schema / properties / goal_id / descriptionAdded value: +"Linked goal ID" - added
Input schema / properties / intention_id / descriptionAdded value: +"Intention ID for dispatch/get/cancel/resolve" - removed
Input schema / properties / parameters / additionalPropertiesRemoved value: -true - added
Input schema / properties / parameters / descriptionAdded value: +"Runtime behavior parameters" - removed
Input schema / properties / parameters / propertiesRemoved value: -{} - added
Input schema / properties / priority / descriptionAdded value: +"Execution priority (0.0 to 1.0)" - added
Input schema / properties / project / descriptionAdded value: +"Target project slug" - removed
Input schema / properties / result / additionalPropertiesRemoved value: -true - added
Input schema / properties / result / descriptionAdded value: +"Outcome payload for resolve action" - removed
Input schema / properties / result / propertiesRemoved value: -{} - added
Input schema / properties / status / descriptionAdded value: +"Status filter or update" - added
Input schema / properties / trace_id / descriptionAdded value: +"Linked decision trace ID"
- Changed
manage_reasoning_db7 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / action / descriptionAdded value: +"Database maintenance operation: stats, audit, doctor (health diagnostics), snapshot, diff, restore" - removed
Input schema / properties / action / enumRemoved value: -[ - "stats", - "audit", - "snapshot", - "diff", - "restore" -] - added
Input schema / properties / description / descriptionAdded value: +"Description for snapshot" - added
Input schema / properties / name / descriptionAdded value: +"Snapshot name" - added
Input schema / properties / project / descriptionAdded value: +"Target project slug"
- Changed
query_knowledge9 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / action / descriptionAdded value: +"Knowledge query mode" - removed
Input schema / properties / action / enumRemoved value: -[ - "search", - "patterns", - "similar_situations" -] - added
Input schema / properties / context_tags / descriptionAdded value: +"Filter by context tags" - added
Input schema / properties / limit / descriptionAdded value: +"Max patterns to return" - added
Input schema / properties / pattern_type / descriptionAdded value: +"Filter by pattern category" - added
Input schema / properties / project / descriptionAdded value: +"Target project slug" - added
Input schema / properties / query / descriptionAdded value: +"Semantic search query string"
- Changed
replan9 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / action / descriptionAdded value: +"Replanning trigger type" - removed
Input schema / properties / action / enumRemoved value: -[ - "blocker", - "event", - "full" -] - added
Input schema / properties / blocker_description / descriptionAdded value: +"Description of the obstacle or blocker encountered" - added
Input schema / properties / goal_id / descriptionAdded value: +"ID of goal to replan" - added
Input schema / properties / preserve_completed / descriptionAdded value: +"Whether to preserve already completed subgoals" - added
Input schema / properties / project / descriptionAdded value: +"Target project slug" - added
Input schema / properties / trigger_event / descriptionAdded value: +"Event description triggering replanning"
- Changed
set_goal21 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / action / descriptionAdded value: +"The goal management operation to perform" - removed
Input schema / properties / action / enumRemoved value: -[ - "create", - "update", - "decompose", - "get", - "list", - "abandon" -] - added
Input schema / properties / client_request_id / descriptionAdded value: +"Idempotency key to prevent duplicate creation" - added
Input schema / properties / deadline_at / descriptionAdded value: +"ISO-8601 deadline timestamp" - added
Input schema / properties / description / descriptionAdded value: +"Detailed goal description" - added
Input schema / properties / id / descriptionAdded value: +"Goal ID (required for update, get, abandon)" - added
Input schema / properties / limit / descriptionAdded value: +"Max items to return for list action" - added
Input schema / properties / parent_id / descriptionAdded value: +"Parent goal ID for hierarchical sub-goals" - added
Input schema / properties / priority / descriptionAdded value: +"Goal priority (0.0 to 1.0)" - added
Input schema / properties / progress / descriptionAdded value: +"Completion progress (0.0 to 1.0)" - added
Input schema / properties / project / descriptionAdded value: +"Target project slug" - added
Input schema / properties / status / descriptionAdded value: +"Goal status" - added
Input schema / properties / subgoals / descriptionAdded value: +"Array of sub-goals for decompose action" - removed
Input schema / properties / subgoals / items / additionalPropertiesRemoved value: -true - added
Input schema / properties / success_criteria / descriptionAdded value: +"List of verifiable conditions" - added
Input schema / properties / title / descriptionAdded value: +"Goal title or objective summary" - removed
Input schema / properties / utility_weights / additionalPropertiesRemoved value: -true - added
Input schema / properties / utility_weights / descriptionAdded value: +"Goal-specific utility weight overrides" - removed
Input schema / properties / utility_weights / propertiesRemoved value: -{}
- Changed
set_utility_weights11 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / action / descriptionAdded value: +"Profile operation" - removed
Input schema / properties / action / enumRemoved value: -[ - "configure", - "get", - "list", - "activate" -] - added
Input schema / properties / description / descriptionAdded value: +"Profile description" - added
Input schema / properties / is_active / descriptionAdded value: +"Whether to set as currently active profile" - added
Input schema / properties / name / descriptionAdded value: +"Profile name (e.g. \"aggressive\", \"cautious\", \"explorer\")" - added
Input schema / properties / project / descriptionAdded value: +"Target project slug" - removed
Input schema / properties / weights / additionalPropertiesRemoved value: -true - added
Input schema / properties / weights / descriptionAdded value: +"Key-value map of weight values (0.0 to 1.0)" - removed
Input schema / properties / weights / propertiesRemoved value: -{}
10 tool updates
v0.1.2- First observed
assess_risk - First observed
evaluate_situation - First observed
get_decision_trace - First observed
manage_beliefs - First observed
manage_intentions - First observed
manage_reasoning_db - First observed
query_knowledge - First observed
replan - First observed
set_goal - First observed
set_utility_weights
TDQS
Scored across 15 tools
Most tools target distinct concepts, and the descriptions actively disambiguate with explicit 'use X instead of Y' guidance (e.g. assess_risk vs evaluate_situation, set_goal vs manage_intentions). However there is a dense cluster of risk/utility evaluation tools (assess_risk, evaluate_situation, ask_score, gate_intention) that an agent could still misselect among.
The set consistently uses snake_case verb_noun patterns (manage_beliefs, query_knowledge, get_decision_trace, set_utility_weights), with coherent families like ask_* and manage_*. Minor deviations are bare verbs (replan, classify) and an unusual token (ask_noul), but overall the convention is predictable.
15 tools sit at the upper edge of the comfortable 3-15 range but are justified by a genuinely broad cognitive-architecture domain (risk, beliefs, goals, intentions, knowledge, DB ops). It is slightly heavy but each tool maps to a real capability rather than being filler.
Coverage is strong: beliefs (update/query/expire/reconcile), goals (create/update/decompose/get/list/abandon), intentions (create/dispatch/get/list/cancel/resolve) and traces/knowledge all form coherent lifecycles. Gaps are minorβno obvious delete or bulk-export for some resourcesβbut core reasoning workflows are well covered.
Maintenance
Related MCP Connectors
Deterministic reasoning stack for AI agents: simulate, decide & compute, plus cross-domain tools.
Durable, user-controlled goals and governed plans for AI agents.
Goal and task planning MCP for Codex and AI agents, with evidence-backed completion.
Deterministic multi-criteria decision analysis for AI agents β score, rank & explain options.
Related MCP Servers
- AlicenseAqualityCmaintenanceCampaign orchestration MCP server for AI agents β dependency DAGs, parallel fan-out, failure policies, and embedded notes.21GPL 3.0
- AlicenseNot gradedqualityDmaintenanceObjective-driven cognitive architecture engine that builds single experts, councils, or full autonomous organizations from user goals, generating deployment-ready superprompts and configurations.8 npmMIT
- -licenseNot gradedqualityNot gradedmaintenanceGoal-oriented narrative state machine for AI agents, exposing live world/scene context as MCP tools and resources.-
- FlicenseAqualityCmaintenanceEnables AI agents to hand high-level objectives to a decision-making core that autonomously reasons, plans, enforces deterministic policy, executes capabilities, evaluates outcomes, and persists semantic memory over stdio or Streamable HTTP.6-