Skip to main content
Glama
putervision

agent-reasoning-mcp

by putervision

@putervision/agent-reasoning-mcp

npm version version CI Node License: MIT

Strategic BDI Reasoning, Multi-Attribute Expected Utility Theory & Decision Intelligence for Autonomous AI Agents

@putervision/agent-reasoning-mcp is a formal Model Context Protocol (MCP) server that provides strategic belief-desire-intention (BDI) reasoning, hierarchical goal decomposition, multi-attribute expected utility calculation ((E[U] = \sum w_i u_i)), exponential belief decay, quantitative risk evaluation, and reactive replanning across multi-modal memory bridges.

🌐 Official Documentation: putervision.com β€’ Interactive Web Docs


⚑ 15-Second Quick Start

# 1. Initialize reasoning database & seed default utility profiles
npx @putervision/agent-reasoning-mcp init

# 2. Run health diagnostics and Merkle audit checks
npx @putervision/agent-reasoning-mcp doctor

# 3. Inspect active goals, intentions, and belief states
npx @putervision/agent-reasoning-mcp inspect

Related MCP server: Summon-MCP

πŸ› οΈ 15 Core MCP Tools

BDI Strategic Deliberation (10 Tools)

Tool

Actions

Purpose

set_goal

create, update, decompose, get, list, abandon

Manage goal hierarchy, task DAGs, and success criteria

evaluate_situation

snapshot, quick

Score and rank candidate actions from environment snapshots

replan

blocker, event, full

Adaptively reconstruct subgoals upon obstacles and abort stale intentions

assess_risk

action, plan, compare

Quantitative threat and risk calculation across candidate actions

query_knowledge

search, patterns, similar_situations

Search learned heuristics, tactical knowledge, and past decision patterns

set_utility_weights

configure, get, list, activate

Configure utility weights (aggression, caution, greed, efficiency, exploration)

get_decision_trace

latest, get, list, explain

Explainable chain-of-thought rationale and latency telemetry

manage_beliefs

update, query, expire, reconcile

Structured belief state with exponential confidence decay ($C = C_0 e^{-\lambda t}$)

manage_intentions

create, dispatch, get, list, cancel, resolve

Wire contract directives queue for runtime execution engines

manage_reasoning_db

stats, audit, snapshot, restore

Reasoning database statistics, SHA-256 Merkle audit, and snapshot rollback

System 1 Fast Decision Layer (5 Tools)

Inspired by the typed System 1 pattern pioneered by TypeSafe's Jev (evaluating typed Choice, Score, and Noul primitives over compact state without token generation), implemented locally via in-memory LRU caches and deterministic heuristics (<2ms) without external API calls.

Tool

Purpose

Latency Target

L1 Cache (p50)

Throughput

classify

Low-latency categorical labeling over multi-modal StatePacks

<2ms

0.0075 ms

~90,000 ops/s

ask_noul

Typed probabilistic hypothesis and Boolean verification ($p \in [0.0, 1.0]$)

<2ms

0.0049 ms

~127,000 ops/s

ask_choice

Discrete $1$-of-$N$ choice selection ($N \le 16$) with probability simplex

<2ms

0.0138 ms

~64,000 ops/s

ask_score

Bounded numeric scalar scoring and calibrated utility rating

<2ms

0.0057 ms

~129,000 ops/s

gate_intention

Pre-dispatch blast-radius audit gate issuing signed HMAC dispatch tokens

<1ms

0.0709 ms

~12,000 ops/s

See docs/benchmarks.md for full benchmark reproduction commands, latency percentiles (p50/p95/p99), and multi-tier caching architecture details.


πŸ›οΈ PuterVision Pentad Multi-Modal Ecosystem

agent-reasoning-mcp coordinates the closed-loop PuterVision Super-Loop:

  • 🧠 agent-reasoning-mcp: Decides what to do (BDI Strategic Reasoning, Utility Theory, Replanning)

  • ⚑ behavior-mcp: Executes how to act at ~60Hz in browser runtimes

  • πŸ“Š state-memory-mcp: Durable workflow memory, tasks, blockers, decisions

  • πŸ‘οΈ vision-memory-mcp: Perceptual caching, visual grounding, video timelines

  • 🌐 world-model-mcp: 3D/2D spatial layout, entity permanence, collision simulation


πŸ“š Deep Documentation Guides


πŸ”— Client Configuration & Environment

Add to .cursor/mcp.json or .vscode/mcp.json:

{
  "mcpServers": {
    "agent-reasoning-mcp": {
      "command": "agent-reasoning-mcp",
      "args": ["run"],
      "env": {
        "PENTAD_HMAC_SECRET": "your-secure-shared-secret-here",
        "DISPATCH_TOKEN_TTL_MS": "30000"
      }
    }
  }
}

Key Environment Variables

  • PENTAD_HMAC_SECRET: 256-bit shared key for cryptographic intention dispatch token signing.

  • DISPATCH_TOKEN_TTL_MS: Dispatch token expiration window (default: 30,000ms).

  • SKIP_MODEL_LOAD: Set to 1 (or OFFLINE=1) to force air-gapped L1/L2 deterministic evaluation.


πŸ§ͺ Testing & Benchmarks

# Run full unit and integration test suites
npm test

# Run System 1 fast decision layer benchmark suite (throughput & latency percentiles)
npm run benchmark

# Run air-gapped verification
OFFLINE=1 SKIP_MODEL_LOAD=1 npm test

πŸ“„ License

MIT Β© PuterVision

Available Tools

15 tools
ask_choiceA
Read-onlyIdempotent

Select 1 option from a discrete set of alternatives (N <= 16) with probability distribution and utility margin. Use ask_choice instead of classify when choosing the best action or alternative under active utility profiles rather than categorizing an entity.

Returns selected option ID, probability distribution, and utility margin.

ParametersJSON Schema
NameRequiredDescriptionDefault
optionsYesMutually exclusive options (strictly capped at 16)
projectYesTarget project slug
questionYesDecision prompt
state_packNoOptional explicit StatePack
utility_profileNoNamed utility profile override

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, closed-world behavior, so the safety profile is covered. The description adds real behavioral context beyond them: selection is probabilistic, produces a probability distribution and a utility margin, and depends on active utility profiles. It does not explain how state_pack or utility_profile resolution affects determinism, which is the one remaining ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, correctly front-loaded with the core action and the sibling-differentiation rule. The return values are stated twice ('with probability distribution and utility margin' and 'Returns selected option ID, probability distribution, and utility margin'), which is redundant but not damaging.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates return values, which is necessary. However, for a tool with nested objects and a 5-parameter signature involving an opaque 'StatePack' and 'utility_profile override', the description never explains what those are or how they influence the choice, leaving conceptual gaps for a non-obvious domain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already carries parameter meaning and the baseline is 3. The description restates the 16-option cap (already expressed as maxItems and in the schema text) and adds nothing further about 'project', 'state_pack', or 'utility_profile'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Select 1 option from a discrete set of alternatives') and adds the scope constraint 'N <= 16' plus what it produces. It explicitly distinguishes itself from the sibling 'classify' by contrasting decision-under-utility with entity categorization.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit selection rule: use ask_choice rather than classify 'when choosing the best action or alternative under active utility profiles rather than categorizing an entity.' That covers the most confusable sibling, but other decision-oriented siblings (ask_score, ask_noul, evaluate_situation) are not addressed, so the routing guidance is clear yet incomplete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ask_noulA
Read-onlyIdempotent

Evaluate whether a specific proposition is true given the current state pack with calibrated probability and abstain safeguards. Use ask_noul instead of ask_score when evaluating binary truth/falsehood rather than scoring an entity on a continuous scale.

Returns boolean answer, confidence probability, and abstain flag.

ParametersJSON Schema
NameRequiredDescriptionDefault
priorNoOptional Bayesian prior [0.0 - 1.0]
projectYesTarget project slug
statementYesProposition to test
state_packNoOptional explicit StatePack

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), so the bar is lower. The description adds genuine behavioral context beyond them: the result is calibrated rather than a raw score, it carries abstain safeguards, and it includes a confidence probability. It does not say when abstention triggers or how the prior influences the result, which keeps it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action and scope, followed by the sibling-routing rule and the return shape. Every sentence carries distinct information with no repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by naming the return fields (boolean answer, confidence, abstain flag), and it covers the key sibling alternative. Minor gaps remain around how abstention is decided and how the optional explicit state_pack relates to the 'current' one, but nothing blocking for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents project, statement, prior and state_pack. The description adds only the notion of a 'current state pack' and nothing about the prior's role or the expected statement phrasing, so it sits at the baseline for a fully documented schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: evaluating whether a proposition is true against the current state pack, with calibrated probability and abstain safeguards. It also differentiates itself from the sibling ask_score by naming the binary-vs-continuous distinction, so an agent can select correctly without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative (ask_score) and the exact condition that selects this tool over it: binary truth/falsehood rather than scoring an entity on a continuous scale. That is a direct when-to-use routing rule, not an implied one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ask_scoreA
Read-onlyIdempotent

Evaluate an entity, plan, or action on a bounded continuous scale against weighted criteria. Use ask_score instead of ask_noul when evaluating continuous numeric quality or fitness rather than binary truth.

Returns normalized score within scale bounds, criterion breakdown, and evaluation confidence.

ParametersJSON Schema
NameRequiredDescriptionDefault
scaleNo[min, max] range (defaults to [0.0, 1.0])
metricYesMetric name
targetYesSubject to evaluate
projectYesTarget project slug
criteriaNoEvaluation criteria
state_packNoOptional explicit StatePack

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and non-destructive, so the safety profile is covered. The description adds valuable context beyond them by disclosing the return shape (normalized score within scale bounds, criterion breakdown, evaluation confidence), which matters since no output schema exists. It does not mention auth or rate-limit behavior, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two well-formed sentences, with the core purpose front-loaded and the sibling comparison second. No filler, though the return-value sentence could be marginally tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only evaluation tool with a rich 6-parameter schema and no output schema, the description covers purpose, routing, and return contents. Minor gaps remain around the optional state_pack and criteria weighting semantics, but nothing needed to call it correctly is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters are self-documented; the baseline is 3. The description reinforces 'weighted criteria' and 'scale bounds' conceptually but adds no syntax, format, or default details beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (evaluate) and resource (entity, plan, or action) with the exact evaluation mode: a bounded continuous scale against weighted criteria. It explicitly names the sibling it differs from (ask_noul), so an agent can distinguish them without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit routing rule: use ask_score instead of ask_noul when the question is continuous numeric quality/fitness rather than binary truth. This is a when-to-use plus an alternative, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

assess_riskA
Read-onlyIdempotent

Compute quantitative risk and threat assessment for candidate actions, plans, or 3D spatial rollouts against active utility weights (actions: action, plan, compare, spatial_rollout). Use assess_risk instead of evaluate_situation when estimating failure probability and threat exposure rather than ranking overall utility.

Returns risk score (0.0-1.0), threat breakdown, and comparative risk ratings.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesRisk assessment mode: action, plan, compare, spatial_rollout
projectNoTarget project slug
obstaclesNoObstacles with position/bounding_box and affordance_mask
parametersNoAction parameters
trajectoryNoArray of [x, y, z] waypoints for spatial rollout risk assessment
candidate_actionNoAction name to evaluate
candidate_actionsNoMultiple actions to compare risk scores
situation_contextNoCurrent environment telemetry & vitals
clearance_thresholdNoMinimum clearance threshold in meters (default: 1.0)

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive, so safety is covered. The description adds real behavioral context: the assessment is computed 'against active utility weights' (a state dependency) and it returns a 0.0-1.0 score with threat breakdown and comparative ratings. It does not describe cost, latency, or what happens when no weights are set, so it stops short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly packed sentences with the purpose and routing rule front-loaded. The parenthetical enumeration of modes is mildly redundant with the schema enum, but it aids scanning and costs little.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter analyzer with nested objects and no output schema, the description covers purpose, routing, state dependency, and the shape of the return value (score range, threat breakdown, comparative ratings). Only edge behaviors, such as missing utility weights or invalid trajectories, are unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all nine parameters including the mode enum, trajectory waypoints, and clearance_threshold default. The description's parenthetical mode list and mention of '3D spatial rollouts' restate the schema rather than adding format or constraint detail. Baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Compute quantitative risk and threat assessment') and scopes the accepted modes. It explicitly names the sibling it is not, so an agent can separate assess_risk from evaluate_situation without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit routing rule: 'Use assess_risk instead of evaluate_situation when estimating failure probability and threat exposure rather than ranking overall utility.' This names the alternative and the condition that selects it, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

classifyA
Read-onlyIdempotent

Assign semantic categorical labels to an entity, visual state, task, or state snapshot using deterministic System One calculus. Use classify instead of ask_choice when assigning predefined taxonomy labels rather than selecting among runtime decision alternatives.

Returns top class label, probability distribution, and classification margin.

ParametersJSON Schema
NameRequiredDescriptionDefault
classesYesCandidate classes (capped at 16)
projectYesTarget project slug
target_idNoIdentifier of the target
state_packNoOptional explicit StatePack
target_typeYesType of target to classify

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive, closed-world behavior. The description adds that the operation is deterministic (System One calculus) and specifies return values (top label, probability distribution, margin), which is valuable since no output schema exists. It does not cover error behavior or rate limits, but the annotation set carries the safety profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with purpose, followed by routing guidance and return values. Every sentence contributes new information without repetition or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema, the description usefully states what is returned (top class label, probability distribution, classification margin). Annotations cover safety and idempotency, and the schema covers parameters fully. Minor gaps around state_pack semantics and error handling remain, but the core needs are met.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema itself documents all five parameters. The description mentions the target types broadly but adds no syntax, constraints, or meaning beyond what the schema already provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (assign semantic categorical labels) and resource scope (entity, visual_state, task, snapshot) and names the sibling tool it replaces (ask_choice). An agent can distinguish it from ask_choice and ask_score without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit when-to-use rule: assign predefined taxonomy labels rather than selecting among runtime decision alternatives, and names the alternative tool ask_choice. This is clear routing guidance with an implicit when-not condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evaluate_situationA

Ingest multi-modal situation snapshot, compute expected utilities against active weights, and output prioritized action recommendations (actions: snapshot, quick). Use evaluate_situation instead of assess_risk when ranking candidate actions across multi-attribute utility dimensions rather than calculating isolated threat probabilities.

Returns ranked candidate actions, expected utility scores, and top recommendation.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesSnapshot evaluation mode or quick text context: snapshot, quick
projectNoTarget project slug
snapshotNoNormalized SituationSnapshot with world, vision, state, and vitals
trace_idNoDistributed trace ID
session_idNoLinked state-memory session ID
quick_contextNoText summary of current situation for quick evaluation
lookahead_depthNoBounded heuristic lookahead plies (e.g. 2-3 plies, discount gamma=0.85)
utility_profileNoNamed utility profile to score against (defaults to active)
candidate_actionsNoCandidate actions to score and rank

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, idempotentHint=false, non-destructive and closed-world, so the safety profile is already supplied. The description adds useful return-shape context (ranked candidate actions, expected utility scores, top recommendation), but says nothing about what state is persisted or whether trace/session records are written, which matters for a non-read-only tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is front-loaded with the core verb sequence, then the routing rule, then the return values, with no filler sentences. The utility/decision-theory terminology is dense but each sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool with nested objects, an enum mode switch and no output schema, the description does cover the return values and the key routing decision. It leaves the role of trace_id, session_id and project unaddressed, but those are documented in the schema, so the gap is minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% across 9 parameters, so the schema already explains action, snapshot, quick_context, lookahead_depth, utility_profile and candidate_actions. The description adds only the two enum mode names and the notion of scoring against active weights, which is marginal beyond the structured fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific pipeline of verbs and resources: ingest a multi-modal snapshot, compute expected utilities against active weights, and output prioritized action recommendations, with the modes (snapshot, quick) called out. It also explicitly distinguishes itself from the sibling assess_risk, so an agent can route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states an explicit routing rule: use evaluate_situation rather than assess_risk when ranking candidate actions across multi-attribute utility dimensions instead of computing isolated threat probabilities. That is a genuine when-to-use/when-not rule, though it covers only one sibling out of many and gives no prerequisites or preconditions for the modes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gate_intentionA
Read-onlyIdempotent

Evaluate an intention before dispatching to behavior-mcp and issue a cryptographic HMAC dispatch token if approved. Use gate_intention instead of manage_intentions when verifying precondition safety and issuing execution authorization rather than tracking intention state.

Returns gate verdict (approved/rejected), risk evaluation, and HMAC dispatch token.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectYesTarget project slug
state_packNoOptional explicit StatePack
intention_idNoTarget intention ID
context_goal_idNoActive goal being pursued
proposed_actionYesProposed behavior action to execute

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint=false, openWorldHint=false), and the description adds real context beyond them: this is a precondition gate whose approval produces a cryptographic HMAC token, and it returns a verdict plus risk evaluation. It stops short of saying what a rejection looks like to the caller or whether the token has a lifetime/one-time-use constraint, so it is not fully transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, roughly fifty words, with the core purpose front-loaded in the first clause and no filler. The return-value sentence earns its place because no output schema exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does so adequately (verdict, risk evaluation, HMAC token). Given a five-parameter nested schema and a gating role in a behavior-dispatch pipeline, it could say more about rejection handling and token use, but nothing essential to calling it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters (project, state_pack, intention_id, context_goal_id, proposed_action and its nested fields) are already documented structurally. The description references the intention and proposed action conceptually but adds no syntax, format, or interaction detail beyond the schema, which is the baseline-3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Evaluate an intention before dispatching to behavior-mcp') plus the concrete output ('issue a cryptographic HMAC dispatch token if approved'). It also explicitly names the sibling it is not, manage_intentions, so an agent can separate the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit routing rule: 'Use gate_intention instead of manage_intentions when verifying precondition safety and issuing execution authorization rather than tracking intention state.' This names the alternative and the condition that selects it, which is exactly the when/when-not guidance an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_decision_traceA
Read-onlyIdempotent

Retrieve explainable step-by-step chain-of-thought rationale, candidate utilities, and risk assessment for past decisions (actions: latest, get, list, explain). Use get_decision_trace instead of query_knowledge when performing deep forensic analysis of a specific historical decision.

Returns step-by-step reasoning trace, utility breakdown, candidate rankings, and explanation text.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax traces to list
actionYesTrace retrieval operation: latest, get, list, explain
goal_idNoFilter traces by linked goal ID
projectNoTarget project slug
trace_idNoTrace ID for get/explain action

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive and closed-world, so the safety profile is covered. The description adds value by enumerating the return structure (trace, utility breakdown, candidate rankings, explanation text), which matters since no output schema exists. It does not discuss filtering semantics or limits, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, purpose front-loaded and the sibling call-out second. The parenthetical action list duplicates the enum and the return enumeration slightly overlaps, but nothing is buried or wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description helpfully summarizes return values, and annotations carry the safety profile. What is missing is guidance on how action values interact with the other params (which params apply to get vs list), leaving a small gap for a 5-param tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented, including the action enum. The description only restates the action values parenthetically and adds no syntax or interaction detail (e.g., trace_id required for get/explain), so baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Retrieve) and resource (decision trace) plus the payload it contains (chain-of-thought rationale, candidate utilities, risk assessment). It explicitly contrasts itself with query_knowledge, so an agent can distinguish it from siblings without schema inspection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative (query_knowledge) and the precise condition that selects this tool: deep forensic analysis of a specific historical decision. The action set (latest, get, list, explain) further signals usage modes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_beliefsA
Destructive

Maintain structured belief state with TTL expiration sweeps, exponential confidence decay, and category filtering (actions: update, query, expire, reconcile, reconcile_spatial). Use manage_beliefs instead of query_knowledge when managing dynamic agent epistemic state rather than static heuristic patterns.

Returns belief record, query matches, expired belief count, or reconciliation report.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesBelief operation: update, query, expire, reconcile, reconcile_spatial
objectNoBelief value / state payload
sourceNoBelief provenance source
projectNoTarget project slug
subjectNoBelief subject (e.g. "north_gate", "enemy_patrol")
categoryNoBelief category
belief_idNoBelief ID for specific lookup
predicateNoPredicate relationship (e.g. "is_locked", "status")
confidenceNoConfidence score (0.0 to 1.0)
decay_rateNoExponential decay rate lambda per hour
expires_atNoISO-8601 expiration timestamp
novel_entitiesNoNovel entities for reconcile_spatial action
matched_entitiesNoMatched entities for reconcile_spatial action
missing_entitiesNoMissing entities for reconcile_spatial action
client_request_idNoIdempotency key
unexpected_entitiesNoUnexpected entities for reconcile_spatial action

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, and the description adds real behavioral context beyond them: TTL expiration sweeps, exponential confidence decay, and category filtering, plus the shape of the return value. It still doesn't say what the destructive expire path actually removes or how update overwrites prior state, so it falls short of full disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the capability and the sibling comparison, then closes with the return shapes. Two focused sentences with little waste; the return sentence earns its place because there is no output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 16-parameter, five-action mutation tool with no output schema, the description compensates by listing return values (belief record, query matches, expired count, reconciliation report) and the action set. The remaining gap is that it never indicates which parameters are required by which action, which matters at this parameter count.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 16 parameters are already documented in the schema and the baseline is 3. The description adds the action vocabulary but no per-action parameter mapping or format detail (e.g., which fields reconcile_spatial requires) beyond what the schema states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("Maintain") and resource ("structured belief state") and enumerates the five supported actions, so the scope is unambiguous. It also explicitly distinguishes itself from the sibling query_knowledge, letting an agent route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative tool (query_knowledge) and the condition that selects this one instead (dynamic agent epistemic state vs static heuristic patterns). That is a direct when-to-use/when-not-to-use rule rather than implied guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_intentionsA

Queue, dispatch, track, cancel, or resolve behavior directives for behavior-mcp runtime execution (actions: create, dispatch, get, list, cancel, resolve). Use manage_intentions instead of set_goal when dispatching immediate execution instructions to runtime behaviors rather than managing abstract objectives.

Returns intention record, dispatch status, wire contract payload, or intention list.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesIntention operation: create, dispatch, get, list, cancel, resolve
resultNoOutcome payload for resolve action
statusNoStatus filter or update
goal_idNoLinked goal ID
projectNoTarget project slug
priorityNoExecution priority (0.0 to 1.0)
trace_idNoLinked decision trace ID
parametersNoRuntime behavior parameters
deadline_atNoISO-8601 completion deadline
intention_idNoIntention ID for dispatch/get/cancel/resolve
behavior_nameNoTarget behavior tree name (e.g. "combat_kite", "gather_loop")
abort_conditionsNoAuto-abort trigger conditions
client_request_idNoIdempotency key

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), so the agent knows this is a non-destructive, non-idempotent mutation surface. The description adds the multi-action nature and the return categories, but does not disclose side effects of cancel/resolve, permission needs, or why the operation is non-idempotent despite having a client_request_id idempotency key.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action verbs and resource, followed by the routing rule and the return-value summary. Tight and well-ordered, with only slight redundancy between the action list in prose and the enum in the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter multi-action tool with nested objects and no output schema, the description covers purpose, routing, and return categories, and the fully-described schema fills in the per-action parameter requirements. Minor gap: it does not map which parameters apply to which action (e.g., intention_id for dispatch/get/cancel/resolve).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented in the schema; the description adds no per-parameter meaning beyond what is structured. Baseline 3 is appropriate when the schema carries the semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States specific verbs (queue, dispatch, track, cancel, resolve) applied to a clear resource (behavior directives / intentions) and enumerates the six actions. It explicitly distinguishes itself from the sibling set_goal, so an agent can tell the two apart without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative tool (set_goal) and the exact condition that selects this one instead: dispatching immediate execution instructions to runtime behaviors vs. managing abstract objectives. This is explicit when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_reasoning_dbA
Destructive

Database maintenance, diagnostics, SHA-256 Merkle audit verification, and snapshot management (actions: stats, audit, doctor, snapshot, diff, restore). Use manage_reasoning_db instead of manage_beliefs when performing SQLite storage integrity verification or database snapshot restore.

Returns database diagnostics, Merkle audit tree, snapshot metadata, or diff reports.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoSnapshot name
actionYesDatabase maintenance operation: stats, audit, doctor, snapshot, diff, restore
projectNoTarget project slug
descriptionNoDescription for snapshot

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false and idempotentHint=false, so the safety profile is covered by structured data. The description adds useful return-shape context ('Returns database diagnostics, Merkle audit tree, snapshot metadata, or diff reports'), but never warns that 'restore' overwrites existing database state or that actions carry differing risk β€” information an agent would want beyond the blunt destructive flag.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the capability summary followed by routing and return information; no filler. The parenthetical action list is slightly redundant with the schema enum but keeps the definition self-contained.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by stating what categories of results are returned, and annotations carry the safety profile. It is close to complete for a six-action multiplexed tool, though per-action behavior (especially restore semantics) remains unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the enum is fully documented in the schema, so the baseline is 3. The description only restates the action list and adds no syntax, format, or dependency details (e.g., that 'name'/'description' are only relevant to snapshot) beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb+resource ('Database maintenance, diagnostics, SHA-256 Merkle audit verification, and snapshot management') and enumerates the six concrete actions, so an agent knows exactly what domain this covers. It also explicitly contrasts itself with the sibling manage_beliefs, making it distinguishable without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit routing rule: 'Use manage_reasoning_db instead of manage_beliefs when performing SQLite storage integrity verification or database snapshot restore.' That is a clear when-to-use-this-vs-alternative condition. It stops short of covering the other actions (stats, diff, doctor) or any when-not/exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_knowledgeA
Read-onlyIdempotent

Search learned heuristic patterns, tactics, and past decision traces by context similarity (actions: search, patterns, similar_situations). Use query_knowledge instead of get_decision_trace when retrieving generalized patterns across sessions rather than inspecting a single execution trace.

Returns matching heuristics, anti-patterns, tactics, and similarity scores.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax patterns to return
queryNoSemantic search query string
actionYesKnowledge query mode: search, patterns, similar_situations
projectNoTarget project slug
context_tagsNoFilter by context tags
pattern_typeNoFilter by pattern category

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, non-destructive, and closed-world behavior. The description adds useful return context by naming the kinds of matches returned (heuristics, anti-patterns, tactics, similarity scores) and the cross-session scope, though it does not discuss pagination or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose and routing guidance, then states the return contents. Every sentence adds useful information without repetition or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter read-only search tool with rich annotations, the description supplies purpose, usage routing, supported action modes, and the shape of returned results. With no output schema, the return summary is especially valuable and sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters are already documented in the schema, including the action and pattern_type enums. The description repeats the action modes but adds no syntax, format, or interaction details beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Search) and resources (learned heuristic patterns, tactics, past decision traces) and names the supported actions. It also explicitly distinguishes the tool from get_decision_trace by contrasting generalized cross-session patterns with a single execution trace.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit routing guidance: use query_knowledge instead of get_decision_trace when retrieving generalized patterns across sessions, rather than inspecting one execution trace. This covers when to use it, when to prefer an alternative, and the alternative itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

replanA

Regenerate sub-task DAG and adjust intentions upon unexpected obstacles or state changes (actions: blocker, event, full). Use replan instead of set_goal when recovering from execution blockers or environment shifts rather than creating new goals.

Returns replanned goal DAG, invalidated intentions, and newly synthesized subgoals.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesReplanning trigger type: blocker, event, full
goal_idYesID of goal to replan
projectNoTarget project slug
trigger_eventNoEvent description triggering replanning
preserve_completedNoWhether to preserve already completed subgoals
blocker_descriptionNoDescription of the obstacle or blocker encountered

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, and idempotentHint=false, so the mutation profile is partly covered. The description adds real context beyond them by enumerating the trigger actions and describing what is produced/side-effected (replanned DAG, invalidated intentions, synthesized subgoals). It stops short of explaining permissions or whether invalidation is reversible, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three front-loaded sentences with clear routing guidance and no filler. Repeating the action list is mildly redundant with the schema enum, which keeps it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter mutation tool with no output schema, the description usefully names the return payload and the alternative tool, and annotations cover safety. It is nearly complete, lacking only error/precondition behavior for when replanning fails.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters including the action enum are already documented in the schema. The description only restates the action values ('blocker, event, full') without adding semantics such as how 'full' differs from 'event'. Baseline 3 applies when the schema carries the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource β€” 'Regenerate sub-task DAG and adjust intentions' β€” and explicitly scopes the trigger to 'unexpected obstacles or state changes'. It also names and contrasts against the sibling set_goal, so an agent can distinguish it without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit when-to-use rule plus the alternative: 'Use replan instead of set_goal when recovering from execution blockers or environment shifts rather than creating new goals.' The condition that selects this tool over the sibling is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_goalA

Register, update, decompose, inspect, or abandon hierarchical goals and task DAGs (actions: create, update, decompose, get, list, abandon). Use set_goal instead of manage_intentions when formulating multi-step objectives and sub-goal DAGs rather than queueing concrete execution directives.

Returns goal record, sub-goal hierarchy, completion progress, or goal lists.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoGoal ID (required for update, get, abandon)
limitNoMax items to return for list action
titleNoGoal title or objective summary
actionYesThe goal management operation to perform: create, update, decompose, get, list, abandon
statusNoGoal status
projectNoTarget project slug
priorityNoGoal priority (0.0 to 1.0)
progressNoCompletion progress (0.0 to 1.0)
subgoalsNoArray of sub-goals for decompose action
parent_idNoParent goal ID for hierarchical sub-goals
deadline_atNoISO-8601 deadline timestamp
descriptionNoDetailed goal description
utility_weightsNoGoal-specific utility weight overrides
success_criteriaNoList of verifiable conditions
client_request_idNoIdempotency key to prevent duplicate creation

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, and idempotentHint=false, so the mutation profile is covered. The description adds an enumeration of return shapes (goal record, sub-goal hierarchy, progress, lists), which is useful since there is no output schema, but it says nothing about permissions, whether 'abandon' is reversible, or side effects of decompose. It adds some context but not rich behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the action set and the sibling routing rule, then closes with a brief returns sentence. It is appropriately sized for a multi-action tool, with no filler, though the opening verb list is slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 15-parameter, nested-object, no-output-schema tool, the description covers the action space, sibling routing, and return shapes, which is most of what an agent needs. It stops short of per-action parameter guidance (e.g. which fields apply to decompose vs update), leaving a modest gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every one of the 15 parameters is already documented, including which actions require `id` and the action enum. The description only restates the action names and does not enrich parameter meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (hierarchical goals and task DAGs) and enumerates the exact verb set (create, update, decompose, get, list, abandon), so an agent knows precisely what the tool does. It also distinguishes itself from the sibling manage_intentions, making it identifiable without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit routing rule: use set_goal rather than manage_intentions 'when formulating multi-step objectives and sub-goal DAGs rather than queueing concrete execution directives.' The alternative and the condition that selects it are both stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_utility_weightsA

Configure, inspect, activate, or incrementally nudge multi-attribute utility weight profiles (actions: configure, get, list, activate, nudge). Use set_utility_weights instead of evaluate_situation when defining decision preferences (aggression, caution, greed, exploration) rather than evaluating actions.

Returns configured utility profile, active weight map, or profile directory.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoProfile name (e.g. "aggressive", "cautious", "explorer")
deltaNoKey-value map of weight deltas for nudge action
actionYesProfile operation: configure, get, list, activate, nudge
projectNoTarget project slug
weightsNoKey-value map of weight values (0.0 to 1.0)
is_activeNoWhether to set as currently active profile
descriptionNoProfile description

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a non-read-only, non-destructive, non-idempotent mutation, so the safety profile is covered. The description adds the return surface (configured profile, active weight map, or profile directory), which is genuinely useful given no output schema exists, but says nothing about side effects of activate/nudge or auth requirements. Adequate, not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and action list, then the routing rule, then returns β€” a sensible ordering with no filler sentences. The action list is restated from the enum, a minor redundancy that keeps it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter multi-action tool with nested objects and no output schema, the definition covers purpose, action routing, alternative selection, and the shape of returns. It is close to complete; the only gap is per-action behavioral detail for the mutating actions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter including the action enum is already documented in the schema; baseline 3 applies. The description adds no syntax, range, or interaction detail (e.g. how delta composes with weights) beyond what the schema states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (multi-attribute utility weight profiles) with an explicit verb set and enumerates the five supported actions, so the agent knows exactly what operations this tool performs. It also names the sibling it must not be confused with (evaluate_situation) and the distinction between them, which is more than most definitions provide.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly gives the selection rule: use this instead of evaluate_situation when defining decision preferences rather than evaluating actions. That is a clear when-to-use-vs-alternative statement. It does not, however, give per-action guidance (e.g. when to nudge vs configure), which keeps it short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.4.0
    • Changedassess_risk5 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Risk assessment mode: action, plan, compare"New value: +"Risk assessment mode: action, plan, compare, spatial_rollout"
      • changedInput schema / properties / action / enum
        Previous value: -[
        -  "action",
        -  "plan",
        -  "compare"
        -]New value: +[
        +  "action",
        +  "plan",
        +  "compare",
        +  "spatial_rollout"
        +]
      • addedInput schema / properties / clearance_threshold
        Added value: +{
        +  "description": "Minimum clearance threshold in meters (default: 1.0)",
        +  "type": "number"
        +}
      • addedInput schema / properties / obstacles
        Added value: +{
        +  "description": "Obstacles with position/bounding_box and affordance_mask",
        +  "items": {},
        +  "type": "array"
        +}
      • addedInput schema / properties / trajectory
        Added value: +{
        +  "description": "Array of [x, y, z] waypoints for spatial rollout risk assessment",
        +  "items": {
        +    "items": {
        +      "type": "number"
        +    },
        +    "type": "array"
        +  },
        +  "type": "array"
        +}
    • Changedmanage_beliefs6 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Belief operation: update, query, expire, reconcile"New value: +"Belief operation: update, query, expire, reconcile, reconcile_spatial"
      • changedInput schema / properties / action / enum
        Previous value: -[
        -  "update",
        -  "query",
        -  "expire",
        -  "reconcile"
        -]New value: +[
        +  "update",
        +  "query",
        +  "expire",
        +  "reconcile",
        +  "reconcile_spatial"
        +]
      • addedInput schema / properties / matched_entities
        Added value: +{
        +  "description": "Matched entities for reconcile_spatial action",
        +  "items": {},
        +  "type": "array"
        +}
      • addedInput schema / properties / missing_entities
        Added value: +{
        +  "description": "Missing entities for reconcile_spatial action",
        +  "items": {},
        +  "type": "array"
        +}
      • addedInput schema / properties / novel_entities
        Added value: +{
        +  "description": "Novel entities for reconcile_spatial action",
        +  "items": {},
        +  "type": "array"
        +}
      • addedInput schema / properties / unexpected_entities
        Added value: +{
        +  "description": "Unexpected entities for reconcile_spatial action",
        +  "items": {},
        +  "type": "array"
        +}
    • Changedset_utility_weights3 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Profile operation: configure, get, list, activate"New value: +"Profile operation: configure, get, list, activate, nudge"
      • changedInput schema / properties / action / enum
        Previous value: -[
        -  "configure",
        -  "get",
        -  "list",
        -  "activate"
        -]New value: +[
        +  "configure",
        +  "get",
        +  "list",
        +  "activate",
        +  "nudge"
        +]
      • addedInput schema / properties / delta
        Added value: +{
        +  "description": "Key-value map of weight deltas for nudge action",
        +  "type": "object"
        +}
  2. 15 tool updatesv0.3.1
    • Addedask_choice
    • Addedask_noul
    • Addedask_score
    • Changedassess_risk2 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Risk assessment mode"New value: +"Risk assessment mode: action, plan, compare"
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "action",
        +  "plan",
        +  "compare"
        +]
    • Addedclassify
    • Changedevaluate_situation2 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Snapshot evaluation mode or quick text context"New value: +"Snapshot evaluation mode or quick text context: snapshot, quick"
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "snapshot",
        +  "quick"
        +]
    • Addedgate_intention
    • Changedget_decision_trace2 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Trace retrieval operation"New value: +"Trace retrieval operation: latest, get, list, explain"
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "latest",
        +  "get",
        +  "list",
        +  "explain"
        +]
    • Changedmanage_beliefs2 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Belief operation"New value: +"Belief operation: update, query, expire, reconcile"
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "update",
        +  "query",
        +  "expire",
        +  "reconcile"
        +]
    • Changedmanage_intentions2 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Intention operation"New value: +"Intention operation: create, dispatch, get, list, cancel, resolve"
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "create",
        +  "dispatch",
        +  "get",
        +  "list",
        +  "cancel",
        +  "resolve"
        +]
    • Changedmanage_reasoning_db2 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Database maintenance operation: stats, audit, doctor (health diagnostics), snapshot, diff, restore"New value: +"Database maintenance operation: stats, audit, doctor, snapshot, diff, restore"
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "stats",
        +  "audit",
        +  "doctor",
        +  "snapshot",
        +  "diff",
        +  "restore"
        +]
    • Changedquery_knowledge2 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Knowledge query mode"New value: +"Knowledge query mode: search, patterns, similar_situations"
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "search",
        +  "patterns",
        +  "similar_situations"
        +]
    • Changedreplan2 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Replanning trigger type"New value: +"Replanning trigger type: blocker, event, full"
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "blocker",
        +  "event",
        +  "full"
        +]
    • Changedset_goal2 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"The goal management operation to perform"New value: +"The goal management operation to perform: create, update, decompose, get, list, abandon"
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "create",
        +  "update",
        +  "decompose",
        +  "get",
        +  "list",
        +  "abandon"
        +]
    • Changedset_utility_weights2 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Profile operation"New value: +"Profile operation: configure, get, list, activate"
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "configure",
        +  "get",
        +  "list",
        +  "activate"
        +]
  3. 10 tool updatesv0.2.1
    • Changedassess_risk16 fields changed
      • removedInput schema / $schema
        Removed value: -"http://json-schema.org/draft-07/schema#"
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / action / description
        Added value: +"Risk assessment mode"
      • removedInput schema / properties / action / enum
        Removed value: -[
        -  "action",
        -  "plan",
        -  "compare"
        -]
      • addedInput schema / properties / candidate_action / description
        Added value: +"Action name to evaluate"
      • addedInput schema / properties / candidate_actions / description
        Added value: +"Multiple actions to compare risk scores"
      • removedInput schema / properties / candidate_actions / items / additionalProperties
        Removed value: -true
      • removedInput schema / properties / candidate_actions / items / properties / parameters / additionalProperties
        Removed value: -true
      • removedInput schema / properties / candidate_actions / items / properties / parameters / properties
        Removed value: -{}
      • removedInput schema / properties / parameters / additionalProperties
        Removed value: -true
      • addedInput schema / properties / parameters / description
        Added value: +"Action parameters"
      • removedInput schema / properties / parameters / properties
        Removed value: -{}
      • addedInput schema / properties / project / description
        Added value: +"Target project slug"
      • removedInput schema / properties / situation_context / additionalProperties
        Removed value: -true
      • addedInput schema / properties / situation_context / description
        Added value: +"Current environment telemetry & vitals"
      • removedInput schema / properties / situation_context / properties
        Removed value: -{}
    • Changedevaluate_situation17 fields changed
      • removedInput schema / $schema
        Removed value: -"http://json-schema.org/draft-07/schema#"
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / action / description
        Added value: +"Snapshot evaluation mode or quick text context"
      • removedInput schema / properties / action / enum
        Removed value: -[
        -  "snapshot",
        -  "quick"
        -]
      • addedInput schema / properties / candidate_actions / description
        Added value: +"Candidate actions to score and rank"
      • removedInput schema / properties / candidate_actions / items / additionalProperties
        Removed value: -true
      • removedInput schema / properties / candidate_actions / items / properties / parameters / additionalProperties
        Removed value: -true
      • removedInput schema / properties / candidate_actions / items / properties / parameters / properties
        Removed value: -{}
      • addedInput schema / properties / lookahead_depth
        Added value: +{
        +  "description": "Bounded heuristic lookahead plies (e.g. 2-3 plies, discount gamma=0.85)",
        +  "type": "number"
        +}
      • addedInput schema / properties / project / description
        Added value: +"Target project slug"
      • addedInput schema / properties / quick_context / description
        Added value: +"Text summary of current situation for quick evaluation"
      • addedInput schema / properties / session_id / description
        Added value: +"Linked state-memory session ID"
      • removedInput schema / properties / snapshot / additionalProperties
        Removed value: -true
      • addedInput schema / properties / snapshot / description
        Added value: +"Normalized SituationSnapshot with world, vision, state, and vitals"
      • removedInput schema / properties / snapshot / properties
        Removed value: -{}
      • addedInput schema / properties / trace_id / description
        Added value: +"Distributed trace ID"
      • addedInput schema / properties / utility_profile / description
        Added value: +"Named utility profile to score against (defaults to active)"
    • Changedget_decision_trace8 fields changed
      • removedInput schema / $schema
        Removed value: -"http://json-schema.org/draft-07/schema#"
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / action / description
        Added value: +"Trace retrieval operation"
      • removedInput schema / properties / action / enum
        Removed value: -[
        -  "latest",
        -  "get",
        -  "list",
        -  "explain"
        -]
      • addedInput schema / properties / goal_id / description
        Added value: +"Filter traces by linked goal ID"
      • addedInput schema / properties / limit / description
        Added value: +"Max traces to list"
      • addedInput schema / properties / project / description
        Added value: +"Target project slug"
      • addedInput schema / properties / trace_id / description
        Added value: +"Trace ID for get/explain action"
    • Changedmanage_beliefs15 fields changed
      • removedInput schema / $schema
        Removed value: -"http://json-schema.org/draft-07/schema#"
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / action / description
        Added value: +"Belief operation"
      • removedInput schema / properties / action / enum
        Removed value: -[
        -  "update",
        -  "query",
        -  "expire",
        -  "reconcile"
        -]
      • addedInput schema / properties / belief_id / description
        Added value: +"Belief ID for specific lookup"
      • addedInput schema / properties / category / description
        Added value: +"Belief category"
      • addedInput schema / properties / client_request_id / description
        Added value: +"Idempotency key"
      • addedInput schema / properties / confidence / description
        Added value: +"Confidence score (0.0 to 1.0)"
      • addedInput schema / properties / decay_rate / description
        Added value: +"Exponential decay rate lambda per hour"
      • addedInput schema / properties / expires_at / description
        Added value: +"ISO-8601 expiration timestamp"
      • addedInput schema / properties / object / description
        Added value: +"Belief value / state payload"
      • addedInput schema / properties / predicate / description
        Added value: +"Predicate relationship (e.g. \"is_locked\", \"status\")"
      • addedInput schema / properties / project / description
        Added value: +"Target project slug"
      • addedInput schema / properties / source / description
        Added value: +"Belief provenance source"
      • addedInput schema / properties / subject / description
        Added value: +"Belief subject (e.g. \"north_gate\", \"enemy_patrol\")"
    • Changedmanage_intentions22 fields changed
      • removedInput schema / $schema
        Removed value: -"http://json-schema.org/draft-07/schema#"
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / abort_conditions / description
        Added value: +"Auto-abort trigger conditions"
      • removedInput schema / properties / abort_conditions / items / additionalProperties
        Removed value: -true
      • removedInput schema / properties / abort_conditions / items / properties
        Removed value: -{}
      • addedInput schema / properties / action / description
        Added value: +"Intention operation"
      • removedInput schema / properties / action / enum
        Removed value: -[
        -  "create",
        -  "dispatch",
        -  "get",
        -  "list",
        -  "cancel",
        -  "resolve"
        -]
      • addedInput schema / properties / behavior_name / description
        Added value: +"Target behavior tree name (e.g. \"combat_kite\", \"gather_loop\")"
      • addedInput schema / properties / client_request_id / description
        Added value: +"Idempotency key"
      • addedInput schema / properties / deadline_at / description
        Added value: +"ISO-8601 completion deadline"
      • addedInput schema / properties / goal_id / description
        Added value: +"Linked goal ID"
      • addedInput schema / properties / intention_id / description
        Added value: +"Intention ID for dispatch/get/cancel/resolve"
      • removedInput schema / properties / parameters / additionalProperties
        Removed value: -true
      • addedInput schema / properties / parameters / description
        Added value: +"Runtime behavior parameters"
      • removedInput schema / properties / parameters / properties
        Removed value: -{}
      • addedInput schema / properties / priority / description
        Added value: +"Execution priority (0.0 to 1.0)"
      • addedInput schema / properties / project / description
        Added value: +"Target project slug"
      • removedInput schema / properties / result / additionalProperties
        Removed value: -true
      • addedInput schema / properties / result / description
        Added value: +"Outcome payload for resolve action"
      • removedInput schema / properties / result / properties
        Removed value: -{}
      • addedInput schema / properties / status / description
        Added value: +"Status filter or update"
      • addedInput schema / properties / trace_id / description
        Added value: +"Linked decision trace ID"
    • Changedmanage_reasoning_db7 fields changed
      • removedInput schema / $schema
        Removed value: -"http://json-schema.org/draft-07/schema#"
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / action / description
        Added value: +"Database maintenance operation: stats, audit, doctor (health diagnostics), snapshot, diff, restore"
      • removedInput schema / properties / action / enum
        Removed value: -[
        -  "stats",
        -  "audit",
        -  "snapshot",
        -  "diff",
        -  "restore"
        -]
      • addedInput schema / properties / description / description
        Added value: +"Description for snapshot"
      • addedInput schema / properties / name / description
        Added value: +"Snapshot name"
      • addedInput schema / properties / project / description
        Added value: +"Target project slug"
    • Changedquery_knowledge9 fields changed
      • removedInput schema / $schema
        Removed value: -"http://json-schema.org/draft-07/schema#"
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / action / description
        Added value: +"Knowledge query mode"
      • removedInput schema / properties / action / enum
        Removed value: -[
        -  "search",
        -  "patterns",
        -  "similar_situations"
        -]
      • addedInput schema / properties / context_tags / description
        Added value: +"Filter by context tags"
      • addedInput schema / properties / limit / description
        Added value: +"Max patterns to return"
      • addedInput schema / properties / pattern_type / description
        Added value: +"Filter by pattern category"
      • addedInput schema / properties / project / description
        Added value: +"Target project slug"
      • addedInput schema / properties / query / description
        Added value: +"Semantic search query string"
    • Changedreplan9 fields changed
      • removedInput schema / $schema
        Removed value: -"http://json-schema.org/draft-07/schema#"
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / action / description
        Added value: +"Replanning trigger type"
      • removedInput schema / properties / action / enum
        Removed value: -[
        -  "blocker",
        -  "event",
        -  "full"
        -]
      • addedInput schema / properties / blocker_description / description
        Added value: +"Description of the obstacle or blocker encountered"
      • addedInput schema / properties / goal_id / description
        Added value: +"ID of goal to replan"
      • addedInput schema / properties / preserve_completed / description
        Added value: +"Whether to preserve already completed subgoals"
      • addedInput schema / properties / project / description
        Added value: +"Target project slug"
      • addedInput schema / properties / trigger_event / description
        Added value: +"Event description triggering replanning"
    • Changedset_goal21 fields changed
      • removedInput schema / $schema
        Removed value: -"http://json-schema.org/draft-07/schema#"
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / action / description
        Added value: +"The goal management operation to perform"
      • removedInput schema / properties / action / enum
        Removed value: -[
        -  "create",
        -  "update",
        -  "decompose",
        -  "get",
        -  "list",
        -  "abandon"
        -]
      • addedInput schema / properties / client_request_id / description
        Added value: +"Idempotency key to prevent duplicate creation"
      • addedInput schema / properties / deadline_at / description
        Added value: +"ISO-8601 deadline timestamp"
      • addedInput schema / properties / description / description
        Added value: +"Detailed goal description"
      • addedInput schema / properties / id / description
        Added value: +"Goal ID (required for update, get, abandon)"
      • addedInput schema / properties / limit / description
        Added value: +"Max items to return for list action"
      • addedInput schema / properties / parent_id / description
        Added value: +"Parent goal ID for hierarchical sub-goals"
      • addedInput schema / properties / priority / description
        Added value: +"Goal priority (0.0 to 1.0)"
      • addedInput schema / properties / progress / description
        Added value: +"Completion progress (0.0 to 1.0)"
      • addedInput schema / properties / project / description
        Added value: +"Target project slug"
      • addedInput schema / properties / status / description
        Added value: +"Goal status"
      • addedInput schema / properties / subgoals / description
        Added value: +"Array of sub-goals for decompose action"
      • removedInput schema / properties / subgoals / items / additionalProperties
        Removed value: -true
      • addedInput schema / properties / success_criteria / description
        Added value: +"List of verifiable conditions"
      • addedInput schema / properties / title / description
        Added value: +"Goal title or objective summary"
      • removedInput schema / properties / utility_weights / additionalProperties
        Removed value: -true
      • addedInput schema / properties / utility_weights / description
        Added value: +"Goal-specific utility weight overrides"
      • removedInput schema / properties / utility_weights / properties
        Removed value: -{}
    • Changedset_utility_weights11 fields changed
      • removedInput schema / $schema
        Removed value: -"http://json-schema.org/draft-07/schema#"
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / action / description
        Added value: +"Profile operation"
      • removedInput schema / properties / action / enum
        Removed value: -[
        -  "configure",
        -  "get",
        -  "list",
        -  "activate"
        -]
      • addedInput schema / properties / description / description
        Added value: +"Profile description"
      • addedInput schema / properties / is_active / description
        Added value: +"Whether to set as currently active profile"
      • addedInput schema / properties / name / description
        Added value: +"Profile name (e.g. \"aggressive\", \"cautious\", \"explorer\")"
      • addedInput schema / properties / project / description
        Added value: +"Target project slug"
      • removedInput schema / properties / weights / additionalProperties
        Removed value: -true
      • addedInput schema / properties / weights / description
        Added value: +"Key-value map of weight values (0.0 to 1.0)"
      • removedInput schema / properties / weights / properties
        Removed value: -{}
  4. 10 tool updatesv0.1.2
    • First observedassess_risk
    • First observedevaluate_situation
    • First observedget_decision_trace
    • First observedmanage_beliefs
    • First observedmanage_intentions
    • First observedmanage_reasoning_db
    • First observedquery_knowledge
    • First observedreplan
    • First observedset_goal
    • First observedset_utility_weights

TDQS

A4.1/5.0

Scored across 15 tools

Disambiguation4/5

Most tools target distinct concepts, and the descriptions actively disambiguate with explicit 'use X instead of Y' guidance (e.g. assess_risk vs evaluate_situation, set_goal vs manage_intentions). However there is a dense cluster of risk/utility evaluation tools (assess_risk, evaluate_situation, ask_score, gate_intention) that an agent could still misselect among.

Naming Consistency4/5

The set consistently uses snake_case verb_noun patterns (manage_beliefs, query_knowledge, get_decision_trace, set_utility_weights), with coherent families like ask_* and manage_*. Minor deviations are bare verbs (replan, classify) and an unusual token (ask_noul), but overall the convention is predictable.

Tool Count4/5

15 tools sit at the upper edge of the comfortable 3-15 range but are justified by a genuinely broad cognitive-architecture domain (risk, beliefs, goals, intentions, knowledge, DB ops). It is slightly heavy but each tool maps to a real capability rather than being filler.

Completeness4/5

Coverage is strong: beliefs (update/query/expire/reconcile), goals (create/update/decompose/get/list/abandon), intentions (create/dispatch/get/list/cancel/resolve) and traces/knowledge all form coherent lifecycles. Gaps are minorβ€”no obvious delete or bulk-export for some resourcesβ€”but core reasoning workflows are well covered.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Objective-driven cognitive architecture engine that builds single experts, councils, or full autonomous organizations from user goals, generating deployment-ready superprompts and configurations.
    8 npm
    MIT
  • -
    license
    Not graded
    quality
    Not graded
    maintenance
    Goal-oriented narrative state machine for AI agents, exposing live world/scene context as MCP tools and resources.
    -