behavior-mcp
This MCP server runs and manages high-frequency (~60Hz) in-browser behavior tree execution with safety guardrails, auditing, and state coordination.
Load, unload, or hot-swap behavior tree instances in the browser runtime.
Set, get, or reset runtime execution parameters.
Query current execution status, history, and full tree node traversal state.
Abort, pause, resume, or unstick behavior execution and disengage inputs.
Register or list reactive interrupt triggers with priority preemption and cooldown guards.
Capture or list deterministic browser action sequence recordings.
Retrieve runtime telemetry, tick durations, stuck events, category stats, and action outcome spool data.
Register, list, get, or synthesize immutable behavior tree JSON definitions with SHA-256 hash verification.
Read, write, delete, lease, list, or ingest shared blackboard state slices across behavior nodes.
Manage the SQLite runtime database: stats, SHA-256 Merkle audit, doctor diagnostics, snapshots, diff, and restore.
Provides tools for managing the SQLite runtime database, including diagnostics, SHA-256 Merkle audit, snapshots, and checkpoint rollback.
@putervision/behavior-mcp
High-Frequency (~60Hz) In-Browser Behavior Tree Tactical Runtime Engine for AI Agents
@putervision/behavior-mcp is a Model Context Protocol (MCP) server that executes deterministic behavior trees directly in browser runtimes at ~60Hz with reactive trigger preemption, frame recordings, SHA-256 Merkle audit chains, and a 5-layer safety guardrail stack.
๐ Official Documentation: putervision.com โข Interactive Web Docs
โก 15-Second Quick Start
# 1. Initialize runtime & seed default behavior trees in your workspace
npx @putervision/behavior-mcp init
# 2. Run system diagnostic and health checks
npx @putervision/behavior-mcp doctor
# 3. Inspect active executions, blackboard state, and reactive triggers
npx @putervision/behavior-mcp inspectRelated MCP server: agent-reasoning-mcp
๐ ๏ธ 10 Core MCP Tools
Tool | Actions | Purpose |
|
| Inject, initialize, or hot-swap a behavior tree in browser runtime |
|
| Dynamically update or query runtime execution parameters |
|
| Query active status, node traversal path, tick count, duration |
|
| Halt, pause, or resume behavior execution safely |
|
| Configure reactive interrupt triggers with priority preemption |
|
| Capture or inspect deterministic browser action sequences |
|
| Retrieve runtime telemetry, tick durations, and stuck events |
|
| CRUD for immutable behavior tree JSON with SHA-256 tree hashes |
|
| Read, write, or query shared behavior tree blackboard variables |
|
| SQLite diagnostics, SHA-256 Merkle audit, and checkpoint rollback |
๐ก๏ธ 5-Layer Safety Guardrail Stack & System 1 Invariants
Fail-Closed Evaluator: Unrecognized node definitions throw fatal exceptions immediately.
60Hz Tick Invariant & Synchronous
semantic_check: The loop never blocks on external network calls; semantic condition checks resolve synchronously against blackboard caches.HMAC Intention Gate Verification: Intentions dispatched to execution nodes require unexpired, cryptographically signed dispatch tokens (
PENTAD_HMAC_SECRET).Action Rate Limiter & Leaf Policy Gate: Strict 60 actions/sec maximum throughput ceiling and policy guardrails blocking irreversible mutations.
Emergency Kill Switch & Watchdog Heartbeat: Atomic safety latch and watchdog timers halting execution if tick stalls beyond 5,000ms.
๐ Deep Documentation Guides
๐ Formal API Reference: Full parameter tables, type definitions, and tool schemas.
๐ก Core Architecture & Concepts: Behavior tree execution model, ~60Hz loop, and trigger lifecycle.
๐ฅ๏ธ CLI Usage Guide: Complete CLI command reference (
init,doctor,inspect,run).๐พ Database Schema: SQLite table structures, indexes, and Merkle audit ledger.
โ๏ธ Configuration Reference:
.behavior-mcp.jsonparameters and environment variables.
๐ Client Configuration & Environment
Add to .cursor/mcp.json or .vscode/mcp.json:
{
"mcpServers": {
"behavior-mcp": {
"command": "behavior-mcp",
"args": ["run"],
"env": {
"PENTAD_HMAC_SECRET": "your-secure-shared-secret-here"
}
}
}
}๐งช Testing
# Run full unit and integration test suite across 38 test files (266 tests)
npm test๐ License
MIT ยฉ PuterVision
Available Tools
10 toolsabort_behaviorA
Immediately halt, pause, resume, or unstick behavior execution and disengage active inputs (actions: abort, pause, resume, unstick). Use abort_behavior instead of register_trigger when manually halting or recovering execution rather than configuring automatic condition interrupts.
Returns transition status, disengaged inputs, and unstick diagnostics.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Execution control operation: abort, pause, resume, unstick (resets stuck score and disengages inputs) | |
| reason | No | Reason for abort or pause | |
| project | No | Target project slug | |
| trace_id | No | Distributed trace ID | |
| execution_id | No | Target execution ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare non-readOnly, non-idempotent, non-destructive, so safety profile is covered. The description adds real behavioral value beyond them: it discloses the side effect of disengaging active inputs and enumerates what comes back (transition status, disengaged inputs, unstick diagnostics). It does not mention authorization requirements or whether abort is recoverable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the operations and kept to three tight sentences. The action set is repeated in the parenthetical after already being named, which is mild redundancy, but nothing else is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter mutation tool with no output schema, the description covers purpose, the sibling alternative, and the return payload, which is enough to invoke it correctly. It omits prerequisites (need for an existing execution/trace context) and the relationship between the four actions, leaving minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the action enum is already documented with per-value semantics, so the schema carries the load. The description only restates the action list and adds no format or interaction detail for reason/project/trace_id/execution_id, making the baseline 3 correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names the concrete operations (abort, pause, resume, unstick) against a specific resource (behavior execution) and states the side effect (disengaging active inputs). It also names the sibling it is not (register_trigger), so an agent can distinguish it without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: 'Use abort_behavior instead of register_trigger when manually halting or recovering execution rather than configuring automatic condition interrupts.' That gives a clear alternative plus the selecting condition. It stops short of 5 because it offers no guidance on choosing among the four actions (abort vs pause vs unstick vs resume) at this level.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_metricsARead-onlyIdempotent
Retrieve runtime execution telemetry, tick durations, stuck events, category statistics, and 60Hz action outcome spool entries (actions: current, history, aggregate, compare, spool, drain_spool). Use get_metrics instead of get_status when evaluating tick performance, reading spooled action outcomes, or inspecting aggregated statistics rather than active node traversal.
Returns telemetry metrics, duration percentiles, stuck event counts, spooled outcome records, or drained entry counts.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max records | |
| action | Yes | Metrics query mode: current, history, aggregate, compare, spool, drain_spool | |
| project | No | Target project slug | |
| trace_id | No | Distributed trace ID | |
| execution_id | No | Filter metrics by execution ID | |
| behavior_name | No | Filter metrics by behavior tree name | |
| unsynced_only | No | Filter only unsynced spool entries (action: spool) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and idempotentHint=true, but the tool exposes a drain_spool action returning 'drained entry counts' โ draining consumes/removes spool entries, which is a state mutation and is not idempotent. The description thus contradicts the read-only/idempotent safety profile for at least one action, and it never warns that one action differs in side effects from the rest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the resource list, then routes to the alternative, then summarizes returns โ a sensible order. The 'Returns telemetry metrics...' sentence largely re-states the opening list, and the action enum is repeated from the schema, adding minor redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-param, 6-mode read/mutate hybrid with no output schema, the description does cover purpose, alternatives, and return categories. However it fails to disambiguate the side-effect profile across actions (drain_spool vs read-only modes), which is exactly the information an agent needs before invoking, so it is adequate but with a material gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented, and the description merely repeats the action enum verbatim. It adds no syntax, format, or meaning (e.g., how trace_id interacts with execution_id filter, or what behavior_name expects) beyond the schema, matching the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Retrieve) plus the concrete resource set: runtime telemetry, tick durations, stuck events, category statistics, and 60Hz action outcome spool entries. It also names the sibling it differs from (get_status) and the axis of difference (performance/spooled outcomes vs active node traversal), so an agent can separate the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Contains an explicit when-to-use/when-not clause: 'Use get_metrics instead of get_status when evaluating tick performance, reading spooled action outcomes, or inspecting aggregated statistics rather than active node traversal.' The alternative tool and the selecting condition are both named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_statusARead-onlyIdempotent
Query active behavior execution status, history, or full tree node traversal state (actions: current, history, tree_state). Use get_status instead of get_metrics when inspecting active execution state and node traversal paths rather than aggregated runtime performance telemetry.
Returns execution status, current node path, tick counters, duration, and error states.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max history entries | |
| action | Yes | Status query mode: current, history, tree_state | |
| project | No | Target project slug | |
| execution_id | No | Target execution ID (or latest if omitted) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive and closed-world, so the safety profile is covered. The description adds value beyond that by enumerating what is returned (execution status, node path, tick counters, duration, error states), which matters since there is no output schema. It stops short of noting auth requirements, rate limits, or pagination behavior for history.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the action modes and the routing rule front-loaded, followed by the return summary. Very little waste, though the parenthetical enum listing is mildly redundant with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only query tool with full schema coverage and safety annotations, the definition covers purpose, alternatives, and return fields, compensating for the absent output schema. Minor gaps remain around history pagination and default-to-latest execution behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents action, project, execution_id, and limit. The description's parenthetical restates the action enum but adds no new syntax or format detail. Baseline 3 is appropriate when the schema carries the parameter burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Query) and resource (behavior execution status/history/tree traversal), and enumerates the three action modes. It also names the sibling get_metrics and explains the distinction, so an agent can differentiate without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use get_status instead of get_metrics and gives the selecting condition: inspecting active execution state and node traversal paths vs. aggregated runtime performance telemetry. This is a full when/when-not/alternative statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
load_behaviorADestructive
Load, unload, or hot-swap a behavior tree instance in the browser runtime (actions: load, unload, swap). Use load_behavior instead of manage_behaviors when executing an active behavior tree instance at runtime rather than registering or inspecting definitions.
Returns execution handle, runtime state, active node, and session/intention bindings.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Behavior loading operation: load, unload, swap | |
| project | No | Target project slug | |
| trace_id | No | Distributed trace ID for cross-server correlation | |
| parameters | No | Initial execution parameters | |
| session_id | No | Linked state-memory session ID | |
| intention_id | No | Linked agent-reasoning intention ID | |
| behavior_name | Yes | Name of the behavior tree to load | |
| behavior_version | No | Version of behavior tree (defaults to latest) | |
| client_request_id | No | Idempotency key |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the safety profile is covered structurally. The description usefully adds what the call returns (handle, runtime state, active node, session/intention bindings) and the hot-swap semantics, but it never explains what unload or swap destroys or what happens to an in-flight tree, which is the key behavioral gap for a destructive tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the operation and its action modes, followed by the sibling-disambiguation rule and the return summary. No padding or repetition of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing returns, which it does compactly, and it routes between sibling tools. The remaining gap is the unaddressed destructive side of unload/swap, which matters for a nine-parameter mutation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented in the schema, establishing a baseline of 3. The description adds no syntax or format detail beyond naming the actions and the return fields, so it does not raise the score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb set (load, unload, hot-swap) on a concrete resource (a behavior tree instance in the browser runtime) and enumerates the three action modes. It also explicitly contrasts itself with the sibling manage_behaviors, so an agent can distinguish the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use rule: use load_behavior rather than manage_behaviors when executing an active behavior tree instance at runtime instead of registering or inspecting definitions. The alternative and the selecting condition are both named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_behaviorsA
Manage immutable JSON behavior tree definitions with SHA-256 hash verification (actions: register, list, get, synthesize). Use manage_behaviors instead of load_behavior when defining, versioning, or synthesizing behavior tree structures rather than executing them.
Returns tree definition, version metadata, SHA-256 content hash, or synthesis DAG.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Behavior tree name | |
| tree | No | Behavior tree JSON object | |
| steps | No | Steps or actions to synthesize into a behavior tree | |
| action | Yes | Behavior definition management operation: register, list, get, synthesize | |
| project | No | Target project slug | |
| version | No | Tree version | |
| strategy | No | Root composition strategy for synthesize action (default: sequence) | |
| tree_json | No | Raw behavior tree JSON string | |
| description | No | Tree description | |
| client_request_id | No | Idempotency key |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and idempotentHint=false, and the description adds meaningful context: definitions are 'immutable' and content is SHA-256 verified. However, it never distinguishes which of the four actions mutate state (register/synthesize) versus which are reads (list/get), nor mentions the idempotency key parameter despite idempotentHint=false. With annotations covering the safety profile, the added value is real but partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with purpose and action enumeration followed by the sibling rule and return values. Slight redundancy in restating the tool name in the second sentence, but otherwise efficient and well ordered.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Compensates for the absent output schema by naming return shapes ('tree definition, version metadata, SHA-256 content hash, or synthesis DAG'). For a 10-parameter, nested-object, multi-action tool, the remaining gap is the lack of per-action parameter guidance, since different actions clearly need different subsets of the 10 fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across 10 parameters, so the schema documents each field's type, enum, and meaning. The description adds no per-action parameter mapping (e.g., which of tree/tree_json/name/steps are required for register vs. synthesize), leaving the schema to do all the work. Baseline 3 is appropriate when the schema carries this burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Manage immutable JSON behavior tree definitions with SHA-256 hash verification') and enumerates the four supported actions, so the agent understands this is a multi-action dispatcher rather than a single operation. It also explicitly contrasts itself with the sibling load_behavior, making it distinguishable without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternative tool (load_behavior) and the exact condition that selects this one: 'when defining, versioning, or synthesizing behavior tree structures rather than executing them.' That is an explicit when-to-use vs. alternative statement, which is the top of the scale.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_blackboardADestructive
Read, write, delete, lease, list, or ingest structured state slices into shared behavior tree blackboard state (actions: get, set, delete, lease, list, ingest_slice). Use manage_blackboard instead of set_parameters when coordinating state across behavior nodes, projecting perception/spatial slices with staleness protection, or acquiring agent mutex leases.
Returns blackboard value, lease acquisition status, key listings, or slice ingestion summary with expiry metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| key | No | Blackboard variable key | |
| mode | No | Lease action mode: acquire or release (default: acquire) | |
| value | No | Blackboard variable value | |
| action | Yes | Blackboard operation: get, set, delete, lease, list, ingest_slice | |
| ttl_ms | No | Time-to-live before slice expires and is marked stale (default: 5000ms) | |
| payload | No | Key-value dictionary of slice data to project into blackboard (action: ingest_slice) | |
| project | No | Target project slug | |
| agent_id | No | Agent identifier for lease actions | |
| slice_type | No | Slice category for ingest_slice (default: spatial) | |
| execution_id | No | Target execution ID | |
| duration_seconds | No | Lease duration in seconds |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and non-idempotent/non-read-only, so the safety profile is covered structurally. The description still adds non-obvious behavior beyond that: TTL-based slice staleness ('expires and is marked stale') and lease acquisition semantics with returned status/expiry metadata. It stops short of saying what a delete irreversibly removes or how leases behave on contention.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the verb/resource and action set, with the usage routing second. Slightly wasteful in that the inline action enumeration duplicates the enum already in the schema, but there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 11 parameters, nested payload objects, and no output schema, the description usefully enumerates the return shapes ('blackboard value, lease acquisition status, key listings, or slice ingestion summary with expiry metadata'). It does not map parameters to specific actions (e.g., lease needing agent_id/duration_seconds), but the fully documented schema covers that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across all 11 parameters, so the schema already carries parameter meaning; the description adds essentially nothing beyond restating the action list. This is the correct baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (shared behavior tree blackboard state) and enumerates the concrete operations (get, set, delete, lease, list, ingest_slice), which map exactly to the action enum. It also names the sibling set_parameters it displaces, so an agent can distinguish it without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use criteria and names the alternative: 'Use manage_blackboard instead of set_parameters when coordinating state across behavior nodes, projecting perception/spatial slices with staleness protection, or acquiring agent mutex leases.' The selection condition is stated rather than inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manage_runtime_dbADestructive
Database maintenance, diagnostics, SHA-256 Merkle audit verification, and snapshot management (actions: stats, audit, doctor, snapshot, diff, restore). Use manage_runtime_db instead of get_metrics when auditing SQLite integrity and Merkle proofs or restoring database snapshots.
Returns maintenance diagnostics, Merkle audit trees, snapshot metadata, or diff reports.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Snapshot name | |
| action | Yes | Database maintenance operation: stats, audit, doctor, snapshot, diff, restore | |
| project | No | Target project slug | |
| description | No | Snapshot description |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the safety profile is covered externally. The description adds the return categories (diagnostics, audit trees, snapshot metadata, diff reports), which is useful since there is no output schema, but it never says which of the six actions is the destructive one or that restore overwrites the current database โ the most important behavioral fact is left implicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, capabilities front-loaded and the routing hint second; nothing is padding. The middle clause enumerating six action names is somewhat redundant with the schema enum, keeping it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-mode tool with no output schema, the description gives only a one-line summary of possible returns and no per-action behavioral detail. An agent knows the tool's domain and the get_metrics alternative but not how the actions differ in effect or which parameters each consumes, so it is minimally adequate rather than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the action enum is fully documented in the schema, so the baseline of 3 applies. The description repeats the action names but adds no mapping of which parameters are required for which action (e.g., that 'name' matters only for snapshot/restore/diff), which is the genuinely missing information.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete resource (runtime database) and enumerates the six operations it bundles (stats, audit, doctor, snapshot, diff, restore). It also distinguishes itself from the sibling get_metrics, so an agent can route between the two without opening schemas. It stops short of a 5 only because the six actions are listed rather than individually characterized.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit selection rule against a named alternative: use this instead of get_metrics when auditing SQLite integrity/Merkle proofs or restoring snapshots. That is clear context plus an alternative, but it offers no when-not guidance (e.g., that read-only diagnostic needs should prefer a lighter tool) and no prerequisites such as required permissions for restore.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
register_triggerA
Configure or list reactive interrupt triggers with priority preemption and cooldown guards (actions: register, list). Use register_trigger instead of abort_behavior when defining automatic condition-based interrupts rather than manually pausing or stopping a tree.
Returns registered trigger configuration, ID, or trigger list.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Trigger name | |
| action | Yes | Trigger operation: register, list | |
| project | No | Target project slug | |
| priority | No | Preemption priority | |
| trigger_id | No | Trigger ID | |
| cooldown_ms | No | Minimum cooldown interval between fires | |
| behavior_name | No | Behavior tree to activate when condition fires | |
| condition_type | No | Condition type (e.g. hp_threshold, enemy_proximity) | |
| condition_params | No | Condition evaluation parameters |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare the write/non-idempotent/non-destructive profile, so the description's extra value is the domain behavior it adds: preemption-by-priority and cooldown guarding between fires. It does not describe persistence, permission requirements, or what a failed registration does, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose and the sibling-routing rule before the trailing return-value note. The final sentence is partly redundant with the follow-on description text but is short enough to justify itself.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the two actions, the key configuration concepts, the return shape (configuration, ID, or list), and the sibling routing for a 9-parameter tool with no output schema. Nested condition_params is left to the schema, which is acceptable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so every parameter already carries its own description (priority, cooldown_ms, condition_type, etc.). The description repeats the register/list action split but adds no format or constraint detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Configure or list reactive interrupt triggers') and names the distinctive behavioral features (priority preemption, cooldown guards). It clearly differentiates itself from the sibling abort_behavior, so an agent can pick between them without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to prefer this tool over abort_behavior: 'defining automatic condition-based interrupts rather than manually pausing or stopping a tree.' That is a named alternative plus the condition that selects it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
replay_recordingA
Capture or list deterministic browser action sequences with adaptive timing (actions: capture, list). Use replay_recording instead of load_behavior when recording or inspecting fixed action sequences rather than running a dynamic behavior tree.
Returns recording metadata, frame sequences, or list of stored recordings.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Recording name | |
| action | Yes | Recording operation: capture, list | |
| frames | No | Captured frame sequence | |
| project | No | Target project slug | |
| execution_id | No | Target execution ID | |
| recording_id | No | Recording ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and idempotentHint=false, so the safety profile is covered. The description adds that the tool returns recording metadata, frame sequences, or a list of recordings, which is valuable given there is no output schema, though it does not say whether a capture overwrites an existing recording or which of project/execution_id are needed for capture.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with purpose and the sibling contrast, and the return-value note last. The action enumeration is arguably redundant with the schema enum, but nothing else is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the sentence describing return values (metadata, frame sequences, list of recordings) is a genuine completion of the picture, and the alternative-routing sentence covers selection. Minor gap: capture side effects and required IDs per action are left unstated for a non-idempotent tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters (name, action, frames, project, execution_id, recording_id) are already documented in the schema. The description only names the action values and the notion of 'adaptive timing', adding little beyond what the structured fields provide.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb pair and resource ('Capture or list deterministic browser action sequences') and explicitly enumerates the two supported actions, so an agent knows exactly what the tool does despite the name 'replay_recording' implying a replay action that doesn't exist. It also names the sibling it differs from, load_behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit routing rule: use this instead of load_behavior 'when recording or inspecting fixed action sequences rather than running a dynamic behavior tree.' The alternative and the selecting condition are both stated, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_parametersA
Dynamically set, inspect, or reset execution parameters for the active behavior runtime (actions: set, get, reset). Use set_parameters instead of manage_blackboard for tuning tree-level execution variables and thresholds rather than sharing cross-node data keys.
Returns updated parameter map and execution ID.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Parameter operation: set, get, reset | |
| project | No | Target project slug | |
| parameters | No | Key-value parameters map to apply | |
| execution_id | No | Target execution instance ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false and idempotentHint=false; the description adds the reset capability and the return payload (updated parameter map + execution ID), which is useful behavioral context. It does not flag that reset is non-idempotent/potentially lossy, keeping it short of a 5, but it exceeds the annotation baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core verb+resource and the action list, followed by the sibling routing and return note. Density is high and little is wasted, though the parenthetical action list slightly duplicates the enum.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully states the return value (parameter map + execution ID), and it covers the actions and the sibling distinction. Enough for an agent to invoke correctly, missing only finer detail on reset semantics and project/execution_id targeting.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents action, project, parameters and execution_id, making the baseline 3. The description restates the three action values and implies the parameters map but adds no syntax or format detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States specific verbs (set, inspect, reset) and a precise resource (execution parameters for the active behavior runtime), and it explicitly distinguishes itself from the sibling manage_blackboard by domain (tree-level execution variables vs cross-node data keys). An agent can differentiate this from its siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the alternative (manage_blackboard) and the condition that selects this tool over it: tuning execution variables/thresholds rather than sharing cross-node keys. It lacks guidance on when to choose set vs get vs reset beyond the enum, so it's clear context but not a full when/when-not treatment.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.4.0- Changed
get_metrics3 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Metrics query mode: current, history, aggregate, compare"New value: +"Metrics query mode: current, history, aggregate, compare, spool, drain_spool" - changed
Input schema / properties / action / enumPrevious value: -[ - "current", - "history", - "aggregate", - "compare" -]New value: +[ + "current", + "history", + "aggregate", + "compare", + "spool", + "drain_spool" +] - added
Input schema / properties / unsynced_onlyAdded value: +{ + "description": "Filter only unsynced spool entries (action: spool)", + "type": "boolean" +}
- Changed
manage_blackboard5 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Blackboard operation: get, set, delete, lease, list"New value: +"Blackboard operation: get, set, delete, lease, list, ingest_slice" - changed
Input schema / properties / action / enumPrevious value: -[ - "get", - "set", - "delete", - "lease", - "list" -]New value: +[ + "get", + "set", + "delete", + "lease", + "list", + "ingest_slice" +] - added
Input schema / properties / payloadAdded value: +{ + "description": "Key-value dictionary of slice data to project into blackboard (action: ingest_slice)", + "type": "object" +} - added
Input schema / properties / slice_typeAdded value: +{ + "description": "Slice category for ingest_slice (default: spatial)", + "enum": [ + "spatial", + "visual", + "task", + "vitals" + ], + "type": "string" +} - added
Input schema / properties / ttl_msAdded value: +{ + "description": "Time-to-live before slice expires and is marked stale (default: 5000ms)", + "type": "number" +}
10 tool updates
v0.3.1- Changed
abort_behavior1 field changed- added
Input schema / properties / action / enumAdded value: +[ + "abort", + "pause", + "resume", + "unstick" +]
- Changed
get_metrics2 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Metrics query mode"New value: +"Metrics query mode: current, history, aggregate, compare" - added
Input schema / properties / action / enumAdded value: +[ + "current", + "history", + "aggregate", + "compare" +]
- Changed
get_status2 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Status query mode"New value: +"Status query mode: current, history, tree_state" - added
Input schema / properties / action / enumAdded value: +[ + "current", + "history", + "tree_state" +]
- Changed
load_behavior2 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Behavior loading operation"New value: +"Behavior loading operation: load, unload, swap" - added
Input schema / properties / action / enumAdded value: +[ + "load", + "unload", + "swap" +]
- Changed
manage_behaviors1 field changed- added
Input schema / properties / action / enumAdded value: +[ + "register", + "list", + "get", + "synthesize" +]
- Changed
manage_blackboard1 field changed- added
Input schema / properties / action / enumAdded value: +[ + "get", + "set", + "delete", + "lease", + "list" +]
- Changed
manage_runtime_db2 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Database maintenance operation: stats, audit, doctor (health diagnostics), snapshot, diff, restore"New value: +"Database maintenance operation: stats, audit, doctor, snapshot, diff, restore" - added
Input schema / properties / action / enumAdded value: +[ + "stats", + "audit", + "doctor", + "snapshot", + "diff", + "restore" +]
- Changed
register_trigger2 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Trigger operation"New value: +"Trigger operation: register, list" - added
Input schema / properties / action / enumAdded value: +[ + "register", + "list" +]
- Changed
replay_recording2 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Recording operation"New value: +"Recording operation: capture, list" - added
Input schema / properties / action / enumAdded value: +[ + "capture", + "list" +]
- Changed
set_parameters2 fields changed- changed
Input schema / properties / action / descriptionPrevious value: -"Parameter operation"New value: +"Parameter operation: set, get, reset" - added
Input schema / properties / action / enumAdded value: +[ + "set", + "get", + "reset" +]
10 tool updates
v0.2.1- Changed
abort_behavior8 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / action / descriptionAdded value: +"Execution control operation: abort, pause, resume, unstick (resets stuck score and disengages inputs)" - removed
Input schema / properties / action / enumRemoved value: -[ - "abort", - "pause", - "resume" -] - added
Input schema / properties / execution_id / descriptionAdded value: +"Target execution ID" - added
Input schema / properties / project / descriptionAdded value: +"Target project slug" - added
Input schema / properties / reason / descriptionAdded value: +"Reason for abort or pause" - added
Input schema / properties / trace_id / descriptionAdded value: +"Distributed trace ID"
- Changed
get_metrics9 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / action / descriptionAdded value: +"Metrics query mode" - removed
Input schema / properties / action / enumRemoved value: -[ - "current", - "history", - "aggregate", - "compare" -] - added
Input schema / properties / behavior_name / descriptionAdded value: +"Filter metrics by behavior tree name" - added
Input schema / properties / execution_id / descriptionAdded value: +"Filter metrics by execution ID" - added
Input schema / properties / limit / descriptionAdded value: +"Max records" - added
Input schema / properties / project / descriptionAdded value: +"Target project slug" - added
Input schema / properties / trace_id / descriptionAdded value: +"Distributed trace ID"
- Changed
get_status7 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / action / descriptionAdded value: +"Status query mode" - removed
Input schema / properties / action / enumRemoved value: -[ - "current", - "history", - "tree_state" -] - added
Input schema / properties / execution_id / descriptionAdded value: +"Target execution ID (or latest if omitted)" - added
Input schema / properties / limit / descriptionAdded value: +"Max history entries" - added
Input schema / properties / project / descriptionAdded value: +"Target project slug"
- Changed
load_behavior14 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / action / descriptionAdded value: +"Behavior loading operation" - removed
Input schema / properties / action / enumRemoved value: -[ - "load", - "unload", - "swap" -] - added
Input schema / properties / behavior_name / descriptionAdded value: +"Name of the behavior tree to load" - added
Input schema / properties / behavior_version / descriptionAdded value: +"Version of behavior tree (defaults to latest)" - added
Input schema / properties / client_request_id / descriptionAdded value: +"Idempotency key" - added
Input schema / properties / intention_id / descriptionAdded value: +"Linked agent-reasoning intention ID" - removed
Input schema / properties / parameters / additionalPropertiesRemoved value: -true - added
Input schema / properties / parameters / descriptionAdded value: +"Initial execution parameters" - removed
Input schema / properties / parameters / propertiesRemoved value: -{} - added
Input schema / properties / project / descriptionAdded value: +"Target project slug" - added
Input schema / properties / session_id / descriptionAdded value: +"Linked state-memory session ID" - added
Input schema / properties / trace_id / descriptionAdded value: +"Distributed trace ID for cross-server correlation"
- Changed
manage_behaviors15 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / action / descriptionAdded value: +"Behavior definition management operation: register, list, get, synthesize" - removed
Input schema / properties / action / enumRemoved value: -[ - "register", - "list", - "get" -] - added
Input schema / properties / client_request_id / descriptionAdded value: +"Idempotency key" - added
Input schema / properties / description / descriptionAdded value: +"Tree description" - added
Input schema / properties / name / descriptionAdded value: +"Behavior tree name" - added
Input schema / properties / project / descriptionAdded value: +"Target project slug" - added
Input schema / properties / stepsAdded value: +{ + "description": "Steps or actions to synthesize into a behavior tree", + "items": { + "type": "object" + }, + "type": "array" +} - added
Input schema / properties / strategyAdded value: +{ + "description": "Root composition strategy for synthesize action (default: sequence)", + "enum": [ + "sequence", + "selector", + "parallel" + ], + "type": "string" +} - removed
Input schema / properties / tree / additionalPropertiesRemoved value: -true - added
Input schema / properties / tree / descriptionAdded value: +"Behavior tree JSON object" - removed
Input schema / properties / tree / propertiesRemoved value: -{} - added
Input schema / properties / tree_json / descriptionAdded value: +"Raw behavior tree JSON string" - added
Input schema / properties / version / descriptionAdded value: +"Tree version"
- Changed
manage_blackboard11 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / action / descriptionAdded value: +"Blackboard operation: get, set, delete, lease, list" - removed
Input schema / properties / action / enumRemoved value: -[ - "get", - "set" -] - added
Input schema / properties / agent_idAdded value: +{ + "description": "Agent identifier for lease actions", + "type": "string" +} - added
Input schema / properties / duration_secondsAdded value: +{ + "description": "Lease duration in seconds", + "type": "number" +} - added
Input schema / properties / execution_id / descriptionAdded value: +"Target execution ID" - added
Input schema / properties / key / descriptionAdded value: +"Blackboard variable key" - added
Input schema / properties / modeAdded value: +{ + "description": "Lease action mode: acquire or release (default: acquire)", + "enum": [ + "acquire", + "release" + ], + "type": "string" +} - added
Input schema / properties / project / descriptionAdded value: +"Target project slug" - added
Input schema / properties / value / descriptionAdded value: +"Blackboard variable value"
- Changed
manage_runtime_db7 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / action / descriptionAdded value: +"Database maintenance operation: stats, audit, doctor (health diagnostics), snapshot, diff, restore" - removed
Input schema / properties / action / enumRemoved value: -[ - "stats", - "audit", - "snapshot", - "diff", - "restore" -] - added
Input schema / properties / description / descriptionAdded value: +"Snapshot description" - added
Input schema / properties / name / descriptionAdded value: +"Snapshot name" - added
Input schema / properties / project / descriptionAdded value: +"Target project slug"
- Changed
register_trigger14 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / action / descriptionAdded value: +"Trigger operation" - removed
Input schema / properties / action / enumRemoved value: -[ - "register", - "list" -] - added
Input schema / properties / behavior_name / descriptionAdded value: +"Behavior tree to activate when condition fires" - removed
Input schema / properties / condition_params / additionalPropertiesRemoved value: -true - added
Input schema / properties / condition_params / descriptionAdded value: +"Condition evaluation parameters" - removed
Input schema / properties / condition_params / propertiesRemoved value: -{} - added
Input schema / properties / condition_type / descriptionAdded value: +"Condition type (e.g. hp_threshold, enemy_proximity)" - added
Input schema / properties / cooldown_ms / descriptionAdded value: +"Minimum cooldown interval between fires" - added
Input schema / properties / name / descriptionAdded value: +"Trigger name" - added
Input schema / properties / priority / descriptionAdded value: +"Preemption priority" - added
Input schema / properties / project / descriptionAdded value: +"Target project slug" - added
Input schema / properties / trigger_id / descriptionAdded value: +"Trigger ID"
- Changed
replay_recording11 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / action / descriptionAdded value: +"Recording operation" - removed
Input schema / properties / action / enumRemoved value: -[ - "capture", - "list" -] - added
Input schema / properties / execution_id / descriptionAdded value: +"Target execution ID" - added
Input schema / properties / frames / descriptionAdded value: +"Captured frame sequence" - removed
Input schema / properties / frames / items / additionalPropertiesRemoved value: -true - removed
Input schema / properties / frames / items / propertiesRemoved value: -{} - added
Input schema / properties / name / descriptionAdded value: +"Recording name" - added
Input schema / properties / project / descriptionAdded value: +"Target project slug" - added
Input schema / properties / recording_id / descriptionAdded value: +"Recording ID"
- Changed
set_parameters9 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / action / descriptionAdded value: +"Parameter operation" - removed
Input schema / properties / action / enumRemoved value: -[ - "set", - "get", - "reset" -] - added
Input schema / properties / execution_id / descriptionAdded value: +"Target execution instance ID" - removed
Input schema / properties / parameters / additionalPropertiesRemoved value: -true - added
Input schema / properties / parameters / descriptionAdded value: +"Key-value parameters map to apply" - removed
Input schema / properties / parameters / propertiesRemoved value: -{} - added
Input schema / properties / project / descriptionAdded value: +"Target project slug"
10 tool updates
v0.1.2- First observed
abort_behavior - First observed
get_metrics - First observed
get_status - First observed
load_behavior - First observed
manage_behaviors - First observed
manage_blackboard - First observed
manage_runtime_db - First observed
register_trigger - First observed
replay_recording - First observed
set_parameters
TDQS
Scored across 10 tools
Most tools target distinct runtime concerns, but several pairs (manage_behaviors vs load_behavior, get_metrics vs get_status, abort_behavior vs register_trigger, manage_blackboard vs set_parameters) are close enough that descriptions must explicitly disambiguate them. After reading those descriptions an agent can choose correctly, so boundaries are clear but not self-evident.
All names use a consistent snake_case verb_noun pattern: manage_behaviors, load_behavior, get_status, set_parameters, etc. The mixed verbs (manage, get, set, load, abort) are appropriate and used predictably, with no camelCase or stylistic inconsistency.
Ten tools is well-scoped for a behavior runtime server covering definitions, execution, state, telemetry, triggers, recordings, and database maintenance. Each tool has a clear place and the set does not feel bloated or thin.
The surface covers most lifecycle needs: behavior definitions, runtime load/unload/swap, status, metrics, blackboard state, parameters, triggers, abort/pause/resume, recordings, and DB maintenance. However, replay_recording exposes only capture/list despite its name, and behavior definitions lack obvious delete/update operations, leaving minor gaps.
Maintenance
Related MCP Connectors
Real-time planetary signal engine and Model Context Protocol (MCP) server for autonomous AI agents.
Control plane for autonomous software labor. Agents claim objectives over MCP with audit trail.
- DazbenchOAuthapp.dazbench
Task management your AI agents can actually run. One line becomes a context-ready task over MCP.
- memnodeOAuthdev.memnode
Persistent, inspectable memory for AI agents with lineage, correction, and a hosted MCP endpoint.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceAn MCP server that enables AI agents to execute sandboxed JavaScript code to control external web applications in real-time. It utilizes WebSockets to synchronize script execution with frontend animations, allowing complex UI interactions to be handled in a single agent turn.-
- AlicenseAqualityAmaintenanceStrategic BDI cognitive reasoning engine for AI agents. Goal decomposition DAGs, multi-factor utility scoring, risk evaluation, and adaptive replanning over MCP.159,039 npm19MIT

genpark-jev-system1official
AlicenseNot gradedqualityBmaintenanceEnables AI agents to resolve typed routing, loop-stall arbitration, and tool safety checks locally at sub-millisecond latency, avoiding unnecessary frontier LLM calls. It enforces reversible execution checkpoints, token budgets, and decayed episodic memory for safe, privacy-first autonomous operation.7MIT- AlicenseNot gradedqualityBmaintenanceEnables AI agents to combine ambient multimodal perception, continuous episodic memory, and recency-decay context briefing for fast deterministic decisions and safe tool routing. It integrates with MCP clients to reduce unnecessary frontier LLM calls while enforcing token budgets and local-first memory retention.7MIT