Skip to main content
Glama

@putervision/behavior-mcp

npm version version CI Node License: MIT

High-Frequency (~60Hz) In-Browser Behavior Tree Tactical Runtime Engine for AI Agents

@putervision/behavior-mcp is a Model Context Protocol (MCP) server that executes deterministic behavior trees directly in browser runtimes at ~60Hz with reactive trigger preemption, frame recordings, SHA-256 Merkle audit chains, and a 5-layer safety guardrail stack.

๐ŸŒ Official Documentation: putervision.com โ€ข Interactive Web Docs


โšก 15-Second Quick Start

# 1. Initialize runtime & seed default behavior trees in your workspace
npx @putervision/behavior-mcp init

# 2. Run system diagnostic and health checks
npx @putervision/behavior-mcp doctor

# 3. Inspect active executions, blackboard state, and reactive triggers
npx @putervision/behavior-mcp inspect

Related MCP server: agent-reasoning-mcp

๐Ÿ› ๏ธ 10 Core MCP Tools

Tool

Actions

Purpose

load_behavior

load, unload, swap

Inject, initialize, or hot-swap a behavior tree in browser runtime

set_parameters

set, get, reset

Dynamically update or query runtime execution parameters

get_status

current, history, tree_state

Query active status, node traversal path, tick count, duration

abort_behavior

abort, pause, resume

Halt, pause, or resume behavior execution safely

register_trigger

register, list

Configure reactive interrupt triggers with priority preemption

replay_recording

capture, list

Capture or inspect deterministic browser action sequences

get_metrics

current, history

Retrieve runtime telemetry, tick durations, and stuck events

manage_behaviors

register, get, list

CRUD for immutable behavior tree JSON with SHA-256 tree hashes

manage_blackboard

get, set

Read, write, or query shared behavior tree blackboard variables

manage_runtime_db

stats, audit, snapshot, restore

SQLite diagnostics, SHA-256 Merkle audit, and checkpoint rollback


๐Ÿ›ก๏ธ 5-Layer Safety Guardrail Stack & System 1 Invariants

  1. Fail-Closed Evaluator: Unrecognized node definitions throw fatal exceptions immediately.

  2. 60Hz Tick Invariant & Synchronous semantic_check: The loop never blocks on external network calls; semantic condition checks resolve synchronously against blackboard caches.

  3. HMAC Intention Gate Verification: Intentions dispatched to execution nodes require unexpired, cryptographically signed dispatch tokens (PENTAD_HMAC_SECRET).

  4. Action Rate Limiter & Leaf Policy Gate: Strict 60 actions/sec maximum throughput ceiling and policy guardrails blocking irreversible mutations.

  5. Emergency Kill Switch & Watchdog Heartbeat: Atomic safety latch and watchdog timers halting execution if tick stalls beyond 5,000ms.


๐Ÿ“š Deep Documentation Guides


๐Ÿ”— Client Configuration & Environment

Add to .cursor/mcp.json or .vscode/mcp.json:

{
  "mcpServers": {
    "behavior-mcp": {
      "command": "behavior-mcp",
      "args": ["run"],
      "env": {
        "PENTAD_HMAC_SECRET": "your-secure-shared-secret-here"
      }
    }
  }
}

๐Ÿงช Testing

# Run full unit and integration test suite across 38 test files (266 tests)
npm test

๐Ÿ“„ License

MIT ยฉ PuterVision

Available Tools

10 tools
abort_behaviorA

Immediately halt, pause, resume, or unstick behavior execution and disengage active inputs (actions: abort, pause, resume, unstick). Use abort_behavior instead of register_trigger when manually halting or recovering execution rather than configuring automatic condition interrupts.

Returns transition status, disengaged inputs, and unstick diagnostics.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesExecution control operation: abort, pause, resume, unstick (resets stuck score and disengages inputs)
reasonNoReason for abort or pause
projectNoTarget project slug
trace_idNoDistributed trace ID
execution_idNoTarget execution ID

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare non-readOnly, non-idempotent, non-destructive, so safety profile is covered. The description adds real behavioral value beyond them: it discloses the side effect of disengaging active inputs and enumerates what comes back (transition status, disengaged inputs, unstick diagnostics). It does not mention authorization requirements or whether abort is recoverable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the operations and kept to three tight sentences. The action set is repeated in the parenthetical after already being named, which is mild redundancy, but nothing else is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter mutation tool with no output schema, the description covers purpose, the sibling alternative, and the return payload, which is enough to invoke it correctly. It omits prerequisites (need for an existing execution/trace context) and the relationship between the four actions, leaving minor gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the action enum is already documented with per-value semantics, so the schema carries the load. The description only restates the action list and adds no format or interaction detail for reason/project/trace_id/execution_id, making the baseline 3 correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names the concrete operations (abort, pause, resume, unstick) against a specific resource (behavior execution) and states the side effect (disengaging active inputs). It also names the sibling it is not (register_trigger), so an agent can distinguish it without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: 'Use abort_behavior instead of register_trigger when manually halting or recovering execution rather than configuring automatic condition interrupts.' That gives a clear alternative plus the selecting condition. It stops short of 5 because it offers no guidance on choosing among the four actions (abort vs pause vs unstick vs resume) at this level.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_metricsA
Read-onlyIdempotent

Retrieve runtime execution telemetry, tick durations, stuck events, category statistics, and 60Hz action outcome spool entries (actions: current, history, aggregate, compare, spool, drain_spool). Use get_metrics instead of get_status when evaluating tick performance, reading spooled action outcomes, or inspecting aggregated statistics rather than active node traversal.

Returns telemetry metrics, duration percentiles, stuck event counts, spooled outcome records, or drained entry counts.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax records
actionYesMetrics query mode: current, history, aggregate, compare, spool, drain_spool
projectNoTarget project slug
trace_idNoDistributed trace ID
execution_idNoFilter metrics by execution ID
behavior_nameNoFilter metrics by behavior tree name
unsynced_onlyNoFilter only unsynced spool entries (action: spool)

TDQS

A3.6/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and idempotentHint=true, but the tool exposes a drain_spool action returning 'drained entry counts' โ€” draining consumes/removes spool entries, which is a state mutation and is not idempotent. The description thus contradicts the read-only/idempotent safety profile for at least one action, and it never warns that one action differs in side effects from the rest.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the resource list, then routes to the alternative, then summarizes returns โ€” a sensible order. The 'Returns telemetry metrics...' sentence largely re-states the opening list, and the action enum is repeated from the schema, adding minor redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-param, 6-mode read/mutate hybrid with no output schema, the description does cover purpose, alternatives, and return categories. However it fails to disambiguate the side-effect profile across actions (drain_spool vs read-only modes), which is exactly the information an agent needs before invoking, so it is adequate but with a material gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented, and the description merely repeats the action enum verbatim. It adds no syntax, format, or meaning (e.g., how trace_id interacts with execution_id filter, or what behavior_name expects) beyond the schema, matching the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Retrieve) plus the concrete resource set: runtime telemetry, tick durations, stuck events, category statistics, and 60Hz action outcome spool entries. It also names the sibling it differs from (get_status) and the axis of difference (performance/spooled outcomes vs active node traversal), so an agent can separate the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Contains an explicit when-to-use/when-not clause: 'Use get_metrics instead of get_status when evaluating tick performance, reading spooled action outcomes, or inspecting aggregated statistics rather than active node traversal.' The alternative tool and the selecting condition are both named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_statusA
Read-onlyIdempotent

Query active behavior execution status, history, or full tree node traversal state (actions: current, history, tree_state). Use get_status instead of get_metrics when inspecting active execution state and node traversal paths rather than aggregated runtime performance telemetry.

Returns execution status, current node path, tick counters, duration, and error states.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax history entries
actionYesStatus query mode: current, history, tree_state
projectNoTarget project slug
execution_idNoTarget execution ID (or latest if omitted)

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, non-destructive and closed-world, so the safety profile is covered. The description adds value beyond that by enumerating what is returned (execution status, node path, tick counters, duration, error states), which matters since there is no output schema. It stops short of noting auth requirements, rate limits, or pagination behavior for history.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the action modes and the routing rule front-loaded, followed by the return summary. Very little waste, though the parenthetical enum listing is mildly redundant with the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only query tool with full schema coverage and safety annotations, the definition covers purpose, alternatives, and return fields, compensating for the absent output schema. Minor gaps remain around history pagination and default-to-latest execution behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents action, project, execution_id, and limit. The description's parenthetical restates the action enum but adds no new syntax or format detail. Baseline 3 is appropriate when the schema carries the parameter burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Query) and resource (behavior execution status/history/tree traversal), and enumerates the three action modes. It also names the sibling get_metrics and explains the distinction, so an agent can differentiate without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use get_status instead of get_metrics and gives the selecting condition: inspecting active execution state and node traversal paths vs. aggregated runtime performance telemetry. This is a full when/when-not/alternative statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

load_behaviorA
Destructive

Load, unload, or hot-swap a behavior tree instance in the browser runtime (actions: load, unload, swap). Use load_behavior instead of manage_behaviors when executing an active behavior tree instance at runtime rather than registering or inspecting definitions.

Returns execution handle, runtime state, active node, and session/intention bindings.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesBehavior loading operation: load, unload, swap
projectNoTarget project slug
trace_idNoDistributed trace ID for cross-server correlation
parametersNoInitial execution parameters
session_idNoLinked state-memory session ID
intention_idNoLinked agent-reasoning intention ID
behavior_nameYesName of the behavior tree to load
behavior_versionNoVersion of behavior tree (defaults to latest)
client_request_idNoIdempotency key

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the safety profile is covered structurally. The description usefully adds what the call returns (handle, runtime state, active node, session/intention bindings) and the hot-swap semantics, but it never explains what unload or swap destroys or what happens to an in-flight tree, which is the key behavioral gap for a destructive tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the operation and its action modes, followed by the sibling-disambiguation rule and the return summary. No padding or repetition of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of describing returns, which it does compactly, and it routes between sibling tools. The remaining gap is the unaddressed destructive side of unload/swap, which matters for a nine-parameter mutation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented in the schema, establishing a baseline of 3. The description adds no syntax or format detail beyond naming the actions and the return fields, so it does not raise the score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb set (load, unload, hot-swap) on a concrete resource (a behavior tree instance in the browser runtime) and enumerates the three action modes. It also explicitly contrasts itself with the sibling manage_behaviors, so an agent can distinguish the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-to-use rule: use load_behavior rather than manage_behaviors when executing an active behavior tree instance at runtime instead of registering or inspecting definitions. The alternative and the selecting condition are both named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_behaviorsA

Manage immutable JSON behavior tree definitions with SHA-256 hash verification (actions: register, list, get, synthesize). Use manage_behaviors instead of load_behavior when defining, versioning, or synthesizing behavior tree structures rather than executing them.

Returns tree definition, version metadata, SHA-256 content hash, or synthesis DAG.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoBehavior tree name
treeNoBehavior tree JSON object
stepsNoSteps or actions to synthesize into a behavior tree
actionYesBehavior definition management operation: register, list, get, synthesize
projectNoTarget project slug
versionNoTree version
strategyNoRoot composition strategy for synthesize action (default: sequence)
tree_jsonNoRaw behavior tree JSON string
descriptionNoTree description
client_request_idNoIdempotency key

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, and idempotentHint=false, and the description adds meaningful context: definitions are 'immutable' and content is SHA-256 verified. However, it never distinguishes which of the four actions mutate state (register/synthesize) versus which are reads (list/get), nor mentions the idempotency key parameter despite idempotentHint=false. With annotations covering the safety profile, the added value is real but partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with purpose and action enumeration followed by the sibling rule and return values. Slight redundancy in restating the tool name in the second sentence, but otherwise efficient and well ordered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Compensates for the absent output schema by naming return shapes ('tree definition, version metadata, SHA-256 content hash, or synthesis DAG'). For a 10-parameter, nested-object, multi-action tool, the remaining gap is the lack of per-action parameter guidance, since different actions clearly need different subsets of the 10 fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% across 10 parameters, so the schema documents each field's type, enum, and meaning. The description adds no per-action parameter mapping (e.g., which of tree/tree_json/name/steps are required for register vs. synthesize), leaving the schema to do all the work. Baseline 3 is appropriate when the schema carries this burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Manage immutable JSON behavior tree definitions with SHA-256 hash verification') and enumerates the four supported actions, so the agent understands this is a multi-action dispatcher rather than a single operation. It also explicitly contrasts itself with the sibling load_behavior, making it distinguishable without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative tool (load_behavior) and the exact condition that selects this one: 'when defining, versioning, or synthesizing behavior tree structures rather than executing them.' That is an explicit when-to-use vs. alternative statement, which is the top of the scale.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_blackboardA
Destructive

Read, write, delete, lease, list, or ingest structured state slices into shared behavior tree blackboard state (actions: get, set, delete, lease, list, ingest_slice). Use manage_blackboard instead of set_parameters when coordinating state across behavior nodes, projecting perception/spatial slices with staleness protection, or acquiring agent mutex leases.

Returns blackboard value, lease acquisition status, key listings, or slice ingestion summary with expiry metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNoBlackboard variable key
modeNoLease action mode: acquire or release (default: acquire)
valueNoBlackboard variable value
actionYesBlackboard operation: get, set, delete, lease, list, ingest_slice
ttl_msNoTime-to-live before slice expires and is marked stale (default: 5000ms)
payloadNoKey-value dictionary of slice data to project into blackboard (action: ingest_slice)
projectNoTarget project slug
agent_idNoAgent identifier for lease actions
slice_typeNoSlice category for ingest_slice (default: spatial)
execution_idNoTarget execution ID
duration_secondsNoLease duration in seconds

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and non-idempotent/non-read-only, so the safety profile is covered structurally. The description still adds non-obvious behavior beyond that: TTL-based slice staleness ('expires and is marked stale') and lease acquisition semantics with returned status/expiry metadata. It stops short of saying what a delete irreversibly removes or how leases behave on contention.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the verb/resource and action set, with the usage routing second. Slightly wasteful in that the inline action enumeration duplicates the enum already in the schema, but there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 11 parameters, nested payload objects, and no output schema, the description usefully enumerates the return shapes ('blackboard value, lease acquisition status, key listings, or slice ingestion summary with expiry metadata'). It does not map parameters to specific actions (e.g., lease needing agent_id/duration_seconds), but the fully documented schema covers that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% across all 11 parameters, so the schema already carries parameter meaning; the description adds essentially nothing beyond restating the action list. This is the correct baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (shared behavior tree blackboard state) and enumerates the concrete operations (get, set, delete, lease, list, ingest_slice), which map exactly to the action enum. It also names the sibling set_parameters it displaces, so an agent can distinguish it without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use criteria and names the alternative: 'Use manage_blackboard instead of set_parameters when coordinating state across behavior nodes, projecting perception/spatial slices with staleness protection, or acquiring agent mutex leases.' The selection condition is stated rather than inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

manage_runtime_dbA
Destructive

Database maintenance, diagnostics, SHA-256 Merkle audit verification, and snapshot management (actions: stats, audit, doctor, snapshot, diff, restore). Use manage_runtime_db instead of get_metrics when auditing SQLite integrity and Merkle proofs or restoring database snapshots.

Returns maintenance diagnostics, Merkle audit trees, snapshot metadata, or diff reports.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoSnapshot name
actionYesDatabase maintenance operation: stats, audit, doctor, snapshot, diff, restore
projectNoTarget project slug
descriptionNoSnapshot description

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the safety profile is covered externally. The description adds the return categories (diagnostics, audit trees, snapshot metadata, diff reports), which is useful since there is no output schema, but it never says which of the six actions is the destructive one or that restore overwrites the current database โ€” the most important behavioral fact is left implicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, capabilities front-loaded and the routing hint second; nothing is padding. The middle clause enumerating six action names is somewhat redundant with the schema enum, keeping it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-mode tool with no output schema, the description gives only a one-line summary of possible returns and no per-action behavioral detail. An agent knows the tool's domain and the get_metrics alternative but not how the actions differ in effect or which parameters each consumes, so it is minimally adequate rather than complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the action enum is fully documented in the schema, so the baseline of 3 applies. The description repeats the action names but adds no mapping of which parameters are required for which action (e.g., that 'name' matters only for snapshot/restore/diff), which is the genuinely missing information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete resource (runtime database) and enumerates the six operations it bundles (stats, audit, doctor, snapshot, diff, restore). It also distinguishes itself from the sibling get_metrics, so an agent can route between the two without opening schemas. It stops short of a 5 only because the six actions are listed rather than individually characterized.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit selection rule against a named alternative: use this instead of get_metrics when auditing SQLite integrity/Merkle proofs or restoring snapshots. That is clear context plus an alternative, but it offers no when-not guidance (e.g., that read-only diagnostic needs should prefer a lighter tool) and no prerequisites such as required permissions for restore.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

register_triggerA

Configure or list reactive interrupt triggers with priority preemption and cooldown guards (actions: register, list). Use register_trigger instead of abort_behavior when defining automatic condition-based interrupts rather than manually pausing or stopping a tree.

Returns registered trigger configuration, ID, or trigger list.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoTrigger name
actionYesTrigger operation: register, list
projectNoTarget project slug
priorityNoPreemption priority
trigger_idNoTrigger ID
cooldown_msNoMinimum cooldown interval between fires
behavior_nameNoBehavior tree to activate when condition fires
condition_typeNoCondition type (e.g. hp_threshold, enemy_proximity)
condition_paramsNoCondition evaluation parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare the write/non-idempotent/non-destructive profile, so the description's extra value is the domain behavior it adds: preemption-by-priority and cooldown guarding between fires. It does not describe persistence, permission requirements, or what a failed registration does, which keeps it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with purpose and the sibling-routing rule before the trailing return-value note. The final sentence is partly redundant with the follow-on description text but is short enough to justify itself.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the two actions, the key configuration concepts, the return shape (configuration, ID, or list), and the sibling routing for a 9-parameter tool with no output schema. Nested condition_params is left to the schema, which is acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so every parameter already carries its own description (priority, cooldown_ms, condition_type, etc.). The description repeats the register/list action split but adds no format or constraint detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Configure or list reactive interrupt triggers') and names the distinctive behavioral features (priority preemption, cooldown guards). It clearly differentiates itself from the sibling abort_behavior, so an agent can pick between them without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to prefer this tool over abort_behavior: 'defining automatic condition-based interrupts rather than manually pausing or stopping a tree.' That is a named alternative plus the condition that selects it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

replay_recordingA

Capture or list deterministic browser action sequences with adaptive timing (actions: capture, list). Use replay_recording instead of load_behavior when recording or inspecting fixed action sequences rather than running a dynamic behavior tree.

Returns recording metadata, frame sequences, or list of stored recordings.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoRecording name
actionYesRecording operation: capture, list
framesNoCaptured frame sequence
projectNoTarget project slug
execution_idNoTarget execution ID
recording_idNoRecording ID

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, and idempotentHint=false, so the safety profile is covered. The description adds that the tool returns recording metadata, frame sequences, or a list of recordings, which is valuable given there is no output schema, though it does not say whether a capture overwrites an existing recording or which of project/execution_id are needed for capture.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with purpose and the sibling contrast, and the return-value note last. The action enumeration is arguably redundant with the schema enum, but nothing else is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the sentence describing return values (metadata, frame sequences, list of recordings) is a genuine completion of the picture, and the alternative-routing sentence covers selection. Minor gap: capture side effects and required IDs per action are left unstated for a non-idempotent tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters (name, action, frames, project, execution_id, recording_id) are already documented in the schema. The description only names the action values and the notion of 'adaptive timing', adding little beyond what the structured fields provide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb pair and resource ('Capture or list deterministic browser action sequences') and explicitly enumerates the two supported actions, so an agent knows exactly what the tool does despite the name 'replay_recording' implying a replay action that doesn't exist. It also names the sibling it differs from, load_behavior.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit routing rule: use this instead of load_behavior 'when recording or inspecting fixed action sequences rather than running a dynamic behavior tree.' The alternative and the selecting condition are both stated, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_parametersA

Dynamically set, inspect, or reset execution parameters for the active behavior runtime (actions: set, get, reset). Use set_parameters instead of manage_blackboard for tuning tree-level execution variables and thresholds rather than sharing cross-node data keys.

Returns updated parameter map and execution ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesParameter operation: set, get, reset
projectNoTarget project slug
parametersNoKey-value parameters map to apply
execution_idNoTarget execution instance ID

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false and idempotentHint=false; the description adds the reset capability and the return payload (updated parameter map + execution ID), which is useful behavioral context. It does not flag that reset is non-idempotent/potentially lossy, keeping it short of a 5, but it exceeds the annotation baseline.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core verb+resource and the action list, followed by the sibling routing and return note. Density is high and little is wasted, though the parenthetical action list slightly duplicates the enum.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully states the return value (parameter map + execution ID), and it covers the actions and the sibling distinction. Enough for an agent to invoke correctly, missing only finer detail on reset semantics and project/execution_id targeting.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents action, project, parameters and execution_id, making the baseline 3. The description restates the three action values and implies the parameters map but adds no syntax or format detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States specific verbs (set, inspect, reset) and a precise resource (execution parameters for the active behavior runtime), and it explicitly distinguishes itself from the sibling manage_blackboard by domain (tree-level execution variables vs cross-node data keys). An agent can differentiate this from its siblings without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names the alternative (manage_blackboard) and the condition that selects this tool over it: tuning execution variables/thresholds rather than sharing cross-node keys. It lacks guidance on when to choose set vs get vs reset beyond the enum, so it's clear context but not a full when/when-not treatment.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.4.0
    • Changedget_metrics3 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Metrics query mode: current, history, aggregate, compare"New value: +"Metrics query mode: current, history, aggregate, compare, spool, drain_spool"
      • changedInput schema / properties / action / enum
        Previous value: -[
        -  "current",
        -  "history",
        -  "aggregate",
        -  "compare"
        -]New value: +[
        +  "current",
        +  "history",
        +  "aggregate",
        +  "compare",
        +  "spool",
        +  "drain_spool"
        +]
      • addedInput schema / properties / unsynced_only
        Added value: +{
        +  "description": "Filter only unsynced spool entries (action: spool)",
        +  "type": "boolean"
        +}
    • Changedmanage_blackboard5 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Blackboard operation: get, set, delete, lease, list"New value: +"Blackboard operation: get, set, delete, lease, list, ingest_slice"
      • changedInput schema / properties / action / enum
        Previous value: -[
        -  "get",
        -  "set",
        -  "delete",
        -  "lease",
        -  "list"
        -]New value: +[
        +  "get",
        +  "set",
        +  "delete",
        +  "lease",
        +  "list",
        +  "ingest_slice"
        +]
      • addedInput schema / properties / payload
        Added value: +{
        +  "description": "Key-value dictionary of slice data to project into blackboard (action: ingest_slice)",
        +  "type": "object"
        +}
      • addedInput schema / properties / slice_type
        Added value: +{
        +  "description": "Slice category for ingest_slice (default: spatial)",
        +  "enum": [
        +    "spatial",
        +    "visual",
        +    "task",
        +    "vitals"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / ttl_ms
        Added value: +{
        +  "description": "Time-to-live before slice expires and is marked stale (default: 5000ms)",
        +  "type": "number"
        +}
  2. 10 tool updatesv0.3.1
    • Changedabort_behavior1 field changed
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "abort",
        +  "pause",
        +  "resume",
        +  "unstick"
        +]
    • Changedget_metrics2 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Metrics query mode"New value: +"Metrics query mode: current, history, aggregate, compare"
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "current",
        +  "history",
        +  "aggregate",
        +  "compare"
        +]
    • Changedget_status2 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Status query mode"New value: +"Status query mode: current, history, tree_state"
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "current",
        +  "history",
        +  "tree_state"
        +]
    • Changedload_behavior2 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Behavior loading operation"New value: +"Behavior loading operation: load, unload, swap"
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "load",
        +  "unload",
        +  "swap"
        +]
    • Changedmanage_behaviors1 field changed
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "register",
        +  "list",
        +  "get",
        +  "synthesize"
        +]
    • Changedmanage_blackboard1 field changed
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "get",
        +  "set",
        +  "delete",
        +  "lease",
        +  "list"
        +]
    • Changedmanage_runtime_db2 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Database maintenance operation: stats, audit, doctor (health diagnostics), snapshot, diff, restore"New value: +"Database maintenance operation: stats, audit, doctor, snapshot, diff, restore"
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "stats",
        +  "audit",
        +  "doctor",
        +  "snapshot",
        +  "diff",
        +  "restore"
        +]
    • Changedregister_trigger2 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Trigger operation"New value: +"Trigger operation: register, list"
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "register",
        +  "list"
        +]
    • Changedreplay_recording2 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Recording operation"New value: +"Recording operation: capture, list"
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "capture",
        +  "list"
        +]
    • Changedset_parameters2 fields changed
      • changedInput schema / properties / action / description
        Previous value: -"Parameter operation"New value: +"Parameter operation: set, get, reset"
      • addedInput schema / properties / action / enum
        Added value: +[
        +  "set",
        +  "get",
        +  "reset"
        +]
  3. 10 tool updatesv0.2.1
    • Changedabort_behavior8 fields changed
      • removedInput schema / $schema
        Removed value: -"http://json-schema.org/draft-07/schema#"
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / action / description
        Added value: +"Execution control operation: abort, pause, resume, unstick (resets stuck score and disengages inputs)"
      • removedInput schema / properties / action / enum
        Removed value: -[
        -  "abort",
        -  "pause",
        -  "resume"
        -]
      • addedInput schema / properties / execution_id / description
        Added value: +"Target execution ID"
      • addedInput schema / properties / project / description
        Added value: +"Target project slug"
      • addedInput schema / properties / reason / description
        Added value: +"Reason for abort or pause"
      • addedInput schema / properties / trace_id / description
        Added value: +"Distributed trace ID"
    • Changedget_metrics9 fields changed
      • removedInput schema / $schema
        Removed value: -"http://json-schema.org/draft-07/schema#"
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / action / description
        Added value: +"Metrics query mode"
      • removedInput schema / properties / action / enum
        Removed value: -[
        -  "current",
        -  "history",
        -  "aggregate",
        -  "compare"
        -]
      • addedInput schema / properties / behavior_name / description
        Added value: +"Filter metrics by behavior tree name"
      • addedInput schema / properties / execution_id / description
        Added value: +"Filter metrics by execution ID"
      • addedInput schema / properties / limit / description
        Added value: +"Max records"
      • addedInput schema / properties / project / description
        Added value: +"Target project slug"
      • addedInput schema / properties / trace_id / description
        Added value: +"Distributed trace ID"
    • Changedget_status7 fields changed
      • removedInput schema / $schema
        Removed value: -"http://json-schema.org/draft-07/schema#"
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / action / description
        Added value: +"Status query mode"
      • removedInput schema / properties / action / enum
        Removed value: -[
        -  "current",
        -  "history",
        -  "tree_state"
        -]
      • addedInput schema / properties / execution_id / description
        Added value: +"Target execution ID (or latest if omitted)"
      • addedInput schema / properties / limit / description
        Added value: +"Max history entries"
      • addedInput schema / properties / project / description
        Added value: +"Target project slug"
    • Changedload_behavior14 fields changed
      • removedInput schema / $schema
        Removed value: -"http://json-schema.org/draft-07/schema#"
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / action / description
        Added value: +"Behavior loading operation"
      • removedInput schema / properties / action / enum
        Removed value: -[
        -  "load",
        -  "unload",
        -  "swap"
        -]
      • addedInput schema / properties / behavior_name / description
        Added value: +"Name of the behavior tree to load"
      • addedInput schema / properties / behavior_version / description
        Added value: +"Version of behavior tree (defaults to latest)"
      • addedInput schema / properties / client_request_id / description
        Added value: +"Idempotency key"
      • addedInput schema / properties / intention_id / description
        Added value: +"Linked agent-reasoning intention ID"
      • removedInput schema / properties / parameters / additionalProperties
        Removed value: -true
      • addedInput schema / properties / parameters / description
        Added value: +"Initial execution parameters"
      • removedInput schema / properties / parameters / properties
        Removed value: -{}
      • addedInput schema / properties / project / description
        Added value: +"Target project slug"
      • addedInput schema / properties / session_id / description
        Added value: +"Linked state-memory session ID"
      • addedInput schema / properties / trace_id / description
        Added value: +"Distributed trace ID for cross-server correlation"
    • Changedmanage_behaviors15 fields changed
      • removedInput schema / $schema
        Removed value: -"http://json-schema.org/draft-07/schema#"
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / action / description
        Added value: +"Behavior definition management operation: register, list, get, synthesize"
      • removedInput schema / properties / action / enum
        Removed value: -[
        -  "register",
        -  "list",
        -  "get"
        -]
      • addedInput schema / properties / client_request_id / description
        Added value: +"Idempotency key"
      • addedInput schema / properties / description / description
        Added value: +"Tree description"
      • addedInput schema / properties / name / description
        Added value: +"Behavior tree name"
      • addedInput schema / properties / project / description
        Added value: +"Target project slug"
      • addedInput schema / properties / steps
        Added value: +{
        +  "description": "Steps or actions to synthesize into a behavior tree",
        +  "items": {
        +    "type": "object"
        +  },
        +  "type": "array"
        +}
      • addedInput schema / properties / strategy
        Added value: +{
        +  "description": "Root composition strategy for synthesize action (default: sequence)",
        +  "enum": [
        +    "sequence",
        +    "selector",
        +    "parallel"
        +  ],
        +  "type": "string"
        +}
      • removedInput schema / properties / tree / additionalProperties
        Removed value: -true
      • addedInput schema / properties / tree / description
        Added value: +"Behavior tree JSON object"
      • removedInput schema / properties / tree / properties
        Removed value: -{}
      • addedInput schema / properties / tree_json / description
        Added value: +"Raw behavior tree JSON string"
      • addedInput schema / properties / version / description
        Added value: +"Tree version"
    • Changedmanage_blackboard11 fields changed
      • removedInput schema / $schema
        Removed value: -"http://json-schema.org/draft-07/schema#"
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / action / description
        Added value: +"Blackboard operation: get, set, delete, lease, list"
      • removedInput schema / properties / action / enum
        Removed value: -[
        -  "get",
        -  "set"
        -]
      • addedInput schema / properties / agent_id
        Added value: +{
        +  "description": "Agent identifier for lease actions",
        +  "type": "string"
        +}
      • addedInput schema / properties / duration_seconds
        Added value: +{
        +  "description": "Lease duration in seconds",
        +  "type": "number"
        +}
      • addedInput schema / properties / execution_id / description
        Added value: +"Target execution ID"
      • addedInput schema / properties / key / description
        Added value: +"Blackboard variable key"
      • addedInput schema / properties / mode
        Added value: +{
        +  "description": "Lease action mode: acquire or release (default: acquire)",
        +  "enum": [
        +    "acquire",
        +    "release"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / project / description
        Added value: +"Target project slug"
      • addedInput schema / properties / value / description
        Added value: +"Blackboard variable value"
    • Changedmanage_runtime_db7 fields changed
      • removedInput schema / $schema
        Removed value: -"http://json-schema.org/draft-07/schema#"
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / action / description
        Added value: +"Database maintenance operation: stats, audit, doctor (health diagnostics), snapshot, diff, restore"
      • removedInput schema / properties / action / enum
        Removed value: -[
        -  "stats",
        -  "audit",
        -  "snapshot",
        -  "diff",
        -  "restore"
        -]
      • addedInput schema / properties / description / description
        Added value: +"Snapshot description"
      • addedInput schema / properties / name / description
        Added value: +"Snapshot name"
      • addedInput schema / properties / project / description
        Added value: +"Target project slug"
    • Changedregister_trigger14 fields changed
      • removedInput schema / $schema
        Removed value: -"http://json-schema.org/draft-07/schema#"
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / action / description
        Added value: +"Trigger operation"
      • removedInput schema / properties / action / enum
        Removed value: -[
        -  "register",
        -  "list"
        -]
      • addedInput schema / properties / behavior_name / description
        Added value: +"Behavior tree to activate when condition fires"
      • removedInput schema / properties / condition_params / additionalProperties
        Removed value: -true
      • addedInput schema / properties / condition_params / description
        Added value: +"Condition evaluation parameters"
      • removedInput schema / properties / condition_params / properties
        Removed value: -{}
      • addedInput schema / properties / condition_type / description
        Added value: +"Condition type (e.g. hp_threshold, enemy_proximity)"
      • addedInput schema / properties / cooldown_ms / description
        Added value: +"Minimum cooldown interval between fires"
      • addedInput schema / properties / name / description
        Added value: +"Trigger name"
      • addedInput schema / properties / priority / description
        Added value: +"Preemption priority"
      • addedInput schema / properties / project / description
        Added value: +"Target project slug"
      • addedInput schema / properties / trigger_id / description
        Added value: +"Trigger ID"
    • Changedreplay_recording11 fields changed
      • removedInput schema / $schema
        Removed value: -"http://json-schema.org/draft-07/schema#"
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / action / description
        Added value: +"Recording operation"
      • removedInput schema / properties / action / enum
        Removed value: -[
        -  "capture",
        -  "list"
        -]
      • addedInput schema / properties / execution_id / description
        Added value: +"Target execution ID"
      • addedInput schema / properties / frames / description
        Added value: +"Captured frame sequence"
      • removedInput schema / properties / frames / items / additionalProperties
        Removed value: -true
      • removedInput schema / properties / frames / items / properties
        Removed value: -{}
      • addedInput schema / properties / name / description
        Added value: +"Recording name"
      • addedInput schema / properties / project / description
        Added value: +"Target project slug"
      • addedInput schema / properties / recording_id / description
        Added value: +"Recording ID"
    • Changedset_parameters9 fields changed
      • removedInput schema / $schema
        Removed value: -"http://json-schema.org/draft-07/schema#"
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / action / description
        Added value: +"Parameter operation"
      • removedInput schema / properties / action / enum
        Removed value: -[
        -  "set",
        -  "get",
        -  "reset"
        -]
      • addedInput schema / properties / execution_id / description
        Added value: +"Target execution instance ID"
      • removedInput schema / properties / parameters / additionalProperties
        Removed value: -true
      • addedInput schema / properties / parameters / description
        Added value: +"Key-value parameters map to apply"
      • removedInput schema / properties / parameters / properties
        Removed value: -{}
      • addedInput schema / properties / project / description
        Added value: +"Target project slug"
  4. 10 tool updatesv0.1.2
    • First observedabort_behavior
    • First observedget_metrics
    • First observedget_status
    • First observedload_behavior
    • First observedmanage_behaviors
    • First observedmanage_blackboard
    • First observedmanage_runtime_db
    • First observedregister_trigger
    • First observedreplay_recording
    • First observedset_parameters

TDQS

A4.1/5.0

Scored across 10 tools

Disambiguation4/5

Most tools target distinct runtime concerns, but several pairs (manage_behaviors vs load_behavior, get_metrics vs get_status, abort_behavior vs register_trigger, manage_blackboard vs set_parameters) are close enough that descriptions must explicitly disambiguate them. After reading those descriptions an agent can choose correctly, so boundaries are clear but not self-evident.

Naming Consistency5/5

All names use a consistent snake_case verb_noun pattern: manage_behaviors, load_behavior, get_status, set_parameters, etc. The mixed verbs (manage, get, set, load, abort) are appropriate and used predictably, with no camelCase or stylistic inconsistency.

Tool Count5/5

Ten tools is well-scoped for a behavior runtime server covering definitions, execution, state, telemetry, triggers, recordings, and database maintenance. Each tool has a clear place and the set does not feel bloated or thin.

Completeness4/5

The surface covers most lifecycle needs: behavior definitions, runtime load/unload/swap, status, metrics, blackboard state, parameters, triggers, abort/pause/resume, recordings, and DB maintenance. However, replay_recording exposes only capture/list despite its name, and behavior definitions lack obvious delete/update operations, leaving minor gaps.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    An MCP server that enables AI agents to execute sandboxed JavaScript code to control external web applications in real-time. It utilizes WebSockets to synchronize script execution with frontend animations, allowing complex UI interactions to be handled in a single agent turn.
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to resolve typed routing, loop-stall arbitration, and tool safety checks locally at sub-millisecond latency, avoiding unnecessary frontier LLM calls. It enforces reversible execution checkpoints, token budgets, and decayed episodic memory for safe, privacy-first autonomous operation.
    7
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to combine ambient multimodal perception, continuous episodic memory, and recency-decay context briefing for fast deterministic decisions and safe tool routing. It integrates with MCP clients to reduce unnecessary frontier LLM calls while enforcing token budgets and local-first memory retention.
    7
    MIT