Skip to main content
Glama

blockbench-mcp

An MCP server that gives an AI agent full programmatic control of Blockbench: 3D modeling, procedural textures, UV, animation, codecs/export, files, plugin management — and arbitrary JavaScript inside Blockbench when the built-in tools aren't enough.

Transport: agent ←(stdio/JSON-RPC)→ this server ←(localhost WebSocket)→ a plugin inside Blockbench. Zero dependencies: just Node.js and Blockbench itself.

┌────────────┐   stdio / MCP    ┌──────────────────┐   WebSocket   ┌─────────────────────┐
│  MCP client│ ───────────────► │ blockbench-mcp   │ ◄──────────── │ Blockbench + plugin │
│ (opencode, │ ◄─────────────── │  (this project)  │  127.0.0.1    │ Blockbench MCP Bridge│
│  Claude…)  │   JSON-RPC       │  port + token    │  connection.json             │
└────────────┘                  └──────────────────┘               └─────────────────────┘

Requirements

  • Node.js ≥ 18 (tested on 22).

  • Blockbench Desktop ≥ 4.9 (tested on 5.2.1). The web build won't work — the plugin uses the desktop app's file layer and Node modules.

Related MCP server: BlockbenchMCP

Quick start

From the project folder:

node install.js --register

This will:

  1. Copy the plugin blockbench_mcp.js into Blockbench's plugins folder.

  2. Add it to StateMemory.installed_plugins (Blockbench does not auto-scan the folder) and load it immediately.

  3. If Blockbench isn't running (or is running without a debug port), relaunch it with --remote-debugging-port and inject the plugin over CDP.

Prefer doing it by hand? Copy plugin/build/blockbench_mcp.js into the Blockbench plugins folder and drag the file into the Blockbench window once, confirming the install. --force lets the installer restart a running Blockbench (unsaved changes would be lost).

Then register the server with your client:

node install.js opencode      # or cursor / claude / windsurf / vscode / gemini / cline
node install.js --all         # all of them
node install.js --print-config   # show configs without writing anything

Check everything is in place:

node install.js --status

Open (or restart) your MCP client and ask the agent to call bb_status — it should return the Blockbench version and a project summary.

The agent's workflow (baked into the server instructions)

  1. bb_status — learn the format, mode and contents.

  2. bb_new_project / bb_open_model — create or open a project.

  3. Build: bb_add_cube, bb_add_group, bb_add_mesh. Repeat parts with bb_array_elements, symmetry with bb_mirror_elements (not hand-placed copies).

  4. Textures: bb_create_texture / bb_generate_texture (seeded ops), refine with bb_draw_texture / bb_paint_pixels, assign with bb_set_face_texture, verify with bb_get_texture_pixel.

  5. Look at the result: bb_review renders the model from 4–6 angles into a single contact sheet (the path is returned; the agent reads the image and fixes proportions/textures). bb_set_view + bb_screenshot are available too.

  6. bb_validate → fix findings → bb_export_model (bbmodel, java_block, bedrock, gltf, obj, fbx, stl, collada, skin…).

  7. bb_execute_js for arbitrary JS inside Blockbench; bb_step batches several calls in one round trip.

Tools (69)

Utility

bb_status, bb_execute_js, bb_step (batch; later steps can reference earlier results with "$0.element.uuid")

Project and codecs

bb_new_project, bb_open_model, bb_save_project, bb_export_model, bb_project_info, bb_set_project

Files

bb_read_file, bb_write_file, bb_list_dir, bb_glob, bb_file_info, bb_mkdir, bb_delete_path, bb_request_fs

Reads and writes go through Blockbench's file layer and need no permission. Directory listing, stat, mkdir and delete need a one-time plugin permission for the filesystem — Blockbench asks on first use; click "Always allow for this plugin".

Model

bb_list_elements, bb_add_cube, bb_add_group, bb_add_mesh, bb_edit_mesh, bb_add_element, bb_set_element, bb_transform_elements, bb_array_elements, bb_mirror_elements, bb_duplicate_elements, bb_delete_elements, bb_reparent_elements, bb_select_elements, bb_group_elements, bb_set_face_texture, bb_set_face_uv, bb_auto_uv, bb_validate

Textures

bb_list_textures, bb_create_texture, bb_generate_texture, bb_draw_texture, bb_paint_pixels, bb_get_texture_pixel, bb_import_texture, bb_export_texture, bb_set_texture_properties, bb_resize_texture, bb_delete_texture

Material presets: wood, planks, stone, cobble, metal, dirt, grass, leaves, bricks, fabric, skin, gem, noise, gradient — set with the preset field; your own ops run on top. Drawing ops (ops) support: fill, noise, cells, gradient, radial, rect, circle, ellipse, line, checker, stripes, border, vignette, scatter, pixel, pixels, text, adjust, replace, blend. Everything is deterministic for a given seed.

Animation

bb_list_animations, bb_create_animation, bb_set_animation, bb_delete_animation, bb_add_keyframe, bb_delete_keyframe, bb_play_animation

UI / rendering

bb_set_mode, bb_set_view, bb_run_action, bb_list_actions, bb_notify, bb_screenshot, bb_review

Plugins and settings

bb_list_plugins, bb_install_plugin, bb_uninstall_plugin, bb_reload_plugin, bb_list_settings, bb_set_setting

Host tools (work even with Blockbench closed)

bb_bridge_status, bb_setup, bb_reconnect

Examples

// procedural 64×64 texture
{ "tool": "bb_generate_texture", "arguments": {
  "name": "stone", "width": 64, "height": 64, "seed": 7,
  "ops": [
    { "op": "fill",  "color": "#3a3f46" },
    { "op": "noise", "color": "#22262b", "color2": "#6b7480", "scale": 3, "octaves": 5 },
    { "op": "cells", "cell_size": 10, "color": "#000000", "color2": "#ffffff", "edge": true },
    { "op": "vignette", "strength": 0.4 }
  ] } }
// the same, with a material preset
{ "tool": "bb_generate_texture", "arguments": { "preset": "wood", "name": "wood", "width": 64, "height": 64 } }
// a deterministic row of cubes
{ "tool": "bb_array_elements", "arguments": {
  "targets": ["step-original-uuid"], "axis": "x", "count": 5, "offset": 16, "names": "step_{i}" } }
// batch with a result reference: add a cube, then texture it by its new uuid
{ "tool": "bb_step", "arguments": { "steps": [
  { "tool": "bb_add_cube", "arguments": { "name": "spike", "from": [0,20,0], "size": [4,4,4] } },
  { "tool": "bb_set_face_texture", "arguments": { "targets": ["$0.element.uuid"], "texture": "stone" } }
] } }
// "do anything": raw JavaScript inside Blockbench
{ "tool": "bb_execute_js", "arguments": {
  "code": "return Cube.all.map(c => c.name + ' @ ' + JSON.stringify(c.from));" } }

Sample output lives in examples/: a quadruped mob (toxin_beast) built entirely through this MCP — .bbmodel, a 6-angle review sheet and a hero render. Open the .bbmodel in Blockbench or look at the PNGs to see what comes out of the box.

Security

  • The server listens only on 127.0.0.1 on a random port and accepts connections only with a matching token (regenerated on every start). The endpoint is loopback-only and token-gated.

  • bb_execute_js, bb_run_action and bb_install_plugin can run arbitrary code and install third-party plugins — that is the intended functionality. Don't point this server at a Blockbench instance you don't trust, and don't install plugins from untrusted sources.

Project layout

blockbench-MCP/
├─ plugin/
│  ├─ src/                  plugin sources (concatenated into one file)
│  │  ├─ 00-core.js         helpers, ref resolution, undo, fs, sanitize
│  │  ├─ 10-textures.js     procedural texture engine (seeded ops) + presets
│  │  ├─ 20-tools-core.js   status, execute_js, projects, files, screenshots
│  │  ├─ 30-tools-model.js  elements, transforms, arrays, UV, validate
│  │  ├─ 40-tools-texture.js
│  │  ├─ 50-tools-animation.js
│  │  ├─ 60-tools-plugins.js
│  │  └─ 99-boot.js         WS client, plugin registration, tool registry
│  └─ build/blockbench_mcp.js   built plugin (this is what Blockbench loads)
├─ src/
│  ├─ index.js              MCP server (stdio) + host tools
│  ├─ mcp.js                MCP protocol implementation (zero-dep)
│  ├─ ws.js                 WebSocket server (RFC 6455, zero-dep)
│  ├─ bridge.js             bridge to the plugin + connection.json
│  ├─ setup.js              plugin install + client configs
│  └─ register.js           CDP-based plugin registration
├─ install.js               installer / config CLI
├─ examples/                a model built through the MCP
└─ test/  ws.js · offline.js · live.js · call.js · eval.js

Development

node plugin/build.js     # rebuild the plugin (validates syntax and schemas)
node test/ws.js          # WebSocket server test
node test/offline.js     # MCP protocol without Blockbench
node test/live.js        # end-to-end test in a live Blockbench (must be running)
node test/call.js bb_status       # one-off tool call
node test/eval.js "Plugins.registered"   # evaluate JS in Blockbench over CDP

After editing the plugin: node plugin/build.js, then node install.js --plugin and node install.js --register --no-plugin (hot reload without restarting Blockbench — only if it's running with a debug port).

Troubleshooting

  • bb_bridge_status → connected: false. The plugin isn't loaded or Blockbench is closed. Run node install.js --register and make sure "Blockbench MCP Bridge" is listed among Blockbench's plugins.

  • The plugin file is in the folder but Blockbench doesn't see it. A file in plugins/ alone isn't loaded — it must be recorded in StateMemory.installed_plugins, which is what --register (or dragging the file into the window) does.

  • FS tools ask for permission. That's expected and one-time; or call bb_request_fs.

  • Format/version mismatch. See available formats and codecs in bb_status and bb_project_info.

License: MIT.

Available Tools

72 tools
bb_add_cubeAdd a cubeB

Create a cube element. Provide from+to, or from+size (Blockbench cubes use from/to where to is the exclusive upper bound). origin defaults to from. Optionally assign a texture to all faces, set per-face uv, and enable auto UV.

ParametersJSON Schema
NameRequiredDescriptionDefault
toNo
fromNo
nameNo
sizeNo
colorNo
facesNoPer-face overrides: {north:{uv:[x1,y1,x2,y2], texture, enabled, rotation, tint}, ...}
autouvNo0 off, 1 auto per face, 2 relative
box_uvNo
originNo
parentNo
inflateNo
textureNoTexture name/uuid to assign to every face.
rotationNo
uv_offsetNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses that "to is the exclusive upper bound" and that origin defaults to from, which are non-obvious semantics. However, it says nothing about prerequisites (an open project), which scope the cube is added to, or undo behavior for a mutating operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then parameter guidance. Two efficient sentences with the coordinate-system caveat placed where it matters. Little wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 14-parameter mutation tool with no annotations and no output schema, the description covers the geometry essentials but leaves many parameters undocumented in both schema and text. Adequate but with clear gaps given the complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 21%, so the description must compensate. It explains from, to, size, origin default, and texture/faces/autouv behavior, but omits name, color, box_uv, parent, inflate, rotation, and uv_offset entirely. Partial compensation over a low-coverage schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Create a cube element" states a specific verb and resource, so the agent can distinguish this from mesh/group/element creation siblings. It doesn't explicitly name an alternative (e.g., bb_add_mesh or bb_add_element), so sibling differentiation is implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no alternative named. The only routing-like text is "Provide from+to, or from+size," which is a parameter constraint rather than usage context. An agent gets no help deciding between this and bb_add_element/bb_add_mesh.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_add_elementAdd a special elementC

Create a locator, null_object, bounding_box, texture_mesh, armature or armature_bone. Pass its properties directly.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
typeYes
parentNo
propertiesNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. 'Pass its properties directly' hints that the properties object is passed through unvalidated, but nothing is said about required project/parent context, side effects, undo integration, or return values (no output schema exists either).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the actionable type list front-loaded and no wasted wording. The trailing 'Pass its properties directly' is under-specified rather than verbose, but it is brief enough not to harm readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a creation tool with a nested properties object, no annotations, no output schema, and zero schema-level parameter docs, the description is too thin. An agent still lacks guidance on required context, valid property keys, and what is returned or how the new element is identified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It substantially helps for the required 'type' parameter by enumerating the six valid values (the schema has no enum), but 'name', 'parent', and the nested 'properties' object remain unexplained beyond the vague 'pass its properties directly'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb (Create) and enumerates the specific element types it handles (locator, null_object, bounding_box, texture_mesh, armature, armature_bone), which distinguishes it from siblings like bb_add_cube and bb_add_mesh. The scope is specific, though it doesn't explain how these 'special elements' relate to the mesh/cube creation siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says what can be created but never states when to choose this tool over bb_add_cube, bb_add_mesh, or bb_add_group, nor any prerequisites (e.g. an open project or an existing parent). Usage is only implied by the type list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_add_groupAdd a groupB

Create an empty group (bone) to hold elements or drive animation. Position it with origin.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
originNo
parentNo
rotationNo

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that the created group is empty, but says nothing about undoability, persistence, name uniqueness, parent-child semantics, or what happens when a parent is supplied.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the creation purpose followed by the one positional hint. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter mutation tool with zero annotation coverage, no output schema, and 0% schema descriptions, the definition covers only one parameter and none of the behavioral risks. Substantially incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across four parameters. The description only clarifies origin ('Position it with origin'), leaving name, parent, and rotation entirely undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (create) and resource (empty group/bone) and clarifies the resource's role: holding elements or driving animation. It implicitly distinguishes from bb_group_elements (which groups existing elements), though it never names that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'to hold elements or drive animation,' and it hints that origin sets position, but there is no when-to-use/when-not guidance and no reference to alternatives such as bb_add_element or bb_group_elements.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_add_keyframeAdd a keyframeB

Add a keyframe for a bone on rotation/position/scale. x/y/z are Molang expressions (default "0"), or plain numbers. The bone is referenced by group name or uuid.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
zNo
boneYes
timeYes
channelNo
animationNo
interpolationNolinear

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations at all, the description carries the full burden, and it does add genuinely useful behavioral detail: x/y/z accept Molang expressions (default "0") or plain numbers, and the bone can be addressed by group name or uuid. However it says nothing about whether existing keyframes are overwritten, what happens if no animation is supplied, error conditions, or persistence — significant gaps for a mutation tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with no filler, front-loading the core action. It is efficient, though arguably too sparse for an 8-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter mutation tool with no annotations, no output schema, and 0% schema coverage, the description addresses only about half the parameters and omits how the animation/time context is resolved and whether the write is additive or destructive. It is under-specified for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the semantics of x, y, z (Molang expressions with a default) and bone (group name or uuid) and implies the channel enum values, but leaves time, animation, and interpolation (including its default) completely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (add) and resource (keyframe) plus the target (a bone) and the channels affected (rotation/position/scale), which cleanly separates it from the sibling bb_delete_keyframe. It does not name or contrast any alternative tool, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus bb_delete_keyframe, bb_set_animation, or bb_create_animation, nor any prerequisite such as needing an existing animation. Usage is only inferable from the verb.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_add_meshAdd a meshB

Create a free-form mesh. vertices is a list of [x,y,z] (or a {key:[x,y,z]} map). faces is a list where each face is either [i0,i1,i2] indices into vertices, or {vertices:[i..], uv:[[u,v]..], texture}. UVs are stored per vertex.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
facesYes
originNo
parentNo
textureNo
rotationNo
verticesYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the full burden. It does usefully disclose the internal data formats (per-vertex UVs, keyed vertex maps, face variants), which is real behavioral value. However, it says nothing about edit-mode requirements, permissions, return values, or what happens to any existing mesh.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the core action front-loaded, then format detail. No filler. Slightly dense prose but every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter mutation tool with no annotations and no output schema, the description covers the two critical data-format params well but leaves the remaining five params and all behavioral aspects (mode, side effects, return) unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the schema for vertices/faces is empty ({}), so the description genuinely compensates by fully specifying those two formats — a real contribution. But five parameters (name, origin, parent, texture, rotation) are undocumented in both the schema and the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Create a free-form mesh.' The word 'free-form' implicitly separates it from bb_add_cube, but it never names that or any sibling explicitly, so the distinction is left for the agent to infer.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance at all. With siblings like bb_add_cube, bb_add_element, and bb_edit_mesh available, the description should say when a free-form mesh is preferred, but it offers nothing about context or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_array_elementsArray elementsB

Create repeated copies of elements along a grid — the correct way to build rows of windows, treads, spokes and so on instead of placing copies by hand. names may contain {i} and starts at 1.

ParametersJSON Schema
NameRequiredDescriptionDefault
axisNo
countNoNumber of copies to create.
namesNoName template, e.g. "step_{i}".
offsetNoSpacing along axis.
vectorNo
targetsYes
select_originalsNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full behavioral burden. It discloses the naming/index behavior ('names may contain {i} and starts at 1'), which the schema does not, but omits whether originals are retained, permission prerequisites, reversibility, and what select_originals does for this mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core purpose before the example list. Some minor filler ('and so on') and a slightly awkward trailing clause about names, but overall efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter mutation tool with 43% schema coverage, no annotations, and no output schema, the description leaves a majority of parameters and the create/retain semantics undocumented. It is not complete enough for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 43%, so the description is expected to compensate but largely does not. It clarifies the names template ('{i}', starting at 1), yet axis, vector, targets, and select_originals are unexplained, and 'along a grid' only vaguely implies how axis/offset/vector interact.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create repeated copies of elements') plus the mechanism ('along a grid'), with concrete examples (windows, treads, spokes) that make the intent unmistakable. It edges toward distinguishing itself from sibling duplication tools, but never names the closest alternatives like bb_duplicate_elements.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a real when-to-use scenario ('build rows of windows, treads, spokes ... instead of placing copies by hand'), which is genuine context. However, it never routes the agent against sibling tools (bb_duplicate_elements, bb_mirror_elements, bb_transform_elements) that overlap, and gives no when-not conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_auto_uvRegenerate UVsC

Set automatic UV mode on cubes: mode "faces" (autouv 1, per-face), "relative" (autouv 2), "box" (box_uv), or "off".

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNofaces
targetsYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses the mode mapping but not whether existing UVs are overwritten, whether the operation is undoable (a bb_undo sibling exists), or what happens to targets not eligible for auto-UV. The title 'Regenerate UVs' also conflicts mildly with the description's 'Set ... mode' framing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact sentence with the operation front-loaded and the mode options enumerated with their internal mappings. No filler, though the semicolon-dense format slightly hurts scanability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A mutation tool with no annotations, no output schema, an undocumented required 'targets' object, and no mention of reversibility or side effects. The description covers only the mode vocabulary, not enough for an agent to invoke it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there is no enum on mode, so the description's enumeration of faces/relative/box/off with their underlying autouv/box_uv values is genuinely valuable. However, the required 'targets' parameter is an empty, undocumented object that the description never explains, leaving a critical gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Set automatic UV mode on cubes') that is distinguishable from siblings such as bb_set_face_uv (manual per-face UV) and bb_generate_texture. It does not explicitly name a sibling, which keeps it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No statement of when to use this tool versus bb_set_face_uv or other UV tools, and no prerequisites (e.g. selection state, cube-only scope). The mode list describes options but not when this tool is the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_bridge_statusBridge statusA

Check whether this server is talking to Blockbench, the port it listens on, the connection file path and where the tool schemas came from. Call this first when any bb_* tool reports "not connected".

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses the information surfaced (port, file path, schema provenance) and the word 'Check' implies a non-mutating read, but it never explicitly states that the call has no side effects, requires no authentication, or is safe to repeat.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, no filler. The diagnostic scope comes first and the activation condition is front-loaded at the end, so the agent gets both the 'what' and the 'when' immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does the necessary work of naming the returned fields (port, connection path, schema source). It is close to complete for a zero-param diagnostic, missing only a note on how it differs from bb_status/bb_setup.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing in the schema the description needs to explain or compensate for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific diagnostic verb ('Check whether') and enumerates the resources inspected: connection state, listen port, connection file path, and schema origin. The scope is clear enough to call correctly, though it never distinguishes itself from the closely named sibling bb_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger: 'Call this first when any bb_* tool reports "not connected".' That is a clear when-to-use condition. It does not, however, compare against bb_setup or bb_reconnect, which an agent might reasonably reach for in the same failure scenario.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_create_animationCreate an animationC

Create a new animation. loop is "once", "loop" or "hold"; length is in seconds.

ParametersJSON Schema
NameRequiredDescriptionDefault
loopNoonce
nameNo
lengthNo
selectNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose allowed loop values and that length is in seconds, but it never says whether creation requires an active project, whether the animation is auto-selected, what happens to existing animations, or what the mutation returns. For a write operation with zero annotation coverage this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the purpose front-loaded and no filler. The parameter details are compactly attached, though they are somewhat terse and inline rather than structured, which keeps it just short of ideal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 0% schema description coverage, the description is too thin for a four-parameter mutation tool. It omits what name and select do, whether a project context is required, and what the call returns, so an agent lacks enough information to invoke this confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It usefully documents two of four parameters: loop’s allowed values ('once', 'loop', 'hold') and length’s unit (seconds). But name and select receive no explanation in either the schema or the description, leaving half the parameters semantically opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Create a new animation.' This clearly distinguishes it from siblings like bb_delete_animation, bb_list_animations, bb_play_animation, and bb_set_animation. However, it does not explicitly name or contrast against those alternatives, which is what a 5 would require.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites (e.g., must a project be open?), and no mention of alternatives or conditions for choosing this tool over bb_set_animation or bb_add_keyframe. The description only states what the tool does, leaving usage entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_create_textureCreate a textureC

Create a blank texture, optionally filled with a colour. Size defaults to the project texture resolution.

ParametersJSON Schema
NameRequiredDescriptionDefault
fillNo
nameNo
widthNo
heightNo
selectNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses that the texture is blank by default and that size falls back to the project texture resolution, but it says nothing about the side effect implied by select (default true) potentially changing the current selection, nor about undo behavior or permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action, and every clause carries information the schema does not. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating creation tool with no annotations, no output schema, and 0% parameter coverage, the description is too thin. It omits the meaning of name and select, the effect on selection state, and any indication of what is returned or how to undo.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It covers fill ('optionally filled with a colour') and size/width/height (defaults to project texture resolution), but leaves name and select entirely undocumented, so three of five parameters still have no semantic explanation anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a blank texture') and adds the distinguishing qualifier 'blank', which separates it from siblings like bb_generate_texture, bb_draw_texture, and bb_import_texture. It does not name those siblings, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus bb_generate_texture or bb_draw_texture, and no prerequisites or exclusions are stated. The only usage hint is 'optionally filled with a colour', which is weak routing information.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_delete_animationDelete an animationC

Remove an animation from the project.

ParametersJSON Schema
NameRequiredDescriptionDefault
animationYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It implies a destructive removal but does not say whether the deletion is undoable via bb_undo, what happens to the animation's keyframes, or what occurs if the named animation does not exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the action front-loaded and no filler. It is appropriately sized for a one-parameter tool, though it is arguably too terse to earn a top mark.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive operation with no annotations, no output schema, and an undocumented identifier parameter, the description omits the information an agent needs to call it safely and correctly. Undo behavior and error conditions are unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single required 'animation' parameter is an undocumented string. The description adds nothing about whether this expects an animation name, an index, or an ID, leaving the caller to guess the identifier format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (remove) and resource (animation) and scopes it to 'the project'. However, it does not differentiate itself from the closely related sibling bb_delete_keyframe, so an agent must infer the distinction from the name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus alternatives such as bb_delete_keyframe or bb_set_animation, nor any stated preconditions (e.g. animation must exist, animation must not be playing). The description is a bare statement of effect.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_delete_elementsDelete elementsC

Delete elements (groups delete their children).

ParametersJSON Schema
NameRequiredDescriptionDefault
targetsYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose one important trait: the cascade behavior where deleting a group also deletes its children. However, it says nothing about irreversibility, whether an undo is available, permission requirements, or what happens to references to deleted elements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the destructive scope note is placed immediately after the verb. It is efficient, though terse enough that brevity shades into under-specification.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation tool with no annotations, no output schema, and 0% parameter coverage, the definition is far too thin. An agent cannot determine accepted target formats, whether the operation is reversible, or what the tool returns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single required parameter 'targets' has an empty schema with 0% description coverage (no type, no structure), and the description never mentions it. The parenthetical at best hints that targets may be groups or elements, which is a thin semantic contribution against a completely undocumented parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource ('Delete elements') that distinguishes it from sibling delete tools like bb_delete_path, bb_delete_texture and bb_delete_keyframe. The parenthetical scope note sharpens the operation further. It does not, however, explicitly route the agent away from overlapping tools such as bb_set_element or bb_undo.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as bb_undo or how this differs from other destructive siblings. The agent must infer all usage context from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_delete_keyframeDelete a keyframeB

Delete a keyframe by uuid, or the one nearest to a given time on a bone/channel.

ParametersJSON Schema
NameRequiredDescriptionDefault
boneYes
timeNo
uuidNo
channelNorotation
animationNo

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It implies a deletion but does not state whether the change is reversible, whether undo is available, what permissions are needed, or what side effects occur on the owning animation or channel.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler. It communicates the action and the two target-selection modes efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter deletion tool with no annotations and no output schema, the description is under-specified. It leaves out the animation parameter, safety/reversibility details, and the interaction between uuid and time, which an agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across five parameters. The description mentions uuid, time, bone, and channel, but it omits the animation parameter entirely and does not explain channel default behavior, uuid/time precedence, or expected formats.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: delete a keyframe. It also distinguishes two identification modes, by uuid or by nearest time on a bone/channel, which cleanly separates it from siblings like bb_add_keyframe and bb_delete_animation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains two ways to identify the target keyframe, but it does not state when to use this tool versus alternatives such as bb_delete_animation, nor does it provide prerequisites or exclusions. Usage is only implied by the operation itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_delete_pathDelete a file or directoryA

Delete a file, or a directory when recursive is true. Needs file-system permission.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
recursiveNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden; it usefully discloses the file-system permission requirement, which an agent must know. However, it omits key destructive-tool behavior: whether deletion is permanent/recoverable, what happens on a nonexistent path, and whether an undo path exists despite the sibling bb_undo.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with zero waste; the condition for directory deletion immediately follows the core action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive two-param tool with no annotations and no output schema, the description covers the action, the recursive branch, and the permission prerequisite, but leaves irreversibility, error behavior, and undo/alternative tooling unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains 'recursive' in behavioral terms (directory deletion), but adds no format, path-resolution, or default semantics beyond the schema's own default:false, leaving 'path' undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Delete') and resource ('file', 'directory') and clarifies the branching behavior for directories. It is clearly distinct from file-reading siblings like bb_read_file/bb_list_dir, though it does not explicitly name any sibling it should be chosen over.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies the recursive=true condition for directories, which is useful conditional guidance, but offers no when-to-use/when-not-to-use framing, no references to alternatives (e.g. bb_undo or bb_delete_elements), and no prerequisites beyond permission.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_delete_textureDelete texture(s)C

Remove textures from the project.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetsYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It doesn't disclose whether deletion is destructive/permanent, whether it requires confirmation, how textures in use are handled, or what happens to assignments referencing deleted textures.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence with no waste and the action front-loaded. Concise, though arguably too terse given the gaps elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A destructive mutation tool with no annotations, no output schema, and an undocumented required parameter. The description is far too thin to safely and correctly invoke this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single required parameter 'targets' is completely undocumented in both schema and description. The description adds no meaning about what a target is or what form it takes.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb (remove) and resource (textures) with scope (from the project). Sibling tools like bb_delete_elements, bb_delete_path, and bb_delete_animation make the resource noun necessary to distinguish, and it does that, though it's a fairly generic statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus alternatives like bb_delete_elements or bb_set_texture_properties, and no mention of prerequisites or cautions. The description leaves the agent to infer context entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_draw_textureDraw on an existing textureB

Apply a list of ops onto an existing texture (adds to what is already there). Same ops as bb_generate_texture.

ParametersJSON Schema
NameRequiredDescriptionDefault
opsYes
seedNo
targetYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It usefully discloses that the operation is additive and mutates existing content, but says nothing about permissions, reversibility (bb_undo exists), rate limits, or whether ops can overwrite prior content. This is thin for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly worded sentences with the key scope constraint front-loaded and zero filler. Appropriately sized for the operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-param mutation tool with no annotations and no output schema, the description conveys the high-level behavior and a useful cross-reference, but leaves the seed parameter and the op payload structure unexplained. Adequate minimum but with clear gaps an agent would need to resolve.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It loosely describes 'ops' (a list) and 'target' (an existing texture) but never explains the op object format or the 'seed' parameter at all. The pointer to bb_generate_texture for op format is the only help, leaving two of three params effectively undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (apply ops) and resource (existing texture) and clarifies the additive semantics with '(adds to what is already there)'. It implicitly distinguishes itself from bb_generate_texture, though it does not name that sibling as the 'not this' alternative as explicitly as a 5 would require.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the parenthetical 'adds to what is already there' hints this is for modifying rather than creating, pointing the agent toward bb_generate_texture for new textures. There is no explicit when-to-use/when-not statement or named alternative, so the agent must infer the split.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_duplicate_elementsDuplicate elementsC

Copy elements (groups copy their children too). Optionally repeat and offset each copy.

ParametersJSON Schema
NameRequiredDescriptionDefault
axisNo
countNo
offsetNo
vectorNo
targetsYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden and does disclose one real trait: groups bring their children along. But it says nothing about how axis and vector interact, whether copies are linked/instanced, whether the operation is undoable, or what is returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action and the group caveat. No filler, though it is arguably too sparse for the parameter surface it fronts.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A five-parameter mutation tool with no annotations, no output schema, and zero schema descriptions requires far more from the description than this. The group-copy note is the only substantive detail; axis/vector/targets handling is entirely absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% across five parameters, so the description must compensate and largely does not. 'Repeat and offset' loosely gestures at count/offset, but axis vs vector, the numeric meaning of offset, and the shape of 'targets' are all undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (copy/duplicate) and resource (elements), plus a useful nuance that groups copy their children. However, it does not distinguish itself from closely related siblings like bb_array_elements or bb_mirror_elements, which plausibly cover the same repeat/offset territory.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no alternatives named. The phrase 'optionally repeat and offset each copy' describes capability, not when an agent should pick this tool over bb_array_elements or bb_mirror_elements.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_edit_meshEdit a meshA

Modify an existing mesh in place. With vertices+faces it redefines the geometry; with merge:true the new vertices/faces are appended. delete_faces / delete_vertices remove by key. Keys are shown by bb_list_elements {include_geometry:true}.

ParametersJSON Schema
NameRequiredDescriptionDefault
facesNoArray of indices / {vertices,uv,texture}, or a {key:{...}} map.
mergeNo
targetYes
verticesNoArray of [x,y,z], or a {key:[x,y,z]} map.
delete_facesNo
delete_verticesNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It usefully discloses in-place mutation, the append-vs-overwrite distinction of merge, and key-based deletion. However, it does not say whether redefining geometry discards existing faces not listed, whether faces indices/UVs are preserved, or whether the operation is reversible (bb_undo exists but is not referenced).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences, front-loaded with the core action and mode semantics, with no filler. The key-discovery pointer is efficiently packed into the final sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter mutation tool with no annotations and no output schema, the description covers the mode logic and key lookup well, but leaves 'target' unspecified and omits failure/destructive semantics. Adequate but leaving meaningful inference to the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Despite only 33% schema description coverage, the description supplies meaning for nearly every parameter: vertices, faces, merge, delete_faces, and delete_vertices are all explained in prose. The required 'target' parameter is left unexplained, which is the one clear gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Modify an existing mesh in place', which clearly distinguishes editing from siblings like bb_add_mesh or bb_add_cube. It stops short of explicitly naming the alternative tools, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the two operating modes clearly: vertices+faces redefine geometry, while merge:true appends, and delete_faces/delete_vertices remove by key. It also routes the agent to bb_list_elements {include_geometry:true} to discover keys. No explicit when-not-to-use or sibling comparison, but the mode selection guidance is solid.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_execute_jsExecute JavaScript in BlockbenchA

Run arbitrary async JavaScript inside Blockbench with full access to every global (Project, Cube, Group, Mesh, Texture, Canvas, Codecs, Formats, Modes, Undo, Outliner, Preview, Menu, Action, BarItems, Plugins, StateMemory, ...) and the Node modules a plugin may use. Use this for anything the structured tools do not cover. The code may use await and must return a value; console output is captured and returned. An api argument is injected with helpers: api.find(ref), api.findAll(refs), api.textures(), api.animations(), api.sleep(ms), api.readFile(path), api.writeFile(path,text), api.toBase64(bytes), api.fromBase64(b64), api.toast(text).

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesJavaScript source. May use await. The value of the last expression / an explicit return is returned.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does well: it discloses async execution, full global access, Node module availability, mandatory return values, console capture, and injected api helpers. It still omits details such as sandboxing, error handling, and side-effect persistence.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose and usage before moving into behavioral details and the api helper list. It is appropriately sized for a powerful, complex tool, though the helper enumeration is dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the high-complexity execution tool, no annotations, and no output schema, the description is largely complete: it explains what the code can access, how return values work, and how console output is handled. It could still say more about error behavior or execution constraints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for the single code parameter, so the schema already explains the parameter. The description adds meaningful context beyond the schema by describing the injected api helpers and the return-value behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: run arbitrary async JavaScript inside Blockbench. It explicitly names the tool as the escape hatch for anything the structured tools do not cover, which clearly distinguishes it from its many siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear usage context: use this for anything the structured tools do not cover. This implicitly says when not to use it, but it does not name specific sibling alternatives or provide more detailed routing conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_export_modelExport modelB

Compile the project into another format and write it to disk. Formats: bbmodel/project, java_block, bedrock, bedrock_old, modded_entity, optifine_entity (jem), optifine_part (jpm), collada (dae), fbx, gltf/glb, obj, stl, skin_model.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
formatNo
optionsNoCodec options, passed to compile().

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses that output is written to disk (a mutation) but says nothing about overwrite behavior, required permissions, failure modes, or whether the project state is modified. For a disk-writing tool this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences: the action is front-loaded and the format list is compact and information-dense. Nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No annotations, no output schema, and only partial parameter documentation for a tool that writes files to disk. An agent still lacks overwrite semantics, error behavior, and whether a return value (e.g., written path) exists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% — only `options` is documented. The description partially compensates by enumerating valid `format` values (which are not constrained by an enum in the schema), but `path` semantics (extension, relative vs absolute, overwrite) are left entirely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (compile/write) and resource (the project/model) and names the destination (disk). The enumerated format list makes the tool's output scope unambiguous and distinguishes it from bb_save_project and bb_export_texture.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The format list implies when this tool applies (converting the project to an external representation), but the description never says when to prefer it over bb_save_project or what preconditions exist. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_export_textureExport texture(s) to PNGB

Write texture PNGs to disk. Give a single target + path, or several targets + a directory.

ParametersJSON Schema
NameRequiredDescriptionDefault
dirNo
pathNo
targetsNo
include_data_urlNo

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses that it writes files to disk, but omits whether existing files are overwritten, whether directories are created, and what permissions or active-project state are required. For a filesystem-writing tool this is a notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the effect (writes PNGs) followed immediately by the two invocation patterns. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, no annotations, and 0% parameter coverage mean the description is the only source of truth. It covers the main flow but leaves include_data_url, overwrite behavior, and error conditions unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains the target/path/dir relationship well, but ignores the include_data_url parameter entirely, leaving one of four parameters with no semantics anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: writes texture PNGs to disk. Clearly distinguishable from siblings like bb_export_model and bb_save_project, though it doesn't name them. The purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains the two invocation modes (single target+path vs. several targets+dir), which is implied usage guidance. However, it never states when to choose this over alternatives like bb_save_project or bb_export_model, nor any prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_file_infoFile infoB

Stat a path: exists, type, size, modified time. Needs file-system permission.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses the permission requirement and the shape of the result, but says nothing about failure modes (e.g. whether a nonexistent path errors or simply reports exists=false) or whether it follows symlinks – gaps that matter for a stat operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short clauses, front-loaded with the action and its outputs, with no filler. It is terse rather than verbose, which is appropriate but leaves no room for the missing usage detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read-only tool with no output schema, disclosing the returned fields inline is a reasonable substitute. Still, no error semantics and no path-format guidance leave it only minimally complete for an agent deciding between it and bb_read_file/bb_list_dir.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%: the single "path" parameter has no description in the schema. The description only says "a path," without clarifying whether it expects an absolute path, a path relative to the project root, or a file vs directory, and the sibling tools mix project-scoped and filesystem concepts, so this ambiguity is real.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Stat a path" gives a specific verb and resource, and the enumeration of returns (exists, type, size, modified time) marks it clearly as metadata-only, distinguishing it in practice from bb_read_file and bb_list_dir. It never names those siblings explicitly, so the differentiation is inferential rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description supplies one useful precondition ("Needs file-system permission"), which tells the agent this may fail without a grant, but it gives no when-to-use guidance relative to bb_read_file, bb_list_dir, or bb_request_fs. Usage is implied by the verb "stat" rather than explained.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_generate_textureGenerate a procedural textureB

Paint a new texture from a preset and/or an ordered list of ops. Deterministic for a given seed. Presets: wood, planks, stone, cobble, metal, dirt, grass, leaves, bricks, fabric, skin, gem, noise, gradient. Ops: fill, noise, cells, gradient, radial, rect, circle, ellipse, line, checker, stripes, border, vignette, scatter, pixel, pixels, text, adjust, replace, blend. Example: {preset:"wood"} or {ops:[{op:"fill",color:"#2b3a55"},{op:"noise",color:"#101820",color2:"#8fb8de",scale:2,octaves:4},{op:"vignette",strength:0.5}]}.

ParametersJSON Schema
NameRequiredDescriptionDefault
opsNo
nameNo
seedNo
widthNo
heightNo
presetNoA named material preset; user ops run on top of it.
selectNo
include_data_urlNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses that output is deterministic for a given seed and that ops are an ordered list, but it does not explain side effects such as whether a new texture is added to the project, selected, or returned, nor does it cover permissions or reversibility.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose, then determinism, presets, ops, and an example. The long preset and op lists consume space but are necessary reference material rather than fluff. The structure is dense but efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 8 parameters, no annotations, no output schema, and low schema description coverage, the description supplies the core procedural recipe syntax. However, it omits return behavior, side effects, and meanings for several parameters, leaving gaps an agent must infer before calling the tool reliably.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 13%, so the description must compensate, and it does so for the most complex parameters. It lists all preset names, enumerates op names, and gives a concrete example showing op payload fields such as color, color2, scale, octaves, and strength. It still leaves width, height, name, select, and include_data_url largely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: "Paint a new texture from a preset and/or an ordered list of ops." This clearly communicates procedural texture generation. It does not explicitly differentiate itself from sibling tools such as bb_create_texture or bb_draw_texture, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an example invocation, which shows how to use the tool, but it never says when to choose this tool over alternatives like bb_create_texture, bb_draw_texture, or bb_paint_pixels. There is no when-to-use, when-not-to-use, or alternative-routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_get_texture_pixelRead texture pixelsB

Read a rectangle of pixels from a texture. Use this to verify what was drawn: it returns rows of [r,g,b,a].

ParametersJSON Schema
NameRequiredDescriptionDefault
hNo
wNo
xNo
yNo
targetsYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses the return shape ('rows of [r,g,b,a]'), which is real behavioral value, but omits coordinate origin/convention, out-of-bounds behavior, and texture-target semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero waste; the core action is front-loaded and the return format is attached where it matters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and while the return format is described, the five undocumented parameters (especially the required, untyped 'targets') and the absence of any annotation coverage leave significant gaps for calling this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across five parameters, and the description does not compensate: x, y, w, h are only implied by 'rectangle', and the required 'targets' parameter is never explained at all. An agent cannot determine valid values for targets from this text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Read a rectangle of pixels from a texture') and adds a downstream purpose ('verify what was drawn'). It is distinguishable from siblings like bb_draw_texture and bb_paint_pixels, though it never names them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides one implied usage context ('verify what was drawn'), which is better than nothing, but gives no when-not conditions, no prerequisites, and no mention of alternative read paths (e.g. bb_screenshot).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_globFind files by globA

Recursively find files under a directory matching a glob such as "**/.json" or ".png". Needs file-system permission.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
limitNo
patternYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses recursion and the file-system permission requirement, but says nothing about the default limit of 200, truncation/pagination behavior, symlink handling, or what happens when nothing matches.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences; the core action is front-loaded and the permission caveat is a short trailing clause. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter tool with no output schema and no annotations, the description covers the essentials of what and how, but omits result-limiting behavior and error/empty-result semantics that an agent would need to call it reliably.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies 'path' as a directory and 'pattern' as a glob with examples, but leaves the 'limit' parameter (default 200) entirely undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('recursively find files') plus the matching semantics, with concrete glob examples ("**/*.json", "*.png"). It is distinguishable from bb_list_dir and bb_read_file, though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The glob examples imply the intended use case (pattern-based discovery vs. simple listing), and the permission prerequisite is stated. However, there is no explicit when-to-use guidance or comparison against bb_list_dir/bb_file_info for the same directory.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_group_elementsGroup elementsB

Create a new group around the given elements, positioned at their centre, and move the elements into it.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
targetsYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose real behavioral traits: it creates a group, positions it at the elements' centre, and moves the elements into it, indicating a mutation with positioning side effects. However, it omits permission/auth requirements, what happens to pre-existing groups, and error behavior, so it is only partially complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single well-formed sentence with the core action front-loaded and zero filler; every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, no output schema, and 0% parameter coverage, the description explains the core operation adequately but leaves the 'name' parameter and error/edge-case behavior unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description does not mention the 'name' parameter at all — it is only faintly implied by 'Create a new group'. 'targets' is gestured at via 'the given elements', but neither parameter is meaningfully documented beyond what the schema already shows.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: creates a new group around given elements and moves them into it. It implicitly differentiates from bb_add_group (which likely adds to an existing group), but never names that sibling, so the differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative guidance. The agent must infer from the description alone that this is for creating a fresh group versus bb_add_group, with no explicit routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_import_textureImport a textureC

Load an image from disk or a data URL as a project texture.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
pathNo
data_urlNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It implies a data-loading mutation but does not state whether the originating file must exist, what happens if both 'path' and 'data_url' are provided, the supported image formats, or where the resulting texture is placed. This is a significant gap for an unannotated tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, efficient sentence with no wasted words. It is front-loaded with the core action and format, though it is arguably under-specified rather than concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter tool with 0% schema coverage, no annotations, and no output schema, the description is incomplete. It omits parameter behavior, requiredness (0 required means the agent must guess), error conditions, and how the imported texture interacts with siblings, leaving the agent inadequately equipped to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate for all three parameters. It only vaguely hints at 'disk or a data URL', mapping to 'path' and 'data_url', but does not explain the 'name' parameter or the mutual exclusivity/priority between path and data_url. None of the parameters are meaningfully documented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('load an image as a project texture'), which is clearer than the tautological title 'Import a texture'. However, it doesn't distinguish this from sibling tools like bb_create_texture or bb_generate_texture, leaving the agent to infer where imported images differ from created ones.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance or mention of alternatives. The agent is not told how this differs from bb_create_texture, bb_generate_texture, or bb_draw_texture, all of which appear to produce textures. No prerequisites or exclusions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_install_pluginInstall a pluginA

Install a Blockbench plugin from source code, a local file, or a URL. The code is written to the plugins folder, recorded so it loads on startup, and loaded immediately. Executes third-party code inside Blockbench.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoPlugin id (required unless it can be read from the code).
urlNoURL of a .js plugin file.
codeNoFull plugin JavaScript source.
pathNoLocal .js file to install.
versionNo1.0.0

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does disclose important side effects: the code is written to the plugins folder, recorded for startup loading, loaded immediately, and executes third-party code inside Blockbench. It does not cover overwrite behavior, permissions, or error handling, but the core mutation behavior is clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly written sentences with no filler. The primary purpose is front-loaded, followed by persistence and execution behavior, and the security-relevant note is clearly placed at the end.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-annotation, no-output-schema install tool, the description gives enough context to understand the main side effects and security implication. It could be more complete about overwrite behavior and prerequisites, but it covers the essential install semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80%, so most parameters are already documented in the schema. The description reinforces the code/path/url source distinction but adds little beyond the schema for id and version.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: install a Blockbench plugin. It also distinguishes the supported source forms (source code, local file, URL), which separates it from siblings like bb_uninstall_plugin, bb_reload_plugin, and bb_list_plugins.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool by naming install sources, but it does not explicitly say when to prefer it over bb_execute_js, bb_reload_plugin, or bb_uninstall_plugin, nor does it state prerequisites or exclusions. Usage is inferable but not well guided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_list_actionsList toolbar actionsC

List runnable Blockbench actions (BarItems) by id and name, optionally filtered.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
filterNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It confirms the listing returns id and name, but says nothing about the limit parameter's truncation behavior (default 60 implies a cap that isn't explained), result ordering, or whether the list is live from the editor state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the resource and scope front-loaded and zero filler. It is appropriately short for a simple list tool, though brevity here comes at the cost of the missing details noted above.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity, zero-required-parameter read tool with no output schema, the description is minimally adequate: it says what is listed and what fields come back. The unexplained limit semantics and absent tie-in to bb_run_action leave real gaps for a 0%-coverage schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and neither parameter is described in the schema. The description mentions filtering but not whether 'filter' matches id, name, or both, and 'limit' is never referenced at all despite its 60 default suggesting truncation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List runnable Blockbench actions (BarItems) by id and name'), and clarifies what is returned. It does not name the obvious sibling bb_run_action, so the agent must infer that this enumerates what that tool executes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Only 'optionally filtered' hints at usage; there is no statement of when to call this versus bb_run_action or any prerequisite/context. An agent can infer the discovery-then-execute pattern, but the description never says it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_list_animationsList animationsA

List the project animations with loop mode, length and per-bone keyframe counts.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It helpfully discloses what the output contains (loop mode, length, per-bone keyframe counts) and implies a safe read, but says nothing about permissions, scoping to the active project, or result size/pagination for large rigs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the returned fields are packed in efficiently and every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple parameterless list tool with no output schema, describing the returned fields (loop mode, length, per-bone keyframe counts) covers most of what an agent needs. Only scoping and pagination caveats are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. The description correctly implies it operates on the whole current project with no filtering arguments.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('List the project animations') and even previews the returned fields, which distinguishes it from mutation siblings like bb_create_animation, bb_set_animation and bb_play_animation. It does not explicitly name those siblings, but the read-only list intent is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no mention of alternatives (e.g., use bb_play_animation to preview instead of list, or bb_add_keyframe to modify), and no prerequisites. The agent is left to infer that this is a read-only inventory call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_list_dirList a directoryA

List the entries of a directory. Needs the plugin file-system permission (the user is asked once; bb_request_fs can trigger the prompt explicitly). Supports recursion and a glob filter.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
limitNo
filterNo
recursiveNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose a key behavioral prerequisite: the plugin file-system permission, the one-time user prompt, and the role of bb_request_fs. However, it omits other behavioral traits such as read-only nature, default limit behavior, output shape, and error handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three short, front-loaded sentences with no filler. The core action is first, followed by the permission prerequisite and the supported options.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and 0% schema description coverage, the description is only partially complete. It addresses the permission requirement and two optional features, but does not explain the limit parameter, return value structure, or edge-case behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain all four parameters. It covers 'recursive' and 'filter' (as a glob filter) but leaves 'path' and especially 'limit' without additional semantic detail beyond their names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('List the entries of a directory') and adds two capability hints (recursion and glob filter). It is clear, but it does not explicitly differentiate itself from siblings such as bb_glob or bb_file_info.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by mentioning the file-system permission requirement and that bb_request_fs can trigger the prompt. It does not say when to use this tool instead of alternatives like bb_glob or bb_file_info, nor does it state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_list_elementsList elementsA

List outliner elements with their ids, transforms and (optionally) geometry. Filter by type or parent.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNocube, mesh, group, locator, null_object, bounding_box, texture_mesh, armature, armature_bone
parentNo
include_facesNo
include_geometryNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden; it does disclose the return payload (ids, transforms, optional geometry) which is genuinely useful. However it does not state that this is a read-only/non-mutating operation, nor anything about result size or limits, leaving the safety profile implied by the word "List".

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the resource and return contents come first and the filter capability last. It is efficient, though the terseness leaves some ambiguity that a slightly longer treatment could resolve.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 4 params, no annotations, and no output schema, the description covers the core operation adequately but leaves include_faces undefined and says nothing about the expected shape or volume of results. For an unpaginated list tool this is a minor but real gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25%, so the description needs to compensate. It maps reasonably to type, parent, and include_geometry ("optionally geometry"), but include_faces is never referenced and its distinction from include_geometry is left unaddressed, and the value format for parent is unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (outliner elements) plus what is returned (ids, transforms, optionally geometry). It clearly distinguishes itself from adjacent list tools like bb_list_textures or bb_list_animations, which target different resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Filter by type or parent" implies the intended usage but does not say when to prefer this over alternatives such as bb_select_elements or bb_glob, nor does it state that no filtering returns everything. Usable guidance, but no explicit when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_list_pluginsList pluginsA

List every Blockbench plugin that is present, with id, version, source, install state and whether it can be reloaded.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden. It discloses the returned fields (id, version, source, install state, reloadability) but never states that this is a side-effect-free read with no mutations, nor any ordering or scope limits. For a zero-param query tool the residual risk is low, so a 3 is fair rather than punitive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the verb and resource, with the return attributes appended compactly. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must describe the result itself — and it enumerates exactly which fields come back. For a zero-parameter listing tool this is sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline of 4 applies. The description correctly implies the tool operates on all plugins with no filtering arguments.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('Blockbench plugin') and even enumerates the attributes returned. Sibling write operations (bb_install_plugin, bb_uninstall_plugin, bb_reload_plugin) are clearly distinguished by the enumeration-only verb.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: an agent infers this is the discovery step before install/reload/uninstall. The mention of 'whether it can be reloaded' hints at the reload workflow but no explicit when-to-use, prerequisites, or exclusions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_list_settingsList Blockbench settingsA

List user settings by id with their current values. Handy before bb_set_setting. filter matches the id.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
filterNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that values returned are the 'current values' and that filter matches ids, which is useful, but says nothing about pagination/limit behavior, result ordering, or the read-only safety profile an agent would want confirmed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the purpose, followed by a usage hint and one parameter note. No filler; each clause contributes something, though it is terse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter list tool with no annotations or output schema, the description gives purpose, a usage cue, and partial filter semantics. It still omits limit behavior and any hint about the returned shape beyond 'current values', leaving gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains that filter matches the id (adding real meaning), but the limit parameter is left completely undocumented in both the description and the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (user settings with their current values), which distinguishes it from the many other bb_list_* siblings and from bb_set_setting. The phrase 'by id' is slightly confusing since neither declared parameter is an id, but the core purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Handy before bb_set_setting' gives a concrete usage context and points at the complementary write tool. It lacks any when-not guidance or explanation of how to use limit, so it is clear but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_list_texturesList texturesB

List the project textures with size, UV resolution, layers and saved state.

ParametersJSON Schema
NameRequiredDescriptionDefault
include_data_urlNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral burden. It usefully discloses what each texture record contains (size, UV resolution, layers, saved state), which establishes this as a non-destructive read, but it says nothing about the cost or side effects of requesting embedded data, or about result size/pagination.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence with no filler, and the resource is stated before the payload detail. It is appropriately sized for a simple list tool, though the space spent enumerating return fields arguably belonged to the parameter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter, no-output-schema listing tool the description is nearly sufficient — it conveys the record shape — but with no annotations and no output schema the agent still cannot tell what include_data_url does or how large the response can get.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the sole parameter, include_data_url, is not mentioned anywhere in the description even though it materially changes the response (embedding full image data). With only one optional boolean, the omission is the single biggest gap in the definition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('List the project textures') and even enumerates the fields returned (size, UV resolution, layers, saved state). It is clearly distinguishable from mutation siblings like bb_create_texture or bb_delete_texture, but it never contrasts itself with the other read-oriented listers (bb_list_elements, bb_list_animations), so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance and no named alternative. The agent can infer it is a read-only listing from the verb, but nothing tells it when this is preferable to bb_project_info or a targeted query.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_mirror_elementsMirror elementsA

Reflect elements across an axis-aligned plane (default x=0). copy=true duplicates first, which is the usual way to build a symmetrical half. Meshes keep correct winding.

ParametersJSON Schema
NameRequiredDescriptionDefault
axisNox
copyNo
planeNo
targetsYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden. It usefully notes that copy=true duplicates first and that 'meshes keep correct winding', but says nothing about whether copy=false mutates existing geometry in place, whether the operation is undoable, or how targets must be selected.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place, with the core action front-loaded ahead of the copy-mode caveat and the winding note. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter mutation tool with no annotations and no output schema, the description covers the geometric behavior well but leaves the required targets parameter and the mutation/undo semantics unaddressed. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it partially does: it clarifies the x-axis/plane=0 defaults and the semantics of copy. However the required 'targets' parameter is entirely undocumented (untyped in the schema, unexplained in prose).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource: 'Reflect elements across an axis-aligned plane (default x=0)'. The reflection action is inherently distinct from neighbors like bb_transform_elements, bb_duplicate_elements and bb_array_elements, so an agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'copy=true ... is the usual way to build a symmetrical half' gives a useful usage cue for one parameter, but the description never states when to prefer this tool over bb_array_elements or bb_duplicate_elements, nor any prerequisites. Usage is implied rather than routed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_mkdirCreate a directoryB

Create a directory (recursively by default). Needs file-system permission.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
recursiveNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses the default recursion behavior and the permission requirement, but omits what happens if the directory already exists and how failures surface.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, with the core action front-loaded and the recursion default immediately following. No filler; nothing to trim for a tool this simple.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter mutation with no output schema and no annotations, the essentials (action, recursion default, permission need) are present, but existing-directory behavior and error semantics are missing, leaving some uncertainty for the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 2 parameters. The description clarifies the 'recursive' parameter via 'recursively by default' and its default value, but 'path' is left entirely to the schema; partial compensation only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a directory') and adds the recursive-by-default scope, which distinguishes it from path-listing siblings like bb_list_dir and from bb_delete_path. It does not explicitly name alternatives, but the operation is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides one prerequisite ('Needs file-system permission') but gives no guidance on when to use this over siblings such as bb_write_file (which may implicitly create parents) or bb_request_fs for permissions. Usage is implied by the tool name rather than explained.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_new_projectNew projectA

Create a new empty Blockbench project. format is a Blockbench format id (free, java_block, bedrock, bedrock_block, modded_entity, skin, image, ...). Always call this (or bb_open_model) before building, unless a project is already open.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
box_uvNo
formatNofree
texture_widthNo
create_textureNoAlso create a blank texture sized texture_width x texture_height.
texture_heightNo

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses the empty-state nature and the required ordering, but does not say what happens if a project is already open (replace vs. error) or what the call returns. That gap is significant for a state-creating tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action, then the format enumeration, then the sequencing constraint. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter, no-annotation, no-output-schema tool, the description covers format and sequencing well but omits behavior on an already-open project and most parameter meanings. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 17%, so the description must compensate. It does explain the most important parameter, format, with concrete id values, but leaves name, box_uv, texture_width, and texture_height entirely undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Create a new empty Blockbench project") and implicitly distinguishes itself from bb_open_model by naming it as the alternate entry point. An agent can tell it apart from sibling creation tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly gives the when-to-use rule ("Always call this ... before building, unless a project is already open") and names the alternative (bb_open_model). Nothing about sequencing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_notifyNotify the userB

Show a toast / quick message inside Blockbench. Use it to tell the human what the agent is doing or to ask them to look at the screen.

ParametersJSON Schema
NameRequiredDescriptionDefault
iconNolink
textYes
expireNo

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It implies a transient, non-blocking UI display ('toast / quick message') but says nothing about whether the call waits for the user, what it returns, or how it behaves if Blockbench is unfocused. Key behavioral traits are left to inference.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with what the tool does before how to use it. No filler, though the second sentence could be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple three-parameter UI tool with no output schema, the description adequately conveys purpose and usage. However, in the absence of annotations and with zero parameter documentation, an agent still lacks the details needed to use icon/expire correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Three parameters with 0% schema description coverage: 'icon', 'expire', and 'text' are undocumented. 'Text' is inferable from 'toast / quick message', and 'quick' faintly hints at the auto-dismiss behavior of 'expire', but icon type and duration units/semantics are not addressed at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Show a toast / quick message inside Blockbench'), which is unambiguous. It does not explicitly contrast itself with siblings, but no sibling tool overlaps with user-facing notifications, so the intent is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives positive guidance on when to call it: to tell the human what the agent is doing, or to ask them to look at the screen. No when-not or alternatives are offered, but the use case is narrow enough that none are needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_open_modelOpen / import a model fileA

Load a model from disk, replacing the current project (or importing into it). Codec inferred from the extension (.bbmodel, .json, .obj, .gltf, .fbx, .stl, .dae, .jem, .jpm), or set with format (ids from bb_status).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
formatNoOptional codec id / alias, e.g. java_block, bedrock, project.
import_to_currentNoMerge into the open project instead of replacing it.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations at all, the description carries the full burden, and it does disclose the key behavioral trait: operation replaces the current project unless importing. It also explains codec inference from the extension list and points to bb_status for format ids. It stops short of stating failure behavior for unknown codecs or whether unsaved changes are preserved or discarded.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with the destructive effect front-loaded before the codec detail. No filler, no restating of the name or title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema and no annotations, so the description must stand alone, and for a 3-parameter tool it covers load semantics, replace/merge behavior, and codec resolution adequately. It lacks any statement of what is returned or how errors (bad path, unsupported codec) surface.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, and the description compensates: it enumerates the recognized extensions that drive inference for path, points to bb_status as the source of valid format ids, and restates the merge-vs-replace meaning of import_to_current. Format remains an open string with no enum, so ids are still not resolvable from this definition alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Load a model from disk") plus the scope of the effect (replacing the current project or importing into it), which separates it from bb_new_project and bb_save_project. It does not name any sibling explicitly, so sibling differentiation is inferential rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear selection context: replace by default, or set import_to_current to merge into the open project. It also explains how the codec is chosen (extension inference vs. the format parameter), but offers no exclusions or prerequisites, such as what happens to unsaved work or unsupported extensions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_paint_pixelsPaint pixelsB

Set individual pixels on a texture — the precise way to hand-draw pixel art. pixels is a list of [x, y, color] with color "#rrggbb" or [r,g,b] or [r,g,b,a].

ParametersJSON Schema
NameRequiredDescriptionDefault
pixelsYes
targetYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, yet it only describes the pixel data shape. It says nothing about prerequisites (must the target texture already exist?), reversibility/undo, whether existing pixels are overwritten, or performance on large lists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, purpose front-loaded before the data-format detail. No waste, though it could have surfaced the format as structured detail rather than a trailing clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter mutation tool with no annotations and no output schema, the description covers the core action and one parameter well, but omits target semantics, side effects, and undo behavior, leaving meaningful gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It thoroughly documents the 'pixels' entry format ([x, y, color] with hex/rgb/rgba variants), which is genuinely useful, but leaves 'target' entirely unexplained (texture name vs id vs handle). Partial compensation warrants a middling score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Set individual pixels on a texture') and frames itself as 'the precise way to hand-draw pixel art', which hints at its distinction from coarser texture tools like bb_draw_texture. It does not name a sibling explicitly, so it stops short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'the precise way to hand-draw pixel art' implies when this is preferable to bulk operations, but there is no explicit when-to-use, when-not-to-use, or named alternative. Usage is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_play_animationTimeline playbackC

Control the animation timeline: set the time, play, pause or stop.

ParametersJSON Schema
NameRequiredDescriptionDefault
timeNo
actionNoset_time
animationNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and falls short: it never says whether playback is blocking or asynchronous, whether it mutates project state permanently, which parameter is required for which action, or what happens on error. The 'action' string has no enum, so valid values (only 'set_time' is visible via the schema default) are left to guesswork.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the resource and operations front-loaded and no filler. It is efficient, though its brevity contributes to the specification gaps elsewhere rather than being purely a virtue.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No annotations, no output schema, and 0% schema coverage leave the agent without the essentials for a state-mutating timeline tool: valid action values, which params each action needs, default vs explicit animation targeting, and return/async behavior are all absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description only loosely maps to the params: 'set the time' hints at 'time', and the verbs hint at 'action' values. The 'animation' parameter is entirely unexplained — no indication whether it is a name, ID, or defaults to the active animation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('control') and resource ('the animation timeline') plus the concrete operations (set time, play, pause, stop). It is clearly distinct from create/set/delete_animation siblings, though it never explicitly contrasts itself with the adjacent bb_step or bb_list_animations tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It enumerates the operations but gives no when-to-use context, no prerequisites, and no routing away from alternatives like bb_step or bb_add_keyframe. The agent must infer from the tool name when this is the right call versus manipulating the animation state directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_project_infoProject overviewA

Dump the current project: format, mode, sizes, and optionally the element tree, textures and animations. Call this to re-orient after a large edit.

ParametersJSON Schema
NameRequiredDescriptionDefault
elementsNo
geometryNoInclude mesh vertices and faces (verbose).
texturesNo
animationsNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses one useful trait (geometry is 'verbose') and implies a read-only, non-mutating dump, but says nothing about cost, output size, permissions, or side effects of a large dump.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences with the payload scope front-loaded and the usage cue second. No filler or repetition of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless-required, read-only inspection tool with no output schema, the description covers what the tool returns and how to toggle it. Missing only the default-on behavior of elements/textures/animations and any note on output volume for large projects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (just geometry), but the description compensates by enumerating the three optional payload families — element tree, textures, animations — which map onto the elements/textures/animations booleans. It adds real meaning beyond the bare boolean names, though it never mentions that all three default to true.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a concrete verb and resource plus the exact payload categories (format, mode, sizes, elements, textures, animations), so an agent knows what comes back. It does not explicitly distinguish itself from near-neighbors like bb_status or bb_list_elements, which is the only thing keeping it off a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Call this to re-orient after a large edit' gives one positive usage trigger. There is no statement of when NOT to use it and no named alternative (e.g., bb_list_elements for elements alone, bb_status for connection state), so the routing guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_read_fileRead a fileA

Read a file from disk. encoding "auto" (default) returns text for text files and base64 for binary; force it with "text" or "base64". Reading uses Blockbench's own file layer and needs no plugin permission.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
encodingNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose non-obvious behavior: encoding 'auto' auto-detects text vs binary, the operation uses Blockbench's own file layer, and it requires no plugin permission. It stops short of covering failure modes (missing file, path resolution, size limits), which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, zero filler, and the encoding behavior — the one non-obvious detail — is front-loaded after the one-line purpose. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read tool with no annotations and no output schema, the description covers the encoding dimension well but leaves gaps: what path is relative to, what happens on failure, and how large files are handled. Adequate but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does: encoding's default ('auto') and the meaning of each enum value ('text'/'base64') are explained, which is more than the bare enum list in the schema provides. The path parameter is self-explanatory and needs no elaboration.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Read a file from disk'), and the read direction clearly separates it from bb_write_file and bb_delete_path. It does not explicitly contrast with other file-adjacent siblings like bb_file_info or bb_glob, but the core purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no exclusions, and no mention of alternatives such as bb_file_info (metadata) or bb_glob (discovery). The agent must infer that this is the tool for retrieving file contents.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_reconnectReconnect the bridgeA

Drop the current Blockbench connection so the plugin dials back in. Use after reloading the plugin in Blockbench.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose the core behavior (drops the connection, plugin reconnects), which is real value, but says nothing about whether the drop is immediate, whether in-flight requests are aborted, or what happens if no plugin is listening.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero waste, with the primary action front-loaded and the usage condition following. Nothing to trim.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, no-output-schema maintenance tool, the description covers what it does and when to invoke it. Only minor edge-case behavior (failure mode when no plugin is connected) is unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies; there is nothing for the description to disambiguate beyond what the empty schema already conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action ('drop the current Blockbench connection') plus the resulting behavior ('the plugin dials back in'), which distinguishes it from sibling bb_bridge_status (a read) and bb_reload_plugin (a different remedy). An agent can tell exactly what this tool does without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly gives the triggering condition: 'Use after reloading the plugin in Blockbench.' That is clear when-to-use guidance, though it names no when-not case or chalk alternative (e.g., bb_reload_plugin) to disambiguate further.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_redoRedoC

Redo the last undone edit.

ParametersJSON Schema
NameRequiredDescriptionDefault
stepsNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it discloses almost nothing behavioral. It does not say whether this mutates the document, whether the redo is itself reversible, what occurs when nothing is undone, or whether history state is affected beyond the edit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short, front-loaded sentence with no wasted words, but brevity here comes at the cost of under-specification rather than genuine information density. It is appropriately sized for what it says, not for what an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and an undocumented parameter, the definition is too thin for a state-mutating history tool. An agent cannot know the preconditions, the effect of a failed redo, or the meaning of steps from this description alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single 'steps' parameter has no description anywhere. The phrase 'the last undone edit' actually implies a single step, giving no hint that multi-step redo is supported, so the description adds little and slightly misleads on the default behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (redo) and resource (last undone edit), which an agent can distinguish from bb_undo by the undo/redo pairing. However, it never names the sibling it mirrors or clarifies that it applies to the editor's history stack rather than some other edit context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied by the undo/redo convention; there is no explicit statement of when redo is applicable (e.g. only after an undo) or what happens when the redo stack is empty. No alternatives or exclusions are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_reload_pluginReload a pluginA

Reload a dev/URL plugin without restarting Blockbench. Store plugins are not reloadable in place.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses that no host restart is required and that store plugins cannot be reloaded in place, but says nothing about whether reload discards plugin state, what happens on failure, or any permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero waste; the core action is front-loaded and the caveat follows. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with no output schema, the description covers the key behavioral caveat but leaves the required argument's meaning unexplained. Adequate but with a clear gap on the only parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the description never mentions the single 'id' parameter, so it does not clarify whether it expects a plugin id, a file path, or a URL. Given low coverage, the description should have compensated but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (reload a plugin) and immediately scopes it to dev/URL plugins, explicitly excluding store plugins. An agent can distinguish this from bb_install_plugin and bb_uninstall_plugin without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear when-not (store plugins are not reloadable in place) and implies the use case of iterating on a dev/URL plugin without restarting Blockbench. It does not name a sibling alternative, but the exclusion rule is the operative guidance here.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_reparent_elementsReparent elementsC

Move elements under a new parent group (or "root").

ParametersJSON Schema
NameRequiredDescriptionDefault
parentYes
targetsYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: whether existing transforms are preserved, what happens to children of the moved elements, how conflicts with an existing parent are resolved, or what errors occur. "Move" clarifies mutation intent but no side effects or preconditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short, front-loaded sentence with no filler; the action and destination are stated immediately. It is under-specified rather than bloated, so conciseness is fine but it is arguably too terse for a required-parameter mutation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A mutation tool with two required, wholly undocumented parameters, no annotations, and no output schema should do more work in the description. Callers cannot determine target format, parent reference format, or the effect on the hierarchy from the text provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for both required parameters. The description partially clarifies "parent" by indicating it accepts a group or the literal "root", but says nothing about the format of "targets" (IDs, names, list, single element) or whether parent accepts an ID versus a name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ("Move elements") and names the destination concept (a new parent group or "root"), which is enough to distinguish it from sibling mutators like bb_transform_elements, bb_group_elements, or bb_duplicate_elements. It does not, however, explain how reparenting differs from grouping or how the hierarchy is affected.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, when not to, or which sibling (e.g. bb_group_elements, bb_set_element) handles adjacent hierarchy tasks. The agent must infer the use case entirely from the name and one sentence.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_request_fsRequest file-system permissionA

Ask the user for the plugin file-system permission (needed only by bb_list_dir, bb_glob, bb_file_info, bb_mkdir, bb_delete_path; reads and writes work without it). An optional scope limits the permission to one directory.

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeNoOptional directory to limit access to.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does most of it: it discloses that this is an interactive user-consent prompt, that the permission is narrowly scoped, and that existing read/write tools bypass it entirely. It stops short of saying what happens if the user declines or whether the call blocks.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler. The parenthetical listing dependent tools is placed immediately after the core action, and the scope note comes last, which is the correct front-loading order.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-param, zero-required permission request with no output schema, the description covers the essentials: trigger, dependents, and scope limitation. It omits the outcome of a denial or what the call returns, which is the one remaining ambiguity for an agent deciding whether to call it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100% for the single optional scope param. The description adds a genuine nuance beyond 'Optional directory to limit access to': that scope restricts the permission to exactly one directory, clarifying the bounded effect on all five dependent tools.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: 'Ask the user for the plugin file-system permission.' It goes further by scoping exactly which sibling tools require it (bb_list_dir, bb_glob, bb_file_info, bb_mkdir, bb_delete_path) and which do not, so the agent can place it precisely among ~60 siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States both when to use it ('needed only by' the five named tools) and when not to ('reads and writes work without it'), naming the alternative path explicitly. No inference required from the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_resize_textureResize a textureB

Resize the texture image (nearest neighbour). uv_width/uv_height follow the new size unless keep_uv is true.

ParametersJSON Schema
NameRequiredDescriptionDefault
widthYes
heightYes
targetYes
keep_uvNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden. It usefully adds behavioral detail the schema cannot express: nearest-neighbour resampling and the coupling of UV dimensions to the new size unless keep_uv is set. However, it says nothing about whether the resize is destructive to source resolution, permission requirements, or undo behavior, leaving the mutation profile largely opaque.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with the core action front-loaded and the conditional caveat trailing. No filler, every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a texture-mutating tool with no annotations and no output schema, the definition covers the essential mechanics (interpolation, UV handling) but omits the 'target' semantics and any note on persistence, reversibility, or return value. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It partially does: keep_uv's effect is explained and width/height are self-evident, but 'target' (the identifier of which texture to resize) is never clarified, and uv_width/uv_height are mentioned despite not appearing in the schema, which may confuse rather than clarify.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Resize the texture image') and even names the resampling method (nearest neighbour), which distinguishes it from sibling texture tools like bb_create_texture or bb_set_texture_properties. It stops short of explicitly routing the agent away from those siblings, but the purpose is unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to resize versus alternatives such as bb_set_texture_properties, nor any prerequisites or side-effect warnings. The only conditional ('unless keep_uv is true') describes a parameter effect, not when to choose this tool over another.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_reviewReview the model from several anglesA

Render the model from several camera angles into one contact-sheet PNG and return its path. Use this to LOOK at your own work and catch proportion and texture problems before declaring a model finished. The agent can then read the image file.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo
widthNoSize of each cell in pixels.
anglesNoDefault: front, right, back, left, isometric_right, top.
backgroundNo#454c57

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose the output format (contact-sheet PNG), that a path is returned, and that the agent can then read the image file, which is useful behavior beyond the schema. However it says nothing about permissions, whether an existing file at 'path' is overwritten, or cost/time of rendering.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences that front-load the core action; no filler or repetition. Slightly more verbose than strictly necessary but every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with no annotations and no output schema, the description covers purpose and workflow intent but omits parameter meaning and file-write semantics. It is adequate to call correctly in the common case but leaves real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: 'width', 'angles', and 'background' defaults are partly documented in the schema, but 'path' and 'background' have no meaning supplied anywhere. The description adds zero parameter guidance, so the coverage gap is not compensated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (render) plus resource (the model from several camera angles) and the concrete artifact produced (one contact-sheet PNG path). This is clearly distinct from the nearby bb_screenshot and bb_export_model siblings, which an agent can tell apart without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit motivation and timing: 'LOOK at your own work and catch proportion and texture problems before declaring a model finished.' That is clear when-to-use context, but it names no alternative (e.g. bb_screenshot) and offers no when-not condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_run_actionRun a toolbar actionB

Trigger any Blockbench action by its BarItems id (see bb_list_actions). This reaches every built-in command the UI exposes.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it does not say what happens on an invalid id, whether triggered actions are undoable, or that these commands may mutate/destroy project state. 'Reaches every built-in command the UI exposes' hints at broad, potentially destructive scope but stops short of warning about it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero waste, with the essential mechanic (id-driven dispatch) front-loaded before the scope note.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an unannotated, output-schema-less dispatcher that can invoke any UI command, the description omits error behavior, undo/reversibility, and side-effect expectations. The agent has enough to attempt a call but not enough to anticipate its consequences.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single 'id' parameter, so the description must compensate. It does explain that the value is a BarItems id obtainable via bb_list_actions, but adds no format, casing, or example beyond that. Partial compensation only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ('Trigger') plus resource ('any Blockbench action') and it names the mechanism (BarItems id). It also routes to bb_list_actions, cleanly separating the enumeration tool from the execution tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical '(see bb_list_actions)' tells the agent where to get a valid id, which is useful context. But there is no guidance on when to prefer this over bb_execute_js, nor any exclusions or preconditions for the call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_save_projectSave project (bbmodel)B

Compile the project to .bbmodel and write it to disk. Without a path, the current project path is reused; pass one to save a copy.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo
embed_texturesNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose one genuinely useful trait: the default-path-reuse vs save-a-copy distinction. But it says nothing about whether an existing file is overwritten, what happens to unsaved state, permissions needed, or what is returned, so the disclosure is partial for a disk-writing tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero waste, purpose front-loaded before the parameter nuance. Nothing needs trimming and nothing is buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No annotations, no output schema, and 0% schema description coverage means the description is the only documentation available. It leaves `embed_textures` unexplained and omits overwrite/return behavior, which is thin for a tool that writes a file to disk.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for both parameters. It explains the `path` parameter's default/save-as semantics well, but `embed_textures` (boolean, default true) is never mentioned anywhere, leaving half the surface undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource: compile the project to .bbmodel and write to disk. That is concrete and actionable. It does not, however, distinguish this tool from plausible siblings like bb_export_model or bb_write_file, so an agent cannot tell from the text alone why it would pick save over export.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only conditional given ('without a path... pass one to save a copy') is parameter behavior, not when-to-use guidance. There is no statement of when to prefer this over bb_export_model, bb_set_project, or bb_write_file, and no exclusions or prerequisites are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_screenshotScreenshot the viewportA

Render the current preview to a PNG on disk and return the path. This is how an agent looks at its own work. Set crop=true (default) to auto-crop to the model, or crop=false for the full viewport. The transparent viewport is composited over a solid background (override with "background", or disable with background:false). Returns a data URL too when include_data_url is true.

ParametersJSON Schema
NameRequiredDescriptionDefault
cropNo
pathNoOutput PNG. Defaults to <temp>/blockbench_mcp_shot_<timestamp>.png
widthNo
heightNo
backgroundNoCSS colour to composite over (default #262b33), or false to keep transparency.
include_data_urlNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the on-disk side effect (writes a PNG), the return value (path, plus optional data URL), and transparency compositing behavior. It omits any failure/error behavior, but coverage of the core traits is strong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with what the tool does and why it exists. Minor redundancy with the schema on background, but essentially no wasted prose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter, no-annotation tool, the description covers the main behavior, output, and two key flags, but leaves width/height unexplained and says nothing about limitations or failures. Adequate, but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%, so the description must compensate. It usefully explains crop semantics (auto-crop vs full viewport) and include_data_url, but width and height are undocumented in both schema and description, and the background explanation largely repeats the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+effect: render the preview to a PNG on disk and return the path. No sibling performs screenshots, so it is unambiguously distinguishable from the surrounding tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"This is how an agent looks at its own work" gives a clear purpose/context for reaching for this tool, and it explains the crop=true default choice vs crop=false and the background override. It stops short of naming explicit alternatives or when-not-to-use conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_select_elementsSelect elementsC

Change the outliner selection. mode: replace (default), add, remove, toggle, none.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoreplace
targetsNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the mode options (replace, add, remove, toggle, none), which is useful, but it omits what targets should contain, whether empty targets are valid, how selection interacts with existing state, and whether any permissions are needed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one front-loaded sentence followed by an inline mode list with zero filler. It is appropriately sized for the content it chooses to include; the problem is omission, not verbosity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 0% schema description coverage, the description needs to do more. It never explains the critical targets parameter, so an agent cannot confidently invoke the tool with correct arguments.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description documents only the mode parameter (values and default). The targets parameter is entirely unspecified in both schema and description, so half the parameters remain semantically opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (change) and resource (outliner selection), so the core action is clear. It does not explicitly differentiate itself from sibling selection-related tools like bb_list_elements or bb_set_element, but the scope is still understandable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or when-not-to-use guidance is provided. The description lists mode values but never says when this tool should be chosen over alternatives such as listing, setting, or transforming elements.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_set_animationEdit an animationC

Change name, loop mode or length of an animation.

ParametersJSON Schema
NameRequiredDescriptionDefault
loopNo
nameNo
lengthNo
animationNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden but only lists editable fields. It does not say whether the target is an existing animation identifier, whether unspecified fields are left unchanged (important since 0 params are required), whether edits are undoable, or what the response contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no waste, but it is terse to the point of under-specifying rather than optimally concise for a 4-parameter mutation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, no output schema, 0% schema description coverage, and no required parameters, the description leaves key questions (target identification, partial-update semantics, return behavior) unanswered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It names 'name', 'loop' and 'length' but leaves the 'animation' parameter (presumably the target identifier) undocumented, and gives no valid values for the 'loop' string nor a unit/format for 'length'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (change) and resource (animation) plus the editable fields, which distinguishes it from the create/delete/list/play animation siblings. It is clear but never explicitly names those siblings to confirm the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this versus bb_create_animation, bb_delete_animation, or bb_add_keyframe, and no preconditions. Usage is only implied by the verb 'change'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_set_elementEdit element propertiesB

Change properties of one or more elements: name, origin, rotation, from/to/size (cubes), position (locators), inflate, visibility, export, locked, box_uv, autouv, shade, mirror_uv, uv_offset, color, and per-face overrides via faces:{north:{uv,texture,enabled,tint,rotation}}. Renaming more than one element at once is refused.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetsYes
propertiesYes
apply_to_childrenNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description must carry the behavioral burden. It discloses that this is a mutation tool and that multi-element renaming is refused, but remains silent on permissions, reversibility, side effects on unspecified properties, and return behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the verb and resource, and the long property list is necessary for parameter semantics. It is dense but largely free of filler, though a bulleted list would improve readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter mutation tool with nested objects, no annotations, and no output schema, the description leaves critical gaps. It omits how to specify targets, what apply_to_children does, and other behavioral details needed to invoke the tool confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It thoroughly documents the `properties` object keys and the nested `faces` override structure, but gives no explanation of the `targets` format or the `apply_to_children` parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Change properties of one or more elements,' then enumerates the property set. It is clear what the tool edits, but does not explicitly differentiate from overlapping siblings such as bb_transform_elements or bb_set_face_uv.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or when-not-to-use guidance is provided, and no alternatives are named. The only constraint, that renaming multiple elements is refused, is operational rather than a usage guideline.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_set_face_textureAssign a texture to facesA

Assign a texture (by name/uuid) to cube or mesh faces, or clear it with null. For cubes, pass faces to limit which sides are changed.

ParametersJSON Schema
NameRequiredDescriptionDefault
facesNo
targetsYes
textureYesTexture name/uuid, or null to clear.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden; it usefully discloses the null-to-clear semantics and the cube-only meaning of 'faces'. It does not say whether an existing texture is overwritten, whether the change is undoable, or what happens on invalid texture names.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core action and followed by the one conditional detail. No filler or restatement of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter mutation tool with no annotations and no output schema, the description covers the texture and faces parameters adequately but leaves 'targets' entirely unexplained, which is a real gap given it is required and untyped.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% and the 'targets' parameter is both undocumented and untyped in the schema, yet the description never explains it. The description does add value for 'texture' (name/uuid, or null to clear) and 'faces' (limits sides on cubes), but that value largely duplicates the schema's own note.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Assign a texture ... to cube or mesh faces') plus the clearing behavior with null. It is clearly distinguishable from UV-oriented siblings like bb_set_face_uv and property-oriented ones like bb_set_texture_properties, though it never names a sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'For cubes, pass faces to limit which sides are changed' gives concrete usage context for one parameter. There is no statement of when not to use it, no prerequisites, and no mention of alternatives such as bb_set_face_uv for coordinate work.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_set_face_uvSet face UVsB

Set UVs on one face. For cubes give uv=[x1,y1,x2,y2] in texture pixels. For meshes give uvs aligned to the face vertices, or a {vertex_key:[u,v]} map.

ParametersJSON Schema
NameRequiredDescriptionDefault
uvNo
uvsNo
faceYesCube face name, or mesh face key (use bb_list_elements include_geometry:true).
targetYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It never states whether existing UVs are overwritten, whether the operation is undoable, what permissions or selection state it needs, or how failures behave. For a mutation tool with zero annotation coverage, this is a substantial disclosure gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the purpose before the format rules. No filler; each sentence adds a distinct invocation detail, though the terse style could label the cube/mesh cases more explicitly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, no output schema, and 25% schema coverage, the description handles the uv/uvs formats adequately but omits the meaning of the required 'target' parameter and any behavioral context (overwrite semantics, undo). It is minimally sufficient but leaves clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25% (only 'face' is documented), so the description must compensate. It does add real meaning for 'uv' (cube pixel rect) and 'uvs' (mesh alignment or vertex_key map), which is valuable, but 'target' remains completely unexplained in both schema and description, leaving a required parameter ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource with a scope constraint: 'Set UVs on one face.' The cube-vs-mesh format split further clarifies the operation, but it never names or contrasts itself with the obvious sibling bb_auto_uv, so the agent must infer when this manual tool is preferred over automatic UV generation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly guides invocation by telling the caller which parameter to use for cubes ([x1,y1,x2,y2]) versus meshes (aligned uvs or a vertex_key map), which is useful. However, there is no explicit when-to-use/when-not guidance and no routing to alternatives like bb_auto_uv, so selection guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_set_modeSwitch editor modeA

Switch Blockbench mode: edit, paint, animate, display or pose. Paint/animate are needed for texture and animation tools on some formats.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose a useful behavioral fact: certain tools are gated behind paint/animate modes on some formats. However, it says nothing about side effects of switching (e.g. whether current selections or unsaved state are affected) or mode availability per format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the action and its valid values, then the practical rationale. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-required-param tool with no output schema and no annotations, the description covers purpose, allowed values, and why the tool matters. Mild gaps remain around format-specific availability and the effect of switching away from a mode, but nothing critical for invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the schema declares only 'mode: string' with no enum, so the description compensates well by listing the five permitted values. It could go further by noting what happens on an invalid mode, but the enum enumeration is the key missing piece and it is supplied.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Switch Blockbench mode') and enumerates the five valid targets (edit, paint, animate, display, pose). An agent can immediately distinguish this state-setting tool from the many creation/edit siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence gives real context for when mode matters ('Paint/animate are needed for texture and animation tools on some formats'), effectively telling the agent to switch modes before using bb_create_texture/bb_paint_pixels or animation tools. It stops short of explicit when-not conditions or named alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_set_projectEdit project metadataC

Change the open project: name, texture resolution, the box_uv default and the view mode. Use it to rename a project or change the texture canvas the model is authored against.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
box_uvNo
view_modeNo
texture_widthNo
texture_heightNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It says 'change the open project' but omits whether edits persist to disk, whether they are undoable via bb_undo, whether the texture resolution change resizes/destroys existing texture data, and what happens to unspecified fields — all material for a metadata mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the resource and the field list before the use case. Nothing is padded, though the second sentence is somewhat redundant with the first's enumeration of fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 5 parameters at 0% schema coverage, no annotations, and no output schema, the description leaves too much unsaid: view_mode's legal values, texture dimension constraints, persistence and undo behavior, and return/confirmation semantics are all missing for a write tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 5 parameters, so the description must compensate. It maps four of five parameters to meaning (name, texture_width/height as 'texture resolution', box_uv, view_mode), but adds no format, default, or allowed-value information — notably view_mode is a bare string with no enumerants and box_uv's semantics are unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific mutation verb ('Change') on a specific resource ('the open project') and enumerates the four editable facets (name, texture resolution, box_uv default, view mode). It reads clearly against bb_new_project and bb_project_info, though it never names those siblings or resolves the overlap with bb_set_view / bb_set_mode, which also touch view/mode state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use it to rename a project or change the texture canvas the model is authored against' gives two concrete intents, which is better than nothing. But there is no when-not guidance, no mention of prerequisites (an open project must exist), and no routing away from bb_new_project or bb_set_view; usage is implied rather than delineated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_set_settingChange a settingC

Set a Blockbench user setting by id (e.g. viewport_zoom_speed, default_cube_size, shading).

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
valueYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not state whether settings persist across sessions, whether the change is reversible, what happens on an invalid id, or how 'value' is validated — significant gaps for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the verb and resource front-loaded and examples trailing. No wasted words, though it is arguably under-specified rather than optimally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter mutation tool with no annotations, no output schema, and 0% schema coverage, the description is too thin. It leaves the value parameter undocumented and omits persistence, error, and permission behavior an agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It usefully enumerates example 'id' values, which adds real meaning for that parameter, but says nothing about the 'value' parameter, whose schema type is unconstrained ({}). Partial compensation only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Set) and resource (Blockbench user setting by id) and supplies concrete id examples like viewport_zoom_speed and default_cube_size. It does not explicitly differentiate itself from the bb_list_settings sibling, but the write-vs-read distinction is clear from 'Set'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus alternatives (e.g. bb_list_settings to discover valid ids first) or any prerequisites. Usage is only implied by the verb 'Set', leaving the agent to infer the workflow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_set_texture_propertiesEdit texture propertiesC

Change texture metadata: name, folder, render_mode, pbr_channel, particle, fps, layers_enabled, saved.

ParametersJSON Schema
NameRequiredDescriptionDefault
fpsNo
nameNo
folderNo
targetsYes
particleNo
keep_sizeNo
pbr_channelNo
render_modeNo
layers_enabledNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Change' implies mutation, but nothing is said about whether changes are persisted/undoable (bb_undo exists as a sibling), what permissions or active context are required, or what happens to unspecified fields. Listing 'saved' as a settable field is also odd and unexplained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence that front-loads the verb and resource, which is efficient. However, the field list is a bare enumeration that both misses a schema parameter and includes a non-existent one, so the brevity comes at the cost of accuracy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter mutation tool with no annotations, no output schema, and 0% schema description coverage, the description is far too thin. It should at minimum explain the required 'targets' argument, the omitted 'keep_size' flag, and persistence/undo behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 9 parameters, so the description is the only source of parameter meaning. It enumerates 8 fields but omits 'keep_size' and 'targets' (the sole required param) and mentions 'saved', which is not in the schema — a partial, partly inaccurate mapping.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Change') and resource ('texture metadata') and enumerates the editable fields, so the agent knows what the tool does. It does not distinguish itself from related siblings like bb_set_face_texture or bb_create_texture, which is the only thing keeping it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus bb_set_face_texture, bb_create_texture, or bb_draw_texture, nor any stated prerequisites (e.g. whether a texture must already exist). The agent must infer all routing from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_setupInstall bridge and configure clientsA

Install or refresh the Blockbench plugin file and print the MCP client configuration for opencode, Claude, Cursor, Windsurf, VS Code, Gemini and Cline. Pass clients:["write"] to actually write the config files.

ParametersJSON Schema
NameRequiredDescriptionDefault
clientsNoClient ids to configure, or ["write"] for all file-based clients.
install_pluginNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose the important safe-default behavior (printing rather than writing unless clients:["write"] is passed), which is genuinely valuable. It omits whether the plugin install overwrites an existing file, what permissions or paths are involved, and what happens on failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the primary action and followed by the critical write-mode caveat. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter setup tool with no annotations and no output schema, the description covers the main behavioral quirk (print vs write) but leaves install_plugin semantics and side effects of the install step unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: clients has a schema description, install_plugin does not. The description adds real meaning for clients by explaining the special ["write"] sentinel value, but says nothing about install_plugin's effect or the default-true behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource pair: install/refresh the Blockbench plugin file and print MCP client config for named clients. It is clearly distinct from most siblings, though it does not explicitly distinguish itself from the similarly named bb_install_plugin.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The line 'Pass clients:["write"] to actually write the config files' implies the default is a print-only/dry-run, which is useful usage guidance. However, it gives no guidance on when to use this versus bb_install_plugin or bb_bridge_status, and no prerequisites are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_set_viewSet view mode and camera angleB

Change the viewport view mode (textured, solid, wireframe, normal, uv) and/or move the camera to a preset angle (front, back, left, right, top, bottom, isometric_right, isometric_left, true_isometric_right, true_isometric_left, north, east, south, west, up, down). Useful before bb_screenshot.

ParametersJSON Schema
NameRequiredDescriptionDefault
zoomNoDistance factor relative to the preset (0.5 = twice as close).
angleNo
view_modeNo

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It never says whether these changes are viewport-only (non-destructive, unsaved) or persist into the project, whether 'move the camera' affects the scene camera or just the view, or whether the action is undoable — all material for a no-annotation mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core capability and its enumerated values are front-loaded in the first sentence, with the workflow hint trailing as a short second sentence. The inline value lists are dense but every token carries information; no filler sentences are present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description needn't cover return values, and the enumerated inputs are complete. It is still thin on the call itself: with all three parameters optional it never states what happens when they are omitted, nor whether the call returns confirmation, leaving an agent guessing about defaults.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% (just 'zoom'), so the description must compensate, and it does: it enumerates every legal value for both `angle` and `view_mode`, which the schema leaves as bare strings. It adds nothing beyond the schema for `zoom`, and does not note accepted casing or whether values are case-sensitive, so it stops short of a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (change) and a concrete resource (viewport view mode + preset camera angle), and enumerates the exact accepted values for both, which is well beyond a tautology. It does not, however, distinguish itself from the nearby sibling bb_set_mode, whose name invites confusion with 'view mode' — an agent must infer the difference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Useful before bb_screenshot' gives one genuine contextual cue for when to reach for this tool. There are no exclusions, no statement of alternatives (e.g., when to use bb_set_mode instead), and 'useful' is a soft rather than directive recommendation, so usage remains implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_statusBridge statusA

Report the connection to the MCP server, the Blockbench version, the open project summary and whether file-system permission has been granted. Call this first in a session.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the report's contents, which implies a non-mutating status read, but never explicitly confirms it is read-only, nor does it mention side effects, auth requirements, or whether it can fail when the bridge is disconnected. Adequate but with clear gaps for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with no filler: the first lists what is reported, the second gives the invocation timing. The most important operational cue ('call this first') is placed at the end but is still prominent and unmissable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly takes on the job of enumerating return contents (connection, version, project summary, fs permission), which is what an agent needs to decide to call it. It stops short of describing failure behavior when the bridge is down, a realistic scenario for a status tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies. There is nothing for the description to clarify beyond what the empty schema already conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Report') and enumerates the four things reported: MCP server connection, Blockbench version, project summary, and file-system permission. This is concrete and unambiguous. However, it does not differentiate itself from the sibling bb_bridge_status, which appears to cover overlapping ground, leaving an agent to guess which status tool to call.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Call this first in a session' gives an explicit when-to-use trigger, which is more than most definitions offer. It does not address when NOT to use it or name bb_bridge_status as an alternative, so the routing guidance is incomplete rather than absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_stepRun several tools in one requestA

Execute a list of tool calls in order inside a single round trip. Each step is {tool, arguments}. Later steps may reference earlier results with "$N.path" (N = step index, or -1/"last" for the previous step), e.g. "$0.element.uuid". Failures are reported per step and do not stop the batch unless stop_on_error is true.

ParametersJSON Schema
NameRequiredDescriptionDefault
stepsYes
stop_on_errorNo

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and does well: it discloses execution order, per-step failure reporting, and stop_on_error semantics. It also explains result referencing via $N.path. It does not discuss side effects, permissions, or transactional behavior, but for a meta-tool it adds substantial useful context beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly written sentences, front-loaded with the core action. Each sentence adds distinct value: execution model, step shape and reference syntax, and failure behavior. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter meta-tool with no annotations and no output schema, the description covers what an agent needs: how to build steps, how to reference earlier results, and how errors are handled. Return values need not be explained because no output schema exists. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It defines the steps structure as {tool, arguments}, explains the reference syntax with $N.path, clarifies N indexing including -1/'last', and gives an example. It also explains stop_on_error. This fully compensates for the missing schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: execute a list of tool calls in order inside one round trip. It clearly distinguishes this batch/orchestration tool from ordinary single-action siblings. The purpose is unambiguous without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for use: batching multiple tool calls into a single round trip. It implies when this is preferable but does not explicitly name alternatives or when-not-to-use conditions, such as direct calls or bb_execute_js. A 4 is appropriate: clear context, no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_transform_elementsTransform elementsA

Move, rotate or scale elements. Use vector [x,y,z], or axis+amount. Move translates from/to/origin (cubes) or origin/position (others). Rotate adds degrees around the element origin. Scale multiplies size about the element origin (cubes, meshes, texture_mesh).

ParametersJSON Schema
NameRequiredDescriptionDefault
axisNo
amountNo
originNoOptional pivot for scale (defaults to each element origin).
vectorNoFor move/scale: [x,y,z] (scale is a factor). For rotate: degrees per axis.
targetsYes
operationYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does disclose key semantics: rotate ADDS degrees about the element origin (relative, not absolute) and scale MULTIPLIES size about the origin. That is genuinely useful and not derivable from the enum. However, it says nothing about permissions, reversibility (a bb_undo exists), failure behavior, or whether the transform is applied to targets atomically.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with the core verb list and then parameter semantics in order. Every sentence conveys information, though the move/rotate/scale clause chain is dense enough to require a second read.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter mutation tool with no annotations, no output schema, and 33% schema coverage, the description covers operations and pivot semantics reasonably but leaves the targets parameter (required!) unexplained, does not indicate whether transforms are undoable, and says nothing about return or error behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% (axis and amount are undocumented in the schema), so the description must compensate, and it does: it explains that vector is [x,y,z] for move/scale as a factor and degrees-per-axis for rotate, that axis+amount is an alternative form, and that origin is an optional pivot defaulting to each element origin. The only unaddressed parameter is targets, whose format is not explained anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb triple and resource: 'Move, rotate or scale elements.' That clearly separates it from sibling mutation tools like bb_mirror_elements, bb_array_elements and bb_duplicate_elements, each of which performs a different geometric operation. It stops short of naming an alternative explicitly, so it lands at 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives applicability rules for each operation ('Move translates from/to/origin (cubes) or origin/position (others)' and scale applies to 'cubes, meshes, texture_mesh'), which is real guidance on which element types each mode supports. It does not say when to prefer this tool over bb_array_elements or bb_mirror_elements for compound transforms, so guidance remains implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_undoUndoC

Undo the last edit (one or more steps).

ParametersJSON Schema
NameRequiredDescriptionDefault
stepsNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it does not say what happens when the history is exhausted, whether undo can be re-applied, or whether the operation is limited to edits in the current session. The only disclosed trait is that multiple steps can be undone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence, front-loaded with the action and with no filler. It is arguably too terse for the behavioral gaps it leaves, but there is no wasted content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations, no output schema, and 0% parameter documentation, the description should do more work. It omits history limits, error behavior, and any relationship to bb_redo, leaving the agent with an incomplete picture of the operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single 'steps' parameter, and the parenthetical '(one or more steps)' is the only indication of what that parameter means. It confirms the parameter controls undo depth but does not specify valid ranges or behavior when steps exceeds the available history.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Undo) and its resource (the last edit), plus the scope of a variable number of steps. It does not, however, explicitly distinguish itself from the sibling bb_redo, which is the natural alternative an agent would weigh.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance, and no mention of the related sibling bb_redo. The usage is only implied by the word 'Undo' itself, leaving the agent to infer the context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_uninstall_pluginUninstall a pluginC

Uninstall a Blockbench plugin by id and optionally delete its file.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
delete_fileNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden for a destructive mutation. It does not say whether uninstalling is reversible, what happens to the plugin's in-memory state, whether delete_file permanently removes files, or what happens if the id is unknown.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the verb and resource first and the optional behavior trailing. No wasted words, though the brevity is close to under-specification.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with no annotations, no output schema, and 0% schema coverage, the description is too thin. It never covers side effects, reversibility, or error conditions that an agent needs before invoking an uninstall operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does partially: it explains that id identifies the plugin and that delete_file controls optional file removal. However, it omits the default (false), the id format, and the consequence of setting delete_file to true.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (uninstall) and resource (Blockbench plugin) with a scoping qualifier (by id), which cleanly separates it from bb_install_plugin, bb_reload_plugin, and bb_list_plugins. It stops short of explicitly naming those siblings, so a 4 rather than a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of alternatives (e.g. bb_reload_plugin for refreshing versus removing), and no stated prerequisites such as whether the plugin must be inactive first. The agent must infer all of this.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_validateValidate the modelA

Run a static audit of the project and return a score plus findings: empty project, zero-size or degenerate cubes, meshes with too few vertices, untextured faces, UVs outside the texture, duplicate cubes, and out-of-bounds UVs. Fix and re-run before declaring a model finished.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does most of it: 'static audit' signals a non-mutating read, and it discloses the return shape (score plus findings) and the specific finding categories. It never explicitly states that the project is left unmodified or how the score is computed, so a 4 rather than a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, purpose front-loaded, and the long check list is dense but each item earns its place by telling the agent what a finding can mean. Slightly list-heavy, but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must supply return semantics, and it does ('a score plus findings' plus the finding taxonomy). Given zero parameters and no annotations, the main remaining gap is that the score's scale/meaning is not described.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. The enumerated check categories serve as useful output semantics rather than parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Run a static audit of the project') and enumerates exactly what it inspects, so an agent knows this is the model-quality checker rather than a general file or element tool. It does not explicitly differentiate itself from the similarly named sibling 'bb_review', which would have pushed this to a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Fix and re-run before declaring a model finished' gives a clear trigger context for invocation as a final QA gate. It offers no explicit when-not guidance or named alternatives (e.g., bb_review), so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bb_write_fileWrite a fileA

Write text or base64-decoded bytes to a file, creating parent directories is the caller's job. Uses Blockbench's file layer (no permission prompt).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
contentYes
encodingNotext

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses the file-layer/no-permission behavior and the parent-directory constraint, but omits overwrite semantics (does an existing file get replaced?), error behavior, and return value — material gaps for a write operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact clauses, front-loaded with verb+resource and followed by caveats. The 'creating parent directories is the caller's job' phrasing is slightly clunky but every sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-param write tool with no annotations and no output schema, the description covers the key operational constraints (no permission prompt, parent dirs) but leaves overwrite behavior and failure modes unspecified, which an agent would need to call it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains the non-obvious encoding dimension ('text or base64-decoded bytes'), which maps to the enum, but adds nothing for path or content beyond what is self-evident. It partially rather than fully covers the gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('write') and resource ('file') plus the content types handled ('text or base64-decoded bytes'), which cleanly distinguishes it from siblings like bb_read_file and bb_delete_path. An agent can tell what this tool does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The note that 'creating parent directories is the caller's job' implicitly routes the agent to bb_mkdir first, and 'no permission prompt' hints at a contrast with bb_request_fs. However, no alternative is named explicitly and there is no clear when-to-use vs when-not statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 72 tool updatesv2.0.0
    • First observedbb_add_cube
    • First observedbb_add_element
    • First observedbb_add_group
    • First observedbb_add_keyframe
    • First observedbb_add_mesh
    • First observedbb_array_elements
    • First observedbb_auto_uv
    • First observedbb_bridge_status
    • First observedbb_create_animation
    • First observedbb_create_texture
    • First observedbb_delete_animation
    • First observedbb_delete_elements
    • First observedbb_delete_keyframe
    • First observedbb_delete_path
    • First observedbb_delete_texture
    • First observedbb_draw_texture
    • First observedbb_duplicate_elements
    • First observedbb_edit_mesh
    • First observedbb_execute_js
    • First observedbb_export_model
    • First observedbb_export_texture
    • First observedbb_file_info
    • First observedbb_generate_texture
    • First observedbb_get_texture_pixel
    • First observedbb_glob
    • First observedbb_group_elements
    • First observedbb_import_texture
    • First observedbb_install_plugin
    • First observedbb_list_actions
    • First observedbb_list_animations
    • First observedbb_list_dir
    • First observedbb_list_elements
    • First observedbb_list_plugins
    • First observedbb_list_settings
    • First observedbb_list_textures
    • First observedbb_mirror_elements
    • First observedbb_mkdir
    • First observedbb_new_project
    • First observedbb_notify
    • First observedbb_open_model
    • First observedbb_paint_pixels
    • First observedbb_play_animation
    • First observedbb_project_info
    • First observedbb_read_file
    • First observedbb_reconnect
    • First observedbb_redo
    • First observedbb_reload_plugin
    • First observedbb_reparent_elements
    • First observedbb_request_fs
    • First observedbb_resize_texture
    • First observedbb_review
    • First observedbb_run_action
    • First observedbb_save_project
    • First observedbb_screenshot
    • First observedbb_select_elements
    • First observedbb_set_animation
    • First observedbb_set_element
    • First observedbb_set_face_texture
    • First observedbb_set_face_uv
    • First observedbb_set_mode
    • First observedbb_set_project
    • First observedbb_set_setting
    • First observedbb_set_texture_properties
    • First observedbb_set_view
    • First observedbb_setup
    • First observedbb_status
    • First observedbb_step
    • First observedbb_transform_elements
    • First observedbb_undo
    • First observedbb_uninstall_plugin
    • First observedbb_validate
    • First observedbb_write_file

TDQS

B3.1/5.0

Scored across 72 tools

Disambiguation4/5

Most tools have distinct resource+action targets, and the descriptions are detailed enough to guide selection. A few overlaps remain, notably bb_set_element vs bb_transform_elements for transforms, bb_set_face_uv/bb_auto_uv/bb_set_element for UVs, and bb_generate_texture/bb_draw_texture/bb_paint_pixels for texture editing, but they are not severe enough to make the set unusable.

Naming Consistency5/5

All tools use the same bb_ prefix and snake_case convention, with predictable verb_noun patterns such as bb_list_elements, bb_add_cube, bb_delete_texture, and bb_set_setting. The few non-verb-noun names (bb_undo, bb_status, bb_validate) are still consistent in style and easily understood.

Tool Count1/5

72 tools is an extreme over-provisioning for an MCP server, well beyond the 50+ threshold for severe mismatch. While Blockbench is a large domain, this many tools creates unnecessary cognitive load and makes discoverability poor for an agent.

Completeness4/5

The surface is very broad, covering project lifecycle, elements, meshes, UVs, textures, animations, plugins, settings, file-system operations, screenshots, validation, and arbitrary JS escape hatches. Minor gaps exist, such as updating or reading individual keyframes and texture layer management, but bb_execute_js can fill most of these.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers