Skip to main content
Glama

Stormworks MCP

Describe a vehicle to your AI assistant and get a file you can load in Stormworks: Build and Rescue. Stormworks MCP is a local MCP server for Claude Desktop and the ChatGPT / Codex Windows app. It gives the assistant tools to shape boat hulls, lay out interiors, assemble road vehicles from your installed parts, wire and plumb them, and find faults. The assistant checks its own work in rendered previews, then saves the vehicle to your Stormworks vehicles folder, ready to load at a workbench.

Everything runs on your PC. The server reads part definitions from your own game install and never modifies the game.

Download for Windows · Run from source · Building walkthrough · What is verified in game

What you can build

Every change comes back as a picture. The assistant compares it with your request, fixes what is off, and shows you the next version before anything is saved.

Boats and ships

Pick one of 10 hull archetypes (runabout, rowboat, fishing trawler, tugboat, motor yacht, landing craft, lifeboat, patrol boat, catamaran or barge) and adjust its length, beam, deadrise, bow shape, sheer and flare. Add superstructure from boxes, cylinders, domes and lattice masts; each shape can be pitched, yawed, mirrored to both sides, repeated in rows, or stacked on another by name. Bulbous bows, skegs, hull numbers, deck markings and painted panels are also available. To model a real ship, enter its dimensions with scale: "1:4", or use scale: "fit" to size a design to your workbench.

The battleship above took four passes:

Four 3/4 previews of the same battleship: bare hull, first superstructure, tower and turrets, final

Hulls are plain blocks by default. smoothing: "wedges" skins them with the game's slope pieces (Wedge, Pyramid and Inverse Pyramid in every size). smoothing: "wedges_v2" adds a refinement pass and keeps the original fit when the refined one adds seams or drifts too far from the shape. Neither mode has been confirmed in game yet; see hull smoothing.

A fishing boat hull smoothed with wedges

Interiors

Add decks, bulkheads and rooms with 0.75 m × 2 m doorways. preview_interior draws a section down the centreline and a labelled plan of every level. It reports each room's clear size, floor width, headroom and doors, and warns when something does not fit, such as an engine too tall for its room. In the access stage, the builder fits real manual doors and complete hatch-and-ladder assemblies wherever the whole part fits. A request that cannot fit leaves the wall or floor sealed; the builder never cuts an unfinished hole for you to fill later.

Interior cutaway of the battleship: a centreline section with 20 labelled rooms and six deck plans

Road vehicles

Start a land vehicle from a preset or a bare chassis, assembled from the parts installed in your game:

Preset

What you get

utility_buggy

Saddle-seat buggy with four suspension wheels, lights, an engine, premade fuel tanks, a battery and a radiator. The parts are mounted but not connected.

utility_4x4

Four-seat 4x4 with sliding doors and a prebuilt diesel, radiator, fuel tank and battery, all piped and wired. Includes steering, throttle, brake, starter, light and reverse controls.

humvee_4x4

Four-seat, Humvee-shaped 4x4 with the same connected systems, open bays for your own doors, and enclosed pipes under the floor.

chassis

A bare chassis at your own dimensions, to extend part by part.

The assistant can also read your saved vehicles for wheel and track layouts to learn from. Previews use the game's own part models, and can hide the bodywork to check seats, tanks and the powertrain.

The Humvee preset seen from the rear: hardtop, rear windows, tail lights and a mounted spare wheel

Wiring, pipes and repairs

Connect signal and electric wires between part ports, and route real pipes between fluid faces, either exposed or enclosed in blocks. Four reusable control assemblies (engine start and idle, clutch, brake and reverse, lighting) bind to the parts you choose. preflight_vehicle checks the fuel, air, exhaust and cooling paths, the driveline, electrical power and controls. plan_vehicle_repairs turns each finding into a fix you can preview before applying, with the problem drawn over the vehicle:

Diagnostic overlay on the Humvee: the front-left wheel with a miswired steering input in red, wheel and steering arrows in blue, and the proposed repair wire in green

Before you test, prepare_vehicle_validation exports a separate copy with a seven-point in-game checklist tied to that exact file. See diagnose, repair and verify.

Edits to new and existing vehicles

Edit single blocks or whole regions in batches: add, fill, move, rotate, paint, copy, mirror or repeat. Each batch can be previewed before you commit it, and undone afterwards. Import one of your own vehicles into a separate draft to repaint or extend it; the original file is never changed. check_seal traces each compartment to the outside and highlights any leak path.

Inspection from any angle

inspect_view points a camera at any detail, at any angle and zoom. open_in_viewer opens the vehicle in an interactive 3D viewer in your browser. The viewer, viewer/index.html, also works on its own with any vehicle file.

Related MCP server: SolidWorks MCP Server

Install

You need Windows 10 or 11 (x64), Stormworks installed through Steam, and Claude Desktop or the ChatGPT / Codex Windows app on the same PC. The download includes Python and every dependency; you do not need to install Python, uv or Git.

  1. Download stormworks-mcp-0.1.0-windows-x64.exe from Releases.

  2. Move it to a permanent folder and double-click it. There is no installer.

  3. Close your AI client. In the app, open MCP > Set up Claude Desktop or MCP > Set up ChatGPT / Codex, check the configuration path it shows, and press Add to client. Your other settings are kept, and the file is backed up first. To replace an existing Stormworks entry, tick the replacement checkbox.

  4. Press Start server, then open your client again.

The window shows the server status and a live log. Keep it open while you build: every client connects to this one shared server, and closing the window stops it and disconnects them. After you start the server again, reconnect or restart your client. A second copy of the app refuses to start.

For other clients, MCP > Custom client / raw configuration copies MCP JSON, Codex TOML or a PowerShell add command. The View menu has diagnostics and log copy/clear. Setup records the EXE's current location, so if you move the file, reopen it and replace its client entries.

The same configuration serves Codex CLI and the IDE extension on that PC. ChatGPT on the web and cloud tasks use a separate connection path that this app does not set up; see OpenAI's MCP guide. To check that a download was built from this repository, see verify a download.

Upgrading from an older setup: close your clients and any older server, replace the EXE, and use the replacement checkbox to update each client entry to --connect. Remove any duplicate Stormworks entry or extension. Entries that launch server.py directly start their own separate server and must be replaced to use the shared one.

Run from source

You need Windows with Stormworks installed through Steam (tested on Windows 11), a supported client, and uv, which installs Python and the dependencies for you.

git clone https://github.com/macery12/Stormworks-MCP.git
cd Stormworks-MCP
uv sync --locked

Open the desktop window from the repository root, set up your client from the MCP menu, and press Start server:

uv run python launcher.py

To configure a client yourself, print the exact configuration for this checkout:

uv run python launcher.py --config json      # MCP JSON
uv run python launcher.py --config toml      # Codex TOML
uv run python launcher.py --config command   # codex mcp add command
uv run python launcher.py --doctor           # check paths and game assets without changing files

For Claude Desktop, open Settings > Developer > Edit Config and add a stormworks-hulls entry under mcpServers, using the folder you cloned into:

{
  "mcpServers": {
    "stormworks-hulls": {
      "command": "C:\\path\\to\\Stormworks-MCP\\.venv\\Scripts\\python.exe",
      "args": ["C:\\path\\to\\Stormworks-MCP\\launcher.py", "--connect"]
    }
  }
}

--connect relays to the running desktop server and never starts one, so keep the window open with the server running. Quit and reopen the client after changing its configuration. For Claude Desktop, right-click the tray icon and choose Quit; in the ChatGPT / Codex app, check the entry under Settings > MCP servers and restart its connection.

Build your first vehicle

  1. Start a new chat and ask:

    List the Stormworks hull presets.

    The assistant calls list_hull_presets and lists 10 archetypes, from runabout to barge. If the tool is missing, check the server entry under Claude's Settings > Developer or the ChatGPT / Codex app's Settings > MCP servers, and check the desktop app's log.

  2. Ask for a boat:

    Design me a chunky fishing boat for bench size S. Show me a preview, then save it.

    The assistant reads the design guide, previews and refines the hull, and saves it.

  3. In Stormworks, open a workbench, choose Load, and pick the name the assistant gave the vehicle.

Load it and check it against the in-game checklist: many features have not been verified in game yet. The building walkthrough covers stages, units and exact edits.

Example prompts

Boats

  • "Give me three different 10 m rescue boat hulls and show them side by side."

  • "Take the tugboat preset but make it longer and sleeker, like a pilot boat."

  • "Add an interior: engine room aft with a large engine, crew quarters amidships, and a bridge."

  • "Try it with wedge smoothing" or "open it in the viewer".

  • "Build the tugboat through the propulsion stage, show every automatic choice, then save it."

Road vehicles

  • "Build the Humvee preset for bench S, then show me the equipment with the bodywork hidden."

  • "Find the wheeled vehicles in my saves and compare their wheelbases."

  • "Run preflight on my 4x4 and preview the repairs before applying any."

  • "Export a test copy of the 4x4 and give me the in-game checklist."

Existing vehicles

  • "Look at my vehicle 'Old Trawler' and design a new hull in the same style."

  • "Import Old Trawler into a draft, repaint the bridge, and save a separate copy."

  • "Add a block-built diesel tank, check its enclosure, and show the outlet and vent connections."

Tools

You do not call these yourself: describe what you want, and the assistant picks the tools. Before building, it reads hull_design_guide for boats (topic="workflow" returns a short starting guide) or land_vehicle_guide for road vehicles. Hull dimensions are in metres; exact edits use integer blocks of 0.25 m. Saving refuses to overwrite a vehicle file the server did not create, and replacing one it did create needs overwrite=true.

Area

Tools

Guide

Guides and setup

hull_design_guide, land_vehicle_guide, get_runtime_status, list_workbenches

Building

Hull design

list_hull_presets, get_hull_spec, preview_hull, preview_interior, deck_profile, save_hull

Building

Designs and drafts

store_design, list_designs, load_design, import_vehicle, preview_vehicle, save_vehicle

Staged builder

Part editing

query_parts, edit_parts, undo_edits

Staged builder

Road vehicles

create_land_vehicle, find_land_vehicles, search_land_parts, analyze_land_vehicle

Land vehicles

Wiring and pipes

query_connections, edit_connections, route_connections, list_vehicle_assemblies, apply_vehicle_assembly

Vehicle repair

Faults and repairs

preflight_vehicle, plan_vehicle_repairs, repair_vehicle, preview_vehicle_diagnostics, check_seal

Vehicle repair

In-game test records

prepare_vehicle_validation, record_vehicle_validation, get_vehicle_validation, get_calibration_observations

In-game testing

Part catalogue

search_parts, get_part_definition, get_part_orientation

Staged builder

Inspection and analysis

inspect_view, open_in_viewer, list_game_vehicles, preview_game_vehicle, analyze_vehicle, analyze_hull, suggest_hull_blocks

Hull smoothing

Issue reports

complaint, list_complaints, get_complaint

Saved locally as JSON and Markdown

Configuration

A standard Windows and Steam install needs no configuration. To override a default, set these environment variables before starting the desktop app:

Variable

Default

Used for

SW_VEHICLES_DIR

%APPDATA%\Stormworks\data\vehicles

Where vehicles are read and saved.

SW_GAME_DIR

Found through your Steam libraries

Stormworks install, for part definitions and workbench locations.

SW_DEFINITIONS_DIR

<game>\rom\data\definitions

Part definitions directly.

SW_WORKSHOP_DIR

Found through your Steam libraries

Installed workshop vehicles used as references.

SW_DESIGNS_DIR

%APPDATA%\stormworks-hull-mcp\designs

Stored designs and drafts.

SW_COMPLAINTS_DIR

%APPDATA%\stormworks-hull-mcp\complaints

Local JSON and Markdown issue reports.

SW_TOOL_TIMEOUT

300

Seconds a build or render may run before the server stops it.

SW_BUILD_CACHE

1

Set to 0 to turn off geometry caching.

SW_BUILD_CACHE_DIR

%TEMP%\stormworks-mcp-builds

Shared geometry cache, capped at 20 entries / 512 MiB.

SW_MCP_PORT

38473

Loopback port of the shared desktop server (1024 to 65535).

SW_MCP_RUNTIME_DIR

%LOCALAPPDATA%\stormworks-hull-mcp\runtime

Runtime files of the shared desktop server.

Without game definitions, plain blocks, slopes and structural editing still work. Installed components need definitions: automatic fit-out reports what it skipped, and a seal check that meets unknown component geometry reports the result as indeterminate.

Status and limitations

This is an early release (0.1). The server writes files without running the game, so the game is the only real test. In-game testing tracks the status of every feature.

  • Verified in game: vehicles load centred in the workbench, blocks-only hulls float, and Wedge 1x1/1x2/1x4, Pyramid and Inverse Pyramid rotate correctly in all 40 tested orientations.

  • Player-tested, with a fix awaiting a check: the connected Humvee's controls, engine and steering work, but W drove it backward because the gearbox's off ratio was reverse. The replacement cube gearbox has not been checked in game yet.

  • Not yet verified in game: both smoothing modes and the larger Pyramid sizes; interiors, doors and hatches; placed components and block-built tanks; seal checks; round and angled shapes, bulbous bows, skegs and paint; and operating the utility buggy and utility 4x4.

  • Placed parts are not a working system. The boat fit-out stages and the utility buggy place engines, tanks, propellers and rudders without connecting them. Add pipes and wires with the connection tools, then test in game.

  • Nothing is simulated. Preflight checks that paths and links exist; it does not simulate physics, fluid flow, engine load or driving. Control logic comes from the connected 4x4 presets and the four control assemblies.

  • Imports accept single-body, version-3 vehicles only. Some configured original parts can only be repainted. Save the result under a new name.

  • Paint is one colour per block, so hull blocks show their outside colour inside rooms.

  • Large designs: a 118 m, 184k-part ship previews in about 15 s, or about 35 s with wedge smoothing. The game appears to drop components past 131,072 on spawn; this is not yet confirmed.

  • The 3D viewer loads three.js from a CDN, so it needs internet access.

How it works

A hull spec describes a continuous shape: plan-view taper, keel rise, sections with deadrise and bilge radius, sheer, and superstructure. The server voxelizes that shape at 0.25 m per block, hollows it to a watertight skin, adds interior structure, and can skin it with slope pieces. Once stored, any design (hull, land or imported) can be changed part by part: the editing, connection and repair tools apply revision-checked batches with undo. The result is written as Stormworks vehicle XML. Previews are drawn in Python with Pillow from the same geometry, using the installed component meshes.

Heavy builds and renders run in a child process that the server stops when a call is cancelled or exceeds SW_TOOL_TIMEOUT, so one slow ship cannot stall later calls. docs/vehicle-format.md documents the file format, including the details that are easy to get wrong: a missing r attribute is not the identity rotation, x in the paint string means unpainted, and the game's axes are left-handed.

Documentation

Guide

Read it to

Building walkthrough

Make a first build, keep units straight, edit exactly and diagnose a result.

Staged builder

Build in stages, edit parts, import vehicles, and add components, access, seals and tanks.

Land vehicles

Use the land presets, the chassis spec, parts, connections and preflight.

Diagnose, repair and verify

Repair faults, apply control assemblies and record in-game tests.

Hull smoothing

Choose slope pieces and check hull depth.

In-game testing

Check a vehicle in game and see what is verified.

Vehicle file format

Read or write Stormworks vehicle XML from your own tools.

Releases

Verify a download, build the EXE locally or publish a release.

Advanced testing

Run placement suites, analyze reference vehicles and generate calibration exhibits.

Development

uv sync --locked                  # install locked dependencies, including pytest and ruff
uv run pytest                     # unit tests; engine tests skip when the game is not installed
uv run ruff check .               # lint
uv run tools/smoke_test.py out    # build and render every preset in both modes into ./out
uv run tools/client_test.py       # drive the server over MCP stdio, like Claude Desktop does
uv run tools/staged_builder_test.py out/staged-builder --benchmark
uv run tools/diagnostic_milestone.py out/diagnostic-review  # three-fault repair milestone
uv run python tools/package_test.py # two clients, shared host, guide assets, workers, export and stop

On Windows, build the portable desktop executable:

uv sync --locked --group build
uv run --group build python tools/build_windows.py
uv run python tools/package_test.py --exe dist/stormworks-mcp.exe
uv run python tools/desktop_test.py --exe dist/stormworks-mcp.exe

Tagging v0.1.0 (or v0.1) runs the tests, builds and tests the executable, and uploads it to a draft GitHub release with checksums and signed build provenance; see releasing and verification. Changes must pass lint and tests. Geometry changes also need a rendered review and a specific in-game check, as described in CLAUDE.md.

Design notes and history: feature requests and bugs, project audit, hull smoothing research and surface design plan.

Path

Contents

server.py

MCP tool definitions

launcher.py

Desktop app entry point and --connect client relay

swhull/hull.py, presets.py

Hull spec to solid voxels; the 10 archetypes

swhull/smooth.py

Hollow, watertight skin and slope-piece fitting

swhull/interior.py, access.py

Decks, rooms and doorways; complete doors and hatches

swhull/components.py, tanks.py

Component fit-out; block-built fluid tanks

swhull/land_build.py, land_presets.py

Land chassis and presets

swhull/networks.py, routing.py

Typed wiring and preflight; pipe routing

swhull/repairs.py, assemblies.py, validation.py

Repair suggestions, control assemblies and in-game test records

swhull/editing.py, drafts.py

Transactional part edits and lossless imports

swhull/seal.py

Compartment air connectivity

swhull/pieces.py, vehicle.py

Piece geometry and the rotation convention; vehicle XML

swhull/render.py, meshes.py

PNG previews using installed component meshes

swhull/definitions.py, benches.py

Game install detection, part definitions and workbench sizes

swhull/jobs.py, cache.py

Killable worker processes; shared geometry cache

swhull/setup_gui.py, desktop_runtime.py

Desktop window and shared server

swhull/guide.md

Design guide served by hull_design_guide

viewer/

Standalone 3D viewer

tools/

Smoke, client, package and desktop tests; in-game calibration vehicles

License

MIT. This is an unofficial fan project, not affiliated with or endorsed by Geometa, the developer of Stormworks. It modifies no game files and ships no game assets; part definitions are read from your own install at runtime.

Available Tools

50 tools
analyze_hullA

Measure actual hull material, floor gaps, and partial-face seams at z stations. x/stations are integer BLOCK coordinates. name reads a saved v3 vehicle (one body at a time); otherwise build spec/preset/design in the fixed build frame. Procedural reports include floor height above keel and drop from rim, so verify depth before saving. Imported saved vehicles use their original body-local frame; choose body_id explicitly when the largest structural body isn't the hull. source="backups" reads data/backups/vehicles, e.g. name="autosave9". This is read-only geometry, not a seal test.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNo
nameNo
specNo
patchNo
designNo
presetNo
sourceNovehicles
body_idNo
stationsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it declares read-only geometry, distinguishes procedural vs imported frame behavior, and warns about depth before saving. It stops short of describing the return payload, but an output schema exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded, leading with the measurement scope before frame and input-mode details. Every sentence adds operational content; no filler, though the run-on style is less scannable than ideal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists, return values need not be explained, and the description covers the coordinate system, input modes, frame caveats, and source override. The one gap is the undocumented 'patch' parameter for a 9-param tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% with 9 params, so the description must compensate and largely does: it explains x/stations as integer block coordinates, name/spec/preset/design as input modes, source='backups' semantics, and when to pick body_id. Only 'patch' is left unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource - measuring hull material, floor gaps, and partial-face seams - at a defined scope (z stations). It explicitly distinguishes itself from the seal-checking sibling with 'This is read-only geometry, not a seal test.'

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the two main modes: name reads a saved vehicle one body at a time, otherwise build spec/preset/design in the fixed build frame. Also notes source='backups' and warns to verify depth before saving. It doesn't fully route the agent against analyze_vehicle or preview_hull, so no explicit exclusion set.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_land_vehicleA

Analyze a saved v3 vehicle or draft: body-local axles, wheelbase, mounts and light axes. Sections: wheels, lights, controls, equipment, issues. All component rows are paginated. forward is a horizontal unit vector; otherwise infer each body's driver-seat facing, falling back to +z explicitly. Axle span measures mounting origins, not tyre-centre track. Multi-body references stay separate. Suspension/steering sweep and mesh edits need game checks.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
limitNo
designNo
offsetNo
sourceNovehicles
body_idNo
forwardNo
sectionNowheels
wheel_rolesNo
exclude_wheel_idsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations the description carries the full burden, and it does disclose real behavior: component rows are paginated, 'forward' must be a horizontal unit vector with a documented inference fallback to +z, and axle span measures mounting origins rather than tyre-centre track. It also flags that suspension/steering sweep and mesh edits require in-game verification. It stops short of stating read-only versus mutating intent or auth/permission needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first clause, then sections, then the geometry caveats. Sentences are dense and technical but each earns its place. Slightly heavy on vector/geometry detail relative to the unexplained parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return formatting need not be described. However, for a 10-parameter, zero-coverage, annotation-free analysis tool, the description leaves several parameters unexplained and gives no sibling routing versus 'analyze_vehicle'. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 10 parameters, so the description must compensate and only partially does. It explains 'forward' semantics thoroughly, enumerates the 'section' values, and covers pagination (limit/offset) and multi-body separation (body_id). But 'design', 'source', 'wheel_roles' and 'exclude_wheel_ids' get no explanation beyond their names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Analyze') and resource ('a saved v3 vehicle or draft') and enumerates the analytical dimensions (body-local axles, wheelbase, mounts, light axes) plus output sections, so the agent knows exactly what this produces. It does not, however, explicitly distinguish itself from the sibling 'analyze_vehicle'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the section list tells the agent what can be queried, and the closing note ('Suspension/steering sweep and mesh edits need game checks') implies a validation boundary. There is no explicit when-to-use versus 'analyze_vehicle' or 'land_vehicle_guide', and no stated prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_vehicleA

Read saved-vehicle examples, body groups, link coverage and rudder mounting/motion issues. Supports multi-body references without editing them. Observations are not game verification. Search rudder/propeller/engine/trans to focus evidence. Sections: parts, links, controllers, bodies, placement_issues, connection_candidates, open_transmission_ports. All are paginated.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
limitNo
offsetNo
searchNo
sourceNovehicles
sectionNoparts

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does fairly well: 'without editing them' discloses non-destructive read-only behavior, 'Observations are not game verification' discloses a reliability limitation on output, and 'All are paginated' discloses pagination behavior. It omits auth requirements and rate limits, but for a read/analysis tool the key behavioral traits are covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The content is dense and front-loaded with purpose, followed by caveats, search guidance, and the section list. Every sentence adds information, though the run-on structure of the middle sentences could be tightened slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary, and the description covers scope, section inventory, search focus, pagination, and key caveats. The remaining gap is an explicit routing statement versus the parallel analysis siblings, which would fully round it out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate and largely does: 'Sections: parts, links, controllers, bodies, placement_issues, connection_candidates, open_transmission_ports' enumerates the valid values for the section parameter, the search terms hint at the search parameter's semantics, and 'All are paginated' explains the limit/offset pair. It leaves the name and source parameters (including source default 'vehicles') undefined.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource: it reads saved-vehicle examples and reports on body groups, link coverage, and rudder mounting/motion issues. It enumerates the concrete output sections (parts, links, controllers, bodies, etc.), so the agent knows exactly what it produces. It does not explicitly differentiate itself from the close sibling analyze_land_vehicle, which slightly weakens disambiguation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives actionable usage context ('Search rudder/propeller/engine/trans to focus evidence', 'Supports multi-body references without editing them') and a caveat that observations aren't game verification. However, it never states when to prefer this over sibling analysis tools like analyze_land_vehicle or preview_vehicle_diagnostics, so the when-to-use guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

apply_vehicle_assemblyA

Configure/wire reusable controls using explicit current part IDs. Existing target drivers are replaced atomically. Clutch needs its own gate, separate from the idle throttle gate. Starter is manual/hold; idle is open-loop, and reverse must be selected while stopped. Preview includes diagnostics. commit=true creates one undo step; tune behavior in game.

ParametersJSON Schema
NameRequiredDescriptionDefault
commitNo
designYes
optionsNo
assemblyYes
bindingsYes
revisionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
assemblyYes
bindingsYes
revisionYes
committedYes
preflightYes
operationsYes
verificationYes
base_revisionYes

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses atomic replacement of existing target drivers, commit=true creating one undo step, preview-includes-diagnostics, and control semantics (starter manual/hold, idle open-loop). It omits permissions/authorization or failure modes, but the mutation and reversibility profile is clearly conveyed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then dense constraint and side-effect sentences with zero filler. Every sentence adds information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values needn't be explained, and the behavioral constraints are valuable. But for a 6-parameter, 0%-coverage tool in a dense sibling ecosystem, the description leaves key parameters and the scope of 'assembly' (engine_start_idle, clutch_engagement, brake_reverse, lighting) unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 6 parameters, so the description must compensate. It does explain commit's undo behavior and the clutch_gate vs throttle_gate distinction and that bindings take current part IDs, but it never explains design, revision, the assembly enum values, or the AssemblyOptions fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('configure/wire') and a resource ('reusable controls') and distinguishes intent from the many sibling preview/edit tools by naming 'explicit current part IDs'. However, 'reusable controls' is somewhat abstract and the description never names the vehicle-assembly domain explicitly, so an agent must infer the resource from the tool name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives behavioral constraints ('clutch needs its own gate', 'reverse while stopped') but no when-to-use or when-not guidance against siblings like edit_connections, route_connections, or preview_vehicle. Nothing tells the agent when this tool is the right choice over alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_sealA

Trace finished geometry from interior to outside: sealed, leaking or indeterminate. Generated drafts choose room/hull seeds automatically. Imported drafts require seeds: [[x,y,z],...] or [{name,position:[x,y,z]}], in integer blocks in their original body frame. door_state=closed/open models supported doors. Returns connected compartments and a highlighted escape path. Unsupported nearby sealing geometry prevents a confident pass.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedsNo
designYes
door_stateNoclosed

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does reasonably well: it discloses the return content (connected compartments and a highlighted escape path), a modeling option (door_state closed/open), and a failure mode ("unsupported nearby sealing geometry prevents a confident pass"). It does not state whether the operation is read-only or mention any side effects or permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and outcomes are front-loaded in the first sentence, followed by tightly packed operational details. Every sentence carries information, though the seed-format enumeration is dense and slightly jargon-heavy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter tool with no annotations and no output schema, the description supplies the missing pieces: what it returns, the seeds contract, the door modeling, and a known failure condition. It falls just short of complete because the 'design' parameter and any permission/precondition details are unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate and largely does: it specifies the exact seeds formats ([[x,y,z],...] or [{name,position:[x,y,z]}]) in integer blocks within the original body frame, and explains when seeds are auto-chosen versus required. door_state is given its closed/open values. Only 'design' is left unexplained, which is a minor gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource ("Trace finished geometry from interior to outside") and enumerates the three possible verdicts (sealed, leaking, indeterminate), so the agent understands this is a hull-seal verification tool. It is clear but does not explicitly distinguish itself from adjacent tools like preflight_vehicle or analyze_hull.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the seed requirements for generated vs imported drafts and the door_state option hint at when the tool applies, but there is no explicit 'use this instead of X' guidance or listed prerequisites. No alternative tool is named or excluded.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

complaintA

Report an encountered problem or improvement request; save local JSON and Markdown reports. Required: title and description. Include failing tool/arguments, expected vs actual behavior, error messages, steps, affected design/vehicle/definition and a suggested improvement if known. Categories: placement, rotation, smoothing, connections, definitions, performance, tool_error, usability, missing_feature, other. Severity: low, medium, high, blocker. Returns a report id and file paths; no external issue is published.

ParametersJSON Schema
NameRequiredDescriptionDefault
toolNo
stepsNo
titleYes
actualNo
designNo
contextNo
vehicleNo
categoryNoother
expectedNo
severityNomedium
definitionNo
suggestionNo
descriptionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden and does so reasonably: it discloses that files (JSON and Markdown) are written locally and that 'no external issue is published', which is the key scope boundary. It does not state permissions, failure behavior, or whether files are overwritten, so it is not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and side effect, then the required fields, then enumerations. Slightly dense and it restates enum values already machine-readable in the schema, but every line carries actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter mutation tool with no annotations, the definition covers required inputs, content expectations, enum vocabularies, and the local-only output boundary; the output schema covers return values. The missing piece is any note on the 'context' parameter or where files land.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate across 13 parameters, and it largely does: it names the two required fields and maps nearly every other parameter (tool, arguments, expected/actual, steps, design/vehicle/definition, suggestion, category, severity). Only 'context' is left unexplained, keeping it below 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Report') plus the resource ('an encountered problem or improvement request') and the side effect ('save local JSON and Markdown reports'). 'complaint' alone is ambiguous, and the description resolves it well; however it never names the create-vs-read distinction against siblings like list_complaints/get_complaint.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for when to invoke: an encountered problem or an improvement request, with the expected contents (failing tool/arguments, expected vs actual, steps, suggestion). It stops short of any exclusion or explicit alternative, so it cannot reach 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_land_vehicleA

Create a land draft: default utility_buggy, or preset=chassis for a custom bare layout. Read land_vehicle_guide for the spec. Dimensions/positions are game metres on the 0.25 m grid. Buggy includes bodywork, four suspension wheels, saddle, premade fuel tanks, engine, battery, radiator and actual lights. Parts must fit, mount and leave driver/service access. preset=humvee_4x4 has open custom-door bays, compact glass and enclosed chassis pipes. humvee_4x4 and utility_4x4 include radiator/fuel/air/exhaust/driveline pipes and typed links. Use query_connections/edit_connections/route_connections/preflight_vehicle to complete custom layouts. Extend through query_parts/edit_parts/preview_vehicle, then save_vehicle.

ParametersJSON Schema
NameRequiredDescriptionDefault
specNo
designYes

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does a fair job: it discloses what the default buggy contains, that parts must fit/mount/leave driver and service access, and that some presets come with pipes and typed links. It does not state whether this persists anything (draft vs saved), what happens on invalid layouts, or permissions, so it is informative but not complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose and preset distinction are front-loaded, which is good. However, the inventory of buggy parts ('bodywork, four suspension wheels, saddle, premade fuel tanks, engine, battery, radiator and actual lights') and the humvee_4x4 detail are verbose relative to the decision the agent needs to make, pushing the length beyond what earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema and no annotations, so the description must cover behavior, parameters, and returns. It covers behavior and the tool pipeline well, but omits what the tool returns (draft handle/id?) and leaves the meaning of 'design' and 'spec' unresolved, which is a real gap for a required-parameter creation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and neither 'design' nor 'spec' is explained in the description. Worse, the description discusses preset=chassis and preset=humvee_4x4 while the schema exposes no 'preset' parameter, leaving unclear whether those values belong in the required 'design' string or the optional 'spec' object. Some vocabulary is added (grid units of 0.25 m, game metres), but the two actual parameters remain semantically opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Create a land draft') and immediately qualifies it with the default (utility_buggy) and the alternative custom path (preset=chassis). It also sketches the rest of the workflow, which separates it from sibling tools like save_vehicle or preview_vehicle. It stops short of explicitly contrasting with load_design/import_vehicle, which also produce vehicle structures.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for the two entry modes (default buggy vs preset=chassis bare layout) and names the exact follow-up tools for each branch: query_connections/edit_connections/route_connections/preflight_vehicle for custom layouts, then query_parts/edit_parts/preview_vehicle, then save_vehicle. This is a strong pipeline, but there is no explicit when-not-to-use or 'use X instead' statement against siblings such as import_vehicle.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

deck_profileB

Deck height above the keel, the y a box gets when left out, and the deck half-beam, every step metres from the transom. Use it to place turrets and deckhouses on a sheered deck. Lengths are in spec units (real-world metres when the spec has a scale).

ParametersJSON Schema
NameRequiredDescriptionDefault
specNo
stepNo
patchNo
designNo
presetNo
spec_pathNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses units ('spec units ... real-world metres when the spec has a scale') and output semantics, but does not state side-effect profile, permissions, or whether it mutates the spec.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with no filler; the returned quantities are front-loaded. It could be slightly clearer about the tool's operation, but it is well-sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, but the input schema has no parameter descriptions and the description covers only one of six parameters. For a 6-parameter computational tool, an agent lacks enough guidance on choosing between spec/design/preset/spec_path/patch.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 6 parameters. The description only clarifies `step` ('every step metres from the transom') and indirectly `spec` via the scale note; it leaves `patch`, `design`, `preset`, and `spec_path` unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource and its returned quantities (deck height above keel, box y, deck half-beam along step from transom) and gives a concrete use case (placing turrets/deckhouses). It does not explicitly contrast with sibling hull/preview tools, so it is clear but not fully differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use it to place turrets and deckhouses on a sheered deck,' giving a clear context for invocation. It provides no when-not conditions or named alternatives, so it misses the top level.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_connectionsA

Preview/commit typed control and electrical wires. Uses the query_connections revision. Ops: {op:'connect',from:{part_id,port},to:{part_id,port}} or {op:'disconnect',link_id}. Signal outputs connect to inputs; electricity is bidirectional. One driver per signal input. Power/fluid use route_connections. Original XML/settings remain lossless; commit has undo.

ParametersJSON Schema
NameRequiredDescriptionDefault
commitNo
designYes
revisionYes
operationsYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses preview-vs-commit (commit defaults false), the revision/concurrency dependency, and critically that original XML/settings remain lossless and commit is undoable. It also surfaces the 'one driver per signal input' rule and bidirectionality. What preview returns is not described, but the mutation semantics and reversibility are well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense, front-loaded, and largely waste-free: the core purpose leads, followed by op syntax and then scope/behavioral notes. The terse shorthand is efficient, though a few phrases require domain familiarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter mutation tool with no annotations and no output schema, the description supplies enough to invoke it correctly: op grammar, preview/commit behavior, revision source, and undo. It does not describe the return payload of a preview, which is the main remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It documents the complex 'operations' parameter with concrete discriminated shapes for connect and disconnect, and covers commit and revision semantics. Only 'design' is left entirely unexplained, so coverage is strong but not complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (edit/preview/commit) and resource (typed control and electrical wires) up front. It explicitly carves out the scope from route_connections, so an agent can distinguish the two connection tools without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit exclusion with a named alternative: 'Power/fluid use route_connections.' It also signals the workflow dependency on query_connections for the revision. It stops short of stating when to preview vs commit or the ordering prerequisites, so it is clear but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_partsA

Atomic part edits; defaults to preview only. Pass the revision from query_parts. Ops: add(part/parts), fill(bounds,color), remove/replace/move/rotate/paint/copy/mirror/repeat. Selection ops need select={ids/bounds/name/definition}. move/copy/repeat use delta in blocks; repeat count is additional copies; rotate uses a local-to-world matrix or r string and pivot; mirror uses axis and plane. replace needs part; paint needs color. Added parts use definition, position in blocks, rotation, optional mirror (1=x, 2=y, 4=z), color, name and scalar settings. Placement rotation accepts an effective mirrored matrix; use that or proper rotation plus mirror. Prebuilt engines default to max_force_scale=1 (100%). commit=true keeps one undo step.

ParametersJSON Schema
NameRequiredDescriptionDefault
commitNo
designYes
revisionYes
operationsYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose real behavioral traits: preview-by-default, single-undo-step semantics on commit, and the max_force_scale=1 default for prebuilt engines. It omits permission/error behavior, but the transaction and default-mode disclosures are substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose and default mode before the op enumeration, and every sentence carries semantic content. Density is high and it reads as a run-on telegraphic block, but little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 10-op discriminated-union mutation tool with no annotations, no output schema, and 0% schema coverage, the description supplies the key op semantics, selection fields, and commit behavior an agent needs. It is close to complete, with minor gaps only in return/error behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it largely does: select={ids/bounds/name/definition}, delta in blocks, count as additional copies, rotate via matrix or r-string plus pivot, mirror via axis/plane, and the Part fields (definition, position, rotation, mirror 1/2/4, color, name, settings). The telegraphic op-list covers most operation parameters, though top-level design/commit/revision get only brief treatment.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Atomic part edits') and immediately pins down its operating mode ('defaults to preview only'), so an agent knows this modifies parts in a design. It names the full op set but does not differentiate itself from siblings like edit_connections beyond the resource name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives workflow context ('Pass the revision from query_parts', 'commit=true keeps one undo step'), which implies the preview-then-commit flow. However it never states when to use this tool versus edit_connections or other edit siblings, and gives no explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_land_vehiclesB

Read the player's saves for wheel/track examples, including unnamed experiments. kind=wheeled/tracked/all. Counts only vehicle body components, not controller internals. Wheel-based boats/aircraft can match; this does not establish vehicle purpose or quality.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNowheeled
limitNo
offsetNo
searchNo
sourceNovehicles

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must carry the behavioral burden. It does add useful context: it reads saves, includes unnamed experiments, counts only vehicle body components (not controller internals), and warns that wheel-based boats/aircraft can match and that it does not establish purpose or quality. This is strong behavioral context beyond the name, though it doesn't cover auth needs or performance.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences, with the core action front-loaded. Each sentence adds a distinct caveat (unnamed experiments, component-scope, false positives). No wasted words, though the caveats could be slightly more structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the key behavioral quirks. But with 5 undocumented parameters and no annotations, it falls short of being complete enough for an agent to call it correctly without inspecting the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and 5 parameters are undocumented. The description only explains the 'kind' parameter (wheeled/tracked/all) and implies filtering by saves, leaving limit, offset, search, and source completely undiscussed. With low schema coverage, the description should compensate but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a verb (find) and resource (land vehicles) via the name, and the description clarifies it reads the player's saves for wheel/track examples. However, the purpose is somewhat muddled by the caveats about boats/aircraft and the mention of 'unnamed experiments', and it doesn't clearly distinguish itself from siblings like search_land_parts or list_designs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies usage by mentioning kind filtering and that wheel-based boats/aircraft can match, but it never says when to use this tool versus alternatives like search_land_parts or analyze_land_vehicle. No explicit when/when-not guidance is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_calibration_observationsB

Read recorded in-game checks for a part; expose stale evidence and contradictory reports.

ParametersJSON Schema
NameRequiredDescriptionDefault
definitionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden; 'Read' does signal a non-mutating operation and the mention of stale evidence and contradictory reports hints at analytical value in the results. It says nothing about permissions, freshness thresholds, or scope of the records, so the disclosure is partial rather than complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the verb and resource come first and the second clause adds useful intent. It is compact to the point of leaving gaps, but those gaps are coverage issues rather than wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the tool is simple with one parameter. Against that low complexity bar the description is only adequate: the unexplained 'definition' parameter and absence of any routing guidance versus siblings leave real holes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single required parameter is a bare string named 'definition' with 0% schema description coverage. The description says the checks are 'for a part', which hints at the value's meaning but never resolves whether 'definition' is a part ID, a name, or a definition document, leaving the most important input ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Read) and resource (recorded in-game checks for a part), which is concrete enough to distinguish it from siblings like get_part_definition or query_parts. It stops short of explicitly contrasting itself with any sibling, but the resource named is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The clause 'expose stale evidence and contradictory reports' implies the tool is for auditing the validity of prior calibration work, which gives an agent a sense of when to reach for it. However, no alternatives are named and there is no explicit when-not guidance against the many get_* siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_complaintB

Read a complaint's complete reproduction evidence and Markdown report by its id.

ParametersJSON Schema
NameRequiredDescriptionDefault
complaint_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral burden alone. 'Read' usefully signals a non-mutating, single-record fetch, but nothing is said about permissions, error behavior for a missing id, or payload size. An output schema exists, so return-format disclosure is not required here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. The verb, resource, lookup key, and returned content are all packed in without waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation, and for a one-parameter read-by-id tool this is largely sufficient. The only real omission is guidance on how an agent obtains a valid complaint id, which list_complaints presumably supplies.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the schema does not explain complaint_id. The phrase 'by its id' at least establishes that the parameter is the record identifier, but adds no format, validation, or sourcing details (e.g., where ids come from).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Read') and resource ('a complaint') and pins the lookup key ('by its id'). It is clearly distinct from list_complaints, but it never names or contrasts with the sibling tools, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus list_complaints (e.g., 'use list_complaints first to obtain the id'). Usage is only implied by the name and the 'by its id' phrasing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_hull_specB

Full resolved spec (every parameter) for a preset, or the defaults. Use as a starting point.

ParametersJSON Schema
NameRequiredDescriptionDefault
presetNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the burden. It usefully discloses that the returned spec is fully resolved ('every parameter') and that omitting a preset yields defaults, but it says nothing about failure behavior for an unknown preset or whether the result is read-only vs. persisted. Adequate but thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with what is returned and followed by the fallback behavior. 'Use as a starting point' is the only sentence that is more filler than substance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return format need not be explained, and with one optional param the surface is small. Still, the description omits how to discover valid preset names and what a resolved spec contains at a high level, leaving an agent to infer the surrounding workflow.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does: it explains that the optional `preset` argument selects a named preset while its absence falls back to 'the defaults,' which is the key semantic the schema cannot convey. It stops short of describing valid preset identifiers.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: it returns the full resolved hull spec for a preset or the defaults. It implicitly distinguishes itself from list_hull_presets (which enumerates presets) by resolving one into every parameter, though it never names that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use as a starting point' gestures at a workflow role but never says when to prefer this over list_hull_presets, hull_design_guide, or preview_hull, nor any prerequisite such as obtaining a valid preset name first. No when-not guidance at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_part_definitionC

Installed footprint, attachment/sealing surfaces and relevant part settings.

ParametersJSON Schema
NameRequiredDescriptionDefault
definitionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it only enumerates returned content. It does not state that this is a read-only lookup, whether the definition must already exist, whether an ID or name is required, or any error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short fragment with no waste, but it is not front-loaded around an action and reads as an under-specified return-field list rather than a purposeful definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, yet the description still omits how to supply a valid 'definition', its relationship to sibling retrieval tools, and any prerequisites. Thin for even a simple one-parameter lookup.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single required 'definition' string parameter is completely unexplained. The phrase 'part definition' in the description loosely hints at the argument's meaning, but neither its format (ID vs name) nor source is clarified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a noun phrase listing what the part definition contains (installed footprint, attachment/sealing surfaces, settings), so an agent can infer the resource is a part definition, but there is no verb and no differentiation from siblings like get_hull_spec, get_part_orientation, or query_parts. Purpose is recognizable but vague in verb and scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this tool versus query_parts, search_parts, get_part_orientation, or get_hull_spec, and no prerequisites or context are given. Only an implicit 'retrieve part info' meaning is conveyed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_part_orientationB

Explain local mounting/motion/function axes and solve their requested world directions. Example Fin Rudder targets: mount_normal=[0,0,1], span_axis=[0,1,0]. Road wheel targets should include axle_axis, wheel_reference_up and wheel_forward; right placements can require mirror. Returns effective rotation plus proper r/mirror for XML. Arrow signs still need game checks.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetsNo
definitionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses output content (effective rotation plus r/mirror for XML) and a real caveat ('Arrow signs still need game checks'), which is useful. It does not state whether this is a pure read/compute operation, its permission needs, or limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: the core purpose leads, followed by examples and the output/caveat notes. Every sentence carries information, though the example and caveat lines run long without much structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be fully explained. Against a two-parameter tool with 0% schema coverage and no annotations, the description leaves the required 'definition' parameter and its accepted values or format unexplained, which is a real gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description should compensate. The examples do add semantic value by naming valid target keys (mount_normal, span_axis, axle_axis, wheel_reference_up, wheel_forward) and their value form. However, the required 'definition' parameter is never explained, so compensation is only partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action (explain mounting/motion/function axes and solve requested world directions) on a specific resource (part orientation). It is recognizable as an axis-orientation solver distinct from the many vehicle/hull/design siblings. It lacks explicit sibling differentiation, but no sibling covers this function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Concrete examples (Fin Rudder, Road wheel) imply when the tool applies, including the mirror case for right placements, but there is no explicit 'use this when…' guidance or statement of when another tool is preferred. Usage is implied rather than directed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_runtime_statusB

Diagnose version, runtime paths, game definitions and timeout without writing files.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses that the operation does not write files, indicating a non-mutating read, but adds nothing about permissions, side effects, or freshness of the reported status.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the key content (what is diagnosed) leads. It is efficient, though quite terse given the absence of other guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and with no parameters the description only needs to convey scope and non-mutating behavior, which it does. The main gap is the missing when-to-use guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to clarify. The baseline of 4 applies since no parameter semantics are needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource scope (version, runtime paths, game definitions, timeout) and frames it as a diagnostic read. It's clear what information the tool surfaces, though the verb 'diagnose' is slightly odd for what is effectively a status getter. No sibling overlap exists, so differentiation is not needed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this tool versus alternatives, no stated prerequisites, and no exclusions. The agent must infer usage entirely from the one-line summary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_vehicle_validationB

Read game-test evidence and report whether the export bytes and draft revision still match.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose the core behavior: it reads game-test evidence and performs a match comparison between export bytes and draft revision. It does not state what happens on a mismatch, missing run_id, permissions, or whether the check is cached/idempotent, so coverage is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single well-formed sentence with the core purpose front-loaded and no filler. It is efficient, though the extreme brevity contributes to the underspecification noted elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description adequately conveys the read/compare intent. However, the undefined run_id and absent usage guidance leave gaps for an agent deciding when and how to invoke it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description never mentions run_id, so the single parameter's meaning (which validation run to look up, format, scoping) is left entirely to inference. The description should have compensated for the coverage gap but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action (read game-test evidence and report whether export bytes and draft revision match), which is concrete and distinguishable from write-oriented siblings like record_vehicle_validation. It does not, however, explicitly differentiate itself from prepare_vehicle_validation or record_vehicle_validation, leaving the agent to infer the get/prepare/record split.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to call this versus prepare_vehicle_validation or record_vehicle_validation, nor any prerequisites or exclusions. Usage is only weakly implied by the word 'Read' and the validation theme.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hull_design_guideA

Read first. Topics: full, workflow, units, spec, interior, archetypes, style, limits, staged, building, testing, smoothing. Focused topics avoid resending the entire spec reference.

ParametersJSON Schema
NameRequiredDescriptionDefault
topicNofull

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose one useful trait: that focused topics avoid re-sending the entire spec reference, i.e., a token/payload-size consideration when choosing a topic. It does not state that the tool is read-only and side-effect-free, nor describe output volume for the 'full' default, so the disclosure is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded imperative ('Read first'), then the topic vocabulary, then a one-clause rationale for choosing focused topics. No filler or repetition, appropriate for a small lookup tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-shape explanation is correctly out of scope, and the topic list is nearly exhaustive. But for a 'read me first' entry point it never says what the guide actually covers, how to interpret topics like 'edits' or 'topics', or how it relates to the similarly-named get_hull_spec / land_vehicle_guide siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the lone 'topic' parameter. It lists most enum values but explains none of them and omits 'edits' and 'topics' entirely, leaving the agent to guess. It does add one semantic note (focused topics are lighter than 'full'), which lifts it above a pure enum restatement.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The name plus 'topics' enumeration make it clear this is a documentation/reference reader for hull design, and the topic vocabulary (spec, interior, archetypes, staged, building) distinguishes it from action siblings like save_hull or edit_parts. It stops short of stating outright 'returns hull design documentation', and it does not distinguish itself from the parallel land_vehicle_guide or get_hull_spec sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Read first' is an explicit sequencing instruction that tells the agent when to invoke this before other operations, which is strong usage guidance. However, it names no alternatives and gives no when-not condition (e.g., versus get_hull_spec, which sounds like it may overlap).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

import_vehicleB

Import an existing single-body version-3 vehicle into a new draft; never changes the source. Existing configured/wired parts are protected. Use query_parts/edit_parts and save_vehicle.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
designYes
sourceNovehicles

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses two behavioral traits: 'never changes the source' (non-destructive read of source) and 'Existing configured/wired parts are protected' (side-effect scoping on the destination draft). It omits permission requirements, failure modes (e.g., non-version-3 or multi-body source), and what 'protected' means operationally.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose in a compact two-sentence form with no filler; the semicolon and trailing tool references are efficient. Slight structure roughness from the fragment-like final sentence, but nothing wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers non-destructiveness plus the follow-on workflow. However, for a draft-creating tool with zero annotations it leaves prerequisites and error conditions unaddressed, so it is adequate but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for name, design, and source — and it does not. It never explains what 'design' identifies or what the 'source' collection default ('vehicles') selects. Only the word 'vehicle' loosely hints at the resource being referenced.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (import) and resource (vehicle) with a meaningful scope qualifier: 'existing single-body version-3 vehicle into a new draft.' That constraint meaningfully separates it from siblings like create_land_vehicle and load_design, though it doesn't name any sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use query_parts/edit_parts and save_vehicle' provides post-import workflow guidance, which implies context but is not a when-to-use rule or an explicit alternative to import_vehicle itself. It never says when NOT to import, nor contrasts with load_design or apply_vehicle_assembly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inspect_viewA

Render one large view of a hull design or saved vehicle from any angle, like a camera you can point. Use it to check details the fixed previews hide.

name: a vehicle in the player's vehicles folder, or pass spec/preset/design(+patch). yaw: 0 = side view with bow to the right, 90 = from the bow, 180 = other side, 270 = stern. pitch: degrees above (+) or below (-) the horizon; -30 shows the hull bottom. zoom: 1 = whole vehicle, 2-6 = close-up. focus: [fx, fy, fz] fractions of the bounding box to centre on (x across, y up, z stern to bow), e.g. [0.5, 0.3, 0.9] = low on the bow. highlight: a superstructure box name (designs only): painted magenta and labelled, and its position is reported in spec units and as game-metre heights above the keel. Views at pitch 0 or +-90 get metre rulers in spec units (z from the transom, y from the keel, x from the centreline).

ParametersJSON Schema
NameRequiredDescriptionDefault
yawNo
nameNo
specNo
zoomNo
focusNo
patchNo
pitchNo
designNo
presetNo
sourceNovehicles
body_idNo
highlightNo
spec_pathNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does substantial work: it explains the camera model, the meaning of each angle/zoom/focus input, and what the render adds (magenta-labeled superstructure, spec/metre rulers, keel/transom reference). It never explicitly states the operation is side-effect free, which is the main gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, then parameter semantics follow compactly. Dense but each line adds real meaning; only the trailing ruler note is slightly tangential.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter render tool with no annotations and no output schema, the description covers the viewing behavior and output additions well but omits several input parameters and any explicit statement of side effects or prerequisite for the various source modes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

At 0% schema description coverage the description must compensate, and it documents the key viewing parameters richly (yaw orientation, pitch sign, zoom ranges, focus fractions, highlight). However source, body_id, and spec_path are left undocumented and spec/patch/design/preset are only name-dropped, so coverage is partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Render one large view of a hull design or saved vehicle from any angle') and contrasts it with the fixed previews, letting an agent distinguish it from siblings like preview_hull/preview_vehicle without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use it to check details the fixed previews hide' gives a clear context for selecting this over the fixed-preview siblings. It stops short of naming the specific alternatives or stating exclusions, so it is not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

land_vehicle_guideC

Road vehicle workflow, installed parts, chassis spec, units, layout limits and game checks.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure, but it only lists topics covered. It does not state whether the tool is read-only, what it returns, or any side effects, permissions, or limits. This is a significant gap for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence, but it is structured as a topic list rather than a front-loaded purpose statement. It avoids waste but does not clearly orient the agent to what the tool does.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter informational tool with an output schema, the description lists the covered subjects, which is somewhat complete. However, it does not clarify how to use the guide or its relationship to other vehicle tools, leaving ambiguity among many siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there are no parameter semantics to document. The empty schema and 100% description coverage mean the baseline of 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (road/land vehicle guide) and enumerates covered topics, but it does not use a clear verb to say what the tool actually does (returns documentation, performs checks, etc.). It is distinguishable from operational siblings like find_land_vehicles or create_land_vehicle, but the purpose remains vague without a defined action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives. There is no mention of prerequisites, context, or sibling tools such as hull_design_guide. The agent is left to infer usage entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_complaintsB

List locally recorded complaints newest first; filter by text, category or severity.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
offsetNo
searchNo
categoryNo
severityNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose two useful traits: local-only scope and newest-first ordering. However, it says nothing about pagination behavior, default limit of 50, or that this is a non-destructive read, leaving notable behavioral gaps for an unannotated tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that states purpose, ordering, and filterable dimensions with zero filler. Nothing is wasted and the key information (list, ordering, filters) leads.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation, and the description covers scope, ordering, and filters. But for a 5-parameter list tool with 0% schema coverage and no annotations, the absence of any pagination/default-limit context leaves it only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and it only partially does: 'filter by text, category or severity' maps to the search/category/severity params, but limit and offset (pagination) are never explained. Three of five parameters gain meaning; two remain opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ('List ... complaints') plus a meaningful scope qualifier ('locally recorded') and an ordering ('newest first'). It clearly differs from the singular get_complaint/complaint siblings, but it never names them, so sibling differentiation is left implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus get_complaint, complaint, or other list_* tools, and no stated prerequisites or exclusions. The agent is left to infer that this is the bulk-retrieval counterpart to the single-complaint getter.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_designsC

Names of stored hull, land and imported drafts.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It implies a read operation by returning names, but does not state side effects, authentication needs, ordering, pagination, or whether any mutation occurs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short and front-loads the output content, which is efficient. However, it is an incomplete noun phrase rather than a structured sentence, and its brevity leaves important context unstated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple zero-parameter list tool with an output schema, the description covers what is returned at a high level. Still, it omits usage guidance and cannot be fully distinguished from sibling list tools, leaving meaningful gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there are no parameter semantics to explain. The baseline score for a parameterless tool is 4, and the description does not need to add parameter details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource and output ('Names of stored hull, land and imported drafts'), which suggests a listing tool. However, it lacks an explicit verb and does not clearly distinguish itself from sibling list tools such as list_hull_presets or list_game_vehicles.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus alternatives. It does not mention prerequisites, filtering, or when a caller should prefer load_design, list_hull_presets, or other listing tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_game_vehiclesC

Vehicles in the player's Stormworks vehicles folder, optionally filtered by a substring.

ParametersJSON Schema
NameRequiredDescriptionDefault
searchNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, yet it never states that this is a non-destructive read, whether results reflect live game state, or how results are ordered. It only discloses scope (the player's vehicles folder) and the optional substring filter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with no filler. It leads with the resource rather than the action, which is a minor front-loading weakness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the tool is a simple single-parameter list. Still, absent annotations, the description should say something about read-only semantics and result behavior to be complete for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate; it does explain that 'search' acts as an optional substring filter. However, it omits details like case sensitivity or matching fields, so it only partially fills the gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

It names a specific resource (vehicles in the player's Stormworks vehicles folder) and its filtering behavior, so an agent can tell what data comes back. The verb is only implied via the name, and no sibling (e.g. list_designs) is explicitly contrasted, so it falls short of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this over list_designs or the other list_* siblings, and no prerequisites or exclusions are given. Usage is only inferable from the name and description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_hull_presetsA

List built-in hull archetypes with a one-line description and main dimensions.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. The verb 'List' implies a read-only, non-destructive operation, and the description discloses what each result contains. It does not discuss pagination, permissions, ordering, or other operational details, but for a zero-parameter built-in preset listing these gaps are relatively minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no waste. It tells the agent what is listed and what the listing includes, which is appropriate for this simple tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter listing tool with an output schema, the description is nearly complete: it identifies the resource and previews the key result fields. The main missing context is explicit routing versus sibling tools, but return values are handled by the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero input parameters, so there are no parameter semantics to document. Baseline credit applies; the description correctly does not invent or misstate any parameter behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: list built-in hull archetypes. It also previews the returned content (one-line description and main dimensions). It is clear, though it does not explicitly differentiate itself from nearby tools such as get_hull_spec or list_designs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: an agent would call this when it needs available built-in hull presets. However, the description gives no explicit when-to-use guidance, no when-not-to-use guidance, and no alternative-tool routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_vehicle_assembliesB

Reusable controls, required ID bindings and gate definitions. Place any required gates using edit_parts, then bind queried IDs with apply_vehicle_assembly; geometry is not preset-specific.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description bears the full behavioral burden. It notes only that 'geometry is not preset-specific' and gives an ordering constraint, but says nothing about read-only nature, permissions, idempotency, or scope of the listing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with no padding, but the leading fragment is not front-loaded with a clear statement of the operation, which hurts scannability. Reasonably sized but structurally awkward.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be explained, and the description does sketch the resource and a follow-on workflow. It still leaves gaps about what the listing actually returns and when this tool is the right entry point.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there are no parameter semantics to document; the baseline for 0 params is 4. The schema coverage is 100% yet empty, so nothing more is required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The name carries the verb ('list') but the description opens with a bare noun phrase, 'Reusable controls, required ID bindings and gate definitions,' without ever stating that it returns a listing of vehicle assemblies. It conveys what an assembly contains but does not clearly frame the operation or distinguish it from siblings like list_workbenches.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides workflow guidance for adjacent tools ('Place any required gates using edit_parts, then bind queried IDs with apply_vehicle_assembly'), which implies a sequence. However, it never states when to call list_vehicle_assemblies itself versus alternatives, so usage is only inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_workbenchesA

Bench size keywords (S, M, L, XL, XXL, MAX) with their size in metres and the edit areas in the player's game (including workbench mods) that are at least that big. Set one as bench in a spec; preview summaries then say whether the design fits and by how much not.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the returned content (size keywords, sizes in metres, qualifying edit areas) and the behavioral consequence of setting a bench in a spec (preview summaries report fit). It doesn't mention auth, rate limits, or whether it's strictly read-only, but for a simple reference list this is good transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with what the tool provides and then how to use the result. It is appropriately sized with no filler, though the first sentence is a dense noun phrase rather than a clean statement of action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained in depth. The description still adds useful context about what the list contains and how it feeds into spec previews. For a parameter-free reference tool, this is nearly complete, with only minor ambiguity about whether it lists workbenches themselves or just bench size categories.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies. There are no parameters whose semantics need explanation, and the description does not need to compensate for schema gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource: bench size keywords (S, M, L, XL, XXL, MAX) with their metre sizes and the edit areas that meet or exceed each size. It doesn't explicitly state the verb 'list' or differentiate from sibling tools, but the tool name supplies the verb and the content is specific enough that an agent knows what it returns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear downstream context: 'Set one as `bench` in a spec; preview summaries then say whether the design fits and by how much not.' This tells the agent what to do with the output. It does not name alternative tools or state when not to use it, but for a reference-listing tool the guidance is solid.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

load_designD

Hull spec, or land/import draft metadata and revision. Use part tools to edit land/imports.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, yet it never says whether this operation mutates state, loads from disk, or is read-only, and gives no permissions, side effects, or reversibility notes. 'Land/import draft metadata and revision' hints at what is operated on but discloses no behavioral trait an agent could act on.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is brief, but the brevity reflects under-specification rather than economy; the leading noun phrase 'Hull spec' does not front-load the tool's action, and the second sentence is a redirect rather than useful instruction.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, but the agent still lacks the essential facts: what this tool loads, from where, and when to choose it over the many hull/design/import siblings. The definition is not sufficient to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter 'name' has 0% schema description coverage and the description never explains what name refers to (a hull spec name, design name, or file name) or its format. Nothing compensates for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a fragment ('Hull spec, or land/import draft metadata and revision') with no clear verb or resource. An agent cannot tell whether load_design reads, creates, or restores a design, nor how it differs from siblings like list_designs, store_design, save_hull, or get_hull_spec.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It redirects editing work to 'part tools' but never says when to call load_design itself or what prerequisite state it expects. No alternative is named for the load/spec case, and no when-not condition is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

open_in_viewerC

Open a vehicle in the interactive 3D viewer in the player's web browser.

name: saved vehicle or installed numeric workshop item (source=workshop). design supports hull/land/imported drafts. Uses installed component meshes, with explicit fallbacks. For articulated references choose body_id; otherwise shows the largest body alone.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
specNo
patchNo
designNo
presetNo
sourceNovehicles
body_idNo
spec_pathNo
diagnosticsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does add some behavioral context: rendering uses installed component meshes with explicit fallbacks, and articulated references collapse to the largest body if body_id is omitted. However, it omits whether the action is read-only, whether it requires a running game/browser session, and any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence and the follow-up notes are compact. The second paragraph is slightly run-on but every line carries information; nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, but with 9 parameters, 0% schema coverage, and no annotations, the definition should do far more. It leaves half the parameters and the safety/session profile unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 9 parameters, and the description only clarifies name, source=workshop, design, and body_id. spec, patch, preset, spec_path, and diagnostics remain entirely undocumented in both schema and description, leaving a substantial gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with a clear modality: 'Open a vehicle in the interactive 3D viewer in the player's web browser.' This tells an agent exactly what happens, though it never explicitly differentiates itself from sibling preview/inspect tools such as preview_vehicle or inspect_view.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no alternatives named among the many preview/inspect siblings. The only conditional ('For articulated references choose body_id') is parameter selection, not tool selection, so an agent gets no routing help.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plan_vehicle_repairsA

Diagnose faults and test suggested edits/wires/pipes without changing the draft. Returns severity, affected IDs, explanations, concrete operations and manual-review reasons. Use the returned revision, plan_id and available finding IDs with repair_vehicle.

ParametersJSON Schema
NameRequiredDescriptionDefault
designYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
plan_idYes
findingsYes
revisionYes
verificationYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it delivers: it declares the operation is non-mutating ('without changing the draft') and enumerates the return content (severity, affected IDs, explanations, operations, manual-review reasons). It omits auth/permission requirements and any rate or cost limits, but the safety-relevant behavior is well disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the primary action and its key constraint, then the return contents, then the handoff instruction. Every sentence carries information; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be spelled out (the description nonetheless summarizes them helpfully). The non-mutating nature and the repair_vehicle handoff make it complete enough to call for a diagnostic tool, though the undocumented 'design' parameter is the notable remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single required parameter ('design') with 0% schema description coverage, yet the description never defines it, its format, or whether it takes an ID versus raw design content. The only hint is the oblique reference to 'the draft' and the returned plan_id/revision, which is inference rather than documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair and resource ('Diagnose faults and test suggested edits') and immediately scopes it with 'without changing the draft'. This cleanly separates it from the mutating sibling repair_vehicle, which the description explicitly names, so an agent can distinguish the two without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the downstream alternative explicitly ('Use the returned revision, plan_id and available finding IDs with repair_vehicle'), giving a clear condition for the handoff. However, it does not distinguish this tool from other diagnostics-adjacent siblings like preview_vehicle_diagnostics or analyze_vehicle, so exclusions are incomplete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preflight_vehicleB

Check engine fuel/air/exhaust/radiator paths, driveline, electrical power and controls. Reports configuration_checks (explicit engine power), gearbox_configuration_checks (saved off/on ratio indices, including reverse-off warnings) and wheel_direction_checks (drive arrows and steering signs for supported direct/inverting paths). Unknown logic stays unknown. Includes blocked transmission exits. Separates physical pipes from typed wires and keeps functional component circuits distinct. Connected geometry still needs in-game testing.

ParametersJSON Schema
NameRequiredDescriptionDefault
designYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
checksYes
issuesYes
statusYes
findingsYes
revisionYes
wire_countYes
verificationYes
missing_countYes
physical_face_pairsYes
configuration_checksYes
non_driven_wheel_facesYes
wheel_direction_checksYes
open_transmission_facesYes
capped_unused_tank_facesYes
gearbox_configuration_checksYes

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and largely meets it: it names the check outputs (configuration_checks, gearbox_configuration_checks, wheel_direction_checks), warns about reverse-off warnings and blocked transmission exits, and sets expectations with 'Unknown logic stays unknown.' It does not discuss permissions, cost, or mutation risk, but for an inspection tool that omission is minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The scope is front-loaded in the first sentence and the remaining sentences each add a distinct behavioral fact rather than padding. It is dense and runs long, and some return-content detail risks overlapping with the available output schema, but nothing is pure filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with an output schema, the description is substantively complete on behavior — it explains the check categories and the unknown/blocked-edge semantics. The gaps are the undocumented parameter and the absence of routing guidance against near-duplicate siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one required parameter ('design') with 0% schema description coverage, and the description never mentions it at all — it does not say whether this is a design ID, name, or handle, nor its format. The schema does none of the work here, so the description should have compensated and doesn't.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Check engine fuel/air/exhaust/radiator paths, driveline, electrical power and controls'), which is far more concrete than the bare name. However, it never distinguishes itself from very close siblings such as preview_vehicle_diagnostics, preview_vehicle, analyze_vehicle, or prepare_vehicle_validation, so an agent cannot route between them from this text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance, and no sibling is named as an alternative. The only contextual cue is 'Connected geometry still needs in-game testing,' which is a limitation caveat, not a usage rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

prepare_vehicle_validationC

Export a separate test vehicle and create a seven-check in-game checklist tied to its exact XML SHA-256 and draft revision. Starts pending; topology never counts as game evidence.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
designYes
overwriteNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
checksYes
designYes
run_idYes
statusYes
testerYes
revisionYes
created_atYes
part_countYes
game_versionYes
vehicle_nameYes
vehicle_pathYes
export_sha256Yes
evidence_levelYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose meaningful traits: the checklist starts pending, it is tied to a specific SHA-256 and draft revision, and topology never counts as game evidence. But it omits permissions, reversibility, and what happens to an existing artifact when overwrite is set, which matters for a create/export operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences, front-loaded with the core action and followed by key state constraints. No filler or restatement of the title. Slightly jargon-heavy but appropriately sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema covers return values, so that burden is lifted, but the tool is a stateful export/create operation with no annotations, fully undocumented parameters, and no usage routing. An agent lacks enough to know when to invoke it and what name/design/overwrite should contain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for all three parameters. The description references 'draft revision' and 'XML SHA-256' but never maps these to name, design, or overwrite, and gives no clue what a 'design' value or overwrite semantics are. It fails to compensate for the zero-coverage schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: exports a separate test vehicle and creates a seven-check in-game checklist tied to XML SHA-256 and draft revision. This is concrete and non-tautological. However, it does not distinguish itself from nearby siblings like preflight_vehicle, get_vehicle_validation, or record_vehicle_validation, which an agent could easily confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance, and no alternative tool is named. The only usage hint is 'Starts pending', which implies a workflow state but does not tell the agent when to call this versus preparing/previewing/recording validation elsewhere.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preview_game_vehicleA

Render any vehicle from the player's vehicles folder, e.g. to study their existing boats.

Uses installed component meshes, with reported footprint fallbacks. source=workshop reads an installed numeric item ID. For articulated vehicles choose body_id; by default show the largest body alone, without pretending unrelated local frames line up.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
sourceNovehicles
body_idNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does add real behavior: mesh sourcing, footprint fallbacks, and the default of showing the largest articulated body alone. However, it never says what the render produces or where it goes (image, viewer, etc.), which is a notable gap for a tool with no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action followed by behavioral and parameter detail. The second paragraph is dense but each sentence carries distinct information (mesh sourcing, source semantics, body selection default). No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Parameter guidance is solid and behavior is partly covered, but with no annotations and no output schema the description should explain what a 'render' returns and how the result is surfaced. That return-side information is missing, leaving the agent uncertain about the tool's output.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it largely does: it clarifies that source=workshop reads an installed numeric item ID and that body_id selects a body for articulated vehicles. 'name' is left implicit, but the two ambiguous parameters gain meaningful context absent from the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Render any vehicle from the player's vehicles folder.' The scope (game/player vehicles) distinguishes it from generic siblings like preview_vehicle, and the 'study existing boats' example clarifies intent. It never explicitly contrasts against preview_vehicle or list_game_vehicles, so differentiation is implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a use case ('to study their existing boats') that implies when to reach for it, but offers no explicit when-not guidance and doesn't name the alternative (preview_vehicle) or the condition that selects it. Usage is inferable but not spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preview_hullA

Build a hull and return a preview image (3/4 above and below, front, side, top) plus stats.

spec: partial or full hull spec (see hull_design_guide); merged over the preset if given. preset: optional preset name to start from. design: a stored design (store_design / save_hull) to start from instead of a preset. patch: JSON-Patch ops applied last, e.g. [{"op": "replace", "path": "/superstructure/bridge/height", "value": 2.5}]. List items can be named instead of indexed. spec_path: a .json file holding a spec (or a stored design), read before spec and patch.

ParametersJSON Schema
NameRequiredDescriptionDefault
specNo
patchNo
designNo
presetNo
spec_pathNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so the description carries full burden. It discloses the merge/precedence order of inputs (spec_path read before spec and patch; spec merged over preset), which is useful behavioral context. However, it doesn't state cost, latency, whether the build persists, or error behavior for an operation that constructs geometry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Opens with the outcome (preview image + stats), then itemizes inputs tersely. The wrapped JSON-Patch example is well-placed. Slightly dense with formatting quirks (pedantic code-span spacing) but every line earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-param, no-annotation, no-output-schema tool, the description covers the input model thoroughly including precedence ordering. Gaps remain: no output schema means the 'stats' return is uncharacterized, and the spec structure is outsourced to hull_design_guide without saying what's in it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: each parameter gets a meaning, the merge/precedence rules are explained, and the patch param includes a concrete example. Loses a point because list-item-by-name behavior is mentioned only in passing and spec format is deferred to an external guide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Build a hull and return a preview image') and enumerates exactly what the preview contains (3/4 above/below, front, side, top) plus stats. This distinguishes it from sibling preview_interior and preview_game_vehicle, though it doesn't name them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description lists five input sources (spec, preset, design, patch, spec_path) implying these are alternative starting points, but never states when to pick one over another or that they're mutually exclusive/merged. No exclusions or alternative tools named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preview_interiorA

Cutaway of the interior: a section on the centreline plus a labelled deck plan for every room level, and a per-room report (clear size, floor width, doors with sill heights and floor steps, engines, warnings). Use after adding interior to a spec (see hull_design_guide, "Interior design"). Floors above the main deck are drawn cropped to their rooms. design/patch/spec_path: as in preview_hull.

ParametersJSON Schema
NameRequiredDescriptionDefault
specNo
patchNo
designNo
presetNo
spec_pathNo

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses output behavior (floors above the main deck are cropped to their rooms) and the content of the per-room report, but never states that this is a non-mutating read/preview, nor covers failure modes when interior data is absent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the output description and tight overall; every clause adds information (output anatomy, prerequisite, cropping rule, param delegation). Slightly dense but no filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Rich output description compensates for the missing output schema, and the workflow prerequisite is stated. However, with 5 zero-coverage params, no annotations, and one parameter (`preset`) unaddressed, the definition leaves real gaps for an agent to guess around.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% across 5 params, so the description must compensate. It partially does so by deferring design/patch/spec_path to preview_hull ('as in preview_hull'), but leaves `preset` completely unexplained and gives no format details for any parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the exact artifact (a cutaway interior render) and enumerates what it contains: a centreline section, a labelled deck plan per room level, and a per-room report with specific fields. This is distinguishable from sibling preview_hull without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit trigger ('Use after adding `interior` to a spec') and points to hull_design_guide's 'Interior design' section for workflow context. It stops short of stating when not to use it or naming a competing alternative beyond the implicit contrast with preview_hull.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preview_vehicleA

Preview a draft using installed meshes. yaw/pitch/zoom/focus gives a close-up. layer=all/components/structure; components hides bodywork to inspect seats/tanks/powertrain. Paint and wheel neutral poses are approximate; the image is not an in-game screenshot.

ParametersJSON Schema
NameRequiredDescriptionDefault
yawNo
zoomNo
focusNo
layerNoall
pitchNo
designYes
door_stateNoclosed

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does so well by disclosing the rendering source (installed meshes) and fidelity caveats: paint/wheel poses are approximate and it is not an in-game screenshot. This is exactly the kind of expectation-setting an agent needs. It does not mention door_state behavior or whether anything is persisted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action and never padded. The layer and camera semantics are packed densely, though the camera sentence is slightly telegraphic.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must characterize the return; it does say an image is rendered ('not an in-game screenshot'), but says nothing about format, size, or how the required design is consumed. With no annotations and 0% schema coverage, missing design/door_state semantics leave gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% across 7 parameters, so the description must compensate. It adds meaning for yaw/pitch/zoom/focus (close-up framing) and layer values (all/components/structure, with components hiding bodywork), but leaves the required 'design' parameter and the door_state enum entirely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Preview a draft using installed meshes'), which is clear and distinguishable from read-only inspection siblings. However, 'draft' is ambiguous (a design?) and it never explicitly distinguishes itself from preview_game_vehicle or preview_vehicle_diagnostics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains how to get a close-up via camera params and what layer=components is for, which implicitly guides usage. But it gives no explicit when-to-use versus preview_game_vehicle/preview_hull, nor any prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

preview_vehicle_diagnosticsB

Show red fault parts, amber blocked exits and blue wheel/steering arrows over geometry. Repair and assembly previews additionally show proposed green wires/pipes. Uses draft block coordinates.

ParametersJSON Schema
NameRequiredDescriptionDefault
yawNo
layerNocomponents
pitchNo
designYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
overlayYes
revisionYes
preflightYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it does disclose real behavioral content: the color-coded overlay semantics and the fact that it reads draft block coordinates. However, it never confirms the operation is read-only/non-mutating, says nothing about permissions, rendering cost, or how the draft geometry is resolved.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the primary visual behavior and followed by the conditional extra overlay. No padding, though the second sentence mixes preview-type context with a data-source note rather than cleanly routing usage.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the visual semantics are covered. But with four undocumented parameters and zero annotation coverage, the definition leaves key invocation details (especially layer and the required design input) unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across four parameters (design, layer, yaw, pitch), and the description mentions none of them. Since the schema does not document them and the description does not compensate, an agent gets no meaning for 'layer' enum values or what yaw/pitch control.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb-plus-resource ('Show ... over geometry') and enumerates exactly what the diagnostic overlay contains: red fault parts, amber blocked exits, blue wheel/steering arrows. That is a clear, non-tautological purpose, but it never distinguishes itself from sibling previews such as preview_vehicle or preview_game_vehicle.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance and no named alternative. The phrase about repair/assembly previews implies a context but does not tell the agent which sibling to pick for which task, leaving selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_connectionsC

Installed/configured node indices, transmission surface indices, wires and draft revision. Wire endpoints: {part_id,port}; pipe endpoints: {part_id,surface_index}. Coordinates are uncentred integer blocks. Paginate components; links retain stable IDs for disconnect.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
designYes
offsetNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does supply some real behavioral context: pagination over components, stable link IDs preserved for later disconnect, and the coordinate convention ('uncentred integer blocks'). However, it never states that the operation is read-only or what the response shape is beyond the output schema, leaving gaps for a mutation-adjacent tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is terse and free of filler, but it reads like shorthand notes rather than a front-loaded statement of purpose; the reader must assemble meaning from fragments. Reasonably sized, but the most important information (what it does) is not placed first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because an output schema exists, return values need not be enumerated in prose, which excuses some brevity. Still, the required input, usage context, and read-only status are missing, so the definition is only partially complete for the calling decision.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description should compensate but largely does not. 'Paginate components' loosely covers limit/offset, but the required 'design' parameter is never explained, and no parameter formats or defaults are given in prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a noun-phrase list ('Installed/configured node indices, transmission surface indices, wires and draft revision') rather than a clear verb+resource, so the agent must infer from the name that this reads connection topology. It conveys the domain (node/wire/pipe connections) but never states outright what the tool returns or does. The intended read-only nature is only implied by contrast with siblings like edit_connections.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no mention of alternatives such as edit_connections or route_connections, which an agent must distinguish this from. The only procedural hint is 'Paginate components,' which touches mechanics but not selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

query_partsC

Parts and revision for precise editing. select: ids, name, definition or inclusive bounds [[min_x,min_y,min_z],[max_x,max_y,max_z]]. Coordinates are integer blocks (0.25 m), in the uncentred build frame for generated hulls/land drafts and original body-local frame for imports. Rows include scalar settings, proper rotation plus a local mirror bitmask, and the effective transform matrix. Partial multi-voxel selections are rejected; select an id to target the whole component.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
designYes
offsetNo
selectNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present the description carries the full behavioral burden, and it does add real value: coordinate frames (uncentred build frame vs. original body-local frame), units (0.25 m integer blocks), the contents of returned rows, and the rejection of partial multi-voxel selections. It stops short of describing pagination/limit behavior, error conditions, or permissions, so coverage is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The body is dense but each sentence carries technical content; the problem is front-loading: the vague 'Parts and revision for precise editing' leads while the concrete coordinate-frame and selection details follow. Size is reasonable for the parameter complexity, so waste is low.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values (rows, transform matrix, mirror bitmask) need not be re-explained, and the description reasonably covers the select/coordinate semantics. It is still incomplete on pagination (limit/offset defaults 100/0), the required 'design' parameter, and how this differs from search_parts, which matters given the crowded sibling set.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the most complex parameter well ('select: ids, name, definition or inclusive bounds [[min_x,min_y,min_z],[max_x,max_y,max_z]]'), including the nested bounds shape. However, 'design', 'limit', and 'offset' receive no explanation at all, leaving three of four parameters undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Parts and revision for precise editing,' which identifies the resource (parts of a design) but the phrasing is oblique rather than a clean verb+resource statement. It never distinguishes itself from close siblings such as search_parts or get_part_definition, so an agent cannot route confidently from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied via 'for precise editing.' There is no explicit when-to-use, no when-not-to-use, and no named alternative among the many part-related siblings. The one actionable guideline ('select an id to target the whole component') is a constraint, not a routing rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_vehicle_validationA

Record human-observed pass/fail/skipped results with evidence and measurements. Refuses a changed export or mismatched hash. All seven checks need pass evidence for a passed run.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes
testerYes
game_versionYes
observationsYes
export_sha256Yes

Output Schema

ParametersJSON Schema
NameRequiredDescription
checksYes
designYes
run_idYes
statusYes
testerYes
revisionYes
created_atYes
part_countYes
game_versionYes
vehicle_nameYes
vehicle_pathYes
export_sha256Yes
evidence_levelYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does add real value: it discloses a guard behavior ("Refuses a changed export or mismatched hash") and an integrity rule (all seven checks need pass evidence). What is missing is idempotency/re-call behavior (can a run be overwritten?), permission requirements, and partial-submission semantics, so it is not fully complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with zero filler, front-loaded with the core action and followed by the two hard constraints. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description covers the key rejection and evidence rules. However, for a complex nested-observation mutation tool with 0% schema coverage and no annotations, the absence of any explanation for run_id/tester/game_version and the overwrite/idempotency behavior leaves gaps an agent may need.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it only partially does: "evidence and measurements" maps to the Observation fields, "checks" maps to the enum check list, and "mismatched hash" alludes to export_sha256. It says nothing about run_id, tester, or game_version, leaving three of five required parameters undocumented anywhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ("Record") and resource ("human-observed pass/fail/skipped results with evidence and measurements"), which clearly distinguishes it from siblings like prepare_vehicle_validation and get_vehicle_validation. It stops short of explicitly naming those siblings, but an agent can identify the tool's role without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by "human-observed" results and the rule that a passed run needs pass evidence, but there is no explicit guidance on when to use this versus prepare_vehicle_validation or get_vehicle_validation. No prerequisites about calling prepare first are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repair_vehicleB

Preview selected repairs with overlays and before/after preflight. commit=true applies the whole batch as one undo step. Revision/plan guards reject stale suggestions and conflicting repairs.

ParametersJSON Schema
NameRequiredDescriptionDefault
commitNo
designYes
plan_idYes
revisionYes
finding_idsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
afterYes
beforeYes
revisionYes
committedYes
base_revisionYes
resolved_finding_idsYes
remaining_finding_idsYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses atomicity ('whole batch as one undo step') and validation behavior ('guards reject stale suggestions and conflicting repairs'), which are real behavioral traits. It still omits permissions, exactly what gets mutated, and failure modes for rejected findings, so it is only partially complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the preview default and then the commit behavior. Every sentence adds information and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema exists, so return format need not be described. But for a 5-param, 4-required mutation tool with no annotations and zero schema coverage, the description under-documents the inputs even though it handles the commit/undo semantics well.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% across 5 params, so the description must compensate. It explains only 'commit' (the optional param) and hints at 'revision' and 'plan_id' via the guard sentence; 'design' and 'finding_ids' are left entirely undefined, leaving a large gap for a tool where 4 params are required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('repair_vehicle' previewing/applying selected repairs) and clarifies the preview-vs-commit distinction. It does not name the nearby siblings plan_vehicle_repairs or preflight_vehicle, so an agent cannot fully separate this from the planning and preflight tools without inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage through commit=false preview vs commit=true apply, which is useful. However, it never states when to pick this over plan_vehicle_repairs (generate suggestions) or preflight_vehicle (validate), leaving the workflow routing implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

route_connectionsA

Preview/commit actual pipes between transmission faces, preserving access and structure. Each route: {from:{part_id,surface_index},to:{part_id,surface_index},bounds?:[[lo],[hi]], waypoints?:[[x,y,z],...],name?:str,pipe_style?:auto|exposed|enclosed,through_blocks?:[part_id,...]}. Bounds/waypoints use integer blocks. Routes minimize length, then bends. Default auto uses enclosed pipes for explicitly selected through_blocks; other cells use exposed pipes. through_blocks replaces only selected unconfigured/unlinked 01_block parts, retaining paint. All selected blocks must lie on the route (use waypoints if needed); removal restores them. Other structure and access are preserved; routes never join unrelated pipe networks. Clutch/gearbox ports may use several routes; create an explicit T-piece for a branch. Remove an added route with {op:'remove',route_id} from query_connections, then reroute.

ParametersJSON Schema
NameRequiredDescriptionDefault
commitNo
designYes
routesYes
revisionYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries full disclosure burden and does so well: route optimization order (length then bends), auto pipe-style resolution, that through_blocks replaces only selected unconfigured/unlinked 01_block parts while retaining paint, that removal restores them, that access and structure are preserved, and that routes never merge unrelated networks. It omits permission/auth requirements and return/output behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first clause, and the dense technical sentences each convey a concrete rule rather than restating the name. The run-on accumulation of pipe-style, block-replacement, and branch rules makes it harder to scan than it could be, but little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity nested-schema tool with no annotations and no output schema, the description covers routing behavior, block replacement side effects, preservation guarantees, and the removal workflow convincingly. The main completeness gaps are revision/concurrency semantics and any indication of success/failure output, but the core call-correctness information is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and it does so substantially for the complex 'routes' object (from/to face endpoints, bounds, waypoints in integer blocks, pipe_style semantics, through_blocks behavior, and the remove op with route_id). However the top-level params 'design' and 'revision' are never explained and 'commit' is only implied by 'Preview/commit', leaving a real gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Preview/commit actual pipes between transmission faces,' immediately telling the agent this creates pipe routes rather than editing parts or querying links. It names the removal workflow via query_connections, but does not clearly differentiate itself from edit_connections or query_connections for the creation case.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Contains embedded guidance (use waypoints to force blocks onto a route, create an explicit T-piece for a branch, remove via query_connections then reroute), which is useful procedural context. However it never states when to prefer this tool over siblings like edit_connections or query_connections, leaving the selection choice largely inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_hullA

Save a hull into the Stormworks vehicles folder so it can be loaded at a workbench.

name: vehicle name shown in the game's Load menu. The spec is also stored as a design of the same name. overwrite: replace an earlier design saved by this tool under the same name. Vehicles not made by this tool are never overwritten. design/patch: as in preview_hull, e.g. save_hull(name="Iowa", design="Iowa").

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
specNo
patchNo
designNo
presetNo
overwriteNo
spec_pathNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose useful behavior: the save location, that the spec is stored as a design of the same name, and the safety guarantee that non-tool vehicles are never overwritten. It doesn't cover permissions, failure modes, or return format, leaving some gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose in one sentence, then lists a few parameter notes without padding. The 'as in preview_hull' pointer is efficient, though the parameter section feels lightly sketched rather than fully structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be explained, and the tool's core purpose is clear. However, for a 7-parameter mutation tool with 0% schema coverage and no annotations, leaving three parameters undocumented is a notable gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 7 params, so the description must compensate and only partially does. It documents name, overwrite, and design/patch (by reference to preview_hull), but spec, preset, and spec_path are undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Save a hull') plus destination and downstream purpose ('into the Stormworks vehicles folder so it can be loaded at a workbench'). This clearly distinguishes it from siblings like store_design, load_design, and preview_hull.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the overwrite condition (only replaces designs saved by this tool; vehicles not made by this tool are never overwritten) and defers design/patch semantics to preview_hull, implying a preview-then-save workflow. No explicit 'when-not-to-use' or named alternative is given, keeping it below a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_vehicleC

Save any draft. Imported originals are never overwritten, even with overwrite=true.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
designYes
overwriteNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose one genuinely valuable non-obvious trait: imported originals are protected from overwrite even when overwrite=true. However, it is silent on permissions, whether an existing same-named draft is overwritten by default, and whether the operation is destructive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with no filler; the save action comes first and the exception second. It is efficient, though the brevity comes partly at the cost of missing information rather than pure concision.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists, so return values need not be explained, but for a mutation tool with no annotations and 0% parameter coverage the description is too thin. It leaves the identity of a "draft," the meaning of name/design, and default overwrite behavior unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and none of the three parameters are documented in the schema. The description adds meaning only to overwrite (its limits on imported originals), while name and design remain completely undescribed, so it does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb "Save" plus the object "any draft" gives an approximate purpose, but the tool is named save_vehicle and the description never says "vehicle" or clarifies what constitutes a draft. It also fails to differentiate from close siblings such as save_hull, store_design, or import_vehicle, so an agent cannot tell which save path applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus save_hull/store_design, no prerequisites, and no distinction between saving a new draft and updating one. The overwrite caveat hints at one boundary case but is a behavioral note rather than usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_land_partsB

Installed components useful for land builds, with footprints and actual connection ports. Categories: wheels, tracks, lights, controls, propulsion, transmission, power, fuel, cooling, body, logic, utility. Use get_part_definition/get_part_orientation for details. Family filters are a convenience; search_parts exposes the full installed catalogue.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
offsetNo
searchNo
categoryNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that results include footprints and actual connection ports, which is useful behavioral context, but it says nothing about pagination, result limits, or the read-only nature of the operation. Adequate but thin for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the what and how in the first sentence, then follows with scope and routing. The category enumeration is somewhat redundant with the schema enum, but overall the text is efficient and each sentence contributes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the purpose is clear. However, for a search tool with four undocumented parameters and no annotations, the description omits pagination and filtering semantics that an agent would need to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 4 parameters. The description lists categories, but that list merely repeats the enum already present in the schema and adds no new meaning. The limit, offset, and search parameters are left entirely undocumented, so the description does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: it searches installed components for land builds, and it differentiates itself from siblings by naming search_parts (full catalogue) and get_part_definition/get_part_orientation (details). The scope 'land builds' and the category list make the resource concrete. It lacks an explicit statement of the return shape but the resource is unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It routes the agent to alternatives: get_part_definition/get_part_orientation for details, and it clarifies the relationship to search_parts (family filters vs the full catalogue). It conveys the intended context (land builds) without spelling out explicit when-not-to-use conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_partsB

Search installed part definitions by name/id; returns actual sizes, mass and pagination.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
offsetNo
searchNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose a meaningful behavioral constraint ('installed' parts only, not all definitions) and implies a safe read via 'search', but is silent on permissions, error behavior, or whether results are fuzzy vs exact matches.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the action and key, then appends scope and return information. No filler, though mentioning returned fields is slightly redundant given an output schema exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a search tool with an output schema present, the description need not explain return values, and it covers scope and the search key. It still leaves sibling disambiguation, pagination semantics, and read-only confirmation unaddressed, which matters in this crowded sibling set.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the schema documents none of the limit/offset/search params. The description compensates partially by naming the search key ('name/id') and flagging pagination, which effectively maps to all three parameters conceptually, but gives no formats, defaults, or matching behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clear verb+resource ('Search installed part definitions') with the lookup key ('by name/id') specified, so an agent knows exactly what the tool retrieves. However, it offers no differentiation from close siblings like query_parts, get_part_definition, and search_land_parts, which an agent could easily confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No statement of when to use this tool versus query_parts, get_part_definition, or search_land_parts, all of which appear to overlap on part lookup. The 'installed' qualifier hints at scope but no explicit context or exclusion is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

store_designA

Keep a design spec on the server without writing a vehicle, so later calls can send only changes. Typical loop: store_design("Iowa", spec=...) once, then preview_hull(design="Iowa", patch=[...]) to try a change and store_design("Iowa", design="Iowa", patch=[...]) to keep it. Replaces any stored design of that name; save_hull(name, design=name) writes the vehicle when ready. spec_path: read the spec from a .json file instead of sending it. Stored designs are plain JSON files that every call re-reads, so editing one on disk takes effect at once.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
specNo
patchNo
designNo
presetNo
spec_pathNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose key behavioral traits: it replaces any stored design of that name (destructive overwrite), and stored designs are plain JSON files re-read on every call so on-disk edits take effect immediately. It stops short of permissions, error behavior, or concurrency concerns, so not a full 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then proceeds through the loop example and the two footnotes (spec_path, on-disk editing). The inline examples are dense but each earns its place by showing the calling pattern; a small amount of repetition in the loop could be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the description covers the stateful persistence model that is the crux of this tool. The only real omission is the 'preset' parameter's role and any note on name collisions beyond the overwrite warning.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and largely does: it explains spec (initial payload), patch (only changes), design (reference to a stored design), spec_path (read spec from a .json file instead of sending it), and name via the examples. Only 'preset' is left unexplained, and the unique semantics of patch/design/spec_path go well beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: storing a design spec on the server so later calls can send only changes. It explicitly separates this tool's role from siblings by naming preview_hull (try a change), save_hull (write the vehicle), and load_design (read it back), so an agent can route correctly without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit typical loop with concrete invocations: store_design once, then preview_hull with a patch to try changes, then store_design with a patch to persist, and save_hull when ready. When-to-use is demonstrated rather than implied, and alternatives are named with their distinct roles.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

suggest_hull_blocksB

Rank real block slope families by an outward x/y/z normal. Read the selection rules: the angle alone cannot choose pyramid vs inverse, location, or a watertight joint.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
normalYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose a meaningful behavioral caveat (the normal angle alone is insufficient for several decisions, so the tool only ranks candidates), but it says nothing about read-only vs mutating behavior, side effects, or what the ranked output represents.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core action front-loaded and no filler. The only weakness is the dangling reference to 'the selection rules,' which assumes external context the agent may not have.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. The remaining gap is that the description defers to unnamed 'selection rules' and omits any treatment of the limit parameter, leaving an agent without enough to use the tool confidently beyond the basic call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies 'normal' as an outward x/y/z direction, which adds real meaning, but the optional 'limit' parameter (default 6) is never mentioned in the description or schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Rank real block slope families') and names the input that drives the ranking (an outward x/y/z normal). It is clear what the tool produces, though it does not explicitly distinguish itself from near neighbors like preview_hull or list_hull_presets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The sentence about the angle alone being insufficient to choose pyramid vs inverse, location, or watertight joint gives useful context on the tool's limits. However, it points to 'the selection rules' without providing them and never names an alternative tool or a concrete when-to-use condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

undo_editsC

Restore the previous committed edit batch. Keeps up to ten batches of undo history.

ParametersJSON Schema
NameRequiredDescriptionDefault
designYes
revisionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does add one genuinely useful trait — undo history is bounded to ten batches, implying failure after enough edits — but says nothing about permissions, whether the restore is itself reversible, or what happens when no history exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with no filler, and the core action is front-loaded before the history-limit detail. Brevity here tips toward under-specification, but structurally the text is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a state-mutating undo tool with no annotations and two wholly undocumented required parameters, the description leaves key facts missing: which design's edits are affected, what a revision argument means, and failure conditions. The output schema covers return values, but the input and behavioral sides are under-described.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters ('design', 'revision') are required with 0% schema description coverage, and the description never mentions them. It is unclear whether 'revision' identifies the batch to undo, the target revision, or a revision string format, so it fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Restore the previous committed edit batch'), which is more than a tautology and lets an agent grasp the operation. However, it never ties itself to the sibling edit tools (edit_parts, edit_connections) whose batches it presumably reverts, so differentiation from siblings depends on inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to invoke this versus alternatives, no prerequisite (e.g., an existing edit batch to revert), and no mention of edit_parts/edit_connections as the operations it undoes. Usage is only implied by the word 'undo'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 39 tool updatesv0.1.1
    • Addedanalyze_hull
    • Addedanalyze_land_vehicle
    • Addedanalyze_vehicle
    • Addedapply_vehicle_assembly
    • Addedcheck_seal
    • Addedcomplaint
    • Addedcreate_land_vehicle
    • Addededit_connections
    • Addededit_parts
    • Addedfind_land_vehicles
    • Addedget_calibration_observations
    • Addedget_complaint
    • Addedget_part_definition
    • Addedget_part_orientation
    • Addedget_runtime_status
    • Addedget_vehicle_validation
    • Changedhull_design_guide1 field changed
      • addedInput schema / properties / topic
        Added value: +{
        +  "default": "full",
        +  "enum": [
        +    "full",
        +    "workflow",
        +    "units",
        +    "spec",
        +    "interior",
        +    "archetypes",
        +    "style",
        +    "limits",
        +    "staged",
        +    "building",
        +    "testing",
        +    "smoothing",
        +    "edits",
        +    "topics"
        +  ],
        +  "title": "Topic",
        +  "type": "string"
        +}
    • Addedimport_vehicle
    • Changedinspect_view2 fields changed
      • addedInput schema / properties / body_id
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Body Id"
        +}
      • addedInput schema / properties / source
        Added value: +{
        +  "default": "vehicles",
        +  "title": "Source",
        +  "type": "string"
        +}
    • Addedland_vehicle_guide
    • Addedlist_complaints
    • Addedlist_vehicle_assemblies
    • Changedopen_in_viewer3 fields changed
      • addedInput schema / properties / body_id
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Body Id"
        +}
      • addedInput schema / properties / diagnostics
        Added value: +{
        +  "default": false,
        +  "title": "Diagnostics",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / source
        Added value: +{
        +  "default": "vehicles",
        +  "title": "Source",
        +  "type": "string"
        +}
    • Addedplan_vehicle_repairs
    • Addedpreflight_vehicle
    • Addedprepare_vehicle_validation
    • Changedpreview_game_vehicle2 fields changed
      • addedInput schema / properties / body_id
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Body Id"
        +}
      • addedInput schema / properties / source
        Added value: +{
        +  "default": "vehicles",
        +  "title": "Source",
        +  "type": "string"
        +}
    • Addedpreview_vehicle
    • Addedpreview_vehicle_diagnostics
    • Addedquery_connections
    • Addedquery_parts
    • Addedrecord_vehicle_validation
    • Addedrepair_vehicle
    • Addedroute_connections
    • Addedsave_vehicle
    • Addedsearch_land_parts
    • Addedsearch_parts
    • Addedsuggest_hull_blocks
    • Addedundo_edits
  2. 15 tool updatesv0.1.0
    • First observeddeck_profile
    • First observedget_hull_spec
    • First observedhull_design_guide
    • First observedinspect_view
    • First observedlist_designs
    • First observedlist_game_vehicles
    • First observedlist_hull_presets
    • First observedlist_workbenches
    • First observedload_design
    • First observedopen_in_viewer
    • First observedpreview_game_vehicle
    • First observedpreview_hull
    • First observedpreview_interior
    • First observedsave_hull
    • First observedstore_design

TDQS

C2.9/5.0

Scored across 50 tools

Disambiguation3/5

Several tool clusters overlap: analyze_vehicle vs analyze_land_vehicle vs analyze_hull, preview_vehicle vs preview_game_vehicle vs inspect_view vs open_in_viewer, and search_parts vs search_land_parts vs query_parts. The detailed descriptions clarify contexts, but an agent could still misselect within these families, especially for preview and analysis tasks.

Naming Consistency4/5

Most tools follow a snake_case verb_noun pattern (list_*, get_*, query_*, edit_*, preview_*, save_*). A few deviate by using noun phrases or unusual forms (complaint, deck_profile, hull_design_guide, land_vehicle_guide, open_in_viewer), but the overall convention is readable and predictable.

Tool Count2/5

50 tools is heavy for a single MCP server, even for a complex domain like Stormworks vehicle design. Many tools could be consolidated or split into sub-servers (e.g., hull design, land design, part editing, validation, previews). The count exceeds the practical sweet spot and increases selection burden.

Completeness4/5

The surface covers hull and land vehicle design, part editing, connection routing, validation lifecycle, previews, and complaint reporting very thoroughly. Minor gaps exist such as no explicit delete/rename design tool and no aircraft-specific design guide, but core workflows are fully supported.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables Claude to create 3D-printable CAD models using build123d, with tools for modeling, modification, analysis, and publishing to platforms like Thingiverse and GitHub.
    14
    Creative Commons Attribution Non Commercial No Derivatives 4.0 International
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables Claude to create and manipulate SolidWorks CAD models through natural language commands, automating part creation, sketching, and extrusion.
    10
    -
  • F
    license
    B
    quality
    B
    maintenance
    Enables designing Minecraft structures through Claude Code and exporting them as Sponge Schematic v3 (.schem) files.
    14
    -