Skip to main content
Glama

carmaker-mcp

An MCP server that lets an AI agent drive IPG CarMaker on your computer: start a session, run test runs, read the results, change vehicle data and Simulink controller parameters, and undo what it changed.

Unofficial community project, not affiliated with or endorsed by IPG Automotive or MathWorks. You need your own CarMaker and MATLAB/Simulink licences. This repository contains no IPG or MathWorks code or data.

Status: alpha. Developed and tested on one setup (CarMaker 14.1.1, MATLAB R2024b, Windows 11). docs/verification.md lists what has been checked against real software and what has not.

Contents: What it does · What it does not cover · Requirements · Setup · With the MATLAB MCP server · Settings · Tools · Safety and privacy · Troubleshooting · How it works

What it does

Things you can ask an agent once the server is connected:

  • "Run the braking test run and tell me the stopping distance."

  • "Set Kp_yaw to 1.5 in the controller model, run the slalom again and compare the yaw rate."

  • "Try the lane change with a vehicle mass of 280, 300 and 320 kg and tabulate the lateral acceleration."

  • "The last run aborted. Why?"

  • "Undo everything you changed."

It works in two ways, and the agent picks the one that fits the vehicle:

MATLAB-connected session

Standalone run

Use it for

vehicles whose controller is a Simulink model (CarMaker for Simulink)

vehicles with CarMaker's built-in controllers

Needs

MATLAB with your model, and the CarMaker GUI (the server can start both)

only CarMaker; no MATLAB, no window

More worked examples: docs/examples.md.

Related MCP server: COMSOL MCP Server

What it does not cover

The server was written for vehicle-dynamics and controller work in a Formula Student team, where every run was either a standalone run or a CarMaker for Simulink run. CarMaker can do much more than that, and the rest is outside what this server does or has been tried with.

Not supported (there are no tools for it):

  • Other platforms and products: Linux; TruckMaker and MotorcycleMaker; CarMaker HIL and CarMaker Office Extended (real-time hardware, Fail Safe Tester, bus interfaces such as CAN, FlexRay, SOME/IP, XCP).

  • CarMaker's own test automation: Test Manager and test series, test reports, parallel (HPC) execution, batch mode with start-up files. The server runs its own parameter studies instead, one run after another.

  • Building scenarios: the Scenario Editor and roads, traffic, sensors, the environment, OpenSCENARIO import. cm_edit changes keys in files that exist; it does not author these things.

  • Most visualisation: Movie NX, IPGControl, Instruments, video export, Model Check. Of IPGMovie, only single pictures.

  • Other result formats: only .erg files are read, not MDF or ASCII.

  • Rebuilding the simulation program: models in C code or generated with Simulink Coder.

  • Remote machines: CarMaker and MATLAB must run on the computer the server runs on.

May work, never tested:

  • FMUs, CarMaker exported as an FMU, third-party tyre models and other co-simulation tools.

  • Encrypted or protected data files.

  • CarMaker and MATLAB versions other than the ones named above; more than one MATLAB session or CarMaker GUI at a time.

  • Several standalone runs in parallel, pausing and resuming a standalone run, parameter studies on standalone runs.

  • Test runs with traffic, sensors or driver-assistance functions. Nothing stands in their way, but every check so far used vehicle-dynamics runs without them.

The state of each tool is in docs/verification.md. Reports from other setups are welcome.

Requirements

  • Windows 10 or 11

  • IPG CarMaker with its Python API (developed on 14.1.1)

  • For the MATLAB-connected session: CarMaker for Simulink and MATLAB R2022b or newer, as far as your CarMaker supports it (see the table in step 1; only R2024b has been tested)

  • uv (it fetches the server and its Python environment; pip also works, see docs/setup.md)

  • An MCP client: Claude Code, Claude Desktop, VS Code with GitHub Copilot, Codex, Antigravity, Cursor, ...

Setup

1. Check your machine

uvx carmaker-mcp doctor

This finds your CarMaker and MATLAB installs, tells you what does not fit and how to fix it, and prints the uvx arguments for your machine. They look like this:

uvx --python 3.12 --with matlabengine==24.2.* carmaker-mcp

The two extra arguments matter, which is why the tool works them out for you:

  • --python 3.12 selects a Python version that both your CarMaker and your MATLAB release work with. Without it uvx may pick a newer Python that neither can use.

  • --with "matlabengine==24.2.*" adds the MATLAB engine package for your MATLAB release. Leave it out if you only use standalone runs.

Both depend on the MATLAB release, because MathWorks publishes one engine package per release and each installs only on some Python versions:

MATLAB

Python

uvx arguments

Bundle on the release page

R2024b

3.12

--python 3.12 --with "matlabengine==24.2.*"

...-R2024b-py312.mcpb

R2024a

3.11

--python 3.11 --with "matlabengine==24.1.*"

...-R2024a-py311.mcpb

R2023b

3.11

--python 3.11 --with "matlabengine==23.2.*"

...-R2023b-py311.mcpb

R2023a

3.10

--python 3.10 --with "matlabengine==9.14.*"

...-R2023a-py310.mcpb

R2022b

3.10

--python 3.10 --with "matlabengine==9.13.*"

...-R2022b-py310.mcpb

none (standalone runs only)

3.12

--python 3.12

...-standalone-py312.mcpb

Only the R2024b row has been run against real software. The others follow from the engine packages' published requirements and are untested; reports are welcome. MATLAB R2022a and older cannot be used for the MATLAB-connected session, because their engine packages need a Python older than this server supports (standalone runs still work).

To look at the tools without CarMaker installed: uvx carmaker-mcp --mock.

2. Register the server in your MCP client

Every client needs the same three things: the command uvx, the arguments from step 1, and your settings as environment variables. The settings you will normally set:

  • CM_PROJECT: your CarMaker project folder.

  • CM_MATLAB_INIT: if you normally start work by running a MATLAB script of your project (one that adds folders to the path and opens the model), name it here so that the agent runs it too.

  • CM_MODEL: the Simulink model the agent opens when it starts a session. A name on the MATLAB path, or a path relative to the project's src_cm4sl folder. If you leave it out, the agent is shown the models it finds and picks one or asks you.

Pick your client below. The examples assume MATLAB R2024b and Python 3.12; replace the arguments with the ones doctor printed. Or let the server write the entry for you, with every CM_ variable set in your shell filled in:

uvx carmaker-mcp config --client vscode     # or claude-code, claude-desktop, codex, antigravity, cursor

Run in a terminal:

claude mcp add carmaker --env CM_PROJECT="C:\CM_Projects\my-project" --env CM_MODEL="MyModel" -- uvx --python 3.12 --with "matlabengine==24.2.*" carmaker-mcp

Add further settings with more --env NAME="value" before the --. With --scope project the entry goes into a .mcp.json in the current folder, where you can edit its env block later; otherwise remove and add it again (claude mcp remove carmaker). claude mcp list shows whether the server connects.

One-click bundle. Download the .mcpb file for your MATLAB release from the latest release (the table in step 1 names them) and drag it into Settings > Extensions. Fill in the project folder, the model and, if your project has one, its MATLAB setup script in the dialog; Configure changes them later. The bundle holds the server but not its Python environment: Claude Desktop builds that with uv on first use, so uv must be installed.

By hand. Settings > Developer > Edit Config opens claude_desktop_config.json. Add the entry and restart Claude Desktop:

{
  "mcpServers": {
    "carmaker": {
      "command": "uvx",
      "args": ["--python", "3.12", "--with", "matlabengine==24.2.*", "carmaker-mcp"],
      "env": {
        "CM_PROJECT": "C:\\CM_Projects\\my-project",
        "CM_MODEL": "MyModel"
      }
    }
  }
}

Create .vscode/mcp.json in your workspace (for all workspaces: MCP: Open User Configuration in the Command Palette):

{
  "servers": {
    "carmaker": {
      "type": "stdio",
      "command": "uvx",
      "args": ["--python", "3.12", "--with", "matlabengine==24.2.*", "carmaker-mcp"],
      "env": {
        "CM_PROJECT": "C:\\CM_Projects\\my-project",
        "CM_MODEL": "MyModel"
      }
    }
  }
}

Click Start above the entry, then use Copilot Chat in agent mode.

Run in a terminal:

codex mcp add carmaker --env CM_PROJECT="C:\CM_Projects\my-project" --env CM_MODEL="MyModel" -- uvx --python 3.12 --with "matlabengine==24.2.*" carmaker-mcp

or edit C:\Users\<username>\.codex\config.toml:

[mcp_servers.carmaker]
command = "uvx"
args = ["--python", "3.12", "--with", "matlabengine==24.2.*", "carmaker-mcp"]
env_vars = ["WINDIR"]

[mcp_servers.carmaker.env]
CM_PROJECT = "C:\\CM_Projects\\my-project"
CM_MODEL = "MyModel"

Codex starts servers with a reduced environment; env_vars = ["WINDIR"] passes on a Windows variable that MATLAB's libraries need.

In the agent panel open the ... menu, then MCP Servers > Manage MCP Servers > View raw config. Add the entry to the mcp_config.json that opens, save, and press Refresh in the server list:

{
  "mcpServers": {
    "carmaker": {
      "command": "uvx",
      "args": ["--python", "3.12", "--with", "matlabengine==24.2.*", "carmaker-mcp"],
      "env": {
        "CM_PROJECT": "C:\\CM_Projects\\my-project",
        "CM_MODEL": "MyModel"
      }
    }
  }
}

Cursor: create .cursor/mcp.json in your project (or ~/.cursor/mcp.json for all projects) with the same content as shown for Antigravity. Most other clients use that mcpServers format too.

In JSON and TOML, write Windows paths with double backslashes or with forward slashes. So far only Claude Code has been used with this server; the other entries follow each client's documented format.

3. Start a session

Ask the agent, for example "start CarMaker and run the braking test run". It opens MATLAB, your model and the CarMaker GUI as far as they are not open yet, and never closes anything.

If you would rather open MATLAB and CarMaker yourself, run this once per MATLAB session (or put it into startup.m) so that the server can attach:

matlab.engine.shareEngine('cm_mcp')

Standalone runs need neither.

Using it with the MATLAB MCP server

For CarMaker for Simulink work it is worth connecting MathWorks' MATLAB MCP Core Server as well. It is optional, and standalone runs do not need it. The two servers do different jobs:

Server

Use it for

carmaker-mcp

The CarMaker side: session, test runs, run control, results and project data, and controller parameters with a change log and undo

MATLAB MCP server

The MATLAB side: writing, checking and running MATLAB code (post-processing scripts, plots) and looking into the Simulink model

Things to know when you use both:

  • Let both work on the same MATLAB session (set the MATLAB server up to use your open MATLAB, see its README), so that they see the same model and workspace. This is how the pair was used during development.

  • MATLAB does one thing at a time. While a model compiles or a simulation runs, calls through the MATLAB server wait. This server's run tools are built around that (time-limited waits that are simply repeated).

  • Start and stop runs with this server's tools, so that it can follow the run (result file, errors, storage mode).

  • What is changed through the MATLAB server is not in this server's change log, and cm_revert_all does not undo it.

Settings

All settings are environment variables in the client entry, set like CM_PROJECT and CM_MODEL above.

Common

Variable

Default

Meaning

CM_PROJECT

read from the running CarMaker GUI

CarMaker project folder. The server writes nowhere else. Needed to start a session from scratch

CM_MODEL

none

Simulink model to open when a session is started

CM_MATLAB_INIT

none

Your project's own MATLAB setup script (in src_cm4sl, or a full path), run once when a session is started. Use it if you normally run a script that adds folders to the path and opens the model: without it the model may not compile

CM_POPUP_TIMEOUT

not set

Seconds after which CarMaker pop-ups answer themselves with their default choice, so that a question cannot block a run. Not set: they wait for your click. See Safety

CM_DISABLE

none

Tool groups to switch off, to give the model a shorter tool list: standalone, study, matlab, movie

CM_ENABLE

none

Optional tools to switch on: tcl (raw Tcl in the CarMaker GUI), experimental (add and delete Simulink blocks)

Advanced

Variable

Default

Meaning

CM_HOME

newest C:\IPG\carmaker\win64-*

CarMaker install folder

CM_MATLAB_EXE

newest installed MATLAB that your CarMaker supports

MATLAB executable used to start a session

CM_MATLAB_DIR

<project>/src_cm4sl

Folder MATLAB starts in (where cmenv.m is)

CM_MATLAB_SESSION

cm_mcp

Name under which MATLAB shares its engine

CM_RESULT_DIRS

none

Extra folders to search for result files

CM_STATE_DIR

%LOCALAPPDATA%\carmaker-mcp

Where backups, the change log and server.log are kept (never inside the project)

CM_ENGINE_TIMEOUT

30

Seconds before a call into MATLAB is given up

CM_LOG_LEVEL

INFO

Detail of server.log

Tools

Each tool tells the client whether it is read-only, changes state, or is destructive, so that clients can ask for confirmation where it matters. Every parameter is described in docs/tools.md.

Session and runs

Tool

What it does

cm_session_start

Start MATLAB, the model and the CarMaker GUI as far as they are missing

cm_doctor

Check the setup and say how to fix what is wrong

cm_status

Simulation state, active model, project folder

cm_load_testrun

Load a test run into the CarMaker GUI

cm_start_sim, cm_stop_sim

Start and stop the simulation; a run writes a result file by default

cm_wait_end

Wait for the end of the run; returns end status, simulation time, distance and the result file

cm_live

Read quantities while the simulation runs

cm_dva_write, cm_dva_release

Overwrite a quantity during a run (Direct Variable Access)

cm_log

Read CarMaker's session log, for example after an aborted run

cm_popups, cm_popup_timeout

See what the CarMaker GUI asked or reported, and let pop-ups answer themselves

Results

Tool

What it does

cm_results_list

Newest result files

cm_results_summary

First, last, min, max and mean of quantities in a result file; search for quantity names

cm_results_read

Time series from a result file

cm_output_quantities, cm_output_quantities_edit

See which quantities runs write to result files, and add or remove some

cm_movie_open, cm_movie_snapshot

Open IPGMovie and get a picture of the 3D view: any moment of the last run, or of a result file

Project data and parameters

Tool

What it does

cm_list, cm_read

List and read test runs, vehicles, drivers, tyres and other project files

cm_edit, cm_clone

Change keys in a project file (only the edited lines change), or copy a file

cm_list_workspace_vars, cm_get_workspace_var, cm_set_workspace_var

MATLAB base workspace and Simulink model workspace, where controller parameters usually are

cm_model_get, cm_model_set, cm_model_save

Simulink block and model parameters; save the model

cm_model_logs_save

After a run, save what Simulink logged (logged signals, To Workspace blocks) and both workspaces to a MAT file. save_logs on cm_start_sim and cm_study_start does it automatically for every run

cm_study_start, cm_study_status, cm_study_cancel

Run one test run with several parameter sets and tabulate the results, without changing any file

Standalone runs (no MATLAB)

Tool

What it does

cm_standalone_launch

Start an independent CarMaker process and run a test run on it

cm_standalone_status, cm_standalone_wait_end, cm_standalone_results

Follow the run and get its result files

cm_standalone_control, cm_standalone_dva_write

Pause, resume or stop; overwrite a quantity

cm_standalone_servers, cm_standalone_attach

Find and attach to CarMaker programs that are already running

cm_standalone_close

Stop the process

History

Tool

What it does

cm_changelog

What this session changed, with old and new values

cm_revert_all

Undo this session's changes

cm_restore

Restore the files of an earlier session from its backups

The server also provides resources (carmaker://guide, carmaker://status, carmaker://changelog, carmaker://log) and three prompts (run_and_summarise, compare_settings, undo_session).

Safety and privacy

This server can start simulations and edit your project files and the live Simulink model in place. Connect it only to agents you trust.

What protects your work:

  • Every file is backed up before its first change, every change is logged with its old value, and cm_revert_all undoes a session.

  • Files are only written inside the project folder, and only while the simulation is idle. No tool deletes project files.

  • Raw Tcl and Simulink structure edits are off unless you switch them on with CM_ENABLE.

  • CarMaker pop-ups wait for your click by default. If you set CM_POPUP_TIMEOUT, a question such as "Vehicle not saved. All changes will be lost. OK to continue?" is answered with its default, which discards unsaved changes in the CarMaker GUI.

Runs write result and log files into the project's SimOutput folder; a revert does not remove those.

The server collects no telemetry and makes no network connections of its own. It only talks to MATLAB and CarMaker on your machine. Details: docs/safety.md.

Troubleshooting

Run uvx carmaker-mcp doctor first: it names most setup problems together with their fix.

Symptom

Cause and fix

"No shared MATLAB session"

MATLAB is closed or has not shared its engine. Let the agent start the session, or run matlab.engine.shareEngine('cm_mcp') in MATLAB

uvx fails while installing matlabengine

The pin must match your MATLAB release, and that MATLAB must be installed. Use the arguments doctor prints

"CarMaker ships no cmapi for Python 3.x"

Add --python 3.12 (or another version the message lists)

A tool call hangs at load or start

The CarMaker GUI is showing a question. Answer it, or see CM_POPUP_TIMEOUT

"too many licenses in use" on a standalone run

An open MATLAB with CarMaker for Simulink holds the licence. Close that model or MATLAB

The wait tool returns finished: false

Not an error: waits are limited to 45 s per call and are simply repeated

More: docs/troubleshooting.md. The server's own log is server.log in the state folder.

How it works

  • MATLAB-connected session: the server attaches to your MATLAB through the MATLAB Engine for Python and sends commands to the CarMaker GUI through CarMaker for Simulink's own cmguicmd (the GUI's Tcl / ScriptControl interface).

  • Standalone runs: it uses IPG's Python API cmapi, loaded from your CarMaker install, to start and control a separate CarMaker process.

  • Project files are edited by a byte-preserving editor for CarMaker's infofile format, and results are read by a NumPy reader for .erg files (checked against MATLAB's cmread).

Nothing of IPG's or MathWorks' is bundled or redistributed.

Documentation and contributing

Related work: pycarmaker talks to the command port of a standalone CarMaker program. This project also covers CarMaker for Simulink, uses IPG's own interfaces, and shares no code with it.

Licence

MIT, see LICENSE.

Available Tools

46 tools
cm_changelogChange logB
Read-only

Recorded changes (file keys, workspace variables, model parameters, GUI settings) with old and new values.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoNewest entries to return
sessionNoSession id (default: this session)

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds useful context on what kinds of changes are tracked and that old/new values are both present, but it says nothing about ordering beyond the schema's 'newest entries' note, pagination, or whether the log is session-scoped.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single informative sentence with no filler, and the content scope is front-loaded. It is a noun fragment without a verb, which is efficient but slightly less directive than a full statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary, and both optional parameters are documented in the schema. The main omission is usage context and the relationship to the revert/restore siblings, but for a read-only viewer the description is nearly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters (limit, session), so the schema already carries the parameter meaning. The description adds no syntax or format detail beyond what the schema provides, meeting the baseline-3 expectation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (recorded changes) and enumerates the four content categories it covers (file keys, workspace variables, model parameters, GUI settings), plus the old/new value pairing. It is clear what data the tool surfaces, but it never states an explicit retrieval verb and does not differentiate itself from near-siblings like cm_log or cm_revert_all.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as cm_log or the revert/restore tools. The agent must infer usage from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_cloneCopy a project data fileA

Copy a project file to a new name in the same kind's folder. Never overwrites.

ParametersJSON Schema
NameRequiredDescriptionDefault
baseYesExisting file to copy
kindYesKind of project data; selects the folder below <project>/Data
new_nameYesName of the copy (must not exist)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds genuinely new behavioral content — "Never overwrites" tells the agent the collision behavior (failure rather than mutation) and "same kind's folder" constrains where the copy lands. It does not explain what happens on a name collision or any permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler, and the location constraint is front-loaded ahead of the non-overwrite guarantee. Nothing could be removed without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and 100% schema coverage, the description need not explain returns. It supplies the key non-obvious facts (folder scoping, non-overwrite) for a low-complexity three-parameter tool; only collision/error semantics remain unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the enum and all three parameters are already documented inline. The description only reinforces the folder-selection meaning of "kind" and the destination role of "new_name"; it adds no format or constraint detail beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Copy a project file") plus the destination scope ("new name in the same kind's folder"). No sibling tool performs a copy, so the operation is immediately distinguishable from cm_read, cm_edit, cm_list, and cm_model_get.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage context is only implied: the agent can infer this is for duplicating existing project data, but the description never states when to prefer it over alternatives or any preconditions beyond the inline "must not exist" note in the schema. Adequate but with clear gaps.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_doctorCheck the setupA
Read-only

Check the setup: CarMaker install, Python version, MATLAB release and engine, shared MATLAB session, GUI, project folder. Each failed check comes with its fix.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds useful behavioral context beyond that: it names exactly what gets inspected and states that failed checks are accompanied by remediation, which tells the agent the output is actionable diagnostics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the purpose before the enumeration, and the fix note last. The check list is dense but each item earns its place by telling the agent the diagnostic scope. Minimal waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless diagnostic tool with rich annotations and an output schema, the description covers what an agent needs: scope of checks and the actionable nature of failures. Return formatting is handled by the output schema, so nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, which is the baseline-4 case. The description correctly does not invent or discuss parameters, and with 100% schema coverage there is nothing further the schema leaves ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb+resource ('Check the setup') and then enumerates the specific items checked (CarMaker install, Python version, MATLAB release/engine, shared session, GUI, project folder). This is far more specific than a tautology and lets an agent know it's a diagnostic tool. It stops short of naming a sibling it is distinct from, which would push it to a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: a 'doctor' tool is understood to be run when troubleshooting setup/environment problems, and 'each failed check comes with its fix' hints at the diagnostic workflow. However, there is no explicit when-to-use guidance or reference to alternatives among the many cm_* tools, so it remains implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_dva_releaseRelease DVA overridesA

Release all quantities from DVA control.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare that this is a non-read-only, non-destructive operation, so the safety profile is covered. The description adds the scope ('all quantities from DVA control') but does not disclose reversibility, side effects, or required state beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no wasted words. It communicates the core action immediately and is appropriately sized for a no-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters, existing annotations, and an output schema, the description is nearly complete for invoking the tool. The main missing piece is usage guidance relative to sibling tools, which is a separate dimension.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema and description both need no parameter explanation. This matches the baseline of 4 for parameter-free tools.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Release') and resource ('all quantities from DVA control'), making the action clear. It does not explicitly differentiate from sibling tools such as cm_dva_write, so it falls short of the top score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like cm_dva_write, nor are prerequisites or post-conditions described. The usage context is only implied by the tool name and title.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_dva_writeOverwrite a quantity (DVA)C

Overwrite a quantity during a run via Direct Variable Access.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoAbsolute value, offset or factor; Ramp variants blend inAbs
nameYesQuantity to overwrite, e.g. 'DM.Gas'
valueYesValue, offset or factor, depending on mode
duration_msNoHow long the override lasts in ms (-1: until released or the run ends)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=false, and the description's 'overwrite' is consistent with that. However, it adds essentially nothing beyond the annotations and the schema: no statement about run-scoped lifetime, whether the override persists until released, or whether it can be applied to unknown quantities.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence with no filler and no redundancy. It is efficient, though arguably too terse given the tool's stateful side effects.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, but the description omits the critical operational context that this override typically must be undone via cm_dva_release (signalled by duration_ms defaulting to -1). For a mutation tool in a family with a dedicated release sibling, that omission leaves an agent without enough to use it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents mode, name, value and duration_ms in detail. The description adds no parameter meaning beyond that, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (overwrite a quantity) plus the mechanism (Direct Variable Access) and timing (during a run). It is clear what the tool does, but it never names the counterpart cm_dva_release or cm_standalone_dva_write, so sibling differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites tier (e.g. requires an active run/session), and no mention of the paired release tool. The only timing information is implicit in the words 'during a run'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_editEdit a project data fileA
Destructive

Change keys in a project file in place (backed up first, logged; only the edited lines change). Requires an idle simulation; load the test run again afterwards.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYesKind of project data; selects the folder below <project>/Data
nameYesFile name relative to the kind's folder
overridesYesKeys to set, e.g. {'Road.FName': 'SKIDPAD.rd5'}; a null value removes the key

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive=true/readOnly=false, but the description adds valuable behavioral detail beyond them: the file is 'backed up first', the edit is 'logged', and 'only the edited lines change'. Because it is destructive, the backup disclosure is important context that reframes the annotation. It adds no return-format or permission information, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences, front-loaded with the action ('Change keys in a project file in place'), with the backup/logging behavior and the idle-simulation prerequisite each occupying a distinct clause. Nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive 3-param edit tool with an output schema present, the description covers the critical behavioral facts an agent needs: backup, logging, minimal-diff, and the idle precondition. It does not spell out reversibility beyond backup or the post-edit reload mechanics, leaving a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents kind, name, and overrides (including the null-removes-key semantics). The description's 'change keys' loosely maps to overrides but adds no syntax or format detail beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Change keys in a project file in place', which clearly denotes an in-place mutation of project data. It is distinct from read-oriented siblings like cm_read/cm_list, but it does not name any alternative editor (e.g. cm_model_set, cm_restore) to disambiguate among the many mutation tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Requires an idle simulation; load the test run again afterwards' supplies a concrete precondition and a required follow-up action, which is real usage context. It stops short of explicit when-not-to-use or named alternatives, so it is clear context without exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_get_workspace_varRead a workspace variableA
Read-only

Value of a variable or field. MATLAB objects are described, not converted: read their fields.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesVariable or field, e.g. 'Kp_12' or 'Table.Value'
scopeNo'base' for the MATLAB base workspace, or the name of a loaded Simulink model for that model's workspacebase

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description still earns credit by disclosing return semantics: MATLAB objects come back as descriptions rather than converted values, and their fields must be read. That is non-obvious behavior an agent needs before calling, and it is not in the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with purpose and followed by the object-handling caveat; nothing is padded. The first sentence is a bare noun fragment, which is minor but slightly reduces polish rather than clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only two-parameter getter with a fully documented schema and no output schema, the description covers purpose, the object-field reading pattern, and the safety profile is carried by annotations. Return-value formatting beyond the object note is not described, but that gap is small given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both parameters (name, scope) carry their own descriptions with examples and defaults, so the schema does the heavy lifting. The description adds no syntax or format detail beyond that, which is the baseline-3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Combined with the title 'Read a workspace variable', the description states a clear verb+resource: it retrieves the value of a workspace variable or one of its fields. However it never distinguishes itself from the nearby cm_list_workspace_vars (enumerate) or cm_set_workspace_var (write), leaving differentiation to inference from the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance and no named alternative. The agent is not told that cm_list_workspace_vars enumerates names or that cm_set_workspace_var is the write counterpart, so selection relies entirely on the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_listList project data filesB
Read-only

Names of project data files of one kind (test runs, vehicles, drivers, tyres...).

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYesKind of project data; selects the folder below <project>/Data
patternNoWildcard filter on the name, e.g. '*Brak*'
include_hiddenNoAlso list names starting with a dot

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds one useful behavioral detail beyond the schema: it returns names only (not paths or contents). It says nothing about ordering or result limits, so it does not go further than a minimum.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the scope constraint precedes the examples. It is efficient, though the phrasing 'Names of...' is a fragment rather than an explicit action statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list tool with full schema coverage and an output schema, the description covers what is returned and how the kind narrows the result. The main gap is the absence of usage routing against sibling listers, which is not severe given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the enum/filters are fully documented there, so the baseline is 3. The description restates the 'one kind' selector with examples that map to enum values but adds no syntax or format guidance beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear resource (project data files) and an implicit listing verb, scoped to 'one kind' with concrete examples (test runs, vehicles, drivers, tyres). It is distinguishable from siblings like cm_results_list or cm_list_workspace_vars, though it never names or acknowledges them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of alternatives despite numerous sibling tools. The agent must infer that this is the catalogue-style lister for project data rather than results or workspace variables.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_list_workspace_varsList workspace variablesA
Read-only

Variables (name, class, size) of the MATLAB base workspace or of a Simulink model's workspace. Controller parameters usually live in the model workspace.

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeNo'base' for the MATLAB base workspace, or the name of a loaded Simulink model for that model's workspacebase
patternNoWildcard filter on the name, e.g. 'Kp_*'

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safe-read nature is covered. The description adds that the result includes name/class/size and that it can target either base or model workspace, which is useful context. It does not describe filter/pagination behavior, so this is a modest add over the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with no filler, and the core action (listing variables and where) is front-loaded before the domain tip. Efficient and well ordered, though the parenthetical output list arguably duplicates what the output schema already provides.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail need not be repeated, and the description covers both scopes plus a usage hint. Combined with full schema coverage and clean annotations, an agent has enough to invoke it correctly; only explicit sibling disambiguation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both the 'scope' and 'pattern' parameters, making 3 the baseline. The description's remark that controller parameters live in the model workspace does add interpretive value to the scope choice, but the wildcard-filter behavior of 'pattern' is left entirely to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb-and-resource (listing variables of a workspace) and even previews the returned fields (name, class, size), so the agent knows exactly what it gets. It clarifies the two possible scopes (MATLAB base vs. Simulink model workspace). It stops short of explicitly naming a sibling like cm_get_workspace_var to disambiguate, so it earns a 4 rather than a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The closing sentence 'Controller parameters usually live in the model workspace' gives contextual guidance on which scope to query. However, it never states when to use this tool versus alternatives (e.g. cm_get_workspace_var for a single variable) or names any exclusion. Usage is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_liveRead live quantitiesA
Read-only

Current values of CarMaker quantities during a run (or the last values after it).

ParametersJSON Schema
NameRequiredDescriptionDefault
quantitiesYesCarMaker quantity names, e.g. ['Time', 'Car.v']

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds real behavioral context — the values are live during a run and freeze to last-known values afterward — but says nothing about failure modes (unknown quantity names) or session prerequisites.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that states the resource and the temporal behavior with no wasted words. Nothing is padded or restated from the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and annotations cover the read-only/open-world profile. The description supplies the key temporal semantics; the only minor gap is whether an active session or run is required to obtain values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%; the single 'quantities' parameter is documented with a format example ('Time', 'Car.v'). The description adds no extra meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific resource ('CarMaker quantities') and its temporal scope ('current values during a run, or the last values after it'), which clearly separates it from batch result readers. It does not name or contrast any sibling such as cm_read or cm_output_quantities, so an agent must infer differentiation from the temporal qualifier alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical 'during a run (or the last values after it)' implies when the tool is applicable, but there is no explicit when-to-use/when-not-to-use statement and no named alternative for post-run data (e.g., cm_results_read). Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_load_testrunLoad a test runA

Load a test run into the CarMaker GUI. Requires an idle simulation. Find names with cm_list('testrun').

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesTest run relative to Data/TestRun, e.g. 'Examples/BasicFunctions/Driver/BackAndForth'
forceNoDo not let the GUI ask about unsaved data; that data is discarded

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnly=false, destructive=false, openWorld=false). The description adds a behavioral precondition the annotations don't carry: the simulation must be idle. It doesn't restate the force-discard behavior, but that is already documented in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action, then the precondition, then the name-discovery hint. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. The description supplies the key precondition and a discovery path, leaving only minor details like failure behavior when the simulation is busy unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both 'name' (with a path example) and 'force' (with discard semantics) are already documented. The description only hints at name discovery via cm_list and adds no syntax beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Load a test run into the CarMaker GUI.' Combined with the title, an agent can distinguish this from siblings like cm_list or cm_read, which read rather than load.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete precondition ('Requires an idle simulation') and routes the agent to cm_list('testrun') to discover valid names. It stops short of stating what to do when the simulation is not idle, but the primary when-to-use context is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_logRead the CarMaker session logA
Read-only

Last lines of the newest CarMaker session log of the project. Use it when a run aborts or does not start.

ParametersJSON Schema
NameRequiredDescriptionDefault
levelNowarning: warnings and errors only; error: errors onlyall
linesNoNumber of lines from the end

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds one genuinely useful behavioral fact — that it reads the *newest* log, implying older logs are not accessible here — but says nothing about truncation behavior, log rotation, or what happens when no log exists. With annotations carrying the safety burden, this is adequate but thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, and the scoping fact (last lines, newest log) is front-loaded before the usage trigger. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and full schema coverage, the description does not need to explain return values or parameters. It supplies purpose and trigger, which is close to sufficient for this simple two-param read tool; only the absence of failure-mode context (no log found, log size limits) keeps it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: the enum semantics for 'level' and the 'lines from the end' meaning are fully documented in the schema. The description adds no parameter-level detail beyond what the schema already provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource and scope — 'last lines of the newest CarMaker session log of the project' — which distinguishes it from the model/log-save siblings (cm_model_logs_save) and the generic cm_read. It lacks an explicit verb of its own (the title supplies 'Read'), so it is clear but not fully self-contained.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a concrete trigger: 'Use it when a run aborts or does not start.' That is a real when-to-use condition rather than vague guidance. However, it names no alternative (e.g. cm_status or cm_doctor for diagnosing a failed run), so the routing decision is left partly to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_model_getRead a Simulink parameterB
Read-only

get_param on a loaded Simulink model or block.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesModel or block path, e.g. 'MyModel/Controller/Gain'
paramYesParameter name, e.g. 'Gain'

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safe-read profile is covered elsewhere. The description adds only the 'loaded' precondition; it says nothing about what happens if the model is not loaded, whether specific parameter names are valid, or the error/return behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single terse fragment with no wasted words, but it is arguably stripped below the point of usefulness, offering no structure or additional sentence to orient the agent. Efficient, but under-specified rather than well-shaped.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read tool with full schema coverage and a readOnlyHint annotation, the description is minimally adequate, and no output schema means return values need not be explained. However, it leaves the relationship to cm_model_set and the failure mode for unloaded models unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters ('path', 'param') are documented with examples in the schema, so the baseline is 3. The description adds no syntax or format detail beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('get_param') and resource ('a loaded Simulink model or block'), and the title reinforces it as a read of a Simulink parameter. It is clear what the tool returns, though it never names its natural counterpart cm_model_set to sharpen the read-vs-write distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The qualifier 'loaded' implies a precondition (the model must already be loaded) and the domain is implied, but there is no explicit when-to-use/when-not guidance and no routing to alternatives such as cm_model_set or cm_get_workspace_var. Guidance is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_model_logs_saveSave Simulink's logged data to a MAT fileA

After a CarMaker for Simulink run: write what Simulink logged (signals marked for logging and To Workspace blocks) to a MAT file, with the base workspace and the model workspace as one struct each (variables: signals, base_workspace, model_workspace, info). Requires an idle simulation. The result file (.erg) holds CarMaker's quantities; this holds the Simulink side.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileNoMAT file to write, inside the project (default: next to the last result file, with the same name)
logsNoinspector: every signal the Simulation Data Inspector holds for the last run (logged signals and To Workspace blocks, complete). workspace: the logging variables in the base workspace as they areinspector
modelNoLoaded Simulink model (default: the active one)
formatNodataset: a Simulink.SimulationData.Dataset. struct: plain arrays (name, time, values, units, block) that load without Simulinkdataset
overwriteNoReplace the file if it exists
timeout_sNoLongest time MATLAB may take to write the file
base_workspaceNoInclude the base workspace's variables
model_workspaceNoInclude the model workspace's variables

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare it is a non-read-only, non-destructive, closed-world write. The description adds genuine behavioral context beyond that: the idle-simulation prerequisite and the exact file contents (variables: signals, base_workspace, model_workspace, info). It does not address overwrite semantics beyond the schema param, but the added context is substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action, then prerequisites, then sibling differentiation. The parenthetical variable list is dense but earns its place by describing file output; overall efficient with little waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, annotations covering the safety profile, and 100% schema description coverage across 8 params, the description supplies the remaining needed context: idle-simulation requirement and file-content semantics. Complete enough to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter's enum/default semantics are already documented in the schema. The description restates workspace content grouping but adds no syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('write what Simulink logged ... to a MAT file') and clarifies scope against the sibling save tools by contrasting with the .erg result file ('The result file (.erg) holds CarMaker's quantities; this holds the Simulink side'). An agent can distinguish this from cm_model_save and cm_dva_write without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear triggering context ('After a CarMaker for Simulink run') and a prerequisite ('Requires an idle simulation'). It does not name an explicit alternative tool to use instead when the simulation is not idle, so it stops short of full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_model_saveSave a Simulink modelA
Destructive

Save a loaded Simulink model. The file must be inside the project; it is backed up first.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYesName of the loaded model

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=true, so the safety profile is covered. The description adds genuinely new behavior beyond them: the file must reside inside the project, and the target is backed up before being overwritten — valuable context for an overwrite operation that could otherwise look risky.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, no filler, and the primary action is front-loaded ahead of the constraints. Every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description covers the key precondition and the backup behavior. What remains unstated — failure modes, whether the model must have unsaved changes, or effect on the backup location — is minor for a one-parameter save tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With a single parameter at 100% schema description coverage, the schema already explains that 'model' is the name of the loaded model. The description adds no further semantics (e.g., whether a path is acceptable or whether the name must match an in-memory model), so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Save a loaded Simulink model'), so the agent immediately knows the operation. It does not, however, distinguish itself from siblings such as cm_model_get, cm_model_set, or cm_model_logs_save, which share the same model-oriented naming.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the verb 'save' and the description does add one precondition ('The file must be inside the project'), which narrows when the call is valid. It never states when to use this versus alternative persistence tools (cm_model_logs_save, cm_dva_write, cm_restore) or what happens if the model is unmodified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_model_setSet a Simulink parameterA
Destructive

set_param on a loaded Simulink model or block. Logged; cm_revert_all restores the old value.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesModel or block path
saveNoAlso back up and save the model file
paramYesParameter name
valueYesNew value (numbers are converted to text as Simulink expects)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, but the description adds two facts not in the structured data: the operation is logged, and cm_revert_all restores the old value. That reversibility/audit context is genuinely valuable for a mutation tool. It does not state permission requirements or what happens on invalid paths/params.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two telegraphic fragments with zero filler, and the core action plus the restore hint are front-loaded. The extreme terseness borders on cryptic for 'set_param' as a verb, keeping it just short of a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and annotations cover the safety profile; the description supplies the prerequisite (model loaded) and the recovery path (cm_revert_all). It is essentially complete for the call decision, missing only edge-case behavior such as failure on nonexistent parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so path, param, value, and save are all documented in the schema itself; the baseline of 3 applies. The description adds no param-level detail such as value typing or save-side effects beyond what the schema already says.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (set_param) and resource (loaded Simulink model or block), which is enough for an agent to distinguish it from cm_model_get and the cm_results/cm_workspace read tools. It stops short of explicitly contrasting with near-neighbors like cm_set_workspace_var or cm_output_quantities_edit, so it is clear but not maximally differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'on a loaded Simulink model or block' implies the prerequisite that the model must already be loaded, and it names cm_revert_all as the undo route. However, there is no explicit when-to-use/when-not guidance or routing against the many sibling mutation tools, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_movie_openOpen IPGMovieA

Open IPGMovie, CarMaker's 3D animation window (or take over one that is open). Open it BEFORE a run: it records a run only while it is open, and cm_movie_snapshot can then show any moment of it.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds genuinely new behavior: the window records a run only while open, and calling it on an already-open window takes over rather than failing. It does not state what happens on close or whether the call blocks, but that is a minor gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero waste, front-loaded with the identity of the window and immediately followed by the operational constraint. Every clause carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with an output schema, the description covers identity, take-over semantics, the required call ordering relative to a run, and the downstream tool that consumes the recording. Nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. The description correctly spends no words on inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (open) and resource (IPGMovie, CarMaker's 3D animation window), and clarifies the take-over case when a window is already open. It also names the related sibling cm_movie_snapshot, so an agent can distinguish it from the snapshot tool without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit timing condition ('Open it BEFORE a run') and explains the consequence of not doing so ('it records a run only while it is open'), plus the follow-on tool that depends on it (cm_movie_snapshot). This is a clear when-to-use plus named alternative relationship.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_movie_snapshotTake a picture in IPGMovieA

A picture of IPGMovie's 3D view, returned as an image. After a run it shows any moment of the last run (time_s), provided IPGMovie was open during that run. During a run, without time_s, it shows the IPGMovie window as it is at that moment (in the window's own size). With erg it replays a saved result file, which must hold the vehicle's motion (cm_output_quantities_edit(preset='movie') before that run). Do not pass time_s during a run: that makes IPGMovie stop following the run.

ParametersJSON Schema
NameRequiredDescriptionDefault
ergNoReplay this result file instead of the last run (absolute path, or relative to the project)
widthNoPicture width in pixels (not during a run)
cameraNoCamera view: 'DEFAULT', "Bird's Eye View", 'Left Side', 'Right Side', 'Sitting In Vehicle (A)', or one from the project (default: keep the view)
heightNoPicture height in pixels (not during a run)
time_sNoSimulation time of the picture in s (default: the latest moment IPGMovie holds)

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already state readOnlyHint=false, destructiveHint=false, openWorldHint=false, and the description adds non-obvious behavioral context: passing time_s during a run makes IPGMovie stop following the run, and erg requires the result file to contain vehicle motion prepared via cm_output_quantities_edit. It does not cover rate limits or exact output format, so 4 rather than 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action and then explains the three modes and the side-effect warning in compact sentences. Every sentence carries a condition or prerequisite, so it earns its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With five optional parameters fully documented in the schema and safety annotations present, the description fills the remaining behavioral and usage gaps: return type, run-context rules, erg prerequisite, and the stop-following side effect. Nothing critical for correct invocation appears missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description nevertheless adds semantic constraints beyond the schema: time_s is only valid after a run, width/height are ignored during a run, and erg must reference a result file with vehicle motion. This is useful contextual meaning, yielding a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool captures IPGMovie's 3D view and returns it as an image, giving a clear verb and resource. It does not explicitly differentiate itself from sibling tools such as cm_movie_open or cm_live, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit conditional guidance: after a run use time_s for any moment (if IPGMovie was open), during a run omit time_s and use the window as-is, and with erg replay a saved result. It also names a clear when-not: do not pass time_s during a run because it breaks run-following.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_output_quantitiesRead the output quantitiesB
Read-only

The quantities the next run writes to its result file, per storage rate (names may hold wildcards), with the sample times. movie_replay says whether such a result file can be replayed in IPGMovie. What a finished result file holds: cm_results_summary(erg, search='*').

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and openWorldHint=false, so the safety profile is already covered. The description does add that it reports what the NEXT run will write (not current results) and that movie_replay indicates replayability in IPGMovie, which is useful context beyond the annotations. However, it doesn't disclose behavior about wildcards in names or how results are returned when the tool has an output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, somewhat run-on sentence with awkward nesting ('What a finished result file holds: ...'). It is front-loaded with the core subject, but the trailing reference to cm_results_summary is syntactically disjoint and would benefit from clearer punctuation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description needn't enumerate return fields, which is appropriate. But for a no-arg, openWorld-disabled read tool, it leaves gaps: what 'storage rate' means in context, how wildcards behave, and the relationship to cm_output_quantities_edit are not explained. It is minimally complete but not thorough.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the baseline is 4. The description correctly focuses on return content rather than parameter semantics, which is appropriate for a no-arg introspection tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description reads 'the quantities the next run writes to its result file' but never states a clear verb or that this is a read operation. It is distinguishable from cm_output_quantities_edit only loosely, and the 'movie_replay' field mention muddies what the tool actually returns. Purpose is inferable but not stated crisply.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus siblings like cm_results_summary or cm_results_read. The only routing evidence is a trailing pointer to cm_results_summary for finished files, which is a contrast but not an explicit usage condition. An agent gets no 'use this when...' clause.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_output_quantities_editChange the output quantitiesA
Destructive

Add quantities to, or remove them from, the project's output-quantities file (backed up first, logged, restored by cm_revert_all). Applies from the next start. Requires an idle simulation. Find names with cm_live (it reads any quantity) or in another result file.

ParametersJSON Schema
NameRequiredDescriptionDefault
addNoQuantities to store as well, e.g. ['Car.az', 'PT.Motor*.Trq']; wildcards * and ? allowed
rateNoStorage rate for added quantities (sample times: cm_output_quantities)normal
presetNomovie: add what IPGMovie needs to replay the vehicle from a result file
removeNoEntries to take out, spelled as they are listed

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial context beyond the annotations: the file is backed up first, edits are logged, and changes can be undone via cm_revert_all. It also discloses the timing behavior (applies from next start) and the idle-simulation precondition, which the destructiveHint alone would not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three information-dense sentences with the mutation action front-loaded, followed by safety/undo and preconditions. Nearly every clause earns its place, though the name-discovery sentence is slightly tangential to the core operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool, the description covers backup, logging, revert path, timing (next start), and the idle-simulation prerequisite. With an output schema present, return values need not be explained, so nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents add/remove/rate/preset in detail, including wildcards and the movie preset. The description contributes only the name-discovery source (cm_live) and does not elaborate on individual parameters, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair (add/remove) and the exact resource modified (the project's output-quantities file), which cleanly distinguishes it from the read-only sibling cm_output_quantities. An agent can tell immediately what changes and where.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operating context: requires an idle simulation, applies from the next start, and tells the agent how to source names (via cm_live or another result file). It does not explicitly contrast with the sibling cm_output_quantities for reading the current list, leaving one alternative inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_popupsRead GUI pop-upsA

Pop-up messages the CarMaker GUI raised for this server's commands (type info/warn/err/question, text, index of the answer given) and the current pop-up timeout. Reading empties the GUI's buffer (it keeps five). cm_load_testrun, cm_start_sim and cm_wait_end already return new ones as 'popups'.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, and the description justifies that non-obvious flag by disclosing that reading EMPTIES the GUI's buffer, which retains only five entries. That consumption/ring-buffer behavior is exactly the kind of side effect annotations cannot express, and it explains why a 'read' is flagged as non-read-only.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the tool's output, then the destructive-read caveat, then the sibling routing. Despite the parenthetical field list, no clause is redundant and nothing essential is buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, yet the description still summarises the payload (type/text/answer index/timeout) and the buffer semantics, so an agent has everything needed to call it and interpret results. Nothing relevant is missing for a zero-parameter read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4; there is no argument surface the description could clarify, and it correctly spends no words on inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Pop-up messages the CarMaker GUI raised for this server's commands') and enumerates the returned fields (type, text, answer index, timeout), which cleanly separates it from cm_popup_timeout and from the cm_*_results/cm_*_logs readers. An agent knows exactly what this tool returns without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The closing sentence routes agents away from this tool when other calls already surface popups, naming cm_load_testrun, cm_start_sim and cm_wait_end as alternatives. It stops short of an explicit 'call this when you suspect missed or queued popups', but the context for usage is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_popup_timeoutLet GUI pop-ups answer themselvesA
Destructive

Let GUI pop-ups raised by this server's commands answer themselves with their DEFAULT choice, so that a question such as 'Vehicle not saved. All changes will be lost. OK to continue?' cannot block a run. The default answer there discards the unsaved changes. Read what was answered with cm_popups. Stays until set back, cm_revert_all or a GUI restart.

ParametersJSON Schema
NameRequiredDescriptionDefault
secondsYesSeconds until a pop-up answers itself; 0 = never shown; -1 = wait for a click in the GUI (the normal setting)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the destructiveHint annotation by disclosing the actual consequence: the default answer discards unsaved changes, and a concrete 'All changes will be lost' example makes the risk tangible. It also discloses persistence ('Stays until set back, cm_revert_all or a GUI restart'), which annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose, then supplies the destructive consequence, the sibling to read results, and the persistence scope in a compact run of sentences. Slightly dense, but every sentence carries distinct, useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-param toggle with annotations and an output schema present, the description supplies everything needed: purpose, risk, companion tool for verifying results, and lifetime. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents the seconds parameter including the meaningful 0/-1 sentinel values. The description adds no additional parameter nuance beyond the persistence note, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action (make GUI pop-ups auto-answer with their default) and the resource (this server's pop-ups), with a concrete example of the blocking question it addresses. It also names the related sibling cm_popups, letting an agent separate this from other popup tools without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the motivating scenario clearly ('cannot block a run') and points to cm_popups to read the answers, which routes the agent correctly. It stops short of stating explicit exclusions or when an agent should avoid setting a timeout, but the use context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_readRead a project data fileB
Read-only

Key / value pairs of a CarMaker infofile. Large files are truncated: filter by keys or prefix.

ParametersJSON Schema
NameRequiredDescriptionDefault
keysNoOnly these keys
kindYesKind of project data; selects the folder below <project>/Data
nameYesFile name relative to the kind's folder, as cm_list returns it
prefixNoOnly keys starting with this, e.g. 'Vehicle' or 'DrivMan.'

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and openWorldHint=false, so safety is already covered, and the description adds a genuinely useful trait annotations cannot express: large files are truncated, with the mitigation (filter by keys or prefix). It still doesn't say what an error looks like for a missing file or whether truncation is silent versus flagged, so it is strong but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero padding, and the truncation warning is placed where it matters. The opening noun fragment is slightly clipped — a verb would make it read better — but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation, and the read-only annotation carries the safety profile; the description covers the one thing neither structured field expresses (truncation). The gap is the missing link to cm_list, which the schema alludes to with "as cm_list returns it" but the description never states.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description goes one step further by tying the optional keys/prefix parameters to a concrete consequence (truncation avoidance), which explains why an agent should use them. The example prefix formats themselves live in the schema, so no further compensation is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource precisely — CarMaker infofile key/value pairs — but uses a noun fragment with no verb, so the action (reading a file's contents) is only inferable from the title. It also never distinguishes itself from siblings like cm_list or cm_edit, leaving the agent to work out that this one fetches content rather than file names or writes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Large files are truncated: filter by keys or prefix" tells the agent how to avoid truncation, but there is no guidance on when to choose this tool over cm_list (to obtain the file name) or cm_edit (to modify it). No prerequisites, no exclusions, and no named alternative — usage is left entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_restoreRestore files of an earlier sessionB
Destructive

Restore project files from the backups of an earlier session. Files only.

ParametersJSON Schema
NameRequiredDescriptionDefault
sessionYesSession id as shown by cm_changelog

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, so the agent knows this is a risky write. The description adds useful scope context ('files only', restoring from backups of an earlier session), but does not say whether current files are overwritten or whether the operation is reversible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the verb and scope front-loaded; the terse 'Files only' clause earns its place by narrowing scope. Nothing is bloated, though it is arguably under-specified rather than optimally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the destructive annotation covers risk. Still missing for a destructive restore: what happens to existing files, prerequisites, and how this relates to cm_revert_all.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter and schema description coverage is 100%, so the schema already explains it (session id as shown by cm_changelog). 'From the backups of an earlier session' adds only marginal meaning beyond the schema, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (restore), resource (project files) and source (backups of an earlier session), and 'Files only' scopes what is restored. It does not distinguish itself from the similar-sounding sibling cm_revert_all, which is the main gap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Once again, no guidance on when and why to use cm_restore versus cm_revert_all or any other restore-like tool is provided. The 'an earlier session' phrasing combined with the schema pointer to cm_changelog weakly implies how to obtain the required session id, but there are no explicit conditions or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_results_listList result filesB
Read-only

Newest .erg result files under the project's SimOutput folder, CM_RESULT_DIRS and the folders this session's runs wrote to.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum number of files

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds genuine value by disclosing the search scope (project SimOutput, CM_RESULT_DIRS, session-written folders), which is behavior the annotations can't convey, but it says nothing about ordering guarantees beyond the implied 'newest', result cap behavior, or what happens when no sources exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the most decision-relevant information (file type) front-loaded. No filler, though it is a fragment rather than a statement of action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the search scope is documented. However, for a tool embedded among many results-oriented siblings, the absence of any routing or sequencing guidance leaves the description only minimally adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and there is a single optional limit parameter, so the schema already carries the semantics. The description's only contribution is the implicit newest-first ordering, which is not formalized anywhere else but is not tied to the limit parameter either.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description specifies the resource precisely: newest .erg result files, and enumerates the search roots (SimOutput, CM_RESULT_DIRS, session run folders). It distinguishes itself from cm_read/cm_list by naming the file type and location, but reads as a noun phrase rather than a verb+resource statement and never differentiates from close siblings like cm_results_read or cm_results_summary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance at all. Given siblings cm_results_read, cm_results_summary, cm_standalone_results, and cm_load_testrun all operate on results, the agent gets no signal about which to pick or in what order to call them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_results_readRead time series from a result fileB
Read-only

Time series from a result file, decimated to at most max_points samples.

ParametersJSON Schema
NameRequiredDescriptionDefault
ergYesResult file (.erg): absolute path, or relative to the project
t_maxNoEnd of the time window in s
t_minNoStart of the time window in s
max_pointsNoSamples are thinned to at most this many per quantity
quantitiesYesCarMaker quantity names, e.g. ['Time', 'Car.v']

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a safe read (readOnlyHint=true, openWorldHint=false). The description adds one genuine behavioral trait beyond annotations: output is decimated to at most max_points samples. It does not state what happens to missing quantities or invalid windows, so it adds only modest value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted words. It is efficient, though its terseness borders on under-specification rather than optimal information density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the full 100% schema coverage relieves parameter documentation pressure. It is minimally adequate, but the absence of any when-to-use or windowing guidance leaves the definition thin for a multi-tool results workflow.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five parameters. The description reinforces the max_points decimation semantics but adds no new syntax or format detail beyond the schema; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (time series from a result file) with an action implied by the tool name 'read'. It is distinguishable from siblings like cm_results_summary or cm_results_list, though the description is a noun phrase rather than an explicit verb+object statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to prefer this over cm_results_summary, cm_results_list, or cm_output_quantities, nor any prerequisite or context for its use. The agent must infer the appropriate scenario entirely from the name and schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_results_summarySummarise a result fileB
Read-only

First / last / min / max / mean and unit of quantities in a result file, plus rows and duration.

ParametersJSON Schema
NameRequiredDescriptionDefault
ergYesResult file (.erg): absolute path, or relative to the project
searchNoInstead of a summary, list the quantity names that contain this text ('PT.') or match this wildcard pattern ('Car.v*'; '*' lists all)
quantitiesNoQuantities to summarise (default: Time, Car.v)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safe-read profile is covered. The description adds the set of aggregate statistics produced, which is useful, but discloses nothing about file-access constraints, error cases, or cost despite operating on an arbitrary path.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with no padding, and the output enumeration is front-loaded. It is slightly telegraphic but wastes nothing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no prose, and all three params are documented. For a read-only summary tool the definition is essentially complete, with only the missing sibling/disambiguation context as a gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with all three parameters fully documented, so the baseline is 3. The description adds no syntax or format detail for `erg`, `search`, or `quantities` beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Summarise) and resource (result file), and enumerates the computed statistics (first/last/min/max/mean, unit, rows, duration). It is clear what the tool produces, but it never names or distinguishes itself from sibling result tools like cm_results_read or cm_results_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no when-to-use guidance and no comparison to the many sibling result/output tools. The only routing hint (that `search` switches behavior to name-listing) lives in the schema, not the description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_revert_allUndo this session's changesA
Destructive

Undo every change made in this server session: workspace values and model parameters in reverse order, backed-up files put back (created files go to a trash folder), GUI settings restored.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true and readOnlyHint=false, but the description adds meaningful behavioral detail: reverse-order reversion, backed-up files being restored, created files moved to a trash folder, and GUI settings restored. This goes beyond the annotations and helps the agent understand the concrete effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that covers the core action and its main effects without redundancy. Every clause contributes useful detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema and destructive annotations, the description is largely complete: it explains scope, ordering, file restoration, and trash behavior. The main gap is the lack of comparison to sibling tools like cm_restore.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool accepts zero parameters, so there are no parameter semantics to explain. The baseline for a parameterless tool is 4, and the empty schema is consistent with the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and scope: undoing every change made in this server session, including workspace values, model parameters, files, and GUI settings. It is clear about what the tool does, but does not distinguish itself from related sibling tools such as cm_restore.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'this server session' implies the context for use, but there is no explicit guidance on when to choose this tool over alternatives like cm_restore or when not to use it. The usage is inferable but not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_session_startStart the CarMaker for Simulink sessionA

Bring up a CarMaker for Simulink session as far as it is missing: start MATLAB (visible, in the project's src_cm4sl folder, engine shared), run cmenv, open the Simulink model and open the CarMaker GUI. Nothing is restarted or closed. Returns ready=true, or ready=false with what it is waiting for: call again until ready. Without a model it opens none and lists models_loaded and models_found to choose from. Many projects need their own setup script before the model compiles: if init_scripts_found lists one, pass it as init.

ParametersJSON Schema
NameRequiredDescriptionDefault
initNoThe project's own MATLAB setup script, run once per MATLAB session before the model is opened (default: CM_MATLAB_INIT). A name from init_scripts_found, or a full path
modelNoSimulink model to open: a name on the MATLAB path, a path as listed in models_found, or a full path (default: CM_MODEL)
max_wait_sNoLongest time to wait in seconds; keep it below your client's tool time-out and call again if not finished

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (destructiveHint=false), the description makes explicit that 'Nothing is restarted or closed', that only missing pieces are brought up (idempotency), and what the ready=true/false return signals plus the waiting condition mean. This adds meaningful behavioral context the annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but front-loads the primary action and each sentence carries information. It is a fairly packed paragraph with some long compound sentences, but there is little outright waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a multi-step session-startup tool, the description covers side effects, idempotency, model-optional behavior, init-script handling, wait/timeout guidance, and return semantics, and an output schema exists to absorb return details. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it ties 'init' to workflow ('if init_scripts_found lists one, pass it as init') and explains the no-model case that drives the 'model' parameter choice. It stops short of detailing max_wait_s beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Bring up a CarMaker for Simulink session') and enumerates the concrete steps (start MATLAB, run cmenv, open Simulink model, open CarMaker GUI). This clearly separates it from siblings like cm_start_sim (run a simulation) and cm_standalone_launch (a different server type).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives strong operational context: bring the session up 'as far as it is missing', call again until ready, and pass a script from init_scripts_found as init when a project needs its own setup. It does not explicitly name alternative tools or state when NOT to use it, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_set_workspace_varSet a workspace variableA
Destructive

Set a variable or field. The old value is logged and cm_revert_all restores it. Model-workspace changes live in memory until cm_model_save. Requires an idle simulation.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesVariable or field, e.g. 'Kp_12' or 'Table.Value'
scopeNo'base' for the MATLAB base workspace, or the name of a loaded Simulink model for that model's workspacebase
valueYesNumber, string, list or nested list (matrix)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as a destructive write, but the description adds genuinely non-obvious behavior: the old value is logged, the change is revertible via cm_revert_all, model-workspace edits stay in memory until cm_model_save, and an idle simulation is required. These are exactly the traits an agent needs before mutating state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action, then the revert/persistence semantics, then the precondition. No filler and every clause carries operational weight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no prose, and the description covers the remaining gaps: mutation durability, undo path, and the idle-simulation prerequisite. Nothing an agent must know to invoke this safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so name, scope and value are already documented with examples ('Kp_12', 'Table.Value') and accepted types. The description adds no parameter-level detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Set a variable or field') with enough precision to distinguish it from the read-side sibling cm_get_workspace_var. However, it does not explicitly contrast itself with near-neighbors like cm_model_set, leaving the agent to infer the boundary from the scope parameter.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives one real precondition ('Requires an idle simulation') and names follow-on tools (cm_revert_all, cm_model_save), which is useful implied usage. But it never states when to choose this over cm_model_set or cm_edit for the same underlying change, so selection guidance is only partial.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_standalone_attachAttach to a running CarMaker programA

Attach to a CarMaker simulation program that is already running, for example one started from a CarMaker Office window.

ParametersJSON Schema
NameRequiredDescriptionDefault
pidYesProcess id from cm_standalone_servers
startNoStart the run at once (needs testrun)
testrunNoTest run to configure (needed to start a run)
overridesNoTest-run keys to change in memory only, e.g. {'Vehicle.DriverTemplate.FName': 'X'}
quantitiesNoQuantities to store
stop_after_sNoSafety stop at this simulation time in s
realtime_factorNoSpeed relative to real time

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false), so the agent knows this is a non-destructive state-changing, closed-world operation. The description adds the 'already running' context but says nothing about whether attaching blocks, how it interacts with session state, or what happens to the existing run.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that carries the core action, a scope qualifier, and a concrete example with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists and all parameters are fully documented, so return values and argument syntax need no prose. However, for a session-oriented tool sitting among many cm_standalone_* siblings, the description does little to situate it in the lifecycle (e.g., relationship to cm_standalone_servers for the pid or cm_standalone_launch).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all seven parameters (pid, start, testrun, overrides, quantities, stop_after_s, realtime_factor) in detail. The description adds no parameter meaning beyond that, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Attach) and resource (a CarMaker simulation program that is already running), with the 'already running / started from an Office window' qualifier effectively distinguishing it from launching a new program. It stops short of naming the sibling it complements (cm_standalone_launch), so an agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'already running' condition and the Office-window example give implied context for when to attach rather than launch, but no alternative tool is named and no exclusions are given. Guidance is inferable rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_standalone_closeClose a standalone instanceA
Destructive

Disconnect; for instances this server launched, also stop the CarMaker process.

ParametersJSON Schema
NameRequiredDescriptionDefault
instanceNoStandalone instance id such as 'sa1' (may be omitted while there is only one)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

This is the description's real contribution: it discloses that closing is a plain disconnect, except for instances this server launched, where it also terminates the CarMaker process. That conditional destruction detail goes beyond the destructiveHint=false/readOnlyHint=false annotations. It stops short of describing irreversibility, error behavior, or state loss, so not a full 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One tight sentence, front-loaded with the core action and immediately qualified by the conditional behavior. No filler, no restating of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and annotations covering the lock/open-world/destructive profile, the description supplies the one missing nuance (conditional process kill). It could still mention what happens to instance state or the omitted-parameter case, but it is close to complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'instance' parameter is fully documented there, including the 'may be omitted while there is only one' note. The description adds nothing about the parameter, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The title and description state a specific verb ('disconnect'/'close') on a specific resource (a standalone instance), and the text distinguishes two modes of closing. It doesn't name any sibling (e.g., cm_standalone_attach or cm_standalone_launch) to sharpen the distinction, so it lands just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied by the name/title; the description gives no explicit when-to-use or when-not-to-use guidance and never points at alternatives among the many cm_standalone_* siblings. Adequate for an obvious lifecycle operation, but nothing is added beyond the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_standalone_controlControl a standalone runC
Destructive

Start, pause, resume or stop the simulation of a standalone instance.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesWhat to do
instanceNoStandalone instance id such as 'sa1' (may be omitted while there is only one)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false, and openWorldHint=false, so the safety profile is covered. The description adds nothing further: it does not clarify that 'stop' terminates the run, what state survives a pause, or any prerequisite (e.g., an attached instance), leaving the behavioral burden entirely on the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the action list front-loaded and no wasted words. Slightly more context could be added without bloat, but the structure is clean.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, but for a destructive multi-verb control tool with an optional instance target, the description omits lifecycle detail (what each action does to the instance) and prerequisites. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with the action enum and instance parameter well documented ('sa1', may be omitted while only one exists). The description merely enumerates the same actions the enum already lists, adding no syntax or semantic value, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb ('control') and resource (a standalone instance's simulation) and enumerates the four actions aligned with the enum. However, it does not distinguish itself from siblings like cm_standalone_launch or cm_standalone_close, so an agent cannot tell from the text alone which control entry point to use.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative-tool guidance is given. The agent is not told how this differs from cm_standalone_launch, cm_standalone_close, or cm_start_sim/cm_stop_sim, all of which plausibly overlap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_standalone_dva_writeOverwrite a quantity in a standalone run (DVA)B

Overwrite a quantity in a running standalone simulation (Direct Variable Access).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesQuantity to overwrite
valueYesAbsolute value
instanceNoStandalone instance id such as 'sa1' (may be omitted while there is only one)
duration_msNoHow long the override lasts in ms (-1: until the run ends)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the mutation/safety profile is covered. The description only adds that the target is a running standalone run and that the change is an overwrite; it does not disclose that the effect is time-bound (per duration_ms) or what happens when the run ends. With annotations carrying the safety bar, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no waste, stating the action and the target domain immediately. It is efficient, though the brevity is partly under-specification rather than true conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values needn't be explained, and the schema fully documents the 4 parameters. However, for a state-mutating tool the description omits the relationship to cm_dva_write/cm_dva_release and the time-bounded nature of the override, leaving meaningful gaps for an agent choosing among the DVA siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all four parameters are documented in the schema, including the duration_ms semantics ('-1: until the run ends'), so the baseline is 3. The description adds nothing about parameters beyond restating the overwrite action.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Overwrite') and resource ('a quantity in a running standalone simulation'), and expanding DVA signals the domain. It distinguishes itself from the plain cm_dva_write sibling only implicitly via 'standalone' in the title/name; the description never names the non-standalone alternative, so an agent must infer the split.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or when-not guidance. The counterpart cm_dva_release (which presumably removes the override) is never referenced, and the temporary nature of the write is not framed as usage guidance. The only context ('running simulation') is implied rather than stated as a precondition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_standalone_launchLaunch a standalone CarMaker runA

Start a NEW independent CarMaker simulation program (not connected to MATLAB) and run a test run on it. Only for test runs whose vehicle does not need a Simulink controller. Results go to the project's SimOutput; see cm_standalone_results.

ParametersJSON Schema
NameRequiredDescriptionDefault
startNoStart the run at once
testrunYesTest run relative to Data/TestRun
overridesNoTest-run keys to change in memory only, e.g. {'Vehicle.DriverTemplate.FName': 'X'}
executableNoCarMaker executable (default: the stock one of the install)
quantitiesNoQuantities to store (default: Time, Car.v, Car.YawRate, Car.ax, Car.ay, Car.Distance)
stop_after_sNoSafety stop at this simulation time in s (null: none)
realtime_factorNoSimulation speed relative to real time (larger is faster)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a non-read-only, non-destructive, closed-world operation. The description adds real context beyond that: it spawns a NEW independent process disconnected from MATLAB and writes results into the project's SimOutput. It stays silent on process lifecycle (that the run persists and must be closed/queried via cm_standalone_close/status/wait_end).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the primary action and the gating constraint, then the results destination. No padding, and the pointer to cm_standalone_results earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a full output schema and annotations covering the safety profile, the description is largely complete, and it correctly defers return values to cm_standalone_results. The one gap is the surrounding lifecycle workflow (attach/status/wait_end/close) implied by the sibling set, which a launching agent would benefit from knowing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all seven parameters including start, overrides, quantities and stop_after_s are already documented in the schema. The description adds no additional parameter meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Start a NEW independent CarMaker simulation program... and run a test run on it') and explicitly scopes it against the MATLAB-connected path. It also names the results sibling, so an agent can distinguish it from cm_start_sim without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-to-use constraint: 'Only for test runs whose vehicle does not need a Simulink controller.' That cleanly excludes the Simulink case, but it never names the alternative tool (cm_start_sim) to use instead, leaving the routing half-inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_standalone_resultsResult files of a standalone runA
Read-only

Result .erg files of the instance's run (read them with cm_results_summary / cm_results_read).

ParametersJSON Schema
NameRequiredDescriptionDefault
instanceNoStandalone instance id such as 'sa1' (may be omitted while there is only one)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint=true annotation already establishes this as a safe, non-mutating read, so the bar is lowered. The description adds that the artifacts are .erg result files and that companion tools consume them, but says nothing about ordering, volume, error states (e.g. no run present), or filtering behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence that is entirely front-loaded with the resource and then the next-step tools. No filler, no redundancy, and every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the single optional parameter is covered by the schema. What is missing for a listing tool in a crowded results family is disambiguation from cm_results_list and any precondition about the run state; adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single optional parameter with 100% schema description coverage, so the schema fully documents 'instance' and the default-when-only-one behavior. The description adds no parameter-level detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource: the result .erg files produced by the instance's run, which is a clear noun+scope. It also implicitly differentiates itself from the reading tools by directing the agent to cm_results_summary / cm_results_read. It does not, however, distinguish itself from the close sibling cm_results_list, which an agent could easily confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides a useful forward-routing hint ('read them with cm_results_summary / cm_results_read'), telling the agent what to do after obtaining results. But there is no guidance on when to choose this over the near-identical cm_results_list, and no exclusion or precondition (e.g. that a run must have completed) is stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_standalone_serversList running CarMaker programsB
Read-only

CarMaker simulation programs running on this machine (pid, description, whether this server manages it). A CarMaker Office window you opened shows up once its application is started.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety and scope are covered. The description adds real value beyond that by disclosing the fields returned (pid, description, whether this server manages it) and the timing nuance that a CarMaker Office window only appears once its application has started. It still says nothing about ordering, staleness guarantees, or whether the list can be empty.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler, and the core enumeration is front-loaded before the qualifying note about Office windows. The second sentence is slightly tangential but does earn its place by warning that entries may not appear immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless read-only enumeration with an output schema available to explain the return shape, the description covers what the tool returns, what kind of objects appear, and the latency caveat. The remaining gap is navigational rather than behavioral: it does not tell the agent how this differs from cm_standalone_status or when to prefer one.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies; the schema is a closed empty object and there is nothing for the description to clarify. The parenthetical field list describes outputs, not inputs, and does not confuse that distinction.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific resource and scope: CarMaker simulation programs currently running on this machine, with the returned fields named in parentheses. It is clear what an agent gets, but it never uses an explicit listing verb and draws no boundary against the adjacent cm_standalone_status, which an agent could easily confuse with this enumeration tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance, and no alternative is named despite several plausible siblings (cm_standalone_status, cm_list, cm_status). The only usage-adjacent statement is the note that an Office window appears once its application is started, which explains data latency rather than tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_standalone_statusStandalone instance statusC
Read-only

State of a standalone instance plus live values.

ParametersJSON Schema
NameRequiredDescriptionDefault
instanceNoStandalone instance id such as 'sa1' (may be omitted while there is only one)
quantitiesNoLive quantities to read (default: Time, Car.v)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish this as a safe read (readOnlyHint=true, openWorldHint=false), so the bar is low, yet the description adds almost nothing: it does not say what 'state' comprises, whether results are cached or live-sampled, or how the default quantities get resolved. 'Plus live values' is the only extra hint and it is not elaborated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short, front-loaded sentence with no waste, but it is arguably too terse for a status tool with two optional parameters, leaving the reader under-informed rather than efficiently informed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and annotations cover the safety profile, but the description still omits sibling differentiation and what 'state' versus 'live values' actually covers, which matters given the crowded standalone_* tool family.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters (instance, quantities) are already documented with defaults, so the baseline is 3. The description adds no syntax, format, or default-resolution detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

Identifies the resource (standalone instance) and idea of reporting state plus live values, but the phrasing 'State of... plus live values' is vague about what is actually returned and never distinguishes this from siblings like cm_status, cm_live, or cm_standalone_servers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No indication of when to call this versus cm_status, cm_live, or the other standalone_* tools, no prerequisites, and no exclusions. The agent has to guess based on the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_standalone_wait_endWait for a standalone run to endA
Read-only

Wait until the instance's simulation has finished. finished=true: simulation time and distance, error flag and log errors. finished=false: call again to keep waiting.

ParametersJSON Schema
NameRequiredDescriptionDefault
instanceNoStandalone instance id such as 'sa1' (may be omitted while there is only one)
timeout_sNoLongest time to wait in seconds; keep it below your client's tool time-out and call again if not finished

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds genuinely useful behavior beyond that: the branching on finished, and what is returned on completion (simulation time, distance, error flag, log errors), which tells the agent this is a poll-until-done loop rather than a one-shot query.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the outcome branching front-loaded after the purpose; every clause carries information. The elliptical fragments ('finished=true: simulation time and distance...') are terse but readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return shape need not be explained in depth, and annotations cover the read-only nature. The description still supplies the essential control-flow contract (poll again if not finished) and the success payload, making the definition adequate for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both instance and timeout_s are already documented in the schema, including the advice on socket timeouts. The description adds only the framing that the wait targets 'the instance's simulation', which is marginal beyond the schema baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: wait until the instance's simulation has finished. It is clear what the tool does, but it never names or contrasts with the sibling cm_wait_end, so an agent cannot tell from the text alone which wait tool applies to non-standalone runs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage via the polling loop ('finished=false: call again to keep waiting') and the schema advises keeping timeout_s below the client tool timeout. However, it gives no explicit when-to-use guidance relative to cm_standalone_status or cm_wait_end, leaving routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_start_simStart the simulationA

Start the loaded test run. Returns once it is running; fails at once, with Simulink's error, if the model does not compile. Then call cm_wait_end until it reports finished.

ParametersJSON Schema
NameRequiredDescriptionDefault
saveNosave: write a result file (GUI storage mode becomes 'Save all', cm_revert_all puts it back). collect: buffer only. keep: leave the GUI setting alonesave
save_logsNoAlso save what Simulink logged in this run to a MAT file next to the result file, as cm_model_logs_save does; cm_wait_end returns it as logs_file
start_timeout_sNoHow long the start may take (a Simulink model may compile first)

Output Schema

ParametersJSON Schema
NameRequiredDescription
noteNo
stateNo
popupsNo
after_sNo
startedNo
save_modeNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover only the safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false), and the description adds genuinely useful behavior: it is synchronous, returns as soon as the run is up, and fails fast with Simulink's own error text if compilation fails. It stops short of disclosing the file-writing side effect of the default save mode, which matters for a non-read-only tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, zero padding, with the primary action, its failure mode, and the next step ordered by importance. Nothing could be cut without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and full parameter documentation, the description only needs to explain the call's lifecycle, and it does: start, fail-fast on compile error, then poll with cm_wait_end. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents save/collect/keep, save_logs, and start_timeout_s in detail. The description adds no parameter-level meaning beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb + resource ('Start the loaded test run') with the scope qualifier 'loaded' that separates it from cm_load_testrun. An agent can distinguish it from cm_stop_sim and cm_wait_end purely from this sentence.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly sequences the workflow by naming cm_wait_end as the required follow-up call until it reports finished, which is real routing guidance. It does not state the prerequisite explicitly (that a test run must already be loaded) or when NOT to call it, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_statusSimulation statusA
Read-only

Simulation state, active Simulink model, project directory, last end status, the GUI's storage mode (save_mode) and gui_all_saved (false: the GUI holds unsaved data and will ask before a load; true is no guarantee, an edit in an open editor window is not counted).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
detailsNo
connectedNo
save_modeNo
end_statusNo
session_idNo
sim_statusNo
project_dirNo
active_modelNo
gui_all_savedNo
matlab_simstateNo
sim_status_codeNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description goes further by disclosing a genuine nuance not in annotations: gui_all_saved=false means the GUI holds unsaved data and will prompt before a load, while true is explicitly not a guarantee because edits in an open editor window are not counted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with no filler, and the primary payload (simulation state) is front-loaded. The nested parenthetical on gui_all_saved is dense and slightly hard to parse, but every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description is not obliged to enumerate return fields, yet the enumeration plus the gui_all_saved caveat covers the semantics an agent needs for a zero-parameter status read. The only gap is routing versus the other status siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description enumerates exactly what state is retrieved: simulation state, active Simulink model, project directory, last end status, save_mode and gui_all_saved. That is concrete and unambiguous, though it is phrased as a noun list rather than a verb and never distinguishes this from siblings like cm_study_status or cm_standalone_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use statement, no prerequisites, and no named alternative. That several sibling status tools exist (cm_study_status, cm_standalone_status, cm_doctor) makes the omission noticeable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_stop_simStop the simulationA
Destructive

Stop the running simulation. With wait_s > 0 it returns what cm_wait_end returns.

ParametersJSON Schema
NameRequiredDescriptionDefault
wait_sNoWait this long for the simulation to become idle (0: do not wait)

Output Schema

ParametersJSON Schema
NameRequiredDescription
liveNo
noteNo
stateNo
popupsNo
testrunNo
finishedNo
waited_sNo
elapsed_sNo
logs_fileNo
distance_mNo
end_statusNo
log_errorsNo
logs_errorNo
sim_time_sNo
poll_errorsNo
result_fileNo
simulink_errorNo
stop_requestedNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, readOnlyHint=false and openWorldHint=false, so the safety profile is covered. The description adds genuinely new behavior for the wait_s > 0 case (blocking until idle and mirroring cm_wait_end's return), but does not say what is discarded or what happens when no simulation is running.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, the primary action front-loaded and the conditional behavior second. No filler and nothing that restates the name or title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be documented. However, for a destructive operation the description says nothing about what state is destroyed, whether results are retained, or error behavior when no simulation is active, leaving a meaningful gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, and the description adds meaning beyond the schema by explaining what the wait actually yields ('returns what cm_wait_end returns') rather than just repeating the idle-wait wording.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Stop the running simulation'), which cleanly contrasts with the sibling cm_start_sim. It stops short of explicit sibling routing, but the reference to cm_wait_end makes the boundary with that tool inferable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'Stop the running simulation' rather than stated as a condition. The note linking wait_s > 0 to cm_wait_end implies a choice between the two tools, but no when-not or prerequisites (e.g. whether a simulation must be running) are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_study_cancelCancel a parameter studyA
Destructive

Stop the run in progress, skip the remaining ones and restore the changed workspace values.

ParametersJSON Schema
NameRequiredDescriptionDefault
studyNoStudy id (default: the latest)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare destructiveHint=true and readOnlyHint=false, and the description usefully explains what that destruction concretely means: the current run is stopped, queued runs are skipped, and changed workspace values are restored. It stops short of stating whether already-completed runs' results survive or whether the study can be resumed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One tight sentence that front-loads the core action and then enumerates the side effects, with no redundant restatement of the name or title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations carrying the safety profile, a single optional parameter, and an output schema for return values, the description is nearly sufficient. The only real gap is the ambiguous fate of results from runs that already completed before cancellation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single optional parameter with 100% schema description coverage ('Study id (default: the latest)'), so the schema already does the work. The description adds nothing about the study identifier or the default-to-latest behavior, which is the correct baseline for a fully documented schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description pairs a concrete verb sequence (stop, skip, restore) with the resource (an in-progress parameter study), so it is clearly distinguishable from cm_study_start and cm_study_status. It does not, however, explicitly name those siblings to sharpen the contrast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied by the name and title: an agent can infer this is the tool for aborting an in-progress study. There is no statement of when to prefer it over alternatives (e.g., cm_study_status to check state first) or of any prerequisite conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_study_startStart a parameter studyA

Run one test run several times with different parameter values, one after another, in the background. Returns a study id at once; poll cm_study_status. Nothing is written to project files.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNocm4sl: the MATLAB-connected session. standalone: a separate CarMaker process per run (no 'workspace' values)cm4sl
scopeNo'base' for the MATLAB base workspace, or the name of a loaded Simulink model for that model's workspacebase
testrunYesTest run relative to Data/TestRun
max_run_sNoA run is stopped after this wall time in s
save_logsNoAfter every run, save what Simulink logged and the workspaces as that run used them to a MAT file next to its result file (mode cm4sl; logs_file per row)
quantitiesNoQuantities to summarise per run (default: Time, Car.v)
variationsYesOne entry per run: {'label': 'heavy', 'keys': {'Body.mass': 320}, 'workspace': {'Kp_yaw': 1.5}}. 'keys' are infofile keys applied in memory for that run only (prefix a place such as 'Vehicle:' if needed); 'workspace' are MATLAB variables in 'scope', restored after the run. Both are optional

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false; the description usefully adds that execution is backgrounded, that a study id is returned immediately, and that nothing is written to project files. It does not cover limits such as the maxItems=50 cap or how failures mid-study are surfaced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight clauses: what it does, the async/polling contract, and the no-write guarantee. Front-loaded with the core action and free of filler, though slightly telegraphic.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return details need not be explained; the description still notes the immediate study id and the polling path. Combined with full schema coverage and safety annotations, it is essentially complete, missing only the run-count ceiling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so each of the 7 parameters is already documented in the schema with defaults and constraints. The description adds only the notion of sequential variation and no extra syntax or semantic detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: run one test run repeatedly with varied parameter values, sequentially and in the background. That cleanly separates it from single-run siblings such as cm_start_sim, and it names its companion tool cm_study_status explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear operating context (background execution, poll cm_study_status for progress) and implies the async workflow. It does not, however, state when to prefer this over alternatives like a single cm_start_sim or how to abort, so exclusions are absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_study_statusParameter study progress and resultsB
Read-only

State of a study (running, done, cancelled, failed), the run in progress and one row per finished run: end status, simulation time, distance, result file, statistics of the requested quantities.

ParametersJSON Schema
NameRequiredDescriptionDefault
studyNoStudy id (default: the latest)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered without the description. The description adds the study state vocabulary (running, done, cancelled, failed) and the per-run field set, but this is largely restatement of return shape that the output schema already carries. With annotations doing the safety work, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the study state and then enumerates the finish-run fields with no filler. It is appropriately sized for the tool's scope, though the list of return fields is somewhat packed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because an output schema exists, the description needn't explain return values, yet the surrounding context (read-only annotations, single optional param, output schema) is well covered. The primary gap is the absence of when-to-use guidance relative to sibling status/results tools, which keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With a single optional parameter at 100% schema description coverage, the schema fully documents the 'study' argument and its default. The description never mentions the parameter, so it adds no meaning beyond the schema. This is the baseline 3 case where the schema does all the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete resource (a parameter study) and enumerates exactly what state and data are surfaced: overall state, the in-progress run, and one row per finished run. It clearly distinguishes this read of study status from siblings like cm_study_start, cm_study_cancel, and cm_results_list. It stops short of explicitly naming those siblings, so it is clear but not sibling-routed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit statement of when to use this tool versus alternatives such as cm_status, cm_results_list, or cm_results_summary, and no prerequisites or exclusions are given. Usage can only be inferred from the tool name and the return contents. No conditional routing guidance is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cm_wait_endWait for the simulation to endA
Read-only

Wait until the simulation is idle again. finished=true: end status, simulation time and distance, result_file, new log errors and pop-ups. finished=false: current state and live values; call again to keep waiting.

ParametersJSON Schema
NameRequiredDescriptionDefault
poll_sNoSeconds between status polls
timeout_sNoLongest time to wait in seconds; keep it below your client's tool time-out and call again if not finished

Output Schema

ParametersJSON Schema
NameRequiredDescription
liveNo
noteNo
stateNo
popupsNo
testrunNo
finishedNo
waited_sNo
elapsed_sNo
logs_fileNo
distance_mNo
end_statusNo
log_errorsNo
logs_errorNo
sim_time_sNo
poll_errorsNo
result_fileNo
simulink_errorNo
stop_requestedNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint and openWorldHint already covering the safety profile, the description adds real behavioral context: the operation may return incomplete, the finished=false path requires re-invocation, and the completed path surfaces new log errors and pop-ups. This goes beyond what the annotations declare.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose in one clause, then splits the two return states cleanly with zero filler. Every sentence earns its place for a tool of this complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description needn't restate return shapes, and it correctly focuses on the waiting/re-call contract instead. The only real gap is explicit sibling routing against cm_status, which is minor given how clearly the blocking semantics come through.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so poll_s and timeout_s are already fully documented in the schema, and the description adds no syntax or format detail for them. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource combination ('wait until the simulation is idle') that is clearly a blocking operation, which an agent can distinguish from instantaneous tools like cm_status or cm_live. However, it never names those siblings explicitly, so the differentiation is inferred rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives the key operating instruction for the not-finished case ('call again to keep waiting') and implies usage via the two finished branches. But it never states when to prefer this over cm_status or how it relates to cm_stop_sim, leaving the routing decision implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 46 tool updatesv0.1.0
    • First observedcm_changelog
    • First observedcm_clone
    • First observedcm_doctor
    • First observedcm_dva_release
    • First observedcm_dva_write
    • First observedcm_edit
    • First observedcm_get_workspace_var
    • First observedcm_list
    • First observedcm_list_workspace_vars
    • First observedcm_live
    • First observedcm_load_testrun
    • First observedcm_log
    • First observedcm_model_get
    • First observedcm_model_logs_save
    • First observedcm_model_save
    • First observedcm_model_set
    • First observedcm_movie_open
    • First observedcm_movie_snapshot
    • First observedcm_output_quantities
    • First observedcm_output_quantities_edit
    • First observedcm_popup_timeout
    • First observedcm_popups
    • First observedcm_read
    • First observedcm_restore
    • First observedcm_results_list
    • First observedcm_results_read
    • First observedcm_results_summary
    • First observedcm_revert_all
    • First observedcm_session_start
    • First observedcm_set_workspace_var
    • First observedcm_standalone_attach
    • First observedcm_standalone_close
    • First observedcm_standalone_control
    • First observedcm_standalone_dva_write
    • First observedcm_standalone_launch
    • First observedcm_standalone_results
    • First observedcm_standalone_servers
    • First observedcm_standalone_status
    • First observedcm_standalone_wait_end
    • First observedcm_start_sim
    • First observedcm_status
    • First observedcm_stop_sim
    • First observedcm_study_cancel
    • First observedcm_study_start
    • First observedcm_study_status
    • First observedcm_wait_end

TDQS

B3.4/5.0

Scored across 46 tools

Disambiguation4/5

Most tools target distinct resources and actions, and the descriptions are unusually detailed, which helps separate e.g. cm_live from cm_results_read or cm_model_get from cm_get_workspace_var. However, the parallel main-session and standalone-session families, plus several read/get-style tools across different data sources, can still require careful description reading to avoid misselection.

Naming Consistency4/5

Almost all names use a consistent cm_ prefix and snake_case, with clear resource-oriented prefixes such as cm_model_, cm_results_, cm_standalone_, and cm_study_. There is minor inconsistency in verb ordering (e.g. cm_model_get vs cm_get_workspace_var), but the pattern remains highly readable.

Tool Count2/5

46 tools is far above the typical well-scoped range and exceeds the 25+ threshold for a heavy surface. Although CarMaker is a broad domain, many tools could likely be consolidated or grouped, especially the parallel main and standalone lifecycle families.

Completeness4/5

The surface covers simulation lifecycle, project file CRUD via list/read/edit/clone, workspace variables, model parameters, results, studies, standalone instances, logs, GUI popups, and movie snapshots. Minor gaps remain, such as no explicit project-file delete/cleanup operation and limited direct GUI-setting management beyond popup timeout.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    C
    maintenance
    Enables AI agents to automate COMSOL Multiphysics simulations, including model management, geometry building, physics configuration, meshing, solving, and results visualization through the MCP protocol.
    78
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to control MATLAB Simulink models through natural language, providing tools for model creation, block management, wiring, simulation, and more via a local MCP backend.
    3
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI agents to automate COMSOL Multiphysics simulations, including model management, geometry building, physics configuration, meshing, solving, and results visualization via the MCP protocol.
    MIT