Skip to main content
Glama

lsd-mcp

An MCP server that lets an agent work with LSD (Laboratory for Simulation Development, Valente and Pereira): inspect and edit models, compile and run them headless, and run sensitivity analyses (meta-models with Sobol indices, elementary effects). It uses LSD's own engine, command-line utilities (lsd_confgen, lsd_getsaved, ...), design code and R package LSDsensitivity; it does not re-implement them. Built against release tag 8.1-stable-5.

It runs LSD either on the host or, with the Docker backend below, inside the container of Docker_LSD_setup.

Install

Needs Python 3.10+, uv, a C++ compiler (c++, g++ or clang++), zlib, and git (to fetch the LSD source on first use). Rscript with the R package LSDsensitivity is only needed for sa_analyze.

uv sync
uv run pytest -q

Related MCP server: JCGEAgentInterface

Environment variables (all optional)

Variable

Default

Meaning

LSDROOT

unset

An LSD folder containing src/. If unset, the tag is fetched (sparse clone of src, Example, Rpkg).

LSD_TAG

8.1-stable-5

Tag to fetch.

LSD_MCP_HOME

~/.cache/lsd-mcp

Fetched source and all build output.

LSD_MODELS

~/lsd-models

Your models. The only place the edit and run tools write.

LSD_MCP_RSCRIPT

Rscript

R interpreter for sa_analyze.

Tools

Inspect: lsd_status, list_models, read_equations, describe_configuration. Edit: copy_model, create_model, write_equations, edit_structure, set_values, set_run_settings, set_saved. Compile and run: compile_model, run_configuration, read_results. Sensitivity analysis: sa_create_design, sa_run_design, sa_analyze. set_values and the factors of sa_create_design take "Name -k" for the k-th lag of a variable.

create_model makes a new model folder as LSD's model manager (LMM) does: the equation file fun_<name>.cpp from LSD's template, model_options.txt, modelinfo.txt, description.txt and a Sim1.lsd holding only Root. edit_structure applies a list of operations (add object, parameter, variable or function, rename, delete, set the number of instances, set the value of each instance, set a description) to a configuration with LSD's own functions, through lsd_edit (src/lsd_mcp/structure.cpp), and saves with LSD's save_configuration; if one operation fails nothing is written. With these, a model can be built from nothing: create_model, edit_structure, write_equations, run_configuration.

sa_create_design makes one of four designs. lhs and random are sampled in Python and written through lsd_confgen. nolh (near-orthogonal Latin hypercube, LSD's own tables, optionally extended) and ee (elementary effects, Morris; trajectories, levels, jump, pool) are made by LSD's own design code (design and sensitivity_doe in set_all.cpp) through lsd_doe, a small program of ours (src/lsd_mcp/doe.cpp) that includes LSD's file unmodified and makes the calls the interface makes. A nolh design gets an out-of-sample set from LSD's Monte Carlo range sampling, as the interface offers after NOLH; an ee design has none. Every design also gets <config>_design.json (method and parameters). sa_analyze fits a Kriging or polynomial meta-model with Sobol indices, or, for an ee design, runs LSD's elementary effects analysis and returns mu, mu_star, sigma, se and p_value per factor. For an ee design made in LSD's interface (no design file) pass metamodel="ee" with its levels and jump.

Register with Claude Code

claude mcp add lsd -- uv run --project /path/to/lsd-mcp lsd-mcp

Add -e LSD_MODELS=/path/to/models to change the models folder.

Docker backend

With LSD_MCP_BACKEND=docker the server still runs on the host, but every tool call is executed inside a running LSD container by the same standard-library code, which lsd-mcp copies into the container (docker cp, re-copied when the source changes). It never starts, stops or removes a container.

The container comes from Docker_LSD_setup, a separate repository that runs LSD in Docker with a browser desktop, R and LSD's R packages. This backend needs nothing installed on the host besides Docker and uv: the compiler, LSD and R are the container's.

Variable

Default

Meaning

LSD_MCP_BACKEND

local

docker forwards tool calls into the container.

LSD_CONTAINER

lsd

Name of the running container.

LSD_MCP_CONTAINER_LSDROOT

/home/lsd/LSD

LSD folder inside the container.

LSD_MCP_CONTAINER_MODELS

/home/lsd/LSD/Work

Models folder inside the container.

claude mcp add lsd-docker -e LSD_MCP_BACKEND=docker -- uv run --project /path/to/lsd-mcp lsd-mcp

Start the container with ./run.sh in your Docker_LSD_setup folder first. Two things to know: models live in the container's Work folder, which is shared with the host and shown in LMM as "Work in Progress" (copy_model adds a modelinfo.txt so LMM lists the copy); and the build cache inside the container is lost when the container is recreated, so the next call rebuilds it.

Limits

  • Kriging can fail with "the leading minor ... is not positive" (covariance matrix not positive definite), typically when design points nearly coincide, for example integer factors with few levels. sa_analyze explains this; try the polynomial meta-model or more spread-out points. sa_create_design warns about integer factors with fewer levels than samples / 4.

  • LSD's polynomial meta-model weights design points by mean/SD of the response and fails when a point has a negative mean (a bug in LSDsensitivity, not worked around here). sa_analyze says so; use metamodel="kriging".

  • LSD's polynomial meta-model also needs at least two factors (its package builds a broken formula for one); sa_analyze refuses it for a single factor.

  • LSD cannot load a configuration with SEED below 1 (27 shipped examples have SEED 0); the tools say so and set_run_settings(seed=1) fixes it.

  • Totals files (.tot.gz) have no header and one row per run; read_results does not read them and names the .res.gz files instead.

  • Sensitivity analysis covers Latin hypercube, uniform random and NOLH designs with a Kriging or polynomial meta-model, and elementary effects designs.

  • Factors are parameters or the initial values of variables ("X" is the first lag, "X -2" the second). One element can be one factor only: LSD's design table names a factor by its label alone. In lsd_confgen's own files a negative lag always means the first lag (confgen.cpp, change_configuration), so the tools send it the positive number it does honour.

  • An elementary effects design made on macOS differs from one made on Linux (including in Docker) for the same seed, because LSD shuffles trajectories with the C++ standard library's shuffle, which differs between libc++ and libstdc++. Both are valid designs. On Linux the design is byte-identical to what LSD's interface writes for the same settings and seed; NOLH and Monte Carlo range designs are identical on both platforms (tests/data/doe_gui holds the interface's files, tests/test_doe.py and the Docker tests compare them).

  • edit_structure saves with LSD's save_configuration, so the file takes LSD's current layout. Of the 151 shipped example configurations LSD can load, 32 come back byte-identical from an empty edit. The rest differ only in layout (no data): descriptions without text were written as "(no description available)" in older files and are now empty, MODELREPORT lost a leading space, an empty EQ_FILE section is added, and the note LSD generates about the initial values of a variable with no lags is dropped (LSD's save_description). If the equation file is not in the model folder LSD would blank the EQUATION name; the tool puts it back. New instances copy the object's first instance. Objects always keep at least one instance.

  • Models made by create_model use LMM's compiler options (SWITCH_CC=-O0 -ggdb3), so they compile without optimisation; edit model_options.txt for speed.

  • NOLH tables cover 1 to 100 factors (17 to 512 points); samples does not apply.

  • Without the Docker backend the compiler and utilities run on the host.

  • Models are compiled headless (-D_NW_). Eight of the 45 example models use GUI-only features (Tcl calls, the debugger variables) and do not compile this way.

  • lsd_confgen in 8.1-stable-5 generates at most as many configurations per call as the CSV has rows; sa_create_design splits the work accordingly.

  • lsd_confgen drops everything after MODELREPORT (descriptions, embedded equations). set_values appends it again from the original file; the numbered design configurations are left as LSD writes them.

  • sa_analyze is tested on Linux with R 4.3.3 and LSDsensitivity 1.2.3 (the tests are skipped where R is missing).

Available Tools

17 tools
compile_modelB

Compile the model's equation file into a headless LSD program (built into the cache, not the model folder). Returns ok and build time, or the first compiler errors as file:line: message. Nothing is rebuilt if the equation file is unchanged.

ParametersJSON Schema
NameRequiredDescriptionDefault
groupNomodels
modelYes

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses where output lands (cache, not model folder), the return shape (ok and build time, or file:line: message errors), and the no-op-if-unchanged caching behavior. It does not mention permissions or whether any artifacts can be invalidated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action and artifact, followed by return values and caching semantics. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description rightly explains the return values and error format. For a two-parameter tool with no annotations, the only real gap is documentation of the 'group' parameter, which leaves the definition slightly short of fully self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description never mentions the parameters. 'model' is arguably implied by 'the model's equation file', but 'group' is completely undocumented in both schema and description, leaving the agent to guess its role.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Compile) and resource (the model's equation file) with a concrete output artifact (headless LSD program). It is not confusable with siblings like write_equations or run_configuration, though it does not explicitly name an alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use guidance or prerequisites. It does not say when compilation is needed relative to write_equations or run_configuration, nor when to skip it. The incremental-build note hints that repeated calls are safe but is behavioral, not usage routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

copy_modelB

Copy a model folder (source files only: no binaries, results, backups, numbered design configurations, design tables or _sa folders) into the models folder as name, so it can be edited. Fails if name exists.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
sourceYes
source_groupNoexamples

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it enumerates excluded artifacts (binaries, results, backups, design configurations, sa folders) and states the failure mode when `name` exists. It omits permissions, side effects on the source, and any rate/limits, but the mutation semantics are reasonably disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The action is front-loaded and packed into a single tight sentence with a scoping parenthetical; nothing is wasted, though the nested exclusion list slightly buries the destination and failure clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No annotations and no output schema, so the description must cover behavior and returns. It covers what gets copied and the failure condition, but never states what the tool returns (new path/name) or explains source_group, leaving gaps for a 3-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains `name` (the destination name) and implies `source`, but `source_group` (default 'examples') is never mentioned, leaving a required-understanding parameter undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource (copy a model folder) and adds precise scope detail about what is and isn't copied, plus the destination and name. It implicitly distinguishes itself from create_model ('copy ... so it can be edited') but never names a sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'so it can be edited' implies the intent context, and 'Fails if `name` exists' hints at a precondition, but there is no explicit when-to-use vs alternatives (e.g., versus create_model) or when-not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_modelA

Create a new, empty model folder in the models folder, as LSD's model manager (LMM) does: equation file fun_.cpp from LSD's template, model_options.txt, modelinfo.txt (LMM lists a folder only with it), description.txt, and a configuration Sim1.lsd holding only Root. name is the folder (letters, digits, underscores; it must not exist or lie inside another model). Then use edit_structure to add objects, parameters and variables, write_equations for the equations, set_saved if needed, and run_configuration. An empty model compiles and runs (it computes nothing and saves nothing).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
titleNo
descriptionNo

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses the side effects (which files are created: fun_<name>.cpp, model_options.txt, modelinfo.txt, description.txt, Sim1.lsd), the naming constraint (must not exist or lie inside another model), and that an empty model compiles and runs without computing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and the generated-file list, then the workflow. Dense and mostly waste-free, though the file enumeration and workflow step run somewhat long for a creation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema or annotations exist, but the description compensates by covering the created artifacts, naming constraints, and downstream workflow. A brief note on the return value or error behavior on name collision would make it fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It fully explains `name` (the folder name, allowed characters, existence/nesting constraint) but says nothing about `title` or `description` parameters; description.txt is mentioned as a file but not linked to the `description` parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Create') and resource ('new, empty model folder in the models folder') and enumerates exactly what files are generated. It is clearly distinguishable from sibling creation tools like copy_model or sa_create_design.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly chains the next steps ('Then use edit_structure to add objects, parameters and variables, write_equations for the equations, set_saved if needed, and run_configuration'), naming the exact siblings and the order, so an agent knows this is only the first step of a pipeline.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

describe_configurationA

Describe a .lsd configuration: the object tree with instance counts, and for each element (variable, parameter or function) its type, number of lags, whether it is saved to the result files, and its value (a single value if all instances are equal, otherwise min, max and count), plus the run settings SIM_NUM, SEED, MAX_STEP and EQUATION. config is the file name with or without .lsd. Set only when they apply: computed=false on objects (not computed), and debug, plot, parallel, saved_separately on elements. For big models pass object='Name' to describe one object, or detail='names' for the tree with instance counts and just the element names grouped by type.

ParametersJSON Schema
NameRequiredDescriptionDefault
groupNomodels
modelYes
configYes
detailNofull
objectNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does reasonably well: it discloses that values collapse to a single number when all instances are equal and otherwise report min/max/count, and that certain attributes appear only conditionally. It never explicitly says the operation is read-only or side-effect free, but the detail about returned content is above the annotation-free baseline.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but front-loads the core purpose (what gets described) before the parameter guidance, and most sentences carry real information. The conditional-attributes sentence is somewhat confusing and could be trimmed, but there is little outright waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must explain the return shape, which it does thoroughly for a complex config-inspection tool. Gaps remain around the group/model parameters and the read-only nature, but an agent has enough to call it correctly for most cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it largely does: config is explained as the file name with or without '.lsd', object selects a single object, and detail='names' alters output granularity. group and model are left unexplained, and the 'Set only when they apply: computed=false ... debug, plot, parallel, saved_separately' sentence ambiguously mixes output attributes with settable fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Describe') and resource ('.lsd configuration') and enumerates exactly what is returned: the object tree with instance counts, per-element type/lags/saved flag/value summaries, and run settings. This is clearly distinguishable from siblings like read_equations, read_results, or list_models.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives practical guidance for large models ('pass object=Name ... or detail=names'), which is useful invocation context, but never states when to prefer this tool over alternatives such as read_equations or read_results, nor any exclusions. Usage is implied through parameter selection rather than explicit routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edit_structureA

Change a configuration's structure (objects, parameters, variables, functions, instance counts) with LSD's own code. operations is a list of objects applied in order; if any fails (the message names it) nothing is written. Writes new_config.lsd, or replaces config.lsd (old file kept as .bak). Names must be valid LSD labels (letter or '' first, then letters, digits, '') and unique across the whole model, objects and elements alike. Operations, one example each: {"op": "add_object", "parent": "Root", "name": "Firm", "instances": 10} {"op": "add_parameter", "object": "Firm", "name": "alpha", "value": 0.5} {"op": "add_variable", "object": "Firm", "name": "K", "lags": 1, "initial": 1.0, "saved": true} (initial: a number, or one per lag) {"op": "add_function", "object": "Firm", "name": "f"} {"op": "rename", "name": "old", "new_name": "new"} (object or element) {"op": "delete", "name": "K"} (element; an object with elements or child objects needs "force": true and goes with all of it) {"op": "set_instances", "object": "Firm", "instances": 50} {"op": "set_instance_values", "name": "alpha", "values": [0.1, 0.2], "lag": 2} (one value per instance in file order; lag only for variables) {"op": "describe", "name": "alpha", "text": "adjustment speed"} Parameters and variables are added to every instance of the object, with the value (default 0) in each. set_instances gives the object that number under every parent instance; new instances are copies of the object's first instance (its values and its child objects), extra ones are removed from the end; at least 1. Equations are not touched: add or change them with write_equations. LSD's save rewrites the file in its own layout, so old files may change in layout (empty descriptions, the notes LSD generates for initial values of variables with no lags are dropped), never in values.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYes
configYes
new_configNo
operationsYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so richly: atomic all-or-nothing application, failure naming, file write behavior with .bak backup, label validation, uniqueness constraints, instance copying/removal semantics, and LSD save layout changes. These are exactly the mutation and side-effect details an agent needs before invoking it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded, and the remaining length is justified by the complex operation vocabulary and file side effects. Examples are compact and each sentence adds necessary detail rather than repeating schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity mutation tool with no annotations, no output schema, and no schema descriptions, the description is nearly complete. It omits an explicit success-return description and does not define the `model` parameter, but otherwise covers operation semantics, failure behavior, naming rules, and file side effects well.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain all four parameters. It exhaustively documents the `operations` parameter with examples, and indirectly clarifies `config`/`new_config` file behavior, but it never defines the required `model` parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Change a configuration's structure' and then enumerates the exact structural elements affected. It also distinguishes itself from write_equations by saying equations are not touched and must be handled separately.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context for structural changes and explicitly routes equation edits to write_equations. However, it does not compare against several sibling tools such as set_values or set_saved, leaving some routing decisions to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_modelsA

List models. group='models' is the user's LSD_MODELS folder; group='examples' is the Example folder of the LSD distribution (read-only). A model is a folder with an equation file (fun_*.cpp) and at least one .lsd configuration. Returns relative path, title, equation file, source files (*.cpp, *.h, *.hpp in the folder) and configuration names for each.

ParametersJSON Schema
NameRequiredDescriptionDefault
groupNomodels

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose meaningful behavior: the 'examples' group is read-only, while 'models' is the user's own folder, and it enumerates what is returned for each entry. It stops short of covering edge cases such as empty folders or malformed models.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the action, with no filler. The return-field enumeration is somewhat long, but each element is informative and the group semantics come before it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the explicit list of returned fields is partly redundant, but the description still supplies the domain framing (what a model is, what the two groups mean) that the schema cannot. For a simple read-only list tool this is close to sufficient, missing only failure or empty-result behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single 'group' parameter has no enum, yet the description fully documents both meaningful values and their backing folders, effectively supplying the enumeration the schema lacks. An agent can select a value correctly without any additional source.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List models') and goes further by defining what a model actually is (a folder with a fun_*.cpp equation file and at least one .lsd configuration). That definition usefully separates it from siblings like describe_configuration or read_equations, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the group values tell the agent which folder it is browsing, and the '(read-only)' note hints that 'examples' is for inspection only. There is no explicit guidance on when to reach for this versus describe_configuration or read_equations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lsd_statusA

Show the setup: LSD source root and tag, models folder, whether a C++ compiler was found, whether LSD's command-line utilities are built yet (they are built automatically on first use), and whether Rscript and the R package LSDsensitivity are available.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It does add real behavioral context by noting that LSD's command-line utilities are built automatically on first use, which tells the agent about a side effect of other operations. However, it does not state that lsd_status itself is read-only/non-mutating, nor how quickly it returns or what it does if components are missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single front-loaded sentence beginning 'Show the setup:' followed by a clean enumerated list of checks, so an agent gets the essentials immediately. The awkward line wrapping and slightly run-on structure cost a point but no sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description effectively stands in for one by enumerating the fields the agent will see (source root/tag, models folder, compiler status, utility build state, R availability). That is sufficient for a zero-parameter diagnostic tool, though finer detail on result format or failure signaling is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate and the baseline is 4. Schema coverage is 100% for an empty object, consistent with no per-argument documentation being needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('Show the setup') and enumerates exactly what is reported: LSD source root/tag, models folder, C++ compiler presence, CLI utility build state, and R/Rscript availability. This makes the tool's function unmistakable. It stops short of naming sibling tools like describe_configuration or list_models for explicit differentiation, though the enumerated content is distinct enough in practice.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the mention that utilities are 'built automatically on first use' hints this is a pre-flight/diagnostic check, but the description never states when to call it (e.g. before compiling or running a configuration) or what alternatives exist. Adequate but leaves the agent to infer the trigger condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_equationsA

Return the text of the model's equation file (the C++ source with the EQUATION(...) blocks). Models whose equations are spread over several files list them as source_files in list_models; pass such a plain file name (.cpp, .h or .hpp, no path) as file to read it instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileNo
groupNomodels
modelYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. The verb 'Return' and the file-resolution explanation make the read-only nature evident, and it explains the multi-file lookup behavior. However, it says nothing about failure modes (missing model, invalid file name) or any access constraints, so a middle score is warranted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the primary behavior and followed by the conditional multi-file case. The parenthetical detailing EQUATION(...) blocks earns its space by telling the agent what the returned text actually looks like. Slightly long second sentence but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the main invocation path (model, plus file for multi-file models) is fully covered. The one gap is the `group` parameter, which is never mentioned and has 0% schema documentation, leaving the agent to guess at its role.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the parameters. It adds real meaning for `file` (accepted extensions .cpp/.h/.hpp, no path allowed), which the schema does not convey. But `group` (default 'models') and `model` are left entirely undocumented, so it only compensates for a third of the parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Return the text of the model's equation file') and pins down exactly what kind of content comes back (the C++ source with EQUATION(...) blocks). It is clearly distinguishable from siblings like list_models, write_equations, and read_results.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: multi-file models list their equations as source_files in list_models, and those plain file names should be passed as `file`. That is a concrete when-to-use rule tied to a sibling tool's output, though the read-vs-write relationship with write_equations is only implied by the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_resultsA

Read time series from one result file (.res.gz) in the model folder. variables filters by name ('Mean') or name with instance ('Mean 1'); start/end select time steps (the row number in the file is the time step; row 0 holds initial values; start must not exceed end; max_points >= 1); the series are thinned evenly to max_points. At most 50 series are returned.

ParametersJSON Schema
NameRequiredDescriptionDefault
endNo
modelYes
startNo
variablesNo
max_pointsNo
results_fileYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the 50-series cap, the even thinning behavior driven by max_points, the row-0-is-initial-values convention, and the constraints start <= end and max_points >= 1. It omits error behavior (missing/uncompiled model) and confirms nothing explicit about read-only safety, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then dense per-parameter guidance with no filler. Every clause carries usable information; the only minor cost is that the parameter notes run together as a long semicolon chain rather than being visually separated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-param tool with no annotations and no output schema, the description covers the time-axis semantics, filtering, thinning, and the return-size limit. It does not describe the shape of the returned series, which is the main remaining gap given the absence of an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it explains variables (name or name+instance), start/end (row number = time step, row 0 = initial, start <= end), and max_points (thinning target, >=1). Only 'model' and 'results_file' lack explicit semantics, though both are largely self-evident from the opening sentence.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource: 'Read time series from one result file (.res.gz) in the model folder.' The file type and location are concrete enough to distinguish it from siblings like read_equations or list_models. It stops short of explicitly naming an alternative tool, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied (inspect simulation output stored in a .res.gz file), but there is no explicit when-to-use/when-not guidance and no named alternatives among the many siblings. An agent can infer the context but must reason about it rather than being told.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_configurationA

Compile if needed and run a configuration of a model in the models folder. seed and runs override SEED and SIM_NUM of the file. threads: with runs > 1, the number of runs executed in parallel; otherwise threads for models that use parallel objects. Results are written next to the configuration as .res.gz. Returns the files written and, for the first run, the last value and mean of each saved series (at most 50 series, those whose name has few instances first; the rest are counted). With runs > 1 (sequential) LSD also writes _.tot.gz; with threads set the runs are parallel and no totals file is written. Row 0 of a result file holds initial values. On failure returns the tail of LSD's output.

ParametersJSON Schema
NameRequiredDescriptionDefault
runsNo
seedNo
modelYes
configYes
threadsNo
timeout_sNo

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so richly: compile-if-needed semantics, seed/runs overriding file SEED and SIM_NUM, parallel vs sequential run behavior, exact output filenames, the .tot.gz totals file being written only in sequential mode, row 0 holding initial values, and failure returning the tail of LSD output. This is unusually complete disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded, leading with the primary action and following with behavioral detail. A few clauses (the series-counting parenthetical, totals-file note) are slightly rambling but each earns its place by conveying output behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 6-parameter tool with no output schema, the description explains the return payload (files written, last value and mean of saved series) and failure mode, which is the key gap. The omission of timeout_s and explicit model/config semantics is the only shortfall.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains seed, runs, and threads in meaningful detail, but leaves model, config, and especially timeout_s undocumented, so roughly half the parameters lack semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific compound action (compile if needed and run) on a precise resource (a configuration of a model in the models folder). It clearly separates itself from compile_model (only compiles) and read_results (reads output), so an agent can distinguish it from siblings without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the workflow (compile-if-needed then run) and explains how runs/seed/threads change behavior, but never explicitly states when to choose this over sa_run_design or set_run_settings, nor any preconditions. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sa_analyzeA

Analyse the design results for one saved variable with LSD's R package LSDsensitivity. metamodel 'kriging' (default for lhs, random and nolh designs) or 'polynomial' fits a meta-model and computes its Sobol decomposition: returns the fit quality (Q2 for kriging, R2 for polynomial) and a table with direct effects and interactions per factor. For a design made with method 'ee' the analysis is elementary effects (metamodel 'ee', chosen automatically): returns per factor mu, mu_star, sigma, se and p_value (parameters scaled to [0, 1]; mu_star is the overall effect, sigma non-linear or interaction effects, p_value tests mu_star = 0), sorted by mu_star. Kriging or polynomial on an ee design, or ee on another design, is an error. For an ee design made in LSD's own interface (no design file) pass metamodel='ee' with its levels and jump. ini_drop drops initial time steps, n_keep keeps that many (-1 = all). The response is the mean of the variable over the kept steps, averaged over runs. Variables whose name starts with '_' work; with several instances only the first instance is analysed. ini_drop must be below MAX_STEP and ini_drop + n_keep at most MAX_STEP. r_seed seeds R's random numbers, so identical calls give identical results. A warning is added when a meta-model fit is below 0.5. The polynomial meta-model fails when a design point has a negative mean response (LSD weights points by mean/SD) and needs at least two factors; use kriging then. Kriging can fail numerically when design points nearly coincide (the message says so; try polynomial). Needs Rscript and LSDsensitivity; says so if they are missing.

ParametersJSON Schema
NameRequiredDescriptionDefault
jumpNo
modelYes
configYes
levelsNo
n_keepNo
r_seedNo
ini_dropNo
variableYes
metamodelNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: it discloses determinism via r_seed, dependency failures (Rscript/LSDsensitivity), a sub-0.5 fit warning, polynomial failure on negative-mean responses, kriging numerical failure on coincident points, and edge behavior for '_'-prefixed variables and multiple instances.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core verb and the two metamodel branches before constraints, which is the right ordering. The prose is dense and run-on in places, but nearly every clause carries a behavioral fact an agent needs, so the length is largely earned for a 9-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with no output schema and no annotations, the description supplies the return shapes (Q2/R2, effects/interactions table, mu/mu_star/sigma/se/p_value), the scaling and sorting of ee output, and the operational preconditions. Nothing essential to invoking or interpreting it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

At 0% schema coverage the description must compensate, and it explains metamodel, ini_drop, n_keep, r_seed, levels and jump with real meaning, including the MAX_STEP constraint and -1 sentinel. It leaves model, config and variable essentially defined only by surrounding prose, so one or two parameters remain under-specified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Analyse the design results for one saved variable with LSD's R package LSDsensitivity.' It clearly delineates this as the analysis step, distinct from the design/run siblings (sa_create_design, sa_run_design), so an agent can place it without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives strong conditional guidance: kriging/polynomial for lhs/random/nolh, 'ee' chosen automatically for an ee design, and that mixing metamodels with the wrong design is an error. It also routes around failures ('use kriging then', 'try polynomial'). It stops short of explicitly naming sibling tools or stating the prerequisite ordering relative to sa_run_design.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sa_create_designA

Create a design of experiments for a sensitivity analysis and write it in LSD's own file layout in the model folder: .sa, the design table _1_N.csv (and for meta-model designs the out-of-sample table _N+1_N+V.csv), numbered configurations _1.lsd ..., and _design.json (our file: method and parameters). factors maps an element to [min, max], or [min, max, "int"] for integers. An element is a parameter name, or a variable name for the variable's initial value at its first lag, or "Name -2", "Name -3" ... (name, space, negative lag) for the initial value at that lag. Functions, variables with no lags and lags beyond the variable's lags are refused, and an element can be a factor only once (LSD's design table names a factor by its label, so the result files and the analysis tables show the plain name, for example "X"; the design file _design.json records each factor's lag, 0 for a parameter). method: 'lhs' (Latin hypercube) or 'random': samples points (required, at least 2) plus validation_samples uniform out-of-sample points, for a Kriging or polynomial meta-model. 'nolh': near-orthogonal Latin hypercube made by LSD's own NOLH tables. LSD chooses the table from the number of factors (17 points for 1-7 factors, 33 for 8-11, 65 for 12-16, 129 for 17-22, 257 for 23-29, 512 for 30-100; extended=True uses LSD's extended size for the table: 33, 65, 129, 257, 257, 512), so samples is not used; the result says how many points LSD produced. Plus validation_samples out-of-sample points by LSD's Monte Carlo range sampling. For meta-model analysis. 'ee': elementary effects (Morris) design made by LSD's own code: trajectories (default 10) each of factors + 1 points, chosen from a pool of pool random trajectories (default 100) on levels levels (even, default 4) with jump (default 2). No out-of-sample set; validation_samples and samples are ignored. An ee design made on macOS differs from one made on Linux for the same seed (the C++ library's shuffle differs); both are valid designs. Analyse it with sa_analyze (metamodel 'ee'). Each point runs runs_per_point times (at least 2) with its own seeds. seed seeds the sampling. Refuses to replace an existing design unless overwrite=True. A warning is returned for integer factors with few levels (a hypercube collapses onto them). Variables whose results will be analysed must be saved (set_saved).

ParametersJSON Schema
NameRequiredDescriptionDefault
jumpNo
poolNo
seedNo
modelYes
configYes
levelsNo
methodNolhs
factorsYes
samplesNo
extendedNo
overwriteNo
trajectoriesNo
runs_per_pointNo
validation_samplesNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden and does so: it discloses the exact files written, the refusal of functions/laggless variables/duplicate factors, the overwrite guard, the integer-factor warning, platform-dependent ee designs, and which parameters are ignored per method. This is far beyond what any annotation would supply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and appropriately sized for 14 parameters and four design methods; every clause carries operational detail. Readability suffers from being one dense paragraph without list formatting, but there is little dead weight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, no-annotation, no-output-schema tool with nested parameters, the description covers method selection, output artifacts, failure modes, and per-method parameter behavior. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the description compensates thoroughly — it defines the factors map format ([min,max] / [min,max,'int']), element naming including negative lags, and the meaning and defaults of samples, validation_samples, extended, trajectories, pool, levels, jump, seed, runs_per_point, and overwrite. Nearly every one of the 14 parameters gains semantics found nowhere in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Create a design of experiments for a sensitivity analysis') and immediately distinguishes itself from siblings by naming the artifacts it produces and the analysis tools it feeds (sa_analyze, metamodel 'ee'). An agent can tell this apart from sa_run_design and sa_analyze without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives rich per-method context — 'lhs'/'random' for Kriging or polynomial meta-models, 'nolh' 'for meta-model analysis', 'ee' to be analysed with sa_analyze — plus the overwrite refusal condition. It stops short of an explicit 'use this instead of X when Y' routing against sa_run_design, but the when-to-use guidance for each mode is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sa_run_designB

Run every numbered configuration of the design in parallel processes (default: one per CPU). Points whose result files already exist are skipped. Returns counts of points done, run, failed and not started.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYes
configYes
threadsNo
timeout_sNo

TDQS

B3.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden and does so reasonably: it discloses parallel execution, the default thread count (one per CPU), skip-if-result-exists idempotency, and the returned counts (done/run/failed/not started). It omits any note on timeout behavior, permissions, or side effects of failed points.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with no filler, front-loaded with the primary action and scope before the skip behavior and return values. Efficient and well-ordered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the return values in lieu of an output schema and explains the parallel/skip behavior, which is good. But for an execution tool with no annotations and zero parameter documentation, it leaves the agent without guidance on the required model/config inputs or the timeout parameter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across four parameters, and the description only partially compensates by explaining the threads default (one per CPU). The meaning of model, config, and especially timeout_s (default 3600) is left entirely undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ("Run every numbered configuration of the design") and clarifies scope as the whole design rather than a single configuration. It implicitly contrasts with the sibling run_configuration (singular), but never names it, so the differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance or alternatives are given. With both run_configuration and sa_run_design present as siblings, the agent must guess which to pick; the description only hints via "every numbered configuration" vs a single one, without stating the selection rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_run_settingsC

Edit SIM_NUM (runs), SEED and MAX_STEP (time steps) of a configuration.

ParametersJSON Schema
NameRequiredDescriptionDefault
runsNo
seedNo
modelYes
stepsNo
configYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Edit' implies mutation but nothing is said about persistence (is the config file rewritten?), required permissions, reversibility, or what happens to unspecified settings. For a mutating tool with zero annotation coverage this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the affected settings front-loaded and no filler. Brevity is appropriate, though it comes at the cost of completeness rather than being optimally balanced.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter mutation tool with no annotations, no output schema, and 0% schema description coverage, the description is too thin: it omits the meaning of the two required parameters, prerequisites, and any indication of the effect of the edit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% with five parameters, so the description must compensate. It usefully maps runs→SIM_NUM, steps→MAX_STEP and seed→SEED, but leaves the required 'model' and 'config' parameters entirely unexplained, so half the interface remains opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (Edit) plus resource (run settings of a configuration), and it names the exact settings touched (runs, seed, steps). It does not distinguish itself from siblings like set_values or set_saved, which also mutate configuration data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus set_values, set_saved, or run_configuration, and no prerequisites (e.g., whether the model must be compiled or the config must exist) are stated. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_savedB

Mark the named variables or parameters as saved (or not saved) to the result files. Only saved elements appear in results.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYes
namesYes
savedNo
configYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses the downstream effect ('Only saved elements appear in results'), which is real behavioral context beyond the name. However, it does not say whether setting saved=false removes previously written results, whether the change persists across runs, or what permissions are needed for a mutation operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and followed by the consequence. No filler, though the phrasing 'saved (or not saved)' is slightly awkward given the parameter already carries the polarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter mutation tool with 0% schema description coverage, no annotations, and no output schema, the description is thin. It omits what model/config refer to, whether the operation is reversible, and how it interacts with existing result files, so an agent cannot call it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It hints at the 'names' parameter ('named variables or parameters') and the boolean toggle ('saved (or not saved)'), which maps to the 'saved' parameter and its default. But 'model' and 'config' — two of the three required parameters — are never explained, leaving the highest-risk inputs undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('mark ... as saved') and resource ('named variables or parameters'), and clarifies the scope ('to the result files'). It does not explicitly distinguish itself from close siblings like set_values or set_run_settings, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or when-not-to-use guidance, and no alternative tool is named. The clause 'Only saved elements appear in results' implies the context in which this matters, which is more than nothing but still leaves the agent to infer when to reach for this versus set_values or set_run_settings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_valuesA

Set element values in a configuration, using LSD's own lsd_confgen. values maps element name to a number; every instance of the element gets the value. A variable's name sets its initial value at the first lag; "Name -2", "Name -3" ... (name, space, negative lag) set the second, third ... lag ("Name -1" is the first); a lag beyond the variable's lags, or on a parameter, is an error, and one element can appear once per call. Writes new_config.lsd, or replaces config.lsd (the old file is kept as .bak). The result says whether an existing file was replaced (replaced, backup).

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYes
configYes
valuesYes
new_configNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral burden. It discloses important side effects: writes new_config.lsd or replaces config.lsd (keeping a .bak), and the result indicates whether replacement occurred. It also states error conditions (lag beyond variable's lags, parameter lag, repeated element per call). Missing: permissions required, atomicity, or what happens on partial failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense with useful details but somewhat disorganized; the lag-suffix explanation is verbose and interleaved with error conditions. It could be front-loaded better (e.g., start with the write behavior) but every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description covers key behavioral aspects: write destination, backup behavior, return information, and error conditions. It lacks the return type details (e.g., exact fields of result) and parameter explanations for 'model', 'config', and 'new_config', but is otherwise fairly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description compensates well: it explains that 'values' maps element names to numbers, describes the special lag suffix syntax ('Name -2'), and notes constraints on parameter usage. However, it does not explain 'model', 'config', or 'new_config' parameters clearly (though their names are somewhat self-explanatory).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb+resource: setting element values in a configuration, and names the underlying mechanism (lsd_confgen). It does not explicitly differentiate from siblings like set_run_settings or edit_structure, but the resource (element values) is distinct enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the description of how values map to elements and lags, but there is no explicit when-to-use versus alternatives, no mention of prerequisites, and no guidance on when to choose new_config vs. overwriting config.lsd.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

write_equationsA

Replace the model's equation file with content (complete C++ source), or the source file named by file (plain name, .cpp, .h or .hpp; a new file is created if it does not exist). The version before the latest write is kept as .bak and the one before the first write as .orig. Only models in the models folder can be written.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileNo
modelYes
contentYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and does well: it discloses that writes are destructive ('Replace'), that a new file is created if absent, that backups are kept as <file>.bak and <file>.orig, and that writes are scoped to the models folder. It omits auth/permission requirements and error behavior, but the backup and scoping disclosures are substantial for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action and the content-vs-file choice are front-loaded, followed by the backup semantics and the scope constraint. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive three-parameter write tool with no annotations and no output schema, the description covers the essential behaviors an agent needs: what gets replaced, file naming/creation, backup retention, and scope restriction. It could go further on permissions and the outcome of a failed write, but it is largely self-sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: `content` is defined as complete C++ source, `file` as a plain name with .cpp/.h/.hpp and create-if-missing semantics, and `model` is constrained via the models-folder rule. It does not explain the interaction/default when `file` is omitted (where the equations file lives), leaving a small gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Replace the model's equation file' with either inline content or a named source file. It clearly distinguishes the two write modes. It does not explicitly contrast with the read counterpart (read_equations) among its many siblings, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains the conditions governing the two write paths (content vs file) and adds the constraint that only models in the models folder can be written, which is genuinely useful context. However, it gives no explicit when-to-use-this-vs-alternative guidance (e.g., when to use write_equations versus edit_structure).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 17 tool updatesv0.1.0
    • First observedcompile_model
    • First observedcopy_model
    • First observedcreate_model
    • First observeddescribe_configuration
    • First observededit_structure
    • First observedlist_models
    • First observedlsd_status
    • First observedread_equations
    • First observedread_results
    • First observedrun_configuration
    • First observedsa_analyze
    • First observedsa_create_design
    • First observedsa_run_design
    • First observedset_run_settings
    • First observedset_saved
    • First observedset_values
    • First observedwrite_equations

TDQS

A3.6/5.0

Scored across 17 tools

Disambiguation4/5

Most tools have clearly distinct scopes (model management, configuration editing, execution, and sensitivity analysis). Slight overlap exists between set_values and edit_structure's set_instance_values operation, and between set_run_settings and run_configuration's seed/runs overrides, but descriptions distinguish them well.

Naming Consistency4/5

Predominantly snake_case verb_noun names such as list_models, read_equations, and create_model. Deviations are the sa_ prefix for sensitivity-analysis tools and lsd_status, but these are readable and consistent within their subgroups.

Tool Count4/5

With 17 tools, the set is slightly above the ideal 3-15 range but the domain is broad and each tool covers a distinct workflow step. The count is reasonable, though a couple of configuration-editing tools could potentially be merged.

Completeness4/5

Coverage is strong across the model lifecycle (create, copy, read/write equations, edit structure, compile, run) and sensitivity analysis (design, run, analyze). Minor gaps include no delete_model/delete_configuration and no explicit list_results, which agents can partially work around.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables MCP-compatible agents to generate Qucs circuit schematics, run simulations, and parse results programmatically.
    1
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI hosts to design and simulate complex adaptive systems, including managing agent types, running simulations, and analyzing emergent behavior, feedback loops, and sensitivities.
    1
    MIT
  • A
    license
    C
    quality
    C
    maintenance
    Enables AI agents to read and write .mph model files, control the COMSOL desktop GUI in real time, and run complete modeling, meshing, solving, and post-processing workflows.
    23
    1
    MIT