openfoam-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@openfoam-mcprun the motorBike tutorial with 6 cores and plot the drag coefficient"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
openfoam-mcp
An MCP server that lets AI coding agents drive OpenFOAM end to end — clone a tutorial, edit dictionaries, mesh, run solvers in the background, watch convergence, and look at the flow — from Claude Code, Codex CLI, Gemini CLI, Cursor, VS Code, Windsurf, Claude Desktop or opencode.
The motorBike tutorial, driven entirely through the server's tools: create_case → import_file (geometry) → run surfaceFeatures / blockMesh → run snappyHexMesh np=6 → check_mesh → run foamRun np=6 → solver_progress → read_postprocessing (forceCoeffs) → render.
Surface pressure ( | Wake slice ( |
|
|
Drag/lift history ( | pitzDaily recirculation ( |
|
|
Why another OpenFOAM MCP?
Existing servers wrap a handful of hard-coded scenarios. This one exposes general, composable primitives and lets the model do the engineering, with guard-rails that make long-running CFD practical from a chat:
openfoam-mcp | |
Works with any case | Clone any of the ~260 tutorials (or your own case) and edit it; no canned geometries |
OpenFOAM versions | openfoam.org 11+ ( |
Long runs | Detached jobs that survive client/server restarts; live progress, ETA, graceful write-and-stop |
Convergence intelligence | Incremental log parsing: residual trends, Courant, continuity, converged / diverging / failed / crashed verdicts, plateau detection, ETA, extracted FOAM FATAL ERROR blocks |
Sees the flow | Headless ParaView renders (slices, patches, streamlines, iso-surfaces such as Q-criterion, mesh) framed on any patch, plus residual/force plots — returned as images |
Safe edits | Dictionary edits go through OpenFOAM's own |
Parallel |
|
Results | Field statistics straight from ASCII or binary field files; postProcessing tables with text and vector columns |
Discoverability |
|
Every client |
|
Token-frugal | Compact JSON, elided field lists, sampled tables, grouped outputs; 25 tools |
Related MCP server: Kratos MCP Server
Install
Requirements: Python ≥ 3.10 and OpenFOAM reachable in one of the ways below. Optional: ParaView with Python for images, OpenMPI for parallel runs.
uv tool install git+https://github.com/0xFFD/openfoam-mcp # or: pipx install git+https://…
openfoam-mcp doctor # checks OpenFOAM, MPI, ParaView, workspace
openfoam-mcp install # registers the server with every MCP client it findsPlatforms
The server runs on the machine (or VM) that can launch OpenFOAM, and finds it automatically:
Platform | How OpenFOAM is reached | Notes |
Linux (native) |
| ParaView: |
macOS (native) | OpenFOAM.app ( | ParaView: the official |
macOS / any OS (container) |
| The way to run openfoam.org releases on a Mac. Tutorials are mirrored to |
Windows | Install inside WSL2; | Native Windows OpenFOAM builds are not supported |
# macOS with OpenFOAM.app
brew install gerlero/openfoam/openfoam && openfoam-mcp doctor
# any OS with Docker (example images: opencfd/openfoam-default for openfoam.com,
# or an openfoam.org image); bake the choice into every client config:
openfoam-mcp install --container-image opencfd/openfoam-defaultinstall detects clients on the Linux side and, when run in WSL, on the Windows side (launching the server via wsl.exe -d <distro> --exec …). Restrict it with --client codex,claude-code, preview with --dry-run, undo with openfoam-mcp uninstall. Every modified config gets a one-time .bak-openfoam-mcp backup.
openfoam-mcp config prints a snippet for each client (--windows for Windows clients using WSL). Examples:
Codex CLI — ~/.codex/config.toml
[mcp_servers.openfoam]
command = "wsl.exe" # or the Linux python path when Codex runs inside Linux
args = ["-d", "Ubuntu-22.04", "--exec", "/home/me/.local/share/uv/tools/openfoam-mcp/bin/python", "-m", "openfoam_mcp", "serve"]
startup_timeout_sec = 30
tool_timeout_sec = 600Claude Code
claude mcp add --scope user openfoam -- wsl.exe -d Ubuntu-22.04 --exec /home/me/.local/share/uv/tools/openfoam-mcp/bin/python -m openfoam_mcp serveCursor / Gemini CLI / Windsurf / Claude Desktop — mcpServers JSON
{ "mcpServers": { "openfoam": { "command": "wsl.exe", "args": ["-d", "Ubuntu-22.04", "--exec", "/home/me/…/python", "-m", "openfoam_mcp", "serve"] } } }Tools
Area | Tools |
Environment |
|
Cases |
|
Files & dictionaries |
|
Running |
|
Mesh |
|
Results |
|
Typical loop the agent follows: list_tutorials → create_case → case_summary → set_dict → run blockMesh/snappyHexMesh → check_mesh → run solver → solver_progress → render / post_process.
render modes
auto (whole 2-D domain, mid-plane slice in 3-D) · slice · surface · patches (e.g. pressure on a body) · mesh · contour (iso-surface, e.g. Q-criterion after post_process("Q")) · streamlines — with camera presets, focus on a patch or group, zoom, component, colour range and colormap.
Tested with
OpenFOAM 14 (openfoam.org) on Ubuntu 22.04 / WSL2: pitzDaily (2-D, serial and parallel) and motorBike (3-D,
snappyHexMeshand solver on 6 MPI ranks, 355 k cells).Clients: Claude Code (agent session through
wsl.exe) and Codex CLI (configuration and health check).Python 3.10 and 3.13.
Not yet verified on real hardware: macOS (OpenFOAM.app), Docker/Podman, openfoam.com (ESI) builds — reports welcome.
Configuration
Flag | Env var | Default |
|
|
|
|
| — extra directories tools may access |
|
| auto-detected ( |
|
| auto-detected ( |
|
| — run OpenFOAM in a container started from IMAGE |
|
| — use an existing container (must mount the workspace at the same path) |
|
|
|
|
| auto-detected ( |
|
| scripts allowed |
|
| 45 — how long a tool waits before returning a job id |
| stdio |
Safety model
File access is confined to the workspace (plus
--rootdirectories);..and absolute-path escapes are rejected.runonly executes OpenFOAM applications from the installation's bin directories, or case-localAll*scripts. There is no general shell tool.This is not a sandbox. Case scripts and OpenFOAM's
#codeStream/coded boundary conditions can run arbitrary code by design. Use--no-scriptsand review cases from untrusted sources; run the server under a dedicated user or container if you need isolation.Tools carry MCP annotations (
readOnlyHint,destructiveHint) so clients can auto-approve reads and confirm deletions.
Development
git clone https://github.com/0xFFD/openfoam-mcp && cd openfoam-mcp
uv sync
uv run pytest # unit tests run anywhere; integration tests run when OpenFOAM is found
uv run ruff check src tests
uv run openfoam-mcp serve --workspace /tmp/ws # or point an MCP inspector at itCI runs the unit tests on Linux and macOS. Container mode is tested with a stand-in docker executable (see tests/test_backends.py) that runs the real pitzDaily case through the container code path; reports from real Docker/Podman and OpenFOAM.app setups are very welcome.
Project layout: server.py (tool definitions) · foam.py (installation discovery; native, launcher and container runners; app allow-list) · jobs.py (persistent background jobs) · logs.py (incremental solver-log parser and convergence verdicts) · mesh.py · fields.py · postproc.py · dictparse.py (fast read-only dictionary parser) · render.py + _pv_render.py (ParaView) · install.py (client configuration).
License
MIT. OpenFOAM is a registered trademark of OpenCFD Ltd; this project is not affiliated with or endorsed by OpenCFD or the OpenFOAM Foundation.
Available Tools
25 toolscase_summaryARead-only
One-call overview of a case: solver, run control, mesh & patches, initial/boundary conditions, physics, numerics, results.
| Name | Required | Description | Default |
|---|---|---|---|
| case | Yes | Case name in the workspace (e.g. 'pitzDaily') or an absolute path inside an allowed root. | |
| include_fields | No | Include initial/boundary conditions of every field. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds the meaningful behavioral signal that this is a broad aggregation returning many sections at once (implying a large payload), but says nothing about cost, output size, or failure behavior for a non-existent case. Adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with a colon-delimited list of contents; zero filler. It is compact and scannable, though the list style is terse enough that a slightly expanded clause on scope would still earn its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description effectively enumerates the return contents, and the annotations carry the safety profile, so an agent knows what calling this yields and that it is a safe read. Only minor gaps remain: behavior for missing/invalid cases and any sense of output volume.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (case and include_fields) are already fully documented in the schema, including the default and what include_fields adds. The description's mention of 'initial/boundary conditions' loosely corresponds to include_fields but adds no format, path, or scoping detail beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('overview') and resource ('a case'), and enumerates the concrete sections returned (solver, run control, mesh & patches, conditions, physics, numerics, results). An agent can tell this apart from get_dict/read_file because it is an aggregated, single-call snapshot. The only weakness is that the distinction from siblings is implied by 'one-call' rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'One-call overview' implies the usage context: use this instead of making many separate section reads. However, there is no explicit when-to-use/when-not, no named alternative (e.g. get_dict for a single entry), and no statement of prerequisites such as the case existing in the workspace. Usage must be inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_meshBRead-only
Run checkMesh and return structured mesh statistics, failed checks and practical advice.
| Name | Required | Description | Default |
|---|---|---|---|
| run | No | Run checkMesh now (false: parse the existing log.checkMesh). | |
| args | No | Extra checkMesh options, e.g. ['-allGeometry', '-allTopology'] or ['-latestTime']. | |
| case | Yes | Case name in the workspace (e.g. 'pitzDaily') or an absolute path inside an allowed root. | |
| wait_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds that it returns stats, failures and advice, but says nothing about the execution side effects (a run writes/overwrites log.checkMesh), the wait_seconds timeout behavior, or what happens when no log exists and run=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, stating action first and return content second. It is efficient, though arguably terse given the tool's four parameters and absent output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must carry the return contract; it does name the three categories of returned data, which is the minimum. It omits anything about timeout/failure behavior, the log-parsing fallback path, or argument pass-through semantics, leaving meaningful gaps for a 4-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 75%, and the schema itself explains `run` and `args` well; only `wait_seconds` lacks a description. The description contributes no parameter-level meaning beyond the schema, so the baseline 3 for high coverage with a thin description is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (run checkMesh, an OpenFOAM mesh quality utility) and names the return payload: structured mesh statistics, failed checks, practical advice. It does not, however, distinguish itself from sibling tools like `run` or `read_log`, which an agent could plausibly pick for a similar job.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance at all: nothing says whether to reach for check_mesh versus the generic `run` tool, nor when to invoke it (e.g., before running a solver). The only usage hint lives in the schema (`run=false` parses an existing log), which is structured data rather than description guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clean_caseBDestructive
Remove results: time directories after the first, postProcessing, logs, processor dirs and optionally the mesh.
| Name | Required | Description | Default |
|---|---|---|---|
| case | Yes | Case name in the workspace (e.g. 'pitzDaily') or an absolute path inside an allowed root. | |
| keep_logs | No | ||
| keep_mesh | No | Keep constant/polyMesh. Mesh removal is refused when the case cannot regenerate it. | |
| keep_processor | No | Keep processor* (decomposed) directories. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the risk profile is known. The description usefully itemizes the destructive surface and notes the mesh is only removed optionally, but omits irreversibility, confirmation behavior, and the effect on the retained first time directory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, listing targets in descending importance. Grammar is slightly compressed ('Remove results: time directories...') but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with no output schema, the description covers what is removed but says nothing about error handling, whether the case directory itself survives, or what state the case is left in. The mesh-regeneration refusal is covered only in the schema, not the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, with keep_mesh and keep_processor documented in the schema but keep_logs left bare. The description roughly mirrors the defaults (logs and processor removed, mesh kept by default) but adds no syntax or edge-case meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb ('Remove') and enumerates exactly which artifacts are targeted (time directories after the first, postProcessing, logs, processor dirs, optionally the mesh). It is clearly distinct from delete_case's whole-case removal, though the sibling relationship is never stated explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a cleanup use case but never says when to use this over delete_case or in what workflow (e.g., before a re-run). No prerequisites, no exclusions, no mention that the first time directory is retained as an intentional rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_caseA
Create a case by cloning a tutorial or copying an existing case (results are left behind unless include_results).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | New case name (directory) in the workspace. | |
| source | Yes | Tutorial path from list_tutorials (e.g. 'incompressibleFluid/pitzDaily') or an existing case name/path to copy. | |
| overwrite | No | Replace an existing case with this name. | |
| include_results | No | When copying a case, also copy time directories, logs and postProcessing. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply a safety profile (readOnlyHint=false, destructiveHint=false, openWorldHint=false). The description adds a genuinely behavioral detail beyond that: results are left behind unless include_results is set, which warns the agent about silent data omission. It does not, however, mention the overwrite-replaces-existing-case hazard, which is the riskiest behavior here.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence: purpose first, then the qualifying caveat. No filler, and the exception clause is attached where it matters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a creation tool with no output schema, the description covers the source options and the results default, which is most of what an agent needs. It omits the overwrite/replace behavior and any note about what the call returns or where the new case lands, so it is slightly short of fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are documented in the schema itself (including overwrite and include_results semantics). The description's parenthetical about results partially duplicates include_results' schema wording and adds no syntax or format detail, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (create) and resource (case) plus the two supported sources: cloning a tutorial or copying an existing case. An agent can immediately separate this from list_tutorials, list_cases, clean_case and delete_case without inspecting any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The two creation paths imply usage, but the description never states when this tool is the right choice versus alternatives, nor any prerequisites (e.g. that the source tutorial must first be found via list_tutorials, or that the target workspace must exist). Usage is implied rather than guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_caseADestructive
Permanently delete a case directory from the workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| case | Yes | Case name in the workspace (e.g. 'pitzDaily') or an absolute path inside an allowed root. | |
| confirm | No | Must be true to actually delete. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false and openWorldHint=false, so the safety profile is covered. The description still adds value by disclosing that the deletion is permanent (irreversible, no trash/undo) and that the target is a whole case directory in the workspace, which destructiveHint alone does not convey. It stops short of stating whether anything else (jobs, logs) is touched or what is returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that leads with the permanence qualifier and the resource, with no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter, destructive operation with full schema coverage and annotations carrying the safety hints, the critical missing piece an agent needs is irreversibility, which the word 'Permanently' supplies. Nothing further is mandatory, though it could note that confirm must be set and whether the deletion can be undone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: the schema explains 'case' (name or absolute path inside an allowed root) and 'confirm' (must be true to actually delete). The description adds no syntax, format, or path-restriction detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Permanently delete a case directory') with the target scope ('from the workspace'), so the agent knows exactly what is affected. It does not distinguish itself from the sibling clean_case, which plausibly removes case artifacts while keeping the directory, leaving the boundary between the two implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no statement of prerequisites, and no mention of the obvious alternative clean_case for less aggressive removal. The adverb 'Permanently' hints at the condition under which this tool is appropriate, but the agent must infer it rather than being told.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
field_statsBRead-only
Min/max/mean of a field (internal and per patch) at a time, read directly from ASCII field files.
| Name | Required | Description | Default |
|---|---|---|---|
| case | Yes | Case name in the workspace (e.g. 'pitzDaily') or an absolute path inside an allowed root. | |
| time | No | 'latest', 'first' or a time directory name. | latest |
| field | Yes | Field name, e.g. 'U', 'p', 'k', 'alpha.water'. | |
| region | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds a genuine behavioral detail: it reads directly from ASCII field files, hinting that binary-format cases may not work — useful context the schema does not convey, though return format and failure modes remain unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the operation front-loaded before the qualifiers. Efficient, though the parenthetical about internal/per-patch breaks the flow slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only metric tool with no output schema, saying it returns min/max/mean covers the essentials, and the ASCII-file note is a nice addition. It is adequate but leaves gaps around output shape, units, and what happens for non-ASCII or missing fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75% and the case/time/field params are reasonably documented in the schema itself. The description adds 'internal and per patch' (mapping loosely to the undocumented region parameter) and the time-scoping constraint, but says nothing about case paths or region semantics beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (computes min/max/mean) and resource (a field's values, internal and per patch) at a given time. Clear enough to separate from siblings like case_summary or post_process, though it never names those alternatives explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'at a time' implies the tool is for point-in-time sampling rather than time-series analysis, but there is no explicit when-to-use guidance and no mention of alternative tools (post_process, case_summary, solver_progress) an agent should consider instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
foam_infoARead-only
OpenFOAM installation, server configuration and resources (version, flavour, paths, cores, ParaView).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already establish that this is a safe read-only, non-open-world operation. The description adds valuable context about what the tool surfaces: version, flavour, paths, core counts, and ParaView details. It does not explain return formatting, but for a simple info tool this is a meaningful behavioral supplement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with key terms front-loaded, no filler. It is not a complete sentence, but every word carries information and nothing needs to be removed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no parameters, the description must tell the agent what information it will receive. It does so by enumerating version, flavour, paths, cores, and ParaView, which is enough for an agent to know when this tool is relevant and what to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter-level semantics to convey. The baseline of 4 applies because the description does not need to compensate for any schema gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource and scope: OpenFOAM installation, server configuration, and resources, with concrete examples (version, flavour, paths, cores, ParaView). An agent can distinguish it from siblings like foam_reference, which likely covers reference material rather than installation/config details. It lacks an explicit verb, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is provided and no alternative tools are mentioned. The agent can infer that this tool is for querying installation/config info, but it receives no help deciding between foam_info and foam_reference or other list/introspection tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
foam_referenceARead-only
Look up what OpenFOAM supports: valid boundary-condition types, models, function objects, solvers, app options.
| Name | Required | Description | Default |
|---|---|---|---|
| topic | Yes | One of: solvers, apps, scalarBCs, vectorBCs, functionObjects, functions (configured post-processing templates), fvModels, fvConstraints, tables, 'table:<name>' (e.g. table:RAScompressibleMomentumTransportModel), 'search:<name>' (which tables contain a type), 'help:<app>' (an application's options). | |
| filter | No | Case-insensitive substring to filter output lines. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the agent knows this is a safe, static lookup. The description adds the domain scope of the lookup, which is genuinely useful, but says nothing about e.g. result shape or whether large topic lists are truncated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that leads with the verb and puts the informative domain list after the colon. No filler or repetition, though the enumerated list is somewhat long for the value it adds.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only reference lookup with no output schema and fully covered parameters, the description gives an adequate sense of what the tool returns. Nothing critical is missing, though it could note the lookup is against a packaged/static reference rather than the current case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the top-level topic enum and filter are fully documented in the schema itself. The description echoes several topic domains (models, function objects, solvers, app options) but adds no syntax or format detail beyond what the schema already provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Look up') and resource ('what OpenFOAM supports') and enumerates the domains covered: boundary conditions, models, function objects, solvers, app options. This clearly separates it from case/mutation siblings like create_case or run, though it does not explicitly distinguish itself from the potentially overlapping foam_info.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the enumeration of lookup domains – an agent can infer it should call this when it needs valid type names or app options rather than inspecting a case. However, there is no explicit when-to-use statement and no named alternative or exclusion to route against.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_dictBRead-only
Read a dictionary or one entry through OpenFOAM's foamDictionary (authoritative parsing).
| Name | Required | Description | Default |
|---|---|---|---|
| case | Yes | Case name in the workspace (e.g. 'pitzDaily') or an absolute path inside an allowed root. | |
| file | Yes | Dictionary path relative to the case, e.g. 'system/controlDict', '0/U'. | |
| entry | No | Entry path with '/' separators, e.g. 'boundaryField/inlet' or 'SIMPLE/residualControl'. Omit for the whole file. | |
| expand | No | Expand #include, $macros and #calc first. | |
| keywords | No | Only list the keywords at that level. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and openWorldHint=false, so safety is already covered. The description adds that parsing is authoritative via foamDictionary, which is real value beyond annotations, but it says nothing about failure modes (missing files, malformed dicts) or the shape of returned parsed values.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tight sentence with no waste, front-loading the action. The parenthetical is useful but slightly elliptical ('authoritative parsing' assumes the reader knows foamDictionary's role).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A read-only tool with full schema coverage and annotations needs little more, but the 'authoritative' claim implies a contrast with the sibling read_file that is left unexplored, and there is no output schema to explain the parsed return shape. Adequate but leaving clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all five parameters are already documented in the schema, including entry, expand, and keywords. The description adds no parameter-level meaning beyond the schema baseline, which is adequate but unremarkable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Read) and resource (dictionary or one entry) via foamDictionary. The 'authoritative parsing' qualifier hints at why this tool exists versus sibling read_file, but does not name that sibling explicitly, so sibling differentiation is implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no conditions for choosing this over read_file or foam_reference, and no mention of what the 'authoritative' qualifier excludes. The agent must infer from the name alone that this is the preferred dictionary reader.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_fileB
Copy geometry (STL/OBJ), meshes or other files into a case, e.g. the surface for snappyHexMesh.
| Name | Required | Description | Default |
|---|---|---|---|
| case | Yes | Case name in the workspace (e.g. 'pitzDaily') or an absolute path inside an allowed root. | |
| dest | Yes | Destination inside the case: a directory ending in '/' (e.g. 'constant/geometry/') or a file path. | |
| source | Yes | File or directory to copy: '$FOAM_TUTORIALS/resources/geometry/motorBike.obj.gz', or a path inside the workspace/allowed roots (Windows paths like C:\\models\\car.stl are accepted under WSL). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=false, so the mutation/safety profile is covered. The description adds the copy-into-case framing but says nothing about overwrite behavior when dest already exists or about permissions/errors.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence naming the action and resource, with the example placed at the end where it does not obstruct the core meaning. No padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema the description needn't explain returns, and annotations cover safety, but for a file-copy tool the overwrite/conflict behavior and any size or format restrictions are left unstated, leaving a modest gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so case/source/dest are already fully documented in the schema. The description only adds example file types (STL/OBJ) and does not explain parameter interactions beyond what the schema provides; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Copy') and resource ('geometry (STL/OBJ), meshes or other files') and the destination context ('into a case'). An agent can tell this apart from read_file/list_files, though the boundary with write_file is not made explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'e.g. the surface for snappyHexMesh' implies a typical use case, but there is no explicit when-to-use, no when-not-to-use, and no reference to the sibling write_file as an alternative for in-case content creation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
job_statusARead-only
Status of a job (progress, residuals, errors, log tail) or a list of running and recent jobs.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | No | Job to inspect. Omit to list running and recent jobs. | |
| wait_seconds | No | Wait up to this long for the job to finish before reporting. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds that results include progress, residuals, errors, and a log tail, which is useful content context, but it says nothing about polling behavior, cost, or how wait_seconds affects the result beyond the schema note.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with a tight parenthetical listing the payload; no filler, and the dual-mode behavior is front-loaded so an agent reaches the decision point immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With readOnly annotations and a fully documented two-parameter schema, the main remaining gap is return formatting, but the description itself enumerates what comes back (progress, residuals, errors, log tail), which largely covers that for a tool with no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, including the omit-to-list semantics of job_id and the wait semantics of wait_seconds, so the schema already does the heavy lifting. The description's mention of the list mode aligns with job_id being optional but adds no syntax or format detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (job) and enumerates the returned content (progress, residuals, errors, log tail), and also covers the list mode. That distinguishes it reasonably well from siblings like read_log or solver_progress, though the primary verb is implicit ("Status of...").
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The dual-mode behavior is stated (inspect one job vs. list running/recent jobs), and the schema param says to omit job_id to list. However, no alternative tools are named and there is no guidance on when to prefer this over read_log or solver_progress.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_casesARead-only
List cases in the workspace with solver, mesh/result state and running jobs.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds real value beyond that by disclosing what the listing contains (solver, mesh/result state, running jobs), which is the only signal about return content since no output schema exists. It still omits pagination or ordering behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler, front-loaded with the action and resource before the per-case detail. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only listing tool with no output schema, the description supplies the essential missing piece: what each listed case includes. Only ordering, pagination, and result-size expectations are left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to disambiguate, and it correctly does not invent parameters that do not exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('List') and resource ('cases in the workspace') and even names the fields surfaced per case (solver, mesh/result state, running jobs). It is clearly distinguishable from siblings like list_files or list_tutorials, though it does not explicitly contrast with case_summary or list_postprocessing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer this is the enumeration entry point for cases, but there is no statement of when to prefer it over case_summary, list_files, or job_status. No exclusions, prerequisites, or alternatives are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_filesARead-only
List files in a case (sizes in bytes; long runs of time directories are collapsed).
| Name | Required | Description | Default |
|---|---|---|---|
| case | Yes | Case name in the workspace (e.g. 'pitzDaily') or an absolute path inside an allowed root. | |
| depth | No | ||
| subdir | No | Directory inside the case to list, e.g. 'system' or 'constant/triSurface'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds genuine behavioral detail beyond that: sizes are reported in bytes and long runs of time directories are collapsed, which tells the agent what the listing will actually look like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with the core action front-loaded and two useful output facts in a compact parenthetical. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should carry more of the return-format burden; it covers units and directory collapsing but says nothing about ordering, recursion semantics via depth, or whether subdirs are included. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%: 'case' and 'subdir' are documented in the schema, but 'depth' (default 2) has no description anywhere, and the description does not compensate. Since the schema carries most of the load, baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List files in a case') with a scoping qualifier that separates it from file-content tools like read_file and from list_cases/list_tutorials. It does not explicitly name a sibling alternative, but the resource is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose this over list_cases, read_file, or get_dict, and no mention of prerequisites such as the case needing to exist or be accessible. Usage is only implied by the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_postprocessingBRead-only
List function-object output files under postProcessing/ with their column names.
| Name | Required | Description | Default |
|---|---|---|---|
| case | Yes | Case name in the workspace (e.g. 'pitzDaily') or an absolute path inside an allowed root. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds that results are function-object output files and that column names are surfaced, which is useful, but it says nothing about ordering, empty-result behavior, or return shape beyond column names.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the resource scope and the column-name detail both earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only listing tool with a fully documented single parameter, this is nearly complete. There is no output schema, so a brief note on what the returned entries look like (paths, empty case) would have closed the remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the single 'case' parameter (name or absolute path in an allowed root) is fully documented in the schema. The description adds no further meaning about the parameter, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (function-object output files under postProcessing/) plus an extra detail (their column names). It clearly distinguishes its scope from a content-reading sibling like read_postprocessing, but it never names that sibling or otherwise contrasts itself explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use, when-not-to-use, or alternative-tool guidance. An agent must infer that this is a discovery step preceding read_postprocessing, and nothing addresses the overlapping list_files tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_tutorialsARead-only
Find OpenFOAM tutorial cases to start from (and working examples of any keyword or feature).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | No | Words that must all appear in the tutorial path or solver, e.g. 'pitzDaily' or 'VoF dam'. | |
| contains | No | Regex searched inside the tutorial files, e.g. 'kOmegaSST' or 'forceCoeffs' - find working examples of a feature. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety and scope are covered. The description adds that results are tutorial starting points and feature examples, but says nothing about result format, result count, or the effect of the limit parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with zero filler, front-loaded with the core action and resource, with the secondary use case in a compact parenthetical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, zero-required-param discovery tool with annotations covering safety and no output schema, this is adequate but thin: return shape, result count, and the role of limit are never addressed, and no alternative tools are mentioned.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%; query and contains are already documented in the schema (query: words in path/solver; contains: regex inside files). The description's parenthetical loosely maps to the contains mode but adds little beyond the schema, and the undocumented limit parameter is left entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Find OpenFOAM tutorial cases', plus a secondary scope ('working examples of any keyword or feature'). The resource ('tutorial cases') is inherently distinct from sibling tools like list_cases or foam_reference, though the description never explicitly contrasts with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implied usage is clear from 'to start from' (bootstrap a new case) and 'working examples of any keyword or feature' (look up a feature in real files). However, it names no alternative tools (e.g., list_cases, foam_reference) and gives no when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
post_processB
Run a post-processing function object on saved results (writes fields and/or postProcessing/ data).
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | ||
| case | Yes | Case name in the workspace (e.g. 'pitzDaily') or an absolute path inside an allowed root. | |
| func | Yes | Function, e.g. 'mag(U)', 'vorticity', 'yPlus', 'wallShearStress', "patchAverage(p, patch=outlet)", 'Q'. List options with foam_reference('functions'). | |
| time | No | 'latest', 'all', or a time value / range like '100:200'. | latest |
| with_solver | No | Construct the solver's physical models (needed for yPlus, wallShearStress, forces, turbulenceFields). | |
| wait_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and destructiveHint=false, so the agent already knows this mutates but is non-destructive. The description usefully clarifies what the mutation produces (fields and/or postProcessing/ data), adding value beyond the annotations. It still omits overwrite behavior on existing fields, the role of wait_seconds, and any failure/long-running-job semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the verb and effect lead. It is efficiently sized, though slightly terse given the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter mutation tool with no output schema, the description covers the core action but leaves gaps: async/wait behavior, what 'args' does, and how output is subsequently retrieved (read_postprocessing). Annotations offload the safety profile, but more operational context would be needed for a higher score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67%; case, func, time, and with_solver are well documented in the schema, while args and wait_seconds are undocumented anywhere. The description adds no parameter-level detail, so it does not compensate for the two gaps. Baseline 3 is appropriate given the schema carries most of the burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Run') and resource ('post-processing function object on saved results'), plus the observable effect ('writes fields and/or postProcessing/ data'). It does not explicitly contrast itself with siblings like run or read_postprocessing, but the resource is distinct enough to differentiate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'on saved results' implies this runs after a solve has produced data, which gives implied usage. However, there is no explicit when-to-use/when-not guidance and no mention of alternatives (e.g., read_postprocessing for reading existing output, or foam_reference('functions') for listing options only hinted at inside the schema).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_fileBRead-only
Read a text file from a case (dictionaries, logs, scripts). Large value lists are elided by default.
| Name | Required | Description | Default |
|---|---|---|---|
| case | Yes | Case name in the workspace (e.g. 'pitzDaily') or an absolute path inside an allowed root. | |
| path | Yes | File path relative to the case, e.g. 'system/fvSolution' or '0/U'. | |
| max_lines | No | ||
| start_line | No | ||
| elide_lists | No | Replace large nonuniform value lists with a placeholder. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered structurally. The description adds one genuinely useful behavioral note ('Large value lists are elided by default'), but says nothing about truncation limits or what happens when a file exceeds max_lines, which matters for an operation whose output can be silently cut.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler, the core action front-loaded and the elision caveat immediately after. Nothing is repeated from the title or annotations, and no sentence is dispensable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool whose annotations cover the safety profile and which has no output schema, the description conveys the general return behavior (elided lists) but omits truncation semantics for max_lines/start_line. An agent can call it, but may be surprised by a 400-line cutoff with no warning in the prose.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 60% (case, path, elide_lists documented; max_lines and start_line bare). The description only restates the elision default, which the schema already documents, and adds no meaning for the two undocumented line-window parameters. Baseline 3 is appropriate when the schema carries most of the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read a text file from a case') and lists representative file types (dictionaries, logs, scripts), which orients the agent quickly. It does not, however, distinguish itself from overlapping siblings like get_dict and read_log, which appear to read the same kinds of files with more structure.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance. With siblings such as get_dict, read_log, read_postprocessing and list_files in the same namespace, the agent is left to infer whether read_file is the raw-text fallback or a general reader for those specific file classes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_logBRead-only
Read the tail of a log, or grep it (e.g. 'FOAM Warning|bounding|Courant').
| Name | Required | Description | Default |
|---|---|---|---|
| log | No | Log file, e.g. 'log.foamRun' or just 'foamRun'. Default: most recently modified log. | |
| case | Yes | Case name in the workspace (e.g. 'pitzDaily') or an absolute path inside an allowed root. | |
| grep | No | Regex; return matching lines (with line numbers) instead of the tail. | |
| tail | No | Number of trailing lines to return (ignored with grep). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds mode-specific behavior (grep replaces the tail, example pattern) which is useful, but says nothing about return format, truncation, or what happens when the log is missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence, front-loaded with the primary action, with the grep example appended economically. Very little waste, though the dual-mode phrasing is slightly compressed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only log inspection tool with a fully documented schema and no output schema, the description covers the essentials. It omits return shape and how it interacts with the broader case/file tools, leaving some gaps an agent must infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% – every parameter (log, case, grep, tail) is already well documented in the schema, including defaults. The description's regex example adds only marginal value. Baseline 3 is appropriate since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (read) and resource (log) and distinguishes the two operating modes: tail vs grep, with a concrete regex example. It does not, however, differentiate itself from siblings such as read_file or read_postprocessing, which an agent could easily confuse with it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the two internal modes but never says when to reach for read_log versus read_file, list_files, or read_postprocessing. No prerequisites, no exclusions, no routing guidance for the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_postprocessingARead-only
Read a postProcessing table (forces, coefficients, probes, sampled lines...) with statistics and optional plot.
| Name | Required | Description | Default |
|---|---|---|---|
| case | Yes | Case name in the workspace (e.g. 'pitzDaily') or an absolute path inside an allowed root. | |
| path | Yes | File path from list_postprocessing, e.g. 'postProcessing/forceCoeffs/0/forceCoeffs.dat'. | |
| plot | No | Also return a line plot of the selected columns. | |
| columns | No | Columns to return/plot (default: all). | |
| max_rows | No | Rows to include (evenly sampled). 0 returns statistics only. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds the return profile (statistics and an optional plot), which is useful context, but does not disclose sampling details, sizes, or error behavior beyond what the schema states.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the core action front-loaded and the data families and return extras trailing. No wasted clauses, though the parenthetical list is slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A read-only tool whose safety is carried by annotations and whose parameters are fully documented by the schema. The description's note about statistics and optional plotting gives the agent enough to know what comes back, despite there being no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are already documented in the schema itself (including max_rows sampling and plot). The description's mention of 'optional plot' adds nothing beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (read) and resource (postProcessing table) and enumerates the data family it covers (forces, coefficients, probes, sampled lines). It is clearly distinguishable from the sibling list_postprocessing, which only enumerates paths.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the 'path' schema mentions list_postprocessing as the source of the path, hinting at a list-then-read workflow. The description itself gives no explicit when-to-use or when-not-to-use guidance, nor does it distinguish this from sibling read_file.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
renderBRead-only
Render a picture of the flow or mesh with ParaView (headless) and return it as an image plus metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| case | Yes | Case name in the workspace (e.g. 'pitzDaily') or an absolute path inside an allowed root. | |
| mode | No | auto: whole domain for 2-D cases, mid-plane slice for 3-D. patches: colour boundary patches (e.g. pressure on a body). contour: iso-surface of contour_field at iso_value. streamlines: of vector `field`. | auto |
| time | No | 'latest', 'first' or a time value. | latest |
| view | No | Camera: 'auto', 'iso', '+x', '-x', '+y', '-y', '+z', '-z' (camera on that side, looking back). | auto |
| zoom | No | Zoom factor >1 to magnify (centre of domain). | |
| edges | No | Draw mesh edges. | |
| field | No | Field to colour by (e.g. 'U', 'p'). Omit with mode='mesh'. | |
| focus | No | Patch or group to frame (e.g. 'motorBikeGroup', 'wall'); slices pass through it by default. | |
| range | No | [min, max] colour range. | |
| width | No | ||
| height | No | ||
| patches | No | Patch/group names for mode='patches'. | |
| colormap | No | coolwarm, viridis, jet, turbo, inferno, plasma, rainbow, blue-red, grayscale. | coolwarm |
| component | No | Vector component (default Magnitude). | |
| iso_value | No | ||
| background | No | light | |
| seed_point1 | No | ||
| seed_point2 | No | ||
| slice_normal | No | Slice normal, e.g. 'y' or [0,1,0]. | |
| slice_origin | No | Slice origin (default: domain centre). | |
| contour_field | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description usefully adds that ParaView runs headless and that the result is an image plus metadata, but says nothing about execution cost, latency, or prerequisite state (e.g. requiring completed simulation data).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence, front-loaded with the action and ending with the observable result. No filler or restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 21-parameter rendering tool with no output schema, the description is adequate but thin. It covers what the tool produces but not the prerequisites, the image format/location, or how the mode-dependent parameters combine, leaving the agent to reconstruct the workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 21 parameters and only 67% schema description coverage, the description carries a compensation burden and contributes zero parameter meaning. It does not clarify the mode/field/patches interplay or any of the undocumented parameters (width, height, iso_value, seed points, contour_field).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (render) and resource (flow/mesh picture via headless ParaView) plus the return type (image plus metadata). It is clearly distinguishable from siblings like post_process or field_stats, though it never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance at all: nothing says whether the case must already be solved, how this relates to post_process or field_stats, or when rendering is or isn't appropriate. The agent must infer the workflow context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
runA
Run an OpenFOAM application or case script as a background job; waits briefly and reports the outcome or progress.
| Name | Required | Description | Default |
|---|---|---|---|
| np | No | MPI ranks. >1 runs with mpirun and -parallel, decomposing the case first if needed. | |
| app | Yes | OpenFOAM application (blockMesh, snappyHexMesh, foamRun, simpleFoam, decomposePar, reconstructPar, setFields, ...) or a case script such as Allrun. | |
| log | No | Log file name (default log.<app>; the previous one is kept as .1). | |
| args | No | Command-line arguments, e.g. ['-overwrite'] or ['-dict', 'system/blockMeshDict.fine']. | |
| case | Yes | Case name in the workspace (e.g. 'pitzDaily') or an absolute path inside an allowed root. | |
| wait_seconds | No | How long to wait for completion before returning a job id (default ~45 s). 0 returns immediately. | |
| decompose_method | No | Decomposition method (scotch, hierarchical, simple...) if the case must be decomposed for np>1. Default: the case's decomposeParDict, else scotch. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds real behavioral context beyond that: execution is asynchronous, the call blocks only briefly, and it returns either an outcome or a job id with progress. It does not mention that outputs/logs are written to the case directory, which is a meaningful gap for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that leads with the action and resource before the execution semantics. No filler, no restatement of the tool name, and every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter job-launching tool with no output schema, the description covers the call mechanics but omits how to follow up on a returned job (job_status, read_log, solver_progress) and what side effects occur in the case directory. Adequate but with clear gaps given the tool's complexity and the rich sibling set.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (np, app, log, args, case, wait_seconds, decompose_method) is already documented in the schema, including non-obvious details like mpirun/decomposition behavior and log rotation. The description adds no parameter-level meaning beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run an OpenFOAM application or case script') plus the execution mode ('as a background job'), which is enough to separate it from job_status or stop_job. It stops short of explicitly naming the sibling tools an agent should pair it with, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'as a background job; waits briefly and reports the outcome or progress' implies when this tool fits (launching long-running solver or utility work) but gives no explicit when-not guidance or alternatives. With siblings like job_status, solver_progress, and read_log in the set, an agent must infer the follow-up path itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_dictA
Edit dictionary entries safely via foamDictionary (set/add/remove), then read back the new values.
| Name | Required | Description | Default |
|---|---|---|---|
| set | No | Entries to set or add: {'endTime': '2000', 'boundaryField/inlet/value': 'uniform (5 0 0)', 'RAS/model': 'kOmegaSST', 'boundaryField/outlet': '{ type zeroGradient; }'}. Values use OpenFOAM syntax without the trailing ';'. | |
| case | Yes | Case name in the workspace (e.g. 'pitzDaily') or an absolute path inside an allowed root. | |
| file | Yes | Dictionary path relative to the case, e.g. 'system/controlDict', '0/U', 'constant/momentumTransport'. | |
| remove | No | Entry paths to remove. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=false, lowering the disclosure bar. The description usefully adds that the edit goes through foamDictionary and that new values are read back, but 'safely' is undefined and it doesn't reconcile the explicit 'remove' operation with the non-destructive annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that conveys the mechanism, the operations and the read-back behavior with no filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter mutation tool with fully documented params and no output schema, the description covers the essentials and even notes that values are read back after the edit. Minor gaps remain around prerequisites/permissions, but nothing required to invoke it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with rich per-parameter docs including OpenFOAM syntax examples for 'set', path forms for 'case'/'file' and 'remove'. The description adds nothing beyond the schema, so the baseline 3 for high-coverage schemas is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Edit dictionary entries'), names the underlying mechanism (foamDictionary) and enumerates the operations (set/add/remove), which cleanly separates it from get_dict, read_file and write_file. It stops short of explicitly naming a sibling as the wrong choice, so it lands at 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the usage context (editing OpenFOAM dictionary entries) but gives no explicit when-to-use or when-not-to-use guidance. With siblings like write_file, get_dict and read_file available, an agent would benefit from knowing that this is the targeted dict-editing path, and that guidance is absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
solver_progressBRead-only
Convergence report from a solver log: state (converged/diverging/failed...), residuals and trends, Courant, continuity, ETA.
| Name | Required | Description | Default |
|---|---|---|---|
| log | No | Solver log; default is the most recent log containing time steps. | |
| case | Yes | Case name in the workspace (e.g. 'pitzDaily') or an absolute path inside an allowed root. | |
| plot | No | Also return a residual (and Courant) plot image. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds useful disclosure about what the report contains (state classification, residuals, Courant, continuity, ETA), but says nothing about behavior when no log exists, how the 'most recent log containing time steps' default is chosen, or the cost of requesting a plot.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the key output content front-loaded after the subject. It is efficient, though the parenthetical ellipsis in the state list is slightly loose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing return values and does so by enumerating the report contents. What is missing is failure behavior and the distinction from read_log, but for a read-only diagnostic with fully documented params this is close to sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema itself documents the log default, case naming/path rules, and the plot flag. The description contributes no additional parameter meaning, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (a solver log) and the concrete output of the tool (a convergence report with state, residuals/trends, Courant, continuity, ETA), so an agent knows exactly what it produces. It is distinguishable from generic siblings like read_log because it yields an interpreted report rather than raw text, though it never explicitly contrasts itself with read_log or job_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to reach for this tool versus read_log, job_status, or case_summary, and no mention of prerequisites (e.g. a run must exist) or exclusions. The context is only implied by the field list, which is not the same as guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stop_jobADestructive
Stop a running job, gracefully (results written) or forcibly.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | 'write': solver writes the current time and exits cleanly (controlDict stopAt writeNow); 'now': exit at next step without writing; 'kill': terminate the process group immediately. | write |
| job_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is covered. The description adds real value beyond that by distinguishing a graceful stop that writes results from a forcible one, letting the agent reason about data loss tradeoffs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the key distinction between graceful and forcible stopping is packed into the clause that follows the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no output schema and covered safety annotations, this is minimally adequate. It omits what happens to the job afterward, whether stopping an already-stopped job is safe, and any indication of response, which would be worth a sentence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: the mode enum is fully documented in the schema, while job_id has no description. The description's 'gracefully (results written) or forcibly' loosely maps to the mode values but adds no syntax or default detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (stop) and resource (a running job), and hints at two behavioral variants (graceful vs. forcible). It does not explicitly distinguish itself from siblings like job_status or run, but the intent is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the phrase 'a running job' signals the precondition that a job must be in flight. There is no explicit when-to-use guidance, no mention of alternatives (e.g., job_status to check before stopping), and no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
write_fileA
Create or overwrite a file in a case (new dictionaries, STL-free text inputs, scripts). Prefer set_dict for edits.
| Name | Required | Description | Default |
|---|---|---|---|
| case | Yes | Case name in the workspace (e.g. 'pitzDaily') or an absolute path inside an allowed root. | |
| path | Yes | File path relative to the case. | |
| append | No | ||
| content | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false and openWorldHint=false, so the safety profile is largely pre-supplied. The description usefully discloses that existing content is overwritten rather than merged, though this sits in mild tension with destructiveHint=false, and it says nothing about whether parent directories are created, allowed-root enforcement, or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the primary action and the routing hint second. The parenthetical list is slightly idiosyncratic ('STL-free text inputs') but does real work and costs little space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple write tool with no output schema this covers the essentials: what it writes, that it overwrites, and where to go for edits. Remaining gaps (append behavior, path/root constraints, return value) are minor but not fully closed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: 'case' and 'path' are documented in the schema, while 'append' and 'content' are not. The description hints at content types but never mentions the append flag, so it does not compensate for the uncovered parameters; baseline 3 for the partial coverage is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Create or overwrite a file in a case') and even enumerates the file categories it targets (dictionaries, text inputs, scripts). It explicitly names the sibling it is not (set_dict for edits), so an agent can distinguish it without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Prefer set_dict for edits' gives a clear alternative with the condition that selects it. It does not cover other plausible alternatives such as import_file (for binary/existing assets) or read_file/modify flows, so guidance is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
25 tool updates
v0.1.0- First observed
case_summary - First observed
check_mesh - First observed
clean_case - First observed
create_case - First observed
delete_case - First observed
field_stats - First observed
foam_info - First observed
foam_reference - First observed
get_dict - First observed
import_file - First observed
job_status - First observed
list_cases - First observed
list_files - First observed
list_postprocessing - First observed
list_tutorials - First observed
post_process - First observed
read_file - First observed
read_log - First observed
read_postprocessing - First observed
render - First observed
run - First observed
set_dict - First observed
solver_progress - First observed
stop_job - First observed
write_file
TDQS
Scored across 25 tools
Most tools target clearly distinct resources or actions (case creation, file I/O, dictionary editing, job control, post-processing). However, some overlap exists: read_file also reads dictionaries alongside get_dict, and job_status, solver_progress, and read_log all partially cover log/monitoring use cases. Descriptions help differentiate, but these boundaries are not perfectly crisp.
All names use snake_case, and the dominant pattern is verb_noun (e.g., create_case, list_files, read_log). Deviations are mostly conventional noun phrases for info/status tools (foam_info, job_status) and two bare verbs (run, render), which are readable but not perfectly uniform.
At 25 tools, the surface feels heavy for the domain, sitting at the upper end of the rubric's borderline range. While most tools are individually justified for a complex CFD workflow, some consolidation (e.g., monitoring/log tools, file/dict operations) could reduce the count without losing capability.
The surface covers the main OpenFOAM lifecycle: case setup, file/dictionary editing, running solvers, monitoring, mesh checking, post-processing, and rendering. Minor gaps exist, such as file-level delete/move/rename operations and dedicated mesh-generation tools (though run can invoke them), but these are workable for an agent.
Maintenance
Related MCP Connectors
- OwlCADOAuthcom.owlcad
Parametric 3D CAD for AI agents: build print-ready parts, check them, export STL, 3MF or STEP.
Build, validate, and deploy multi-agent AI solutions from any AI environment.
- mcp-serverOAuthcom.make
Give your AI agents the tools to build, manage, and run automation workflows.
Design, save, and run outcome-aligned AI workflows and verifiers, with reliable image output.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceAutomates OpenFOAM CFD simulations via MCP, enabling AI agents to mesh, run, and post-process cases from natural language prompts without any API keys.MIT
- AlicenseBqualityAmaintenanceEnables AI assistants to drive Kratos Multiphysics finite element simulations end to end, including introspection, scaffolding, execution, and post-processing.522MIT
- AlicenseAqualityAmaintenanceVoice-driven CFD MCP server that enables natural language setup, execution, and analysis of OpenFOAM simulations, with results exportable to ParaView.16MIT
- AlicenseNot gradedqualityCmaintenanceTurns an AI assistant into an OpenFOAM setup and debugging co-pilot, enabling case scaffolding, dictionary edits, mesh sizing, turbulence calculations, and solver log analysis through natural language.MIT



