Skip to main content
Glama

smoke_test

Run BPP briefly to confirm a control file is accepted and its data loads before full analysis.

Instructions

Run BPP briefly to prove it accepts the control file and loads the data.

Call after lint_control_file reports server.status "valid". Runs a copy of the file with a short chain (nsample, burnin) in a scratch folder that is deleted afterwards, from the control file's folder (as the real run will be). The results are NOT an analysis; never report them as findings.

Read in the result:

  • ok: true if BPP finished; false if it failed (see error_line, BPP's fatal message, and output_tail); null if still running at timeout_s without an error, which means BPP read the file and data and started the MCMC. Only when lint is valid AND ok is not false is the file ready. Then call run_command to tell the user how to run the full analysis.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
ctlYes
burninNo
nsampleNo
timeout_sNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and destructiveHint=false, but the description goes further: it discloses that the run happens on a copy in a scratch folder that is deleted afterwards, and that it executes from the control file's folder because the real run will. That is exactly the side-effect and working-directory context an agent needs before invoking a non-read-only tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the one-line purpose, then the precondition, then the mechanics, then the result-reading guide. The bulleted `ok` semantics earn their space, but the result-reading block is somewhat long and could be tightened without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description correctly takes on the return-value burden: it defines `ok` (true/false/null and what null means), and points to `error_line` and `output_tail` for diagnosis. It also closes the loop with the handoff to run_command, so nothing an agent needs to act on the result is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the load, and it partially does: it explains that `nsample` and `burnin` form the short chain and that `timeout_s` governs the still-running state. `ctl` (the only required param) is only implied as 'the file'. Three of four parameters gain meaning the bare schema lacks, but `ctl` itself is left to inference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: run BPP briefly to prove it accepts the control file and loads the data. It is clearly distinguished from siblings lint_control_file (which it depends on) and run_command (which it defers the full analysis to), so an agent can pick it out of the list without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit precondition: 'Call after lint_control_file reports server.status valid.' Explicit exclusion: results are NOT an analysis and must never be reported as findings. Explicit next step: 'Only when lint is valid AND ok is not false... call run_command.' When-to-use, when-not-to-use, and the alternative are all named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.