bpp-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@bpp-mcpguide me from raw sequence data to a validated BPP control file"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
bpp-mcp
An MCP server that guides a researcher, including one new to BPP, from raw sequence data to a validated BPP control file that has passed a short test run. You talk to an AI assistant (Claude Code, Claude Desktop, Codex, Gemini CLI or any other MCP host); the assistant uses this server's tools to inspect your data, build the species tree, write and check the control file, and test it.
What it is not:
It does not run your analysis. It runs BPP for a few seconds to prove the control file and data load, then tells you the command to start the real run yourself.
It does not read BPP results or judge MCMC convergence.
It holds no BPP knowledge of its own. Syntax, defaults and checks come from the bpp command-line tools it wraps (
bpp-seqs,bpp-tree,bpp-lint,bpp-docs,bpp), and the assistant is told to quote the manual rather than answer from memory.
Install
Linux or macOS, no root needed:
curl -fsSL https://raw.githubusercontent.com/bpp/bpp-mcp/main/install.sh | shThis installs uv into ~/.local/bin if you
don't have it (uv fetches a suitable Python if needed), installs the latest
bpp-mcp release with it, and runs bpp-mcp install-tools, which downloads
the tested releases of the BPP tools into ~/.local/share/bpp-mcp and checks
their SHA-256 hashes. Your shell startup files are not changed. At the end it
prints the command for registering the server with Claude Code or Claude
Desktop.
To do the same steps by hand, with any Python >= 3.10, install the release with the tool you prefer and then fetch the BPP tools:
pipx install https://github.com/bpp/bpp-mcp/releases/latest/download/bpp-mcp.tar.gz
bpp-mcp install-toolsuv tool install <that URL> and pip install <that URL> in a virtual
environment work the same way. bpp-mcp is not on PyPI yet, so
pipx install bpp-mcp does not work.
Run bpp-mcp install-tools again at any time; it only downloads what is
missing or out of date.
Prebuilt tool releases currently cover:
Tool | Linux x86_64 | Linux aarch64 | macOS arm64 | macOS x86_64 |
bpp | yes | yes | yes | yes |
bpp-seqs | yes | – | yes | yes |
bpp-tree | yes | – | yes | yes |
bpp-lint | yes | – | yes | yes |
bpp-docs | yes | – | yes | yes |
For anything missing, build the tool from its repository and put it on PATH
(or set BPP_MCP_<TOOL>, e.g. BPP_MCP_BPP_DOCS=/path/to/bpp-docs). The
check_environment tool reports what was found.
Related MCP server: Bio-MCP FastQC Server
Register it with your AI host
The server is one command, bpp-mcp, speaking MCP over stdio. bpp-mcp install-tools prints its full path; use that path below in place of
/path/to/bpp-mcp.
The server only reads and writes inside one project directory:
BPP_MCP_ROOT, if set; otherwise the directory the host starts the server in.With
BPP_MCP_PROJECTS_DIRset instead, the assistant asks which folder under it to use (theset_projecttool). This is for desktop apps, which start servers in no particular directory.
The first thing the assistant does is call check_environment, which reports
the project directory. If it is not the folder holding your data, set
BPP_MCP_ROOT to that folder's absolute path.
Claude Code
claude mcp add --scope user bpp -- /path/to/bpp-mcpStart claude in the folder holding your data; that folder is the project
directory. /mcp inside a session shows whether the server is connected,
and claude mcp get bpp and claude mcp remove bpp inspect and remove it.
To share the server with everyone working in one repository, use
--scope project, which writes .mcp.json there. This registration was
tested with Claude Code on macOS.
Claude Desktop
Settings > Developer > Edit Config opens claude_desktop_config.json
(~/Library/Application Support/Claude/ on macOS, %APPDATA%\Claude\ on
Windows; bpp-mcp itself supports only macOS and Linux). Add:
{
"mcpServers": {
"bpp": {
"command": "/path/to/bpp-mcp",
"env": { "BPP_MCP_PROJECTS_DIR": "/Users/you/bpp-projects" }
}
}
}Quit and reopen the app. Put each analysis in its own folder under that
directory. The app's log for the server is
~/Library/Logs/Claude/mcp-server-bpp.log.
Codex CLI
codex mcp add bpp -- /path/to/bpp-mcpor in ~/.codex/config.toml:
[mcp_servers.bpp]
command = "/path/to/bpp-mcp"/mcp in Codex lists the server. Codex's documentation does not say which
directory the server starts in. If check_environment reports the wrong
project directory, add cwd = "/path/to/project" to the table or pass
--env BPP_MCP_ROOT=/path/to/project.
Gemini CLI
gemini mcp add -s user bpp /path/to/bpp-mcp(no -- before the command), or in ~/.gemini/settings.json:
{
"mcpServers": {
"bpp": { "command": "/path/to/bpp-mcp" }
}
}/mcp in Gemini CLI lists the server. As with Codex, the starting directory
is not documented; add "cwd": "/path/to/project" or
"env": { "BPP_MCP_ROOT": "/path/to/project" } if the project directory is
wrong.
The Codex and Gemini entries follow those hosts' documentation as of October 2026 and have not been tried by hand.
Switching hosts or models
Nothing in the server depends on the host or the model. Register the same
bpp-mcp command with another host and point it at the same project
directory; the data, tree and control files are ordinary files there, and
any assistant can pick up where another stopped by linting the control file
again.
What the assistant can do
The usual path is: check the installation, inspect the data, convert it, build the tree, make the control file, lint it until valid, test it, and hand you the run command.
Tool | What it does |
| Finds the BPP tools, checks their versions, reports the project directory. |
| Chooses the project folder (only with |
| Summarises your data files and Imap without changing anything. |
| Converts alignments, BAM/CRAM or gVCF data to a BPP sequence file. |
| Tiles a genome into candidate loci when you have no BED file. |
| Writes a sequence file with a subset of the loci, e.g. for a trial run. |
| Builds the species or guide tree from joins such as |
| Reads a tree you already have (Newick, or an old control file). |
| Writes the control file for an A00, A01, A10 or A11 analysis, with priors derived from the data. |
| Changes one keyword in a control file and lints it again. |
| Validates a control file and checks it against its data files. |
| Shows, then applies, the fixes that bring a BPP 2.x/3.x file up to date. |
| Quote the BPP manual. |
| Explains a lint diagnostic code. |
| Runs BPP briefly on a copy of the control file to prove it loads. |
| Gives you the command and folder for the real run. Does not start it. |
Three prompts start the common jobs (in Claude Code they appear as slash
commands): novice_setup, upgrade_old_file and check_my_ctl. Two
resources are available to hosts that load them: bpp://manual/{keyword}
and bpp://examples/{name} (the example control files shipped with BPP and
bpp-lint; bpp://examples lists them).
Privacy: what reaches the model
Whatever a tool returns is sent to the AI model your host uses, which is usually a service run by a third party. The server is built so that your sequences are not:
Tools return summaries: counts, file and sample names, species names, per-locus statistics, lint diagnostics, control-file text, and BPP's screen output from the test run.
Any run of more than 30 nucleotide characters in a result is replaced by a note saying how many characters were removed, and long lists and strings are cut, with a note saying so.
The tests check, for every tool, that no sequence from the test data appears in its output.
Sample names, species names, file names and locus coordinates do reach the
model. The server itself makes no network connections while serving; only
bpp-mcp install-tools downloads anything. The assistant can still read
files itself if your host gives it its own file access, which this server
does not control.
Development
python3 -m venv .venv && .venv/bin/pip install -e '.[test]'
.venv/bin/bpp-mcp install-tools
.venv/bin/pytest -qBPP-MCP-BUILD.md is the specification and claude_next.md the current
status. To release, set __version__, tag vX.Y.Z and push the tag; the
release workflow builds, tests and publishes. Its "Run workflow" button
does everything except publish.
Licence: AGPL-3.0-or-later.
Available Tools
16 toolsbuild_species_treeA
Build the species (or guide) tree from a join formula (bpp-tree).
Writes PREFIX.stree (the species&tree block, with individual counts taken
from the Imap, plus any migration block) and PREFIX.nwk. Pass PREFIX.stree
to make_control_file as stree_file.
joins: comma-separated joins, each 'X+Y' or 'X+Y=name', e.g. 'chimp+bonobo=pan, pan+human, gorilla+pan_human'. A label like A_B refers to the clade of A and B. Tip names are the Imap's species names. Ask the user for the tree (or guide tree for delimitation); do not invent one.migration(MSC-M): bands 'SRC->DST, ...'. Mutually exclusive withintrogression. The control file then also needs a wprior.introgression(MSC-I): events 'DONOR->RECIP phi=0.1, ...'. The control file then also needs a phiprior. Look up those priors with lookup_docs; lint_control_file reports what is missing.
Read in the report: newick, taxa, species_counts,
species_and_tree_block, and warnings (e.g. ROOT_AUTO_JOINED: two
clades were joined at the root automatically) and errors.
server.diagram is an ASCII drawing of the tree. ALWAYS show the user the
diagram and ask them to confirm the topology before going on.
| Name | Required | Description | Default |
|---|---|---|---|
| imap | Yes | ||
| joins | Yes | ||
| migration | No | ||
| out_prefix | Yes | ||
| introgression | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false. The description adds critical behavior beyond that: the exact output files written (PREFIX.stree and PREFIX.nwk), the need to inspect warnings and errors, that warnings can include ROOT_AUTO_JOINED, and that server.diagram must be shown to the user for topology confirmation before proceeding.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then presents output files, parameter guidance, and reporting details in tight bullets. Every sentence adds actionable information with no wasted restatement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and several interacting parameters, the description covers file outputs, parameter interactions, prerequisites (wprior/phiprior), warnings and errors, and the required user-confirmation step with the ASCII diagram. Nothing important for correct invocation appears missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the burden. It thoroughly explains joins with examples, migration bands, and introgression events. It also references imap (tip names and individual counts) and out_prefix (PREFIX.stree/PREFIX.nwk), though it does not explicitly name or define those two parameters as directly as it could.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Build the species (or guide) tree from a join formula (bpp-tree)'. This clearly distinguishes it from the sibling read_species_tree, which reads an existing tree rather than constructing one.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to ask the user for the tree and not invent one. It also tells the agent to pass PREFIX.stree to make_control_file as stree_file, notes that migration and introgression are mutually exclusive, and directs the agent to lookup_docs and lint_control_file for needed priors.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_environmentARead-only
Check that the BPP command-line tools are installed and new enough.
Call this FIRST in every session, before any other tool. It runs
<tool> --version for bpp-seqs, bpp-tree, bpp-lint, bpp-docs and bpp.
Output:
ready: true only when every tool is found and new enough AND a project directory is set. If false, readproblemsand help the user fix them before doing anything else; tell them the exactinstall_hintcommand.tools.<name>:found,path,version,minimum_version,ok, andproblem/install_hintwhen something is wrong. Most tools are installed by the user runningbpp-mcp install-toolsin a terminal (no root needed); you cannot run it for them.project_root: the only directory the tools can read or write. Every path you pass to other tools is relative to it. If it is null, no project is set: ask the user which project folder to use and call set_project (if available), or explain that the host configuration must set BPP_MCP_ROOT.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the readOnlyHint/openWorldHint annotations by disclosing the full output contract (`ready`, `tools.<name>` fields, `project_root`), the constraint that the agent cannot run the install command itself, and that project_root is the only readable/writable directory with all other paths relative to it. This is unusually rich operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the imperative ('Check that ...') and the ordering rule ('Call this FIRST'), then organizes the rest into a scannable Output: list. Because no output schema exists, the return-value detail is load-bearing rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no input parameters and no output schema, the description carries the full burden and meets it: it documents the readiness flag, per-tool fields, failure remediation, and the project_root scoping rule. An agent has everything needed to invoke it and act on its result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description does add useful contextual semantics about the environment it inspects (project_root as the write scope, tool version minimums), but there are no parameters whose meaning could be clarified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: verifying that the BPP command-line tools are installed and sufficiently recent, and names the exact tools checked (bpp-seqs, bpp-tree, bpp-lint, bpp-docs, bpp). This clearly distinguishes it from sibling tools that operate on data, trees, or control files.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs 'Call this FIRST in every session, before any other tool', and explains what to do when the check fails (read `problems`, tell the user the exact `install_hint` command). It also names the alternative path for setting a project (set_project or BPP_MCP_ROOT) when project_root is null.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
convert_dataA
Convert the inspected data into BPP input files (bpp-seqs).
Writes PREFIX.txt (the BPP seqfile), PREFIX.imap, PREFIX.stats.tsv and
PREFIX.loci.tsv, where PREFIX is out_prefix relative to the project.
Use the same files (globs allowed) and imap as inspect_data.
Options, all passed to bpp-seqs unchanged:
phasing(BAM/CRAM and gVCF input only): iupac | split | haploid | vcf. Ask the user whether their diploid data are phased; for vcf also givephased_vcf. It does not apply to alignments.reference: designate the reference FASTA explicitly.Locus filters (bpp-seqs defaults apply when omitted):
min_length,max_missing,min_snps,keep_invariant; read-based calling:min_bq,min_mq,min_dp,het_freq.overwrite: existing outputs are refused unless true. Ask the user first.
Read in the report: summary.n_loci_passed (also server.nloci) is the
nloci value for make_control_file; summary.failure_reasons and loci[]
say which loci were dropped and why; output_files names the files
written. Tell the user how many loci passed. Next: build_species_tree.
| Name | Required | Description | Default |
|---|---|---|---|
| imap | Yes | ||
| files | Yes | ||
| min_bq | No | ||
| min_dp | No | ||
| min_mq | No | ||
| phasing | No | iupac | |
| het_freq | No | ||
| min_snps | No | ||
| overwrite | No | ||
| reference | No | ||
| min_length | No | ||
| out_prefix | Yes | ||
| phased_vcf | No | ||
| max_missing | No | ||
| keep_invariant | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false and destructiveHint=false; the description supplies the substantive behavior: which files are created, that existing outputs are refused unless overwrite=true, that the user must be consulted first, and how to read the resulting report (summary.n_loci_passed, failure_reasons, output_files). This is well beyond what the annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: the first sentence states the action and outputs, then bulleted option groups, then the report-reading guidance. No filler sentences; the length is justified by 15 parameters and no schema documentation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex multi-parameter tool with no output schema, the description covers inputs, side effects, user-confirmation requirements, return-field interpretation, and the follow-on tool. The only small gap is units/precise meaning of a few numeric thresholds (min_dp, min_bq, het_freq), though their grouping under 'read-based calling' gives adequate context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 15 parameters and 0% schema description coverage, the description documents essentially every argument: phasing enum values (iupac | split | haploid | vcf), its applicability limits, phased_vcf pairing, reference, the five locus filters, the four read-calling thresholds, overwrite semantics, out_prefix resolution ('relative to the project'), and files accepting globs. It fully compensates for the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: convert inspected data into BPP input files (bpp-seqs), and enumerates the exact artifacts written (PREFIX.txt, .imap, .stats.tsv, .loci.tsv). It is clearly distinguishable from the sibling inspect_data it references.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit prerequisites ('Use the same files and imap as inspect_data'), conditional guidance tied to input type ('BAM/CRAM and gVCF input only' for phasing; 'does not apply to alignments'), a user-interaction requirement ('Ask the user whether their diploid data are phased'; 'Ask the user first' before overwrite), and a next-step route ('Next: build_species_tree').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explain_diagnosticARead-only
Long explanation of a bpp-lint diagnostic code, e.g. 'BPP101' or '101'.
Use it to explain a lint_control_file diagnostic to the user in plain
language. Returns code and text.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety profile is covered. The description adds that it returns 'code' and 'text', which is useful. However, it doesn't disclose whether the explanation is static or dynamic, or whether all codes are supported.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose and followed by usage instruction and return values. No waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, usage context, input format, and return values. The only gap is lack of explicit differentiation from similar documentation tools in the sibling list, but for a simple single-param read tool, this is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%. The description provides format examples for the 'code' parameter ('BPP101' or '101'), which compensates for the lack of schema description and clarifies accepted input formats. Baseline for 1 param is 4; the description adds meaningful detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (explain) and resource (bpp-lint diagnostic code), and explicitly ties itself to the sibling tool lint_control_file as the source of diagnostics. It's clear what it does, though it doesn't distinguish itself from lookup_docs or search_docs which might overlap for documentation-related queries.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context: 'Use it to explain a lint_control_file diagnostic to the user in plain language.' This tells the agent when to use it (after lint_control_file produces a diagnostic). However, it doesn't explicitly rule out alternatives like lookup_docs or search_docs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_dataARead-only
Inspect the user's data files without changing anything (bpp-seqs --dry-run).
Call after check_environment, on whatever the user has: aligned loci
(FASTA/PHYLIP/NEXUS, one or many per file), BAM/CRAM, gVCF, BED, a reference
FASTA, and the Imap (sample -> species table). File types are detected from
content. files may contain glob patterns relative to the project, e.g.
["loci/*.fasta"].
Read in the report:
workflow(e.g. fasta2bpp) andready_to_run: whether convert_data can run with these inputs.missing[]: inputs still needed (often the Imap). Ask the user for them.cross_validation.issues: mismatches between files, e.g. samples in the data but not the Imap. Explain them to the user before converting.files_provided[]: per-file type and counts;server.file_typescounts the types. Next: convert_data with the same files and Imap.
| Name | Required | Description | Default |
|---|---|---|---|
| imap | No | ||
| files | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true/openWorldHint=false, and the description reinforces this with the dry-run framing and 'without changing anything'. It goes well beyond the annotations by enumerating the report contents (workflow, ready_to_run, missing[], cross_validation.issues, files_provided[], server.file_types) despite there being no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and constraint, then structured with a bulleted breakdown of report fields. It is long, but nearly every line carries actionable information; only minor tightening is possible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and 0% param coverage, the description compensates fully: it explains input file types, detection behavior, glob syntax, and the exact report fields the agent must read and act on. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the load. It documents `files` semantics ('glob patterns relative to the project, e.g. ["loci/*.fasta"]') and characterizes the Imap (sample -> species table), but does not state that `imap` is optional/defaulted to null.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Inspect the user's data files') and immediately qualifies scope with 'without changing anything (bpp-seqs --dry-run)'. This clearly distinguishes it from convert_data (which mutates) and check_environment (which precedes it).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit sequencing: 'Call after check_environment' and 'Next: convert_data with the same files and Imap.' It also describes what to do on missing inputs ('Ask the user for them') and on validation issues ('Explain them to the user before converting'), leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lint_control_fileARead-only
Validate a control file (bpp-lint --json --check-priors) and check it against its data.
Call after make_control_file and after every change. Loop until
server.status is "valid": fix each error (with set_keyword, or by
remaking the file with make_control_file and the corrected arguments),
then lint again.
Read in the result:
server.status: "valid" only if bpp-lint reports no errors AND the temporaryserver.data_checksfound none. Use this, not report.status.report.diagnostics[]: each hascode,severity,message,suggestion,suggested_fix. Use explain_diagnostic(code) to explain one to the user.report.prior_check: priors compared with estimates from the data.server.data_checks.issues[](temporary, until bpp-lint checks these itself): missing data files, nloci larger than the data, sequence tags with no Imap line, Imap species that differ from the tree, phase digit count, speciesdelimitation with the wrong number of arguments. A valid lint does not prove BPP will run: smoke_test next.
| Name | Required | Description | Default |
|---|---|---|---|
| ctl | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations give readOnlyHint=true and openWorldHint=false, but the description adds substantial behavioral context beyond that: the loop-until-valid workflow, that server.status (not report.status) is authoritative, that server.data_checks are temporary, and that a clean lint still does not guarantee BPP will run. This is rich disclosure the annotations do not cover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose in the first sentence, then uses bullets for the workflow and return-value semantics. It is somewhat long and detailed, but each section (loop, server.status, diagnostics, data_checks, smoke_test caveat) carries distinct actionable information rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description fully documents the return surface: server.status vs report.status, report.diagnostics fields and how to explain them, prior_check, and data_checks.issues including the specific checks performed. An agent has everything needed to call it and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single ctl parameter has 0% schema description coverage, so the description must carry the load. It implies ctl is the control file to validate, but never states whether it is a path, filename, or inline content, nor the expected format. Marginal value over the bare schema entry.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Validate a control file") and even names the exact underlying command (bpp-lint --json --check-priors). It is clearly distinguishable from siblings like make_control_file, set_keyword, explain_diagnostic, and smoke_test, which the description explicitly routes to.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ("Call after make_control_file and after every change"), the loop condition, the remediation alternatives (set_keyword, remake with make_control_file), and the next step ("smoke_test next"). Nothing about selection or sequencing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lookup_docsARead-only
The BPP manual's entry for one control-file keyword, quoted verbatim (bpp-docs).
Use this instead of your own knowledge whenever you state BPP syntax,
defaults, allowed values or keyword dependencies, and quote it to the user.
keyword is a control-file variable such as 'thetaprior', 'phase' or
'speciesdelimitation'.
Read in the report: found; syntax, values, default,
dependencies, description, and text (the whole section). If
found is false, use search_docs.
| Name | Required | Description | Default |
|---|---|---|---|
| keyword | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds real behavioral context: results are quoted verbatim, and it enumerates the fields returned in the report (`found`, `syntax`, `values`, `default`, `dependencies`, `description`, `text`). It does not discuss limits on keyword matching or failure modes beyond the `found` flag.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then usage, then parameter meaning and return fields. The backtick-heavy formatting is a little busy but each sentence carries information; nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description compensates by listing the returned report fields, including the `found` flag that drives the fallback. Combined with usage and parameter guidance, an agent has enough to call it correctly; only matching semantics are unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single `keyword` parameter, so the description carries the burden — and it does, defining it as a control-file variable and giving three concrete example values ('thetaprior', 'phase', 'speciesdelimitation'). That compensates well, though it does not specify matching rules (case sensitivity, exact vs partial).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource (retrieve the BPP manual's entry for one control-file keyword), and adds the scope qualifier that it is quoted verbatim. It differentiates itself from sibling `search_docs` by naming it as the fallback path, so an agent can tell the two apart without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance (use this instead of your own knowledge whenever stating BPP syntax, defaults, allowed values or keyword dependencies) and an explicit when-not path (if `found` is false, use search_docs). The alternative sibling is named with its selecting condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
make_control_fileA
Create a BPP control file (bpp-lint --template --suggest-priors).
Never write a control file yourself; use this, then lint_control_file.
analysis: A00 = estimate parameters on a fixed species tree; A01 = estimate the species tree; A10 = species delimitation on a fixed guide tree; A11 = joint species delimitation and species tree. Ask the user, explaining the choices in plain language.seqfile,imapfile: from convert_data (PREFIX.txt, PREFIX.imap).stree_file: from build_species_tree (PREFIX.stree).nloci: from convert_data's server.nloci.out: the control file to write. Data paths are written relative to the control file's folder.phase: for unphased diploid data, one digit per species in species&tree order (look up 'phase' with lookup_docs and ask the user).thetaprior,tauprior: leave unset to use priors derived from the data (inverse-gamma, alpha = 3); set only if the user asks.nsample,burnin,sampfreq,seed: chain settings; template defaults apply when unset. Discuss chain length with the user.extra: any other keywords as {"keyword": "value"}, e.g. {"wprior": "...", "threads": "..."}. Each name is checked against the manual's keyword list. Look up the syntax with lookup_docs first.overwrite: an existing file is refused unless true. Ask the user.
Returns the file's text and path. server.workarounds_applied lists
patches for known bpp-lint bugs. Next: lint_control_file.
| Name | Required | Description | Default |
|---|---|---|---|
| out | Yes | ||
| seed | No | ||
| extra | No | ||
| nloci | Yes | ||
| phase | No | ||
| burnin | No | ||
| jobname | No | out | |
| nsample | No | ||
| seqfile | Yes | ||
| analysis | Yes | ||
| imapfile | Yes | ||
| sampfreq | No | ||
| tauprior | No | ||
| overwrite | No | ||
| stree_file | Yes | ||
| thetaprior | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare a non-readOnly write with destructiveHint=false; the description adds genuinely new behavior: existing files are refused unless overwrite=true, and server.workarounds_applied reports patches for known bpp-lint bugs. Return shape (text and path) is also disclosed, though it does not detail file-size or error modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and an imperative, then organized as tight per-parameter bullets. It is long, but for a 16-parameter tool nearly every line earns its place; the density is justified rather than padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex file-generating tool with no output schema, the description covers inputs, defaults, the overwrite guard, the return payload, bug workarounds, and the next step. An agent has everything needed to call it correctly without opening the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the burden, and it does: it explains analysis codes, the source tools for seqfile/imapfile/stree_file/nloci, path relativization, phase digit rules, prior defaults, chain settings, and extra-keyword validation. Only jobname is undocumented, which is minor against 16 parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a BPP control file') plus the underlying mechanism (bpp-lint --template --suggest-priors), and clearly distinguishes itself from siblings like lint_control_file and upgrade_control_file by positioning itself as the creation step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Never write a control file yourself; use this, then lint_control_file' and gives per-parameter when-to-use rules (ask the user for analysis; leave priors unset unless asked). The workflow handoff is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
make_loci_bedA
Make a BED file of candidate loci by tiling a genome into windows (bpp-seqs windows).
Use when the user has BAM/CRAM or gVCF data but no BED file saying which
regions are the loci (inspect_data then lists a BED under missing).
input is anything holding chromosome names and lengths: the reference
FASTA, a BAM/CRAM or a VCF/gVCF. out is the BED file to write; pass it
to inspect_data and convert_data with the other files.
Ask the user for the locus size and spacing; do not choose them yourself. All options are passed to bpp-seqs unchanged:
window_size: locus length in bp.step: distance between window starts (default: window_size, so windows do not overlap).min_spacing: least distance in bp between kept loci on a chromosome.n_loci: sample this many windows at random (seedfixes the sample); omit to keep them all.include_chrom/exclude_chrom: chromosome names.autosomes_only: skip sex chromosomes, mitochondria and unplaced contigs (by name).skip_edges: drop this many bp at both ends of each chromosome.exclude_regions: a BED file of intervals to avoid.overwrite: an existing file is refused unless true. Ask the user.
Read in the report: n_windows_emitted (loci written) and the counts
before it, which show what each filter removed. Next: inspect_data.
| Name | Required | Description | Default |
|---|---|---|---|
| out | Yes | ||
| seed | No | ||
| step | No | ||
| input | Yes | ||
| n_loci | No | ||
| overwrite | No | ||
| skip_edges | No | ||
| min_spacing | No | ||
| window_size | Yes | ||
| exclude_chrom | No | ||
| include_chrom | No | ||
| autosomes_only | No | ||
| exclude_regions | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false, openWorldHint=false, destructiveHint=false. The description adds substantial context beyond them: overwrite refusal of existing files, the meaning of the report counts ('which show what each filter removed'), and the n_windows_emitted output field. Good added value, though it doesn't cover rate limits or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose sentence, then usage condition, then a dense but necessary bulleted parameter list. It is long, but with 13 parameters every bullet earns its place. Minor redundancy in the repeated 'ask the user' instruction for size/spacing and overwrite.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 13 parameters, no output schema and only hint-level annotations, the description covers what the tool does, when to use it, what each argument means, what the report returns, and the next step. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for all 13 parameters — and it does: it documents window_size, step (with default semantics), min_spacing, n_loci/seed sampling, include/exclude_chrom, autosomes_only, skip_edges, exclude_regions, overwrite, and both input and out. This is a strong example of compensating for empty schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Make a BED file of candidate loci by tiling a genome into windows'), names the bpp-seqs concept, and identifies the inputs. An agent can distinguish this from sibling tools like subset_loci or convert_data without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit triggering condition ('when the user has BAM/CRAM or gVCF data but no BED file ... inspect_data then lists a BED under missing') and names the follow-on tool (inspect_data, convert_data). The 'ask the user for locus size and spacing; do not choose them yourself' instruction is an unusually clear behavioral directive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_species_treeA
Read a species tree the user already has (bpp-tree --read).
Use instead of build_species_tree when the user has a tree file: a Newick (or extended Newick with introgression), a species&tree block, or an old control file containing one. Migration and introgression in the file are recovered too.
imap: the Imap from convert_data; fills in the individual counts. Without it the block has '?' for the counts and cannot be used yet.out_prefix: also write PREFIX.stree and PREFIX.nwk, for make_control_file'sstree_file. Give it together withimap.
Read in the report: status, newick, taxa, species_and_tree_block,
individual_counts_filled, migration, introgression, warnings and
errors. server.diagram is an ASCII drawing of the tree: ALWAYS show
it to the user and ask them to confirm it is the tree they meant.
server.stree_file is set when out_prefix was given.
| Name | Required | Description | Default |
|---|---|---|---|
| imap | No | ||
| path | Yes | ||
| out_prefix | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply readOnlyHint=false/destructiveHint=false/openWorldHint=false; the description goes well beyond them by disclosing that out_prefix writes PREFIX.stree and PREFIX.nwk, that imap is needed for counts to be filled (otherwise the block has '?' and is unusable), and that server.diagram must always be shown to the user for confirmation. The readOnlyHint=false is also resolved, since the write behavior is explained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and the sibling comparison, then organized with bullets for the two optional parameters. The enumerated list of report fields is slightly over-long, but it is functional given the absence of an output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming the returned fields (status, newick, taxa, species_and_tree_block, individual_counts_filled, migration, introgression, warnings, errors) and the derived server.diagram / server.stree_file values. Nothing an agent needs to call this correctly or interpret the result is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the load, and it does for two of three parameters: imap's provenance from convert_data and its effect on individual counts, and out_prefix's side effects and its dependency on imap. The required `path` parameter is only implied ('a species tree the user already has') rather than defined, leaving a small gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Read a species tree') plus the underlying operation (bpp-tree --read), and explicitly distinguishes itself from build_species_tree, the closest sibling. An agent can route between the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use rule ('Use instead of build_species_tree when the user has a tree file') and enumerates the accepted file forms (Newick, extended Newick, species&tree block, old control file). It also states the condition coupling imap and out_prefix, which is actionable guidance rather than inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_commandARead-only
How the user starts the full analysis themselves. Does NOT run BPP.
Call last, once lint_control_file reports server.status "valid" and smoke_test did not fail. This server never starts the real run, which can take hours or days; give the user the command to run in their own terminal or job script.
Read in the result:
directory: the control file's folder (absolute). BPP must be started from it, because the data paths in the file are relative to it.command: the BPP command line, with the full path of the same bpp binary smoke_test used.shell: both as one line to paste into a terminal.jobname,threads: the file's current values (null if unset). Output files are named fromjobname. Look up 'threads' with lookup_docs before advising on it, and change it with set_keyword.notes: what to tell the user about long runs and clusters.
| Name | Required | Description | Default |
|---|---|---|---|
| ctl | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and openWorldHint=false, and the description is consistent with that. It goes well beyond annotations by disclosing the exact returned fields, the constraint that BPP must be launched from `directory` because data paths are relative, that long runs take hours or days, and which fields are null when unset.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core behavior and the 'does NOT run BPP' caveat in the first two sentences, then uses a compact field list. The length is justified because there is no output schema, though the return-value breakdown is slightly verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and a single undocumented parameter, the description fully carries the load: it enumerates every returned field (directory, command, shell, jobname, threads, notes) and their meaning, and points to lookup_docs/set_keyword for follow-up work on threads.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter `ctl` has no schema description (0% coverage), so the description must compensate. It refers generically to 'the control file's folder' and 'the file's current values', which implies `ctl` is the control file path but never states it outright or its format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: it produces the command the user runs to start the full analysis, and explicitly says it does NOT run BPP. This cleanly separates it from execution-oriented siblings like smoke_test and from file-authoring tools like make_control_file.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit ordering rule ('Call last, once lint_control_file reports server.status "valid" and smoke_test did not fail') plus a clear exclusion: this server never starts the real run. The agent knows both the precondition and the boundary of the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_docsARead-only
Ranked full-text search of the BPP manual (bpp-docs --search).
Use for concepts rather than single keywords, e.g. 'migration prior',
'unphased diploid', 'species delimitation algorithm'. Returns the best
matching sections with heading, score and snippet. Follow up with
lookup_docs for a keyword, and quote the manual to the user.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds value beyond that by disclosing the return shape (heading, score, snippet), which matters since no output schema exists, plus the expected follow-up workflow and a behavioral instruction to quote the manual to the user.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and scope, then examples, return shape, and follow-up in descending priority. Four short lines with no filler; the CLI hint '(bpp-docs --search)' is a minor but useful anchor.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotation coverage of return values, the description supplies the field names returned and the intended follow-up path, which is enough to call and use the tool correctly. Pagination or result-count limits are not addressed, a minor gap for a search tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter and schema description coverage is 0%, so the description must carry the load. It does so by showing the expected query style via examples ('migration prior', 'unphased diploid', 'species delimitation algorithm'), which tells the agent to pass natural-language concept phrases rather than single tokens.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Ranked full-text search of the BPP manual') and explicitly contrasts itself with the sibling lookup_docs (concept search vs. keyword lookup). An agent can distinguish it from every other sibling without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives both when to use it ('concepts rather than single keywords') with concrete example queries, and when to use the alternative ('Follow up with lookup_docs for a keyword'). The routing rule is explicit, not inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_keywordA
Change one keyword in a control file, then lint the file again.
Use this for every edit to an existing control file; never edit one by hand. Look up the keyword's syntax with lookup_docs first, and ask the user before changing a scientific choice.
keyword: a control-file keyword, checked against the manual's list.value: everything after the=, on one line, e.g. "200000" for nsample or "invgamma 3 0.002" for thetaprior.
The keyword's line is replaced where it stands (layout and comments are kept); a keyword the file does not have yet is added at the end. A value that spans several lines in the file (such as species&tree) is refused: remake the file with make_control_file instead.
Returns the lint_control_file result for the edited file, plus
server.changed (keyword, old, new; old is null if it was unset).
Read server.status as usual, and run smoke_test again once it is "valid".
| Name | Required | Description | Default |
|---|---|---|---|
| ctl | Yes | ||
| value | Yes | ||
| keyword | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only give the safety profile (readOnly=false, destructive=false, openWorld=false); the description adds substantial behavior: in-place line replacement with comments/layout preserved, append-at-end for missing keywords, refusal of multiline values, and the returned lint result plus server.changed fields and follow-up smoke_test step.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded core action in the first sentence, then prerequisites, then parameter notes, then return/next-step behavior. Multi-paragraph length is justified by real content, though the return-value paragraph could be tightened slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the return contract (lint_control_file result, server.changed with keyword/old/new) and the next step (check server.status, re-run smoke_test). It also covers the main failure mode (multiline value) and the redirect to make_control_file.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and there are no enums, so the description must carry parameter meaning: it explains that keyword is validated against the manual's list and gives concrete value examples ('200000' for nsample, 'invgamma 3 0.002' for thetaprior). The ctl parameter, however, is never explicitly described beyond being 'a control file'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Change one keyword in a control file, then lint the file again.' It clearly distinguishes itself from the file-creating/upgrading siblings by scoping to edits of existing control files.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('Use this for every edit to an existing control file; never edit one by hand'), sequences prerequisites (lookup_docs first, ask user before scientific choices), and names an alternative (make_control_file) for the multiline case it refuses.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
smoke_testA
Run BPP briefly to prove it accepts the control file and loads the data.
Call after lint_control_file reports server.status "valid". Runs a copy of
the file with a short chain (nsample, burnin) in a scratch folder that
is deleted afterwards, from the control file's folder (as the real run
will be). The results are NOT an analysis; never report them as findings.
Read in the result:
ok: true if BPP finished; false if it failed (seeerror_line, BPP's fatal message, andoutput_tail); null if still running attimeout_swithout an error, which means BPP read the file and data and started the MCMC. Only when lint is valid AND ok is not false is the file ready. Then call run_command to tell the user how to run the full analysis.
| Name | Required | Description | Default |
|---|---|---|---|
| ctl | Yes | ||
| burnin | No | ||
| nsample | No | ||
| timeout_s | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=false, but the description goes further: it discloses that the run happens on a copy in a scratch folder that is deleted afterwards, and that it executes from the control file's folder because the real run will. That is exactly the side-effect and working-directory context an agent needs before invoking a non-read-only tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the one-line purpose, then the precondition, then the mechanics, then the result-reading guide. The bulleted `ok` semantics earn their space, but the result-reading block is somewhat long and could be tightened without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description correctly takes on the return-value burden: it defines `ok` (true/false/null and what null means), and points to `error_line` and `output_tail` for diagnosis. It also closes the loop with the handoff to run_command, so nothing an agent needs to act on the result is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load, and it partially does: it explains that `nsample` and `burnin` form the short chain and that `timeout_s` governs the still-running state. `ctl` (the only required param) is only implied as 'the file'. Three of four parameters gain meaning the bare schema lacks, but `ctl` itself is left to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: run BPP briefly to prove it accepts the control file and loads the data. It is clearly distinguished from siblings lint_control_file (which it depends on) and run_command (which it defers the full analysis to), so an agent can pick it out of the list without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit precondition: 'Call after lint_control_file reports server.status valid.' Explicit exclusion: results are NOT an analysis and must never be reported as findings. Explicit next step: 'Only when lint is valid AND ok is not false... call run_command.' When-to-use, when-not-to-use, and the alternative are all named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
subset_lociA
Write a new BPP seqfile holding a subset of the loci of an existing one (bpp-seqs extract).
Use to make a small data set for a trial run, or to drop or keep
particular loci. seqfile is a PREFIX.txt from convert_data. Writes
out_prefix.txt, plus .imap and .loci.tsv when the input has them next to
it. The original files are not changed.
Selection (at least one; passed to bpp-seqs unchanged):
first/last: the first or last N loci.range: 1-based positions such as "1-50" or "1-10,41-50". These three add together.loci: locus names.chrom: loci from this chromosome (needs the .loci.tsv).min_sites/max_sites: by alignment length.Different kinds of selection combine with AND.
invert: keep the loci that do NOT match.imap: use this Imap instead of the one next to the seqfile.overwrite: existing outputs are refused unless true. Ask the user.
Read in the report: n_loci_input, n_loci_kept (also server.nloci:
the nloci value for make_control_file with the new seqfile) and
output_files. Next: make_control_file with the new seqfile, or
set_keyword for seqfile and nloci on an existing control file.
| Name | Required | Description | Default |
|---|---|---|---|
| imap | No | ||
| last | No | ||
| loci | No | ||
| chrom | No | ||
| first | No | ||
| range | No | ||
| invert | No | ||
| seqfile | Yes | ||
| max_sites | No | ||
| min_sites | No | ||
| overwrite | No | ||
| out_prefix | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare a non-read-only, non-destructive, closed-world write, and the description goes well beyond them: it states the original files are not changed, that overwrite is refused unless true (with a user-consent instruction), and that auxiliary .imap/.loci.tsv files are written when present next to the input. It also names the report fields for verification.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well front-loaded with the core action first, then selection rules as a scannable bullet list, then outputs and next steps. It is dense and long, but nearly every clause adds actionable detail; the only mildly redundant line is the terse pointer about passing selection flags 'unchanged' to bpp-seqs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 12 parameters, no schema descriptions and no output schema, the description supplies everything needed: inputs, selection grammar, output artifacts, overwrite safety, and the report keys (n_loci_input, n_loci_kept, server.nloci, output_files) that an agent would otherwise have to discover by trial.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 12 parameters, so the description carries the full burden and delivers: it explains first/last/range semantics (1-based, additively combining), loci, chrom (requires .loci.tsv), min_sites/max_sites, invert, imap overriding the adjacent file, and overwrite behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+scope: writes a new BPP seqfile holding a subset of loci from an existing one, and ties itself to the `bpp-seqs extract` command. An agent can distinguish this from convert_data (which produces seqfiles) and make_control_file (which consumes them) without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the use cases ('make a small data set for a trial run, or to drop or keep particular loci'), specifies the required input provenance (a PREFIX.txt from convert_data), and routes the agent forward to make_control_file or set_keyword. Selection kinds are laid out with their AND/invert combination semantics.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upgrade_control_fileA
Bring a control file written for BPP 2.x/3.x up to current syntax (bpp-lint --diff / --fix).
For users who arrive with an old control file. Call first with
apply=false: nothing is written, and server.upgrade.diff shows the
changes bpp-lint can make by itself (renamed keywords, old value
formats). Show the user the diff. Only after they agree, call again with
apply=true: the file is rewritten in place and the original is kept as
.bak.
Returns the lint_control_file result for the file as it now stands (after
the rewrite when apply=true), plus server.upgrade:
fixes_available: false means bpp-lint has nothing to rewrite.diff: unified diff of the automatic fixes.applied, andbackup(the .bak file) when applied. Automatic fixes are only part of an upgrade: errors left inreport.diagnosticsneed a decision from the user. Make those changes with set_keyword, then lint again untilserver.statusis "valid".
| Name | Required | Description | Default |
|---|---|---|---|
| ctl | Yes | ||
| apply | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by disclosing what each mode does: apply=false writes nothing, apply=true rewrites the file in place and preserves the original as <name>.bak. This destructive/backup behavior is exactly the kind of context annotations don't carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose and the workflow ordering, then enumerates return fields compactly. The return-value bullets are somewhat verbose but justified since there is no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so, documenting fixes_available, diff, applied, and backup, plus the recommended follow-up loop with set_keyword and lint_control_file. Complete for this tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does thoroughly for apply (false = nothing written, true = in-place rewrite with .bak). The ctl parameter itself is not described, but its meaning is evident from the domain context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Bring a control file written for BPP 2.x/3.x up to current syntax') and clarifies its scope with the bpp-lint --diff/--fix mechanism. It is easily distinguishable from siblings like make_control_file and lint_control_file.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames the audience ('For users who arrive with an old control file') and prescribes the two-step workflow: call with apply=false first, show the diff, then apply=true only after agreement. It names set_keyword as the tool for the remaining manual fixes, though it does not explicitly rule out alternatives in a when-not sense.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
16 tool updates
v0.1.0- First observed
build_species_tree - First observed
check_environment - First observed
convert_data - First observed
explain_diagnostic - First observed
inspect_data - First observed
lint_control_file - First observed
lookup_docs - First observed
make_control_file - First observed
make_loci_bed - First observed
read_species_tree - First observed
run_command - First observed
search_docs - First observed
set_keyword - First observed
smoke_test - First observed
subset_loci - First observed
upgrade_control_file
TDQS
Scored across 16 tools
Each tool maps to a distinct BPP pipeline stage: environment check, data inspection/conversion, tree building/reading, control-file creation/editing/linting, documentation lookup/search, smoke testing, and run-command generation. The closest overlap is inspect_data vs convert_data, but the dry-run versus write distinction is made explicit in the descriptions. No tools appear to do the same thing.
All tool names use snake_case and follow a mostly consistent verb_noun or verb_object pattern (check_environment, inspect_data, convert_data, build_species_tree, lint_control_file, run_command). The only minor variant is smoke_test as a compound noun, but it still fits the readable convention and does not disrupt predictability.
The server has 16 tools, slightly above the typical 3-15 range, but each corresponds to a real, non-redundant step in the BPP preparation workflow. Given the domain's complexity (data conversion, BED creation, subsetting, tree handling, control-file lifecycle, docs, smoke test, and run command), the count is reasonable rather than bloated.
The surface covers the core lifecycle from environment check through data conversion, tree building, control-file creation/linting, smoke testing, and final run-command generation. However, check_environment references set_project when no project root is configured, yet that tool is absent, leaving a configuration dead end; post-run BPP result parsing is also outside the surface.
Maintenance
Related MCP Connectors
Run protein folding, docking, design and sequence analysis tools, and read their results.
Hosted DNA/RNA/protein tools: primers, oligos, PCR, cloning, CRISPR, alignment, batch & pipelines.
Public read-only MCP for the Y-chromosome haplogroup tree & Y-SNP resolution. 父系单倍群树查询与Y-SNP解析。
Stateless advisor + validator for Conducted Development: kickoff, artifact validation, rule checks.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to perform DNA/RNA sequence alignment using BWA (Burrows-Wheeler Aligner), supporting both short and long read alignment to reference genomes with indexing, BWA-MEM, and BWA-backtrack algorithms.MIT

Bio-MCP FastQC Serverofficial
AlicenseNot gradedqualityDmaintenanceEnables AI assistants to perform quality control analysis on high-throughput sequencing data using FastQC and MultiQC. It supports single-file and batch processing of FASTQ/FASTA files and generates comprehensive, interactive summary reports.MIT- AlicenseAqualityDmaintenanceEnables the generation, mutation, and evolution of DNA and protein sequences using various evolutionary models and phylogenetic algorithms. It supports realistic next-generation sequencing read simulation and population-level evolutionary tracking for bioinformatics research and testing.6BSD 2-Clause "Simplified"
- FlicenseNot gradedqualityDmaintenanceBioOpenMCP enables users to run bioinformatics tools like FastQC, Cutadapt, and STAR with background execution and status checking. It integrates with Claude Desktop to perform quality control, trimming, alignment, and reporting via natural language.1-