Skip to main content
Glama

subset_loci

Create a smaller BPP seqfile by selecting loci from an existing one for trial runs or dropping specific loci. Keeps originals unchanged while copying imap and loci metadata.

Instructions

Write a new BPP seqfile holding a subset of the loci of an existing one (bpp-seqs extract).

Use to make a small data set for a trial run, or to drop or keep particular loci. seqfile is a PREFIX.txt from convert_data. Writes out_prefix.txt, plus .imap and .loci.tsv when the input has them next to it. The original files are not changed.

Selection (at least one; passed to bpp-seqs unchanged):

  • first / last: the first or last N loci. range: 1-based positions such as "1-50" or "1-10,41-50". These three add together.

  • loci: locus names. chrom: loci from this chromosome (needs the .loci.tsv). min_sites / max_sites: by alignment length.

  • Different kinds of selection combine with AND. invert: keep the loci that do NOT match.

  • imap: use this Imap instead of the one next to the seqfile.

  • overwrite: existing outputs are refused unless true. Ask the user.

Read in the report: n_loci_input, n_loci_kept (also server.nloci: the nloci value for make_control_file with the new seqfile) and output_files. Next: make_control_file with the new seqfile, or set_keyword for seqfile and nloci on an existing control file.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
imapNo
lastNo
lociNo
chromNo
firstNo
rangeNo
invertNo
seqfileYes
max_sitesNo
min_sitesNo
overwriteNo
out_prefixYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare a non-read-only, non-destructive, closed-world write, and the description goes well beyond them: it states the original files are not changed, that overwrite is refused unless true (with a user-consent instruction), and that auxiliary .imap/.loci.tsv files are written when present next to the input. It also names the report fields for verification.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well front-loaded with the core action first, then selection rules as a scannable bullet list, then outputs and next steps. It is dense and long, but nearly every clause adds actionable detail; the only mildly redundant line is the terse pointer about passing selection flags 'unchanged' to bpp-seqs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 12 parameters, no schema descriptions and no output schema, the description supplies everything needed: inputs, selection grammar, output artifacts, overwrite safety, and the report keys (n_loci_input, n_loci_kept, server.nloci, output_files) that an agent would otherwise have to discover by trial.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 12 parameters, so the description carries the full burden and delivers: it explains first/last/range semantics (1-based, additively combining), loci, chrom (requires .loci.tsv), min_sites/max_sites, invert, imap overriding the adjacent file, and overwrite behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: writes a new BPP seqfile holding a subset of loci from an existing one, and ties itself to the `bpp-seqs extract` command. An agent can distinguish this from convert_data (which produces seqfiles) and make_control_file (which consumes them) without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the use cases ('make a small data set for a trial run, or to drop or keep particular loci'), specifies the required input provenance (a PREFIX.txt from convert_data), and routes the agent forward to make_control_file or set_keyword. Selection kinds are laid out with their AND/invert combination semantics.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.