Skip to main content
Glama

convert_data

Convert inspected sequence data into BPP input files (bpp-seqs), writing seqfile, imap, stats, and loci tables using specified files and imap.

Instructions

Convert the inspected data into BPP input files (bpp-seqs).

Writes PREFIX.txt (the BPP seqfile), PREFIX.imap, PREFIX.stats.tsv and PREFIX.loci.tsv, where PREFIX is out_prefix relative to the project. Use the same files (globs allowed) and imap as inspect_data.

Options, all passed to bpp-seqs unchanged:

  • phasing (BAM/CRAM and gVCF input only): iupac | split | haploid | vcf. Ask the user whether their diploid data are phased; for vcf also give phased_vcf. It does not apply to alignments.

  • reference: designate the reference FASTA explicitly.

  • Locus filters (bpp-seqs defaults apply when omitted): min_length, max_missing, min_snps, keep_invariant; read-based calling: min_bq, min_mq, min_dp, het_freq.

  • overwrite: existing outputs are refused unless true. Ask the user first.

Read in the report: summary.n_loci_passed (also server.nloci) is the nloci value for make_control_file; summary.failure_reasons and loci[] say which loci were dropped and why; output_files names the files written. Tell the user how many loci passed. Next: build_species_tree.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
imapYes
filesYes
min_bqNo
min_dpNo
min_mqNo
phasingNoiupac
het_freqNo
min_snpsNo
overwriteNo
referenceNo
min_lengthNo
out_prefixYes
phased_vcfNo
max_missingNo
keep_invariantNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and destructiveHint=false; the description supplies the substantive behavior: which files are created, that existing outputs are refused unless overwrite=true, that the user must be consulted first, and how to read the resulting report (summary.n_loci_passed, failure_reasons, output_files). This is well beyond what the annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: the first sentence states the action and outputs, then bulleted option groups, then the report-reading guidance. No filler sentences; the length is justified by 15 parameters and no schema documentation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex multi-parameter tool with no output schema, the description covers inputs, side effects, user-confirmation requirements, return-field interpretation, and the follow-on tool. The only small gap is units/precise meaning of a few numeric thresholds (min_dp, min_bq, het_freq), though their grouping under 'read-based calling' gives adequate context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 15 parameters and 0% schema description coverage, the description documents essentially every argument: phasing enum values (iupac | split | haploid | vcf), its applicability limits, phased_vcf pairing, reference, the five locus filters, the four read-calling thresholds, overwrite semantics, out_prefix resolution ('relative to the project'), and files accepting globs. It fully compensates for the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: convert inspected data into BPP input files (bpp-seqs), and enumerates the exact artifacts written (PREFIX.txt, .imap, .stats.tsv, .loci.tsv). It is clearly distinguishable from the sibling inspect_data it references.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit prerequisites ('Use the same files and imap as inspect_data'), conditional guidance tied to input type ('BAM/CRAM and gVCF input only' for phasing; 'does not apply to alignments'), a user-interaction requirement ('Ask the user whether their diploid data are phased'; 'Ask the user first' before overwrite), and a next-step route ('Next: build_species_tree').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.