Skip to main content
Glama
musharna
by musharna

Infer a phylogenetic tree with bootstrap support

infer_tree
Read-onlyIdempotent

Infer a maximum-likelihood tree from aligned sequences and obtain bootstrap support for every clade to assess branch reliability.

Instructions

Build a maximum-likelihood tree and measure how well the data support it.

Always bootstraps. There is deliberately no option to skip it: an unsupported topology is the failure mode this server exists to prevent.

Args: fasta: Aligned sequences in FASTA, nucleotide or protein as declared by sequence_type. All sequences must be the same length — run align_sequences first if they are not. model: Substitution model, e.g. "JC", "HKY", "GTR+G" (or "LG+G" for protein). Run select_substitution_model first if you do not have a reason to prefer one. replicates: Bootstrap replicates (20-1000). Cost is roughly linear in this, so 100 is a reasonable default and 1000 is for a final answer. seed: 0 to 2**31-1. Fixes the column resampling exactly; the engine's search is seeded too but is not bit-exact (see reproducibility). sequence_type: "dna" (default) or "protein". DECLARED, never sniffed: an alignment of only A/C/G/T is a valid protein alignment too, so guessing would silently fit a nucleotide model to protein data. A protein alignment also needs a protein model — "LG+G" or "WAG", not the nucleotide default — so run select_substitution_model with the same sequence_type first.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
seedNo
fastaYes
modelNoGTR+G
replicatesNo
sequence_typeNodna

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
modelYes
engineYes
newickYes
supportYes
warningsYes
alignmentYes
log_likelihoodYes
reproducibilityYes
branch_length_unitsYes
newick_with_supportYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed4 schema fields changedv0.3.0
    • addedInput schema / properties / sequence_type
      Added value: +{
      +  "default": "dna",
      +  "title": "Sequence Type",
      +  "type": "string"
      +}
    • addedOutput schema / properties / branch_length_units
      Added value: +{
      +  "title": "Branch Length Units",
      +  "type": "string"
      +}
    • addedOutput schema / properties / engine
      Added value: +{
      +  "additionalProperties": true,
      +  "title": "Engine",
      +  "type": "object"
      +}
    • changedOutput schema / required
      Previous value: -[
      -  "newick",
      -  "newick_with_support",
      -  "model",
      -  "log_likelihood",
      -  "alignment",
      -  "support",
      -  "reproducibility",
      -  "warnings"
      -]New value: +[
      +  "newick",
      +  "newick_with_support",
      +  "model",
      +  "log_likelihood",
      +  "alignment",
      +  "support",
      +  "reproducibility",
      +  "engine",
      +  "branch_length_units",
      +  "warnings"
      +]
  2. First observedv0.1.0

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses that bootstrapping is mandatory, that sequence_type is declared rather than sniffed to avoid silent model misfits, and that the engine's search is seeded but not bit-exact. These are meaningful behavioral traits that help the agent anticipate nondeterminism and avoid misuse. The 'not bit-exact' note does not contradict the idempotentHint because it concerns exact numerical reproducibility, not side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is organized with a clear lead sentence, a short behavioral warning, and a labeled Args block. Every sentence adds information; there is no filler. The most important caveats (always bootstraps, declared sequence_type) are front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 parameters, 1 required field, and an output schema already present, the description covers prerequisites, parameter semantics, constraints, defaults, and cost behavior. Nothing critical is missing for an agent to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden, and it succeeds. Every parameter gets a meaningful explanation: fasta alignment requirements, model examples and selection guidance, replicate range and cost tradeoffs, seed range and reproducibility caveats, and the sequence_type default plus protein-specific constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Build a maximum-likelihood tree and measure how well the data support it.' It clearly distinguishes the tool from siblings by explicitly naming align_sequences and select_substitution_model as prerequisites rather than alternatives, and the bootstrap focus is stated upfront.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: run align_sequences first if sequences are not already aligned, and run select_substitution_model first if you lack a preferred model. It also states a hard constraint ('Always bootstraps. There is deliberately no option to skip it') and warns that protein data requires a protein model, leaving little to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.