Skip to main content
Glama
musharna
by musharna

phylokit-mcp

ci PyPI python license Glama DOI

Phylogenetic inference over MCP, driving IQ-TREE 3 through piqtree 0.8, which builds IQ-TREE 3 into its wheel.

A topology without support is not a result. infer_tree always runs a bootstrap and always returns per-clade support. There is no flag to skip it.

6 tools. The test suite runs real IQ-TREE and real MAFFT (no mocked engine), and includes 23 mutation checks and a real-process JSON-RPC handshake test.

Why the rule

A maximum-likelihood tree looks identical whether or not the data support it. Measured here, on alignments simulated from a known 7-taxon tree so the right answer is not in doubt:

sites

informative sites

recovered the true tree?

lowest clade support

300

51

yes, exactly

1.00

60

11

no — RF 2

0.57

At 60 sites the tree contains a clade (C,D,G) that does not exist and omits one that does (E,F,G). Both runs return a fully resolved Newick string of the same shape; nothing about the topology itself distinguishes them. The support values do — and the false clade is the lowest-supported one in the tree.

That is the entire argument for this server. Returning a bare tree returns a result the caller cannot evaluate.

Related MCP server: STRING-db MCP Server

What it reports that a Newick string cannot

  • Conflicting clades — groupings the data support at ≥0.70 that are absent from the reported tree. A support-annotated Newick string has nowhere to attach these, so the standard format silently drops them.

  • fraction_resolved — the share of clades clearing 0.70. The headline number, before any individual grouping is repeated as fact.

  • Model runners-up with ΔAIC — not just a winner. On the 300-site alignment above, simulated under JC, the AIC winner is not JC (TPM2u and TPM2 tie at the top), and several other models, JC among them at ΔAIC 0.7, sit inside the conventional ±2 indistinguishability margin (seed=1, piqtree 0.8.3). A winner without its margin is a claim the numbers do not support.

  • Length versus evidence — n_parsimony_informative alongside n_sites. A 10,000-site alignment of near-identical sequences supports nothing.

Tools

tool

what it does

infer_tree

ML tree plus bootstrap support, per clade. Never one without the other.

select_substitution_model

Ranks 100+ models with ΔAIC/AICc/BIC, and says when the criteria disagree.

compare_trees

Robinson–Foulds distance and the clades that differ. Compares splits, not strings.

simulate_alignment

Generates sequences along a tree you specify — the positive control.

align_sequences

Aligns unaligned FASTA with MAFFT; the output goes straight into infer_tree.

capabilities

piqtree and IQ-TREE versions, 215 substitution models, MAFFT version, limits.

Install

pip install phylokit-mcp

piqtree ships prebuilt wheels, so there is no compiler, no R and no conda step — but it requires Python 3.12+, and so does this package.

align_sequences is the one tool that needs something pip cannot install: the MAFFT binary on PATH (apt install mafft, brew install mafft, or conda install -c bioconda mafft). The other five tools work without it, capabilities reports aligner_version: null, and calling align_sequences returns a refusal that names the install rather than a crash. A MAFFT that is installed but does not answer --version is a different state: aligner_version is still null and aligner_error says what happened.

Configure your MCP client

{
  "mcpServers": {
    "phylokit": {
      "command": "uvx",
      "args": ["phylokit-mcp"]
    }
  }
}

uvx fetches the released package on demand, so this needs no prior install — but it must resolve a Python 3.12+ interpreter, since that is piqtree's wheel floor. If uvx picks an older one, pin it with "args": ["--python", "3.12", "phylokit-mcp"].

If you installed it yourself instead, "command": "phylokit-mcp" works when the executable is on your PATH; give the absolute path to the entry point in the environment you installed into if it is not.

The same file ships as .mcp.json in this repo, which Claude Code picks up automatically when the repo is your working directory.

Reproducibility, stated precisely

Measured, not assumed:

  • Not bit-exact, in any setting. The same request with the same seed — on repeat in one process, or in a fresh process — can return branch lengths and a log-likelihood that differ in the trailing digits (eight fresh processes gave five distinct log-likelihoods, spread ~2e-6). IQ-TREE reads the wall clock during its search: freezing gettimeofday() alone made every run bit-identical. piqtree exposes no option to take the clock out, so the server reports deterministic_across_processes: false rather than promise it.

  • Support can move by a replicate flipping. Over six repeated 50-replicate calls, three of four clades were bit-identical and one moved 0.02, well inside the bootstrap's own sampling error (~0.07 at 50 replicates). The topology and every conclusion were unchanged. The column resampling itself is numpy-seeded and exact.

Compare trees with compare_trees, and numbers with a tolerance — never by string equality.

Threads are pinned to 1 before piqtree is imported: likelihood sums accumulate in thread-completion order, floating-point addition is not associative, and near-tied topologies can flip on the last bits. Pinning is necessary, not sufficient.

Limitations

  • Nucleotide and protein alignments. Pass sequence_type="protein" and a protein model (LG, WAG, …). Codon models are still not exposed. The molecule type is declared, never sniffed: an alignment of only A/C/G/T is a valid protein alignment too (Ala/Cys/Gly/Thr), so guessing would fit a nucleotide model to protein data and return a tree, a likelihood and support values that are all wrong and none of which complain.

  • Bootstrap only — no aLRT, no approximate Bayes, no UFBoot. Support is the nonparametric bootstrap (Felsenstein 1985), computed here rather than read back from IQ-TREE, because piqtree 0.8.3 runs bootstrap_replicates but does not expose the resulting values.

  • Cost is linear in replicates. Each replicate is a full maximum-likelihood search on a resampled alignment, so it grows with taxon count, alignment length and model complexity. Capped at 200 taxa and 1000 replicates.

  • Alignment is MAFFT --auto, single-threaded, and nothing else. No choice of strategy, no profile alignment, no trimming, at most 200 sequences of 100,000 residues, and a 600 s wall-clock cap. Input that already contains gaps is refused rather than silently degapped. The tree tools still refuse ragged input; they do not align it for you.

  • Unrooted trees. No rooting, no dating, no ancestral reconstruction.

Licence

GPL-2.0-only. The "only" is load-bearing: piqtree declares GPL-2.0-only, which is incompatible with GPL-3.0, so the distributed combination cannot be GPL-3. cogent3 is BSD and imposes nothing.

Unofficial. Not affiliated with, endorsed by, or sponsored by the IQ-TREE authors or the cogent3 project. IQ-TREE is academic software and expects to be cited — if results from this server appear in published work, cite IQ-TREE 3 as directed at iqtree.org (currently Wong et al. 2026, doi:10.1093/molbev/msag117, plus the paper for any method used, such as ModelFinder), and piqtree, McArthur et al. 2026, doi:10.1093/molbev/msag061 — not this wrapper. The same holds for MAFFT when align_sequences produced the alignment: Katoh & Standley 2013, doi:10.1093/molbev/mst010. MAFFT is BSD-licensed and is run as a separate program, not linked. See NOTICE.

Available Tools

6 tools
align_sequencesAlign sequences with MAFFT, ready for infer_treeA
Read-onlyIdempotent

Align unaligned sequences with MAFFT, ready for infer_tree.

The returned fasta has every row the same length and can be passed to infer_tree or select_substitution_model unchanged. Each output row, with its gaps removed, is verified to equal the input sequence before it is returned.

Args: fasta: UNALIGNED sequences in FASTA, 2-200 of them. Gap characters are refused: input that is already aligned does not need this tool. sequence_type: "dna" (default) or "protein". Declared, never sniffed, and passed to MAFFT explicitly so it does not guess either.

ParametersJSON Schema
NameRequiredDescriptionDefault
fastaYes
sequence_typeNodna

Output Schema

ParametersJSON Schema
NameRequiredDescription
fastaYes
engineYes
warningsYes
alignmentYes
input_lengthsYes
sequence_typeYes
ready_for_infer_treeYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavior beyond the annotations. It discloses that each output row is verified to equal the input after gap removal, that sequence_type is declared and never sniffed, and that gap characters are refused. These are meaningful constraints not present in the readOnlyHint and idempotentHint annotations, and no contradiction exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise yet information-dense. The opening line states the purpose, the second paragraph covers output guarantees, and the Args section clarifies each parameter. Every sentence adds value; there is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for this tool. It covers input constraints, output behavior, parameter semantics, and how the output integrates with sibling tools. Since an output schema exists, the description need not enumerate return fields, and it does not. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema coverage, the description fully compensates. It explains the fasta parameter: unaligned sequences, 2-200 count, and refusal of gap characters. For sequence_type, it specifies 'dna' or 'protein' and that it is declared and passed explicitly to MAFFT, adding critical semantics beyond the schema's bare type/default.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: 'Align unaligned sequences with MAFFT'. It distinguishes itself from siblings by specifying the input must be unaligned and that the output is ready for infer_tree. It also explicitly notes that already-aligned input does not need this tool, separating it from any alignment-related alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool (unaligned sequences, 2-200) and when not to use it (input with gap characters is refused, and already-aligned input does not need this tool). It also names downstream tools (infer_tree, select_substitution_model) that can consume the output unchanged, giving clear routing context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capabilitiesEngine capabilities and limitsA
Read-onlyIdempotent

What this server can do, and the bounds it enforces.

engine_version is the installed piqtree version (e.g. "0.8.3"); iqtree_version is the IQ-TREE build inside it (e.g. "3.1.2").

Args: include_models: Include the full substitution-model list (long).

ParametersJSON Schema
NameRequiredDescriptionDefault
include_modelsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
engineYes
limitsYes
criteriaYes
aligner_errorYes
engine_versionYes
iqtree_versionYes
threads_pinnedYes
aligner_versionYes
support_thresholdsYes
substitution_modelsNo
n_substitution_modelsYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, and the description adds useful context about the version fields and the behavior of include_models. It also mentions enforced bounds, which goes beyond the structured annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short, front-loaded with the tool's purpose, and each line serves a clear function. The 'Args:' header is slightly redundant with the schema, but there is no wasted prose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity, the existing output schema, and the annotations, the description is complete enough for correct invocation. It covers the version fields and the one behavioral toggle without needing to restate return formats.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the single parameter. The line about include_models adds real meaning by clarifying that it controls whether the full substitution-model list is included and warns that the output is long.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly conveys that the tool reports what the server can do and the limits it enforces, which distinguishes it from the phylogenetic analysis siblings. It lacks an explicit operative verb like 'returns' or 'lists', but the intent is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this is an introspection capability for discovering server capabilities and limits, but it never explicitly says when to call it versus alternatives. It offers no direct 'use this when' guidance or comparison to sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compare_treesCompare two tree topologiesA
Read-onlyIdempotent

Robinson-Foulds distance between two trees, and which clades differ.

Compares SPLITS, not strings: the same topology has many valid Newick representations, so string equality answers a different question.

Args: newick_a: First tree in Newick format. newick_b: Second tree in Newick format.

ParametersJSON Schema
NameRequiredDescriptionDefault
newick_aYes
newick_bYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
warningsYes
only_in_aYes
only_in_bYes
shared_taxaYes
n_shared_cladesYes
robinson_fouldsYes
identical_topologyYes
normalised_robinson_fouldsYes

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and idempotentHint true, so the description adds only the nuance that it compares splits. No contradictions, but minimal extra behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise lines plus Args, no redundancy. Every sentence adds value; front-loaded with purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Tool is simple with two params and has an output schema. Description covers concept, input format, and key distinction (split vs string). Complete for the complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has 0% description coverage, but the description's Args section adds meaning: 'First tree in Newick format' and 'Second tree in Newick format'. This compensates well, though not exhaustive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states it computes Robinson-Foulds distance and identifies differing clades, distinguishing it from string comparison. Specific verb+resource with unique focus.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly compares splits not strings, hinting that string equality is a different task. Sibling tools like infer_tree are distinct, so no confusion. Slightly lacking explicit when-not or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

infer_treeInfer a phylogenetic tree with bootstrap supportA
Read-onlyIdempotent

Build a maximum-likelihood tree and measure how well the data support it.

Always bootstraps. There is deliberately no option to skip it: an unsupported topology is the failure mode this server exists to prevent.

Args: fasta: Aligned sequences in FASTA, nucleotide or protein as declared by sequence_type. All sequences must be the same length — run align_sequences first if they are not. model: Substitution model, e.g. "JC", "HKY", "GTR+G" (or "LG+G" for protein). Run select_substitution_model first if you do not have a reason to prefer one. replicates: Bootstrap replicates (20-1000). Cost is roughly linear in this, so 100 is a reasonable default and 1000 is for a final answer. seed: 0 to 2**31-1. Fixes the column resampling exactly; the engine's search is seeded too but is not bit-exact (see reproducibility). sequence_type: "dna" (default) or "protein". DECLARED, never sniffed: an alignment of only A/C/G/T is a valid protein alignment too, so guessing would silently fit a nucleotide model to protein data. A protein alignment also needs a protein model — "LG+G" or "WAG", not the nucleotide default — so run select_substitution_model with the same sequence_type first.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
fastaYes
modelNoGTR+G
replicatesNo
sequence_typeNodna

Output Schema

ParametersJSON Schema
NameRequiredDescription
modelYes
engineYes
newickYes
supportYes
warningsYes
alignmentYes
log_likelihoodYes
reproducibilityYes
branch_length_unitsYes
newick_with_supportYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses that bootstrapping is mandatory, that sequence_type is declared rather than sniffed to avoid silent model misfits, and that the engine's search is seeded but not bit-exact. These are meaningful behavioral traits that help the agent anticipate nondeterminism and avoid misuse. The 'not bit-exact' note does not contradict the idempotentHint because it concerns exact numerical reproducibility, not side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is organized with a clear lead sentence, a short behavioral warning, and a labeled Args block. Every sentence adds information; there is no filler. The most important caveats (always bootstraps, declared sequence_type) are front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 parameters, 1 required field, and an output schema already present, the description covers prerequisites, parameter semantics, constraints, defaults, and cost behavior. Nothing critical is missing for an agent to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden, and it succeeds. Every parameter gets a meaningful explanation: fasta alignment requirements, model examples and selection guidance, replicate range and cost tradeoffs, seed range and reproducibility caveats, and the sequence_type default plus protein-specific constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Build a maximum-likelihood tree and measure how well the data support it.' It clearly distinguishes the tool from siblings by explicitly naming align_sequences and select_substitution_model as prerequisites rather than alternatives, and the bootstrap focus is stated upfront.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: run align_sequences first if sequences are not already aligned, and run select_substitution_model first if you lack a preferred model. It also states a hard constraint ('Always bootstraps. There is deliberately no option to skip it') and warns that protein data requires a protein model, leaving little to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

select_substitution_modelRank substitution models, with the margin over the runners-upA
Read-onlyIdempotent

Compare substitution models and report how much the winner won by.

A single model name reads as a finding. The ranking, the delta to the next model, and whether AIC/AICc/BIC agree are what make it one.

Args: fasta: Aligned sequences in FASTA, nucleotide or protein as declared by sequence_type. criterion: "AIC", "AICc" or "BIC". BIC penalises parameters more heavily. seed: Fixes the engine's search. top_n: How many ranked models to return. sequence_type: "dna" (default) or "protein". Ranks within that molecule type's model set — nucleotide and protein models are not comparable.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
fastaYes
top_nNo
criterionNoAIC
sequence_typeNodna

Output Schema

ParametersJSON Schema
NameRequiredDescription
seedYes
rankingYes
warningsYes
alignmentYes
criterionYes
best_modelYes
criteria_agreeYes
best_by_criterionYes
n_models_comparedYes
indistinguishable_from_bestYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint and idempotentHint, and the description adds valuable context beyond those: the finding requires ranking plus delta and agreement among criteria, BIC penalizes parameters more heavily, seed fixes the search, and rankings are confined to the declared sequence_type's model set. No contradiction with annotations exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with a clear one-sentence summary, followed by a compact but necessary Args block that compensates for the schema's lack of descriptions. Every sentence adds value, and there is no filler or repetition of annotation data.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity, the presence of an output schema, and read-only/idempotent annotations, the description covers what an agent needs: input requirements, parameter options, behavioral expectations, and the cross-molecule-type limitation. No critical invocation information is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden, and it delivers. It explains all five parameters: fasta as aligned sequences, criterion with its three options and BIC's heavier penalty, seed as a search fix, top_n as the count of ranked models, and sequence_type with the important DNA/protein comparability caveat.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Compare substitution models and report how much the winner won by.' It clearly differentiates this tool from siblings like compare_trees and infer_tree by focusing on substitution-model ranking rather than tree comparison or inference. The title reinforces the same purpose without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool through its purpose and includes meaningful guidance such as the need to report the delta and whether AIC/AICc/BIC agree, plus the warning that nucleotide and protein models are not comparable. However, it never explicitly states when to prefer this tool over alternatives like infer_tree, compare_trees, or align_sequences, nor does it provide exclusions such as 'use this only when...'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

simulate_alignmentSimulate an alignment from a known treeA
Read-onlyIdempotent

Generate sequences along a tree you specify, so the true answer is known.

This is the positive control for everything else here: infer a tree from the output and compare it back with compare_trees. If inference cannot recover a topology you generated from, the problem is the data or the settings, not the biology.

Args: newick: The true tree, with a branch length on every edge. model: Substitution model to simulate under. A protein model ("LG", "WAG", ...) simulates protein; pass the output to infer_tree with sequence_type="protein". alignment.moltype says which it is. length: Number of sites. seed: Fixes the simulation.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
modelNoJC
lengthNo
newickYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
seedYes
fastaYes
modelYes
warningsYes
alignmentYes
true_newickYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only and idempotent behavior. The description adds useful context beyond that: the output has a known true tree, the seed fixes the simulation, protein models produce sequences whose moltype is recorded, and every newick edge needs a branch length. There is no contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose, then gives the practical workflow context, then compactly documents each parameter. There is no filler or repetition of schema defaults; every sentence contributes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simulation tool with an output schema and safety annotations already provided, the description covers the essential usage loop, parameter semantics, and expected result semantics. Nothing an agent needs to call it correctly or interpret its output is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden, and it succeeds. It explains newick (true tree with branch lengths), model (substitution model and moltype implication), length (number of sites), and seed (fixes simulation) in plain terms that add meaning beyond the raw schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Generate sequences along a tree you specify, so the true answer is known.' It also orients the tool as the 'positive control for everything else here,' distinguishing it clearly from inference and comparison siblings like infer_tree and compare_trees.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool: as a positive control for validating inference workflows. It names the follow-up steps (infer_tree, compare_trees) and explains how to interpret failure, but it does not explicitly state when not to use it or name an alternative simulation approach.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.6.0
    • Changedcapabilities2 fields changed
      • addedOutput schema / properties / iqtree_version
        Added value: +{
        +  "title": "Iqtree Version",
        +  "type": "string"
        +}
      • changedOutput schema / required
        Previous value: -[
        -  "engine",
        -  "engine_version",
        -  "n_substitution_models",
        -  "criteria",
        -  "aligner_version",
        -  "aligner_error",
        -  "limits",
        -  "support_thresholds",
        -  "threads_pinned"
        -]New value: +[
        +  "engine",
        +  "engine_version",
        +  "iqtree_version",
        +  "n_substitution_models",
        +  "criteria",
        +  "aligner_version",
        +  "aligner_error",
        +  "limits",
        +  "support_thresholds",
        +  "threads_pinned"
        +]
  2. 2 tool updatesv0.5.0
    • Addedalign_sequences
    • Changedcapabilities4 fields changed
      • addedOutput schema / properties / aligner_error
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "title": "Aligner Error"
        +}
      • addedOutput schema / properties / aligner_version
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "title": "Aligner Version"
        +}
      • removedOutput schema / properties / substitution_models / default
        Removed value: -null
      • changedOutput schema / required
        Previous value: -[
        -  "engine",
        -  "engine_version",
        -  "n_substitution_models",
        -  "criteria",
        -  "limits",
        -  "support_thresholds",
        -  "threads_pinned"
        -]New value: +[
        +  "engine",
        +  "engine_version",
        +  "n_substitution_models",
        +  "criteria",
        +  "aligner_version",
        +  "aligner_error",
        +  "limits",
        +  "support_thresholds",
        +  "threads_pinned"
        +]
  3. 2 tool updatesv0.3.0
    • Changedinfer_tree4 fields changed
      • addedInput schema / properties / sequence_type
        Added value: +{
        +  "default": "dna",
        +  "title": "Sequence Type",
        +  "type": "string"
        +}
      • addedOutput schema / properties / branch_length_units
        Added value: +{
        +  "title": "Branch Length Units",
        +  "type": "string"
        +}
      • addedOutput schema / properties / engine
        Added value: +{
        +  "additionalProperties": true,
        +  "title": "Engine",
        +  "type": "object"
        +}
      • changedOutput schema / required
        Previous value: -[
        -  "newick",
        -  "newick_with_support",
        -  "model",
        -  "log_likelihood",
        -  "alignment",
        -  "support",
        -  "reproducibility",
        -  "warnings"
        -]New value: +[
        +  "newick",
        +  "newick_with_support",
        +  "model",
        +  "log_likelihood",
        +  "alignment",
        +  "support",
        +  "reproducibility",
        +  "engine",
        +  "branch_length_units",
        +  "warnings"
        +]
    • Changedselect_substitution_model1 field changed
      • addedInput schema / properties / sequence_type
        Added value: +{
        +  "default": "dna",
        +  "title": "Sequence Type",
        +  "type": "string"
        +}
  4. 5 tool updatesv0.1.0
    • First observedcapabilities
    • First observedcompare_trees
    • First observedinfer_tree
    • First observedselect_substitution_model
    • First observedsimulate_alignment

TDQS

A4.4/5.0

Scored across 6 tools

Disambiguation5/5

Each tool targets a distinct stage or concern in the phylogenetic workflow: alignment, model selection, tree inference, tree comparison, simulation, and capabilities. There is no functional overlap; even the two 'tree' tools are clearly differentiated by purpose (compare topologies vs. infer from data).

Naming Consistency4/5

Five of six tools follow the consistent verb_noun snake_case pattern (compare_trees, simulate_alignment, infer_tree, select_substitution_model, align_sequences). 'capabilities' breaks the pattern as a bare noun, but it is a standard introspection tool and the deviation is minor.

Tool Count5/5

Six tools is well-scoped for a phylogenetics-focused server; each one earns its place and covers the core pipeline without redundancy or bloat. The count feels deliberately curated rather than padded.

Completeness5/5

The tool surface covers the complete typical workflow: align raw sequences, select a substitution model, infer a bootstrapped tree, compare trees, and simulate data for validation. The descriptions explicitly point to the next step in the pipeline, and there are no obvious dead ends or missing operations.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    B
    quality
    D
    maintenance
    A comprehensive Model Context Protocol (MCP) server for accessing the STRING protein interaction database. This server provides powerful tools for protein network analysis, functional enrichment, and comparative genomics through the STRING API.
    6
    4
    -
  • A
    license
    A
    quality
    D
    maintenance
    Enables the generation, mutation, and evolution of DNA and protein sequences using various evolutionary models and phylogenetic algorithms. It supports realistic next-generation sequencing read simulation and population-level evolutionary tracking for bioinformatics research and testing.
    6
    BSD 2-Clause "Simplified"