Skip to main content
Glama
musharna

plant-genomics-mcp

by musharna

Ensembl Plants: Assembly

ensembl_plants_assembly
Read-onlyIdempotent

Retrieve an organism's Ensembl assembly: name, GCA accession, date, karyotype, and all top-level sequence regions with lengths. Karyotype regions come first to plan region walks.

Instructions

Describe an organism's Ensembl assembly (rest.ensembl.org /info/assembly; free, no key): assembly name, GCA accession and date, the karyotype, and every top-level seq-region with its length. The names are the region values ensembl_region_query takes, and a start past a region's length is refused there, so a region walk can be planned before the first call. Karyotype regions come first, in karyotype order, then unplaced scaffolds and contigs, longest first. Names are Ensembl's: tomato's chromosomes are CM001064.4 and so on, not '1'. coord_system labels differ between assemblies (chromosome, scaffold, supercontig, primary_assembly), so in_karyotype, not coord_system, says which regions are chromosomes. total counts every region before limit; truncated=true when limit cut some off (soybean has over 1,100).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
limitNo
organismYesPlant organism — accepts canonical slug (arabidopsis_thaliana), scientific or common name, or NCBI taxid

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
totalYesHow many top-level regions exist upstream for this query, all pages (pre-cap) (#123)
regionsYesKaryotype regions first, in karyotype order; then the rest, longest first
organismYesCanonical organism slug
returnedYesRows in this payload (#123)
karyotypeYesThe chromosomes, in Ensembl's karyotype order
truncatedYesTrue when limit cut regions off
assembly_dateYesAssembly date as Ensembl gives it (YYYY-MM); null when Ensembl gives none
assembly_nameYesAssembly name, e.g. TAIR10, IRGSP-1.0
upstream_versionNoEnsembl release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='ensembl_plants') reports the release its own endpoint calls current at query time, or why there is none. null means Ensembl did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered.
assembly_accessionYesINSDC assembly accession (GCA_...)

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv1.27.0
    • changedOutput schema / properties / upstream_version / description
      Previous value: -"Ensembl release that produced THIS response, — always null today: this backend states no release on its responses. null means Ensembl did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."New value: +"Ensembl release that produced THIS response, — always null today: this backend states no release on its responses; upstream_release(backend='ensembl_plants') reports the release its own endpoint calls current at query time, or why there is none. null means Ensembl did not state one — never that no release exists, and never inferred from a separate metadata call, which can describe a different release than the one that answered."
  2. Addedv1.26.0

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint, but the description goes well beyond them. It discloses ordering (karyotype first, then scaffolds/contigs longest first), naming conventions (Ensembl names like CM001064.4 vs '1'), the in_karyotype flag vs coord_system for identifying chromosomes, and the total-before-limit counting with truncated=true. These are rich behavioral traits that materially affect how an agent interprets results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four sentences long but each sentence carries distinct weight: API/endpoint overview, integration with ensembl_region_query, ordering/naming rules, and limit/truncation behavior. It is front-loaded with the primary purpose and then layered with specifics. While not minimal, there is no filler and every clause earns its place, so a 4 fits.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool is a read-only metadata endpoint with a simple 2-parameter schema and an output schema, the description covers every critical nuance an agent needs: the exact data returned, ordering rules, naming pitfalls, the distinction between coord_system and in_karyotype, and limit/truncation semantics. It even cites concrete examples (tomato, soybean) to make the behavior tangible. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema describes the organism parameter completely (slug, name, taxid) but leaves limit only with constraints. The description compensates by explaining limit's effect: 'total counts every region before limit; truncated=true when limit cut some off', and ties region lengths to downstream region_query validation. This adds practical meaning beyond the raw schema, so it earns a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Describe an organism's Ensembl assembly' and enumerates the exact output fields (assembly name, GCA accession, date, karyotype, seq-regions with lengths). It also distinguishes the tool from its sibling ensembl_region_query by explaining how the names it returns are used there, which clearly positions this tool as the assembly metadata provider.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly name alternatives or draw exclusions, but it gives a concrete usage scenario: 'a region walk can be planned before the first call' with ensembl_region_query. It implies this tool is the prerequisite for region queries and notes it is 'free, no key', which conveys straightforward access. However, it stops short of saying 'use this when X, not when Y', so it stays at 4.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.