ensembl-mcp-server
Server Details
Look up genes, sequences, variants, homologs, and cross-database xrefs from Ensembl REST.
- Status
- Healthy
- Uptime
- 100.0% over 54 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- cyanheads/ensembl-mcp-server
- GitHub Stars
- 3
- Server Listing
- ensembl-mcp-server
TDQS
Scored across 7 tools
Each tool targets a distinct Ensembl capability: gene resolution, sequence fetching, homology, xrefs, variant prediction, region feature queries, and species discovery. Even the two region-accepting tools are clearly separated by return type (sequence vs. annotated features).
All tools share the consistent ensembl_ prefix and follow verb_object naming: get_homology, get_sequence, get_xrefs, list_species, lookup_gene, predict_variant, query_region. The verb choice varies by action but the pattern is uniform and predictable.
Seven tools is a well-scoped size for an Ensembl data-access server. Each tool covers a major workflow area without redundancy or bloat.
The surface covers the core Ensembl use cases: species discovery, gene lookup, sequence retrieval, homology, cross-references, variant consequence prediction, and region-based feature queries. No obvious dead ends remain for common research workflows.
Available Tools
7 toolsensembl_get_homologyGet Gene HomologsARead-onlyIdempotentInspect
Find orthologs and/or paralogs of a gene across species. Returns each homolog's stable ID, species, homology type (ortholog_one2one, ortholog_one2many, paralog_many2many, etc.), perc_id (percent identity), perc_pos (percent positives), and taxonomy level. Essential for cross-species research — for example, "what is the mouse equivalent of human TP53?" or "how conserved is BRCA2 across mammals?". Provide either symbol + species or a stable gene ID. Target species can be filtered to a single species or left open to return all available homologs.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Ensembl stable gene ID (e.g. ENSG00000139618). Use ensembl_lookup_gene to get the stable ID from a symbol. Cannot be combined with symbol. | |
| type | No | Type of homologs to return. orthologues: genes related by speciation (cross-species equivalents). paralogues: genes related by duplication (within or across species). all: both orthologs and paralogs. | orthologues |
| symbol | No | Gene symbol in the source species (e.g. BRCA2, TP53). Species defaults to homo_sapiens; set species for other organisms. Cannot be combined with id. | |
| species | No | Source species (the species the query gene belongs to) in Ensembl internal format. Default is homo_sapiens. Use ensembl_list_species to discover valid values. | homo_sapiens |
| max_results | No | Maximum number of homologs to return. Broad orthology queries (e.g. BRCA2 across all species) can return 150+ homologs; the default keeps responses focused. Set to 0 to return every homolog uncapped. totalCount always reports the true number available before this cap. | |
| target_species | No | Filter to homologs in a single target species (e.g. mus_musculus for mouse). Omit to return homologs across all available species. Use ensembl_list_species to discover valid values. |
Output Schema
| Name | Required | Description |
|---|---|---|
| cap | No | The max_results limit applied to the homolog list. |
| error | No | Present when the call failed. Absent on success. |
| shown | No | Number of homologs returned after the max_results cap. |
| notice | No | Guidance when no homologs are found or the list was capped. |
| queryId | No | The resolved Ensembl gene ID used for the homology query. |
| homologs | No | Homologous genes found for the query gene, capped to max_results. totalCount reports the full count available before the cap. |
| queryType | No | The homology type queried (orthologues, paralogues, or all). |
| truncated | No | True when the homolog list was capped at max_results. |
| totalCount | No | Total number of homologs available before the max_results cap. Exceeds the returned homologs count when the list was capped. |
| querySpecies | No | The source species used for the query. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, openWorld, and idempotent hints. The description adds valuable behavioral context beyond this: it details the return fields (stable ID, species, homology type, perc_id, etc.), warns about large result sets (150+ homologs), and clarifies the max_results cap and totalCount reporting. This goes beyond simple read-only disclosure and informs the agent of result volume and cap behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured: purpose first, then return details, then usage examples, then input guidance. Every sentence adds information without redundancy. It is appropriately front-loaded and not overlong.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (6 params, output schema present), the description covers purpose, usage scenarios, input flexibility, and filtering options. It fully complements the schema and output schema, including cap behavior and totalCount. An agent can correctly invoke the tool without additional information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and parameter descriptions are detailed, so the baseline is 3. The description adds value by explaining the relationship between id and symbol (mutually exclusive), the default species, and the effect of omitting target_species. It also interprets max_results behavior (cap vs. totalCount), which is not fully inferred from the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific purpose: 'Find orthologs and/or paralogs of a gene across species.' It clearly identifies the verb, resource, and scope, and distinguishes from siblings like ensembl_get_sequence or ensembl_get_xrefs by focusing on homology. The use of examples ('mouse equivalent of human TP53') further clarifies intent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context for when to use the tool (cross-species research, homology queries) and explains the input options (symbol+species or stable ID). It does not explicitly name alternatives or exclusions, but the examples and mention of ensembl_lookup_gene in the schema imply a workflow. Slight deduction for no explicit when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ensembl_get_sequenceGet SequenceARead-onlyIdempotentInspect
Fetch the DNA, cDNA, CDS, or protein sequence for a gene, transcript, protein, or genomic region. Returns a window of the sequence — the first 10,000 characters by default — with its stable ID, molecule type, and full length. When more follows the window, truncated is true and nextOffset is the offset to request next; walking nextOffset reconstructs the whole sequence, and max_length 0 returns everything from offset to the end. The type parameter selects which sequence is fetched: genomic (default, includes introns), cdna (spliced transcript), cds (coding sequence only), protein. For region mode, set id to a region — either species:chr:start-end (e.g. homo_sapiens:13:32315086-32400268) or a bare chr:start-end with species set (e.g. id 13:32315086-32400268, species homo_sapiens), spanning at most 10,000,000 bases; regions return genomic DNA only. Protein sequences require a transcript or protein stable ID (ENST…/ENSP…), not a gene ID — use ensembl_lookup_gene with expand_transcripts=true to get the canonical transcript ID first.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Ensembl stable ID (ENSG…, ENST…, ENSP…) or a genomic region for region mode. Region accepts species:chr:start-end (e.g. homo_sapiens:13:32315086-32400268) or a bare chr:start-end (e.g. 13:32315086-32400268) when the species field is set. A region needs start at or below end, within the sequence region, and spans at most 10,000,000 bases. | |
| type | No | Sequence type to retrieve. genomic: full genomic DNA including introns (default). cdna: spliced transcript sequence (requires ENST… ID). cds: coding sequence only, no UTRs (requires ENST… ID with coding transcript). protein: amino acid sequence (requires ENST… or ENSP… ID). Region ids are genomic-only — request cdna, cds, or protein from a transcript or protein stable ID. | genomic |
| offset | No | 0-based character offset where the returned window starts, counted in the resolved sequence (including any expand_5prime/expand_3prime flank). Default 0. Pass nextOffset from a truncated response to fetch the following window; an offset at or past the end returns an empty window. | |
| species | No | Species in Ensembl internal format (e.g. homo_sapiens). Required for a bare chr:start-end region; optional for the species:chr:start-end form (the embedded species is used when the field is omitted). Optional for stable ID lookups — Ensembl infers species from the ID prefix. | |
| max_length | No | Maximum number of characters in the returned window. Default 10000. Set to 0 to return everything from offset to the end, uncapped. | |
| expand_3prime | No | Number of base pairs to extend downstream (3' direction) of the requested feature. Default 0. Only applies to genomic sequences and region queries. | |
| expand_5prime | No | Number of base pairs to extend upstream (5' direction) of the requested feature. Default 0. Only applies to genomic sequences and region queries. |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | No | The stable ID or region used for the lookup. |
| seq | No | The requested window of the sequence: at most max_length characters starting at offset. DNA sequences use IUPAC nucleotide codes (ACGT + ambiguity codes); protein sequences use single-letter amino acid codes. Empty when offset is at or past the end. |
| type | No | Sequence type returned (genomic, cdna, cds, or protein). |
| error | No | Present when the call failed. Absent on success. |
| length | No | Full sequence length in characters, not the window size — nucleotides for genomic/cdna/cds, amino-acid residues for protein. Includes any expand_5prime/expand_3prime flank. |
| notice | No | Guidance about the window: how to continue when truncated, or why it is empty when the offset is past the end. |
| offset | No | 0-based character offset where this window starts. |
| truncated | No | True when more sequence follows this window; request nextOffset to continue. |
| nextOffset | No | Offset of the first character after this window — pass it as offset to fetch the next window. Present only when truncated. |
| description | No | Sequence description from Ensembl, if provided. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint, idempotentHint), the description discloses truncation and pagination behavior: returns a window of 10,000 characters by default, truncated flag, nextOffset for walking, and max_length 0 returning everything to the end. It also states region mode limits (at most 10,000,000 bases) and that regions return genomic DNA only. This is rich behavioral context that the annotations do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is organized into clear thematic sections: purpose, window/pagination behavior, type selection, and region mode constraints. Every sentence carries operational information, with examples embedded inline. It is long but not verbose; no filler or redundant restatement of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool with 100% schema coverage and an output schema, the description covers all necessary operational aspects: sequence types, identifier requirements, region syntax, pagination, and a prerequisite workflow via ensembl_lookup_gene. It is complete enough for an agent to call the tool correctly without additional information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description adds value beyond the schema by explaining the interaction between type, id, and region mode, giving concrete examples (homo_sapiens:13:32315086-32400268), and clarifying the walking mechanism with nextOffset. It does not add much on expand parameters, but the overall contribution is meaningful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair: 'Fetch the DNA, cDNA, CDS, or protein sequence for a gene, transcript, protein, or genomic region.' It clearly enumerates the sequence types and accepted identifier forms, and the focus on sequence retrieval distinguishes it from sibling tools like ensembl_get_homology or ensembl_predict_variant.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance, including region mode syntax and a crucial alternative: 'Protein sequences require a transcript or protein stable ID (ENST…/ENSP…), not a gene ID — use ensembl_lookup_gene with expand_transcripts=true to get the canonical transcript ID first.' This names the sibling tool and the precise condition under which it should be used instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ensembl_get_xrefsGet Cross-Database ReferencesARead-onlyIdempotentInspect
Retrieve cross-database references for a gene or feature — HGNC, UniProt, EntrezGene, OMIM, RefSeq, Reactome, and others. Returns each xref with its database name, primary ID, display ID, and description. The dbname filter narrows to specific databases; omit to return all xrefs. IDs returned here chain to protein (pubchem via UniProt), literature (pubmed via PubMed IDs), disease (OMIM via MIM_GENE), and pathway (Reactome) resources. Requires an Ensembl stable ID — use ensembl_lookup_gene to get the ENSG… ID first. Common dbname values: HGNC, Uniprot_gn, EntrezGene, MIM_GENE, RefSeq_mRNA, RefSeq_peptide, Reactome, GO (Gene Ontology), ChEMBL.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Ensembl stable gene ID (ENSG…) or transcript ID (ENST…). Use ensembl_lookup_gene to get the stable ID from a gene symbol. xrefs/id returns the full cross-reference set (56+ entries for well-annotated genes like BRCA2). | |
| dbname | No | Filter to a specific external database by its Ensembl internal name. Examples: HGNC (HGNC gene ID), Uniprot_gn (UniProt gene name), EntrezGene (NCBI Gene ID), MIM_GENE (OMIM disease gene), RefSeq_mRNA (NCBI RefSeq transcript), Reactome (pathway IDs), GO (Gene Ontology terms). Omit to return all available xrefs. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present when the call failed. Absent on success. |
| xrefs | No | Cross-database references for the queried Ensembl ID. |
| notice | No | Guidance when no cross-references are found. |
| queriedId | No | The Ensembl stable ID that was queried. |
| totalCount | No | Total number of cross-references returned. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered. The description adds useful behavioral context: it requires an Ensembl stable ID, explains that the returned IDs chain to other resources (protein, literature, disease, pathway), and notes that well-annotated genes return 56+ entries. It doesn't contradict annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph that front-loads the core purpose and then adds usage details. It's slightly long but every sentence earns its place: purpose, return contents, filter behavior, chaining semantics, prerequisite, and common values. The structure could be improved with bullet points, but it's not bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity, the description covers the prerequisite, the filter behavior, the return contents, and the chaining semantics. The output schema exists, so return values don't need to be re-explained. The only minor gap is that it doesn't explicitly state what happens when an invalid ID is provided, but that's a minor omission for a read-only lookup tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters thoroughly. The description adds some value by listing common dbname values and explaining the chaining semantics of returned IDs, but it doesn't go beyond the schema's parameter descriptions in a major way. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Retrieve') and resource ('cross-database references for a gene or feature'), and enumerates the databases involved. It clearly distinguishes this from sibling tools like ensembl_get_homology or ensembl_get_sequence by focusing on xrefs and even names the sibling ensembl_lookup_gene as a prerequisite.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use this tool: after obtaining an Ensembl stable ID via ensembl_lookup_gene, and it explains the dbname filter behavior ('narrow to specific databases; omit to return all xrefs'). It also provides common dbname values, which is actionable guidance for an agent deciding whether to call this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ensembl_list_speciesList Ensembl SpeciesARead-onlyIdempotentInspect
List species supported by Ensembl with display name, common name, assembly, taxon ID, and division. Required discovery step — species names like homo_sapiens are opaque to non-biologists and are the input format every other Ensembl tool expects. Filter by division to select one; use nameContains to find a species by partial name match. With no division, returns the endpoint default division — the vertebrates (~356 species on the default GRCh38 endpoint); pass a division to list that division.
| Name | Required | Description | Default |
|---|---|---|---|
| division | No | Filter to a specific Ensembl division. EnsemblVertebrates includes human, mouse, zebrafish, and other vertebrates. EnsemblPlants covers crop and model plant genomes. EnsemblFungi, EnsemblMetazoa, EnsemblProtists cover non-vertebrate model organisms. Omit to return the endpoint default division (vertebrates). | |
| nameContains | No | Case-insensitive substring filter applied locally after fetching. Matches against species name, display name, and common name. Example: "sapiens" matches homo_sapiens; "mouse" matches mus_musculus. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present when the call failed. Absent on success. |
| notice | No | Guidance when the filter matches no species. |
| species | No | Species matching the filter criteria, sorted by internal name. |
| totalCount | No | Total number of matching species after local filtering. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already communicate read-only and idempotent behavior. The description adds genuinely useful behavioral details beyond that: the default division behavior when no division is passed, the division categories, and the fact that nameContains performs a local case-insensitive substring match. There is no contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient: three sentences cover purpose, workflow role, and filtering behavior. The only slight redundancy is repeating the default-division behavior already stated in the schema, but the rest of the content earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low complexity (2 optional params, 100% schema coverage, output schema present, read-only annotations), the description covers all essential decision points: what the tool returns, why it must be called first, and how to limit results by division or partial name. No meaningful gap remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the input schema already fully explains both division and nameContains with examples. The description adds workflow context—'input format every other Ensembl tool expects'—which helps an agent understand why the parameters matter, but it does not materially change the semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'List species supported by Ensembl' and enumerates the returned fields (display name, common name, assembly, taxon ID, division). It also clearly separates this discovery/list tool from the sibling get/lookup/query/predict tools, which operate on specific species or regions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives strong contextual guidance, calling the tool a 'Required discovery step' and explaining that species names like homo_sapiens are the input format all other Ensembl tools expect. It also explains when to use division and nameContains. It does not explicitly name alternatives to avoid, but the discovery-step framing makes the place in the workflow clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ensembl_lookup_geneLookup GeneARead-onlyIdempotentInspect
Resolve a gene by symbol + species (or by stable ID) to its Ensembl ID, genomic location (chr:start-end:strand), biotype, description, and transcript list. Entry point for most workflows — the stable ID and coordinates returned here are inputs to other tools. Accepts both symbol lookup (BRCA2 + homo_sapiens) and direct ID lookup (ENSG00000139618). Supports batch lookup of up to 20 IDs or symbols in one call via the ids or symbols field. Provide exactly one of symbol, id, ids, or symbols. For symbol lookups species defaults to homo_sapiens (override for other organisms); for ID lookups species is not needed. Use ensembl_list_species to discover valid species names.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Ensembl stable gene ID (e.g. ENSG00000139618 or ENSG00000139618.7 with version). Species is not required for ID lookup. | |
| ids | No | Batch lookup: up to 20 Ensembl stable IDs (ENSG…, ENST…). Returns a succeeded/failed split. Provide exactly one of symbol, id, ids, or symbols. | |
| symbol | No | Gene symbol to look up (e.g. BRCA2, TP53, EGFR). Species defaults to homo_sapiens; set species for other organisms. Case-insensitive in most species. | |
| species | No | Species in Ensembl internal format: lowercase scientific name with underscores (e.g. homo_sapiens, mus_musculus, danio_rerio). Optional for symbol lookups — defaults to homo_sapiens; set it for other organisms. Use ensembl_list_species to discover valid values. | |
| symbols | No | Batch lookup: up to 20 gene symbols. Species defaults to homo_sapiens; set species for other organisms. Returns a succeeded/failed split. Provide exactly one of symbol, id, ids, or symbols. | |
| expand_transcripts | No | When true, include the full transcript list in the response. Each transcript has its ID, biotype, canonical flag, and coordinates. Default is false to keep responses compact. |
Output Schema
| Name | Required | Description |
|---|---|---|
| gene | No | Single gene record. Present for symbol or id lookups. |
| batch | No | Batch results. Present for ids or symbols lookups. |
| error | No | Present when the call failed. Absent on success. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and open-world hints. The description adds useful behavioral detail beyond that: batch results split into succeeded/failed, species defaults, case-insensitivity, compact response behavior, and the one-of constraint. It does not belabor safety since annotations cover it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Six sentences cover purpose, role in workflow, lookup modes, batch limit, one-of constraint, species defaults, and discovery route. Every sentence carries necessary information, and the structure front-loads the core purpose before edge-case guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with six parameters, multiple lookup modes, and an output schema, the description is complete: it explains when to use it, how to choose parameters, defaults, batch behavior, and where to find valid species names. The output schema covers return values, so no further description is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaningful cross-parameter semantics: exactly one of the four lookup fields must be provided, species is only needed for symbol lookups, and ID lookups can skip species. These conditional relationships are not all fully captured by individual property descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Resolve a gene... to its Ensembl ID, genomic location...'), lists concrete output fields, and positions itself as the entry point whose outputs feed other tools. This clearly distinguishes it from siblings like ensembl_get_sequence or ensembl_get_homology even without naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives strong situational context: entry point for most workflows, symbol vs ID lookup, batch mode, default species, and exactly-one-field rule. It stops short of explicitly saying when NOT to use it or naming alternate sibling tools for other tasks, so it is not a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ensembl_predict_variantPredict Variant EffectARead-onlyIdempotentInspect
Predict the functional consequences of a sequence variant using the Ensembl Variant Effect Predictor (VEP). Accepts three input formats: HGVS notation (transcript-relative, e.g. ENST00000380152.8:c.2T>A, or genomic, e.g. 13:g.32316462T>A); region+allele (chr:start:end:strand/allele, e.g. 1:65568:65568:1/T); and a dbSNP rsID (e.g. rs334). Returns the most severe consequence term, affected transcripts and genes, impact level (HIGH/MODERATE/LOW/MODIFIER), and any colocated known variants with clinical significance. HGVS input: provide the full notation including transcript version for best results. Region+allele input: Ensembl normalizes chromosome names and canonical vertebrate output omits the chr prefix (a chr-prefixed name is also accepted). By default the response caps transcript consequences (max_transcript_consequences) and per-variant PubMed IDs (max_pubmed_ids_per_variant) to keep large VEP results compact — well-studied variants like rs334 otherwise carry 60+ consequences and 100+ citations. Truthful totals are always reported; set a cap to 0 (or include_all_colocated_pubmed=true) to retrieve the full set.
| Name | Required | Description | Default |
|---|---|---|---|
| species | No | Species in Ensembl internal format. Default is homo_sapiens. For non-human variants, set the appropriate species (e.g. mus_musculus for mouse). Use ensembl_list_species to discover valid values. | homo_sapiens |
| variant | Yes | Variant in one of three formats: (1) HGVS notation — transcript-relative: ENST00000380152.8:c.2T>A; genomic: 13:g.32316462T>A; (2) Region+allele: chr:start:end:strand/allele — e.g. 1:65568:65568:1/T (strand is 1 for forward or -1 for reverse); (3) dbSNP rsID — e.g. rs334. Ensembl normalizes chromosome names; canonical vertebrate output omits the "chr" prefix, though a chr-prefixed name is also accepted. | |
| max_pubmed_ids_per_variant | No | Maximum PubMed IDs to return per colocated known variant. Well-studied variants (e.g. rs334) cite 100+ papers; the default trims each list. Set to 0 to return every PubMed ID uncapped. pubmedTotal on each colocated variant reports the true pre-cap count. Ignored when include_all_colocated_pubmed is true. | |
| max_transcript_consequences | No | Maximum transcript consequences to return per VEP record. High-impact variants can affect 60+ transcripts; the default keeps the response focused on the top consequences. Set to 0 to return every transcript consequence uncapped. transcriptConsequencesTotal on each record always reports the true pre-cap count. | |
| include_all_colocated_pubmed | No | When true, return every PubMed ID for each colocated variant, overriding max_pubmed_ids_per_variant. Default false to keep responses compact. |
Output Schema
| Name | Required | Description |
|---|---|---|
| cap | No | The max_transcript_consequences limit applied. |
| error | No | Present when the call failed. Absent on success. |
| shown | No | Total transcript consequences returned across all records after the cap. |
| notice | No | Guidance when no results are returned or when caps omitted detail. |
| results | No | VEP consequence records — typically one per input variant. Multiple records appear when a single notation matches multiple genomic positions. |
| truncated | No | True when transcript consequences were capped at max_transcript_consequences. |
| totalCount | No | Number of VEP consequence records returned. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial behavioral context beyond the readOnlyHint/openWorldHint/idempotentHint annotations: it discloses response truncation by default, explains that caps can be set to 0 for full results, and assures that 'truthful totals are always reported.' This is exactly the kind of non-obvious behavior that could affect agent interpretation of results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: purpose, input formats, return highlights, format-specific guidance, and cap behavior are all covered in a compact block. The most important information is front-loaded, and the behavioral caveats about truncation come with concrete remediation instructions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 parameters and an output schema, the description covers all critical aspects an agent needs: supported variant formats, normalization behavior, default truncation, how to retrieve full data, and what response fields to expect. The output schema handles return-value details, so no additional description is necessary for completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description goes beyond the schema by giving concrete examples for every input format, explaining chromosome-name normalization, noting that transcript version improves HGVS results, and clarifying how max_transcript_consequences and max_pubmed_ids_per_variant caps behave (including the 0 and include_all_colocated_pubmed cases). This adds real semantic value for an agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Predict the functional consequences of a sequence variant using the Ensembl Variant Effect Predictor (VEP).' It clearly distinguishes itself from the other ensembl_* tools by focusing on variant consequence prediction, and it enumerates the accepted input formats and return content, leaving no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear contextual guidance: use this tool when you have a sequence variant in HGVS, region+allele, or rsID form and need functional consequences. It does not explicitly name alternative tools or state when not to use it, but the defined purpose and sibling names make the intended use obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ensembl_query_regionQuery Genomic RegionARead-onlyIdempotentInspect
Find genomic features overlapping a chromosomal region: genes, transcripts, variants, regulatory elements, or exons. Returns each feature with its stable ID, type, location, biotype, and name, plus the genome assembly the coordinates are on. Useful for "what's in this locus?" and for seeding follow-up lookups. Region format is chr:start-end (e.g. 13:32315086-32400268 for the BRCA2 locus), spanning at most 5,000,000 bases. Ensembl normalizes chromosome names and canonical vertebrate output omits the chr prefix (13, not chr13); a chr-prefixed name like chr13 is also accepted. The feature parameter defaults to gene only — requesting variation in an 85 kb region matches 44,000+ entries. Explicitly include variation, regulatory, transcript, or exon only when needed. The response returns up to max_results features (default 100) while totalCount always reports the full count; set max_results to 0 for every feature, or query a smaller region to see a different slice. Exon rows carry the parent transcript ID, so the same exon appears once per transcript it belongs to.
| Name | Required | Description | Default |
|---|---|---|---|
| region | Yes | Genomic region in chr:start-end format (e.g. 13:32315086-32400268). Ensembl serves at most 5,000,000 bases per region; split a larger area into smaller windows. Ensembl normalizes chromosome names and canonical vertebrate output omits the chr prefix (13, not chr13); a chr-prefixed name like chr13 is also accepted. For large regions (>100 kb), limit to gene feature type to avoid overwhelming results. | |
| biotype | No | Optional biotype filter (e.g. protein_coding, lncRNA, SNV). Applied server-side by Ensembl. Not all feature types support biotype filtering. | |
| feature | No | Feature types to retrieve — at least one. Default is gene only. Requesting variation in a large region can match tens of thousands of features. Include variation only for targeted small regions (single gene loci or smaller). | |
| species | Yes | Species in Ensembl internal format (e.g. homo_sapiens, mus_musculus). Use ensembl_list_species to discover valid values. | |
| max_results | No | Maximum number of features to return. A gene-length region can hold tens of thousands of variation features; the default keeps the response compact. Set to 0 to return every feature uncapped. totalCount always reports the true number found before this cap. |
Output Schema
| Name | Required | Description |
|---|---|---|
| cap | No | The max_results limit applied to the feature list. |
| error | No | Present when the call failed. Absent on success. |
| shown | No | Number of features returned after the max_results cap. |
| notice | No | Guidance about the result set: empty, large, capped, or missing assembly. |
| region | No | The region queried, as provided. |
| species | No | The species queried. |
| features | No | Genomic features found in the requested region, capped to max_results. totalCount reports the full count found before the cap. |
| truncated | No | True when the feature list was capped at max_results. |
| totalCount | No | Total number of features found in the region before the max_results cap. Exceeds the returned features count when the list was capped. |
| assemblyName | No | Genome assembly the coordinates are on (e.g. GRCh38). Omitted only when it could not be resolved, in which case the notice says so. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the read-only, idempotent annotations by disclosing important behaviors: chromosome-name normalization, the chr-prefix acceptance, the gene-only default, the max_results cap with totalCount always reporting the full count, and the fact that exon rows can be duplicated per parent transcript. This significantly reduces the chance of misinterpreting results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficiently organized: purpose, output shape, use case, region format, chromosome normalization, feature defaults, and response-limit semantics. Each sentence adds distinct value and there is no filler or repetition of annotation fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the tool and the presence of a rich input schema and output schema, the description covers the critical operational details: region constraints, feature selection trade-offs, response cap behavior, and exon duplication. An agent has enough context to call the tool correctly and interpret the response without needing additional external information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all parameters. The description adds valuable parameter-level context beyond the schema, such as the 85 kb variation example, the explicit guidance to include variation only when needed, and the semantics of max_results=0 with totalCount preserving the true count. It does not add much for species or biotype, but those are already well-covered in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly specifies a concrete operation with a verb and resource: 'Find genomic features overlapping a chromosomal region' and enumerates feature types. It is unambiguous about what the tool returns, including stable ID, type, location, biotype, name, and assembly. However, it does not explicitly distinguish itself from sibling tools like ensembl_lookup_gene or ensembl_get_sequence, so the differentiation is implicit rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: it is 'useful for what's in this locus?' and 'for seeding follow-up lookups.' It also provides practical guidance on when to include variation and other feature types versus sticking with the gene default. It does not explicitly name alternatives or state when not to use this tool, but the context is strong enough for an agent to make a reasonable selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
- Changed
ensembl_get_homology1 field changed- added
Input schema / properties / species / minLengthAdded value: +1
- Changed
ensembl_get_sequence14 fields changed- changed
Input schema / properties / id / descriptionPrevious value: -"Ensembl stable ID (ENSG…, ENST…, ENSP…) or a genomic region for region mode. Region accepts species:chr:start-end (e.g. homo_sapiens:13:32315086-32400268) or a bare chr:start-end (e.g. 13:32315086-32400268) when the species field is set."New value: +"Ensembl stable ID (ENSG…, ENST…, ENSP…) or a genomic region for region mode. Region accepts species:chr:start-end (e.g. homo_sapiens:13:32315086-32400268) or a bare chr:start-end (e.g. 13:32315086-32400268) when the species field is set. A region needs start at or below end, within the sequence region, and spans at most 10,000,000 bases." - added
Input schema / properties / id / minLengthAdded value: +1 - added
Input schema / properties / max_lengthAdded value: +{ + "default": 10000, + "description": "Maximum number of characters in the returned window. Default 10000. Set to 0 to return everything from offset to the end, uncapped.", + "maximum": 9007199254740991, + "minimum": 0, + "type": "integer" +} - added
Input schema / properties / offsetAdded value: +{ + "default": 0, + "description": "0-based character offset where the returned window starts, counted in the resolved sequence (including any expand_5prime/expand_3prime flank). Default 0. Pass nextOffset from a truncated response to fetch the following window; an offset at or past the end returns an empty window.", + "maximum": 9007199254740991, + "minimum": 0, + "type": "integer" +} - changed
Input schema / properties / type / descriptionPrevious value: -"Sequence type to retrieve. genomic: full genomic DNA including introns (default). cdna: spliced transcript sequence (requires ENST… ID). cds: coding sequence only, no UTRs (requires ENST… ID with coding transcript). protein: amino acid sequence (requires ENST… or ENSP… ID)."New value: +"Sequence type to retrieve. genomic: full genomic DNA including introns (default). cdna: spliced transcript sequence (requires ENST… ID). cds: coding sequence only, no UTRs (requires ENST… ID with coding transcript). protein: amino acid sequence (requires ENST… or ENSP… ID). Region ids are genomic-only — request cdna, cds, or protein from a transcript or protein stable ID." - changed
Output schema / anyOfPrevious value: -[ - { - "not": { - "required": [ - "error" - ] - }, - "required": [ - "id", - "type", - "seq", - "length" - ] - }, - { - "required": [ - "error" - ] - } -]New value: +[ + { + "not": { + "required": [ + "error" + ] + }, + "required": [ + "id", + "type", + "seq", + "length", + "offset", + "truncated" + ] + }, + { + "required": [ + "error" + ] + } +] - changed
Output schema / properties / error / properties / data / properties / reason / descriptionPrevious value: -"Machine-readable failure mode. Declared by this tool: `not_found`: The stable ID or region was not found in Ensembl. `type_mismatch`: The requested sequence type is incompatible with the provided ID type. `missing_species`: A bare chr:start-end region was given without a species. Other values are possible when a failure originates below the handler."New value: +"Machine-readable failure mode. Declared by this tool: `not_found`: The stable ID or region was not found in Ensembl. `type_mismatch`: A non-genomic type (cdna, cds, or protein) was requested for a region id or a gene ID. `missing_species`: A bare chr:start-end region was given without a species. `invalid_region`: A region id has its start after its end, starts past the end of its sequence region, spans more than the 10,000,000-base maximum, or names a sequence region the species lacks. Other values are possible when a failure originates below the handler." - changed
Output schema / properties / error / properties / data / properties / reason / examplesPrevious value: -[ - "not_found", - "type_mismatch", - "missing_species" -]New value: +[ + "not_found", + "type_mismatch", + "missing_species", + "invalid_region" +] - changed
Output schema / properties / length / descriptionPrevious value: -"Sequence length in characters — nucleotides for genomic/cdna/cds, amino-acid residues for protein. Use this to budget context window usage before processing the sequence."New value: +"Full sequence length in characters, not the window size — nucleotides for genomic/cdna/cds, amino-acid residues for protein. Includes any expand_5prime/expand_3prime flank." - added
Output schema / properties / nextOffsetAdded value: +{ + "description": "Offset of the first character after this window — pass it as offset to fetch the next window. Present only when truncated.", + "type": "number" +} - added
Output schema / properties / noticeAdded value: +{ + "description": "Guidance about the window: how to continue when truncated, or why it is empty when the offset is past the end.", + "type": "string" +} - added
Output schema / properties / offsetAdded value: +{ + "description": "0-based character offset where this window starts.", + "type": "number" +} - changed
Output schema / properties / seq / descriptionPrevious value: -"The full sequence. DNA sequences use IUPAC nucleotide codes (ACGT + ambiguity codes). Protein sequences use single-letter amino acid codes. Large genomic sequences (e.g. 85 kb for BRCA2) are returned in full."New value: +"The requested window of the sequence: at most max_length characters starting at offset. DNA sequences use IUPAC nucleotide codes (ACGT + ambiguity codes); protein sequences use single-letter amino acid codes. Empty when offset is at or past the end." - added
Output schema / properties / truncatedAdded value: +{ + "description": "True when more sequence follows this window; request nextOffset to continue.", + "type": "boolean" +}
- Changed
ensembl_get_xrefs1 field changed- added
Input schema / properties / id / minLengthAdded value: +1
- Changed
ensembl_lookup_gene2 fields changed- added
Input schema / properties / ids / items / minLengthAdded value: +1 - added
Input schema / properties / symbols / items / minLengthAdded value: +1
- Changed
ensembl_predict_variant2 fields changed- added
Input schema / properties / species / minLengthAdded value: +1 - added
Input schema / properties / variant / minLengthAdded value: +1
- Changed
ensembl_query_region14 fields changed- changed
Input schema / properties / feature / descriptionPrevious value: -"Feature types to retrieve. Default is gene only. Requesting variation in a large region can return tens of thousands of features. Include variation only for targeted small regions (single gene loci or smaller)."New value: +"Feature types to retrieve — at least one. Default is gene only. Requesting variation in a large region can match tens of thousands of features. Include variation only for targeted small regions (single gene loci or smaller)." - added
Input schema / properties / feature / minItemsAdded value: +1 - added
Input schema / properties / max_resultsAdded value: +{ + "default": 100, + "description": "Maximum number of features to return. A gene-length region can hold tens of thousands of variation features; the default keeps the response compact. Set to 0 to return every feature uncapped. totalCount always reports the true number found before this cap.", + "maximum": 9007199254740991, + "minimum": 0, + "type": "integer" +} - changed
Input schema / properties / region / descriptionPrevious value: -"Genomic region in chr:start-end format (e.g. 13:32315086-32400268). Ensembl normalizes chromosome names and canonical vertebrate output omits the chr prefix (13, not chr13); a chr-prefixed name like chr13 is also accepted. For large regions (>100 kb), limit to gene feature type to avoid overwhelming results."New value: +"Genomic region in chr:start-end format (e.g. 13:32315086-32400268). Ensembl serves at most 5,000,000 bases per region; split a larger area into smaller windows. Ensembl normalizes chromosome names and canonical vertebrate output omits the chr prefix (13, not chr13); a chr-prefixed name like chr13 is also accepted. For large regions (>100 kb), limit to gene feature type to avoid overwhelming results." - added
Input schema / properties / region / minLengthAdded value: +1 - added
Input schema / properties / species / minLengthAdded value: +1 - added
Output schema / properties / assemblyNameAdded value: +{ + "description": "Genome assembly the coordinates are on (e.g. GRCh38). Omitted only when it could not be resolved, in which case the notice says so.", + "type": "string" +} - added
Output schema / properties / capAdded value: +{ + "description": "The max_results limit applied to the feature list.", + "type": "number" +} - changed
Output schema / properties / error / properties / data / properties / reason / descriptionPrevious value: -"Machine-readable failure mode. Declared by this tool: `invalid_region`: The region string could not be parsed or contains invalid coordinates. `invalid_species`: The species string was not recognized by Ensembl. Other values are possible when a failure originates below the handler."New value: +"Machine-readable failure mode. Declared by this tool: `invalid_region`: The region string could not be parsed, contains invalid coordinates, or spans more than the 5,000,000-base maximum. `invalid_species`: The species string was not recognized by Ensembl. Other values are possible when a failure originates below the handler." - changed
Output schema / properties / features / descriptionPrevious value: -"Genomic features found in the requested region."New value: +"Genomic features found in the requested region, capped to max_results. totalCount reports the full count found before the cap." - changed
Output schema / properties / notice / descriptionPrevious value: -"Warning or guidance about the result set."New value: +"Guidance about the result set: empty, large, capped, or missing assembly." - added
Output schema / properties / shownAdded value: +{ + "description": "Number of features returned after the max_results cap.", + "type": "number" +} - changed
Output schema / properties / totalCount / descriptionPrevious value: -"Number of features returned. Note: very large regions may return truncated results."New value: +"Total number of features found in the region before the max_results cap. Exceeds the returned features count when the list was capped." - added
Output schema / properties / truncatedAdded value: +{ + "description": "True when the feature list was capped at max_results.", + "type": "boolean" +}
7 tool updates
- Changed
ensembl_get_homology6 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Input schema / additionalPropertiesAdded value: +false - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Output schema / anyOfAdded value: +[ + { + "not": { + "required": [ + "error" + ] + }, + "required": [ + "homologs", + "totalCount", + "queryId", + "querySpecies", + "queryType" + ] + }, + { + "required": [ + "error" + ] + } +] - added
Output schema / properties / errorAdded value: +{ + "additionalProperties": {}, + "description": "Present when the call failed. Absent on success.", + "properties": { + "code": { + "description": "JSON-RPC error code for this failure.", + "maximum": 9007199254740991, + "minimum": -9007199254740991, + "type": "integer" + }, + "data": { + "additionalProperties": {}, + "properties": { + "reason": { + "description": "Machine-readable failure mode. Declared by this tool: `not_found`: The gene symbol or stable ID was not found in Ensembl. `no_input`: Neither symbol nor id was provided. `conflicting_input`: Both symbol and id were provided. Other values are possible when a failure originates below the handler.", + "examples": [ + "not_found", + "no_input", + "conflicting_input" + ], + "type": "string" + }, + "recovery": { + "additionalProperties": {}, + "description": "Actionable next step for the caller.", + "properties": { + "hint": { + "type": "string" + } + }, + "required": [ + "hint" + ], + "type": "object" + }, + "retryable": { + "description": "Whether retrying may succeed.", + "type": "boolean" + } + }, + "type": "object" + }, + "message": { + "description": "Human-readable description of what went wrong.", + "type": "string" + } + }, + "required": [ + "code", + "message" + ], + "type": "object" +} - removed
Output schema / requiredRemoved value: -[ - "homologs", - "totalCount", - "queryId", - "querySpecies", - "queryType" -]
- Changed
ensembl_get_sequence6 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Input schema / additionalPropertiesAdded value: +false - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Output schema / anyOfAdded value: +[ + { + "not": { + "required": [ + "error" + ] + }, + "required": [ + "id", + "type", + "seq", + "length" + ] + }, + { + "required": [ + "error" + ] + } +] - added
Output schema / properties / errorAdded value: +{ + "additionalProperties": {}, + "description": "Present when the call failed. Absent on success.", + "properties": { + "code": { + "description": "JSON-RPC error code for this failure.", + "maximum": 9007199254740991, + "minimum": -9007199254740991, + "type": "integer" + }, + "data": { + "additionalProperties": {}, + "properties": { + "reason": { + "description": "Machine-readable failure mode. Declared by this tool: `not_found`: The stable ID or region was not found in Ensembl. `type_mismatch`: The requested sequence type is incompatible with the provided ID type. `missing_species`: A bare chr:start-end region was given without a species. Other values are possible when a failure originates below the handler.", + "examples": [ + "not_found", + "type_mismatch", + "missing_species" + ], + "type": "string" + }, + "recovery": { + "additionalProperties": {}, + "description": "Actionable next step for the caller.", + "properties": { + "hint": { + "type": "string" + } + }, + "required": [ + "hint" + ], + "type": "object" + }, + "retryable": { + "description": "Whether retrying may succeed.", + "type": "boolean" + } + }, + "type": "object" + }, + "message": { + "description": "Human-readable description of what went wrong.", + "type": "string" + } + }, + "required": [ + "code", + "message" + ], + "type": "object" +} - removed
Output schema / requiredRemoved value: -[ - "id", - "type", - "seq", - "length" -]
- Changed
ensembl_get_xrefs6 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Input schema / additionalPropertiesAdded value: +false - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Output schema / anyOfAdded value: +[ + { + "not": { + "required": [ + "error" + ] + }, + "required": [ + "xrefs", + "totalCount", + "queriedId" + ] + }, + { + "required": [ + "error" + ] + } +] - added
Output schema / properties / errorAdded value: +{ + "additionalProperties": {}, + "description": "Present when the call failed. Absent on success.", + "properties": { + "code": { + "description": "JSON-RPC error code for this failure.", + "maximum": 9007199254740991, + "minimum": -9007199254740991, + "type": "integer" + }, + "data": { + "additionalProperties": {}, + "properties": { + "reason": { + "description": "Machine-readable failure mode. Declared by this tool: `not_found`: The Ensembl stable ID was not found or has no cross-references. Other values are possible when a failure originates below the handler.", + "examples": [ + "not_found" + ], + "type": "string" + }, + "recovery": { + "additionalProperties": {}, + "description": "Actionable next step for the caller.", + "properties": { + "hint": { + "type": "string" + } + }, + "required": [ + "hint" + ], + "type": "object" + }, + "retryable": { + "description": "Whether retrying may succeed.", + "type": "boolean" + } + }, + "type": "object" + }, + "message": { + "description": "Human-readable description of what went wrong.", + "type": "string" + } + }, + "required": [ + "code", + "message" + ], + "type": "object" +} - removed
Output schema / requiredRemoved value: -[ - "xrefs", - "totalCount", - "queriedId" -]
- Changed
ensembl_list_species6 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Input schema / additionalPropertiesAdded value: +false - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Output schema / anyOfAdded value: +[ + { + "not": { + "required": [ + "error" + ] + }, + "required": [ + "species", + "totalCount" + ] + }, + { + "required": [ + "error" + ] + } +] - added
Output schema / properties / errorAdded value: +{ + "additionalProperties": {}, + "description": "Present when the call failed. Absent on success.", + "properties": { + "code": { + "description": "JSON-RPC error code for this failure.", + "maximum": 9007199254740991, + "minimum": -9007199254740991, + "type": "integer" + }, + "data": { + "additionalProperties": {}, + "properties": { + "reason": { + "description": "Machine-readable failure mode.", + "type": "string" + }, + "recovery": { + "additionalProperties": {}, + "description": "Actionable next step for the caller.", + "properties": { + "hint": { + "type": "string" + } + }, + "required": [ + "hint" + ], + "type": "object" + }, + "retryable": { + "description": "Whether retrying may succeed.", + "type": "boolean" + } + }, + "type": "object" + }, + "message": { + "description": "Human-readable description of what went wrong.", + "type": "string" + } + }, + "required": [ + "code", + "message" + ], + "type": "object" +} - removed
Output schema / requiredRemoved value: -[ - "species", - "totalCount" -]
- Changed
ensembl_lookup_gene5 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Input schema / additionalPropertiesAdded value: +false - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Output schema / anyOfAdded value: +[ + { + "not": { + "required": [ + "error" + ] + } + }, + { + "required": [ + "error" + ] + } +] - added
Output schema / properties / errorAdded value: +{ + "additionalProperties": {}, + "description": "Present when the call failed. Absent on success.", + "properties": { + "code": { + "description": "JSON-RPC error code for this failure.", + "maximum": 9007199254740991, + "minimum": -9007199254740991, + "type": "integer" + }, + "data": { + "additionalProperties": {}, + "properties": { + "reason": { + "description": "Machine-readable failure mode. Declared by this tool: `not_found`: The gene symbol or stable ID was not found in Ensembl. `invalid_species`: The species string was not recognized by Ensembl. `no_input`: Neither symbol, id, ids, nor symbols was provided. `conflicting_input`: More than one of symbol, id, ids, or symbols was provided. Other values are possible when a failure originates below the handler.", + "examples": [ + "not_found", + "invalid_species", + "no_input", + "conflicting_input" + ], + "type": "string" + }, + "recovery": { + "additionalProperties": {}, + "description": "Actionable next step for the caller.", + "properties": { + "hint": { + "type": "string" + } + }, + "required": [ + "hint" + ], + "type": "object" + }, + "retryable": { + "description": "Whether retrying may succeed.", + "type": "boolean" + } + }, + "type": "object" + }, + "message": { + "description": "Human-readable description of what went wrong.", + "type": "string" + } + }, + "required": [ + "code", + "message" + ], + "type": "object" +}
- Changed
ensembl_predict_variant6 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Input schema / additionalPropertiesAdded value: +false - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Output schema / anyOfAdded value: +[ + { + "not": { + "required": [ + "error" + ] + }, + "required": [ + "results", + "totalCount" + ] + }, + { + "required": [ + "error" + ] + } +] - added
Output schema / properties / errorAdded value: +{ + "additionalProperties": {}, + "description": "Present when the call failed. Absent on success.", + "properties": { + "code": { + "description": "JSON-RPC error code for this failure.", + "maximum": 9007199254740991, + "minimum": -9007199254740991, + "type": "integer" + }, + "data": { + "additionalProperties": {}, + "properties": { + "reason": { + "description": "Machine-readable failure mode. Declared by this tool: `invalid_notation`: The variant notation is malformed or cannot be parsed by VEP. `not_found`: The variant location falls outside any known transcript or assembly region. Other values are possible when a failure originates below the handler.", + "examples": [ + "invalid_notation", + "not_found" + ], + "type": "string" + }, + "recovery": { + "additionalProperties": {}, + "description": "Actionable next step for the caller.", + "properties": { + "hint": { + "type": "string" + } + }, + "required": [ + "hint" + ], + "type": "object" + }, + "retryable": { + "description": "Whether retrying may succeed.", + "type": "boolean" + } + }, + "type": "object" + }, + "message": { + "description": "Human-readable description of what went wrong.", + "type": "string" + } + }, + "required": [ + "code", + "message" + ], + "type": "object" +} - removed
Output schema / requiredRemoved value: -[ - "results", - "totalCount" -]
- Changed
ensembl_query_region6 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Input schema / additionalPropertiesAdded value: +false - changed
Output schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - added
Output schema / anyOfAdded value: +[ + { + "not": { + "required": [ + "error" + ] + }, + "required": [ + "features", + "totalCount", + "region", + "species" + ] + }, + { + "required": [ + "error" + ] + } +] - added
Output schema / properties / errorAdded value: +{ + "additionalProperties": {}, + "description": "Present when the call failed. Absent on success.", + "properties": { + "code": { + "description": "JSON-RPC error code for this failure.", + "maximum": 9007199254740991, + "minimum": -9007199254740991, + "type": "integer" + }, + "data": { + "additionalProperties": {}, + "properties": { + "reason": { + "description": "Machine-readable failure mode. Declared by this tool: `invalid_region`: The region string could not be parsed or contains invalid coordinates. `invalid_species`: The species string was not recognized by Ensembl. Other values are possible when a failure originates below the handler.", + "examples": [ + "invalid_region", + "invalid_species" + ], + "type": "string" + }, + "recovery": { + "additionalProperties": {}, + "description": "Actionable next step for the caller.", + "properties": { + "hint": { + "type": "string" + } + }, + "required": [ + "hint" + ], + "type": "object" + }, + "retryable": { + "description": "Whether retrying may succeed.", + "type": "boolean" + } + }, + "type": "object" + }, + "message": { + "description": "Human-readable description of what went wrong.", + "type": "string" + } + }, + "required": [ + "code", + "message" + ], + "type": "object" +} - removed
Output schema / requiredRemoved value: -[ - "features", - "totalCount", - "region", - "species" -]
2 tool updates
- Changed
ensembl_get_homology7 fields changed- added
Input schema / properties / max_resultsAdded value: +{ + "default": 25, + "description": "Maximum number of homologs to return. Broad orthology queries (e.g. BRCA2 across all species) can return 150+ homologs; the default keeps responses focused. Set to 0 to return every homolog uncapped. totalCount always reports the true number available before this cap.", + "maximum": 9007199254740991, + "minimum": 0, + "type": "integer" +} - added
Output schema / properties / capAdded value: +{ + "description": "The max_results limit applied to the homolog list.", + "type": "number" +} - changed
Output schema / properties / homologs / descriptionPrevious value: -"Homologous genes found for the query gene."New value: +"Homologous genes found for the query gene, capped to max_results. totalCount reports the full count available before the cap." - changed
Output schema / properties / notice / descriptionPrevious value: -"Guidance when no homologs are found."New value: +"Guidance when no homologs are found or the list was capped." - added
Output schema / properties / shownAdded value: +{ + "description": "Number of homologs returned after the max_results cap.", + "type": "number" +} - changed
Output schema / properties / totalCount / descriptionPrevious value: -"Total number of homologs returned."New value: +"Total number of homologs available before the max_results cap. Exceeds the returned homologs count when the list was capped." - added
Output schema / properties / truncatedAdded value: +{ + "description": "True when the homolog list was capped at max_results.", + "type": "boolean" +}
- Changed
ensembl_predict_variant11 fields changed- added
Input schema / properties / include_all_colocated_pubmedAdded value: +{ + "default": false, + "description": "When true, return every PubMed ID for each colocated variant, overriding max_pubmed_ids_per_variant. Default false to keep responses compact.", + "type": "boolean" +} - added
Input schema / properties / max_pubmed_ids_per_variantAdded value: +{ + "default": 10, + "description": "Maximum PubMed IDs to return per colocated known variant. Well-studied variants (e.g. rs334) cite 100+ papers; the default trims each list. Set to 0 to return every PubMed ID uncapped. pubmedTotal on each colocated variant reports the true pre-cap count. Ignored when include_all_colocated_pubmed is true.", + "maximum": 9007199254740991, + "minimum": 0, + "type": "integer" +} - added
Input schema / properties / max_transcript_consequencesAdded value: +{ + "default": 10, + "description": "Maximum transcript consequences to return per VEP record. High-impact variants can affect 60+ transcripts; the default keeps the response focused on the top consequences. Set to 0 to return every transcript consequence uncapped. transcriptConsequencesTotal on each record always reports the true pre-cap count.", + "maximum": 9007199254740991, + "minimum": 0, + "type": "integer" +} - added
Output schema / properties / capAdded value: +{ + "description": "The max_transcript_consequences limit applied.", + "type": "number" +} - changed
Output schema / properties / notice / descriptionPrevious value: -"Guidance when no results are returned."New value: +"Guidance when no results are returned or when caps omitted detail." - changed
Output schema / properties / results / items / properties / colocatedVariants / items / properties / pubmed / descriptionPrevious value: -"PubMed IDs for literature associated with this variant."New value: +"PubMed IDs for literature associated with this variant, capped to max_pubmed_ids_per_variant. pubmedTotal reports the full count when the list was capped." - added
Output schema / properties / results / items / properties / colocatedVariants / items / properties / pubmedTotalAdded value: +{ + "description": "Total PubMed IDs available for this variant before the max_pubmed_ids_per_variant cap. Equals the returned pubmed length when the list was not capped.", + "type": "number" +} - changed
Output schema / properties / results / items / properties / transcriptConsequences / descriptionPrevious value: -"Per-transcript consequence details. High-impact variants may affect many transcripts; focus on canonical transcripts (isCanonical from ensembl_lookup_gene) for primary effect."New value: +"Per-transcript consequence details, capped to max_transcript_consequences. High-impact variants may affect many transcripts; focus on canonical transcripts (isCanonical from ensembl_lookup_gene) for primary effect. transcriptConsequencesTotal reports the full count." - added
Output schema / properties / results / items / properties / transcriptConsequencesTotalAdded value: +{ + "description": "Total transcript consequences available before the max_transcript_consequences cap. Equals the returned transcriptConsequences length when not capped.", + "type": "number" +} - added
Output schema / properties / shownAdded value: +{ + "description": "Total transcript consequences returned across all records after the cap.", + "type": "number" +} - added
Output schema / properties / truncatedAdded value: +{ + "description": "True when transcript consequences were capped at max_transcript_consequences.", + "type": "boolean" +}
1 tool update
- Changed
ensembl_get_sequence2 fields changed- changed
Input schema / properties / id / descriptionPrevious value: -"Ensembl stable ID (ENSG…, ENST…, ENSP…) or region in the format species:chr:start-end (e.g. homo_sapiens:13:32315086-32400268) for region mode. For genomic region queries, species is also required."New value: +"Ensembl stable ID (ENSG…, ENST…, ENSP…) or a genomic region for region mode. Region accepts species:chr:start-end (e.g. homo_sapiens:13:32315086-32400268) or a bare chr:start-end (e.g. 13:32315086-32400268) when the species field is set." - changed
Input schema / properties / species / descriptionPrevious value: -"Species in Ensembl internal format (e.g. homo_sapiens). Required for region mode (when id is a species:chr:start-end string). Optional for stable ID lookups — Ensembl infers species from the ID prefix."New value: +"Species in Ensembl internal format (e.g. homo_sapiens). Required for a bare chr:start-end region; optional for the species:chr:start-end form (the embedded species is used when the field is omitted). Optional for stable ID lookups — Ensembl infers species from the ID prefix."
2 tool updates
- Changed
ensembl_predict_variant1 field changed- changed
Input schema / properties / variant / descriptionPrevious value: -"Variant in one of two formats: (1) HGVS notation — transcript-relative: ENST00000380152.8:c.2T>A; genomic: 13:g.32316462T>A; (2) Region+allele: chr:start:end:strand/allele — e.g. 1:65568:65568:1/T. For region+allele, strand is 1 (forward) or -1 (reverse); chromosome names use no \"chr\" prefix for vertebrates."New value: +"Variant in one of three formats: (1) HGVS notation — transcript-relative: ENST00000380152.8:c.2T>A; genomic: 13:g.32316462T>A; (2) Region+allele: chr:start:end:strand/allele — e.g. 1:65568:65568:1/T (strand is 1 for forward or -1 for reverse); (3) dbSNP rsID — e.g. rs334. Ensembl normalizes chromosome names; canonical vertebrate output omits the \"chr\" prefix, though a chr-prefixed name is also accepted."
- Changed
ensembl_query_region3 fields changed- changed
Input schema / properties / region / descriptionPrevious value: -"Genomic region in chr:start-end format (e.g. 13:32315086-32400268). Chromosome names use Ensembl format — no \"chr\" prefix for vertebrates (13, not chr13). For large regions (>100 kb), limit to gene feature type to avoid overwhelming results."New value: +"Genomic region in chr:start-end format (e.g. 13:32315086-32400268). Ensembl normalizes chromosome names and canonical vertebrate output omits the chr prefix (13, not chr13); a chr-prefixed name like chr13 is also accepted. For large regions (>100 kb), limit to gene feature type to avoid overwhelming results." - added
Output schema / properties / features / items / properties / parentIdAdded value: +{ + "description": "Parent transcript ID (ENST…) for exon features. An exon is reported once per parent transcript it belongs to, so the same exon ID can appear on multiple rows that differ only by this field — not duplicates.", + "type": "string" +} - added
Output schema / properties / features / items / properties / rankAdded value: +{ + "description": "Position (1-based) of an exon within its parent transcript.", + "type": "number" +}
1 tool update
- Changed
ensembl_list_species1 field changed- changed
Input schema / properties / division / descriptionPrevious value: -"Filter to a specific Ensembl division. EnsemblVertebrates includes human, mouse, zebrafish, and other vertebrates. EnsemblPlants covers crop and model plant genomes. EnsemblFungi, EnsemblMetazoa, EnsemblProtists cover non-vertebrate model organisms. Omit to return all divisions."New value: +"Filter to a specific Ensembl division. EnsemblVertebrates includes human, mouse, zebrafish, and other vertebrates. EnsemblPlants covers crop and model plant genomes. EnsemblFungi, EnsemblMetazoa, EnsemblProtists cover non-vertebrate model organisms. Omit to return the endpoint default division (vertebrates)."
3 tool updates
- Changed
ensembl_get_homology2 fields changed- changed
Input schema / properties / id / descriptionPrevious value: -"Ensembl stable gene ID (e.g. ENSG00000139618). Use ensembl_lookup_gene to get the stable ID from a symbol. Cannot be used together with symbol."New value: +"Ensembl stable gene ID (e.g. ENSG00000139618). Use ensembl_lookup_gene to get the stable ID from a symbol. Cannot be combined with symbol." - changed
Input schema / properties / symbol / descriptionPrevious value: -"Gene symbol in the source species (e.g. BRCA2, TP53). Requires species to be set. Cannot be used together with id."New value: +"Gene symbol in the source species (e.g. BRCA2, TP53). Species defaults to homo_sapiens; set species for other organisms. Cannot be combined with id."
- Changed
ensembl_get_sequence3 fields changed- added
Output schema / properties / lengthAdded value: +{ + "description": "Sequence length in characters — nucleotides for genomic/cdna/cds, amino-acid residues for protein. Use this to budget context window usage before processing the sequence.", + "type": "number" +} - removed
Output schema / properties / lengthInBpRemoved value: -{ - "description": "Sequence length in characters (nucleotides or amino acids). Use this to budget context window usage before processing the sequence.", - "type": "number" -} - changed
Output schema / requiredPrevious value: -[ - "id", - "type", - "seq", - "lengthInBp" -]New value: +[ + "id", + "type", + "seq", + "length" +]
- Changed
ensembl_lookup_gene4 fields changed- changed
Input schema / properties / ids / descriptionPrevious value: -"Batch lookup: up to 20 Ensembl stable IDs (ENSG…, ENST…). Returns a succeeded/failed split. Cannot be combined with symbol or id."New value: +"Batch lookup: up to 20 Ensembl stable IDs (ENSG…, ENST…). Returns a succeeded/failed split. Provide exactly one of symbol, id, ids, or symbols." - changed
Input schema / properties / species / descriptionPrevious value: -"Species in Ensembl internal format: lowercase scientific name with underscores (e.g. homo_sapiens, mus_musculus, danio_rerio). Required when using symbol. Default is homo_sapiens for symbol-based lookups. Use ensembl_list_species to discover valid values."New value: +"Species in Ensembl internal format: lowercase scientific name with underscores (e.g. homo_sapiens, mus_musculus, danio_rerio). Optional for symbol lookups — defaults to homo_sapiens; set it for other organisms. Use ensembl_list_species to discover valid values." - changed
Input schema / properties / symbol / descriptionPrevious value: -"Gene symbol to look up (e.g. BRCA2, TP53, EGFR). Requires species to be set. Case-insensitive in most species."New value: +"Gene symbol to look up (e.g. BRCA2, TP53, EGFR). Species defaults to homo_sapiens; set species for other organisms. Case-insensitive in most species." - changed
Input schema / properties / symbols / descriptionPrevious value: -"Batch lookup: up to 20 gene symbols. Requires species to be set. Returns a succeeded/failed split. Cannot be combined with symbol, id, or ids."New value: +"Batch lookup: up to 20 gene symbols. Species defaults to homo_sapiens; set species for other organisms. Returns a succeeded/failed split. Provide exactly one of symbol, id, ids, or symbols."
7 tool updates
- First observed
ensembl_get_homology - First observed
ensembl_get_sequence - First observed
ensembl_get_xrefs - First observed
ensembl_list_species - First observed
ensembl_lookup_gene - First observed
ensembl_predict_variant - First observed
ensembl_query_region
Related MCP Connectors
UCSC Genome Browser REST API — reference genome assemblies for ~250 species, the annotation tracks…
Look up Pokémon, moves, abilities, items, natures, and type matchups from PokéAPI v2.
dbSNP refSNP records and HGVS/SPDI/rsID normalization for human genetic variants, from NCBI…
Related MCP Servers
- AlicenseBqualityDmaintenanceProvides access to the Ensembl genomics REST API with 30+ tools for genomic data including gene lookup, sequence retrieval, genetic variants, cross-species homology, phenotypes, and regulatory features.25116 npmISC
- AlicenseNot gradedqualityBmaintenanceEnables querying Ensembl genomic data including gene lookup, sequence retrieval, homology, variation, and variant effect prediction via MCP tools.252 npmMIT
- AlicenseAqualityCmaintenanceEnables variant annotation and effect prediction using the Ensembl VEP API, with support for batch and single queries.91MIT
- AlicenseNot gradedqualityCmaintenanceEnables querying UCSC Genome Browser assemblies, annotation tracks, track data over genomic intervals, and raw DNA sequences across ~250 species without authentication.70 npmMIT
Glama MCP Gateway
Add one secure layer between your agents and this server.