Skip to main content
Glama

Get Sequence

ensembl_get_sequence
Read-onlyIdempotent

Fetch the DNA, cDNA, CDS, or protein sequence for a gene, transcript, protein, or genomic region. Returns a window of the sequence — the first 10,000 characters by default — with its stable ID, molecule type, and full length. When more follows the window, truncated is true and nextOffset is the offset to request next; walking nextOffset reconstructs the whole sequence, and max_length 0 returns everything from offset to the end. The type parameter selects which sequence is fetched: genomic (default, includes introns), cdna (spliced transcript), cds (coding sequence only), protein. For region mode, set id to a region — either species:chr:start-end (e.g. homo_sapiens:13:32315086-32400268) or a bare chr:start-end with species set (e.g. id 13:32315086-32400268, species homo_sapiens), spanning at most 10,000,000 bases; regions return genomic DNA only. Protein sequences require a transcript or protein stable ID (ENST…/ENSP…), not a gene ID — use ensembl_lookup_gene with expand_transcripts=true to get the canonical transcript ID first.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
idYesEnsembl stable ID (ENSG…, ENST…, ENSP…) or a genomic region for region mode. Region accepts species:chr:start-end (e.g. homo_sapiens:13:32315086-32400268) or a bare chr:start-end (e.g. 13:32315086-32400268) when the species field is set. A region needs start at or below end, within the sequence region, and spans at most 10,000,000 bases.
typeNoSequence type to retrieve. genomic: full genomic DNA including introns (default). cdna: spliced transcript sequence (requires ENST… ID). cds: coding sequence only, no UTRs (requires ENST… ID with coding transcript). protein: amino acid sequence (requires ENST… or ENSP… ID). Region ids are genomic-only — request cdna, cds, or protein from a transcript or protein stable ID.genomic
offsetNo0-based character offset where the returned window starts, counted in the resolved sequence (including any expand_5prime/expand_3prime flank). Default 0. Pass nextOffset from a truncated response to fetch the following window; an offset at or past the end returns an empty window.
speciesNoSpecies in Ensembl internal format (e.g. homo_sapiens). Required for a bare chr:start-end region; optional for the species:chr:start-end form (the embedded species is used when the field is omitted). Optional for stable ID lookups — Ensembl infers species from the ID prefix.
max_lengthNoMaximum number of characters in the returned window. Default 10000. Set to 0 to return everything from offset to the end, uncapped.
expand_3primeNoNumber of base pairs to extend downstream (3' direction) of the requested feature. Default 0. Only applies to genomic sequences and region queries.
expand_5primeNoNumber of base pairs to extend upstream (5' direction) of the requested feature. Default 0. Only applies to genomic sequences and region queries.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
idNoThe stable ID or region used for the lookup.
seqNoThe requested window of the sequence: at most max_length characters starting at offset. DNA sequences use IUPAC nucleotide codes (ACGT + ambiguity codes); protein sequences use single-letter amino acid codes. Empty when offset is at or past the end.
typeNoSequence type returned (genomic, cdna, cds, or protein).
errorNoPresent when the call failed. Absent on success.
lengthNoFull sequence length in characters, not the window size — nucleotides for genomic/cdna/cds, amino-acid residues for protein. Includes any expand_5prime/expand_3prime flank.
noticeNoGuidance about the window: how to continue when truncated, or why it is empty when the offset is past the end.
offsetNo0-based character offset where this window starts.
truncatedNoTrue when more sequence follows this window; request nextOffset to continue.
nextOffsetNoOffset of the first character after this window — pass it as offset to fetch the next window. Present only when truncated.
descriptionNoSequence description from Ensembl, if provided.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed14 schema fields changed
    • changedInput schema / properties / id / description
      Previous value: -"Ensembl stable ID (ENSG…, ENST…, ENSP…) or a genomic region for region mode. Region accepts species:chr:start-end (e.g. homo_sapiens:13:32315086-32400268) or a bare chr:start-end (e.g. 13:32315086-32400268) when the species field is set."New value: +"Ensembl stable ID (ENSG…, ENST…, ENSP…) or a genomic region for region mode. Region accepts species:chr:start-end (e.g. homo_sapiens:13:32315086-32400268) or a bare chr:start-end (e.g. 13:32315086-32400268) when the species field is set. A region needs start at or below end, within the sequence region, and spans at most 10,000,000 bases."
    • addedInput schema / properties / id / minLength
      Added value: +1
    • addedInput schema / properties / max_length
      Added value: +{
      +  "default": 10000,
      +  "description": "Maximum number of characters in the returned window. Default 10000. Set to 0 to return everything from offset to the end, uncapped.",
      +  "maximum": 9007199254740991,
      +  "minimum": 0,
      +  "type": "integer"
      +}
    • addedInput schema / properties / offset
      Added value: +{
      +  "default": 0,
      +  "description": "0-based character offset where the returned window starts, counted in the resolved sequence (including any expand_5prime/expand_3prime flank). Default 0. Pass nextOffset from a truncated response to fetch the following window; an offset at or past the end returns an empty window.",
      +  "maximum": 9007199254740991,
      +  "minimum": 0,
      +  "type": "integer"
      +}
    • changedInput schema / properties / type / description
      Previous value: -"Sequence type to retrieve. genomic: full genomic DNA including introns (default). cdna: spliced transcript sequence (requires ENST… ID). cds: coding sequence only, no UTRs (requires ENST… ID with coding transcript). protein: amino acid sequence (requires ENST… or ENSP… ID)."New value: +"Sequence type to retrieve. genomic: full genomic DNA including introns (default). cdna: spliced transcript sequence (requires ENST… ID). cds: coding sequence only, no UTRs (requires ENST… ID with coding transcript). protein: amino acid sequence (requires ENST… or ENSP… ID). Region ids are genomic-only — request cdna, cds, or protein from a transcript or protein stable ID."
    • changedOutput schema / anyOf
      Previous value: -[
      -  {
      -    "not": {
      -      "required": [
      -        "error"
      -      ]
      -    },
      -    "required": [
      -      "id",
      -      "type",
      -      "seq",
      -      "length"
      -    ]
      -  },
      -  {
      -    "required": [
      -      "error"
      -    ]
      -  }
      -]New value: +[
      +  {
      +    "not": {
      +      "required": [
      +        "error"
      +      ]
      +    },
      +    "required": [
      +      "id",
      +      "type",
      +      "seq",
      +      "length",
      +      "offset",
      +      "truncated"
      +    ]
      +  },
      +  {
      +    "required": [
      +      "error"
      +    ]
      +  }
      +]
    • changedOutput schema / properties / error / properties / data / properties / reason / description
      Previous value: -"Machine-readable failure mode. Declared by this tool: `not_found`: The stable ID or region was not found in Ensembl. `type_mismatch`: The requested sequence type is incompatible with the provided ID type. `missing_species`: A bare chr:start-end region was given without a species. Other values are possible when a failure originates below the handler."New value: +"Machine-readable failure mode. Declared by this tool: `not_found`: The stable ID or region was not found in Ensembl. `type_mismatch`: A non-genomic type (cdna, cds, or protein) was requested for a region id or a gene ID. `missing_species`: A bare chr:start-end region was given without a species. `invalid_region`: A region id has its start after its end, starts past the end of its sequence region, spans more than the 10,000,000-base maximum, or names a sequence region the species lacks. Other values are possible when a failure originates below the handler."
    • changedOutput schema / properties / error / properties / data / properties / reason / examples
      Previous value: -[
      -  "not_found",
      -  "type_mismatch",
      -  "missing_species"
      -]New value: +[
      +  "not_found",
      +  "type_mismatch",
      +  "missing_species",
      +  "invalid_region"
      +]
    • changedOutput schema / properties / length / description
      Previous value: -"Sequence length in characters — nucleotides for genomic/cdna/cds, amino-acid residues for protein. Use this to budget context window usage before processing the sequence."New value: +"Full sequence length in characters, not the window size — nucleotides for genomic/cdna/cds, amino-acid residues for protein. Includes any expand_5prime/expand_3prime flank."
    • addedOutput schema / properties / nextOffset
      Added value: +{
      +  "description": "Offset of the first character after this window — pass it as offset to fetch the next window. Present only when truncated.",
      +  "type": "number"
      +}
    • addedOutput schema / properties / notice
      Added value: +{
      +  "description": "Guidance about the window: how to continue when truncated, or why it is empty when the offset is past the end.",
      +  "type": "string"
      +}
    • addedOutput schema / properties / offset
      Added value: +{
      +  "description": "0-based character offset where this window starts.",
      +  "type": "number"
      +}
    • changedOutput schema / properties / seq / description
      Previous value: -"The full sequence. DNA sequences use IUPAC nucleotide codes (ACGT + ambiguity codes). Protein sequences use single-letter amino acid codes. Large genomic sequences (e.g. 85 kb for BRCA2) are returned in full."New value: +"The requested window of the sequence: at most max_length characters starting at offset. DNA sequences use IUPAC nucleotide codes (ACGT + ambiguity codes); protein sequences use single-letter amino acid codes. Empty when offset is at or past the end."
    • addedOutput schema / properties / truncated
      Added value: +{
      +  "description": "True when more sequence follows this window; request nextOffset to continue.",
      +  "type": "boolean"
      +}
  2. Changed6 schema fields changed
    • changedInput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
    • addedInput schema / additionalProperties
      Added value: +false
    • changedOutput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
    • addedOutput schema / anyOf
      Added value: +[
      +  {
      +    "not": {
      +      "required": [
      +        "error"
      +      ]
      +    },
      +    "required": [
      +      "id",
      +      "type",
      +      "seq",
      +      "length"
      +    ]
      +  },
      +  {
      +    "required": [
      +      "error"
      +    ]
      +  }
      +]
    • addedOutput schema / properties / error
      Added value: +{
      +  "additionalProperties": {},
      +  "description": "Present when the call failed. Absent on success.",
      +  "properties": {
      +    "code": {
      +      "description": "JSON-RPC error code for this failure.",
      +      "maximum": 9007199254740991,
      +      "minimum": -9007199254740991,
      +      "type": "integer"
      +    },
      +    "data": {
      +      "additionalProperties": {},
      +      "properties": {
      +        "reason": {
      +          "description": "Machine-readable failure mode. Declared by this tool: `not_found`: The stable ID or region was not found in Ensembl. `type_mismatch`: The requested sequence type is incompatible with the provided ID type. `missing_species`: A bare chr:start-end region was given without a species. Other values are possible when a failure originates below the handler.",
      +          "examples": [
      +            "not_found",
      +            "type_mismatch",
      +            "missing_species"
      +          ],
      +          "type": "string"
      +        },
      +        "recovery": {
      +          "additionalProperties": {},
      +          "description": "Actionable next step for the caller.",
      +          "properties": {
      +            "hint": {
      +              "type": "string"
      +            }
      +          },
      +          "required": [
      +            "hint"
      +          ],
      +          "type": "object"
      +        },
      +        "retryable": {
      +          "description": "Whether retrying may succeed.",
      +          "type": "boolean"
      +        }
      +      },
      +      "type": "object"
      +    },
      +    "message": {
      +      "description": "Human-readable description of what went wrong.",
      +      "type": "string"
      +    }
      +  },
      +  "required": [
      +    "code",
      +    "message"
      +  ],
      +  "type": "object"
      +}
    • removedOutput schema / required
      Removed value: -[
      -  "id",
      -  "type",
      -  "seq",
      -  "length"
      -]
  3. Changed2 schema fields changed
    • changedInput schema / properties / id / description
      Previous value: -"Ensembl stable ID (ENSG…, ENST…, ENSP…) or region in the format species:chr:start-end (e.g. homo_sapiens:13:32315086-32400268) for region mode. For genomic region queries, species is also required."New value: +"Ensembl stable ID (ENSG…, ENST…, ENSP…) or a genomic region for region mode. Region accepts species:chr:start-end (e.g. homo_sapiens:13:32315086-32400268) or a bare chr:start-end (e.g. 13:32315086-32400268) when the species field is set."
    • changedInput schema / properties / species / description
      Previous value: -"Species in Ensembl internal format (e.g. homo_sapiens). Required for region mode (when id is a species:chr:start-end string). Optional for stable ID lookups — Ensembl infers species from the ID prefix."New value: +"Species in Ensembl internal format (e.g. homo_sapiens). Required for a bare chr:start-end region; optional for the species:chr:start-end form (the embedded species is used when the field is omitted). Optional for stable ID lookups — Ensembl infers species from the ID prefix."
  4. Changed3 schema fields changed
    • addedOutput schema / properties / length
      Added value: +{
      +  "description": "Sequence length in characters — nucleotides for genomic/cdna/cds, amino-acid residues for protein. Use this to budget context window usage before processing the sequence.",
      +  "type": "number"
      +}
    • removedOutput schema / properties / lengthInBp
      Removed value: -{
      -  "description": "Sequence length in characters (nucleotides or amino acids). Use this to budget context window usage before processing the sequence.",
      -  "type": "number"
      -}
    • changedOutput schema / required
      Previous value: -[
      -  "id",
      -  "type",
      -  "seq",
      -  "lengthInBp"
      -]New value: +[
      +  "id",
      +  "type",
      +  "seq",
      +  "length"
      +]
  5. First observed

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint, idempotentHint), the description discloses truncation and pagination behavior: returns a window of 10,000 characters by default, truncated flag, nextOffset for walking, and max_length 0 returning everything to the end. It also states region mode limits (at most 10,000,000 bases) and that regions return genomic DNA only. This is rich behavioral context that the annotations do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is organized into clear thematic sections: purpose, window/pagination behavior, type selection, and region mode constraints. Every sentence carries operational information, with examples embedded inline. It is long but not verbose; no filler or redundant restatement of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with 100% schema coverage and an output schema, the description covers all necessary operational aspects: sequence types, identifier requirements, region syntax, pagination, and a prerequisite workflow via ensembl_lookup_gene. It is complete enough for an agent to call the tool correctly without additional information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description adds value beyond the schema by explaining the interaction between type, id, and region mode, giving concrete examples (homo_sapiens:13:32315086-32400268), and clarifying the walking mechanism with nextOffset. It does not add much on expand parameters, but the overall contribution is meaningful.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair: 'Fetch the DNA, cDNA, CDS, or protein sequence for a gene, transcript, protein, or genomic region.' It clearly enumerates the sequence types and accepted identifier forms, and the focus on sequence retrieval distinguishes it from sibling tools like ensembl_get_homology or ensembl_predict_variant.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance, including region mode syntax and a crucial alternative: 'Protein sequences require a transcript or protein stable ID (ENST…/ENSP…), not a gene ID — use ensembl_lookup_gene with expand_transcripts=true to get the canonical transcript ID first.' This names the sibling tool and the precise condition under which it should be used instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.