Skip to main content
Glama
cyanheads

ensembl-mcp-server

by cyanheads

Version License Docker MCP SDK npm TypeScript Bun

Install in Claude Desktop Install in Cursor Install in VS Code

Framework

Public Hosted Server: https://ensembl.caseyjhand.com/mcp


Overview

Gene, sequence, and variant data for vertebrates and other model organisms from the Ensembl REST API. Look up genes, fetch sequences, predict variant consequences, find orthologs, and cross-reference external databases from any MCP client. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.

Tools

Tool

Description

ensembl_list_species

List species supported by Ensembl with display name, common name, assembly, taxon ID, and division

ensembl_lookup_gene

Resolve a gene by symbol + species or by stable ID to its Ensembl ID, genomic location, biotype, and transcript list

ensembl_get_sequence

Fetch the DNA, cDNA, CDS, or protein sequence for a gene, transcript, protein, or genomic region

ensembl_query_region

Find genomic features (genes, transcripts, variants, regulatory elements, exons) overlapping a chromosomal region

ensembl_predict_variant

Predict functional consequences of a sequence variant using the Ensembl Variant Effect Predictor (VEP)

ensembl_get_homology

Find orthologs and/or paralogs of a gene across species with percent identity and taxonomy level

ensembl_get_xrefs

Retrieve cross-database references for a gene — HGNC, UniProt, EntrezGene, OMIM, RefSeq, Reactome, and others

Resources

Resource

Description

ensembl://gene/{id}

Gene record by stable ID (ENSG…) — location, biotype, description, and transcript list

ensembl://transcript/{id}

Transcript record by stable ID (ENST…) — parent gene, location, biotype, canonical flag, and length

ensembl://species

Supported Ensembl species for the endpoint default division (vertebrates on the default endpoint)

ensembl://species/{division}

Supported species in one division (EnsemblVertebrates, EnsemblPlants, EnsemblFungi, EnsemblMetazoa, EnsemblProtists)

All resource data is also reachable via the ensembl_list_species tool, which additionally filters by name.

Prompts

Prompt

Description

ensembl_gene_dossier

Structured workflow for assembling a complete gene profile: symbol → ID + location → sequence → variants → orthologs → xrefs

Related MCP server: Ensembl MCP Server

Capability reference

ensembl_list_species tool

  • Filter by division (EnsemblVertebrates, EnsemblPlants, EnsemblFungi, EnsemblMetazoa, EnsemblProtists) or nameContains for a local substring match against name, display name, and common name

  • Omit division to return the endpoint default division (vertebrates, ~356 species on the default GRCh38 endpoint)

  • Returns internal name (the value every other tool expects), display name, common name, taxon ID, assembly, and division

  • Required first step — species names like homo_sapiens are opaque to non-biologists


ensembl_lookup_gene tool

  • Exactly one of symbol (+ optional species, default homo_sapiens), id, ids (batch, up to 20), or symbols (batch, up to 20)

  • expand_transcripts (default false) adds the full transcript list with biotype and canonical flag

  • Batch modes (ids/symbols) return a succeeded/failed split with per-item error strings instead of failing the call

  • Errors: not_found, invalid_species, no_input, conflicting_input


ensembl_get_sequence tool

  • type: genomic (default, includes introns), cdna (spliced), cds (coding only), protein

  • Accepts a stable ID (ENSG…/ENST…/ENSP…) or a region — species:chr:start-end, or bare chr:start-end with species set; a region spans at most 10,000,000 bases, with start at or below end

  • expand_5prime / expand_3prime (default 0) extend flanking base pairs for genomic and region queries

  • protein and cds require a transcript or protein ID, not a gene ID; region ids are genomic-only

  • Returns a bounded window: offset (0-based, default 0) and max_length (default 10000, 0 for the rest uncapped) index the resolved sequence, flanks included

  • length is always the full sequence length; truncated and nextOffset say whether more follows and where to resume, so walking nextOffset reconstructs the whole sequence

  • Errors: not_found, type_mismatch, missing_species, invalid_region


ensembl_query_region tool

  • region in chr:start-end format, at most 5,000,000 bases; feature array (at least one) defaults to ["gene"], also accepts transcript, variation, regulatory, exon; optional biotype filter

  • Defaults to genes only — requesting variation on a large locus can match 44,000+ features

  • max_results caps the feature list (default 100, 0 uncapped); totalCount always reports the true count found

  • assemblyName (e.g. GRCh38) names the assembly the coordinates are on

  • Exon rows carry a parentId and rank, since one exon is reported once per parent transcript

  • Errors: invalid_region, invalid_species


ensembl_predict_variant tool

  • variant accepts HGVS (transcript-relative or genomic), region+allele (chr:start:end:strand/allele), or a dbSNP rsID

  • max_transcript_consequences (default 10) and max_pubmed_ids_per_variant (default 10) cap large VEP results; set either to 0 for the full set, or include_all_colocated_pubmed: true for uncapped PubMed IDs

  • Returns most severe consequence term, per-transcript impact (HIGH/MODERATE/LOW/MODIFIER), and colocated known variants with clinical significance

  • Totals (transcriptConsequencesTotal, pubmedTotal) are always reported even when capped

  • Errors: invalid_notation, not_found


ensembl_get_homology tool

  • Exactly one of symbol (+ species, default homo_sapiens) or id; optional target_species filter

  • type: orthologues (default), paralogues, or all

  • max_results caps the homolog list (default 25, 0 uncapped); totalCount always reports the true count available

  • Errors: not_found, no_input, conflicting_input


ensembl_get_xrefs tool

  • id (ENSG…/ENST…) required; optional dbname filter (e.g. HGNC, Uniprot_gn, EntrezGene, MIM_GENE, RefSeq_mRNA, Reactome, GO)

  • Uses the xrefs/id endpoint, returning the full cross-reference set (56+ entries for well-annotated genes like BRCA2)

  • Errors: not_found


ensembl://gene/{id} resource

  • Returns location, biotype, description, and transcript list for a gene stable ID (ENSG…); version suffix optional

  • Errors: not_found


ensembl://transcript/{id} resource

  • Returns parent gene, location, biotype, canonical flag, and length for a transcript stable ID (ENST…); version suffix optional

  • Errors: not_found


ensembl://species resource

  • No parameters — returns the endpoint default division (vertebrates, ~356 species on the default GRCh38 endpoint)

  • For a named division, read ensembl://species/{division} instead


ensembl://species/{division} resource

  • division required: EnsemblVertebrates, EnsemblPlants, EnsemblFungi, EnsemblMetazoa, or EnsemblProtists


ensembl_gene_dossier prompt

  • Arguments: gene_symbol required; species optional (default homo_sapiens)

  • Sequences a 7-step workflow: resolve the gene → fetch the protein sequence → find variants in the locus → predict variant consequences → find cross-species orthologs → get external database IDs → synthesize the dossier

Features

Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.

Ensembl-specific:

  • Keyless REST API — no API key required; Ensembl REST is fully public at 55,000 req/hr

  • Rate-limit-aware service layer: retries 429 honoring Retry-After, and retries transient 5xx and HTML error pages

  • Batch POST endpoints used throughout — POST /lookup/id and POST /lookup/symbol/{species} (up to 1,000 items each upstream) reduce N+1 round trips in multi-gene workflows

  • GRCh37 legacy support via ENSEMBL_BASE_URL — point the entire server at https://grch37.rest.ensembl.org for clinical workflows on the older assembly

  • All coordinate-bearing responses echo the assembly name so agents never see a bare genomic position without assembly context

Agent-friendly output:

  • ensembl_get_sequence returns sequences in bounded windows (10,000 characters by default) with the full length and a nextOffset to continue, so a long gene or locus never lands in one response unasked

  • ensembl_list_species is explicitly the discovery step — tool descriptions call out the opaque internal-name format and direct agents to it before using species-dependent tools

  • Cross-tool chaining made explicit: xref IDs from ensembl_get_xrefs are described as inputs for protein and literature servers; the ensembl_gene_dossier prompt sequences all 6 tools into one research workflow

Getting started

Public Hosted Instance

A public instance is available at https://ensembl.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:

{
  "mcpServers": {
    "ensembl-mcp-server": {
      "type": "streamable-http",
      "url": "https://ensembl.caseyjhand.com/mcp"
    }
  }
}

Self-Hosted / Local

Add the following to your MCP client configuration file.

{
  "mcpServers": {
    "ensembl-mcp-server": {
      "type": "stdio",
      "command": "bunx",
      "args": ["@cyanheads/ensembl-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info"
      }
    }
  }
}

Or with npx (no Bun required):

{
  "mcpServers": {
    "ensembl-mcp-server": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@cyanheads/ensembl-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info"
      }
    }
  }
}

Or with Docker:

{
  "mcpServers": {
    "ensembl-mcp-server": {
      "type": "stdio",
      "command": "docker",
      "args": [
        "run", "-i", "--rm",
        "-e", "MCP_TRANSPORT_TYPE=stdio",
        "ghcr.io/cyanheads/ensembl-mcp-server:latest"
      ]
    }
  }
}

For Streamable HTTP, set the transport and start the server:

MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcp

Prerequisites

  • Bun v1.4.0 or higher (or Node.js v24+).

  • No API key required — Ensembl REST is fully public.

Installation

  1. Clone the repository:

git clone https://github.com/cyanheads/ensembl-mcp-server.git
  1. Navigate into the directory:

cd ensembl-mcp-server
  1. Install dependencies:

bun install
  1. Configure environment:

cp .env.example .env
# edit .env if you need to override ENSEMBL_BASE_URL (e.g. for GRCh37)

Configuration

All configuration is validated at startup via Zod schemas in src/config/server-config.ts.

Variable

Description

Default

ENSEMBL_BASE_URL

Ensembl REST API base URL. Override for GRCh37 (https://grch37.rest.ensembl.org) or a local mirror.

https://rest.ensembl.org

MCP_TRANSPORT_TYPE

Transport: stdio or http

stdio

MCP_HTTP_PORT

HTTP server port

3010

MCP_HTTP_ENDPOINT_PATH

HTTP endpoint path

/mcp

MCP_SESSION_MODE

HTTP session mode: auto, stateful, or stateless. Schema default auto resolves to stateful; this server explicitly uses stateless.

stateless

MCP_AUTH_MODE

Authentication: none, jwt, or oauth

none

MCP_LOG_LEVEL

Log level (debug, info, warning, error, etc.)

info

LOGS_DIR

Directory for log files (Node.js only)

<project-root>/logs

OTEL_ENABLED

Enable OpenTelemetry

false

See .env.example for the full list of optional overrides.

Running the server

Local development

  • Build and run:

    # One-time build
    bun run rebuild
    
    # Run the built server
    bun run start:stdio
    # or
    bun run start:http
  • Run checks and tests:

    bun run devcheck   # Lint, format, typecheck, security
    bun run test       # Vitest test suite
    bun run lint:mcp   # Validate MCP definitions against spec

Docker

docker build -t ensembl-mcp-server .
docker run --rm -p 3010:3010 ensembl-mcp-server

The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/ensembl-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.

Project structure

Directory

Purpose

src/index.ts

createApp() entry point — registers tools/resources/prompts and inits services

src/config

Server-specific environment variable parsing and validation with Zod

src/mcp-server/tools

Tool definitions (*.tool.ts) — 7 tools

src/mcp-server/resources

Resource definitions (*.resource.ts) — gene, transcript, species

src/mcp-server/prompts

Prompt definitions (*.prompt.ts) — gene dossier workflow

src/services/ensembl

Ensembl REST API client — HTTP, rate-limit handling, retry, error normalization

tests/

Unit and integration tests mirroring src/

Development guide

See CLAUDE.md for development guidelines and architectural rules. The short version:

  • Handlers throw, framework catches — no try/catch in tool logic

  • Use ctx.log for request-scoped logging, ctx.state for tenant-scoped storage

  • Register new tools and resources in the createApp() arrays in src/index.ts

  • Wrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields

Contributing

Issues are welcome. Run checks and tests before submitting:

bun run devcheck
bun run test

License

Apache-2.0 — see LICENSE for details.

Related MCP Connectors

Related MCP Servers