Skip to main content
Glama
cyanheads

@cyanheads/gnomad-genetics-mcp-server

by cyanheads

Version License Docker MCP SDK npm TypeScript Bun

Install in Claude Desktop Install in Cursor Install in VS Code

Framework

Public Hosted Server: https://gnomad-genetics.caseyjhand.com/mcp


Overview

Population genetics over gnomAD (Broad Institute), with ClinVar clinical significance joined in from NCBI. Look up per-ancestry allele frequencies, gene loss-of-function constraint, gene variant catalogs, and sequencing coverage, then query large result sets with SQL from any MCP client. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.

Tools

Tool

Description

gnomad_get_variant

Full population record for one or more variants — AC/AN/AF overall and per genetic-ancestry group, homozygote/hemizygote counts, quality flags, transcript consequence, in-silico predictors, and joined ClinVar significance.

gnomad_get_gene_constraint

Gene loss-of-function constraint — pLI, LOEUF (oe_lof_upper) with confidence interval, observed/expected ratios, and Z-scores. By HGNC symbol or Ensembl gene ID.

gnomad_list_gene_variants

Every variant in a gene, transcript, or region with allele frequencies and predicted consequences, filterable by consequence class and max allele frequency.

gnomad_get_coverage

Sequencing coverage across a gene, transcript, or region — mean/median depth and the fraction of samples over depth thresholds, per callset track.

gnomad_search_clinvar

Gene-level ClinVar detail via NCBI E-utilities — classified variants, review status, conditions, submission counts, and gnomAD-compatible variant IDs, paged by offset.

gnomad_dataframe_query

Run a read-only SQL SELECT across canvas tables staged by the list tools.

gnomad_dataframe_describe

List the tables staged on a canvas and their columns before writing SQL.

gnomad_dataframe_drop

Drop a named table from a canvas to reclaim memory. Opt-in via GNOMAD_DATAFRAME_DROP_ENABLED=true — off by default.

Resources

Resource

Description

gnomad://variant/{dataset}/{variantId}

Population record for one variant — mirrors gnomad_get_variant.

gnomad://gene/{dataset}/{gene}/constraint

Gene loss-of-function constraint — mirrors gnomad_get_gene_constraint.

All resource data is also reachable via tools. The list tools (gnomad_list_gene_variants, gnomad_get_coverage, gnomad_search_clinvar) return analytical row sets rather than stable single-URI documents, so they are not exposed as resources — call the tools instead.

Prompts

Prompt

Description

gnomad_variant_triage

Guided rare-disease variant-triage workflow: population frequency → gene constraint → callability check, in order.

Related MCP server: EVEE MCP Server

Capability reference

gnomad_get_variant tool

  • Batch up to 25 IDs per call (default; raise via GNOMAD_MAX_VARIANT_BATCH), each a chrom-pos-ref-alt variantId on chromosome 1–22, X, or Y with an optional chr prefix (e.g. 1-55051215-G-GA) or an rsID (e.g. rs11591147)

  • Per-item partial success — a malformed or absent ID lands in failed[] without failing the others

  • Each failed[] item is { variant, error, reason, recovery }: reason is typed (invalid_variant_id, variant_not_found, upstream_unavailable, …) and recovery is the hint the tool declares for it; an ambiguous rsID adds candidates

  • Mitochondrial IDs (M, MT, chrM) are refused per item as mitochondrial_unsupported with no fetch — gnomAD models mitochondrial variants separately, and they are outside this server

  • Per-ancestry frequency vector is returned in full, never collapsed to a single global AF

  • Reports which callset(s) (exome / genome) carry the variant, quality flags, transcript consequence, in-silico predictor scores, and the ClinVar significance gnomAD joins per variant

  • An empty found[] for a well-formed ID means the variant is not in the chosen dataset — pair with gnomad_get_coverage to confirm the position is callable before concluding true absence


gnomad_get_gene_constraint tool

  • Accepts an HGNC symbol (PCSK9) or an Ensembl gene ID (ENSG00000169174)

  • Returns pLI (>0.9 intolerant), LOEUF / oe_lof_upper with its lower bound, observed/expected ratios for LoF / missense / synonymous, and the three Z-scores

  • constraint_release names the release the metrics come from:

    dataset

    Source

    constraint_release

    LoF-intolerance guidance

    gnomad_r4

    GRCh38 gnomAD constraint

    gnomAD v4.1.2

    LOEUF < 0.45

    gnomad_r3

    GRCh38 gnomAD constraint (gnomAD publishes no v3 constraint)

    gnomAD v4.1.2

    LOEUF < 0.45

    gnomad_r2_1

    GRCh37 gnomAD constraint

    gnomAD v2.1.1

    LOEUF < 0.35

    exac

    GRCh37 ExAC constraint

    ExAC r0.3

    pLI only

  • ExAC publishes pLI, the Z-scores, and observed/expected counts only, so on exac the ratios and LOEUF are null and constraint_flags is empty

  • Many genes have null constraint (sparse upstream) — null fields are reported as such, never fabricated

  • constraint_flags carries the caveat flags gnomAD attaches to a gene's constraint (e.g. no_exp_lof, syn_outlier)


gnomad_list_gene_variants tool

  • Supply exactly one of gene, transcript_id, or region (chrom-start-stop, 1-based inclusive, chromosome 1–22, X, or Y with an optional chr prefix)

  • A region must span less than 2,500,000 bp (stop − start) and hold at most ~30,000 variants; an unserved chromosome or out-of-range coordinate fails as invalid_region and an over-wide span as region_too_large, both before any fetch, while a region over the variant ceiling fails as region_too_large after one request, carrying gnomAD's own message

  • Mitochondrial targets (an M/MT region, or a gene or transcript gnomAD places on chromosome M) fail as mitochondrial_unsupported instead of returning an empty list

  • Optional filters: one consequence_class (lof / missense / synonymous / other) and/or a maximum allele frequency

  • A result too large to inline (the preview holds about 14,000 characters of rows, keeping a response near 24 KB) is staged on a DataCanvas table named gene_variants, returned as canvas_id and table_name beside the preview — inspect it with gnomad_dataframe_describe, then query it with gnomad_dataframe_query to rank by AF, count by consequence, or group across every row. The response notice names the table and both tools

  • A result that fits inline stages no table and uses no canvas (canvas_id is empty) unless you pass a canvas_id

  • Passing a canvas_id always writes the result to gene_variants on that canvas, REPLACING the previous table (it never appends), even when the result fits inline; a result with no variants removes the table

  • When the canvas is disabled (CANVAS_PROVIDER_TYPE != duckdb) the tool returns the same capped inline preview (as many rows as fit about 14,000 characters) and the SQL path is unavailable

  • A blank gene or transcript_id counts as omitted


gnomad_get_coverage tool

  • Supply exactly one of gene, transcript_id, or region; a blank gene or transcript_id counts as omitted

  • region takes the same chrom-start-stop form as gnomad_list_gene_variants — chromosome 1–22, X, or Y, optional chr prefix, a span under 2,500,000 bp — and fails with invalid_region or region_too_large before any fetch otherwise

  • Mitochondrial targets fail as mitochondrial_unsupported instead of returning empty coverage

  • Returns mean and median read depth plus the mean fraction of samples covered at each threshold (1× through 100×), summarized per callset track

  • coverage_source narrows to one track (exome / genome); omit to return every available track

  • A variant missing from a well-covered region is informative; one missing from a poorly-covered region is not


gnomad_search_clinvar tool

  • Returns a gene's classified ClinVar variants — clinical significance, review status with a 0–4 star rating, associated conditions, molecular consequences, and submission counts

  • Each row carries gnomAD-compatible identifiers: canonical_spdi, rsids, and grch38_variant_id (chrom-pos-ref-alt, set for SNVs, MNVs, and delins), which gnomad_get_variant resolves in the GRCh38 datasets

  • Optional filters: clinical_significance (e.g. pathogenic; blank means no filter) and a minimum star rating (min_review_stars, 0–4)

  • Returns one window of up to 500 ClinVar records per call: total_found is ClinVar's candidate count for the gene and filter terms (taken before the significance and star filters narrow each window), truncated and next_offset say whether more remain, and offset / limit (1–500) page through them. limit counts records before the filters, so a window can return fewer rows

  • VariationIDs ClinVar returns no summary for are listed in unavailable_ids rather than returned as blank rows

  • Accepts an HGNC symbol only — ClinVar's gene index doesn't resolve Ensembl gene IDs, unlike the other gnomAD tools. An Ensembl gene ID returns guidance to resolve its symbol and searches nothing, leaving any canvas_id you pass untouched

  • A window too large to inline (the preview holds about 11,000 characters of rows, keeping a response near 24 KB) is staged on the clinvar_variants canvas table — inspect it with gnomad_dataframe_describe, then query it with gnomad_dataframe_query. A window that fits inline stages no table unless you pass a canvas_id; passing one always writes the window to clinvar_variants, REPLACING the previous table, and a window with no rows removes it. With the canvas disabled, the preview is the same capped preview (as many rows as fit about 11,000 characters)

  • Keyless, but honors NCBI_API_KEY for a higher rate limit (10 vs 3 req/s)


gnomad_dataframe_query tool

  • Runs single-statement, read-only SQL SELECTs against a canvas table staged by gnomad_list_gene_variants or gnomad_search_clinvar — writes, DDL, and file/HTTP table functions are rejected by the canvas gate

  • Reference tables by the name the staging tool returned (gene_variants or clinvar_variants)

  • Returns one page of the result: offset (default 0) and limit (default 100, max 500) select it, and a page also ends before its rows pass 10,000 characters of JSON, which keeps a full page under about 24 KB

  • Output: rows (dynamic columns per the SQL projection), columns, offset, returned, total (exact row count; null when the result exceeds the canvas row cap), truncated (rows exist after this page), and next_offset (null on the last page) — follow next_offset until it is null

  • Each page re-runs the SQL: stable paging needs an ORDER BY over a unique key (such as variant_id) and an unchanged table. Paging stops at the canvas row cap (CANVAS_DEFAULT_ROW_LIMIT, 10,000 by default); filter or aggregate in SQL to reach rows past it

  • A row over the 10,000-character budget fails with row_too_large — select fewer or narrower columns

  • Requires CANVAS_PROVIDER_TYPE=duckdb — otherwise fails with a canvas_disabled error


gnomad_dataframe_describe tool

  • Lists every table staged on a canvas with its row count and column schema (name and DuckDB type)

  • Call it before writing SQL for gnomad_dataframe_query

  • Requires CANVAS_PROVIDER_TYPE=duckdb — otherwise fails with a canvas_disabled error


gnomad_dataframe_drop tool

  • Drops a named table from a canvas to reclaim memory — a deliberate mutation (readOnlyHint: false, destructiveHint: true) on an otherwise read-only surface

  • Opt-in via GNOMAD_DATAFRAME_DROP_ENABLED=true; absent from tools/list when off, since per-table TTL already reclaims memory automatically

  • Requires CANVAS_PROVIDER_TYPE=duckdb — otherwise fails with a canvas_disabled error


gnomad://variant/{dataset}/{variantId} resource

  • Population record for one variant as application/json — mirrors gnomad_get_variant; the dataset segment keeps the URI self-describing

  • variantId accepts a chrom-pos-ref-alt ID or an rsID, same grammar as the tool

  • Typed errors: invalid_variant_id (outside the coordinate/rsID grammar), mitochondrial_unsupported (an M/MT/chrM ID), variant_not_found, ambiguous_rsid (with candidates), and the gnomAD failures graphql_error, upstream_build_mismatch, upstream_unavailable, upstream_timeout, upstream_access, and invalid_upstream_response — each with its declared recovery hint


gnomad://gene/{dataset}/{gene}/constraint resource

  • Gene loss-of-function constraint as application/json — mirrors gnomad_get_gene_constraint, including the same constraint_release for each dataset segment (exac serves ExAC r0.3 constraint; gnomad_r3 serves the GRCh38 gnomAD v4.1.2 table)

  • gene accepts an HGNC symbol or Ensembl gene ID

  • Typed errors: gene_not_found when no gene matches in the requested build, invalid_constraint_data when gnomAD's metrics fall outside their valid ranges, and the gnomAD failures graphql_error, upstream_unavailable, upstream_timeout, upstream_access, and invalid_upstream_response — each with its declared recovery hint


gnomad_variant_triage prompt

  • Arguments: variant required (chrom-pos-ref-alt or rsID); gene and dataset optional. dataset is one of gnomad_r4, gnomad_r3, gnomad_r2_1, exac — any other value is rejected; omitted or blank, the emitted calls carry no dataset and use the server default. A blank gene counts as omitted

  • Emits a three-step chain as one user message: population frequency (gnomad_get_variant) → gene constraint (gnomad_get_gene_constraint) → callability check (gnomad_get_coverage on the exact position, not gene-level)

  • Rejects a malformed variant with a validation error before generating the chain

Features

Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.

gnomAD-specific:

  • Single keyless GraphQL source for the entire core surface — ClinVar significance is joined per variant inside gnomAD's own response

  • dataset and reference_genome are distinct, coherence-validated parameters (v4/v3 ⇒ GRCh38, v2.1/ExAC ⇒ GRCh37); both are echoed in every tool's output so a wrong-build coordinate mismatch is visible

  • Polite client — conservative concurrency cap (GNOMAD_MAX_CONCURRENCY, default 2) and exponential backoff against a community-funded, rate-limited API

  • In-conversation SQL analytics: gnomad_list_gene_variants and gnomad_search_clinvar stage results too large to inline on a DuckDB-backed canvas table — gnomad_dataframe_describe lists its columns, gnomad_dataframe_query runs SQL over every row

Agent-friendly output:

  • Per-ancestry allele-frequency vector returned in full, never collapsed to a single global AF — the cross-ancestry contrast is the signal clinical interpretation needs

  • Graceful partial failure — gnomad_get_variant returns per-item failed[] rows, each with a typed reason and its recovery hint, instead of failing the whole batch

  • Provenance on every response — effective dataset and reference_genome echoed back; null upstream fields preserved as null, never fabricated

  • Every error carries a typed reason and the recovery hint its tool or resource declares for that reason, so callers know the next move: input problems (incoherent_build, invalid_target, invalid_variant_id, invalid_region, region_too_large, mitochondrial_unsupported, ambiguous_rsid), absences (gene_not_found, variant_not_found), gnomAD refusals (graphql_error, upstream_build_mismatch, invalid_constraint_data), upstream faults (upstream_unavailable, upstream_timeout, upstream_access, invalid_upstream_response), and the canvas (canvas_disabled, row_too_large). gnomad_get_variant puts the same reason and hint on each failed[] item

Getting started

Public Hosted Instance

A public instance is available at https://gnomad-genetics.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:

{
  "mcpServers": {
    "gnomad-genetics-mcp-server": {
      "type": "streamable-http",
      "url": "https://gnomad-genetics.caseyjhand.com/mcp"
    }
  }
}

Self-Hosted / Local

Add the following to your MCP client configuration file. gnomAD is a free, keyless API — no credentials required.

{
  "mcpServers": {
    "gnomad-genetics-mcp-server": {
      "type": "stdio",
      "command": "bunx",
      "args": ["@cyanheads/gnomad-genetics-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info"
      }
    }
  }
}

Or with npx (no Bun required):

{
  "mcpServers": {
    "gnomad-genetics-mcp-server": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@cyanheads/gnomad-genetics-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info"
      }
    }
  }
}

Or with Docker:

{
  "mcpServers": {
    "gnomad-genetics-mcp-server": {
      "type": "stdio",
      "command": "docker",
      "args": [
        "run", "-i", "--rm",
        "-e", "MCP_TRANSPORT_TYPE=stdio",
        "ghcr.io/cyanheads/gnomad-genetics-mcp-server:latest"
      ]
    }
  }
}

For Streamable HTTP, set the transport and start the server:

MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcp

To enable the SQL analytics path, also set CANVAS_PROVIDER_TYPE=duckdb — @duckdb/node-api ships as a dependency, so nothing extra to install.

Prerequisites

  • Bun v1.4.0 or higher (or Node.js v24+).

  • No API key — gnomAD's GraphQL endpoint is keyless. An optional NCBI_API_KEY raises the gnomad_search_clinvar rate limit.

Installation

  1. Clone the repository:

git clone https://github.com/cyanheads/gnomad-genetics-mcp-server.git
  1. Navigate into the directory:

cd gnomad-genetics-mcp-server
  1. Install dependencies:

bun install
  1. Configure environment:

cp .env.example .env
# all vars are optional — the server runs keyless out of the box

Configuration

All variables are optional; the server runs keyless with the defaults below.

Variable

Description

Default

GNOMAD_API_BASE_URL

gnomAD GraphQL endpoint. Override for a private mirror or testing.

https://gnomad.broadinstitute.org/api

GNOMAD_DEFAULT_DATASET

Dataset used when a tool call omits dataset (gnomad_r4 / gnomad_r3 / gnomad_r2_1 / exac).

gnomad_r4

GNOMAD_REQUEST_TIMEOUT_MS

Per-request timeout against the GraphQL endpoint, in milliseconds.

30000

GNOMAD_MAX_CONCURRENCY

Cap on concurrent upstream requests — politeness against a community-funded API.

2

GNOMAD_MAX_VARIANT_BATCH

Maximum variant IDs accepted per gnomad_get_variant call.

25

CLINVAR_BASE_URL

NCBI E-utilities base URL for gnomad_search_clinvar.

https://eutils.ncbi.nlm.nih.gov/entrez/eutils

NCBI_API_KEY

Optional NCBI key. Raises the E-utilities rate limit from 3 to 10 req/s.

—

CANVAS_PROVIDER_TYPE

Set to duckdb to enable the spill/SQL path behind the list tools. When none, they return a capped inline preview.

none

GNOMAD_DATAFRAME_DROP_ENABLED

Gate for the opt-in gnomad_dataframe_drop tool. Off by default.

false

MCP_TRANSPORT_TYPE

Transport: stdio or http.

stdio

MCP_HTTP_PORT

Port for the HTTP server.

3010

MCP_AUTH_MODE

Auth mode: none, jwt, or oauth.

none

MCP_LOG_LEVEL

Log level (RFC 5424).

info

OTEL_ENABLED

Enable OpenTelemetry instrumentation.

false

See .env.example for the full list of optional overrides.

Running the server

Local development

  • Build and run:

    # One-time build
    bun run rebuild
    
    # Run the built server
    bun run start:stdio
    # or
    bun run start:http
  • Run checks and tests:

    bun run devcheck   # Lint, format, typecheck, security
    bun run test       # Vitest test suite
    bun run lint:mcp   # Validate MCP definitions against spec

Docker

docker build -t gnomad-genetics-mcp-server .
docker run --rm -e MCP_TRANSPORT_TYPE=http -p 3010:3010 gnomad-genetics-mcp-server

The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/gnomad-genetics-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.

Project structure

Directory

Purpose

src/index.ts

createApp() entry point — registers tools/resources/prompts and inits services.

src/config

Server-specific environment variable parsing and validation with Zod.

src/mcp-server/tools

Tool definitions (*.tool.ts) and shared input schemas.

src/mcp-server/resources

Resource definitions (*.resource.ts).

src/mcp-server/prompts

Prompt definitions (*.prompt.ts).

src/services/gnomad

gnomAD GraphQL client, query documents, and domain types.

src/services/clinvar

NCBI E-utilities client for the optional ClinVar tool.

src/services/canvas-accessor.ts

Module-level accessor for the framework's optional DataCanvas.

Development guide

See CLAUDE.md/AGENTS.md for development guidelines and architectural rules. The short version:

  • Handlers throw, framework catches — no try/catch in tool logic

  • Use ctx.log for request-scoped logging, ctx.state for tenant-scoped storage

  • Register new tools and resources in the createApp() arrays in src/index.ts

  • Wrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields

Data attribution

gnomAD data is provided by the Genome Aggregation Database (Broad Institute). ClinVar data is provided by NCBI.

Contributing

Issues are welcome. Run checks and tests before submitting:

bun run devcheck
bun run test

License

Apache-2.0 — see LICENSE for details.

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    C
    maintenance
    Annotate variants by with a deep and rich set of data. Can annotate: genetic change, rsID, CAid, HGVS (g./c./p.), protein change.
    5
    -
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables querying ClinGen curated evidence for gene-disease validity, dosage, actionability, and variant pathogenicity via MCP tools.
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Federates 13 gene-related MCP backends (gnomAD, GTEx, etc.) behind a single Streamable HTTP endpoint with collision-free namespacing and search-based tool discovery.
    6
    MIT