Skip to main content
Glama
cyanheads

@cyanheads/uniprot-mcp-server

by cyanheads

Version License Docker MCP SDK npm TypeScript Bun

Install in Claude Desktop Install in Cursor Install in VS Code

Framework

Public Hosted Server: https://uniprot.caseyjhand.com/mcp


Overview

Protein research over UniProtKB (rest.uniprot.org). Search by function or gene, fetch curated records, translate identifiers across sibling databases, and pull reference proteomes, taxonomy, and sequences from any MCP client. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.

Tools

Tool

Description

uniprot_search_proteins

Search UniProtKB by plain text or a Lucene field query, with the reviewed (Swiss-Prot) filter foregrounded and optional server-side facet counts. Cursor-paginated. The discovery entry point.

uniprot_get_entry

Fetch full curated entries by accession in one batch (up to 20) — function, catalytic activity, disease, variants, isoforms, GO terms, cross-references. Partial-success output; an oversized record returns a section outline.

uniprot_map_ids

Translate identifiers across databases via UniProt's async ID-mapping service — gene names, Ensembl, RefSeq, ChEMBL, PDB, GeneID ↔ UniProtKB accessions. Polls within a budget; running jobs return a ticket and completed pages return a continuation.

uniprot_get_proteome

Fetch a reference proteome by UPID or NCBI taxon ID — protein count, BUSCO completeness, genome assembly inline, plus an opt-in capped page of the proteins.

uniprot_get_taxonomy

Resolve a taxonomy record by NCBI taxon ID or scientific name — name, rank, parent, full lineage, and optionally the immediate children.

uniprot_get_sequence

Fetch the canonical amino-acid sequence (FASTA) for an accession, with length and parsed header — and optionally the isoform sequences. The cheap sequence-only path.

Resources

Resource

Description

uniprot://entry/{accession}

A curated UniProtKB entry by accession — the resource mirror of uniprot_get_entry for a single accession.

uniprot://taxonomy/{taxonId}

A taxonomy record by NCBI taxon ID — name, rank, parent, full lineage. The mirror of uniprot_get_taxonomy by ID.

All resource data is also reachable via tools — tool-only clients lose nothing. UniProtKB is far too large to enumerate, so there is no resource list(); discovery is uniprot_search_proteins's job.

Prompts

Prompt

Description

uniprot_protein_dossier

Guided protein-research workflow — resolve an identifier, fetch the curated entry, pull disease and variants, and surface cross-references for structure, citations, and bioactivity.

Related MCP server: mcp-uniprot

Capability reference

uniprot_search_proteins tool

  • text_search for plain language (the 80% case) or query for full Lucene field syntax (gene, organism_id, keyword, go, reviewed, protein_name, family, length, existence, accession) — exactly one

  • reviewed defaults to true (Swiss-Prot only) so the agent isn't drowned in TrEMBL predictions; set false to include them

  • organism_id convenience filter ANDed onto the query

  • Optional facets for server-side count breakdowns (e.g. reviewed, model_organism)

  • Forward cursor pagination (UniProtKB has no offset paging); totalResults and the effective query echoed back

  • Every hit carries reviewed, annotationScore, and proteinExistence so curation quality is weighable


uniprot_get_entry tool

  • Batch up to 20 accessions in one round trip

  • Sectioned record: function, catalytic activity, cofactors, subcellular location, disease, PTMs, natural variants, isoforms, domains, GO terms, keywords, cross-references

  • Partial-success output — resolved entries in succeeded[], unknown/withdrawn ones in failed[]; the whole batch never aborts on one bad accession

  • fields trims the upstream projection; identity and provenance fields are always retained

  • A single oversized record returns kind: "outline" (a section listing) instead of overflowing context — re-call the same accession with sections: [...] to pull only what's needed

  • Accessions come from uniprot_search_proteins or uniprot_map_ids; strip any -N isoform suffix first


uniprot_map_ids tool

  • from_db / to_db are validated enums (e.g. Gene_Name, Ensembl, RefSeq_Protein, ChEMBL, PDB, GeneID, UniProtKB_AC-ID) so an unsupported pair fails before the upstream call

  • Target UniProtKB-Swiss-Prot for reviewed accessions only (the usual intent), or UniProtKB / UniProtKB_AC-ID to include unreviewed TrEMBL

  • The job runs asynchronously; the tool submits it and polls within a budget. A running job returns status: "running" with a ticket — pass that ticket alone to poll the same job

  • A completed call returns status: "finished" with one upstream page. If continuation is present, pass it alone to fetch the next completed page without polling or re-submitting; its absence marks the terminal page

  • Pair a gene-symbol from_db with tax_id to disambiguate species

  • unmappedIds is populated only from UniProt's failedIds, so identifiers UniProt normalizes in successful result rows are not misclassified as failures


uniprot_get_proteome tool

  • Provide exactly one of upid (e.g. UP000005640) or taxon_id — providing both or neither fails validation

  • Metadata inline: proteome type, total protein count, BUSCO completeness (score, complete/fragmented/missing counts, lineage dataset), genome assembly accession

  • The protein set is opt-in via include_proteins (it is large — human is ~147,506) and returns a capped page with a forward cursor and truncation disclosure

  • Narrow the protein list with the query filter (UniProtKB Lucene syntax) for a subset

  • Resolve an organism name to a taxon ID first with uniprot_get_taxonomy


uniprot_get_taxonomy tool

  • Provide exactly one of taxon_id (NCBI numeric ID) or name (scientific name) — both or neither fails validation

  • Returns scientific and common name, mnemonic, rank, parent, and the full lineage (root → near ancestor)

  • include_children fetches the immediate child taxa via a follow-up search (not inline on the record)

  • Turns an organism name into the taxon ID that uniprot_search_proteins (organism_id) and uniprot_get_proteome (taxon_id) expect


uniprot_get_sequence tool

  • Returns the canonical sequence with its length and parsed FASTA header

  • include_isoforms also returns the alternatively-spliced isoform sequences

  • Accessions come from uniprot_search_proteins or uniprot_map_ids; strip any -N isoform suffix first


uniprot://entry/{accession} resource

  • Same record as uniprot_get_entry for a single accession, addressed by URI instead of a tool call

  • An annotation-heavy entry over the outline budget returns a bounded identity summary plus a section outline instead of the full record — fetch specific sections via the uniprot_get_entry tool (sections: [...]); the resource URI itself takes no sections parameter

  • The overflow summary always keeps identity fields — accession, entry name, protein name, genes, organism, length, reviewed status, annotation score, protein existence


uniprot://taxonomy/{taxonId} resource

  • Addressed by NCBI taxon ID, e.g. uniprot://taxonomy/9606 for human

  • Returns scientific and common name, mnemonic, rank, parent, and full lineage — same shape as uniprot_get_taxonomy without include_children

  • Tool-only clients reach the same data via uniprot_get_taxonomy


uniprot_protein_dossier prompt

  • Arguments: identifier required (gene name, accession, or protein name); organism optional to disambiguate

  • Five-step workflow embedded in the generated message: resolve identifier → fetch entry → summarize function/localization/provenance → pull disease & variants → surface cross-references for structure (PDB), bioactivity (ChEMBL), and citations (PubMed)

  • Returns a single user-role message carrying the full instructions — no separate assistant framing message

Features

Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.

UniProt-specific:

  • Keyless — UniProt REST requires no API key; works against any rest.uniprot.org-compatible base (override UNIPROT_BASE_URL for a private mirror)

  • One thin fetch client over all four REST collections (UniProtKB, ID Mapping, Proteomes, Taxonomy) with retry/backoff and HTML-error-page detection

  • Batch entry fetch — N accessions in one round trip, cross-referenced against the request to flag any missing

  • Async ID-mapping run → poll bounded by a wall-clock budget, with a running-job ticket and separately paginated completed results

Agent-friendly output:

  • Provenance is data, not decoration — reviewed, annotationScore, proteinExistence, and per-field PubMed/ECO evidence ship on every record so the agent can weigh manual vs. predicted annotation

  • Graceful partial failure — uniprot_get_entry returns per-accession succeeded[] / failed[] rows instead of aborting the batch

  • Discriminated output contracts — uniprot_get_entry returns kind: "full" | "outline", while uniprot_map_ids separates a running-job ticket from a finished-page continuation; callers branch on data, not string parsing

  • Sparsity preserved — absent upstream fields stay absent, never fabricated (most curated sections are legitimately missing on TrEMBL entries)

Getting started

Public Hosted Instance

A public instance is available at https://uniprot.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:

{
  "mcpServers": {
    "uniprot-mcp-server": {
      "type": "streamable-http",
      "url": "https://uniprot.caseyjhand.com/mcp"
    }
  }
}

Self-Hosted / Local

Add the following to your MCP client configuration file. UniProt REST is keyless — no API key required.

{
  "mcpServers": {
    "uniprot-mcp-server": {
      "type": "stdio",
      "command": "bunx",
      "args": ["@cyanheads/uniprot-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info"
      }
    }
  }
}

Or with npx (no Bun required):

{
  "mcpServers": {
    "uniprot-mcp-server": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@cyanheads/uniprot-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info"
      }
    }
  }
}

Or with Docker:

{
  "mcpServers": {
    "uniprot-mcp-server": {
      "type": "stdio",
      "command": "docker",
      "args": ["run", "-i", "--rm", "-e", "MCP_TRANSPORT_TYPE=stdio", "ghcr.io/cyanheads/uniprot-mcp-server:latest"]
    }
  }
}

For Streamable HTTP, set the transport and start the server:

MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcp

Prerequisites

  • Bun v1.3.0 or higher (or Node.js v24+).

  • No API key — UniProt REST is keyless and open. Data is UniProt, CC BY 4.0.

Installation

  1. Clone the repository:

git clone https://github.com/cyanheads/uniprot-mcp-server.git
  1. Navigate into the directory:

cd uniprot-mcp-server
  1. Install dependencies:

bun install

Configuration

All configuration is validated at startup via Zod schemas in src/config/server-config.ts. UniProt REST is keyless, so every server-specific variable below is an optional override.

Variable

Description

Default

UNIPROT_BASE_URL

UniProt REST base URL. Override for a private mirror or testing.

https://rest.uniprot.org

UNIPROT_TIMEOUT_MS

Per-request HTTP timeout in ms.

30000

UNIPROT_ID_MAPPING_BUDGET_MS

Wall-clock budget for the inline ID-mapping poll loop before returning a resumable ticket. Must be less than UNIPROT_TIMEOUT_MS.

8000

UNIPROT_DEFAULT_PAGE_SIZE

Default page size for search and proteome protein listing when the caller leaves it unset.

25

MCP_TRANSPORT_TYPE

Transport: stdio or http.

stdio

MCP_HTTP_PORT

Port for HTTP server.

3010

MCP_AUTH_MODE

Auth mode: none, jwt, or oauth.

none

MCP_SESSION_MODE

HTTP session posture: auto, stateful, or stateless. The server declares stateless in code — no tool holds per-session state — so set this only to opt back into a session store.

stateless

MCP_LOG_LEVEL

Log level (RFC 5424).

info

LOGS_DIR

Directory for log files (Node.js only).

<project-root>/logs

STORAGE_PROVIDER_TYPE

Storage backend.

in-memory

OTEL_ENABLED

Enable OpenTelemetry instrumentation (spans, metrics, completion logs).

false

See .env.example for the full list of optional overrides.

Running the server

Local development

  • Build and run:

    # One-time build
    bun run rebuild
    
    # Run the built server
    bun run start:stdio
    # or
    bun run start:http
  • Run checks and tests:

    bun run devcheck   # Lint, format, typecheck, security
    bun run test       # Vitest test suite
    bun run lint:mcp   # Validate MCP definitions against spec

Docker

docker build -t uniprot-mcp-server .
docker run --rm -e MCP_TRANSPORT_TYPE=stdio ghcr.io/cyanheads/uniprot-mcp-server:latest

The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/uniprot-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.

Project structure

Directory

Purpose

src/index.ts

createApp() entry point — registers tools/resources/prompts and inits the UniProt service.

src/config

Server-specific environment variable parsing and validation with Zod.

src/mcp-server/tools

Tool definitions (*.tool.ts). Six tools over UniProtKB, ID mapping, proteomes, taxonomy, and sequences.

src/mcp-server/resources

Resource definitions (*.resource.ts). Entry and taxonomy by-ID mirrors.

src/mcp-server/prompts

Prompt definitions (*.prompt.ts). The protein-dossier workflow prompt.

src/services/uniprot

The rest.uniprot.org REST client — search, batch entries, ID mapping, proteomes, taxonomy, FASTA — plus normalized domain types.

tests/

Unit and integration tests mirroring src/.

Development guide

See CLAUDE.md (and the byte-identical AGENTS.md) for development guidelines and architectural rules. The short version:

  • Handlers throw, framework catches — no try/catch in tool logic

  • Use ctx.log for request-scoped logging, ctx.state for tenant-scoped storage

  • Register new tools and resources in the createApp() arrays

  • Wrap the UniProt API: validate raw → normalize to domain type → return the output schema; never fabricate missing fields

Contributing

Issues are welcome. Run checks and tests before submitting:

bun run devcheck
bun run test

License

Apache-2.0 — see LICENSE for details.

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    Provides seamless access to UniProtKB protein database, enabling queries for protein entries, sequences, Gene Ontology annotations, full-text search, and ID mapping across 200+ database types.
    5
    2
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Provides access to UniProt protein sequence and function knowledge base, enabling search and retrieval of protein entries, proteomes, taxonomy, and feature annotations.
    190 npm
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    An MCP server that grounds protein research in the UniProt SPARQL endpoint, providing tools for querying proteins, sequences, variants, diseases, and more via intent-named tools and raw SPARQL.
    15
    MIT