Skip to main content
Glama
josefdc

UniProt MCP Server

by josefdc

search_uniprot

Search the UniProtKB protein database to find curated protein entries using queries for sequences, annotations, or full-text search.

Instructions

Search UniProtKB and return curated hits.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
queryYes
sizeNo
reviewed_onlyNo
fieldsNo
sortNo
include_isoformNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Implementation Reference

  • The main tool handler for 'search_uniprot'. Decorated with @mcp.tool() for registration. Performs size validation, optional field optimization, calls the UniProt search API via search_json, and parses the response into SearchHit models.
    @mcp.tool()  # type: ignore[misc]
    async def search_uniprot(
        query: str,
        size: int = 10,
        reviewed_only: bool = False,
        fields: list[str] | None = None,
        sort: str | None = None,
        include_isoform: bool = False,
    ) -> list[SearchHit]:
        """Search UniProtKB and return curated hits."""
    
        size = max(1, min(size, 500))
        effective_fields = fields
        if fields is None and os.getenv("UNIPROT_ENABLE_FIELDS"):
            effective_fields = [FIELDS_SEARCH_LIGHT]
        async with new_client() as client:
            payload = await search_json(
                client,
                query=query,
                size=size,
                reviewed_only=reviewed_only,
                fields=effective_fields,
                sort=sort,
                include_isoform=include_isoform,
            )
        return parse_search_hits(payload)
  • Pydantic model defining the structured output for each search hit returned by the tool.
    class SearchHit(BaseModel):
        """Result entry returned by UniProt search endpoints."""
    
        accession: str = Field(description="Primary accession identifier.")
        id: str | None = Field(default=None, description="UniProt entry name/ID.")
        reviewed: bool = Field(description="True for Swiss-Prot hits.")
        protein_name: str | None = Field(
            default=None, description="Recommended protein name when available."
        )
        organism: str | None = Field(
            default=None, description="Scientific name of the source organism."
        )
  • Helper function that parses the raw JSON search response from UniProt into a list of SearchHit Pydantic models.
    def parse_search_hits(js: dict[str, Any]) -> list[SearchHit]:
        """Convert a UniProt search response into SearchHit models."""
    
        hits: list[SearchHit] = []
        for result in js.get("results") or []:
            if not isinstance(result, dict):
                continue
            accession = result.get("primaryAccession") or result.get("accession")
            if not accession:
                continue
            organism_block = result.get("organism", {})
            organism_name = (
                organism_block.get("scientificName") if isinstance(organism_block, dict) else None
            )
            entry_type = result.get("entryType") or ""
            hits.append(
                SearchHit(
                    accession=str(accession),
                    id=result.get("uniProtkbId") or result.get("id"),
                    reviewed=str(entry_type).startswith("UniProtKB reviewed"),
                    protein_name=_extract_protein_name(result),
                    organism=organism_name,
                )
            )
        return hits
  • Low-level HTTP client helper that executes the actual search request to UniProt's REST API (/uniprotkb/search) with retry logic, handles reviewed_only query modification, and returns raw JSON.
    async def search_json(
        client: httpx.AsyncClient,
        *,
        query: str,
        size: int,
        reviewed_only: bool = False,
        fields: Iterable[str] | None = None,
        sort: str | None = None,
        include_isoform: bool | None = None,
    ) -> dict[str, Any]:
        """Search UniProtKB and return the raw JSON response."""
    
        actual_query = query
        if reviewed_only and "reviewed:" not in query.lower():
            actual_query = f"({query}) AND reviewed:true"
    
        bounded_size = max(1, min(size, 500))
    
        params: dict[str, Any] = {
            "query": actual_query,
            "size": bounded_size,
        }
        if fields:
            params["fields"] = ",".join(fields)
        if sort:
            params["sort"] = sort
        if include_isoform is not None:
            params["includeIsoform"] = str(include_isoform).lower()
    
        async with _SEMAPHORE:
            response = await client.get(
                "/uniprotkb/search",
                params=params,
            )
        if response.status_code >= 400:
            if response.status_code in RETRYABLE_STATUS:
                response.raise_for_status()
            else:
                try:
                    response.raise_for_status()
                except httpx.HTTPStatusError as exc:
                    raise UniProtClientError(str(exc)) from exc
        return _ensure_dict(response.json())
  • The @mcp.tool() decorator registers the search_uniprot function as an MCP tool, using the function name as the tool name.
    @mcp.tool()  # type: ignore[misc]

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv1.0.0
    • addedInput schema / title
      Added value: +"search_uniprotArguments"
  2. First observed

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description should disclose behavioral traits. It only says 'search and return curated hits,' offering no details on pagination, rate limits, or mutation (if any). The tool's behavior remains opaque.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise but lacks structure and fails to provide necessary details in a front-loaded manner. It is not overly verbose but under-informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 6 parameters, no annotations, and an output schema exists but is not described, the description is incomplete. It does not explain what 'curated hits' entails or how parameters affect behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 6 parameters with 0% description coverage, and the tool description adds no parameter information. The agent is left to infer meaning from parameter names alone, which is insufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (search) and resource (UniProtKB) and mentions 'curated hits' as output, which differentiates from sibling tools that retrieve specific entries or sequences. However, 'curated hits' is somewhat vague, preventing a perfect score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus siblings like fetch_entry or get_sequence. There is no mention of prerequisites or alternatives, leaving the agent without context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.