Skip to main content
Glama
TakumiY235

UniProt MCP Server

by TakumiY235

UniProt MCP 服务器

一个模型上下文协议 (MCP) 服务器,提供对 UniProt 蛋白质信息的访问。该服务器允许 AI 助手直接从 UniProt 获取蛋白质功能和序列信息。

特征

  • 通过 UniProt 登录号获取蛋白质信息

  • 批量检索多种蛋白质

  • 缓存以提高性能(24 小时 TTL)

  • 错误处理和日志记录

  • 信息包括:

    • 蛋白质名称

    • 功能描述

    • 完整序列

    • 序列长度

    • 生物

Related MCP server: UniProt MCP Server

快速入门

  1. 确保安装了 Python 3.10 或更高版本

  2. 克隆此存储库:

    git clone https://github.com/TakumiY235/uniprot-mcp-server.git
    cd uniprot-mcp-server
  3. 安装依赖项:

    # Using uv (recommended)
    uv pip install -r requirements.txt
    
    # Or using pip
    pip install -r requirements.txt

配置

添加到您的 Claude Desktop 配置文件:

  • Windows: %APPDATA%\Claude\claude_desktop_config.json

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json

  • Linux: ~/.config/Claude/claude_desktop_config.json

{
  "mcpServers": {
    "uniprot": {
      "command": "uv",
      "args": ["--directory", "path/to/uniprot-mcp-server", "run", "uniprot-mcp-server"]
    }
  }
}

使用示例

在 Claude Desktop 中配置服务器后,您可以提出以下问题:

Can you get the protein information for UniProt accession number P98160?

对于批量查询:

Can you get and compare the protein information for both P04637 and P02747?

API 参考

工具

  1. get_protein_info

    • 获取单个蛋白质的信息

    • 必需参数: accession (UniProt 登录号)

    • 响应示例:

      {
        "accession": "P12345",
        "protein_name": "Example protein",
        "function": ["Description of protein function"],
        "sequence": "MLTVX...",
        "length": 123,
        "organism": "Homo sapiens"
      }
  2. get_batch_protein_info

    • 获取多种蛋白质的信息

    • 必需参数: accessions (UniProt 登录号数组)

    • 返回蛋白质信息对象数组

发展

设置开发环境

  1. 克隆存储库

  2. 创建虚拟环境:

    python -m venv .venv
    source .venv/bin/activate  # On Windows: .venv\Scripts\activate
  3. 安装开发依赖项:

    pip install -e ".[dev]"

运行测试

pytest

代码风格

该项目使用:

  • 黑色表示代码格式

  • isort 用于导入排序

  • flake8 用于 linting

  • mypy 用于类型检查

  • 强盗进行安全检查

  • 依赖项漏洞检查的安全性

运行所有检查:

black .
isort .
flake8 .
mypy .
bandit -r src/
safety check

技术细节

  • 使用 MCP Python SDK 构建

  • 使用 httpx 进行异步 HTTP 请求

  • 使用基于 OrderedDict 的缓存实现 24 小时 TTL 的缓存

  • 处理速率限制和重试

  • 提供详细的错误消息

错误处理

服务器处理各种错误情况:

  • 无效的登录号(404 个响应)

  • API 连接问题(网络错误)

  • 速率限制(429 条回复)

  • 格式错误的响应(JSON 解析错误)

  • 缓存管理(TTL 和大小限制)

贡献

欢迎大家贡献代码!欢迎提交 Pull 请求。贡献方式如下:

  1. 分叉存储库

  2. 创建你的功能分支( git checkout -b feature/amazing-feature )

  3. 提交您的更改( git commit -m 'Add some amazing feature' )

  4. 推送到分支( git push origin feature/amazing-feature )

  5. 打开拉取请求

请确保适当更新测试并遵守现有的编码风格。

执照

该项目根据 MIT 许可证获得许可 - 有关详细信息,请参阅LICENSE文件。

致谢

  • UniProt 提供蛋白质数据 API

  • Anthropic 的模型上下文协议规范

  • 帮助改进此项目的贡献者

Available Tools

2 tools
get_batch_protein_infoB

Get protein information for multiple accession No.

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionsYesList of UniProt accession No.

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden but lacks behavioral details. It doesn't disclose whether this is a read-only operation, potential rate limits, authentication needs, or what 'protein information' includes (e.g., format, fields). The description is minimal and adds little beyond the basic action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no wasted words, clearly front-loading the purpose. It is appropriately sized for a simple tool with one parameter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description is incomplete. It doesn't explain what 'protein information' entails, potential errors, or behavioral traits, leaving significant gaps for a tool that presumably returns complex data.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with the parameter 'accessions' documented as 'List of UniProt accession No.' in the schema. The description adds no additional meaning beyond this, such as format examples, constraints, or usage tips, so it meets the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get protein information') and the resource ('multiple accession No.'), making the purpose understandable. It distinguishes from the sibling tool 'get_protein_info' by specifying 'multiple' vs. presumably single, though not explicitly naming the alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when multiple accession numbers are needed, but provides no explicit guidance on when to use this vs. the sibling tool 'get_protein_info' (e.g., for bulk vs. single queries). No exclusions or prerequisites are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_protein_infoB

Get protein function and sequence information from UniProt using an accession No.

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionYesUniProt Accession No. (e.g., P12345)

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It mentions the data source (UniProt) and type of information, but lacks details on behavioral traits like rate limits, error handling, authentication needs, or response format. This is a significant gap for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the purpose without unnecessary words. Every part of the sentence contributes to understanding the tool's function.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description is incomplete. It does not explain what the return values look like (e.g., format of function and sequence information), error cases, or other contextual details needed for effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, with the parameter 'accession' well-documented in the schema. The description adds minimal value by mentioning 'UniProt Accession No.' and providing an example, but does not elaborate beyond what the schema already specifies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get') and resource ('protein function and sequence information from UniProt'), specifying the data source and type of information retrieved. It distinguishes from the sibling tool 'get_batch_protein_info' by implying this is for single proteins, though not explicitly contrasting them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when you have a UniProt accession number and need protein details, but does not explicitly state when to use this versus the sibling batch tool or other alternatives. No exclusions or prerequisites are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updates
    • First observedget_batch_protein_info
    • First observedget_protein_info

TDQS

B3.3/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have clearly distinct purposes: get_protein_info retrieves detailed function and sequence information for a single protein accession, while get_batch_protein_info handles multiple accessions in batch. There is no overlap or ambiguity in their functions.

Naming Consistency5/5

Both tools follow a consistent verb_noun pattern with 'get_' prefix and snake_case naming. The naming clearly indicates the action (get) and target (protein_info), with batch differentiation for the multi-accession tool.

Tool Count2/5

With only two tools, the server feels severely under-scoped for a UniProt domain. While the tools cover basic retrieval, there are obvious gaps for operations like searching, filtering, or accessing related data (e.g., taxonomy, structures), making the surface too thin for comprehensive protein information workflows.

Completeness2/5

The server is severely incomplete for UniProt functionality. It only provides protein information retrieval (single and batch), missing essential operations like search_by_keyword, get_taxonomy, get_structure, or update tracking. This will cause agent failures when trying to perform typical bioinformatics tasks beyond simple lookups.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    An MCP server that enables language models to fetch protein information from the UniProt database, including protein details, sequences, functions, and structures.
    MIT
  • A
    license
    B
    quality
    D
    maintenance
    Provides seamless access to UniProtKB protein database, enabling queries for protein entries, sequences, Gene Ontology annotations, full-text search, and ID mapping across 200+ database types.
    5
    2
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Provides access to UniProt protein sequence and function knowledge base, enabling search and retrieval of protein entries, proteomes, taxonomy, and feature annotations.
    190 npm
    MIT
  • F
    license
    B
    quality
    D
    maintenance
    Provides programmatic access to AlphaFold protein structure predictions and UniProt data, enabling users to retrieve protein structures, summaries, and annotations through natural language.
    3
    -