UniProt MCP Server
UniProt MCP 서버
UniProt 단백질 정보에 대한 접근을 제공하는 모델 컨텍스트 프로토콜(MCP) 서버입니다. 이 서버를 통해 AI 보조자는 UniProt에서 단백질 기능 및 서열 정보를 직접 가져올 수 있습니다.
특징
UniProt 접근 번호로 단백질 정보를 얻으세요
다중 단백질의 일괄 검색
성능 향상을 위한 캐싱(24시간 TTL)
오류 처리 및 로깅
정보는 다음과 같습니다.
단백질 이름
기능 설명
전체 시퀀스
시퀀스 길이
유기체
Related MCP server: UniProt MCP Server
빠른 시작
Python 3.10 이상이 설치되어 있는지 확인하세요.
이 저장소를 복제하세요:
지엑스피1
종속성 설치:
# Using uv (recommended) uv pip install -r requirements.txt # Or using pip pip install -r requirements.txt
구성
Claude Desktop 구성 파일에 다음을 추가합니다.
Windows:
%APPDATA%\Claude\claude_desktop_config.jsonmacOS:
~/Library/Application Support/Claude/claude_desktop_config.json리눅스:
~/.config/Claude/claude_desktop_config.json
{
"mcpServers": {
"uniprot": {
"command": "uv",
"args": ["--directory", "path/to/uniprot-mcp-server", "run", "uniprot-mcp-server"]
}
}
}사용 예
Claude Desktop에서 서버를 구성한 후 다음과 같은 질문을 할 수 있습니다.
Can you get the protein information for UniProt accession number P98160?일괄 쿼리의 경우:
Can you get and compare the protein information for both P04637 and P02747?API 참조
도구
get_protein_info단일 단백질에 대한 정보 얻기
필수 매개변수:
accession(UniProt 접근 번호)응답 예시:
{ "accession": "P12345", "protein_name": "Example protein", "function": ["Description of protein function"], "sequence": "MLTVX...", "length": 123, "organism": "Homo sapiens" }
get_batch_protein_info다양한 단백질에 대한 정보를 얻으세요
필수 매개변수:
accessions(UniProt 접근 번호 배열)단백질 정보 객체의 배열을 반환합니다.
개발
개발 환경 설정
저장소를 복제합니다
가상 환경 만들기:
python -m venv .venv source .venv/bin/activate # On Windows: .venv\Scripts\activate개발 종속성 설치:
pip install -e ".[dev]"
테스트 실행
pytest코드 스타일
이 프로젝트에서는 다음을 사용합니다.
코드 서식을 위한 검정색
수입 정렬을 위한 isort
린팅용 flake8
유형 검사를 위한 mypy
보안 검사를 위한 산적
종속성 취약성 검사를 위한 안전성
모든 검사를 실행합니다.
black .
isort .
flake8 .
mypy .
bandit -r src/
safety check기술적 세부 사항
MCP Python SDK를 사용하여 구축됨
비동기 HTTP 요청에 httpx를 사용합니다.
OrderedDict 기반 캐시를 사용하여 24시간 TTL로 캐싱을 구현합니다.
속도 제한 및 재시도를 처리합니다.
자세한 오류 메시지를 제공합니다
오류 처리
서버는 다양한 오류 시나리오를 처리합니다.
잘못된 접근 번호(404개 응답)
API 연결 문제(네트워크 오류)
속도 제한(429개 응답)
잘못된 응답(JSON 구문 분석 오류)
캐시 관리(TTL 및 크기 제한)
기여하다
여러분의 참여를 환영합니다! 풀 리퀘스트를 제출해 주세요. 참여 방법은 다음과 같습니다.
저장소를 포크하세요
기능 브랜치를 생성합니다(
git checkout -b feature/amazing-feature)변경 사항을 커밋하세요(
git commit -m 'Add some amazing feature')브랜치에 푸시(
git push origin feature/amazing-feature)풀 리퀘스트 열기
적절하게 테스트를 업데이트하고 기존 코딩 스타일을 준수하세요.
특허
이 프로젝트는 MIT 라이선스에 따라 라이선스가 부여되었습니다. 자세한 내용은 라이선스 파일을 참조하세요.
감사의 말
단백질 데이터 API를 제공하는 UniProt
모델 컨텍스트 프로토콜 사양에 대한 Anthropic
이 프로젝트를 개선하는 데 도움을 준 기여자
Available Tools
2 toolsget_batch_protein_infoB
Get protein information for multiple accession No.
| Name | Required | Description | Default |
|---|---|---|---|
| accessions | Yes | List of UniProt accession No. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but lacks behavioral details. It doesn't disclose whether this is a read-only operation, potential rate limits, authentication needs, or what 'protein information' includes (e.g., format, fields). The description is minimal and adds little beyond the basic action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with no wasted words, clearly front-loading the purpose. It is appropriately sized for a simple tool with one parameter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description is incomplete. It doesn't explain what 'protein information' entails, potential errors, or behavioral traits, leaving significant gaps for a tool that presumably returns complex data.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the parameter 'accessions' documented as 'List of UniProt accession No.' in the schema. The description adds no additional meaning beyond this, such as format examples, constraints, or usage tips, so it meets the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Get protein information') and the resource ('multiple accession No.'), making the purpose understandable. It distinguishes from the sibling tool 'get_protein_info' by specifying 'multiple' vs. presumably single, though not explicitly naming the alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when multiple accession numbers are needed, but provides no explicit guidance on when to use this vs. the sibling tool 'get_protein_info' (e.g., for bulk vs. single queries). No exclusions or prerequisites are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_protein_infoB
Get protein function and sequence information from UniProt using an accession No.
| Name | Required | Description | Default |
|---|---|---|---|
| accession | Yes | UniProt Accession No. (e.g., P12345) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions the data source (UniProt) and type of information, but lacks details on behavioral traits like rate limits, error handling, authentication needs, or response format. This is a significant gap for a tool with no annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the purpose without unnecessary words. Every part of the sentence contributes to understanding the tool's function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description is incomplete. It does not explain what the return values look like (e.g., format of function and sequence information), error cases, or other contextual details needed for effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, with the parameter 'accession' well-documented in the schema. The description adds minimal value by mentioning 'UniProt Accession No.' and providing an example, but does not elaborate beyond what the schema already specifies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Get') and resource ('protein function and sequence information from UniProt'), specifying the data source and type of information retrieved. It distinguishes from the sibling tool 'get_batch_protein_info' by implying this is for single proteins, though not explicitly contrasting them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when you have a UniProt accession number and need protein details, but does not explicitly state when to use this versus the sibling batch tool or other alternatives. No exclusions or prerequisites are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
- First observed
get_batch_protein_info - First observed
get_protein_info
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: get_protein_info retrieves detailed function and sequence information for a single protein accession, while get_batch_protein_info handles multiple accessions in batch. There is no overlap or ambiguity in their functions.
Both tools follow a consistent verb_noun pattern with 'get_' prefix and snake_case naming. The naming clearly indicates the action (get) and target (protein_info), with batch differentiation for the multi-accession tool.
With only two tools, the server feels severely under-scoped for a UniProt domain. While the tools cover basic retrieval, there are obvious gaps for operations like searching, filtering, or accessing related data (e.g., taxonomy, structures), making the surface too thin for comprehensive protein information workflows.
The server is severely incomplete for UniProt functionality. It only provides protein information retrieval (single and batch), missing essential operations like search_by_keyword, get_taxonomy, get_structure, or update tracking. This will cause agent failures when trying to perform typical bioinformatics tasks beyond simple lookups.
Maintenance
Related MCP Connectors
Give AI assistants access to real-time data. Search the web, compare flights, find hotels, and more.
Provide AI assistants with real-time access to official SEC EDGAR filings and financial data. Enab…
Enable AI assistants to perform web searches using Perplexity's Sonar Pro.
Live data gateway for AI — 3,300+ tools across 750+ sources, with citations
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceAn MCP server that enables language models to fetch protein information from the UniProt database, including protein details, sequences, functions, and structures.MIT
- AlicenseBqualityDmaintenanceProvides seamless access to UniProtKB protein database, enabling queries for protein entries, sequences, Gene Ontology annotations, full-text search, and ID mapping across 200+ database types.52MIT
- AlicenseNot gradedqualityBmaintenanceProvides access to UniProt protein sequence and function knowledge base, enabling search and retrieval of protein entries, proteomes, taxonomy, and feature annotations.190 npmMIT
- FlicenseBqualityDmaintenanceProvides programmatic access to AlphaFold protein structure predictions and UniProt data, enabling users to retrieve protein structures, summaries, and annotations through natural language.3-