Skip to main content
Glama

UniProt MCP 服务器徽标

UniProt MCP 服务器

全面的模型上下文协议 (MCP) 服务器,提供对 UniProt 蛋白质数据库的高级访问。该服务器提供 26 种专用生物信息学工具,使 AI 助手和 MCP 客户端能够直接通过 UniProt 的 REST API 进行复杂的蛋白质研究、比较基因组学、结构生物学分析和系统生物学研究。

由增强自然开发

特征

核心蛋白质分析(5 种工具)

  • 蛋白质搜索:通过蛋白质名称、关键词或生物体搜索 UniProt 数据库

  • 详细蛋白质信息:检索全面的蛋白质信息,包括功能、结构和注释

  • 基于基因的搜索:通过基因名称或符号查找蛋白质

  • 序列检索:获取 FASTA 或 JSON 格式的氨基酸序列

  • 特征分析:访问功能域、活性位点、结合位点和其他蛋白质特征

比较与进化分析(4 种工具)

  • 蛋白质比较:并排比较多种蛋白质,并进行序列和特征分析

  • 同源物发现:寻找不同物种间的同源蛋白质

  • 直系同源物鉴定:鉴定用于进化研究的直系同源蛋白

  • 系统发育分析:检索进化关系和系统发育数据

结构与功能分析(4 种工具)

  • 3D 结构信息:访问 PDB 参考和结构数据

  • 高级域分析:使用 InterPro、Pfam 和 SMART 注释增强域分析

  • 变异分析:疾病相关变异和突变

  • 序列组成:氨基酸组成、疏水性和其他序列特性

生物学背景分析(4 种工具)

  • 通路整合:来自 KEGG 和 Reactome 的相关生物通路

  • 蛋白质相互作用:蛋白质-蛋白质相互作用网络

  • 功能分类:通过 GO 术语或功能注释进行搜索

  • 亚细胞定位:通过亚细胞定位查找蛋白质

批处理和高级搜索(3 种工具)

  • 批量处理:高效处理多个蛋白质样本

  • 高级搜索:具有多个过滤器的复杂查询(长度、质量、生物体、功能)

  • 分类学分类:按详细分类法搜索

文献与交叉引用(3 种工具)

  • 外部数据库链接:链接到 PDB、EMBL、RefSeq、Ensembl 和其他数据库

  • 参考文献:相关出版物和引文

  • 注释质量:不同注释的质量分数和置信度

数据导出和实用程序(3 个工具)

  • 专业导出:以 GFF、GenBank、EMBL 和 XML 格式导出数据

  • 接入验证:验证 UniProt 接入号的有效性

  • 分类信息:详细的分类和谱系数据

资源模板

  • 通过 URI 模板直接访问蛋白质数据,实现无缝集成

Related MCP server: ChEMBL MCP Server

安装

先决条件

  • Node.js(v16 或更高版本)

  • npm 或 yarn

设置

  1. 克隆存储库:

git clone <repository-url>
cd uniprot-server
  1. 安装依赖项:

npm install
  1. 构建项目:

npm run build

Docker

构建 Docker 镜像

构建 Docker 镜像:

docker build -t uniprot-mcp-server .

使用 Docker 运行

运行容器:

docker run -i uniprot-mcp-server

对于 MCP 客户端集成,您可以直接使用容器:

{
  "mcpServers": {
    "uniprot": {
      "command": "docker",
      "args": ["run", "-i", "uniprot-mcp-server"],
      "env": {}
    }
  }
}

Docker Compose(可选)

创建一个docker-compose.yml以便于管理:

version: "3.8"
services:
  uniprot-mcp:
    build: .
    image: uniprot-mcp-server
    stdin_open: true
    tty: true

运行:

docker-compose up

用法

作为 MCP 服务器

该服务器设计为作为通过 stdio 进行通信的 MCP 服务器运行:

npm start

添加至 MCP 客户端配置

将服务器添加到您的 MCP 客户端配置(例如,Claude Desktop):

{
  "mcpServers": {
    "uniprot": {
      "command": "node",
      "args": ["/path/to/uniprot-server/build/index.js"],
      "env": {}
    }
  }
}

可用工具

1. 搜索蛋白质

按名称、关键字或生物体搜索 UniProt 数据库中的蛋白质。

参数:

  • query (必填):搜索查询(蛋白质名称、关键字或复杂搜索)

  • organism (可选):用于过滤结果的生物体名称或分类 ID

  • size (可选):返回的结果数(1-500,默认值:25)

  • format (可选):输出格式 - json、tsv、fasta、xml(默认值:json)

例子:

{
  "query": "insulin",
  "organism": "human",
  "size": 5
}

2. 获取蛋白质信息

通过 UniProt 登录获取特定蛋白质的详细信息。

参数:

  • accession (必填):UniProt 登录号(例如,P04637)

  • format (可选):输出格式 - json、tsv、fasta、xml(默认值:json)

例子:

{
  "accession": "P01308",
  "format": "json"
}

3. 按基因搜索

通过基因名称或符号搜索蛋白质。

参数:

  • gene (必填):基因名称或符号(例如,BRCA1、INS)

  • organism (可选):用于过滤结果的生物体名称或分类 ID

  • size (可选):返回的结果数(1-500,默认值:25)

例子:

{
  "gene": "BRCA1",
  "organism": "human"
}

4. 获取蛋白质序列

获取蛋白质的氨基酸序列。

参数:

  • accession (必填):UniProt 接入号

  • format (可选):输出格式 - fasta、json(默认值:fasta)

例子:

{
  "accession": "P01308",
  "format": "fasta"
}

5. 获取蛋白质特征

获取蛋白质的功能特征和结构域。

参数:

  • accession (必填):UniProt 接入号

例子:

{
  "accession": "P01308"
}

资源模板

服务器通过 URI 模板提供对 UniProt 数据的直接访问:

1. 蛋白质信息

  • URI : uniprot://protein/{accession}

  • 描述:UniProt 数据库的完整蛋白质信息

  • 例如: uniprot://protein/P01308

2. 蛋白质序列

  • URI : uniprot://sequence/{accession}

  • 描述:FASTA格式的蛋白质序列

  • 例如: uniprot://sequence/P01308

3.搜索结果

  • URI : uniprot://search/{query}

  • 描述:与查询匹配的蛋白质的搜索结果

  • 例如: uniprot://search/insulin

示例

基本蛋白质搜索

在人类中寻找胰岛素蛋白:

// Tool call
{
  "tool": "search_proteins",
  "arguments": {
    "query": "insulin",
    "organism": "human",
    "size": 10
  }
}

获取详细的蛋白质信息

检索有关人类胰岛素的综合信息:

// Tool call
{
  "tool": "get_protein_info",
  "arguments": {
    "accession": "P01308"
  }
}

基于基因的搜索

查找与 BRCA1 基因相关的蛋白质:

// Tool call
{
  "tool": "search_by_gene",
  "arguments": {
    "gene": "BRCA1",
    "organism": "human"
  }
}

检索蛋白质序列

获取人类胰岛素的氨基酸序列:

// Tool call
{
  "tool": "get_protein_sequence",
  "arguments": {
    "accession": "P01308",
    "format": "fasta"
  }
}

分析蛋白质特征

获取人类胰岛素的功能域和特征:

// Tool call
{
  "tool": "get_protein_features",
  "arguments": {
    "accession": "P01308"
  }
}

API 集成

该服务器集成了 UniProt REST API,可通过编程方式访问蛋白质数据。有关 UniProt 的更多信息,请访问:

所有 API 请求包括:

  • 用户代理: UniProt-MCP-Server/1.0.0

  • 超时:30秒

  • 基本网址: https://rest.uniprot.org (仅限编程访问)

错误处理

该服务器包括全面的错误处理:

  • 输入验证:所有参数都使用类型保护进行验证

  • API 错误:捕获网络和 API 错误并返回描述性消息

  • 超时处理:30秒后请求超时

  • 优雅降级:部分故障得到适当处理

发展

构建项目

npm run build

开发模式

在监视模式下运行 TypeScript 编译器:

npm run dev

项目结构

uniprot-server/
├── src/
│   └── index.ts          # Main server implementation
├── build/                # Compiled JavaScript output
├── package.json          # Node.js dependencies and scripts
├── tsconfig.json         # TypeScript configuration
└── README.md            # This file

依赖项

  • @modelcontextprotocol/sdk :用于服务器实现的核心 MCP SDK

  • axios :UniProt API 请求的 HTTP 客户端

  • typescript :用于开发的 TypeScript 编译器

执照

MIT 许可证

贡献

  1. 分叉存储库

  2. 创建功能分支

  3. 进行更改

  4. 如果适用,添加测试

  5. 提交拉取请求

支持

对于问题和疑问:

  1. 查看UniProt API 文档

  2. 查看模型上下文协议规范

  3. 在存储库上打开一个问题

关于增强自然

这款功能全面的 UniProt MCP 服务器由**Augmented Nature**开发,该公司是人工智能驱动的生物信息学和计算生物学解决方案领域的领先创新者。Augmented Nature 专注于开发先进的工具,弥合人工智能与生物学研究之间的差距,使研究人员能够从生物数据中获得更深入的洞察。

完整工具参考

核心蛋白质分析工具

  1. search_proteins - 按名称、关键字或生物体搜索 UniProt 数据库

  2. get_protein_info - 通过访问获取详细的蛋白质信息

  3. search_by_gene - 通过基因名称或符号查找蛋白质

  4. get_protein_sequence - 检索氨基酸序列

  5. get_protein_features - 访问功能特征和域

比较与进化分析工具

  1. compare_proteins - 并排比较多种蛋白质

  2. get_protein_homologs - 查找跨物种的同源蛋白质

  3. get_protein_orthologs - 识别直系同源蛋白

  4. get_phylogenetic_info - 检索进化关系

结构和功能分析工具

  1. get_protein_structure - 从 PDB 获取 3D 结构信息

  2. get_protein_domains_detailed - 增强域分析(InterPro、Pfam、SMART)

  3. get_protein_variants - 疾病相关变异和突变

  4. analyze_sequence_composition - 氨基酸组成分析

生物学背景工具

  1. get_protein_pathways - 相关生物途径(KEGG、Reactome)

  2. get_protein_interactions - 蛋白质-蛋白质相互作用网络

  3. search_by_function - 通过 GO 术语或功能注释进行搜索

  4. search_by_localization - 通过亚细胞定位查找蛋白质

批处理和高级搜索工具

  1. batch_protein_lookup - 高效处理多个蛋白质

  2. advanced_search - 具有多个过滤器的复杂查询

  3. search_by_taxonomy - 按分类法搜索

文献与交叉引用工具

  1. get_external_references - 链接到其他数据库(PDB、EMBL、RefSeq 等)

  2. get_literature_references - 相关出版物和引文

  3. get_annotation_confidence - 注释的质量分数

数据导出和实用工具

  1. export_protein_data - 以专门的格式导出(GFF、GenBank、EMBL、XML)

  2. validate_accession接入号有效性

  3. get_taxonomy_info - 详细的分类信息

变更日志

v1.0.0 - 综合生物信息学平台

  • 主要扩展:增加了 21 种新的专用工具(总计:26 种工具)

  • 比较分析:蛋白质比较、同源物/直系同源物鉴定、系统发育分析

  • 结构生物学:3D结构整合、详细域分析、变体分析

  • 系统生物学:通路整合、蛋白质相互作用、功能分类

  • 高级搜索:批处理、复杂过滤、分类搜索

  • 文献整合:外部数据库链接、引文、注释置信度

  • 数据导出:多种专用格式(GFF、GenBank、EMBL、XML)

  • 增强的 Docker 支持:采用安全最佳实践的多阶段构建

  • 全面的文档:完整的工具参考和示例

  • 由 Augmented Nature 开发:专业生物信息学平台

Available Tools

26 tools
analyze_sequence_compositionC

Amino acid composition, hydrophobicity, and other sequence properties

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionYesUniProt accession number

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions the types of properties analyzed (amino acid composition, hydrophobicity, etc.) but lacks critical details: whether this is a read-only operation, computational requirements, potential rate limits, or what the output looks like (e.g., numerical values, plots). For a tool with no annotation coverage, this is a significant gap in transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that lists key analysis types without unnecessary words. It's front-loaded with the core purpose, though it could be slightly more structured by explicitly mentioning the input or output. Overall, it's concise and avoids redundancy, earning a high score for brevity and clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of sequence analysis and the lack of annotations and output schema, the description is incomplete. It doesn't explain what the tool returns (e.g., a report, JSON object, or visual summary), how properties are calculated, or any limitations (e.g., supported sequence types). For a tool that likely involves computational analysis, this leaves too much unspecified for effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with the single parameter 'accession' documented as a UniProt accession number. The description adds no additional parameter semantics beyond what the schema provides—it doesn't explain how the accession is used to derive the analysis or any constraints (e.g., valid formats). Since schema coverage is high, the baseline score of 3 is appropriate, as the description doesn't compensate but doesn't detract either.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: analyzing amino acid composition, hydrophobicity, and other sequence properties. It specifies the verb 'analyze' and the resource 'sequence properties', distinguishing it from siblings like get_protein_sequence (which retrieves raw sequence) or get_protein_info (which provides general metadata). However, it doesn't explicitly mention the input (accession number) or output format, leaving some ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't specify scenarios where this analysis is needed (e.g., for protein characterization vs. structural prediction) or differentiate it from siblings like get_protein_features (which might include some overlapping properties). Without such context, users must infer usage based on the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

batch_protein_lookupC

Process multiple accessions efficiently

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionsYesArray of UniProt accession numbers (1-100)
formatNoOutput format (default: json)

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. 'Process multiple accessions efficiently' implies a read operation but lacks details on behavior: it doesn't specify what data is returned, any rate limits, error handling for invalid accessions, or performance characteristics. This is inadequate for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no wasted words, making it appropriately concise. However, it's under-specified rather than optimally structured, as it could benefit from front-loading more specific information about the tool's purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description is incomplete for a tool that likely returns protein-related data. It doesn't explain what 'process' yields (e.g., protein info, sequences), leaving gaps in understanding the tool's behavior and output, which is insufficient for effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents both parameters (accessions array with constraints, format enum with default). The description adds no meaning beyond this, as it doesn't explain parameter usage or semantics. Baseline 3 is appropriate since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Process multiple accessions efficiently' states the verb ('process') and resource ('multiple accessions'), but it's vague about what processing entails compared to siblings like 'get_protein_info' or 'get_protein_sequence'. It doesn't specify if this returns protein data, sequences, or annotations, leaving ambiguity in distinguishing its exact function from similar tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. With siblings like 'get_protein_info' and 'get_protein_sequence', it's unclear if this tool is for batch retrieval of general info, sequences, or something else, and there are no explicit when/when-not instructions or named alternatives mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compare_proteinsB

Compare multiple proteins side-by-side with sequence and feature comparison

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionsYesArray of UniProt accession numbers (2-10)
formatNoOutput format (default: json)

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden but only mentions the comparison functionality without behavioral details. It doesn't disclose rate limits, authentication requirements, whether it's read-only or has side effects, or what the comparison output looks like (beyond format options in schema).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that clearly states the tool's purpose without unnecessary words. It's appropriately sized and front-loaded with the core functionality.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter tool with no annotations and no output schema, the description is minimally adequate but lacks important context. It doesn't explain what 'feature comparison' includes, how results are structured, or any limitations beyond the parameter constraints in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so parameters are well-documented in the schema. The description adds no additional parameter semantics beyond what's in the schema (e.g., doesn't explain what 'feature comparison' entails or how accessions should be formatted). Baseline 3 is appropriate when schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('compare') and resource ('multiple proteins') with specific comparison aspects ('sequence and feature comparison'). It distinguishes from siblings like 'get_protein_sequence' or 'get_protein_features' by emphasizing side-by-side comparison of multiple proteins, though it doesn't explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for comparing multiple proteins, but doesn't explicitly state when to use this versus alternatives like 'get_protein_homologs' for evolutionary comparisons or 'batch_protein_lookup' for basic data retrieval. No guidance on prerequisites or exclusions is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_protein_dataC

Export data in specialized formats (GFF, GenBank, etc.)

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionYesUniProt accession number
formatYesExport format

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It states 'export' implies a read operation, but lacks details on permissions, rate limits, output behavior (e.g., file download vs. inline data), or side effects. This is a significant gap for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero waste. It's front-loaded with the core purpose and includes relevant examples, making it appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description is incomplete. It doesn't cover behavioral aspects like output format details, error handling, or usage context, which are critical for an export tool with two required parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents both parameters (accession and format with enum). The description adds minimal value by listing example formats (GFF, GenBank) beyond the schema's enum, but doesn't explain parameter interactions or constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'export' and resource 'data', specifying specialized formats like GFF and GenBank. It distinguishes from siblings by focusing on export functionality rather than search, analysis, or retrieval operations, though it doesn't explicitly contrast with specific tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description mentions formats but doesn't indicate scenarios (e.g., for data sharing, analysis compatibility) or prerequisites, leaving usage context unclear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_annotation_confidenceD

Quality scores for different annotations

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionYesUniProt accession number

TDQS

D1.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. However, it only states 'Quality scores for different annotations' without explaining what 'quality scores' are (e.g., confidence values, metrics), how they are returned, or any behavioral traits like rate limits, permissions, or response format. This is inadequate for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single phrase, 'Quality scores for different annotations', which is concise but under-specified—it lacks necessary detail for clarity. While it is front-loaded and wastes no words, the brevity comes at the cost of usefulness, making it more of a placeholder than an informative description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity implied by the tool name (involving 'annotation confidence') and the lack of annotations and output schema, the description is incomplete. It does not explain what 'quality scores' are, how they are structured, or what annotations are covered, leaving significant gaps for the agent to understand the tool's functionality and output.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with the single parameter 'accession' clearly documented as a 'UniProt accession number'. The description adds no additional meaning beyond this, such as examples or constraints. Since schema coverage is high, the baseline score of 3 is appropriate, as the schema adequately handles parameter semantics without description enhancement.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Quality scores for different annotations' is vague and tautological—it essentially restates the tool name 'get_annotation_confidence' without specifying what resource it acts on or what 'quality scores' entail. It does not clearly distinguish this tool from siblings like 'get_protein_info' or 'get_protein_features', which might also provide annotation-related data. The purpose lacks a specific verb and target resource.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention any context, prerequisites, or exclusions, nor does it refer to sibling tools. This leaves the agent with no information to decide between this tool and others like 'get_protein_info' or 'get_protein_features' for annotation-related queries.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_external_referencesD

Links to other databases (PDB, EMBL, RefSeq, etc.)

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionYesUniProt accession number

TDQS

D1.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. However, it only lists database names without explaining what the tool does (e.g., returns URLs, IDs, or metadata), any rate limits, authentication needs, or output format. This leaves the agent guessing about the tool's behavior, warranting a score of 1.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise phrase, but it is under-specified rather than efficiently informative. It lacks front-loaded clarity (e.g., starting with a verb like 'Retrieve') and wastes space on generic examples ('PDB, EMBL, RefSeq, etc.') without adding actionable context. A score of 3 reflects this balance between brevity and insufficient detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (a read operation with one parameter) and the absence of annotations and output schema, the description is incomplete. It does not explain what the tool returns (e.g., links, identifiers, or metadata), leaving gaps in understanding its functionality. While the schema covers the parameter, the overall context is inadequate, scoring 2.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with the 'accession' parameter clearly documented as a 'UniProt accession number'. The description adds no additional meaning about parameters, such as format examples or constraints. According to the rules, with high schema coverage (>80%), the baseline is 3 even without param info in the description, so this score is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Links to other databases (PDB, EMBL, RefSeq, etc.)' is vague and tautological—it essentially restates the tool name 'get_external_references' without specifying the action (e.g., 'retrieve' or 'fetch') or the resource (e.g., 'for a given protein'). It does not clearly distinguish this tool from siblings like 'get_protein_info' or 'get_protein_sequence', which might also involve external data. A score of 2 reflects this lack of specificity and differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention any context, prerequisites, or exclusions, such as when to prefer this over 'get_protein_info' (which might include references) or 'search_by_function'. With no implied or explicit usage instructions, this is a minimal score of 1.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_literature_referencesC

Associated publications and citations

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionYesUniProt accession number

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It does not mention whether this is a read-only operation, if it requires authentication, rate limits, or what the output format might be. The description is minimal and fails to provide essential behavioral context for a tool with no annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient phrase with no wasted words. It is appropriately sized for a simple tool and front-loaded with the core purpose, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description is incomplete. It does not explain what the tool returns (e.g., list of publications, citation details) or any behavioral traits. For a tool with no structured data beyond the input schema, the description should provide more context to be fully helpful.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with the parameter 'accession' clearly documented as a 'UniProt accession number'. The description adds no additional meaning beyond the schema, so it meets the baseline of 3 for high schema coverage without compensating with extra details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Associated publications and citations' states the purpose but is vague about the action. It implies retrieving references but doesn't specify the verb (e.g., 'retrieve' or 'fetch') or clearly distinguish it from sibling tools like 'get_external_references'. The purpose is understandable but lacks specificity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as 'get_external_references' or other siblings. The description does not mention any context, prerequisites, or exclusions, leaving the agent to infer usage based on the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_phylogenetic_infoC

Retrieve evolutionary relationships and phylogenetic data

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionYesUniProt accession number

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool 'retrieves' data, implying a read-only operation, but doesn't specify whether it's idempotent, has rate limits, requires authentication, or what the return format looks like. This is inadequate for a tool with potential complexity in phylogenetic data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no wasted words. It's front-loaded with the core purpose and appropriately sized for a simple retrieval tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description is incomplete. It doesn't explain what 'phylogenetic data' includes (e.g., tree formats, confidence scores) or behavioral aspects like error handling. For a tool dealing with evolutionary relationships, more context is needed to use it effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with the single parameter 'accession' documented as a 'UniProt accession number'. The description doesn't add any additional meaning beyond this, such as format examples or constraints, so it meets the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verbs ('retrieve') and resources ('evolutionary relationships and phylogenetic data'), making it easy to understand what the tool does. However, it doesn't explicitly differentiate from sibling tools like 'get_taxonomy_info' or 'get_protein_homologs' which might also relate to evolutionary data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites, context, or exclusions, leaving the agent to infer usage from the tool name and parameters alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_protein_domains_detailedC

Enhanced domain analysis with InterPro, Pfam, and SMART annotations

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionYesUniProt accession number

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions 'enhanced domain analysis' but doesn't specify what 'enhanced' entails (e.g., more detailed data, additional sources, or computational intensity), nor does it cover aspects like rate limits, authentication needs, or response format. This is a significant gap for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the core purpose ('Enhanced domain analysis') and key details (annotation sources). There is no wasted verbiage, making it highly concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of domain analysis and the lack of annotations and output schema, the description is incomplete. It doesn't explain what 'enhanced' means, what the output includes, or how it differs from simpler tools, leaving the agent with insufficient context for effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with the single parameter 'accession' clearly documented as a 'UniProt accession number'. The description doesn't add any meaning beyond this, such as format examples or validation details, but the high schema coverage justifies the baseline score of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose as 'Enhanced domain analysis' with specific annotation sources (InterPro, Pfam, SMART), which is a specific verb+resource combination. However, it doesn't explicitly distinguish this from sibling tools like 'get_protein_features' or 'get_protein_info', which might also provide domain-related information, so it doesn't reach the highest differentiation standard.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites, context, or compare it to siblings like 'get_protein_features' or 'get_protein_info', leaving the agent to infer usage based on the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_protein_featuresC

Get functional features and domains for a protein

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionYesUniProt accession number

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions 'Get functional features and domains' but does not specify whether this is a read-only operation, if it requires authentication, what the output format is, or any rate limits. For a tool with no annotation coverage, this leaves significant gaps in understanding its behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It is front-loaded and wastes no space, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description is incomplete for a tool that likely returns complex data (functional features and domains). It does not explain what 'features and domains' entail, the format of the response, or any limitations, leaving the agent with insufficient context for effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with the single parameter 'accession' clearly documented as a 'UniProt accession number'. The description does not add any additional meaning beyond this, such as examples or constraints, so it meets the baseline of 3 where the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get') and target ('functional features and domains for a protein'), making the purpose understandable. However, it does not explicitly differentiate this tool from sibling tools like 'get_protein_domains_detailed' or 'get_protein_info', which might offer overlapping or related functionality, preventing a score of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, such as sibling tools like 'get_protein_domains_detailed' or 'get_protein_info'. It lacks context on prerequisites, exclusions, or specific use cases, leaving the agent to infer usage from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_protein_homologsC

Find homologous proteins across different species

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionYesUniProt accession number
organismNoTarget organism to find homologs in
sizeNoNumber of results to return (1-100, default: 25)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden for behavioral disclosure. It states the action ('Find homologous proteins') but doesn't describe what the tool returns (e.g., list of homologs with scores), performance characteristics, error conditions, or data sources. This is inadequate for a tool with 3 parameters and no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that gets straight to the point with zero wasted words. It's appropriately sized for a straightforward lookup tool and is well-structured for quick comprehension.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 3 parameters, no annotations, and no output schema, the description is insufficient. It doesn't explain what constitutes a 'homolog' in this system, what data is returned, or how results are structured. The agent would be left guessing about the tool's behavior and output format.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents all parameters. The description doesn't add any parameter-specific information beyond what's in the schema, such as explaining how 'organism' should be formatted or what 'homologous' means in this context. Baseline 3 is appropriate when schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('Find') and resource ('homologous proteins'), and specifies the scope ('across different species'). It distinguishes from siblings like 'get_protein_orthologs' by focusing on general homology rather than orthology, but could be more explicit about this distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention when to choose this over 'get_protein_orthologs', 'compare_proteins', or other search tools, nor does it specify prerequisites or exclusions for usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_protein_infoC

Get detailed information for a specific protein by UniProt accession

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionYesUniProt accession number (e.g., P04637)
formatNoOutput format (default: json)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden for behavioral disclosure. It mentions retrieving 'detailed information' but doesn't specify what that includes (e.g., sequence, structure, annotations), whether it's a read-only operation, potential rate limits, or authentication needs. This leaves significant gaps for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the core purpose without unnecessary words. Every element ('Get detailed information', 'specific protein', 'UniProt accession') contributes directly to understanding the tool's function.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of protein data retrieval, no annotations, no output schema, and many sibling tools, the description is insufficient. It doesn't explain what 'detailed information' encompasses, how it differs from specialized sibling tools, or what the return format looks like beyond the parameter options. This leaves too many open questions for effective tool selection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, providing complete parameter documentation. The description adds no additional parameter semantics beyond what's in the schema (e.g., examples of what 'detailed information' includes, format implications). With high schema coverage, the baseline score of 3 is appropriate as the description doesn't compensate but doesn't detract either.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get detailed information') and target resource ('for a specific protein by UniProt accession'), making the purpose immediately understandable. However, it doesn't differentiate from siblings like 'get_protein_sequence' or 'get_protein_structure' that also retrieve protein information but focus on specific aspects.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. With many sibling tools (e.g., 'get_protein_sequence', 'get_protein_structure', 'search_proteins'), the description lacks context about when this general information retrieval is preferred over more specific queries or searches.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_protein_interactionsD

Protein-protein interaction networks

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionYesUniProt accession number

TDQS

D1.8/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It fails to describe any behavioral traits—such as whether this is a read-only query, if it requires authentication, rate limits, or what the output entails (e.g., network data, lists, visualizations). For a tool with no annotation coverage, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single phrase ('Protein-protein interaction networks') that is under-specified, not concise in a helpful way. It lacks structure and front-loading of key information, failing to earn its place with actionable details. This is brevity at the cost of clarity, not effective conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of protein interaction data and the lack of annotations and output schema, the description is incomplete. It does not explain what the tool returns (e.g., network graphs, interaction lists), how results are formatted, or any limitations. For a tool with no structured output information, this leaves critical gaps for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with one parameter ('accession') clearly documented as a 'UniProt accession number'. The description adds no additional meaning beyond the schema, such as format examples or constraints. According to the rules, with high schema coverage (>80%), the baseline is 3, which is appropriate here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Protein-protein interaction networks' is vague and tautological—it essentially restates the tool name 'get_protein_interactions' without specifying the action (e.g., retrieve, analyze, or visualize). It lacks a clear verb and does not distinguish this tool from siblings like 'get_protein_homologs' or 'compare_proteins', which might also involve protein relationships. This falls short of a minimum viable description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. With many sibling tools related to protein data (e.g., 'get_protein_info', 'search_proteins', 'compare_proteins'), the description offers no context, prerequisites, or exclusions. This leaves the agent without direction, making it misleading in a crowded toolset.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_protein_orthologsC

Identify orthologous proteins for evolutionary studies

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionYesUniProt accession number
organismNoTarget organism to find orthologs in
sizeNoNumber of results to return (1-100, default: 25)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool 'identifies' orthologous proteins but doesn't explain how (e.g., algorithm, data sources), what the output looks like, or any limitations (e.g., rate limits, accuracy). This is a significant gap for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence: 'Identify orthologous proteins for evolutionary studies.' It's front-loaded with the core purpose and contains no redundant information, making it highly concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (identifying orthologs involves biological inference) and lack of annotations and output schema, the description is incomplete. It doesn't address behavioral traits, output format, or limitations, which are crucial for an agent to use it effectively in evolutionary studies.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters (accession, organism, size) with descriptions. The description adds no additional parameter semantics beyond what's in the schema, such as explaining relationships between parameters or usage examples. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Identify orthologous proteins for evolutionary studies.' It specifies the action (identify), resource (orthologous proteins), and context (evolutionary studies). However, it doesn't explicitly differentiate from sibling tools like 'get_protein_homologs' or 'get_phylogenetic_info,' which prevents a perfect score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites, exclusions, or compare it to similar tools like 'get_protein_homologs' or 'compare_proteins,' leaving the agent to infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_protein_pathwaysC

Associated biological pathways (KEGG, Reactome)

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionYesUniProt accession number

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden but offers minimal behavioral insight. It implies a read-only operation by mentioning retrieval of pathways, but doesn't disclose rate limits, authentication needs, error handling, or output format (e.g., list vs. detailed data). This is inadequate for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient phrase with no wasted words. It's front-loaded with the core purpose, though it could be more structured (e.g., starting with a verb). Every word earns its place, but it's borderline under-specified rather than concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and a single parameter with full schema coverage, the description is incomplete. It doesn't explain what 'associated' means (e.g., direct vs. inferred pathways), the scope of results, or how KEGG/Reactome data is presented. For a biological data tool, this leaves significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with the single parameter 'accession' documented as a UniProt accession number. The description adds no additional meaning about the parameter (e.g., format examples, validation rules). Baseline 3 is appropriate since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Associated biological pathways (KEGG, Reactome)' states what the tool retrieves (pathways) and mentions specific databases, but it lacks a clear verb and doesn't distinguish from siblings like 'get_external_references' or 'search_by_function'. It's vague about whether this is a lookup or search operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing a valid accession), exclusions, or how it differs from siblings such as 'get_external_references' or 'search_by_function' that might overlap with pathway-related queries.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_protein_sequenceC

Get the amino acid sequence for a protein

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionYesUniProt accession number
formatNoOutput format (default: fasta)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It states what the tool does but reveals nothing about behavioral traits: no information about rate limits, authentication requirements, error conditions, response format details beyond format parameter, or whether this is a read-only operation. The description is minimal and lacks essential operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence that states exactly what the tool does with zero wasted words. It's appropriately sized for a simple retrieval tool and is perfectly front-loaded with the core functionality.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description is insufficiently complete. For a tool with 2 parameters and no structured output documentation, the description should provide more context about what the response contains, error conditions, or usage constraints. It leaves too much undefined for proper agent understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters thoroughly. The description adds no additional parameter semantics beyond what's in the schema - it doesn't explain accession format requirements or when to choose different output formats. This meets the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get') and resource ('amino acid sequence for a protein'), making the purpose immediately understandable. It distinguishes this from siblings like 'get_protein_info' or 'get_protein_structure' by specifying the sequence aspect. However, it doesn't explicitly differentiate from 'batch_protein_lookup' which might also retrieve sequences.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. With siblings like 'get_protein_info' (which might include sequence), 'batch_protein_lookup', and 'search_proteins', there's no indication of when this specific sequence-fetching tool is preferred or what its limitations are.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_protein_structureC

Retrieve 3D structure information from PDB references

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionYesUniProt accession number

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions retrieving information but lacks details on permissions, rate limits, error handling, or response format. For a tool with no annotation coverage, this is a significant gap in transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It is appropriately sized and front-loaded, making it easy to understand quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description is incomplete. It does not explain what '3D structure information' entails (e.g., coordinates, formats) or behavioral aspects like data sources or limitations, leaving gaps for effective tool use in a complex domain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with the parameter 'accession' documented as a 'UniProt accession number'. The description does not add any additional meaning beyond this, such as format examples or constraints, so it meets the baseline for high schema coverage without compensating further.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Retrieve') and resource ('3D structure information from PDB references'), making the purpose understandable. However, it does not explicitly differentiate this tool from siblings like 'get_protein_info' or 'get_protein_features', which might also provide structural data, leaving some ambiguity about uniqueness.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. With many sibling tools available, such as 'get_protein_info' or 'search_proteins', there is no indication of specific contexts, prerequisites, or exclusions for choosing this tool over others.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_protein_variantsC

Disease-associated variants and mutations

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionYesUniProt accession number

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It doesn't disclose behavioral traits such as whether this is a read-only operation, if it requires authentication, rate limits, or what the output format looks like. For a tool with no annotations, this is a significant gap in transparency about how it behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single phrase 'Disease-associated variants and mutations', which is concise and front-loaded with the core purpose. However, it's under-specified rather than efficiently informative, lacking necessary details for a tool with no annotations, which slightly reduces its effectiveness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (retrieving disease-associated variants), lack of annotations, and no output schema, the description is incomplete. It doesn't explain return values, error handling, or behavioral context, leaving gaps that could hinder an AI agent's ability to use it correctly in a broader context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with the single parameter 'accession' clearly documented as a UniProt accession number. The description adds no additional meaning beyond the schema, such as format examples or constraints. With high schema coverage, the baseline score of 3 is appropriate as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Disease-associated variants and mutations' states what the tool retrieves but is vague about the action. It mentions the resource (protein variants/mutations) but lacks a specific verb like 'retrieve', 'fetch', or 'list'. It doesn't distinguish from siblings like 'get_protein_features' or 'get_protein_info', which might also relate to variants.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. It doesn't mention prerequisites, exclusions, or compare to siblings like 'search_by_function' or 'get_protein_homologs', which could also involve variant data. The description implies a specific focus on disease-associated variants but doesn't clarify context or limitations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_taxonomy_infoC

Detailed taxonomic information for organisms

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionYesUniProt accession number

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool provides 'detailed taxonomic information' but doesn't describe what that includes (e.g., lineage, ranks, sources), whether it's a read-only operation, potential rate limits, or error handling. The description is too vague to inform the agent adequately about behavioral traits beyond the basic purpose.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence: 'Detailed taxonomic information for organisms'. It's front-loaded with the core purpose and avoids unnecessary words. However, it could be more structured by including key details like the resource type or usage context, but it earns its place by being clear and concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of a tool with one parameter but no annotations or output schema, the description is incomplete. It doesn't explain what 'detailed taxonomic information' entails, how it's returned, or any behavioral aspects. For a tool that likely returns structured data, the description should provide more context to compensate for the lack of output schema and annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with the parameter 'accession' clearly documented as a 'UniProt accession number'. The description doesn't add any meaning beyond this, as it doesn't explain parameter usage or constraints. With high schema coverage, the baseline score of 3 is appropriate, as the schema handles the parameter documentation effectively.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Detailed taxonomic information for organisms' states what the tool does but is vague about the specific resource and scope. It mentions 'taxonomic information' but doesn't specify that it retrieves this for proteins via UniProt accession numbers, unlike siblings like 'get_phylogenetic_info' or 'search_by_taxonomy' which might overlap in purpose. It distinguishes minimally by focusing on 'detailed' information but lacks explicit differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites like needing a UniProt accession, nor does it compare to siblings such as 'get_phylogenetic_info' or 'search_by_taxonomy', which might offer similar or related data. Usage is implied only by the parameter, but no explicit context or exclusions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_by_functionC

Search proteins by GO terms or functional annotations

ParametersJSON Schema
NameRequiredDescriptionDefault
goTermNoGene Ontology term (e.g., GO:0005524)
functionNoFunctional description or keyword
organismNoOrganism name or taxonomy ID to filter results
sizeNoNumber of results to return (1-500, default: 25)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool searches proteins but doesn't describe how results are returned (e.g., format, pagination), potential limitations (e.g., rate limits, data freshness), or error conditions. For a search tool with zero annotation coverage, this leaves significant gaps in understanding its behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero waste. It's front-loaded with the core purpose and uses clear terminology. Every word contributes directly to understanding the tool's function without redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of a search tool with 4 parameters, no annotations, and no output schema, the description is incomplete. It doesn't explain what the tool returns (e.g., protein IDs, annotations), how results are structured, or any behavioral traits like performance or constraints. The high schema coverage helps with parameters, but overall context is lacking.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters thoroughly. The description adds no additional parameter semantics beyond implying that 'GO terms or functional annotations' map to the 'goTerm' and 'function' parameters. It doesn't clarify parameter interactions or provide examples, so it meets the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Search proteins by GO terms or functional annotations.' It specifies the verb ('Search'), resource ('proteins'), and search criteria ('GO terms or functional annotations'). However, it doesn't explicitly differentiate from sibling tools like 'search_by_gene' or 'search_proteins,' which likely have overlapping purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention sibling tools like 'search_by_gene' or 'search_proteins,' nor does it specify prerequisites, exclusions, or contextual cues for selection. Usage is implied but not articulated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_by_geneC

Search for proteins by gene name or symbol

ParametersJSON Schema
NameRequiredDescriptionDefault
geneYesGene name or symbol (e.g., BRCA1, INS)
organismNoOrganism name or taxonomy ID to filter results
sizeNoNumber of results to return (1-500, default: 25)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. While 'Search' implies a read-only operation, the description doesn't address important behavioral aspects like whether this is a fuzzy or exact match search, what format results are returned in, whether there are rate limits, authentication requirements, or what happens when no matches are found. For a search tool with zero annotation coverage, this is insufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that immediately conveys the core functionality without unnecessary words. It's appropriately sized for a search tool and front-loads the essential information, making it easy for an agent to quickly understand what the tool does.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description is incomplete for a search tool with 3 parameters. It doesn't explain what kind of results are returned (protein IDs, names, sequences?), how results are formatted, whether there's pagination, or what happens with partial/no matches. For a tool that likely returns complex protein data, more context about the output would be helpful.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description mentions 'gene name or symbol' which aligns with the 'gene' parameter in the schema, but doesn't add meaningful semantic context beyond what the 100% schema coverage already provides. The schema descriptions fully document each parameter's purpose, constraints, and examples, so the description adds minimal additional value regarding parameter meaning or usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Search for proteins') and the target resource ('by gene name or symbol'), making the purpose immediately understandable. However, it doesn't explicitly differentiate this tool from sibling tools like 'search_by_function' or 'search_proteins', which would require more specific language about when to use gene-based searching versus other search methods.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. With multiple search-related sibling tools (search_by_function, search_by_localization, search_by_taxonomy, search_proteins), there's no indication of when gene-based searching is appropriate versus other search methods or what distinguishes this from the generic 'search_proteins' tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_by_localizationB

Find proteins by subcellular localization

ParametersJSON Schema
NameRequiredDescriptionDefault
localizationYesSubcellular localization (e.g., nucleus, mitochondria)
organismNoOrganism name or taxonomy ID to filter results
sizeNoNumber of results to return (1-500, default: 25)

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden but only states the basic function. It doesn't disclose behavioral traits such as whether this is a read-only operation, performance characteristics, rate limits, or what the output format looks like (no output schema exists).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero wasted words. It's front-loaded with the core purpose and appropriately sized for a simple search tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a search tool with 3 parameters and 100% schema coverage but no output schema, the description is minimally adequate. It states what the tool does but lacks context about output format, result limitations, or how it differs from sibling tools, leaving gaps for an AI agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters thoroughly. The description doesn't add any meaning beyond what the schema provides, such as examples of localization values beyond 'nucleus, mitochondria' or organism naming conventions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Find') and resource ('proteins') with a specific criterion ('by subcellular localization'). It distinguishes from siblings like 'search_by_function' or 'search_by_taxonomy' by focusing on localization, but doesn't explicitly contrast with 'search_proteins' which might be more general.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like 'search_proteins' or 'search_by_function'. The description implies usage for localization-based queries but doesn't specify exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_by_taxonomyC

Search by detailed taxonomic classification

ParametersJSON Schema
NameRequiredDescriptionDefault
taxonomyIdNoNCBI taxonomy ID
taxonomyNameNoTaxonomic name (e.g., Mammalia, Bacteria)
sizeNoNumber of results to return (1-500, default: 25)

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden for behavioral disclosure. The description only mentions 'search' without specifying what is returned (e.g., protein records, sequences), whether results are paginated, if authentication is required, or any rate limits. For a search tool with zero annotation coverage, this is a significant gap in behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient phrase with no wasted words. It's appropriately sized for the tool's complexity and front-loads the core purpose without unnecessary elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description is incomplete. It doesn't explain what the search returns (e.g., protein data, sequences), how results are structured, or any behavioral traits. For a search tool with 3 parameters and no structured output information, the description should provide more context about the search scope and results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with clear parameter descriptions in the schema (e.g., 'NCBI taxonomy ID', 'Taxonomic name', 'Number of results to return'). The description adds no additional parameter semantics beyond what the schema already provides, so it meets the baseline of 3 for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Search by detailed taxonomic classification' states the action (search) and resource domain (taxonomic classification), but is vague about what exactly is being searched (proteins, sequences, etc.) and doesn't distinguish from sibling tools like 'search_by_function', 'search_by_gene', or 'get_taxonomy_info'. It provides basic purpose but lacks specificity about the search target.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like 'search_by_function', 'search_by_gene', 'get_taxonomy_info', or 'search_proteins'. The description doesn't mention prerequisites, exclusions, or comparative use cases, leaving the agent to infer usage from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_proteinsC

Search UniProt database for proteins by name, keyword, or organism

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesSearch query (protein name, keyword, or complex search)
organismNoOrganism name or taxonomy ID to filter results
sizeNoNumber of results to return (1-500, default: 25)
formatNoOutput format (default: json)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It mentions searching a database but lacks details on rate limits, authentication needs, pagination, error handling, or what the search returns (e.g., list of proteins with basic info). This is a significant gap for a search tool with no structured behavioral hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero waste—it directly states the tool's purpose without unnecessary words. It is appropriately sized and front-loaded, making it easy to understand at a glance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description is incomplete. It doesn't explain what the search returns (e.g., protein IDs, names, sequences), potential limitations, or how results are structured. For a search tool with 4 parameters and many siblings, more context is needed to guide effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents all parameters. The description adds minimal value beyond the schema by hinting at search criteria ('by name, keyword, or organism'), but doesn't explain parameter interactions or provide examples. Baseline 3 is appropriate as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('search') and target resource ('UniProt database for proteins'), specifying search criteria ('by name, keyword, or organism'). It distinguishes from siblings like 'search_by_function' or 'search_by_gene' by mentioning general search terms, though not explicitly contrasting them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives like 'advanced_search' or 'search_by_function' is provided. The description implies usage for basic protein searches but lacks context on prerequisites, exclusions, or comparisons to sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_accessionC

Check if accession numbers are valid

ParametersJSON Schema
NameRequiredDescriptionDefault
accessionYesUniProt accession number to validate

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool checks validity but does not explain what 'valid' means (e.g., format, existence in a database), potential error conditions, rate limits, or authentication needs. For a validation tool with zero annotation coverage, this is a significant gap in transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero waste. It is front-loaded and appropriately sized for a simple tool, making it easy for an agent to parse quickly without unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (1 parameter, no output schema, no annotations), the description is incomplete. It lacks details on what constitutes validity, potential return values (e.g., boolean, error messages), or behavioral context. While concise, it does not provide enough information for an agent to fully understand the tool's operation and outcomes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with the parameter 'accession' documented as 'UniProt accession number to validate'. The description adds no additional meaning beyond this, such as format examples or validation criteria. Since the schema does the heavy lifting, the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Check if accession numbers are valid' clearly states the tool's purpose with a specific verb ('Check') and resource ('accession numbers'), but it does not distinguish this from sibling tools. While siblings like 'batch_protein_lookup' or 'get_protein_info' might involve accession numbers, this tool's specific validation focus is implied but not explicitly contrasted.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention prerequisites, context (e.g., before other operations), or exclusions, and it fails to reference sibling tools like 'batch_protein_lookup' that might overlap in functionality. This leaves the agent without clear usage direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 26 tool updates
    • First observedadvanced_search
    • First observedanalyze_sequence_composition
    • First observedbatch_protein_lookup
    • First observedcompare_proteins
    • First observedexport_protein_data
    • First observedget_annotation_confidence
    • First observedget_external_references
    • First observedget_literature_references
    • First observedget_phylogenetic_info
    • First observedget_protein_domains_detailed
    • First observedget_protein_features
    • First observedget_protein_homologs
    • First observedget_protein_info
    • First observedget_protein_interactions
    • First observedget_protein_orthologs
    • First observedget_protein_pathways
    • First observedget_protein_sequence
    • First observedget_protein_structure
    • First observedget_protein_variants
    • First observedget_taxonomy_info
    • First observedsearch_by_function
    • First observedsearch_by_gene
    • First observedsearch_by_localization
    • First observedsearch_by_taxonomy
    • First observedsearch_proteins
    • First observedvalidate_accession

TDQS

C2.9/5.0

Scored across 26 tools

Disambiguation4/5

Most tools have distinct purposes targeting specific UniProt data aspects, but some overlap exists. For example, 'get_protein_info' and 'get_protein_sequence' could be confused as both retrieve protein data, though their descriptions clarify the distinction. Overall, the set is well-organized with clear boundaries for most tools.

Naming Consistency5/5

Tool names follow a highly consistent verb_noun pattern throughout, such as 'get_protein_info', 'search_by_function', and 'analyze_sequence_composition'. This predictability makes it easy for agents to understand and select tools without confusion, enhancing usability.

Tool Count3/5

With 26 tools, the count is borderline high for a single server, potentially overwhelming for agents. While UniProt is a complex domain, this many tools might indicate over-specialization or fragmentation, making it harder to navigate efficiently.

Completeness5/5

The tool set provides comprehensive coverage for UniProt data access, including search, retrieval, analysis, and export functions. It covers all major aspects like sequences, structures, interactions, and annotations, with no obvious gaps for typical agent workflows in this domain.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers