Skip to main content
Glama
liueic

PubChem Chemical Safety MCP Server

by liueic

PubChem Chemical Safety MCP Server

一个基于 Model Context Protocol (MCP) 的化学安全信息服务器,用于从化合物名称或 CID 自动获取毒理、GHS 安全分类、化学性质等信息。

功能特性

  • 获取化合物基础属性信息(分子式、分子量、IUPAC名称等)

  • 获取 GHS 安全分类信息(信号词、象形图、危害声明)

  • 获取毒性实验数据(LD50、LC50等)

  • 支持批量查询和缓存机制

  • 基于 MCP 协议,可与 Claude Desktop 等 AI 客户端集成

  • 支持代理访问,解决网络连接问题

Related MCP server: pubchem-mcp-server

技术栈

  • 协议: Model Context Protocol (MCP)

  • 语言: Python 3.10+

  • 依赖管理: uv

  • 数据源: PubChem REST API

  • 缓存: 本地文件缓存

  • HTTP客户端: aiohttp (支持代理和重试机制)

安装与运行

1. 安装依赖

uv sync

2. 运行 MCP 服务器

uv run python -m pubchem_mcp.mcp_server

3. 测试服务器

uv run verify_mcp.py

MCP 工具

服务器提供以下 3 个 MCP 工具:

1. get_compound_info

获取化合物基础信息

  • 参数: name (化合物名称)

  • 返回: CID、分子式、分子量、IUPAC名称、SMILES等

  • 示例: get_compound_info("aspirin")get_compound_info("阿司匹林")

2. get_safety_info

获取 GHS 安全分类信息

  • 参数: cid (PubChem化合物ID)

  • 返回: 信号词、GHS象形图、危害声明、预防措施等

  • 示例: get_safety_info(2244) (阿司匹林的CID)

3. get_toxicity_data

获取毒性实验数据

  • 参数: cid (PubChem化合物ID)

  • 返回: 急性毒性、生态毒性、致癌性、生殖毒性等详细数据

  • 示例: get_toxicity_data(2244) (阿司匹林的CID)

使用示例

Python 客户端示例

import asyncio
from mcp.client.stdio import stdio_client
from mcp import ClientSession, StdioServerParameters

async def main():
    server_params = StdioServerParameters(
        command='uv',
        args=['--directory', '/path/to/project', 'run', 'python', '-m', 'pubchem_mcp.mcp_server']
    )
    
    async with stdio_client(server_params) as (stdio, write):
        async with ClientSession(stdio, write) as session:
            await session.initialize()
            
            # 获取化合物信息
            result = await session.call_tool('get_compound_info', {
                'name': 'aspirin'
            })
            print(result.content[0].text)

asyncio.run(main())

Claude Desktop 中使用

在 Claude Desktop 中可以直接使用自然语言查询:

请查询阿司匹林的安全信息
获取咖啡因的毒性数据

调试工具

使用 MCP Inspector 进行调试:

npx -y @modelcontextprotocol/inspector uv run python -m pubchem_mcp.mcp_server

项目结构

pubchem_mcp/
├── mcp_server.py          # MCP 服务器主文件
├── services/
│   ├── pubchem_client.py  # PubChem API 客户端(支持代理和重试)
│   ├── cache_service.py   # 缓存服务
│   └── pubchem_service.py # PubChem 服务层
├── models/
│   └── schemas.py         # 数据模型定义
└── api/
    └── routes.py          # API 路由(可选)

tests/                     # 测试文件
manage_cache.py            # 缓存管理工具
verify_mcp.py             # MCP 验证工具
claude_desktop_config.json # Claude Desktop 配置示例

配置选项

  • CACHE_DIR: 缓存目录路径(默认 .cache

  • PUBCHEM_RATE_LIMIT: API 请求限制(默认 5 req/s)

  • https_proxy: HTTPS 代理设置

  • http_proxy: HTTP 代理设置

缓存管理

服务器使用本地文件缓存来提高性能:

  • 缓存位置: .cache/ 目录

  • 缓存策略:

    • 化合物信息缓存2小时

    • 安全信息和毒性数据缓存1小时

    • 过期文件自动删除

  • 缓存管理: 运行 uv run manage_cache.py 查看缓存统计

网络配置

代理设置

项目支持通过环境变量设置代理:

export https_proxy=http://127.0.0.1:10808
export http_proxy=http://127.0.0.1:10808

重试机制

  • 自动重试503错误(服务器繁忙)

  • 递增等待时间避免过于频繁的请求

  • 最多重试3次

请求头优化

  • 使用完整的浏览器User-Agent

  • 添加必要的HTTP请求头(Accept、Referer等)

  • 支持gzip压缩

测试化合物

以下化合物已测试可用:

  • aspirin (阿司匹林) - CID: 2244

  • caffeine (咖啡因) - CID: 2519

  • water (水) - CID: 962

  • ethanol (乙醇) - CID: 702

  • benzene (苯) - CID: 241

故障排除

常见问题

  1. 503 错误: 服务器繁忙,会自动重试

  2. 网络连接失败: 检查代理设置是否正确

  3. 化合物未找到: 尝试使用英文名称或化学式

日志查看

服务器会输出详细的日志信息,包括:

  • 请求状态

  • 重试信息

  • 错误详情

开发

运行测试

uv run pytest tests/

代码格式化

uv run black .
uv run isort .

类型检查

uv run mypy pubchem_mcp/

许可证

MIT License

贡献

欢迎提交 Issue 和 Pull Request!

Available Tools

3 tools
get_compound_infoC

获取化合物基础信息

Args: name: 化合物名称

Returns: 包含CID、分子式、分子量等基础信息的字典

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It mentions returning a dictionary with CID, molecular formula, and molecular weight, which adds some behavioral context about output format. However, it lacks details on permissions, rate limits, error handling, or whether it's a read-only operation, which is insufficient for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with a clear purpose statement followed by Args and Returns sections. However, the Args section is redundant with the input schema (repeating 'name' without added value), and the structure could be more front-loaded by integrating parameter hints into the main description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema (which should cover return values), the description's mention of return format is somewhat redundant. It provides minimal context for a simple lookup tool but lacks guidance on usage versus siblings and behavioral details, making it incomplete despite the output schema support.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds that the 'name' parameter is a compound name, providing basic semantics beyond the schema's generic 'Name' title. However, it doesn't specify format constraints (e.g., IUPAC name, common name) or examples, leaving gaps in parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool '获取化合物基础信息' (get compound basic information), which is a clear verb+resource combination. However, it doesn't differentiate from sibling tools like get_safety_info or get_toxicity_data, leaving the scope ambiguous about what distinguishes 'basic information' from safety or toxicity data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like get_safety_info or get_toxicity_data. The description implies it's for basic compound information but doesn't specify use cases, prerequisites, or exclusions, leaving the agent to guess based on tool names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_safety_infoB

获取GHS安全分类信息

Args: cid: PubChem化合物ID

Returns: 包含信号词、GHS象形图、危害声明等安全信息的字典

ParametersJSON Schema
NameRequiredDescriptionDefault
cidYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states this is a retrieval operation ('获取' - get) which implies read-only behavior, but doesn't explicitly confirm this or mention any side effects, authentication requirements, rate limits, or error conditions. The description provides basic functional information but lacks important behavioral context for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately concise with clear sectioning: a purpose statement followed by Args and Returns sections. Each sentence serves a specific purpose - stating what the tool does, explaining the parameter, and describing the return value. There's no redundant information, though the structure could be slightly more polished.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema (which handles return value documentation) and only one parameter, the description provides adequate context for the basic operation. However, with no annotations and sibling tools that appear related, the description should ideally provide more guidance on when to use this specific tool versus alternatives. It covers the essentials but leaves gaps in usage context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds significant value beyond the input schema, which has 0% description coverage. It explains that 'cid' represents a 'PubChem化合物ID' (PubChem compound ID), providing crucial semantic context that the schema lacks. With only one parameter and the schema providing no descriptions, the tool description effectively compensates for the schema's deficiencies by explaining what the parameter represents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: '获取GHS安全分类信息' (Get GHS safety classification information). It specifies the resource (GHS safety classification) and the action (get/retrieve). However, it doesn't explicitly differentiate from sibling tools like 'get_compound_info' or 'get_toxicity_data' - all appear to retrieve chemical information but for different aspects.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus the sibling tools. While it mentions the specific parameter (PubChem compound ID), it doesn't explain when one would need safety information versus general compound information or toxicity data. There's no context about prerequisites, alternatives, or exclusion criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_toxicity_dataC

获取毒性实验数据

Args: cid: PubChem化合物ID

Returns: 包含急性毒性、生态毒性、致癌性等毒性数据的字典

ParametersJSON Schema
NameRequiredDescriptionDefault
cidYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the return format ('包含急性毒性、生态毒性、致癌性等毒性数据的字典' - dictionary containing acute toxicity, ecotoxicity, carcinogenicity, etc.) but doesn't address important behavioral aspects like whether this is a read-only operation, potential rate limits, authentication requirements, error conditions, or data freshness. For a data retrieval tool with zero annotation coverage, this leaves significant gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately concise with three clear sections: purpose statement, parameter documentation, and return value description. Each sentence earns its place by providing essential information. The structure with labeled 'Args:' and 'Returns:' sections is helpful, though the formatting could be more consistent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema (which presumably documents the return structure), the description doesn't need to fully explain return values. However, for a toxicity data retrieval tool with no annotations and only basic parameter documentation, the description should provide more context about data sources, limitations, or typical use cases. It's minimally adequate but leaves important questions unanswered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description explicitly documents the single parameter 'cid' as 'PubChem化合物ID' (PubChem compound ID), which adds crucial semantic meaning beyond the schema's generic 'Cid' title and integer type. With 0% schema description coverage, this parameter documentation is essential. However, it doesn't provide format examples, valid ranges, or handling of invalid inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose as '获取毒性实验数据' (get toxicity experiment data), which is a specific verb+resource combination. It distinguishes from sibling tools like 'get_compound_info' and 'get_safety_info' by focusing specifically on toxicity data. However, it doesn't explicitly differentiate itself from 'get_safety_info' which might overlap in scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'get_safety_info'. There's no mention of prerequisites, appropriate contexts, or exclusions. The only implied usage is when toxicity data for a specific compound is needed, but this is too vague for effective tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observedget_compound_info
    • First observedget_safety_info
    • First observedget_toxicity_data

TDQS

B3.2/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: get_compound_info retrieves basic chemical properties, get_safety_info focuses on GHS hazard classifications, and get_toxicity_data provides experimental toxicity data. There is no overlap in functionality, and the descriptions clearly differentiate their domains (basic info vs. safety vs. toxicity).

Naming Consistency5/5

All tools follow a consistent verb_noun pattern with 'get_' prefix and snake_case naming (get_compound_info, get_safety_info, get_toxicity_data). This predictable structure makes it easy for agents to understand and navigate the toolset without confusion.

Tool Count3/5

With only 3 tools, the set feels thin for a chemical safety domain that might benefit from additional operations like searching compounds, batch queries, or regulatory data. While the tools cover core areas, the low count limits comprehensive coverage and could require agents to work around missing functionality.

Completeness3/5

The tools cover key aspects (basic info, safety, toxicity) but leave notable gaps. There is no search or lookup tool to find compounds by criteria, no batch operations for efficiency, and no update/delete capabilities (though less critical here). Agents may struggle to initiate workflows without a way to discover compounds beyond direct name input.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    Enables comprehensive access to PubChem's chemical database with over 110 million compounds. Supports chemical searches, structure analysis, bioactivity data, safety information, and molecular property calculations through 30 specialized tools.
    30
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Enables AI agents and applications to search, retrieve, and analyze chemical compounds, substances, and bioassays from PubChem's vast chemical information database through comprehensive tools for chemical research and discovery.
    10
    133
    9
    Apache 2.0
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to search and retrieve chemical compound information, structures, and physical properties from the PubChem database. It supports querying via compound names, SMILES notation, or CIDs to provide detailed molecular data for chemical analysis.
    5
    -
  • F
    license
    B
    quality
    F
    maintenance
    Extracts basic chemical information and drug data from the PubChem API. It enables users to retrieve molecular details such as SMILES, IUPAC names, molecular formulas, and synonyms for specific compounds.
    3
    11
    -