PubChem Chemical Safety MCP Server
Provides chemical safety information retrieval capabilities through PubChem REST API, enabling access to compound properties, GHS safety classifications, and toxicity data for chemical compounds
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@PubChem Chemical Safety MCP Serverget safety info for aspirin"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
PubChem Chemical Safety MCP Server
一个基于 Model Context Protocol (MCP) 的化学安全信息服务器,用于从化合物名称或 CID 自动获取毒理、GHS 安全分类、化学性质等信息。
功能特性
获取化合物基础属性信息(分子式、分子量、IUPAC名称等)
获取 GHS 安全分类信息(信号词、象形图、危害声明)
获取毒性实验数据(LD50、LC50等)
支持批量查询和缓存机制
基于 MCP 协议,可与 Claude Desktop 等 AI 客户端集成
支持代理访问,解决网络连接问题
Related MCP server: pubchem-mcp-server
技术栈
协议: Model Context Protocol (MCP)
语言: Python 3.10+
依赖管理: uv
数据源: PubChem REST API
缓存: 本地文件缓存
HTTP客户端: aiohttp (支持代理和重试机制)
安装与运行
1. 安装依赖
uv sync2. 运行 MCP 服务器
uv run python -m pubchem_mcp.mcp_server3. 测试服务器
uv run verify_mcp.pyMCP 工具
服务器提供以下 3 个 MCP 工具:
1. get_compound_info
获取化合物基础信息
参数:
name(化合物名称)返回: CID、分子式、分子量、IUPAC名称、SMILES等
示例:
get_compound_info("aspirin")或get_compound_info("阿司匹林")
2. get_safety_info
获取 GHS 安全分类信息
参数:
cid(PubChem化合物ID)返回: 信号词、GHS象形图、危害声明、预防措施等
示例:
get_safety_info(2244)(阿司匹林的CID)
3. get_toxicity_data
获取毒性实验数据
参数:
cid(PubChem化合物ID)返回: 急性毒性、生态毒性、致癌性、生殖毒性等详细数据
示例:
get_toxicity_data(2244)(阿司匹林的CID)
使用示例
Python 客户端示例
import asyncio
from mcp.client.stdio import stdio_client
from mcp import ClientSession, StdioServerParameters
async def main():
server_params = StdioServerParameters(
command='uv',
args=['--directory', '/path/to/project', 'run', 'python', '-m', 'pubchem_mcp.mcp_server']
)
async with stdio_client(server_params) as (stdio, write):
async with ClientSession(stdio, write) as session:
await session.initialize()
# 获取化合物信息
result = await session.call_tool('get_compound_info', {
'name': 'aspirin'
})
print(result.content[0].text)
asyncio.run(main())Claude Desktop 中使用
在 Claude Desktop 中可以直接使用自然语言查询:
请查询阿司匹林的安全信息获取咖啡因的毒性数据调试工具
使用 MCP Inspector 进行调试:
npx -y @modelcontextprotocol/inspector uv run python -m pubchem_mcp.mcp_server项目结构
pubchem_mcp/
├── mcp_server.py # MCP 服务器主文件
├── services/
│ ├── pubchem_client.py # PubChem API 客户端(支持代理和重试)
│ ├── cache_service.py # 缓存服务
│ └── pubchem_service.py # PubChem 服务层
├── models/
│ └── schemas.py # 数据模型定义
└── api/
└── routes.py # API 路由(可选)
tests/ # 测试文件
manage_cache.py # 缓存管理工具
verify_mcp.py # MCP 验证工具
claude_desktop_config.json # Claude Desktop 配置示例配置选项
CACHE_DIR: 缓存目录路径(默认.cache)PUBCHEM_RATE_LIMIT: API 请求限制(默认 5 req/s)https_proxy: HTTPS 代理设置http_proxy: HTTP 代理设置
缓存管理
服务器使用本地文件缓存来提高性能:
缓存位置:
.cache/目录缓存策略:
化合物信息缓存2小时
安全信息和毒性数据缓存1小时
过期文件自动删除
缓存管理: 运行
uv run manage_cache.py查看缓存统计
网络配置
代理设置
项目支持通过环境变量设置代理:
export https_proxy=http://127.0.0.1:10808
export http_proxy=http://127.0.0.1:10808重试机制
自动重试503错误(服务器繁忙)
递增等待时间避免过于频繁的请求
最多重试3次
请求头优化
使用完整的浏览器User-Agent
添加必要的HTTP请求头(Accept、Referer等)
支持gzip压缩
测试化合物
以下化合物已测试可用:
aspirin (阿司匹林) - CID: 2244
caffeine (咖啡因) - CID: 2519
water (水) - CID: 962
ethanol (乙醇) - CID: 702
benzene (苯) - CID: 241
故障排除
常见问题
503 错误: 服务器繁忙,会自动重试
网络连接失败: 检查代理设置是否正确
化合物未找到: 尝试使用英文名称或化学式
日志查看
服务器会输出详细的日志信息,包括:
请求状态
重试信息
错误详情
开发
运行测试
uv run pytest tests/代码格式化
uv run black .
uv run isort .类型检查
uv run mypy pubchem_mcp/许可证
MIT License
贡献
欢迎提交 Issue 和 Pull Request!
Available Tools
3 toolsget_compound_infoC
获取化合物基础信息
Args: name: 化合物名称
Returns: 包含CID、分子式、分子量等基础信息的字典
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It mentions returning a dictionary with CID, molecular formula, and molecular weight, which adds some behavioral context about output format. However, it lacks details on permissions, rate limits, error handling, or whether it's a read-only operation, which is insufficient for a tool with no annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with a clear purpose statement followed by Args and Returns sections. However, the Args section is redundant with the input schema (repeating 'name' without added value), and the structure could be more front-loaded by integrating parameter hints into the main description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has an output schema (which should cover return values), the description's mention of return format is somewhat redundant. It provides minimal context for a simple lookup tool but lacks guidance on usage versus siblings and behavioral details, making it incomplete despite the output schema support.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds that the 'name' parameter is a compound name, providing basic semantics beyond the schema's generic 'Name' title. However, it doesn't specify format constraints (e.g., IUPAC name, common name) or examples, leaving gaps in parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool '获取化合物基础信息' (get compound basic information), which is a clear verb+resource combination. However, it doesn't differentiate from sibling tools like get_safety_info or get_toxicity_data, leaving the scope ambiguous about what distinguishes 'basic information' from safety or toxicity data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like get_safety_info or get_toxicity_data. The description implies it's for basic compound information but doesn't specify use cases, prerequisites, or exclusions, leaving the agent to guess based on tool names alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_safety_infoB
获取GHS安全分类信息
Args: cid: PubChem化合物ID
Returns: 包含信号词、GHS象形图、危害声明等安全信息的字典
| Name | Required | Description | Default |
|---|---|---|---|
| cid | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states this is a retrieval operation ('获取' - get) which implies read-only behavior, but doesn't explicitly confirm this or mention any side effects, authentication requirements, rate limits, or error conditions. The description provides basic functional information but lacks important behavioral context for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately concise with clear sectioning: a purpose statement followed by Args and Returns sections. Each sentence serves a specific purpose - stating what the tool does, explaining the parameter, and describing the return value. There's no redundant information, though the structure could be slightly more polished.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has an output schema (which handles return value documentation) and only one parameter, the description provides adequate context for the basic operation. However, with no annotations and sibling tools that appear related, the description should ideally provide more guidance on when to use this specific tool versus alternatives. It covers the essentials but leaves gaps in usage context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds significant value beyond the input schema, which has 0% description coverage. It explains that 'cid' represents a 'PubChem化合物ID' (PubChem compound ID), providing crucial semantic context that the schema lacks. With only one parameter and the schema providing no descriptions, the tool description effectively compensates for the schema's deficiencies by explaining what the parameter represents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: '获取GHS安全分类信息' (Get GHS safety classification information). It specifies the resource (GHS safety classification) and the action (get/retrieve). However, it doesn't explicitly differentiate from sibling tools like 'get_compound_info' or 'get_toxicity_data' - all appear to retrieve chemical information but for different aspects.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus the sibling tools. While it mentions the specific parameter (PubChem compound ID), it doesn't explain when one would need safety information versus general compound information or toxicity data. There's no context about prerequisites, alternatives, or exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_toxicity_dataC
获取毒性实验数据
Args: cid: PubChem化合物ID
Returns: 包含急性毒性、生态毒性、致癌性等毒性数据的字典
| Name | Required | Description | Default |
|---|---|---|---|
| cid | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the return format ('包含急性毒性、生态毒性、致癌性等毒性数据的字典' - dictionary containing acute toxicity, ecotoxicity, carcinogenicity, etc.) but doesn't address important behavioral aspects like whether this is a read-only operation, potential rate limits, authentication requirements, error conditions, or data freshness. For a data retrieval tool with zero annotation coverage, this leaves significant gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately concise with three clear sections: purpose statement, parameter documentation, and return value description. Each sentence earns its place by providing essential information. The structure with labeled 'Args:' and 'Returns:' sections is helpful, though the formatting could be more consistent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has an output schema (which presumably documents the return structure), the description doesn't need to fully explain return values. However, for a toxicity data retrieval tool with no annotations and only basic parameter documentation, the description should provide more context about data sources, limitations, or typical use cases. It's minimally adequate but leaves important questions unanswered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description explicitly documents the single parameter 'cid' as 'PubChem化合物ID' (PubChem compound ID), which adds crucial semantic meaning beyond the schema's generic 'Cid' title and integer type. With 0% schema description coverage, this parameter documentation is essential. However, it doesn't provide format examples, valid ranges, or handling of invalid inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose as '获取毒性实验数据' (get toxicity experiment data), which is a specific verb+resource combination. It distinguishes from sibling tools like 'get_compound_info' and 'get_safety_info' by focusing specifically on toxicity data. However, it doesn't explicitly differentiate itself from 'get_safety_info' which might overlap in scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'get_safety_info'. There's no mention of prerequisites, appropriate contexts, or exclusions. The only implied usage is when toxicity data for a specific compound is needed, but this is too vague for effective tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
get_compound_info - First observed
get_safety_info - First observed
get_toxicity_data
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: get_compound_info retrieves basic chemical properties, get_safety_info focuses on GHS hazard classifications, and get_toxicity_data provides experimental toxicity data. There is no overlap in functionality, and the descriptions clearly differentiate their domains (basic info vs. safety vs. toxicity).
All tools follow a consistent verb_noun pattern with 'get_' prefix and snake_case naming (get_compound_info, get_safety_info, get_toxicity_data). This predictable structure makes it easy for agents to understand and navigate the toolset without confusion.
With only 3 tools, the set feels thin for a chemical safety domain that might benefit from additional operations like searching compounds, batch queries, or regulatory data. While the tools cover core areas, the low count limits comprehensive coverage and could require agents to work around missing functionality.
The tools cover key aspects (basic info, safety, toxicity) but leave notable gaps. There is no search or lookup tool to find compounds by criteria, no batch operations for efficiency, and no update/delete capabilities (though less critical here). Agents may struggle to initiate workflows without a way to discover compounds beyond direct name input.
Maintenance
Related MCP Connectors
Search PubChem compounds, properties, safety data, bioactivity, and cross-references.
Link compounds to protein targets, rank bioactivity, and look up drug mechanisms and indications.
PubChem MCP — NIH chemistry compound database (no auth)
Biomedical data: compounds, drug info, and molecular targets
Related MCP Servers
- AlicenseBqualityDmaintenanceEnables comprehensive access to PubChem's chemical database with over 110 million compounds. Supports chemical searches, structure analysis, bioactivity data, safety information, and molecular property calculations through 30 specialized tools.30MIT
- AlicenseAqualityAmaintenanceEnables AI agents and applications to search, retrieve, and analyze chemical compounds, substances, and bioassays from PubChem's vast chemical information database through comprehensive tools for chemical research and discovery.101339Apache 2.0
- FlicenseNot gradedqualityDmaintenanceEnables AI assistants to search and retrieve chemical compound information, structures, and physical properties from the PubChem database. It supports querying via compound names, SMILES notation, or CIDs to provide detailed molecular data for chemical analysis.5-
- FlicenseBqualityFmaintenanceExtracts basic chemical information and drug data from the PubChem API. It enables users to retrieve molecular details such as SMILES, IUPAC names, molecular formulas, and synonyms for specific compounds.311-