wan2.2-image-generator
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@wan2.2-image-generatorgenerate a watercolor painting of a mountain landscape"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
wan2.2 MCP 图像生成服务器
本项目全部由AIIDE编写 🎨 基于阿里云通义万相的 MCP (Model Context Protocol) 图像生成服务器,支持在 CherryStudio、VSCode、Cursor 等工具中直接使用文本生成图像功能。
🚀 快速开始(小白必看)
第一步:下载和安装
下载项目
git clone <github项目地址> cd wan2.2MCP安装依赖
npm install构建项目
npm run build
第二步:获取 API 密钥
注册/登录阿里云账号
开通「通义万相」服务
创建 API 密钥(格式:
sk-xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx)确保账户有足够余额(有赠送wan2.2-t2i-plus wan2.2-t2i-flash两种模型各100次)
第三步:配置 API 密钥
复制配置模板
cp data/config.example.json data/config.json编辑配置文件 打开
data/config.json,填入你的 API 密钥:{ "api_key": "sk-你的真实API密钥", "region": "cn-beijing", "default_size": "1024*1024", "default_style": "photography", "default_quality": "standard" }或使用
set-config工具: 通过MCP工具动态配置API密钥,无需手动编辑文件。直接和模型说我要将APIkey设置为sk-XXX,模型会自动将sk-XXX写入配置文件。
Related MCP server: jimeng_visual_generation
使用方法
启动服务器
npm start看到以下输出表示启动成功:
[INFO] MCP Server started successfully
[INFO] Listening on stdio🔌 接入各种工具
CherryStudio 接入
打开 CherryStudio
进入设置 → MCP 服务器
添加新的服务器配置:
{
"mcpServers": {
"wan2.2-image-generator": {
"command": "node",
"args": [
"C:\\Users\\用户名\\wan2.2MCP\\dist\\index.js"
],
"cwd": "C:\\Users\\用户名\\wan2.2MCP"
}
}
}⚠️ 重要提醒:请将路径替换为你的实际项目路径!
VSCode 接入
安装 MCP 扩展
在 VSCode 设置中添加:
{
"mcp.servers": {
"wan2.2-image-generator": {
"command": "node",
"args": [
"C:\\Users\\用户名\\wan2.2MCP\\dist\\index.js"
],
"cwd": "C:\\Users\\用户名\\wan2.2MCP"
}
}
}Cursor 接入
打开 Cursor 设置
找到 MCP 配置选项
添加服务器:
{
"mcpServers": {
"wan2.2-image-generator": {
"command": "node",
"args": [
"C:\\Users\\用户名\\wan2.2MCP\\dist\\index.js"
],
"cwd": "C:\\Users\\用户名\\wan2.2MCP"
}
}
}通用配置模板
如果你使用其他支持 MCP 的工具,可以使用这个通用配置:
{
"mcpServers": {
"wan2.2-image-generator": {
"command": "node",
"args": [
"/path/to/wan2.2MCP/dist/index.js"
],
"cwd": "/path/to/wan2.2MCP"
}
}
}🎯 使用方法
配置完成后,你就可以在对话中使用以下功能:
生成图像
直接在对话中说:
"帮我生成一张可爱小猫的图片"
"画一个科幻风格的机器人"
"生成一张1280x720的风景照片"
"用水彩风格画一朵玫瑰花"
可用的图像风格
📸 photography - 摄影风格
🎭 portrait - 肖像风格
🎮 3d cartoon - 3D卡通
🌸 anime - 动漫风格
🎨 oil painting - 油画
🌊 watercolor - 水彩
✏️ sketch - 素描
🖼️ chinese painting - 中国画
📱 flat illustration - 扁平插画
支持的图像尺寸
1024×1024 - 方形图片
720×1280 - 竖屏图片
1280×720 - 横屏图片
🛠️ 高级配置
修改默认设置
编辑 data/config.json 文件:
{
"api_key": "你的API密钥",
"region": "cn-beijing",
"default_size": "1024*1024", // 默认图片尺寸
"default_style": "photography", // 默认图片风格
"default_quality": "standard" // 默认图片质量 (standard/hd)
}查看生成历史
所有生成的图片记录都保存在 data/history.json 文件中,包含:
生成时间
提示词
图片参数
图片URL
🔧 故障排除
常见问题
Q: 提示 "API key not configured"
A: 检查 data/config.json 文件是否存在且包含正确的 API 密钥
Q: 提示 "Request failed"
A: 检查网络连接和 API 密钥是否有效,确认阿里云账户余额充足
Q: 工具中找不到图像生成功能
A: 确认 MCP 服务器已正确启动,检查配置路径是否正确
Q: 生成的图片无法显示
A: 图片URL有时效性,建议及时保存到本地
Q: 路径配置错误
A: 确保使用绝对路径,Windows 系统注意使用双反斜杠 \\
调试模式
如果遇到问题,可以开启调试模式查看详细日志:
LOG_LEVEL=DEBUG npm start测试连接
运行测试脚本验证配置:
node test_mcp_functions.cjs📋 功能特性
🎨 文本生成图像: 支持中英文提示词,生成高质量图像
🎭 多种风格: 支持摄影、插画、3D卡通、动漫、油画、水彩、素描等风格
📐 多种尺寸: 支持方形、横向、纵向等多种图像尺寸
⚙️ 配置管理: 灵活的API配置和默认参数设置
📚 历史记录: 完整的生成历史记录和统计信息
📊 任务跟踪: 实时任务状态查询和自动等待完成
📝 完善日志: 详细的日志记录和错误处理
🛡️ 错误处理: 健壮的错误处理和恢复机制
🔒 安全说明
⚠️ 请勿将 API 密钥提交到公共仓库
🔐 定期更换 API 密钥
💰 监控 API 使用量和费用
🚫 不要在公共场所展示包含密钥的配置文件
📞 技术支持
如果遇到问题:
📖 查看本文档的故障排除部分
🔍 检查项目的
SECURITY.md文件🧪 运行测试脚本验证配置
📝 查看详细的错误日志
🎓 使用示例
基本使用
用户: 帮我生成一张可爱的小猫图片
AI: 我来为您生成一张可爱的小猫图片...
[生成的图片URL]指定风格和尺寸
用户: 用动漫风格生成一张1280x720的机器人图片
AI: 我来为您生成一张动漫风格的机器人图片...
[生成的图片URL]📄 许可证
MIT License - 详见 LICENSE 文件
🎉 恭喜!现在你可以在各种工具中愉快地使用 AI 图像生成功能了!
如果你觉得这个项目有用,请给个 ⭐ Star 支持一下!
Available Tools
6 toolsdiagnose-apiC
诊断API连接和配置问题
| Name | Required | Description | Default |
|---|---|---|---|
| check_auth | No | 是否检查API认证 | |
| check_quota | No | 是否检查配额限制 | |
| check_network | No | 是否检查网络连接 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description carries full burden for behavioral transparency. It says 'diagnose' but does not explain whether it runs network tests, modifies state, or returns a report. This leaves the agent unsure about side effects or read-only nature.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the purpose. However, given the tool's importance, a bit more context would be beneficial without breaking conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description lacks information about what the tool returns or outputs. For a diagnostic tool, knowing the output format (e.g., success/failure, detailed report) is crucial. It also does not provide context relative to sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptions for all three parameters. The tool description adds no additional meaning beyond the schema, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose as diagnosing API connection and configuration issues. It uses a specific verb 'diagnose' and a clear resource. However, it does not explicitly distinguish from sibling tools like get-config or test-model, which might have overlapping functionality.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no guidance on when to use this tool versus alternatives. It does not mention prerequisites, typical use cases, or when not to use it. The agent must infer from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate-imageC
使用通义万相API生成图像
| Name | Required | Description | Default |
|---|---|---|---|
| n | No | 生成图像数量 | |
| size | No | 图像尺寸 | 1024*1024 |
| style | No | 图像风格 | photography |
| prompt | Yes | 图像描述文本,支持中英文 | |
| quality | No | 图像质量 | standard |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, and the description offers no behavioral details such as rate limits, authentication requirements, or the impact of parameters like 'n' or 'quality'. The tool's operation is opaque.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise (one sentence) but lacks necessary detail. It is not front-loaded with critical context and could benefit from elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 5 parameters and no output schema, the description does not explain return format, error handling, or parameter interactions. It is incomplete for safe invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds no additional meaning beyond what the schema provides for the parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool generates images using the Tongyi Wanxiang API, with a specific verb and resource. It distinguishes from siblings which focus on configuration, diagnostics, and history.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, no prerequisites or exclusions mentioned. The description lacks context for decision-making.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get-configB
获取当前配置
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description must convey behavioral traits. It only states the action without disclosing safety, caching, or authentication requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence, front-loaded and efficient. It could benefit from slight expansion but is appropriately sized for a simple tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a parameterless getter, but considering no output schema or annotations, it could mention whether the config is static or dynamic, or provide return value hints.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With no parameters (schema coverage 100%), the description's job is minimal. It adds meaning by describing the verb and resource, though the schema already captures this.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description '获取当前配置' clearly conveys retrieving current configuration. It distinguishes from the sibling 'set-config' which writes, but doesn't explicitly differentiate from other siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No usage guidance provided. The description does not indicate when to use this tool instead of alternatives, nor any context or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list-historyB
获取生成历史记录
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | 返回记录数量 | |
| offset | No | 偏移量 | |
| date_to | No | 结束日期 (ISO 8601格式) | |
| date_from | No | 开始日期 (ISO 8601格式) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of disclosing behavioral traits. It only states the tool 'gets history' but omits crucial details like pagination behavior, return format, whether data is sorted, or if any destructive actions occur.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no wasted words. However, it is too brief to cover all necessary information, slightly reducing effectiveness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of output schema and minimal description, the tool definition is incomplete. It does not explain return values, ordering, or any side effects, which is insufficient for a complete understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are well-documented in the schema. The description adds no additional meaning beyond what the schema provides, hence a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description '获取生成历史记录' clearly states the tool retrieves generation history records. It distinguishes from siblings like generate-image (image generation) and get-config/set-config (configuration).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No usage guidelines are provided. The description does not indicate when to use this tool versus alternatives or any conditions under which it should or should not be used.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set-configC
设置API配置
| Name | Required | Description | Default |
|---|---|---|---|
| region | No | 服务区域 | cn-beijing |
| api_key | Yes | 阿里云DashScope API密钥 | |
| default_size | No | 默认图像尺寸 | |
| default_style | No | 默认图像风格 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, yet the description fails to disclose behavioral traits such as whether the configuration persists, overwrites existing settings, or requires specific authorization. The description adds no behavioral context beyond the basic action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely brief (one short phrase) but this brevity results in under-specification. It restates the tool's name without adding valuable context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 4 parameters (including enums) and no output schema, the description should explain the overall effect and usage flow. It provides none of this, leaving the agent without sufficient context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% coverage with descriptions for all parameters. The tool description does not add additional meaning beyond what is already in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description '设置API配置' (Set API configuration) indicates the tool's action and resource, but it lacks specificity about which configuration aspects are set beyond the parameter list. It does not differentiate itself from sibling tool 'get-config'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'get-config' or 'diagnose-api'. The description does not mention prerequisites or context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
test-modelB
测试特定模型的可用性
| Name | Required | Description | Default |
|---|---|---|---|
| model | Yes | 要测试的模型名称 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry behavioral transparency. It does not disclose side effects, success/failure indicators, or return format. As a test tool, it likely has no destructive effects, but this is not stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence. It is front-loaded with the key action. However, it lacks any additional structure like use cases or examples, but for a simple tool this is acceptable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of output schema, the description does not explain what the tool returns (e.g., boolean, status message). For a test tool, the notion of 'availability' is ambiguous. The description is minimally complete for a simple tool but leaves gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with one parameter fully described (enum and description). The description adds no additional meaning beyond the schema, which already documents the model parameter completely.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool tests the availability of a specific model (verb+resource), and the name 'test-model' aligns with this. It distinguishes from siblings like 'diagnose-api' which tests API connectivity, not model availability.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives (e.g., diagnose-api, generate-image). The description only states what it does without providing usage context or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v1.0.0- First observed
diagnose-api - First observed
generate-image - First observed
get-config - First observed
list-history - First observed
set-config - First observed
test-model
TDQS
Scored across 6 tools
Each tool has a clearly distinct purpose: API diagnosis, image generation, config management, history listing, config setting, and model testing. No overlap or ambiguity.
All tools follow a consistent verb_noun pattern (e.g., generate-image, set-config) using hyphens, making the set predictable and easy to understand.
With 6 tools, the server is well-scoped for an image generation API, covering configuration, generation, history, diagnostics, and model testing without unnecessary bloat.
Core workflows are covered: generation, config management, history, and diagnostics. Minor gaps exist, such as no explicit tool to list available models, but test-model provides indirect coverage.
Maintenance
Related MCP Connectors
MCP server for Qwen Image 3 AI image generation
Generate AI images and videos from any compatible MCP client.
MCP server for Hailuo (MiniMax) AI video generation
MCP server for Flux AI image generation
Related MCP Servers
- FlicenseNot gradedqualityNot gradedmaintenanceAn MCP server that enables users to generate AI images using the Jimeng AI API with support for automatic status polling and local file saving. It allows for detailed parameter configuration including image dimensions, model selection, and prompt optimization.-
- AlicenseAqualityCmaintenanceMCP server for generating images and videos using Volcengine's Jimeng APIs, supporting text-to-image, image-to-image, multi-image fusion, text-to-video, and image-to-video.31MIT
- AlicenseAqualityDmaintenanceMCP server integrating VolcEngine's image generation capabilities, enabling text-to-image, image-to-image, and image set generation for AI applications.422 npmMIT
- AlicenseAqualityDmaintenanceAn MCP server that generates images using SenseNova U1 Fast API. Supports Chinese prompts, multiple resolutions, and generating up to 4 images per request.1678 npm2MIT