Skip to main content
Glama
Soullqs

wan2.2-image-generator

by Soullqs

wan2.2 MCP 图像生成服务器

本项目全部由AIIDE编写 🎨 基于阿里云通义万相的 MCP (Model Context Protocol) 图像生成服务器,支持在 CherryStudio、VSCode、Cursor 等工具中直接使用文本生成图像功能。

🚀 快速开始(小白必看)

第一步:下载和安装

  1. 下载项目

    git clone <github项目地址>
    cd wan2.2MCP
  2. 安装依赖

    npm install
  3. 构建项目

    npm run build

第二步:获取 API 密钥

  1. 访问 阿里云DashScope控制台

  2. 注册/登录阿里云账号

  3. 开通「通义万相」服务

  4. 创建 API 密钥(格式:sk-xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx)

  5. 确保账户有足够余额(有赠送wan2.2-t2i-plus wan2.2-t2i-flash两种模型各100次)

第三步:配置 API 密钥

  1. 复制配置模板

    cp data/config.example.json data/config.json
  2. 编辑配置文件 打开 data/config.json,填入你的 API 密钥:

    {
      "api_key": "sk-你的真实API密钥",
      "region": "cn-beijing",
      "default_size": "1024*1024",
      "default_style": "photography",
      "default_quality": "standard"
    }
  3. 或使用 set-config 工具: 通过MCP工具动态配置API密钥,无需手动编辑文件。直接和模型说我要将APIkey设置为sk-XXX,模型会自动将sk-XXX写入配置文件。

Related MCP server: jimeng_visual_generation

使用方法

启动服务器

npm start

看到以下输出表示启动成功:

[INFO] MCP Server started successfully
[INFO] Listening on stdio

🔌 接入各种工具

CherryStudio 接入

  1. 打开 CherryStudio

  2. 进入设置 → MCP 服务器

  3. 添加新的服务器配置:

{
  "mcpServers": {
    "wan2.2-image-generator": {
      "command": "node",
      "args": [
        "C:\\Users\\用户名\\wan2.2MCP\\dist\\index.js"
      ],
      "cwd": "C:\\Users\\用户名\\wan2.2MCP"
    }
  }
}

⚠️ 重要提醒:请将路径替换为你的实际项目路径!

VSCode 接入

  1. 安装 MCP 扩展

  2. 在 VSCode 设置中添加:

{
  "mcp.servers": {
    "wan2.2-image-generator": {
      "command": "node",
      "args": [
        "C:\\Users\\用户名\\wan2.2MCP\\dist\\index.js"
      ],
      "cwd": "C:\\Users\\用户名\\wan2.2MCP"
    }
  }
}

Cursor 接入

  1. 打开 Cursor 设置

  2. 找到 MCP 配置选项

  3. 添加服务器:

{
  "mcpServers": {
    "wan2.2-image-generator": {
      "command": "node",
      "args": [
        "C:\\Users\\用户名\\wan2.2MCP\\dist\\index.js"
      ],
      "cwd": "C:\\Users\\用户名\\wan2.2MCP"
    }
  }
}

通用配置模板

如果你使用其他支持 MCP 的工具,可以使用这个通用配置:

{
  "mcpServers": {
    "wan2.2-image-generator": {
      "command": "node",
      "args": [
        "/path/to/wan2.2MCP/dist/index.js"
      ],
      "cwd": "/path/to/wan2.2MCP"
    }
  }
}

🎯 使用方法

配置完成后,你就可以在对话中使用以下功能:

生成图像

直接在对话中说:

  • "帮我生成一张可爱小猫的图片"

  • "画一个科幻风格的机器人"

  • "生成一张1280x720的风景照片"

  • "用水彩风格画一朵玫瑰花"

可用的图像风格

  • 📸 photography - 摄影风格

  • 🎭 portrait - 肖像风格

  • 🎮 3d cartoon - 3D卡通

  • 🌸 anime - 动漫风格

  • 🎨 oil painting - 油画

  • 🌊 watercolor - 水彩

  • ✏️ sketch - 素描

  • 🖼️ chinese painting - 中国画

  • 📱 flat illustration - 扁平插画

支持的图像尺寸

  • 1024×1024 - 方形图片

  • 720×1280 - 竖屏图片

  • 1280×720 - 横屏图片

🛠️ 高级配置

修改默认设置

编辑 data/config.json 文件:

{
  "api_key": "你的API密钥",
  "region": "cn-beijing",
  "default_size": "1024*1024",     // 默认图片尺寸
  "default_style": "photography",  // 默认图片风格
  "default_quality": "standard"    // 默认图片质量 (standard/hd)
}

查看生成历史

所有生成的图片记录都保存在 data/history.json 文件中,包含:

  • 生成时间

  • 提示词

  • 图片参数

  • 图片URL

🔧 故障排除

常见问题

Q: 提示 "API key not configured"

A: 检查 data/config.json 文件是否存在且包含正确的 API 密钥

Q: 提示 "Request failed"

A: 检查网络连接和 API 密钥是否有效,确认阿里云账户余额充足

Q: 工具中找不到图像生成功能

A: 确认 MCP 服务器已正确启动,检查配置路径是否正确

Q: 生成的图片无法显示

A: 图片URL有时效性,建议及时保存到本地

Q: 路径配置错误

A: 确保使用绝对路径,Windows 系统注意使用双反斜杠 \\

调试模式

如果遇到问题,可以开启调试模式查看详细日志:

LOG_LEVEL=DEBUG npm start

测试连接

运行测试脚本验证配置:

node test_mcp_functions.cjs

📋 功能特性

  • 🎨 文本生成图像: 支持中英文提示词,生成高质量图像

  • 🎭 多种风格: 支持摄影、插画、3D卡通、动漫、油画、水彩、素描等风格

  • 📐 多种尺寸: 支持方形、横向、纵向等多种图像尺寸

  • ⚙️ 配置管理: 灵活的API配置和默认参数设置

  • 📚 历史记录: 完整的生成历史记录和统计信息

  • 📊 任务跟踪: 实时任务状态查询和自动等待完成

  • 📝 完善日志: 详细的日志记录和错误处理

  • 🛡️ 错误处理: 健壮的错误处理和恢复机制

🔒 安全说明

  • ⚠️ 请勿将 API 密钥提交到公共仓库

  • 🔐 定期更换 API 密钥

  • 💰 监控 API 使用量和费用

  • 🚫 不要在公共场所展示包含密钥的配置文件

📞 技术支持

如果遇到问题:

  1. 📖 查看本文档的故障排除部分

  2. 🔍 检查项目的 SECURITY.md 文件

  3. 🧪 运行测试脚本验证配置

  4. 📝 查看详细的错误日志

🎓 使用示例

基本使用

用户: 帮我生成一张可爱的小猫图片
AI: 我来为您生成一张可爱的小猫图片...
[生成的图片URL]

指定风格和尺寸

用户: 用动漫风格生成一张1280x720的机器人图片
AI: 我来为您生成一张动漫风格的机器人图片...
[生成的图片URL]

📄 许可证

MIT License - 详见 LICENSE 文件


🎉 恭喜!现在你可以在各种工具中愉快地使用 AI 图像生成功能了!

如果你觉得这个项目有用,请给个 ⭐ Star 支持一下!

Available Tools

6 tools
diagnose-apiC

诊断API连接和配置问题

ParametersJSON Schema
NameRequiredDescriptionDefault
check_authNo是否检查API认证
check_quotaNo是否检查配额限制
check_networkNo是否检查网络连接

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description carries full burden for behavioral transparency. It says 'diagnose' but does not explain whether it runs network tests, modifies state, or returns a report. This leaves the agent unsure about side effects or read-only nature.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that front-loads the purpose. However, given the tool's importance, a bit more context would be beneficial without breaking conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description lacks information about what the tool returns or outputs. For a diagnostic tool, knowing the output format (e.g., success/failure, detailed report) is crucial. It also does not provide context relative to sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions for all three parameters. The tool description adds no additional meaning beyond the schema, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose as diagnosing API connection and configuration issues. It uses a specific verb 'diagnose' and a clear resource. However, it does not explicitly distinguish from sibling tools like get-config or test-model, which might have overlapping functionality.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus alternatives. It does not mention prerequisites, typical use cases, or when not to use it. The agent must infer from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate-imageC

使用通义万相API生成图像

ParametersJSON Schema
NameRequiredDescriptionDefault
nNo生成图像数量
sizeNo图像尺寸1024*1024
styleNo图像风格photography
promptYes图像描述文本,支持中英文
qualityNo图像质量standard

TDQS

C2.6/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, and the description offers no behavioral details such as rate limits, authentication requirements, or the impact of parameters like 'n' or 'quality'. The tool's operation is opaque.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise (one sentence) but lacks necessary detail. It is not front-loaded with critical context and could benefit from elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 5 parameters and no output schema, the description does not explain return format, error handling, or parameter interactions. It is incomplete for safe invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds no additional meaning beyond what the schema provides for the parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool generates images using the Tongyi Wanxiang API, with a specific verb and resource. It distinguishes from siblings which focus on configuration, diagnostics, and history.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives, no prerequisites or exclusions mentioned. The description lacks context for decision-making.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-configB

获取当前配置

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description must convey behavioral traits. It only states the action without disclosing safety, caching, or authentication requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence, front-loaded and efficient. It could benefit from slight expansion but is appropriately sized for a simple tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is adequate for a parameterless getter, but considering no output schema or annotations, it could mention whether the config is static or dynamic, or provide return value hints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With no parameters (schema coverage 100%), the description's job is minimal. It adds meaning by describing the verb and resource, though the schema already captures this.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description '获取当前配置' clearly conveys retrieving current configuration. It distinguishes from the sibling 'set-config' which writes, but doesn't explicitly differentiate from other siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No usage guidance provided. The description does not indicate when to use this tool instead of alternatives, nor any context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-historyB

获取生成历史记录

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo返回记录数量
offsetNo偏移量
date_toNo结束日期 (ISO 8601格式)
date_fromNo开始日期 (ISO 8601格式)

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of disclosing behavioral traits. It only states the tool 'gets history' but omits crucial details like pagination behavior, return format, whether data is sorted, or if any destructive actions occur.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence with no wasted words. However, it is too brief to cover all necessary information, slightly reducing effectiveness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of output schema and minimal description, the tool definition is incomplete. It does not explain return values, ordering, or any side effects, which is insufficient for a complete understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so parameters are well-documented in the schema. The description adds no additional meaning beyond what the schema provides, hence a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description '获取生成历史记录' clearly states the tool retrieves generation history records. It distinguishes from siblings like generate-image (image generation) and get-config/set-config (configuration).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No usage guidelines are provided. The description does not indicate when to use this tool versus alternatives or any conditions under which it should or should not be used.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set-configC

设置API配置

ParametersJSON Schema
NameRequiredDescriptionDefault
regionNo服务区域cn-beijing
api_keyYes阿里云DashScope API密钥
default_sizeNo默认图像尺寸
default_styleNo默认图像风格

TDQS

C2.2/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, yet the description fails to disclose behavioral traits such as whether the configuration persists, overwrites existing settings, or requires specific authorization. The description adds no behavioral context beyond the basic action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely brief (one short phrase) but this brevity results in under-specification. It restates the tool's name without adding valuable context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 parameters (including enums) and no output schema, the description should explain the overall effect and usage flow. It provides none of this, leaving the agent without sufficient context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% coverage with descriptions for all parameters. The tool description does not add additional meaning beyond what is already in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description '设置API配置' (Set API configuration) indicates the tool's action and resource, but it lacks specificity about which configuration aspects are set beyond the parameter list. It does not differentiate itself from sibling tool 'get-config'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like 'get-config' or 'diagnose-api'. The description does not mention prerequisites or context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

test-modelB

测试特定模型的可用性

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYes要测试的模型名称

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry behavioral transparency. It does not disclose side effects, success/failure indicators, or return format. As a test tool, it likely has no destructive effects, but this is not stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence. It is front-loaded with the key action. However, it lacks any additional structure like use cases or examples, but for a simple tool this is acceptable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of output schema, the description does not explain what the tool returns (e.g., boolean, status message). For a test tool, the notion of 'availability' is ambiguous. The description is minimally complete for a simple tool but leaves gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with one parameter fully described (enum and description). The description adds no additional meaning beyond the schema, which already documents the model parameter completely.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool tests the availability of a specific model (verb+resource), and the name 'test-model' aligns with this. It distinguishes from siblings like 'diagnose-api' which tests API connectivity, not model availability.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives (e.g., diagnose-api, generate-image). The description only states what it does without providing usage context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv1.0.0
    • First observeddiagnose-api
    • First observedgenerate-image
    • First observedget-config
    • First observedlist-history
    • First observedset-config
    • First observedtest-model

TDQS

B3.3/5.0

Scored across 6 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: API diagnosis, image generation, config management, history listing, config setting, and model testing. No overlap or ambiguity.

Naming Consistency5/5

All tools follow a consistent verb_noun pattern (e.g., generate-image, set-config) using hyphens, making the set predictable and easy to understand.

Tool Count5/5

With 6 tools, the server is well-scoped for an image generation API, covering configuration, generation, history, diagnostics, and model testing without unnecessary bloat.

Completeness4/5

Core workflows are covered: generation, config management, history, and diagnostics. Minor gaps exist, such as no explicit tool to list available models, but test-model provides indirect coverage.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers