Tongyi Wanxiang MCP Server
This server provides Model Context Protocol (MCP) access to Alibaba Cloud's Tongyi Wanxiang API for AI-powered text-to-image and text-to-video generation.
Capabilities include:
Generate images from text prompts (with optional negative prompts) using the
wanx-t2i-image-generationtoolRetrieve generated images using task IDs via the
wanx-t2i-image-generation-resulttoolGenerate videos from text prompts using the
wanx-t2v-video-generationtoolRetrieve generated videos using task IDs via the
wanx-t2v-video-generation-resulttool
Integrates with Alibaba Cloud's Tongyi Wanxiang AI to provide text-to-image and text-to-video generation capabilities, allowing users to create high-quality AI-generated images and videos through Tongyi Wanxiang's APIs
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Tongyi Wanxiang MCP Servergenerate an image of a serene mountain landscape at sunset"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
通义万相 MCP 服务器
这是一个基于 TypeScript 的 Model Context Protocol (MCP) 服务器,专门提供阿里云通义万相的文生图(Text-to-Image)和文生视频(Text-to-Video)能力。该服务器通过 MCP 协议,允许大语言模型(LLM)直接调用通义万相的图像和视频生成 API。
功能特点
文生图能力集成:接入阿里云通义万相文生图 API,支持高质量的 AI 图像生成
文生视频能力集成:接入阿里云通义万相文生视频 API,支持高质量的 AI 视频生成
异步任务处理:支持长时间运行的图像和视频生成任务,通过异步轮询获取最终结果
MCP 协议支持:符合 Model Context Protocol 规范,可与支持 MCP 的 LLM 无缝协作
Related MCP server: mcp-flux-schnell
环境要求
Node.js >= 16.x
npm >= 8.x 或 pnpm
如何使用
以百炼平台举例
{
"mcpServers": {
"tongyi-wanxiang": {
"command": "npx",
"args": [
"-y",
"tongyi-wanx-mcp-server@latest"
],
"env": {
"DASHSCOPE_API_KEY": "<你的通义万相 API 密钥>"
}
}
}
}如何开发
安装依赖
# 使用 npm
npm install
# 或使用 pnpm
pnpm install构建与运行
# 构建项目
npm run build
# 或
pnpm run build
# 运行服务器
npm start
# 或
pnpm start
# 使用调试工具运行
npm run debug
# 或
pnpm run debugAPI 使用
该服务器提供以下 MCP 工具:
1. 文生图生成(wanx-t2i-image-generation)
启动图像生成任务,返回任务 ID。
参数:
prompt: 图像生成提示词negative_prompt: 负面提示词(不希望在图像中出现的元素)
返回:
包含
task_id的任务信息
2. 获取生成结果(wanx-t2i-image-generation-result)
通过任务 ID 获取图像生成结果。
参数:
task_id: 由文生图生成工具返回的任务 ID
返回:
图像生成结果,包含图像 URL
3. 文生视频生成(wanx-t2v-video-generation)
启动视频生成任务,返回任务 ID。
参数:
prompt: 视频生成提示词
返回:
包含
task_id的任务信息
4. 获取视频生成结果(wanx-t2v-video-generation-result)
通过任务 ID 获取视频生成结果。
参数:
task_id: 由文生视频生成工具返回的任务 ID
返回:
视频生成结果,包含视频 URL
项目结构
project/
├── src/ # 源代码目录
│ ├── index.ts # 主入口文件,MCP 服务器定义
│ ├── wanx-t2i.js # 通义万相文生图 API 集成
│ ├── wanx-t2v.js # 通义万相文生视频 API 集成
│ └── config.ts # 配置文件
├── dist/ # 编译后的代码目录
├── package.json # 项目配置
├── tsconfig.json # TypeScript 配置
└── README.md # 项目说明通义万相 API 参数说明
文生图 API 支持的参数
model: 模型名称,默认为
wanx2.1-t2i-turbosize: 图像尺寸,默认为
1024*1024n: 生成图像数量,默认为 1
seed: 随机种子,用于复现结果
prompt_extend: 是否启用提示词扩展,默认为 true
watermark: 是否添加水印,默认为 false
高级配置
您可以在 src/config.ts 中修改以下配置:
pollingInterval: 轮询任务状态的间隔时间(毫秒)
maxRetries: 最大轮询次数
defaultModel: 默认使用的模型
注意事项
请确保您有有效的通义万相 API 访问权限和密钥
图像生成是一个异步过程,可能需要数秒到数十秒不等
视频生成过程耗时较长,可能需要数分钟到十几分钟不等
视频生成状态查询可能会多次失败,系统会自动重试,请耐心等待
请合理设置轮询间隔和最大重试次数,以适应您的使用场景
对于视频生成任务,建议增加最大重试次数和轮询间隔时间
参考资料
Available Tools
4 toolswanx-t2i-image-generationB
使用阿里云万相文生图大模型的文生图能力,由于图片生成耗时比较久,需要调用 wanx-t2i-image-generation-result 工具获取结果
| Name | Required | Description | Default |
|---|---|---|---|
| negative_prompt | Yes | ||
| prompt | Yes | ||
| seed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the asynchronous behavior (needs result tool) and performance characteristic (takes time), which is valuable. However, it doesn't mention permissions, rate limits, or what happens if generation fails, leaving gaps in behavioral understanding.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences that efficiently convey the core functionality and usage requirement. It's front-loaded with the main purpose and follows with critical behavioral information, with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 3 parameters, 0% schema coverage, no annotations, and no output schema, the description is incomplete. It explains the asynchronous workflow but doesn't cover parameter meanings, error handling, or output format, leaving significant gaps for the agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It provides no information about the three parameters (prompt, negative_prompt, seed), their formats, or examples. This leaves parameters completely undocumented beyond the schema structure.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool uses Alibaba Cloud's Wanxiang text-to-image model for image generation, which provides a basic purpose. However, it doesn't specify what kind of images it generates or differentiate from the video generation sibling tools, making it somewhat vague.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states that image generation takes time and requires calling 'wanx-t2i-image-generation-result' to get results, providing clear context for usage. It doesn't mention when to use this versus the video generation tools, but the time constraint guidance is helpful.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wanx-t2i-image-generation-resultC
获取阿里云万相文生图大模型的文生图结果
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool 'gets' results, implying a read-only operation, but doesn't specify behavioral traits such as whether it requires authentication, rate limits, what happens if the task_id is invalid (e.g., errors, null returns), or the format of the results (e.g., image data, JSON metadata). This leaves significant gaps in understanding how the tool behaves in practice.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence in Chinese that directly states the tool's purpose without unnecessary words. It's appropriately sized for a simple tool, though it could be more front-loaded with key details (e.g., clarifying 'result' as image retrieval). There's no wasted text, earning a high score for conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (a result-retrieval operation with 1 parameter), lack of annotations, no output schema, and 0% schema description coverage, the description is incomplete. It doesn't explain the parameter's semantics, behavioral aspects like error handling or output format, or how it integrates with sibling tools. This makes it inadequate for an AI agent to use the tool correctly without additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 1 parameter (task_id) with 0% description coverage, meaning the schema provides no semantic information. The description doesn't add any meaning beyond the schema—it doesn't explain what 'task_id' is (e.g., an ID from a prior generation request), its format, or how to obtain it. With low schema coverage and no compensation in the description, this falls short of the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool '获取阿里云万相文生图大模型的文生图结果' which translates to 'Get the text-to-image generation results of Alibaba Cloud Wanxiang text-to-image model.' This specifies the action (get results) and resource (text-to-image generation results), but it's somewhat vague about what exactly 'results' entails (e.g., images, metadata, status). It doesn't clearly distinguish from sibling tools like 'wanx-t2i-image-generation' (which likely initiates generation) or 'wanx-t2v-video-generation-result' (which handles video results).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing a task_id from a previous generation request), exclusions, or how it relates to siblings like 'wanx-t2i-image-generation' (presumably for initiating tasks) or 'wanx-t2v-video-generation-result' (for video results). Usage is implied only through the name 'result,' but no explicit context is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wanx-t2v-video-generationA
使用阿里云万相文生视频大模型的文生视频能力,由于视频生成耗时比较久,需要调用 wanx-t2v-video-generation-result 工具获取结果
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It adds useful context about the asynchronous nature (long processing time) and the need to poll with another tool, which is critical for understanding the tool's behavior. However, it doesn't mention other traits like potential rate limits, error conditions, authentication needs, or what happens if the prompt is invalid—leaving gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is highly concise and well-structured in two sentences: the first states the core purpose, and the second provides critical usage guidance. Every sentence earns its place with no wasted words, making it easy to parse and front-loaded with essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (asynchronous video generation with no output schema and no annotations), the description is moderately complete. It covers the basic purpose and the need for a result tool, but lacks details on error handling, output format, or integration with siblings beyond the result tool. For a mutation tool with significant behavioral implications, more context would be beneficial.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 1 parameter with 0% description coverage, so the description must compensate. It implies the 'prompt' parameter is for text input to generate video, but doesn't add specific meaning beyond that (e.g., format, length constraints, or examples). Since schema coverage is low, the description provides minimal semantic value, meeting the baseline but not fully compensating for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: it uses Alibaba Cloud's Wanxiang text-to-video model to generate videos from text prompts. It specifies the exact capability ('text-to-video generation') and mentions the resource (Alibaba Cloud Wanxiang model). However, it doesn't explicitly differentiate from its sibling 'wanx-t2i-image-generation' (text-to-image), though the naming implies the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context by stating that video generation takes a long time and requires calling 'wanx-t2v-video-generation-result' to get results. This gives practical guidance on when to use this tool (for initiating generation) versus the sibling result tool (for retrieving results). It doesn't explicitly mention when not to use it or compare to alternatives like text-to-image, but the context is sufficient for basic usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wanx-t2v-video-generation-resultC
获取阿里云万相文生视频大模型的文生视频结果
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states it 'gets' results, implying a read-only operation, but doesn't describe key behaviors: whether it polls for completion, returns partial results, has rate limits, requires authentication, or what happens if the task_id is invalid. This leaves significant gaps for a tool that likely interacts with an async API.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence in Chinese that directly states the tool's purpose without unnecessary words. It's appropriately sized and front-loaded, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (likely involving async video generation results), lack of annotations, no output schema, and 0% schema description coverage, the description is insufficient. It doesn't explain the return format (e.g., video URL, status), error conditions, or behavioral nuances, leaving the agent with inadequate information to use the tool effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description doesn't mention parameters, and schema description coverage is 0% (no descriptions for the 'task_id' parameter). However, with only one parameter, the baseline is higher. The description implies a 'task_id' is needed to retrieve results, adding minimal context beyond the schema's structure, but doesn't explain what a task_id is or where to get it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('获取' meaning 'get' or 'retrieve') and the resource ('文生视频结果' meaning 'text-to-video generation result'), specifying it's for the Alibaba Cloud Wanx model. It distinguishes from the sibling 'wanx-t2v-video-generation' (which likely initiates generation) by focusing on result retrieval, though it doesn't explicitly mention this distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing a task_id from a prior generation request), when not to use it, or how it relates to sibling tools like 'wanx-t2v-video-generation' (presumably for initiating generation). Usage is implied but not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v1.0.0- First observed
wanx-t2i-image-generation - First observed
wanx-t2i-image-generation-result - First observed
wanx-t2v-video-generation - First observed
wanx-t2v-video-generation-result
TDQS
Scored across 4 tools
The tools have perfectly distinct purposes: two for image generation (one to start, one to get results) and two for video generation (one to start, one to get results). There is no overlap or ambiguity between the image and video domains, and the result tools are clearly tied to their respective generation tools.
All tool names follow a consistent pattern: 'wanx-t2i-' or 'wanx-t2v-' prefixes indicating the model type, followed by a descriptive hyphenated phrase (e.g., 'image-generation', 'video-generation-result'). This pattern is applied uniformly across all four tools, making them predictable and easy to understand.
With 4 tools, the server is well-scoped for its purpose of generating images and videos via the Tongyi Wanxiang models. Each tool earns its place by covering the necessary asynchronous workflow (start generation and retrieve results) for both image and video tasks, avoiding bloat or redundancy.
The tool set is complete for the server's domain of image and video generation. It provides full coverage for both tasks with dedicated tools to initiate generation and retrieve results, ensuring agents can handle the entire lifecycle without gaps or dead ends.
Maintenance
Related MCP Connectors
A Model Context Protocol server for Wix AI tools
MCP server for Qwen Image 3 AI image generation
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
MCP server for AI dialogue using various LLM models via AceDataCloud
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA TypeScript-based Model Context Protocol (MCP) server enabling integration with PiAPI for media content generation using platforms like Midjourney, Flux, and others through MCP-compatible applications.75MIT
- AlicenseBqualityDmaintenanceA TypeScript-based MCP server that enables text-to-image generation using Cloudflare's Flux Schnell model API.15MIT
- -licenseCqualityNot gradedmaintenanceA TypeScript-based Model Context Protocol server that integrates with Volcengine's Jimeng AI image generation service, allowing users to generate AI images through simple tool calls.11 npm2-
- AlicenseBqualityDmaintenanceA Model Context Protocol server that enables agent applications like Cursor and Cline to integrate with Alibaba Cloud Function Compute, allowing them to deploy and manage serverless functions through natural language interactions.1213 npm9MIT