Skip to main content
Glama
Suixinlei

Tongyi Wanxiang MCP Server

by Suixinlei

通义万相 MCP 服务器

这是一个基于 TypeScript 的 Model Context Protocol (MCP) 服务器,专门提供阿里云通义万相的文生图(Text-to-Image)和文生视频(Text-to-Video)能力。该服务器通过 MCP 协议,允许大语言模型(LLM)直接调用通义万相的图像和视频生成 API。

功能特点

  • 文生图能力集成:接入阿里云通义万相文生图 API,支持高质量的 AI 图像生成

  • 文生视频能力集成:接入阿里云通义万相文生视频 API,支持高质量的 AI 视频生成

  • 异步任务处理:支持长时间运行的图像和视频生成任务,通过异步轮询获取最终结果

  • MCP 协议支持:符合 Model Context Protocol 规范,可与支持 MCP 的 LLM 无缝协作

Related MCP server: mcp-flux-schnell

环境要求

  • Node.js >= 16.x

  • npm >= 8.x 或 pnpm

如何使用

以百炼平台举例

{
  "mcpServers": {
    "tongyi-wanxiang": {
      "command": "npx",
      "args": [
        "-y",
        "tongyi-wanx-mcp-server@latest"
      ],
      "env": {
        "DASHSCOPE_API_KEY": "<你的通义万相 API 密钥>"
      }
    }
  }
}

如何开发

安装依赖

# 使用 npm
npm install

# 或使用 pnpm
pnpm install

构建与运行

# 构建项目
npm run build
# 或
pnpm run build

# 运行服务器
npm start
# 或
pnpm start

# 使用调试工具运行
npm run debug
# 或
pnpm run debug

API 使用

该服务器提供以下 MCP 工具:

1. 文生图生成(wanx-t2i-image-generation)

启动图像生成任务,返回任务 ID。

参数

  • prompt: 图像生成提示词

  • negative_prompt: 负面提示词(不希望在图像中出现的元素)

返回

  • 包含 task_id 的任务信息

2. 获取生成结果(wanx-t2i-image-generation-result)

通过任务 ID 获取图像生成结果。

参数

  • task_id: 由文生图生成工具返回的任务 ID

返回

  • 图像生成结果,包含图像 URL

3. 文生视频生成(wanx-t2v-video-generation)

启动视频生成任务,返回任务 ID。

参数

  • prompt: 视频生成提示词

返回

  • 包含 task_id 的任务信息

4. 获取视频生成结果(wanx-t2v-video-generation-result)

通过任务 ID 获取视频生成结果。

参数

  • task_id: 由文生视频生成工具返回的任务 ID

返回

  • 视频生成结果,包含视频 URL

项目结构

project/
├── src/                  # 源代码目录
│   ├── index.ts          # 主入口文件,MCP 服务器定义
│   ├── wanx-t2i.js       # 通义万相文生图 API 集成
│   ├── wanx-t2v.js       # 通义万相文生视频 API 集成
│   └── config.ts         # 配置文件
├── dist/                 # 编译后的代码目录
├── package.json          # 项目配置
├── tsconfig.json         # TypeScript 配置
└── README.md             # 项目说明

通义万相 API 参数说明

文生图 API 支持的参数

  • model: 模型名称,默认为 wanx2.1-t2i-turbo

  • size: 图像尺寸,默认为 1024*1024

  • n: 生成图像数量,默认为 1

  • seed: 随机种子,用于复现结果

  • prompt_extend: 是否启用提示词扩展,默认为 true

  • watermark: 是否添加水印,默认为 false

高级配置

您可以在 src/config.ts 中修改以下配置:

  • pollingInterval: 轮询任务状态的间隔时间(毫秒)

  • maxRetries: 最大轮询次数

  • defaultModel: 默认使用的模型

注意事项

  1. 请确保您有有效的通义万相 API 访问权限和密钥

  2. 图像生成是一个异步过程,可能需要数秒到数十秒不等

  3. 视频生成过程耗时较长,可能需要数分钟到十几分钟不等

  4. 视频生成状态查询可能会多次失败,系统会自动重试,请耐心等待

  5. 请合理设置轮询间隔和最大重试次数,以适应您的使用场景

  6. 对于视频生成任务,建议增加最大重试次数和轮询间隔时间

参考资料

Available Tools

4 tools
wanx-t2i-image-generationB

使用阿里云万相文生图大模型的文生图能力,由于图片生成耗时比较久,需要调用 wanx-t2i-image-generation-result 工具获取结果

ParametersJSON Schema
NameRequiredDescriptionDefault
negative_promptYes
promptYes
seedNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses the asynchronous behavior (needs result tool) and performance characteristic (takes time), which is valuable. However, it doesn't mention permissions, rate limits, or what happens if generation fails, leaving gaps in behavioral understanding.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences that efficiently convey the core functionality and usage requirement. It's front-loaded with the main purpose and follows with critical behavioral information, with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 3 parameters, 0% schema coverage, no annotations, and no output schema, the description is incomplete. It explains the asynchronous workflow but doesn't cover parameter meanings, error handling, or output format, leaving significant gaps for the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It provides no information about the three parameters (prompt, negative_prompt, seed), their formats, or examples. This leaves parameters completely undocumented beyond the schema structure.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool uses Alibaba Cloud's Wanxiang text-to-image model for image generation, which provides a basic purpose. However, it doesn't specify what kind of images it generates or differentiate from the video generation sibling tools, making it somewhat vague.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states that image generation takes time and requires calling 'wanx-t2i-image-generation-result' to get results, providing clear context for usage. It doesn't mention when to use this versus the video generation tools, but the time constraint guidance is helpful.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wanx-t2i-image-generation-resultC

获取阿里云万相文生图大模型的文生图结果

ParametersJSON Schema
NameRequiredDescriptionDefault
task_idYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool 'gets' results, implying a read-only operation, but doesn't specify behavioral traits such as whether it requires authentication, rate limits, what happens if the task_id is invalid (e.g., errors, null returns), or the format of the results (e.g., image data, JSON metadata). This leaves significant gaps in understanding how the tool behaves in practice.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence in Chinese that directly states the tool's purpose without unnecessary words. It's appropriately sized for a simple tool, though it could be more front-loaded with key details (e.g., clarifying 'result' as image retrieval). There's no wasted text, earning a high score for conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (a result-retrieval operation with 1 parameter), lack of annotations, no output schema, and 0% schema description coverage, the description is incomplete. It doesn't explain the parameter's semantics, behavioral aspects like error handling or output format, or how it integrates with sibling tools. This makes it inadequate for an AI agent to use the tool correctly without additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 1 parameter (task_id) with 0% description coverage, meaning the schema provides no semantic information. The description doesn't add any meaning beyond the schema—it doesn't explain what 'task_id' is (e.g., an ID from a prior generation request), its format, or how to obtain it. With low schema coverage and no compensation in the description, this falls short of the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool '获取阿里云万相文生图大模型的文生图结果' which translates to 'Get the text-to-image generation results of Alibaba Cloud Wanxiang text-to-image model.' This specifies the action (get results) and resource (text-to-image generation results), but it's somewhat vague about what exactly 'results' entails (e.g., images, metadata, status). It doesn't clearly distinguish from sibling tools like 'wanx-t2i-image-generation' (which likely initiates generation) or 'wanx-t2v-video-generation-result' (which handles video results).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing a task_id from a previous generation request), exclusions, or how it relates to siblings like 'wanx-t2i-image-generation' (presumably for initiating tasks) or 'wanx-t2v-video-generation-result' (for video results). Usage is implied only through the name 'result,' but no explicit context is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wanx-t2v-video-generationA

使用阿里云万相文生视频大模型的文生视频能力,由于视频生成耗时比较久,需要调用 wanx-t2v-video-generation-result 工具获取结果

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It adds useful context about the asynchronous nature (long processing time) and the need to poll with another tool, which is critical for understanding the tool's behavior. However, it doesn't mention other traits like potential rate limits, error conditions, authentication needs, or what happens if the prompt is invalid—leaving gaps for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is highly concise and well-structured in two sentences: the first states the core purpose, and the second provides critical usage guidance. Every sentence earns its place with no wasted words, making it easy to parse and front-loaded with essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (asynchronous video generation with no output schema and no annotations), the description is moderately complete. It covers the basic purpose and the need for a result tool, but lacks details on error handling, output format, or integration with siblings beyond the result tool. For a mutation tool with significant behavioral implications, more context would be beneficial.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 1 parameter with 0% description coverage, so the description must compensate. It implies the 'prompt' parameter is for text input to generate video, but doesn't add specific meaning beyond that (e.g., format, length constraints, or examples). Since schema coverage is low, the description provides minimal semantic value, meeting the baseline but not fully compensating for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: it uses Alibaba Cloud's Wanxiang text-to-video model to generate videos from text prompts. It specifies the exact capability ('text-to-video generation') and mentions the resource (Alibaba Cloud Wanxiang model). However, it doesn't explicitly differentiate from its sibling 'wanx-t2i-image-generation' (text-to-image), though the naming implies the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage context by stating that video generation takes a long time and requires calling 'wanx-t2v-video-generation-result' to get results. This gives practical guidance on when to use this tool (for initiating generation) versus the sibling result tool (for retrieving results). It doesn't explicitly mention when not to use it or compare to alternatives like text-to-image, but the context is sufficient for basic usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wanx-t2v-video-generation-resultC

获取阿里云万相文生视频大模型的文生视频结果

ParametersJSON Schema
NameRequiredDescriptionDefault
task_idYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states it 'gets' results, implying a read-only operation, but doesn't describe key behaviors: whether it polls for completion, returns partial results, has rate limits, requires authentication, or what happens if the task_id is invalid. This leaves significant gaps for a tool that likely interacts with an async API.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence in Chinese that directly states the tool's purpose without unnecessary words. It's appropriately sized and front-loaded, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (likely involving async video generation results), lack of annotations, no output schema, and 0% schema description coverage, the description is insufficient. It doesn't explain the return format (e.g., video URL, status), error conditions, or behavioral nuances, leaving the agent with inadequate information to use the tool effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description doesn't mention parameters, and schema description coverage is 0% (no descriptions for the 'task_id' parameter). However, with only one parameter, the baseline is higher. The description implies a 'task_id' is needed to retrieve results, adding minimal context beyond the schema's structure, but doesn't explain what a task_id is or where to get it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('获取' meaning 'get' or 'retrieve') and the resource ('文生视频结果' meaning 'text-to-video generation result'), specifying it's for the Alibaba Cloud Wanx model. It distinguishes from the sibling 'wanx-t2v-video-generation' (which likely initiates generation) by focusing on result retrieval, though it doesn't explicitly mention this distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing a task_id from a prior generation request), when not to use it, or how it relates to sibling tools like 'wanx-t2v-video-generation' (presumably for initiating generation). Usage is implied but not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv1.0.0
    • First observedwanx-t2i-image-generation
    • First observedwanx-t2i-image-generation-result
    • First observedwanx-t2v-video-generation
    • First observedwanx-t2v-video-generation-result

TDQS

A3.5/5.0

Scored across 4 tools

Disambiguation5/5

The tools have perfectly distinct purposes: two for image generation (one to start, one to get results) and two for video generation (one to start, one to get results). There is no overlap or ambiguity between the image and video domains, and the result tools are clearly tied to their respective generation tools.

Naming Consistency5/5

All tool names follow a consistent pattern: 'wanx-t2i-' or 'wanx-t2v-' prefixes indicating the model type, followed by a descriptive hyphenated phrase (e.g., 'image-generation', 'video-generation-result'). This pattern is applied uniformly across all four tools, making them predictable and easy to understand.

Tool Count5/5

With 4 tools, the server is well-scoped for its purpose of generating images and videos via the Tongyi Wanxiang models. Each tool earns its place by covering the necessary asynchronous workflow (start generation and retrieve results) for both image and video tasks, avoiding bloat or redundancy.

Completeness5/5

The tool set is complete for the server's domain of image and video generation. It provides full coverage for both tasks with dedicated tools to initiate generation and retrieve results, ensuring agents can handle the entire lifecycle without gaps or dead ends.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers