Vidu MCP Server
Vidu MCP 服务器
用于与 Vidu 视频生成 API 交互的模型上下文协议 (MCP) 服务器。该服务器提供使用 Vidu 强大的 AI 模型从图像生成视频的工具。
特征
图像到视频的转换:使用可自定义的设置从静态图像生成视频
检查生成状态:监控视频生成任务的进度
图像上传:轻松上传与 Vidu API 一起使用的图像
Related MCP server: veo-mcp-server
先决条件
Node.js(v14 或更高版本)
Vidu API 密钥(可从Vidu 网站获取)
TypeScript(用于开发)
安装
通过 Smithery 安装
要通过Smithery自动为 Claude Desktop 安装 Vidu 视频生成服务器:
npx -y @smithery/cli install @el-el-san/vidu-mcp-server --client claude手动安装
克隆此存储库:
git clone https://github.com/el-el-san/vidu-mcp-server.git
cd vidu-mcp-server安装依赖项:
npm install根据
.env.template创建一个.env文件并添加您的 Vidu API 密钥:
VIDU_API_KEY=your_api_key_here用法
构建 TypeScript 代码:
npm run build启动服务器:
npm startMCP 服务器将启动并准备接受来自 MCP 客户端的连接。
工具
1. 图像转视频
使用可自定义的参数将静态图像转换为视频。
参数:
image_url(必填):要转换为视频的图像的 URLprompt(可选):视频生成的文本提示(最多 1500 个字符)duration(可选):输出视频的持续时间(以秒为单位)(4 或 8,默认 4)model(可选):生成的模型名称(“vidu1.0”,“vidu1.5”,“vidu2.0”,默认“vidu2.0”)resolution(可选):输出视频的分辨率(“360p”,“720p”,“1080p”,默认“720p”)movement_amplitude(可选):物体在框架内的运动幅度(“自动”、“小”、“中”、“大”,默认“自动”)seed(可选):用于重复性的随机种子
示例请求:
{
"image_url": "https://example.com/image.jpg",
"prompt": "A serene lake with mountains in the background",
"duration": 8,
"model": "vidu2.0",
"resolution": "720p",
"movement_amplitude": "medium",
"seed": 12345
}2. 检查生成状态
检查正在运行的视频生成任务的状态。
参数:
task_id(必填):图片转视频工具返回的任务ID
示例请求:
{
"task_id": "12345abcde"
}3.上传图片
上传图像以供 Vidu API 使用。
参数:
image_path(必需):图像文件的本地路径image_type(必需):图像文件类型(“png”,“webp”,“jpeg”,“jpg”)
示例请求:
{
"image_path": "/path/to/your/image.jpg",
"image_type": "jpg"
}工作原理
该服务器使用模型上下文协议 (MCP) 为 AI 工具提供标准化接口。启动服务器时,它会通过标准输入/输出通道监听命令,并以结构化格式返回结果。
服务器处理与 Vidu API 交互的所有复杂问题,包括:
使用 API 密钥进行身份验证
文件上传和格式验证
异步任务管理和轮询
错误处理和报告
故障排除
API 密钥问题:确保您的 Vidu API 密钥在
.env文件中正确设置文件上传错误:检查您的图像文件是否有效且大小不超过 10MB
连接问题:确保您可以访问互联网并可以访问 Vidu API 服务器
贡献
欢迎贡献代码!欢迎提交 Pull 请求。
Available Tools
3 toolscheck-generation-statusB
Check the status of a video generation task
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | Task ID returned by the image-to-video tool |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states the tool checks status but doesn't disclose behavioral traits like whether it's read-only, safe to call repeatedly, rate-limited, or what the response format might be (e.g., pending, completed, failed). This leaves significant gaps for an agent to understand how to interact with it effectively.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence that directly states the tool's purpose without any wasted words. It is appropriately sized and front-loaded, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a status-checking tool with no annotations and no output schema, the description is incomplete. It doesn't explain what statuses might be returned, error handling, or usage patterns (e.g., polling intervals), which are crucial for an agent to use this tool correctly in a workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the parameter 'task_id' fully described as 'Task ID returned by the image-to-video tool.' The description adds no additional parameter semantics beyond this, so it meets the baseline of 3 where the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose as checking the status of a video generation task, which is a specific verb (check) and resource (video generation task). However, it doesn't explicitly distinguish this from sibling tools like 'image-to-video' or 'upload-image' beyond the implied relationship through the task_id parameter description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by referencing 'video generation task,' and the parameter description mentions 'task_id returned by the image-to-video tool,' suggesting when to use it (after initiating a generation). However, it lacks explicit guidance on when not to use it or alternatives, such as whether it's for polling or one-time checks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
image-to-videoC
Generate a video from an image using Vidu API
| Name | Required | Description | Default |
|---|---|---|---|
| duration | No | Duration of the output video in seconds (4 or 8) | |
| image_url | Yes | URL of the image to convert to video | |
| model | No | Model name for generation | vidu2.0 |
| movement_amplitude | No | Movement amplitude of objects in the frame | auto |
| prompt | No | Text prompt for video generation (max 1500 chars) | |
| resolution | No | Resolution of the output video | 720p |
| seed | No | Random seed for reproducibility |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool generates a video but lacks details on execution time, rate limits, authentication needs, output format (e.g., video file type), error handling, or whether it's a synchronous/asynchronous operation. For a complex 7-parameter tool with no annotations, this is a significant gap in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without redundancy. It's front-loaded with the core action and resource, and every word earns its place by specifying the API used. No unnecessary details or fluff are included.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (7 parameters, video generation task) and lack of annotations and output schema, the description is incomplete. It doesn't cover behavioral aspects like performance, output details, or error handling, which are critical for an AI agent to use this tool effectively. The description alone is insufficient for a tool of this nature.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all 7 parameters with descriptions, defaults, and constraints. The description adds no parameter-specific information beyond what's in the schema, such as explaining interactions between parameters (e.g., how 'prompt' influences generation). Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Generate a video') and resource ('from an image'), specifying it uses the Vidu API. It distinguishes from sibling tools like 'check-generation-status' and 'upload-image' by focusing on video generation rather than status checking or image uploading. However, it doesn't explicitly differentiate from potential non-sibling alternatives beyond mentioning the API.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing an uploaded image first), when not to use it, or how it relates to sibling tools like 'check-generation-status' for monitoring generation progress. Usage is implied only by the tool name and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload-imageC
Upload an image to use with the Vidu API
| Name | Required | Description | Default |
|---|---|---|---|
| image_path | Yes | Local path to the image file | |
| image_type | Yes | Image file type |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden but only states the basic action. It doesn't disclose behavioral traits such as authentication needs, rate limits, error handling, or what happens after upload (e.g., returns an image ID). This leaves significant gaps for an agent to understand the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero waste. It's front-loaded and appropriately sized for a simple upload tool, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description is incomplete. It lacks details on what the tool returns, error conditions, or integration context with Vidu API. For a tool with two parameters and no structured behavioral data, this leaves the agent under-informed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents the two parameters. The description adds no additional meaning beyond implying the image is for Vidu API use, which is minimal value. Baseline 3 is appropriate as the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('upload') and resource ('an image'), specifying it's for use with the Vidu API. It doesn't differentiate from sibling tools like 'image-to-video' or 'check-generation-status', but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'image-to-video'. The description mentions the Vidu API context but doesn't specify prerequisites, constraints, or typical workflows, leaving usage unclear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.0- First observed
check-generation-status - First observed
image-to-video - First observed
upload-image
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: check-generation-status monitors task progress, image-to-video creates videos from images, and upload-image handles image uploads. There is no overlap in functionality, making tool selection straightforward.
The tools follow a consistent verb-object naming pattern (check-generation-status, image-to-video, upload-image), all using hyphens. However, the pattern is slightly inconsistent as 'image-to-video' uses a preposition 'to' while others do not, but it remains readable and predictable.
With 3 tools, the count is appropriate for a focused video generation API server, covering core operations. It is slightly lean but reasonable for the domain, as it includes upload, generation, and status checking without unnecessary bloat.
The tools cover basic video generation workflows: upload, generate, and check status. However, there are notable gaps such as missing operations for managing or deleting uploaded images, handling video outputs, or supporting other input types beyond images, which could limit agent capabilities.
Maintenance
Related MCP Connectors
MCP server for Google Veo AI video generation
MCP server for Wan AI video generation
MCP server for Luma Dream Machine AI video generation
MCP server for Hailuo (MiniMax) AI video generation
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceAI-powered image and video generation and processing server that supports text-to-image, image-to-image, text/image-to-video generation, image analysis, and comprehensive editing operations (crop, resize, convert, adjust) through providers like Doubao and Aliyun.6MIT
- AlicenseNot gradedqualityCmaintenanceMCP server for 4K video generation using Google VEO 3.1 — text-to-video, image-to-video, video extension, and frame interpolation.2MIT
- AlicenseAqualityDmaintenanceMCP server for Google Veo 3.1 video generation. Supports text/video/image-based generation, extension, and interpolation with cost estimation and batch processing.16433 npm1MIT
- AlicenseAqualityCmaintenanceMCP server for generating, editing, and batch processing videos using xAI's Grok Imagine Video API, with support for text-to-video, image-to-video, and video editing via natural language prompts.477 npm1MIT