Vidu MCP Server
Vidu MCP 서버
Vidu 비디오 생성 API와 상호 작용하는 모델 컨텍스트 프로토콜(MCP) 서버입니다. 이 서버는 Vidu의 강력한 AI 모델을 사용하여 이미지에서 비디오를 생성하는 도구를 제공합니다.
특징
이미지를 비디오로 변환 : 사용자 정의 가능한 설정으로 정적 이미지에서 비디오를 생성합니다.
생성 상태 확인 : 비디오 생성 작업의 진행 상황을 모니터링합니다.
이미지 업로드 : Vidu API와 함께 사용할 이미지를 쉽게 업로드하세요
Related MCP server: veo-mcp-server
필수 조건
Node.js(v14 이상)
Vidu API 키( Vidu 웹사이트 에서 사용 가능)
TypeScript(개발용)
설치
Smithery를 통해 설치
Smithery를 통해 Claude Desktop용 Vidu Video Generation Server를 자동으로 설치하려면:
지엑스피1
수동 설치
이 저장소를 복제하세요:
git clone https://github.com/el-el-san/vidu-mcp-server.git
cd vidu-mcp-server종속성 설치:
npm install.env.template기반으로.env파일을 만들고 Vidu API 키를 추가합니다.
VIDU_API_KEY=your_api_key_here용법
TypeScript 코드를 작성합니다.
npm run build서버를 시작합니다:
npm startMCP 서버가 시작되고 MCP 클라이언트의 연결을 수락할 준비가 됩니다.
도구
1. 이미지를 비디오로
사용자 정의 가능한 매개변수를 사용하여 정적 이미지를 비디오로 변환합니다.
매개변수:
image_url(필수): 비디오로 변환할 이미지의 URLprompt(선택 사항): 비디오 생성을 위한 텍스트 프롬프트(최대 1500자)duration(선택 사항): 출력 비디오의 지속 시간(초)(4 또는 8, 기본값 4)model(선택 사항): 생성을 위한 모델 이름("vidu1.0", "vidu1.5", "vidu2.0", 기본값 "vidu2.0")resolution(선택 사항): 출력 비디오의 해상도("360p", "720p", "1080p", 기본값 "720p")movement_amplitude(선택 사항): 프레임 내 객체의 이동 진폭("auto", "small", "medium", "large", 기본값 "auto")seed(선택 사항): 재현성을 위한 무작위 시드
요청 예시:
{
"image_url": "https://example.com/image.jpg",
"prompt": "A serene lake with mountains in the background",
"duration": 8,
"model": "vidu2.0",
"resolution": "720p",
"movement_amplitude": "medium",
"seed": 12345
}2. 세대 상태 확인
실행 중인 비디오 생성 작업의 상태를 확인합니다.
매개변수:
task_id(필수): 이미지-비디오 도구에서 반환된 작업 ID
요청 예시:
{
"task_id": "12345abcde"
}3. 이미지 업로드
Vidu API와 함께 사용할 이미지를 업로드합니다.
매개변수:
image_path(필수): 이미지 파일의 로컬 경로image_type(필수): 이미지 파일 유형("png", "webp", "jpeg", "jpg")
요청 예시:
{
"image_path": "/path/to/your/image.jpg",
"image_type": "jpg"
}작동 원리
서버는 모델 컨텍스트 프로토콜(MCP)을 사용하여 AI 도구에 표준화된 인터페이스를 제공합니다. 서버를 시작하면 표준 입출력 채널을 통해 명령을 수신하고 구조화된 형식으로 결과를 응답합니다.
서버는 다음을 포함하여 Vidu API와 상호 작용하는 모든 복잡한 작업을 처리합니다.
API 키를 사용한 인증
파일 업로드 및 형식 검증
비동기 작업 관리 및 폴링
오류 처리 및 보고
문제 해결
API 키 문제 : Vidu API 키가
.env파일에 올바르게 설정되어 있는지 확인하세요.파일 업로드 오류 : 이미지 파일이 유효하고 크기가 10MB 이하인지 확인하세요.
연결 문제 : 인터넷에 접속할 수 있고 Vidu API 서버에 접속할 수 있는지 확인하세요.
기여하다
기여를 환영합니다! 풀 리퀘스트를 제출해 주세요.
Available Tools
3 toolscheck-generation-statusB
Check the status of a video generation task
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | Task ID returned by the image-to-video tool |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states the tool checks status but doesn't disclose behavioral traits like whether it's read-only, safe to call repeatedly, rate-limited, or what the response format might be (e.g., pending, completed, failed). This leaves significant gaps for an agent to understand how to interact with it effectively.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence that directly states the tool's purpose without any wasted words. It is appropriately sized and front-loaded, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a status-checking tool with no annotations and no output schema, the description is incomplete. It doesn't explain what statuses might be returned, error handling, or usage patterns (e.g., polling intervals), which are crucial for an agent to use this tool correctly in a workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the parameter 'task_id' fully described as 'Task ID returned by the image-to-video tool.' The description adds no additional parameter semantics beyond this, so it meets the baseline of 3 where the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose as checking the status of a video generation task, which is a specific verb (check) and resource (video generation task). However, it doesn't explicitly distinguish this from sibling tools like 'image-to-video' or 'upload-image' beyond the implied relationship through the task_id parameter description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by referencing 'video generation task,' and the parameter description mentions 'task_id returned by the image-to-video tool,' suggesting when to use it (after initiating a generation). However, it lacks explicit guidance on when not to use it or alternatives, such as whether it's for polling or one-time checks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
image-to-videoC
Generate a video from an image using Vidu API
| Name | Required | Description | Default |
|---|---|---|---|
| duration | No | Duration of the output video in seconds (4 or 8) | |
| image_url | Yes | URL of the image to convert to video | |
| model | No | Model name for generation | vidu2.0 |
| movement_amplitude | No | Movement amplitude of objects in the frame | auto |
| prompt | No | Text prompt for video generation (max 1500 chars) | |
| resolution | No | Resolution of the output video | 720p |
| seed | No | Random seed for reproducibility |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool generates a video but lacks details on execution time, rate limits, authentication needs, output format (e.g., video file type), error handling, or whether it's a synchronous/asynchronous operation. For a complex 7-parameter tool with no annotations, this is a significant gap in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without redundancy. It's front-loaded with the core action and resource, and every word earns its place by specifying the API used. No unnecessary details or fluff are included.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (7 parameters, video generation task) and lack of annotations and output schema, the description is incomplete. It doesn't cover behavioral aspects like performance, output details, or error handling, which are critical for an AI agent to use this tool effectively. The description alone is insufficient for a tool of this nature.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all 7 parameters with descriptions, defaults, and constraints. The description adds no parameter-specific information beyond what's in the schema, such as explaining interactions between parameters (e.g., how 'prompt' influences generation). Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Generate a video') and resource ('from an image'), specifying it uses the Vidu API. It distinguishes from sibling tools like 'check-generation-status' and 'upload-image' by focusing on video generation rather than status checking or image uploading. However, it doesn't explicitly differentiate from potential non-sibling alternatives beyond mentioning the API.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing an uploaded image first), when not to use it, or how it relates to sibling tools like 'check-generation-status' for monitoring generation progress. Usage is implied only by the tool name and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload-imageC
Upload an image to use with the Vidu API
| Name | Required | Description | Default |
|---|---|---|---|
| image_path | Yes | Local path to the image file | |
| image_type | Yes | Image file type |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden but only states the basic action. It doesn't disclose behavioral traits such as authentication needs, rate limits, error handling, or what happens after upload (e.g., returns an image ID). This leaves significant gaps for an agent to understand the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero waste. It's front-loaded and appropriately sized for a simple upload tool, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description is incomplete. It lacks details on what the tool returns, error conditions, or integration context with Vidu API. For a tool with two parameters and no structured behavioral data, this leaves the agent under-informed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents the two parameters. The description adds no additional meaning beyond implying the image is for Vidu API use, which is minimal value. Baseline 3 is appropriate as the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('upload') and resource ('an image'), specifying it's for use with the Vidu API. It doesn't differentiate from sibling tools like 'image-to-video' or 'check-generation-status', but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'image-to-video'. The description mentions the Vidu API context but doesn't specify prerequisites, constraints, or typical workflows, leaving usage unclear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.0- First observed
check-generation-status - First observed
image-to-video - First observed
upload-image
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: check-generation-status monitors task progress, image-to-video creates videos from images, and upload-image handles image uploads. There is no overlap in functionality, making tool selection straightforward.
The tools follow a consistent verb-object naming pattern (check-generation-status, image-to-video, upload-image), all using hyphens. However, the pattern is slightly inconsistent as 'image-to-video' uses a preposition 'to' while others do not, but it remains readable and predictable.
With 3 tools, the count is appropriate for a focused video generation API server, covering core operations. It is slightly lean but reasonable for the domain, as it includes upload, generation, and status checking without unnecessary bloat.
The tools cover basic video generation workflows: upload, generate, and check status. However, there are notable gaps such as missing operations for managing or deleting uploaded images, handling video outputs, or supporting other input types beyond images, which could limit agent capabilities.
Maintenance
Related MCP Connectors
MCP server for Google Veo AI video generation
MCP server for Wan AI video generation
MCP server for Luma Dream Machine AI video generation
MCP server for Hailuo (MiniMax) AI video generation
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceAI-powered image and video generation and processing server that supports text-to-image, image-to-image, text/image-to-video generation, image analysis, and comprehensive editing operations (crop, resize, convert, adjust) through providers like Doubao and Aliyun.6MIT
- AlicenseNot gradedqualityCmaintenanceMCP server for 4K video generation using Google VEO 3.1 — text-to-video, image-to-video, video extension, and frame interpolation.2MIT
- AlicenseAqualityDmaintenanceMCP server for Google Veo 3.1 video generation. Supports text/video/image-based generation, extension, and interpolation with cost estimation and batch processing.16433 npm1MIT
- AlicenseAqualityCmaintenanceMCP server for generating, editing, and batch processing videos using xAI's Grok Imagine Video API, with support for text-to-video, image-to-video, and video editing via natural language prompts.477 npm1MIT