Skip to main content
Glama

Servidor Vidu MCP

insignia de herrería

Un servidor de Protocolo de Contexto de Modelo (MCP) para interactuar con la API de generación de video de Vidu. Este servidor proporciona herramientas para generar videos a partir de imágenes utilizando los potentes modelos de IA de Vidu.

Características

  • Conversión de imagen a vídeo : genere vídeos a partir de imágenes estáticas con configuraciones personalizables

  • Verificar el estado de generación : supervisa el progreso de las tareas de generación de video

  • Carga de imágenes : cargue imágenes fácilmente para usarlas con la API de Vidu

Related MCP server: veo-mcp-server

Prerrequisitos

  • Node.js (v14 o superior)

  • Una clave API de Vidu (disponible en el sitio web de Vidu )

  • TypeScript (para desarrollo)

Instalación

Instalación mediante herrería

Para instalar Vidu Video Generation Server para Claude Desktop automáticamente a través de Smithery :

npx -y @smithery/cli install @el-el-san/vidu-mcp-server --client claude

Instalación manual

  1. Clonar este repositorio:

git clone https://github.com/el-el-san/vidu-mcp-server.git
cd vidu-mcp-server
  1. Instalar dependencias:

npm install
  1. Cree un archivo .env basado en .env.template y agregue su clave API de Vidu:

VIDU_API_KEY=your_api_key_here

Uso

  1. Construya el código TypeScript:

npm run build
  1. Iniciar el servidor:

npm start

El servidor MCP se iniciará y estará listo para aceptar conexiones de clientes MCP.

Herramientas

1. Imagen a vídeo

Convierte una imagen estática en un vídeo con parámetros personalizables.

Parámetros:

  • image_url (obligatorio): URL de la imagen a convertir a vídeo

  • prompt (opcional): Aviso de texto para la generación de vídeo (máximo 1500 caracteres)

  • duration (opcional): Duración del vídeo de salida en segundos (4 u 8, predeterminado 4)

  • model (opcional): Nombre del modelo para la generación ("vidu1.0", "vidu1.5", "vidu2.0", predeterminado "vidu2.0")

  • resolution (opcional): Resolución del vídeo de salida ("360p", "720p", "1080p", predeterminado "720p")

  • movement_amplitude (opcional): Amplitud de movimiento de los objetos en el marco ("auto", "pequeño", "mediano", "grande", predeterminado "auto")

  • seed (opcional): semilla aleatoria para reproducibilidad

Ejemplo de solicitud:

{
  "image_url": "https://example.com/image.jpg",
  "prompt": "A serene lake with mountains in the background",
  "duration": 8,
  "model": "vidu2.0",
  "resolution": "720p",
  "movement_amplitude": "medium",
  "seed": 12345
}

2. Verificar el estado de la generación

Comprueba el estado de una tarea de generación de vídeo en ejecución.

Parámetros:

  • task_id (obligatorio): ID de tarea devuelto por la herramienta de imagen a video

Ejemplo de solicitud:

{
  "task_id": "12345abcde"
}

3. Subir imagen

Sube una imagen para usar con la API de Vidu.

Parámetros:

  • image_path (obligatorio): Ruta local al archivo de imagen

  • image_type (obligatorio): Tipo de archivo de imagen ("png", "webp", "jpeg", "jpg")

Ejemplo de solicitud:

{
  "image_path": "/path/to/your/image.jpg",
  "image_type": "jpg"
}

Cómo funciona

El servidor utiliza el Protocolo de Contexto de Modelo (MCP) para proporcionar una interfaz estandarizada para herramientas de IA. Al iniciar el servidor, este escucha comandos a través de canales de entrada/salida estándar y responde con resultados en un formato estructurado.

El servidor gestiona toda la complejidad de la interacción con la API de Vidu, incluyendo:

  • Autenticación con claves API

  • Carga de archivos y validación de formato

  • Gestión de tareas asincrónicas y sondeo

  • Manejo y reporte de errores

Solución de problemas

  • Problemas con la clave API : asegúrese de que su clave API de Vidu esté configurada correctamente en el archivo .env

  • Errores de carga de archivos : Verifique que sus archivos de imagen sean válidos y tengan un tamaño inferior a 10 MB

  • Problemas de conexión : asegúrese de tener acceso a Internet y poder acceder a los servidores de la API de Vidu

Contribuyendo

¡Agradecemos sus contribuciones! No dude en enviar una solicitud de incorporación de cambios.

Available Tools

3 tools
check-generation-statusB

Check the status of a video generation task

ParametersJSON Schema
NameRequiredDescriptionDefault
task_idYesTask ID returned by the image-to-video tool

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It states the tool checks status but doesn't disclose behavioral traits like whether it's read-only, safe to call repeatedly, rate-limited, or what the response format might be (e.g., pending, completed, failed). This leaves significant gaps for an agent to understand how to interact with it effectively.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence that directly states the tool's purpose without any wasted words. It is appropriately sized and front-loaded, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of a status-checking tool with no annotations and no output schema, the description is incomplete. It doesn't explain what statuses might be returned, error handling, or usage patterns (e.g., polling intervals), which are crucial for an agent to use this tool correctly in a workflow.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with the parameter 'task_id' fully described as 'Task ID returned by the image-to-video tool.' The description adds no additional parameter semantics beyond this, so it meets the baseline of 3 where the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose as checking the status of a video generation task, which is a specific verb (check) and resource (video generation task). However, it doesn't explicitly distinguish this from sibling tools like 'image-to-video' or 'upload-image' beyond the implied relationship through the task_id parameter description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context by referencing 'video generation task,' and the parameter description mentions 'task_id returned by the image-to-video tool,' suggesting when to use it (after initiating a generation). However, it lacks explicit guidance on when not to use it or alternatives, such as whether it's for polling or one-time checks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

image-to-videoC

Generate a video from an image using Vidu API

ParametersJSON Schema
NameRequiredDescriptionDefault
durationNoDuration of the output video in seconds (4 or 8)
image_urlYesURL of the image to convert to video
modelNoModel name for generationvidu2.0
movement_amplitudeNoMovement amplitude of objects in the frameauto
promptNoText prompt for video generation (max 1500 chars)
resolutionNoResolution of the output video720p
seedNoRandom seed for reproducibility

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool generates a video but lacks details on execution time, rate limits, authentication needs, output format (e.g., video file type), error handling, or whether it's a synchronous/asynchronous operation. For a complex 7-parameter tool with no annotations, this is a significant gap in transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose without redundancy. It's front-loaded with the core action and resource, and every word earns its place by specifying the API used. No unnecessary details or fluff are included.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (7 parameters, video generation task) and lack of annotations and output schema, the description is incomplete. It doesn't cover behavioral aspects like performance, output details, or error handling, which are critical for an AI agent to use this tool effectively. The description alone is insufficient for a tool of this nature.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents all 7 parameters with descriptions, defaults, and constraints. The description adds no parameter-specific information beyond what's in the schema, such as explaining interactions between parameters (e.g., how 'prompt' influences generation). Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Generate a video') and resource ('from an image'), specifying it uses the Vidu API. It distinguishes from sibling tools like 'check-generation-status' and 'upload-image' by focusing on video generation rather than status checking or image uploading. However, it doesn't explicitly differentiate from potential non-sibling alternatives beyond mentioning the API.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing an uploaded image first), when not to use it, or how it relates to sibling tools like 'check-generation-status' for monitoring generation progress. Usage is implied only by the tool name and description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload-imageC

Upload an image to use with the Vidu API

ParametersJSON Schema
NameRequiredDescriptionDefault
image_pathYesLocal path to the image file
image_typeYesImage file type

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden but only states the basic action. It doesn't disclose behavioral traits such as authentication needs, rate limits, error handling, or what happens after upload (e.g., returns an image ID). This leaves significant gaps for an agent to understand the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero waste. It's front-loaded and appropriately sized for a simple upload tool, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description is incomplete. It lacks details on what the tool returns, error conditions, or integration context with Vidu API. For a tool with two parameters and no structured behavioral data, this leaves the agent under-informed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents the two parameters. The description adds no additional meaning beyond implying the image is for Vidu API use, which is minimal value. Baseline 3 is appropriate as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('upload') and resource ('an image'), specifying it's for use with the Vidu API. It doesn't differentiate from sibling tools like 'image-to-video' or 'check-generation-status', but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like 'image-to-video'. The description mentions the Vidu API context but doesn't specify prerequisites, constraints, or typical workflows, leaving usage unclear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.0.0
    • First observedcheck-generation-status
    • First observedimage-to-video
    • First observedupload-image

TDQS

B3.3/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: check-generation-status monitors task progress, image-to-video creates videos from images, and upload-image handles image uploads. There is no overlap in functionality, making tool selection straightforward.

Naming Consistency4/5

The tools follow a consistent verb-object naming pattern (check-generation-status, image-to-video, upload-image), all using hyphens. However, the pattern is slightly inconsistent as 'image-to-video' uses a preposition 'to' while others do not, but it remains readable and predictable.

Tool Count4/5

With 3 tools, the count is appropriate for a focused video generation API server, covering core operations. It is slightly lean but reasonable for the domain, as it includes upload, generation, and status checking without unnecessary bloat.

Completeness3/5

The tools cover basic video generation workflows: upload, generate, and check status. However, there are notable gaps such as missing operations for managing or deleting uploaded images, handling video outputs, or supporting other input types beyond images, which could limit agent capabilities.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers