media-mcp
Officialmedia-mcp (Node.js)
中文 | English
Servicio de mejora de vídeo y segmentación de imágenes basado en el protocolo MCP, que actúa como cliente-servidor MCP interactuando con un servidor HTTP backend.
Funcionalidades
Proporciona las siguientes herramientas MCP:
create_task- Crea una tarea de mejora de vídeo (admite URL o carga de archivos locales)get_task_status- Consulta el estado de la tareaenhance_video_sync- Mejora de vídeo síncrona (bloquea hasta completar)sam3_predict- Segmentación de imágenes SAM3 (admite rutas locales, URL o imágenes en Base64)
Related MCP server: Grok Imagine Video MCP Server
Requisitos previos
Node.js >= 18 (verificar:
node --version)API Key (para autenticación, contacte al proveedor del servicio para obtenerla)
Instalación rápida (recomendada)
Si su agente de IA tiene una ruta de configuración MCP definida, simplemente copie y pegue la siguiente frase para enviársela a la IA:
帮我安装 npm 包 @avclabs.ai/media-mcp 作为 MCP server。我的 API Key 是:sk-xxxxxxxx。La IA completará automáticamente:
La detección de su cliente MCP.
La localización de la ruta del archivo de configuración.
La escritura de la configuración correcta.
La solicitud de reinicio del cliente.
Instalación manual
No requiere instalación, ejecútelo directamente en la configuración del cliente MCP usando npx.
1. Claude Code (CLI)
Ejecute en Claude Code:
/mcpRevise la salida para encontrar la ruta del archivo de configuración correspondiente a "User MCPs" y edite dicho archivo.
Rutas comunes (si /mcp no está disponible):
Windows:
%USERPROFILE%\.claude.jsonmacOS:
~/.claude.jsonLinux:
~/.claude.jsonVersión antigua/alternativa:
~/.claude/mcp.json
Pegue el siguiente contenido (reemplace your-api-key con su API Key real):
{
"mcpServers": {
"video-enhancement": {
"command": "npx",
"args": ["-y", "@avclabs.ai/media-mcp@latest"],
"env": {
"API_KEY": "your-api-key"
}
}
}
}Después de guardar, ejecute /mcp para verificar si se cargó correctamente.
2. Cursor
Vaya a Settings > Tools & MCPs > Add New MCP Server:
Name:
video-enhancementType:
commandCommand:
env HTTP_API_KEY=your-api-key npx -y @avclabs.ai/media-mcp@latest
O edite ~/.cursor/mcp.json:
{
"mcpServers": {
"video-enhancement": {
"command": "npx",
"args": ["-y", "@avclabs.ai/media-mcp@latest"],
"env": {
"API_KEY": "your-api-key"
}
}
}
}Verificar la instalación
Tras reiniciar el cliente, confirme si las herramientas se cargaron correctamente:
O pregunte directamente a la IA: "¿Qué herramientas tienes disponibles?"
Debería ver:
create_task,get_task_status,enhance_video_sync,sam3_predict
Parámetros de configuración
Nombre de variable | Obligatorio | Valor predeterminado | Descripción |
| Sí | - | Clave de autenticación API (compartida para mejora de vídeo y SAM3) |
| No |
| URL de la interfaz del servicio de mejora de vídeo |
| No |
| URL de la interfaz del servicio SAM3 |
| No |
| Intervalo de sondeo (ms) |
| No |
| Número máximo de intentos de sondeo |
URL de servicio personalizada
{
"env": {
"HTTP_API_BASE_URL": "https://your-endpoint.com",
"API_KEY": "your-api-key",
"SAM3_API_BASE_URL": "http://localhost:8001"
}
}O mediante argumentos de línea de comandos:
npx -y @avclabs.ai/media-mcp@latest --base-url https://your-endpoint.com --api-key your-api-key --sam3-base-url http://localhost:8001Ejemplos de uso
Una vez configurado, hable con la IA en lenguaje natural:
"Ayúdame a mejorar este vídeo a 1080p: https://example.com/video.mp4"
"Mejora la calidad de video.mp4 de mi escritorio a 2k"
La IA llamará automáticamente a la herramienta correspondiente para completar la tarea.
"Ayúdame a analizar esta imagen y encontrar todos los objetos: C:\Users\xxx\photo.png"
"Segmenta esta imagen usando SAM3, el prompt es 'find all cars'"
Herramientas proporcionadas
create_task
Crea una tarea de mejora de vídeo (asíncrona).
Parámetro | Tipo | Obligatorio | Valor predeterminado | Descripción |
| string | Sí | - | URL del vídeo o ruta del archivo local (la URL debe ser accesible públicamente, no se admiten enlaces que requieran inicio de sesión o firma) |
| string | No |
|
|
| string | No |
|
|
Valor de retorno:
{
"success": true,
"task_id": "xxx",
"status": "wait"
}get_task_status
Consulta el estado de la tarea.
Parámetro | Tipo | Obligatorio |
| string | Sí |
Valor de retorno:
{
"success": true,
"task_id": "xxx",
"status": "completed",
"progress": 100,
"video_url": "https://..."
}enhance_video_sync
Mejora de vídeo síncrona (bloquea hasta completar).
Parámetro | Tipo | Obligatorio | Valor predeterminado | Descripción |
| string | Sí | - | URL del vídeo o ruta del archivo local (la URL debe ser accesible públicamente) |
| string | No |
|
|
| string | No |
| Resolución objetivo |
| number | No |
| Intervalo de sondeo (segundos) |
| number | No |
| Tiempo de espera (segundos) |
sam3_predict
Utiliza la API de segmentación SAM3 para analizar imágenes y generar resultados de inferencia (máscaras, cajas, puntuaciones).
Parámetros:
Entrada de imagen (elija una de las tres, debe proporcionar al menos una):
imagePath(string): Ruta absoluta de la imagen local. Admite formatos comunes (PNG, JPG, JPEG).Ejemplo:
"C:\\Users\\xxx\\photo.png","/home/user/images/cat.jpg"Escenario: El usuario proporciona explícitamente una ruta de archivo local.
imageUrl(string): URL de imagen accesible públicamente.Ejemplo:
"https://example.com/photo.jpg"Escenario: La imagen está en línea y el usuario proporciona el enlace.
Nota: La URL debe ser accesible públicamente; no se admiten enlaces que requieran inicio de sesión o firma.
imageBase64(string): Datos de imagen codificados en Base64.Ejemplo:
"iVBORw0KGgoAAAANSUhEUgAA..."Escenario: El usuario arrastra o carga un archivo adjunto, y el agente codifica la imagen en base64.
Nota: Los datos base64 de imágenes grandes pueden ser pesados y el tiempo de transmisión puede ser mayor.
Otros parámetros:
prompt(string, required): Prompt de texto en inglés para especificar el objeto objetivo a segmentar. Por ejemplo,"person","car","a cat sitting on a sofa". Dado que el modelo SAM3 solo acepta prompts en inglés, se recomienda enviar descripciones en inglés. Si el usuario proporciona texto en otro idioma, el agente lo traducirá automáticamente antes de llamar a la API.
Retorno:
Tras completar la inferencia, devuelve directamente una cadena JSON. Este JSON contiene los siguientes tres campos:
masks: Matriz bidimensional. Cada elemento es una máscara binaria (0 o 1) del mismo tamaño que la imagen de entrada, marcando la posición a nivel de píxel del objeto detectado. La i-ésima máscara corresponde a la i-ésima instancia detectada.boxes: Matriz bidimensional. Cada elemento es una coordenada de caja delimitadora en formato[x1, y1, x2, y2], representando el área rectangular del objeto.x1, y1son las coordenadas superior izquierda,x2, y2son las coordenadas inferior derecha.Explicación del sistema de coordenadas: El origen
(0, 0)está en la esquina superior izquierda de la imagen, el ejexcrece hacia la derecha y el ejeyhacia abajo, en píxeles.scores: Matriz unidimensional. Cada elemento es la puntuación de confianza del resultado de detección, de 0 a 1. Cuanto mayor sea la puntuación, mayor es la confianza del modelo.
Ejemplo de contenido JSON resultante:
{
"masks": [
[[0, 0, 1, ...], [0, 1, 1, ...], ...],
[[0, 0, 0, ...], [0, 0, 1, ...], ...]
],
"boxes": [
[120, 80, 300, 450],
[400, 200, 600, 500]
],
"scores": [0.95, 0.87]
}Preguntas frecuentes
¿Recibo un error de "archivo no encontrado" tras arrastrar un adjunto?
Esta es una limitación conocida de stdio MCP. Al arrastrar o cargar archivos, la ruta a menudo no se pasa automáticamente al servidor MCP.
Solución:
Proporcione la ruta manualmente (recomendado): Después de arrastrar la imagen, añada en el texto la ruta absoluta local:
"Por favor, procesa esta imagen
D:\photos\cat.jpg, encuentra el gato"Espere la codificación automática: Claude puede codificar la imagen en base64 automáticamente. Si tiene éxito, no se requiere acción adicional.
Responda a la solicitud de ruta: Si Claude pregunta por la ruta, responda directamente con la ruta absoluta local.
¿Hay prioridad entre los tres métodos de entrada?
No hay una prioridad estricta. Claude elegirá automáticamente el método más adecuado según el contexto de la conversación.
¿Qué formatos de imagen se admiten?
Formatos comunes: PNG, JPG, JPEG, BMP, WebP, etc. Se recomienda usar PNG o JPG.
¿Qué hacer si la descarga de la imagen por URL falla?
Asegúrese de que la URL sea públicamente accesible y no requiera inicio de sesión, cookies o firmas. Si la imagen está en un servicio privado (como un S3 privado), descárguela primero localmente y use imagePath.
¿Qué hacer si la imagen Base64 es demasiado grande?
Si la imagen es muy grande (ej. resolución 4K), los datos base64 serán muy pesados. Se recomienda:
Usar
imagePathen su lugar.Comprimir la imagen antes de codificarla.
Instrucciones de carga de archivos
Cuando type es "local", el servidor MCP:
Lee el archivo local.
Lo sube directamente al almacenamiento de objetos TOS mediante una URL prefirmada.
Tamaño máximo de archivo: 100MB.
Solución de problemas
"command not found: npx"
Instale Node.js >= 18: https://nodejs.org/
"Error: Se requiere --api-key o configurar API_KEY"
Falta la API Key, verifique env.API_KEY en la configuración.
El servidor MCP aparece en rojo/error en el cliente
Revise los registros:
Claude Desktop macOS:
~/Library/Logs/Claude/mcp*.logClaude Desktop Windows:
%APPDATA%\Claude\logs\mcp*.logCursor: Panel de salida > MCP
"Error de carga en TOS"
Generalmente es un error de firma; confirme que HTTP_API_BASE_URL y HTTP_API_KEY sean correctos y válidos.
Instalación global (opcional)
Si no desea usar npx cada vez:
npm install -g @avclabs.ai/media-mcpLuego, en la configuración, use "command": "media-mcp" junto con "args": ["--api-key", "your-api-key"].
Licencia
Licencia MIT - Ver el archivo LICENSE para más detalles.
Available Tools
4 toolscreate_taskB
创建视频增强任务(异步)
支持两种上传方式:
URL 上传:提供视频 URL
本地上传:提供本地文件路径,MCP Server 自动上传到 TOS 对象存储
参数说明:
video_source: 视频 URL 或本地文件路径
type: "url" 或 "local"
resolution: 目标分辨率
| Name | Required | Description | Default |
|---|---|---|---|
| video_source | Yes | 视频URL地址或本地文件路径(URL必须公网可访问,不支持需要登录或签名的链接) | |
| type | No | 上传类型:url=网络视频,local=本地文件 | url |
| resolution | No | 目标分辨率,默认720p | 720p |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so the description must carry full burden. It notes async behavior and TOS upload but omits side effects, permissions, failure modes, or rate limits. The description only partially discloses behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with three bullet points and front-loaded purpose. Every sentence earns its place, but structure could be slightly improved with clearer differentiation from siblings.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers input parameters well but lacks output schema explanation (e.g., task ID or status). With no annotations and multiple siblings, more context on post-creation steps would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has 100% description coverage, so baseline is 3. The description groups parameters and explains the two upload modes, but does not add new information beyond the schema's existing parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool creates an async video enhancement task and distinguishes between two upload methods (URL and local). It uses specific verbs and resources, and is not a tautology.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use each upload type (URL vs local) but does not explicitly guide when to use this async tool over its sync sibling (enhance_video_sync) or other tools like get_task_status.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
enhance_video_syncA
同步增强视频(阻塞等待完成)
支持两种上传方式:
URL 上传:提供视频 URL
本地上传:提供本地文件路径,MCP Server 自动上传到 TOS 对象存储
参数说明:
video_source: 视频 URL 或本地文件路径
type: "url" 或 "local"
resolution: 目标分辨率
poll_interval: 轮询间隔(秒)
timeout: 超时时间(秒)
| Name | Required | Description | Default |
|---|---|---|---|
| video_source | Yes | 视频URL地址或本地文件路径(URL必须公网可访问,不支持需要登录或签名的链接) | |
| type | No | 上传类型:url=网络视频,local=本地文件 | url |
| resolution | No | 目标分辨率,默认720p | 720p |
| poll_interval | No | 轮询间隔(秒),默认5 | |
| timeout | No | 超时时间(秒),默认600 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears full responsibility for behavioral disclosure. It explicitly states 'blocking wait for completion', explains the automatic upload of local files to TOS storage, and mentions polling parameters, giving good transparency. It does not mention side effects, but given the nature, none are expected.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with bullet points and clear categorization of upload methods and parameters. Every sentence serves a purpose, and there is no redundancy or unnecessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the blocking nature, upload methods, and all parameters thoroughly. It lacks an explicit description of the return value, but given the synchronous nature, it likely returns the enhanced video. Overall, it is fairly complete for a tool without an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds context beyond the schema, such as the automatic upload process for local files and that URLs must be publicly accessible. This extra information enhances understanding of parameter usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: to enhance video synchronously, with a blocking wait. It details two upload methods (URL and local), distinguishing it from sibling tools that handle different operations like task creation or status checking.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the two upload methods and the blocking nature, implicitly indicating when to use the tool. However, it does not explicitly contrast with siblings like create_task (likely async) or provide when-not-to-use guidance, making usage guidelines less explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_task_statusA
查询视频增强任务状态
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | 任务ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, and the description only states the purpose without disclosing any behavioral traits such as polling requirements, rate limits, or expected response behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence with no wasted words; efficient and to the point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple status query with one parameter, the description is mostly complete but could benefit from mentioning possible return statuses or output format since no output schema is provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the description adds no additional meaning beyond what is already in the input schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'query' and the resource 'video enhancement task status', distinguishing from sibling tools 'create_task' and 'enhance_video_sync' which have different actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives, but the context of sibling tools implies it is for checking status after creation or enhancement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sam3_predictA
Analyze an image using the SAM3 segmentation API to generate inference results (masks, boxes, scores). The image can be provided in one of three ways:
imagePath: Absolute path of a local image file (e.g. C:\Users\xxx\photo.png). Use this when the user provides a local file path.
imageUrl: Publicly accessible URL of the image (e.g. https://example.com/photo.jpg). Use this when the user provides a web link.
imageBase64: Base64-encoded image data. Use this when the user uploads or drags-and-drops an image as an attachment and no local path is available. In this case, encode the image content as base64 and pass it via this parameter. If the user mentions an uploaded image but does not provide a path, URL, or base64 data, ask the user for the local absolute path. Prompt must be in English. If the user provides Chinese or other non-English text, translate it to English before calling this tool.
| Name | Required | Description | Default |
|---|---|---|---|
| imagePath | No | Absolute path of a local image file (e.g. C:\\Users\\xxx\\photo.png) | |
| imageUrl | No | Publicly accessible URL of the image to process | |
| imageBase64 | No | Base64-encoded image data. Use this when the image is provided as an attachment without a local path | |
| prompt | Yes | Text prompt for mask generation. Must be in English. If the user provides Chinese or other non-English text, translate it to English before calling this tool |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses that it calls an external API and generates masks, boxes, scores. However, it lacks details on potential side effects, authentication, error handling, or rate limits. It does not contradict annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with bullet points, front-loading the main purpose. Every sentence serves a purpose, explaining input methods and prompt requirements without redundancy. It is concise yet comprehensive.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers what the tool does (segmentation analysis), how to provide input (three methods), and what outputs are generated (masks, boxes, scores). Even without an output schema, it gives sufficient information for an agent to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 100%, the description adds significant value by explaining usage contexts for each image parameter and specifying that the prompt must be in English, requiring translation if needed. This goes beyond schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Analyze an image using the SAM3 segmentation API to generate inference results (masks, boxes, scores).' This specifies the verb (analyze), resource (image via SAM3 API), and output, effectively distinguishing it from siblings like create_task.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use each image input method (imagePath, imageUrl, imageBase64) and includes instructions for handling non-English prompts. However, it does not explicitly mention when not to use this tool or compare it to alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
create_task - First observed
enhance_video_sync - First observed
get_task_status - First observed
sam3_predict
TDQS
Scored across 4 tools
The first three tools are about video enhancement tasks with overlapping functionality (create_task and enhance_video_sync both appear to initiate enhancement), and the fourth tool (sam3_predict) is for image segmentation, a completely different domain. The descriptions are not clear enough to distinguish which tool to use for a given task, causing confusion.
Tool names partially follow a verb_noun pattern (create_task, get_task_status), but 'enhance_video_sync' is awkward and 'sam3_predict' mixes model name with verb, introducing inconsistency.
With 4 tools, the count is reasonable for a focused server, but the server actually combines two unrelated capabilities (video enhancement and image segmentation), making the scope unclear but the number itself is not extreme.
For video enhancement, there are create, sync enhance, and status query, but missing cancel, list, or delete operations. For image segmentation, only a single predict tool exists. The surface is incomplete for both domains.
Maintenance
Related MCP Connectors
MCP server for Wan AI video generation
MCP server for Google Veo AI video generation
MCP server for MiniMax H3 multimodal video generation
MCP server for Luma Dream Machine AI video generation
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceMCP server for 4K video generation using Google VEO 3.1 — text-to-video, image-to-video, video extension, and frame interpolation.1MIT
- AlicenseAqualityBmaintenanceMCP server for generating, editing, and batch processing videos using xAI's Grok Imagine Video API, with support for text-to-video, image-to-video, and video editing via natural language prompts.438 npm1MIT
- AlicenseAqualityDmaintenanceMCP server for AI-powered media generation: images, videos, audio, and upscaling using 99 AI models.6MIT
- AlicenseNot gradedqualityDmaintenanceAn MCP server enabling video processing via natural language: transcription with Whisper, segment cutting with FFmpeg, and file management.MIT