Skip to main content
Glama

Omni-Video Studio MCP

El Omni-Video Studio MCP es un servidor de Model Context Protocol (MCP) autónomo de nivel empresarial que permite a cualquier IDE habilitado para LLM (Cursor, Claude Code, Antigravity) actuar como un editor de vídeo profesional.

Evoluciona la edición basada en transcripciones simples hacia un flujo de trabajo determinista, eficiente en tokens y basado en canalizaciones, que incluye gráficos en movimiento nativos para agentes (Hyperframes), proxies de metadatos visuales y renderizados finales de alta fidelidad.

🌟 Características principales

  1. Ingesta de proxies de metadatos: En lugar de transmitir costosos tokens de vídeo a un LLM, este servidor preprocesa el metraje para extraer un takes_packed.md (mapeo de audio) y un Grafo de Escena Visual. El agente edita utilizando proxies de texto, reduciendo costes y acelerando el razonamiento.

  2. Motor Hyperframes: Olvídate de las complejas dependencias de Node.js (p. ej., Remotion). El agente genera gráficos en movimiento deterministas en HTML/CSS que se renderizan instantáneamente a vídeo transparente utilizando Playwright.

  3. Canalización de renderizado avanzado: Impulsado por robustos grafos de filtros de FFmpeg, la salida final admite cortes EDL (Edit Decision List), renderizado de superposiciones, incrustación de subtítulos, corrección de color LUT y restauración de audio mediante IA DeepFilterNet opcional.

  4. Independiente del IDE: Debido a que cumple con la especificación oficial de MCP, se integra directamente en Cursor, Antigravity o Claude Desktop sin necesidad de complementos personalizados.

Related MCP server: omni-video-mcp

📦 Instalación

Requisitos previos:

  • python 3.10+

  • ffmpeg (debe estar instalado en la ruta de tu sistema)

  • uv (recomendado para la gestión de dependencias)

# Clone the repository
git clone https://github.com/your-org/omni-video-mcp.git
cd omni-video-mcp

# Install dependencies
uv venv
source .venv/bin/activate
uv pip install -e .

# Install Playwright browsers (for Hyperframes)
playwright install chromium

🛠 Configuración

Añade el servidor al archivo de configuración MCP de tu IDE (p. ej., ~/.gemini/antigravity/mcp_config.json, ~/.cursor/mcp.json o la configuración de Claude Desktop):

{
  "mcpServers": {
    "omni-video-mcp": {
      "command": "uv",
      "args": [
        "run",
        "/path/to/omni-video-mcp/server.py"
      ],
      "env": {
        "ELEVENLABS_API_KEY": "your_api_key_here" 
      }
    }
  }
}

Nota: Actualmente se requiere la ELEVENLABS_API_KEY para el mapeo de transcripción a nivel de palabra de alta fidelidad durante la ingesta.

🎬 Cómo funciona (La canalización del agente)

Cuando el agente utiliza este servidor MCP, sigue una arquitectura de 4 fases:

  1. Fase 1: Ingesta (omni_video_ingest) El agente escanea tus archivos .mp4 / .mov sin procesar, extrayendo una transcripción en markdown empaquetada y un Grafo de Escena Visual inicial.

  2. Fase 2: Montaje del director (omni_video_preview) El agente utiliza la transcripción para construir una EDL (Edit Decision List) de las mejores tomas. Los cortes ambiguos pueden verificarse visualmente generando tiras de película en PNG a través de la herramienta de vista previa.

  3. Fase 3: VFX (omni_video_generate_vfx) El agente genera gráficos en movimiento en HTML/CSS (tercios inferiores, diseños de b-roll) y el servidor los renderiza de forma determinista en vídeos .webm transparentes a través de Hyperframes.

  4. Fase 4: Postproducción y renderizado (omni_video_render) El agente pasa la EDL, las marcas de tiempo de VFX y la configuración de renderizado al servidor, el cual construye un grafo complejo de FFmpeg para concatenar el metraje, corregir el color, restaurar el audio y exportar un máster final.

🤝 Contribuciones

¡Las contribuciones son bienvenidas! Si vas a añadir nuevas capacidades a la canalización de renderizado (como seguimiento automático o alternativas locales de whisper), por favor abre un PR. Asegúrate de que cualquier dependencia de Python añadida se incluya en el pyproject.toml usando uv add <package>.

📄 Licencia

Licencia MIT

Available Tools

4 tools
omni_video_generate_vfxC

Renders motion graphics (e.g., lower thirds, titles) using Hyperframes. Returns the path to the rendered transparent .mov or .webm file.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavioral traits. It states the tool renders transparent videos using Hyperframes and returns a file path, but fails to mention whether it is synchronous or async, any size limits, error conditions, or side effects. The minimal detail leaves the agent guessing about important behaviors.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise—two sentences covering the core action and output. It front-loads the primary purpose and wastes no words. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and an output schema that likely exists but is not described, the tool description is incomplete. It does not explain the return format beyond a file path, nor does it cover prerequisites for HTML/CSS validity or duration limits. The agent may need to infer or experiment to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has a single parameter 'request' with nested properties (html_content, css_content, duration_seconds), all with schema descriptions. However, the context indicates 0% schema description coverage, likely because the top-level parameter lacks a description. The tool description itself does not mention any parameters or add meaning beyond what the schema provides. For parameter understanding, the agent gets no extra help from the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool renders motion graphics (lower thirds, titles) using Hyperframes and returns the path to a transparent .mov or .webm file. The verb 'renders' and specific resource 'motion graphics' distinguish it from sibling tools like omni_video_ingest, omni_video_preview, and omni_video_render.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, nor when not to use it. It does not mention any prerequisites or context for invoking it. With sibling tools listed but no differentiation, the agent lacks contextual cues for proper selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

omni_video_ingestA

Ingests a directory of video files, generates word-level audio transcripts, and constructs a semantic Visual Scene Graph for B-Roll searching. Returns the path to the generated project metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the transparency burden. It discloses that transcripts and scene graphs are generated and a metadata path is returned, but does not mention whether ingestion modifies source files, requires permissions, or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences long, front-loaded with the main action, and every sentence provides essential information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is non-trivial (multiple stages), but the output schema exists to clarify return values. The description covers the main inputs and outputs, though it lacks edge-case or error details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the top-level parameter. The tool description adds that 'directory_path' is a directory of video files, but the schema already includes a similar description for 'directory_path'. Thus, the description adds minimal value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verbs: 'ingests', 'generates word-level audio transcripts', 'constructs a semantic Visual Scene Graph'. It distinguishes well from sibling tools which handle VFX, preview, and render.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage context is implied (ingesting video files for transcript and scene graph generation), but no explicit guidance on when to use versus alternatives or prerequisites is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

omni_video_previewA

Generates a filmstrip PNG of the specified video segment. Useful for visually verifying cut boundaries or B-roll placement. Returns the absolute path to the generated PNG file.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that the tool generates a PNG and returns its path, implying a read-only operation. However, it does not detail any side effects, permissions, or safety implications.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, each serving a distinct purpose: what it does, when to use it, and what it returns. No wasted words, front-loaded with core action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the relatively simple operation and the presence of a nested parameter schema with descriptions, plus an output schema (not shown but indicated), the description adequately covers the tool's purpose and return value. It could mention the parameter structure briefly, but the schema covers that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the tool description does not explain any parameters. While the schema itself contains descriptions for file_path, start_time, and end_time, the description adds no additional meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Generates a filmstrip PNG of the specified video segment,' which is a specific verb-resource combination. It distinguishes from siblings (generate_vfx, ingest, render) by focusing on preview.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit use cases: 'visually verifying cut boundaries or B-roll placement.' It does not mention when not to use or alternatives, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

omni_video_renderA

Orchestrates the final render pipeline: applies EDL cuts, overlay graphics, LUT color grading, audio restoration, and subtitle burning. Returns the path to the final rendered video.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries behavioral disclosure. It accurately describes the tool's actions (applying cuts, LUT, etc.) and the return value (path to final video). It does not mention side effects like file creation or potential errors, but the output_path parameter implies file writing. This is adequate but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that efficiently lists all main operations and the return value. It is front-loaded with the core purpose ('Orchestrates the final render pipeline') and contains no fluff or redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the tool's core functionality and return value, and the schema outlines required parameters. It lacks details about error handling or prerequisites, but given the presence of an output schema (implied) and sibling context, it is mostly complete for a final render step.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema properties have detailed descriptions (e.g., edl_json, lut_path, output_path), which already clarify parameter meanings. The tool description does not add additional parameter information beyond listing the tool's capabilities. With good schema coverage, a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool orchestrates the final render pipeline and enumerates specific operations (EDL cuts, overlay graphics, LUT grading, audio restoration, subtitle burning). It distinguishes from siblings through the term 'final render,' differentiating it from ingest, VFX generation, and preview tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this is the last step in a video processing pipeline, but does not explicitly state when to use it versus alternatives or mention prerequisites. The mention of 'final render' provides some context, but more direct guidance would improve this dimension.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedomni_video_generate_vfx
    • First observedomni_video_ingest
    • First observedomni_video_preview
    • First observedomni_video_render

TDQS

A4/5.0

Scored across 4 tools

Disambiguation5/5

Each tool serves a distinct function: ingest creates metadata, generate_vfx produces graphics, preview provides visual verification, and render finalizes output. No overlapping purposes.

Naming Consistency5/5

All tools follow the 'omni_video_' prefix with a verb describing the action (ingest, generate_vfx, preview, render), forming a clear and predictable pattern.

Tool Count5/5

Four tools cover the essential phases of video production—ingestion, effect creation, previewing, and rendering—without being excessive or insufficient.

Completeness4/5

The tool set covers the core pipeline, but lacks explicit tools for timeline editing or asset selection, requiring reliance on the render tool for EDL processing.

Maintenance

ActivitySlowing
ResponsivenessSyncing

Related MCP Connectors

Related MCP Servers