Skip to main content
Glama

Omni-Video Studio MCP

Omni-Video Studio MCP — это сервер корпоративного уровня, работающий по протоколу Model Context Protocol (MCP), который позволяет любой IDE с поддержкой LLM (Cursor, Claude Code, Antigravity) выступать в роли профессионального видеоредактора.

Он превращает простое редактирование на основе транскрипции в детерминированный, эффективный с точки зрения токенов и управляемый конвейером рабочий процесс, включающий агентную анимированную графику (Hyperframes), прокси визуальных метаданных и высококачественный финальный рендеринг.

🌟 Ключевые особенности

  1. Загрузка прокси метаданных: Вместо передачи дорогостоящих видеотокенов в LLM, этот сервер предварительно обрабатывает отснятый материал для извлечения takes_packed.md (сопоставление аудио) и графа визуальных сцен. Агент редактирует видео, используя текстовые прокси, что снижает затраты и ускоряет процесс принятия решений.

  2. Движок Hyperframes: Забудьте о сложных зависимостях Node.js (например, Remotion). Агент генерирует детерминированную HTML/CSS анимированную графику, которая мгновенно преобразуется в видео с прозрачностью с помощью Playwright.

  3. Продвинутый конвейер рендеринга: Благодаря мощным фильтрам FFmpeg, финальный вывод поддерживает монтаж по EDL (Edit Decision List), наложение графики, вшивание субтитров, цветокоррекцию LUT и опциональное восстановление аудио с помощью ИИ DeepFilterNet.

  4. Независимость от IDE: Поскольку сервер соответствует официальной спецификации MCP, он легко интегрируется в Cursor, Antigravity или Claude Desktop без необходимости в специальных плагинах.

Related MCP server: omni-video-mcp

📦 Установка

Предварительные требования:

  • python 3.10+

  • ffmpeg (должен быть установлен в системном пути)

  • uv (рекомендуется для управления зависимостями)

# Clone the repository
git clone https://github.com/your-org/omni-video-mcp.git
cd omni-video-mcp

# Install dependencies
uv venv
source .venv/bin/activate
uv pip install -e .

# Install Playwright browsers (for Hyperframes)
playwright install chromium

🛠 Конфигурация

Добавьте сервер в файл настроек MCP вашей IDE (например, ~/.gemini/antigravity/mcp_config.json, ~/.cursor/mcp.json или файл конфигурации Claude Desktop):

{
  "mcpServers": {
    "omni-video-mcp": {
      "command": "uv",
      "args": [
        "run",
        "/path/to/omni-video-mcp/server.py"
      ],
      "env": {
        "ELEVENLABS_API_KEY": "your_api_key_here" 
      }
    }
  }
}

Примечание: ELEVENLABS_API_KEY в настоящее время требуется для высокоточного сопоставления транскрипции на уровне слов во время загрузки.

🎬 Как это работает (Конвейер агента)

Когда агент использует этот MCP-сервер, он следует 4-фазной архитектуре:

  1. Фаза 1: Загрузка (omni_video_ingest) Агент сканирует ваши исходные файлы .mp4 / .mov, извлекая упакованную транскрипцию в формате markdown и начальный граф визуальных сцен.

  2. Фаза 2: Режиссерская версия (omni_video_preview) Агент использует транскрипцию для создания EDL (Edit Decision List) из лучших дублей. Неоднозначные склейки можно визуально проверить, создав PNG-превью с помощью инструмента предварительного просмотра.

  3. Фаза 3: VFX (omni_video_generate_vfx) Агент генерирует HTML/CSS анимированную графику (нижние трети, макеты B-roll), а сервер детерминированно преобразует их в видео .webm с прозрачностью через Hyperframes.

  4. Фаза 4: Улучшение и рендеринг (omni_video_render) Агент передает EDL, временные метки VFX и настройки рендеринга на сервер, который строит сложный граф FFmpeg для объединения отснятого материала, цветокоррекции, восстановления звука и экспорта финального мастер-файла.

🤝 Участие в разработке

Мы приветствуем вклад в проект! Если вы добавляете новые возможности конвейера рендеринга (например, автотрекинг или локальные резервные копии whisper), пожалуйста, создайте PR. Убедитесь, что любые добавленные зависимости Python внесены в pyproject.toml с помощью uv add <package>.

📄 Лицензия

Лицензия MIT

Available Tools

4 tools
omni_video_generate_vfxC

Renders motion graphics (e.g., lower thirds, titles) using Hyperframes. Returns the path to the rendered transparent .mov or .webm file.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavioral traits. It states the tool renders transparent videos using Hyperframes and returns a file path, but fails to mention whether it is synchronous or async, any size limits, error conditions, or side effects. The minimal detail leaves the agent guessing about important behaviors.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise—two sentences covering the core action and output. It front-loads the primary purpose and wastes no words. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and an output schema that likely exists but is not described, the tool description is incomplete. It does not explain the return format beyond a file path, nor does it cover prerequisites for HTML/CSS validity or duration limits. The agent may need to infer or experiment to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has a single parameter 'request' with nested properties (html_content, css_content, duration_seconds), all with schema descriptions. However, the context indicates 0% schema description coverage, likely because the top-level parameter lacks a description. The tool description itself does not mention any parameters or add meaning beyond what the schema provides. For parameter understanding, the agent gets no extra help from the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool renders motion graphics (lower thirds, titles) using Hyperframes and returns the path to a transparent .mov or .webm file. The verb 'renders' and specific resource 'motion graphics' distinguish it from sibling tools like omni_video_ingest, omni_video_preview, and omni_video_render.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, nor when not to use it. It does not mention any prerequisites or context for invoking it. With sibling tools listed but no differentiation, the agent lacks contextual cues for proper selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

omni_video_ingestA

Ingests a directory of video files, generates word-level audio transcripts, and constructs a semantic Visual Scene Graph for B-Roll searching. Returns the path to the generated project metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the transparency burden. It discloses that transcripts and scene graphs are generated and a metadata path is returned, but does not mention whether ingestion modifies source files, requires permissions, or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences long, front-loaded with the main action, and every sentence provides essential information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is non-trivial (multiple stages), but the output schema exists to clarify return values. The description covers the main inputs and outputs, though it lacks edge-case or error details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the top-level parameter. The tool description adds that 'directory_path' is a directory of video files, but the schema already includes a similar description for 'directory_path'. Thus, the description adds minimal value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verbs: 'ingests', 'generates word-level audio transcripts', 'constructs a semantic Visual Scene Graph'. It distinguishes well from sibling tools which handle VFX, preview, and render.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage context is implied (ingesting video files for transcript and scene graph generation), but no explicit guidance on when to use versus alternatives or prerequisites is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

omni_video_previewA

Generates a filmstrip PNG of the specified video segment. Useful for visually verifying cut boundaries or B-roll placement. Returns the absolute path to the generated PNG file.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that the tool generates a PNG and returns its path, implying a read-only operation. However, it does not detail any side effects, permissions, or safety implications.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, each serving a distinct purpose: what it does, when to use it, and what it returns. No wasted words, front-loaded with core action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the relatively simple operation and the presence of a nested parameter schema with descriptions, plus an output schema (not shown but indicated), the description adequately covers the tool's purpose and return value. It could mention the parameter structure briefly, but the schema covers that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the tool description does not explain any parameters. While the schema itself contains descriptions for file_path, start_time, and end_time, the description adds no additional meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Generates a filmstrip PNG of the specified video segment,' which is a specific verb-resource combination. It distinguishes from siblings (generate_vfx, ingest, render) by focusing on preview.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit use cases: 'visually verifying cut boundaries or B-roll placement.' It does not mention when not to use or alternatives, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

omni_video_renderA

Orchestrates the final render pipeline: applies EDL cuts, overlay graphics, LUT color grading, audio restoration, and subtitle burning. Returns the path to the final rendered video.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries behavioral disclosure. It accurately describes the tool's actions (applying cuts, LUT, etc.) and the return value (path to final video). It does not mention side effects like file creation or potential errors, but the output_path parameter implies file writing. This is adequate but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that efficiently lists all main operations and the return value. It is front-loaded with the core purpose ('Orchestrates the final render pipeline') and contains no fluff or redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the tool's core functionality and return value, and the schema outlines required parameters. It lacks details about error handling or prerequisites, but given the presence of an output schema (implied) and sibling context, it is mostly complete for a final render step.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema properties have detailed descriptions (e.g., edl_json, lut_path, output_path), which already clarify parameter meanings. The tool description does not add additional parameter information beyond listing the tool's capabilities. With good schema coverage, a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool orchestrates the final render pipeline and enumerates specific operations (EDL cuts, overlay graphics, LUT grading, audio restoration, subtitle burning). It distinguishes from siblings through the term 'final render,' differentiating it from ingest, VFX generation, and preview tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this is the last step in a video processing pipeline, but does not explicitly state when to use it versus alternatives or mention prerequisites. The mention of 'final render' provides some context, but more direct guidance would improve this dimension.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedomni_video_generate_vfx
    • First observedomni_video_ingest
    • First observedomni_video_preview
    • First observedomni_video_render

TDQS

A4/5.0

Scored across 4 tools

Disambiguation5/5

Each tool serves a distinct function: ingest creates metadata, generate_vfx produces graphics, preview provides visual verification, and render finalizes output. No overlapping purposes.

Naming Consistency5/5

All tools follow the 'omni_video_' prefix with a verb describing the action (ingest, generate_vfx, preview, render), forming a clear and predictable pattern.

Tool Count5/5

Four tools cover the essential phases of video production—ingestion, effect creation, previewing, and rendering—without being excessive or insufficient.

Completeness4/5

The tool set covers the core pipeline, but lacks explicit tools for timeline editing or asset selection, requiring reliance on the render tool for EDL processing.

Maintenance

ActivitySlowing
ResponsivenessSyncing

Related MCP Connectors

Related MCP Servers