Skip to main content
Glama

Omni-Video Studio MCP

Omni-Video Studio MCP는 모든 LLM 지원 IDE(Cursor, Claude Code, Antigravity)가 전문 영상 편집기처럼 작동할 수 있도록 지원하는 엔터프라이즈급 자율 Model Context Protocol(MCP) 서버입니다.

이 서버는 단순한 전사 기반 편집을 결정론적이고 토큰 효율적이며 파이프라인 중심적인 워크플로우로 발전시키며, 에이전트 네이티브 모션 그래픽(Hyperframes), 시각적 메타데이터 프록시, 고충실도 최종 렌더링 기능을 제공합니다.

🌟 주요 기능

  1. 메타데이터 프록시 수집: 비용이 많이 드는 영상 토큰을 LLM으로 스트리밍하는 대신, 서버가 푸티지를 사전 처리하여 takes_packed.md(오디오 매핑)와 시각적 장면 그래프(Visual Scene Graph)를 추출합니다. 에이전트는 텍스트 프록시를 사용하여 편집하므로 비용을 절감하고 추론 속도를 높입니다.

  2. Hyperframes 엔진: 복잡한 Node.js 의존성(예: Remotion)은 잊으십시오. 에이전트가 결정론적인 HTML/CSS 모션 그래픽을 생성하면 Playwright를 사용하여 즉시 투명 영상으로 렌더링합니다.

  3. 고급 렌더링 파이프라인: 강력한 FFmpeg 필터 그래프를 기반으로 하며, 최종 출력물은 EDL(편집 결정 리스트) 컷, 오버레이 렌더링, 자막 삽입, LUT 색 보정, 선택적 DeepFilterNet AI 오디오 복원을 지원합니다.

  4. IDE 독립적: 공식 MCP 사양을 준수하므로 별도의 사용자 지정 플러그인 없이 Cursor, Antigravity 또는 Claude Desktop에 바로 적용할 수 있습니다.

Related MCP server: omni-video-mcp

📦 설치

사전 요구 사항:

  • python 3.10+

  • ffmpeg (시스템 경로에 설치되어 있어야 함)

  • uv (의존성 관리를 위해 권장)

# Clone the repository
git clone https://github.com/your-org/omni-video-mcp.git
cd omni-video-mcp

# Install dependencies
uv venv
source .venv/bin/activate
uv pip install -e .

# Install Playwright browsers (for Hyperframes)
playwright install chromium

🛠 구성

IDE의 MCP 설정 파일(예: ~/.gemini/antigravity/mcp_config.json, ~/.cursor/mcp.json 또는 Claude Desktop 설정)에 서버를 추가합니다:

{
  "mcpServers": {
    "omni-video-mcp": {
      "command": "uv",
      "args": [
        "run",
        "/path/to/omni-video-mcp/server.py"
      ],
      "env": {
        "ELEVENLABS_API_KEY": "your_api_key_here" 
      }
    }
  }
}

참고: 현재 수집 과정에서 고충실도 단어 단위 전사 매핑을 위해 ELEVENLABS_API_KEY가 필요합니다.

🎬 작동 방식 (에이전트 파이프라인)

에이전트가 이 MCP 서버를 사용할 때 다음과 같은 4단계 아키텍처를 따릅니다:

  1. 1단계: 수집 (omni_video_ingest) 에이전트가 원본 .mp4 / .mov 파일을 스캔하여 압축된 마크다운 전사본과 초기 시각적 장면 그래프를 추출합니다.

  2. 2단계: 감독판 편집 (omni_video_preview) 에이전트가 전사본을 사용하여 최상의 테이크로 구성된 EDL(편집 결정 리스트)을 작성합니다. 모호한 컷은 미리보기 도구를 통해 필름 스트립 PNG를 생성하여 시각적으로 확인할 수 있습니다.

  3. 3단계: VFX (omni_video_generate_vfx) 에이전트가 HTML/CSS 모션 그래픽(하단 자막, B-roll 레이아웃)을 생성하고, 서버는 Hyperframes를 통해 이를 결정론적으로 투명 .webm 영상으로 렌더링합니다.

  4. 4단계: 보정 및 렌더링 (omni_video_render) 에이전트가 EDL, VFX 타임스탬프, 렌더링 설정을 서버에 전달하면, 서버는 복잡한 FFmpeg 그래프를 구축하여 푸티지를 연결하고, 색 보정 및 오디오 복원을 수행한 뒤 최종 마스터 파일을 내보냅니다.

🤝 기여

기여를 환영합니다! 자동 추적이나 로컬 whisper 폴백과 같은 새로운 렌더링 파이프라인 기능을 추가하려면 PR을 열어주세요. 추가된 Python 의존성은 uv add <package>를 사용하여 pyproject.toml에 추가해야 합니다.

📄 라이선스

MIT 라이선스

Available Tools

4 tools
omni_video_generate_vfxC

Renders motion graphics (e.g., lower thirds, titles) using Hyperframes. Returns the path to the rendered transparent .mov or .webm file.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavioral traits. It states the tool renders transparent videos using Hyperframes and returns a file path, but fails to mention whether it is synchronous or async, any size limits, error conditions, or side effects. The minimal detail leaves the agent guessing about important behaviors.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise—two sentences covering the core action and output. It front-loads the primary purpose and wastes no words. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and an output schema that likely exists but is not described, the tool description is incomplete. It does not explain the return format beyond a file path, nor does it cover prerequisites for HTML/CSS validity or duration limits. The agent may need to infer or experiment to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has a single parameter 'request' with nested properties (html_content, css_content, duration_seconds), all with schema descriptions. However, the context indicates 0% schema description coverage, likely because the top-level parameter lacks a description. The tool description itself does not mention any parameters or add meaning beyond what the schema provides. For parameter understanding, the agent gets no extra help from the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool renders motion graphics (lower thirds, titles) using Hyperframes and returns the path to a transparent .mov or .webm file. The verb 'renders' and specific resource 'motion graphics' distinguish it from sibling tools like omni_video_ingest, omni_video_preview, and omni_video_render.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, nor when not to use it. It does not mention any prerequisites or context for invoking it. With sibling tools listed but no differentiation, the agent lacks contextual cues for proper selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

omni_video_ingestA

Ingests a directory of video files, generates word-level audio transcripts, and constructs a semantic Visual Scene Graph for B-Roll searching. Returns the path to the generated project metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the transparency burden. It discloses that transcripts and scene graphs are generated and a metadata path is returned, but does not mention whether ingestion modifies source files, requires permissions, or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences long, front-loaded with the main action, and every sentence provides essential information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is non-trivial (multiple stages), but the output schema exists to clarify return values. The description covers the main inputs and outputs, though it lacks edge-case or error details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the top-level parameter. The tool description adds that 'directory_path' is a directory of video files, but the schema already includes a similar description for 'directory_path'. Thus, the description adds minimal value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verbs: 'ingests', 'generates word-level audio transcripts', 'constructs a semantic Visual Scene Graph'. It distinguishes well from sibling tools which handle VFX, preview, and render.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage context is implied (ingesting video files for transcript and scene graph generation), but no explicit guidance on when to use versus alternatives or prerequisites is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

omni_video_previewA

Generates a filmstrip PNG of the specified video segment. Useful for visually verifying cut boundaries or B-roll placement. Returns the absolute path to the generated PNG file.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that the tool generates a PNG and returns its path, implying a read-only operation. However, it does not detail any side effects, permissions, or safety implications.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, each serving a distinct purpose: what it does, when to use it, and what it returns. No wasted words, front-loaded with core action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the relatively simple operation and the presence of a nested parameter schema with descriptions, plus an output schema (not shown but indicated), the description adequately covers the tool's purpose and return value. It could mention the parameter structure briefly, but the schema covers that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the tool description does not explain any parameters. While the schema itself contains descriptions for file_path, start_time, and end_time, the description adds no additional meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Generates a filmstrip PNG of the specified video segment,' which is a specific verb-resource combination. It distinguishes from siblings (generate_vfx, ingest, render) by focusing on preview.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit use cases: 'visually verifying cut boundaries or B-roll placement.' It does not mention when not to use or alternatives, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

omni_video_renderA

Orchestrates the final render pipeline: applies EDL cuts, overlay graphics, LUT color grading, audio restoration, and subtitle burning. Returns the path to the final rendered video.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries behavioral disclosure. It accurately describes the tool's actions (applying cuts, LUT, etc.) and the return value (path to final video). It does not mention side effects like file creation or potential errors, but the output_path parameter implies file writing. This is adequate but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that efficiently lists all main operations and the return value. It is front-loaded with the core purpose ('Orchestrates the final render pipeline') and contains no fluff or redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the tool's core functionality and return value, and the schema outlines required parameters. It lacks details about error handling or prerequisites, but given the presence of an output schema (implied) and sibling context, it is mostly complete for a final render step.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema properties have detailed descriptions (e.g., edl_json, lut_path, output_path), which already clarify parameter meanings. The tool description does not add additional parameter information beyond listing the tool's capabilities. With good schema coverage, a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool orchestrates the final render pipeline and enumerates specific operations (EDL cuts, overlay graphics, LUT grading, audio restoration, subtitle burning). It distinguishes from siblings through the term 'final render,' differentiating it from ingest, VFX generation, and preview tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this is the last step in a video processing pipeline, but does not explicitly state when to use it versus alternatives or mention prerequisites. The mention of 'final render' provides some context, but more direct guidance would improve this dimension.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedomni_video_generate_vfx
    • First observedomni_video_ingest
    • First observedomni_video_preview
    • First observedomni_video_render

TDQS

A4/5.0

Scored across 4 tools

Disambiguation5/5

Each tool serves a distinct function: ingest creates metadata, generate_vfx produces graphics, preview provides visual verification, and render finalizes output. No overlapping purposes.

Naming Consistency5/5

All tools follow the 'omni_video_' prefix with a verb describing the action (ingest, generate_vfx, preview, render), forming a clear and predictable pattern.

Tool Count5/5

Four tools cover the essential phases of video production—ingestion, effect creation, previewing, and rendering—without being excessive or insufficient.

Completeness4/5

The tool set covers the core pipeline, but lacks explicit tools for timeline editing or asset selection, requiring reliance on the render tool for EDL processing.

Maintenance

ActivitySlowing
ResponsivenessSyncing

Related MCP Connectors

Related MCP Servers