Omni-Video Studio MCP
Omni-Video Studio MCP
Omni-Video Studio MCPは、LLM対応のIDE(Cursor、Claude Code、Antigravityなど)をプロフェッショナルなビデオ編集ツールに変える、エンタープライズグレードの自律型Model Context Protocol (MCP) サーバーです。
これは、単純な文字起こしベースの編集を、決定論的でトークン効率が高く、パイプライン駆動型のワークフローへと進化させます。エージェントネイティブなモーショングラフィックス(Hyperframes)、視覚的メタデータプロキシ、高忠実度な最終レンダリング機能を備えています。
🌟 主な機能
メタデータプロキシの取り込み: 高コストなビデオトークンをLLMにストリーミングする代わりに、このサーバーはフッテージを前処理して
takes_packed.md(音声マッピング)とビジュアルシーングラフを抽出します。エージェントはテキストプロキシを使用して編集を行うため、コストを削減し、推論を高速化できます。Hyperframesエンジン: 複雑なNode.jsの依存関係(Remotionなど)は不要です。エージェントが決定論的なHTML/CSSモーショングラフィックスを生成し、Playwrightを使用して透過ビデオとして即座にレンダリングします。
高度なレンダリングパイプライン: 強力なFFmpegフィルターグラフを搭載しており、EDL(編集決定リスト)カット、オーバーレイレンダリング、字幕焼き込み、LUTカラーグレーディング、およびオプションのDeepFilterNet AI音声復元をサポートしています。
IDE非依存: 公式のMCP仕様に準拠しているため、カスタムプラグインなしでCursor、Antigravity、またはClaude Desktopに直接組み込むことができます。
Related MCP server: omni-video-mcp
📦 インストール
前提条件:
python 3.10+ffmpeg(システムパスにインストールされている必要があります)uv(依存関係管理に推奨)
# Clone the repository
git clone https://github.com/your-org/omni-video-mcp.git
cd omni-video-mcp
# Install dependencies
uv venv
source .venv/bin/activate
uv pip install -e .
# Install Playwright browsers (for Hyperframes)
playwright install chromium🛠 設定
IDEのMCP設定ファイル(例: ~/.gemini/antigravity/mcp_config.json、~/.cursor/mcp.json、またはClaude Desktopの設定)にサーバーを追加します:
{
"mcpServers": {
"omni-video-mcp": {
"command": "uv",
"args": [
"run",
"/path/to/omni-video-mcp/server.py"
],
"env": {
"ELEVENLABS_API_KEY": "your_api_key_here"
}
}
}
}注: 現在、取り込み時の高忠実度な単語レベルの文字起こしマッピングには ELEVENLABS_API_KEY が必要です。
🎬 仕組み(エージェントパイプライン)
エージェントがこのMCPサーバーを使用する場合、以下の4フェーズのアーキテクチャに従います:
フェーズ 1: 取り込み (
omni_video_ingest) エージェントが生の.mp4/.movファイルをスキャンし、パックされたマークダウン形式の文字起こしと初期のビジュアルシーングラフを抽出します。フェーズ 2: ディレクターズカット (
omni_video_preview) エージェントは文字起こしを使用して、最適なテイクのEDL(編集決定リスト)を構築します。曖昧なカットは、プレビューツールを介してフィルムストリップPNGを生成することで視覚的に確認できます。フェーズ 3: VFX (
omni_video_generate_vfx) エージェントがHTML/CSSモーショングラフィックス(下部テロップ、Bロールレイアウトなど)を生成し、サーバーがHyperframesを介してそれらを透過.webmビデオとして決定論的にレンダリングします。フェーズ 4: スウィートニングとレンダリング (
omni_video_render) エージェントがEDL、VFXのタイムスタンプ、レンダリング設定をサーバーに渡すと、サーバーは複雑なFFmpegグラフを構築してフッテージの結合、グレーディング、音声復元を行い、最終的なマスターを出力します。
🤝 コントリビューション
コントリビューションを歓迎します!自動トラッキングやローカルWhisperフォールバックなど、新しいレンダリングパイプライン機能を追加する場合は、プルリクエストを作成してください。追加するPythonの依存関係は、uv add <package> を使用して pyproject.toml に追加するようにしてください。
📄 ライセンス
MITライセンス
Available Tools
4 toolsomni_video_generate_vfxC
Renders motion graphics (e.g., lower thirds, titles) using Hyperframes. Returns the path to the rendered transparent .mov or .webm file.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must fully disclose behavioral traits. It states the tool renders transparent videos using Hyperframes and returns a file path, but fails to mention whether it is synchronous or async, any size limits, error conditions, or side effects. The minimal detail leaves the agent guessing about important behaviors.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise—two sentences covering the core action and output. It front-loads the primary purpose and wastes no words. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and an output schema that likely exists but is not described, the tool description is incomplete. It does not explain the return format beyond a file path, nor does it cover prerequisites for HTML/CSS validity or duration limits. The agent may need to infer or experiment to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has a single parameter 'request' with nested properties (html_content, css_content, duration_seconds), all with schema descriptions. However, the context indicates 0% schema description coverage, likely because the top-level parameter lacks a description. The tool description itself does not mention any parameters or add meaning beyond what the schema provides. For parameter understanding, the agent gets no extra help from the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool renders motion graphics (lower thirds, titles) using Hyperframes and returns the path to a transparent .mov or .webm file. The verb 'renders' and specific resource 'motion graphics' distinguish it from sibling tools like omni_video_ingest, omni_video_preview, and omni_video_render.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, nor when not to use it. It does not mention any prerequisites or context for invoking it. With sibling tools listed but no differentiation, the agent lacks contextual cues for proper selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
omni_video_ingestA
Ingests a directory of video files, generates word-level audio transcripts, and constructs a semantic Visual Scene Graph for B-Roll searching. Returns the path to the generated project metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the transparency burden. It discloses that transcripts and scene graphs are generated and a metadata path is returned, but does not mention whether ingestion modifies source files, requires permissions, or any side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, front-loaded with the main action, and every sentence provides essential information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is non-trivial (multiple stages), but the output schema exists to clarify return values. The description covers the main inputs and outputs, though it lacks edge-case or error details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the top-level parameter. The tool description adds that 'directory_path' is a directory of video files, but the schema already includes a similar description for 'directory_path'. Thus, the description adds minimal value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs: 'ingests', 'generates word-level audio transcripts', 'constructs a semantic Visual Scene Graph'. It distinguishes well from sibling tools which handle VFX, preview, and render.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied (ingesting video files for transcript and scene graph generation), but no explicit guidance on when to use versus alternatives or prerequisites is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
omni_video_previewA
Generates a filmstrip PNG of the specified video segment. Useful for visually verifying cut boundaries or B-roll placement. Returns the absolute path to the generated PNG file.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool generates a PNG and returns its path, implying a read-only operation. However, it does not detail any side effects, permissions, or safety implications.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, each serving a distinct purpose: what it does, when to use it, and what it returns. No wasted words, front-loaded with core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the relatively simple operation and the presence of a nested parameter schema with descriptions, plus an output schema (not shown but indicated), the description adequately covers the tool's purpose and return value. It could mention the parameter structure briefly, but the schema covers that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the tool description does not explain any parameters. While the schema itself contains descriptions for file_path, start_time, and end_time, the description adds no additional meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Generates a filmstrip PNG of the specified video segment,' which is a specific verb-resource combination. It distinguishes from siblings (generate_vfx, ingest, render) by focusing on preview.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit use cases: 'visually verifying cut boundaries or B-roll placement.' It does not mention when not to use or alternatives, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
omni_video_renderA
Orchestrates the final render pipeline: applies EDL cuts, overlay graphics, LUT color grading, audio restoration, and subtitle burning. Returns the path to the final rendered video.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries behavioral disclosure. It accurately describes the tool's actions (applying cuts, LUT, etc.) and the return value (path to final video). It does not mention side effects like file creation or potential errors, but the output_path parameter implies file writing. This is adequate but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that efficiently lists all main operations and the return value. It is front-loaded with the core purpose ('Orchestrates the final render pipeline') and contains no fluff or redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the tool's core functionality and return value, and the schema outlines required parameters. It lacks details about error handling or prerequisites, but given the presence of an output schema (implied) and sibling context, it is mostly complete for a final render step.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema properties have detailed descriptions (e.g., edl_json, lut_path, output_path), which already clarify parameter meanings. The tool description does not add additional parameter information beyond listing the tool's capabilities. With good schema coverage, a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool orchestrates the final render pipeline and enumerates specific operations (EDL cuts, overlay graphics, LUT grading, audio restoration, subtitle burning). It distinguishes from siblings through the term 'final render,' differentiating it from ingest, VFX generation, and preview tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is the last step in a video processing pipeline, but does not explicitly state when to use it versus alternatives or mention prerequisites. The mention of 'final render' provides some context, but more direct guidance would improve this dimension.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
omni_video_generate_vfx - First observed
omni_video_ingest - First observed
omni_video_preview - First observed
omni_video_render
TDQS
Scored across 4 tools
Each tool serves a distinct function: ingest creates metadata, generate_vfx produces graphics, preview provides visual verification, and render finalizes output. No overlapping purposes.
All tools follow the 'omni_video_' prefix with a verb describing the action (ingest, generate_vfx, preview, render), forming a clear and predictable pattern.
Four tools cover the essential phases of video production—ingestion, effect creation, previewing, and rendering—without being excessive or insufficient.
The tool set covers the core pipeline, but lacks explicit tools for timeline editing or asset selection, requiring reliance on the render tool for EDL processing.
Maintenance
Related MCP Connectors
AI video editor for agents and humans: timeline, captions, color, audio and generation as MCP tools.
- VidmoatOAuthcom.vidmoat
AI video editor: create projects, edit timelines, add captions and effects, and render videos.
- tonpitOAuthcom.tonpit
Video production studio for AI agents: AI media, motion graphics as code, timeline and export.
Edit video by talking to your AI — search footage, cut timelines, apply effects, add captions.
Related MCP Servers
- FlicenseAqualityFmaintenanceEnables AI assistants to create and edit professional videos through natural language by automating JianYing (CapCut) video production workflows. Supports adding media segments, effects, transitions, animations, and exporting editable project files.20298-
- FlicenseAqualityDmaintenanceAn MCP server that transforms LLM-enabled IDEs into professional video editors by pre-processing footage into text proxies, generating motion graphics via HTML/CSS, and orchestrating complex FFmpeg renders.4-
- AlicenseBqualityDmaintenanceEnables AI video automation pipeline: ComfyUI image-to-video, FFmpeg processing, After Effects template rendering, review, and multi-platform publishing preparation.9MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to edit videos through natural language, providing tools for timeline editing, audio management, rendering, and more.2MIT