seedance-2-mcp
seedance-2-mcp
オープンソースでローカル実行可能な MCP (Model Context Protocol) サーバーです。火山エンジン (Volcengine ARK) の Seedance 2.0 動画生成機能を3つのstdioツールとして公開し、Codex、Claude Desktop、CursorなどのあらゆるMCPクライアントから利用可能にします。
完全ローカルのstdio、クラウドへのデプロイは一切不要。
Node.js + TypeScriptで記述され、公式の
@modelcontextprotocol/sdkを使用。ユーザーは自身のマシンで
ARK_API_KEYを設定するだけで、Seedance 2.0の テキストから動画 / 画像から動画 / マルチモーダル参照生成 を呼び出せます。
提供されるMCPツール
ツール | 役割 |
| 完全な使用ガイド(標準フロー、モデル選択、パラメータ表、注意事項)を返します。 |
| Seedance 2.0動画生成タスクを送信し、即座に |
|
|
詳細なパラメータは Toolsの詳細 を参照してください。
Related MCP server: Seedance MCP
ローカルエージェント向けのクイック接続プロンプト
Codex、Claude Desktop、CursorなどのMCP対応ローカルエージェントを使用している場合、このリポジトリのリンクを直接送信し、以下のセクションを読ませることができます。
你是一个本地开发 Agent。请帮我把这个仓库提供的 seedance-2-mcp 接入到当前 MCP 客户端中。
目标:
1. 读取仓库 README,理解这是一个 stdio MCP server,用于调用火山方舟 Seedance 2.0 视频生成 API。
2. 优先使用 npx 方式接入:command = "npx",args = ["-y", "seedance-2-mcp"]。
3. 如果 npm 包暂不可用,或我明确想从源码运行,请 clone 本仓库,执行 npm install && npm run build,并将 MCP command 配为 "node",args 配为 ["<仓库绝对路径>/dist/index.js"]。
4. 只向我索要或确认 ARK_API_KEY,不要把真实 API Key 写进仓库、README、示例文件或 git。
5. 根据我当前使用的客户端自动修改对应 MCP 配置:
- Codex:修改 ~/.codex/config.toml
- Claude Desktop:修改 claude_desktop_config.json
- Cursor 或其他客户端:使用它们支持的 stdio MCP 配置格式
6. 配置完成后,提醒我重启或刷新 MCP 客户端,然后先调用 seedance_usage_guide,再按 create -> wait -> check 的流程生成视频。
7. 如果本机没有 Node.js >= 18 或 npx 不可用,请先指出缺失项,并给出最小安装建议。
重要约束:
- stdout 是 MCP JSON-RPC 通道,不要让 server 在 stdout 打调试日志。
- ARK_API_KEY 只能放在 MCP 客户端 env 配置或本机环境变量里。
- 生成的 video_url 通常会过期,任务成功后应提示我尽快下载。また、エージェントに対して直接次のように指示することもできます:
このリポジトリを読み、READMEの「ローカルエージェント向けのクイック接続プロンプト」に従ってSeedance MCPを設定してください。
ARK_API_KEYは私が提供します。
インストール
1. npx を使用(推奨 - 手動インストール不要)
MCPクライアントの設定で直接使用します:
npx -y seedance-2-mcp起動のたびに最新バージョンがダウンロード(またはキャッシュを利用)されます。
2. グローバルインストール
npm install -g seedance-2-mcpその後、クライアントの設定で seedance-2-mcp を使用します。
3. ソースコードから実行(開発者向け)
git clone https://github.com/seedance/seedance-2-mcp.git
cd seedance-2-mcp
npm install
npm run build
node dist/index.jsNode.js >= 18 が必要です(ネイティブ fetch に依存)。
環境変数
変数 | 必須 | 説明 |
| はい | 火山エンジンARKのAPIキー。https://console.volcengine.com/ark から取得してください。 |
| いいえ | デフォルトは |
.env.example をローカルの .env にコピーして開発の参考にしてください。実際に有効な設定場所はMCPクライアント設定内の env です。クライアントがサブプロセスとしてMCPを起動する際、環境変数を注入するためです。
ツール呼び出し時に ARK_API_KEY が設定されていない場合、APIを実際に呼び出す2つのツールは明確なエラーを返します:
Missing ARK_API_KEY environment variable. Please set ARK_API_KEY to your Volcengine ARK API key.
クライアント設定例
Codex(~/.codex/config.toml)
[mcp_servers.seedance-2-mcp]
command = "npx"
args = ["-y", "seedance-2-mcp"]
env = { "ARK_API_KEY" = "your_key_here" }Claude Desktop(claude_desktop_config.json)
macOSパス: ~/Library/Application Support/Claude/claude_desktop_config.json
Windowsパス: %APPDATA%\Claude\claude_desktop_config.json
{
"mcpServers": {
"seedance-2-mcp": {
"command": "npx",
"args": ["-y", "seedance-2-mcp"],
"env": {
"ARK_API_KEY": "your_key_here"
}
}
}
}Cursor / その他のMCPクライアント
stdio MCPをサポートするあらゆるクライアントの共通設定:
{
"command": "npx",
"args": ["-y", "seedance-2-mcp"],
"env": { "ARK_API_KEY": "your_key_here" }
}Toolsの詳細
seedance_usage_guide
引数なし。Markdown形式の完全な使用ガイドを返します。seedance_create_task を初めて呼び出す前に、一度呼び出すことを推奨します。
seedance_create_task
動画生成タスクを送信し、即座に task_id を返します。
入力パラメータ:
フィールド | 型 | デフォルト | 説明 |
| string | —(必須) | 自然言語による記述。参照素材がある場合は、プロンプト内で |
| enum |
|
|
| integer |
| 動画の長さ(秒)。 |
| enum |
|
|
| enum |
|
|
| boolean |
| 同期オーディオ(セリフ / 効果音 / BGM)を同時に生成するかどうか。 |
| boolean |
| プラットフォームの透かしを追加するかどうか。アカウントによってはオフにできない場合があります。 |
| boolean |
| プロンプトに対してWeb検索による強化を有効にするか。純粋なテキスト入力でのみ使用可能。画像/動画/音声と同時に使用不可。 |
| boolean |
| 複数のセグメントを結合するために、最終フレームの画像URLを返すかどうか。 |
| array | — | 最大9項目。各項目は |
| array | — | 最大3項目。各項目は |
| array | — | 最大3項目。各項目は |
検証ルール:
durationは[4, 15]の整数である必要があります。image_urls≤ 9,video_urls≤ 3,audio_urls≤ 3。純粋なテキスト +
audio_urlsは拒否されます(Seedanceはサポートしていません)。web_search=trueと参照素材を同時に指定すると拒否されます(Web検索強化は純粋なテキストのみサポート)。
戻り値:
{
"task_id": "cgt-2026xxxx-xxxxxx",
"model": "doubao-seedance-2-0-260128",
"duration": 5,
"ratio": "16:9",
"resolution": "720p",
"raw": { /* 火山原始响应 */ }
}seedance_check_task
入力 { task_id: string }。可能な状態:
running/queued/pending— 処理中。30〜90秒待ってから再度呼び出すことを推奨します。succeeded—video_urlを返します。return_last_frame=trueの場合はlast_frame_urlも返されます。failed—fail_reasonを返します(存在する場合)。cancelled/expired— タスクがキャンセルされたか、期限切れです。その他 —
statusと元のペイロードをそのまま返します。
標準的な呼び出しフロー
client → seedance_usage_guide ← 阅读规则
client → seedance_create_task { prompt, ... } ← 提交任务
↓
task_id: cgt-...
↓
client → seedance_check_task { task_id } ← 30-90s 后轮询
↓
status: running (继续等待)
↓
status: succeeded ← 返回 video_url
↓
立刻下载 video_url(约 24h 内会过期)15秒の標準モデルタスクは通常 2〜5分 で完了します。高速版はより短くなります。
セキュリティに関する注意事項
ARK_API_KEYをgitリポジトリに書き込まないでください。MCPクライアント設定(claude_desktop_config.json、~/.codex/config.tomlなど)のenvフィールド、またはシェルの環境変数に設定してください。本ツールはログや戻り値の中に
ARK_API_KEYを出力しません。火山エンジンが生成する
video_urlおよびlast_frame_urlは 署名付きの一時URL であり、火山エンジンの公式仕様に基づきデフォルトで24時間有効です。タスク完了後は速やかにダウンロードし、リンク切れを防いでください。指定する
image_urls/video_urls/audio_urlsはすべて パブリックアクセス可能な HTTPS(またはHTTP)アドレスである必要があります。ローカルパス、イントラネットアドレス、ログインが必要なリソースは火山エンジンのサーバーから取得できません。火山エンジンおよびSeedanceモデルの利用規約を遵守し、違法、未成年者に不適切なコンテンツ、または権利を侵害するコンテンツを生成しないでください。
開発
npm install
npm run typecheck
npm run dev # 用 tsx 直接跑 src/index.ts
npm run build # 输出到 dist/
npm start # node dist/index.jsstdio MCPのデバッグには以下を推奨します:
npx -y @modelcontextprotocol/inspector npx -y seedance-2-mcpまたはローカルソースコード:
npx -y @modelcontextprotocol/inspector node dist/index.jsプロジェクト構造
.
├── src/
│ ├── index.ts # stdio MCP 入口(带 shebang)
│ ├── server.ts # 注册 McpServer 和三个 tools
│ ├── seedance.ts # 火山方舟 Seedance 2.0 REST API 客户端
│ ├── schema.ts # zod 输入 schema
│ └── usageGuide.ts # seedance_usage_guide 返回的文本
├── package.json
├── tsconfig.json
├── .env.example
├── .gitignore
├── LICENSE
└── README.mdライセンス
MIT © seedance-2-mcp contributors
本プロジェクトは、ByteDance、火山エンジン、Volcengine ARKとは 公式な関連はありません。 「Seedance」、「Doubao」、「火山エンジン」などの名称の著作権はそれぞれの所有者に帰属します。
Available Tools
3 toolsseedance_check_taskCheck Seedance 2.0 task statusARead-onlyIdempotent
Query the status of a Seedance 2.0 task by task_id. Returns running / succeeded / failed / other. On success returns video_url (and last_frame_url if return_last_frame was true). Generated URLs expire within ~24h - download promptly.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | Seedance task id returned by seedance_create_task (e.g. cgt-...). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate readOnly and idempotent. Description adds URL expiration (~24h) and return format, which isn't in annotations. No contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with purpose, no redundant or extraneous content. Efficient and clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Describes statuses and URL expiry without output schema. Adequate for a status check tool; could mention error handling but not required.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 100% coverage with description for task_id. Description adds context that it's the ID from seedance_create_task, improving usability.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it queries status of a Seedance 2.0 task by task_id, lists possible statuses and success returns. Distinguishes from sibling tools (create and guide).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage after task creation but does not explicitly state when to use vs alternatives or mention polling patterns. No when-not or alternatives provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
seedance_create_taskCreate Seedance 2.0 video generation taskA
Submit a Seedance 2.0 video generation task to the Volcengine ARK API and return the task_id immediately. Does NOT wait for the video to render - poll seedance_check_task afterwards. Reference media URLs must be publicly reachable.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Natural-language description of the desired video. Reference images / videos / audios with [Image1], [Video1], [Audio1] in 1-based order if you provided any. | |
| model | No | Seedance 2.0 model id. doubao-seedance-2-0-260128 is the standard, highest-quality model. doubao-seedance-2-0-fast-260128 trades quality for latency. | doubao-seedance-2-0-260128 |
| duration | No | Video length in seconds. Must be an integer in [4, 15]. | |
| ratio | No | Aspect ratio. Use 9:16 for vertical short-video, 16:9 for landscape. 'adaptive' lets the model pick the best fit when reference media is provided. | 16:9 |
| resolution | No | Output resolution. 720p is recommended; 480p is faster/cheaper. | 720p |
| generate_audio | No | Whether to generate synchronized audio (dialogue, SFX, music). Set false for silent video. | |
| watermark | No | Whether to add the platform watermark. Some accounts cannot disable this. | |
| web_search | No | Enable prompt enhancement via web search. Text-only input is required when this is true. | |
| return_last_frame | No | Return the last frame as an image URL alongside the video URL. Useful for chaining segments. | |
| image_urls | No | Up to 9 reference images. Each item is { url, role? }. role defaults to 'reference_image'. Use 'first_frame' (and optionally 'last_frame') for image-to-video animation. | |
| video_urls | No | Up to 3 reference videos. Each item is { url, role? }. role defaults to 'reference_video'. | |
| audio_urls | No | Up to 3 reference audios. Each item is { url, role? }. role defaults to 'reference_audio'. Audio MUST be paired with at least one image or video reference. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate readOnlyHint=false, and the description confirms a mutation (creates a task). It adds context beyond annotations: the async nature ('Does NOT wait for the video to render') and the requirement for public URLs. However, it does not mention rate limits, error handling, or response structure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences containing no fluff: first states core function, second explains async behavior with action, third adds constraint. Front-loaded with the most important information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 12 parameters and no output schema, the description covers the essential workflow and a key constraint (public URLs). It lacks details on error cases or the exact format of the returned task_id, but the async instruction is clear. Overall adequate but could be slightly more detailed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds value beyond schema by emphasizing the async behavior and public URL requirement, which are not evident from individual parameter descriptions. It also clarifies the workflow (submit then poll).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Submit' and the resource 'Seedance 2.0 video generation task'. It specifies that it returns a task_id immediately and does not wait for rendering, distinguishing it from sibling tools like seedance_check_task that are used for polling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly instructs to poll seedance_check_task afterwards and notes that reference media URLs must be publicly reachable. It provides clear guidance on when to use this tool (to initiate a task) and what to do next (poll for results).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
seedance_usage_guideSeedance 2.0 usage guideARead-onlyIdempotent
Returns the canonical usage guide for the Seedance 2.0 MCP: standard create -> wait -> check workflow, model choices, parameter reference, and important caveats. Call this before your first seedance_create_task.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations (readOnlyHint, idempotentHint) already declare safe, idempotent behavior. The description adds that it returns a guide, but no additional behavioral traits beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no wasted words. Front-loaded with key purpose and usage instruction. Highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema guide tool, the description fully explains its purpose, content, and when to use it. No gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters exist, so baseline is 4. The description adds meaning by explaining the tool's purpose and content, compensating for lack of parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it returns the canonical usage guide for Seedance 2.0 MCP, detailing workflow, model choices, parameters, and caveats. It distinguishes itself from sibling tools (check, create) by being the preliminary guide.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit advice to call before first seedance_create_task provides clear context for use. No exclusions or alternatives are needed given its unique role as a guide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
seedance_check_task - First observed
seedance_create_task - First observed
seedance_usage_guide
TDQS
Scored across 3 tools
Each tool has a unique and clear purpose: creating a task, checking its status, and providing usage instructions. There is no overlap in functionality.
All tools share the 'seedance_' prefix and follow a noun-based pattern (check_task, create_task, usage_guide), but 'usage_guide' is not a verb_noun like the others, causing minor inconsistency.
With only 3 tools, the set is tightly scoped to the core workflow of generating and monitoring a video task, plus a guide. This is appropriate for the narrow domain.
The tools cover the essential create-and-check workflow, and the guide complements them. However, missing operations like task cancellation or listing are minor gaps but not critical for the basic use case.
Maintenance
Related MCP Connectors
MCP server for ByteDance Seedance AI video generation
AI image, video, voice and music generation over MCP, routed to Veo 3.1, Seedance 2.5 and more.
Build, run, schedule, and publish AI video pipelines to YouTube and TikTok from any MCP client.
MCP server for Google Veo AI video generation
Related MCP Servers
- AlicenseCqualityDmaintenanceIntegrates Volcengine's Ark API to provide image and video generation and editing capabilities. It supports both local stdio and remote HTTP transport modes for flexible deployment and use.824 npm3MIT
- FlicenseAqualityDmaintenanceEnables video generation using the Seedance 2.0 model through MCP, supporting both OpenAI and Volcengine API formats with tools for creating, monitoring, and downloading videos.6-
- FlicenseAqualityDmaintenanceProvides a audio/video creation toolbox via MCP protocol, enabling natural language-based video editing tasks such as image-to-video, video merging, subtitle extraction, and more.93-
- AlicenseNot gradedqualityBmaintenanceMCP server for generating Seedance 2.5 videos via MuAPI, enabling text-to-video, image-to-video, first/last frame, and omni reference workflows with 720p/480p resolution selection.7MIT