video-workflow-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@video-workflow-mcp帮我把产品演示.mp4 配上中文旁白,音色用云希"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
video-workflow-mcp — 视频工作流 MCP Server
演示


一条命令出片:Whisper 转写 → SRT 字幕 → 中文 TTS 配音 → 合成新视频。 全本地、CPU、离线,无 GPU 也能跑。
Related MCP server: Video Transcript MCP Server
为什么做这个
调研确认(7 轮实测):"转写+配音+合成"一条链路的 MCP 协议层 = 0 竞争。
whisper+mcp 572个但全是单步工具(会议纪要/语音对话),无一人做"视频重新配音"
dubbing+mcp 30个,全是付费API包装,本地免费配音 MCP = 0
需求被非MCP工具验证:VideoLingo 18.5k★、pyvideotrans 19.2k★ —— 但没人做成可插拔的 MCP
本项目的价值:把"视频本地化/加配音/加字幕"变成 AI 工作流里的一条命令。
核心能力
工具 | 说明 |
| 视频/音频 → Whisper 转写 → SRT(带时间戳,自动中文检测) |
| 视频 → 转写 → 中文TTS配音 → 合成新视频(一条命令) |
| 多音色对话配音:转写 → 每段自动分配男女音色 → 配音(影视翻译/解说) |
| 复用 tts-mcp-server 的聚合音色目录 |
典型用途
1. 给无声视频加中文旁白:
dub_video_tool(input_path="产品演示.mp4", voice_id="edge:zh-CN-YunxiNeural")
2. 外语视频翻译配音成中文:
转写 → 翻译 → 配音 → 合成
3. 批量给短视频加配音(自媒体矩阵)性能(无 GPU,CPU 实测)
操作 | 耗时 |
16s 视频转写 | ~5s(RTF 0.3,faster-whisper int8) |
转写+配音+合成 | ~12s |
安装
pip install video-workflow-mcp-server # 依赖 faster-whisper/moviepy
# 需要 ffmpeg(Windows: winget install Gyan.FFmpeg)使用(Claude Desktop)
{
"mcpServers": {
"video-workflow": {
"command": "video-workflow-mcp",
"args": []
}
}
}生态协同(四个 MCP 组合)
tts-mcp-server 文本 → 语音 + 逐句字幕
audio-post-mcp 语音 → 响度规范/降噪/静音切割
idphoto-mcp 人像 → 证件照
video-workflow-mcp 视频 → 转写/配音/合成 ← 汇聚层License
MIT。Whisper 用 faster-whisper(MIT),TTS 用 edge-tts(GPL,独立适配层)。
系列(Chinese MCP Suite)
中文内容创作/工具 MCP 全家桶,全本地、CPU、离线:
tts-mcp-server — TTS聚合+字幕对齐
idphoto-mcp — 本地证件照
audio-post-mcp — 音频后期
video-workflow-mcp(本仓库)— 视频工作流
poetry-mcp — 古诗文
script-mcp — 口播文案
asset-mcp — 素材库
idiom-mcp — 成语典故
weather-mcp — 天气/空气
This server cannot be deployed
Maintenance
Related MCP Connectors
- RendobarOAuthcom.rendobar
Transform video, audio and images, and generate media from prompts. FFmpeg, captions, models.
Transcribe, subtitle and dub videos into 100+ languages, and translate text.
Transcribe audio & video: diarization, timed SRT/VTT, podcasts, paste-a-link, whole-feed batch.
Video, audio, and image processing for AI agents: convert, transcribe, upscale - 150+ operations.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables video text extraction using multiple speech recognition providers including local Whisper, JianYing/CapCut, and Bilibili Cut services. Supports video downloading, audio extraction, and automatic speech-to-text transcription with configurable providers.7MIT
- AlicenseAqualityAmaintenanceEnables transcription of videos and audio from 1000+ platforms (YouTube, Bilibili, TikTok, etc.) using subtitle extraction first, then local Whisper transcription, with support for long videos, async tasks, and Chinese ASR optimization.42MIT
- AlicenseNot gradedqualityAmaintenanceEnables high-performance, offline transcription of videos from 1000+ platforms and local files using whisper.cpp, with support for multiple model sizes, languages, and output formats over stdio or HTTP.18Apache 2.0
- AlicenseAqualityCmaintenanceEnables MCP clients to transcribe audio/video files locally, generate SRT subtitles, and burn captions into videos via tool calls, without a cloud API.3MIT