zoom-editor
Processes Zoom meeting recordings and their accompanying audio_transcript_*.vtt transcripts to produce short captioned videos, including clips, trailers, longform best-ofs, and social-sized versions with burned-in captions, waveform, spectrum bars, and headlines. Agents can create projects from Zoom recordings, inspect transcript cues, select moments, render videos, and retrieve completed job files.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@zoom-editorTurn my Zoom recording and transcript into a captioned 30-second trailer"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
zoom-editor
Turn a meeting recording and its transcript into short, captioned videos ready for social media: a single clip, a trailer of about 30 seconds, or a longform best-of. Each comes as a plain cut and, optionally, as 1:1, 9:16 and 16:9 versions with burned-in captions, a waveform, spectrum bars and a headline.
It runs as one service on port 8093, serving an MCP server (streamable HTTP at /mcp) for AI agents and a REST API (/api/v1) for scripts. It encodes on the GPU with NVIDIA's encoder (h264_nvenc) when one is available and falls back to the CPU (libx264) otherwise.
Zoom is the motivating case because Zoom cloud recordings come with an audio_transcript_*.vtt. Any video with a WebVTT or SRT transcript works.
How an agent uses it
create_project: point it at the video (URL, or a path under the media root) and the transcript.get_transcript: read the cues and decide which moments are worth sharing. The service doesn't pick highlights itself; the calling agent (or person) does, so no LLM key is needed here.Render:
render_clip: one moment.render_trailer: one whole sentence from each moment, the one that best matches itshook, played in recording order.render_longform: every moment in full, with the headline changing for each.render_social: no cut; the social sizes of the whole video, for a clip you already have.caption_startsays where it sits on the transcript's timeline, andheadline_timelinecan change the headline over time.
get_jobuntil the job'sstatusisdone, then download the file URLs.
Times can be given in seconds or as HH:MM:SS.mmm, measured on the source video's timeline. Captions never cut a sentence mid-way, and trailer snippets always start and end on sentence boundaries.
Related MCP server: atsurae
Run it
docker pull ghcr.io/isleprince/zoom-editor # or build: docker compose up -d --build
docker compose up -d # GPU passthrough requested; works without one too
curl http://localhost:8093/api/v1/health/api/v1/health reports which encoder was chosen and why. The choice comes from a short test encode at startup, not from checking nvidia-smi. For NVENC inside Docker, the NVIDIA container runtime must pass the video capability, which the compose file requests.
Connect an MCP client to http://<host>:8093/mcp (streamable HTTP).
Variable | Default | |
| 8093 | |
|
| projects and job outputs |
| (empty, off) | local files callers may reference; nothing outside it |
|
|
|
| (empty) | if set, required as |
| request host | base URL for file links in MCP results |
| 1 | concurrent jobs (one GPU encodes one job best) |
| 3 with NVENC, else 1 | social sizes rendered at once within a job (the CPU filter graph is the bottleneck on long videos) |
| 14 | finished jobs are deleted after this; projects stay |
| 8 | cap on downloads from |
REST
GET /api/v1/health
POST /api/v1/projects {video_url|video_path, transcript_url|transcript_path|transcript_text, name}
POST /api/v1/projects/upload multipart: video, transcript?, name?
GET /api/v1/projects[/{id}] DELETE /api/v1/projects/{id}
GET /api/v1/projects/{id}/transcript ?start=&end=
POST /api/v1/jobs {project_id, type: clip|trailer|longform|social, moments:[{start,end,label,hook?}],
sizes:["1x1","9x16","16x9"], subtitles, headline, accent, bg, quality,
caption_start, headline_timeline:[{t0,t1,text}]} (last two: social)
(a clip may instead give start/end/headline at the top level)
GET /api/v1/jobs[/{id}] GET /api/v1/jobs/{id}/files/{name}Develop
pip install -e '.[test]'
pytest -q # generates its own test video; needs an ffmpeg that has drawtextThe layout and transcript rules were ported from the nokemo.com Zoom-to-social pipeline (render_social.py, render_compilations.py) and behave the same.
This server cannot be deployed
Maintenance
Related MCP Connectors
Clip videos into captioned shorts, add captions, and schedule posts from AI agents.
Turn long videos into short, captioned viral clips from your AI assistant. 28 tools, OAuth.
Turn any video or livestream into scored, captioned, ready-to-post vertical clips.
Turns long videos into captioned vertical clips, cut at the moments that stand on their own.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceTurns long-form videos into short-form clips (TikTok/Reels) by reasoning over word-timestamped transcripts, with silence-aware rendering, STT-based validation, and optional reframing/captions.-
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to edit videos through natural language, providing tools for timeline editing, audio management, rendering, and more.2MIT
- AlicenseAqualityDmaintenanceEnables AI agents to edit video assemblies from A-roll and B-roll, add captions, and publish to social media platforms.278 npmMIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to turn natural-language creative direction, transcripts, and source media into fully structured, editable video projects, then verify and render delivery files.2MIT