engine
Adds optional ElevenLabs text-to-speech voiceovers for video renders via an API key.
Produces platform-ready Facebook video packages with video, cover, captions, post copy, and QA report, formatted for Facebook's UI safe zones.
Produces platform-ready Instagram Reels video packages with video, cover, captions, post copy, and QA report, formatted for Instagram's UI safe zones.
Produces platform-ready TikTok video packages with video, cover, captions, post copy, and QA report, formatted for TikTok's UI safe zones.
Produces platform-ready YouTube Shorts video packages with video, cover, captions, post copy, and QA report, formatted for YouTube Shorts' UI safe zones.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@engineCreate a 30-second TikTok reel from this README with captions."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
https://github.com/user-attachments/assets/d860c1b2-8653-455c-9bfa-640ad8a0b849
Turn a README, a paper, a web page or a folder of clips into a finished, captioned short video, one package per platform, without leaving Claude Code.
Features · See what it makes · Quick start · Requirements · How it works · Commands · Limits · What's new
Why
Turning knowledge into short videos usually means a timeline editor, a caption tool, a thumbnail tool and a checklist for every platform's safe zones and limits. video-studio compiles instead:
Grounded: every claim on screen cites a line in your sources.
verifyshows what is covered.Platform-ready: one
dist/<platform>/package each for TikTok, Instagram Reels, YouTube Shorts, LinkedIn and Facebook, with the video, cover, captions, post copy and a QA report.lintchecks captions and text against each app's UI.Made to be watched: plans follow a story arc with timed "reads", each list item, step or number appears as the voice says it, and Claude reviews contact sheets of its own render before handing it over.
Works with real footage: ingest recordings or YouTube/Vimeo/Loom links, transcribe them locally in ~99 languages, find the best short clips, keep a moving speaker in frame for vertical video, cut away to graphics while they talk, and trim pauses and filler words.
You stay in control: Claude proposes and waits for your approval; paid voices, model downloads and screen recordings ask you through Claude Code's approval dialog and follow your
policy.yamlspend limits; renders can be cancelled.Reproducible:
video.lockpins every tool, font, renderer and asset hash. Re-renders are cached scene by scene, anddiffandtestcatch regressions.Local-first: ffmpeg, system text-to-speech and local whisper.cpp. Claude writes the plan, so the plugin needs no LLM API key, and none of the features below need a paid service.
Related MCP server: MotionKit MCP Server
Features
Highlights
🎬 Motion written as code. Claude can write a scene as an HTML page drawn by a pure
seek(t)function: springs, morphs, match cuts and kinetic type at the level of hand-made motion design. Every page runs under a strict security policy, is checked for unsafe code before it renders, and is proven deterministic (the same time always draws the same frame).📏 Pacing you can measure. QA counts big visual changes per second, the longest still stretch and frozen time, and fails a slideshow-paced reel. Vague asks like "make it pop" become acceptance numbers the render must meet, and
comparescores your render against a reference video you like.🎯 Style from a reference reel.
analyzemeasures how a reel you like moves (entrance times, easing, stagger, holds) without keeping anything from it, andwrite_styleturns that into a style pack in your project.🚀 Launch videos in one command.
/video-studio:launchturns your repo or site into an 18–22 s reel of the product in use, in its own colours, fonts and logo (drafted from the source for you to accept), with a synthesized score, sound effects and a poster on frame 0 for chat previews. One approval, after the preview.🎵 Music that drives the cut. Beat analysis finds beats, downbeats and the drop; cuts snap to them and sound effects land on their peak. No track?
synth:scores are composed locally, CC0, with an exact beat grid.📱 One source, every platform. Per-platform packages (video, cover, captions, post copy, QA) for Instagram Reels, YouTube Shorts, TikTok, LinkedIn and Facebook, with text and captions kept clear of each app's UI.
✅ Grounded and reproducible. Every on-screen claim cites your sources (
verify),video.lockpins every tool and asset, and scenes re-render only when something they use changes.🔒 Local-first, no keys needed. ffmpeg, system voices and local whisper. Claude writes the plan, so there is no LLM API key; paid providers are optional placeholders until you add keys.
Everything it does
Area | Features |
Sources | Markdown, text, PDF, DOCX, PPTX, web pages (JavaScript-built pages too, with your approval), local repos, video/audio files, clip folders, video URLs (YouTube, Vimeo, Loom via your |
Planning | Story-arc plans with hooks and a hook-strength check, 24 templates that ask for the inputs they need first, a beat-level plan at the approval step, series bibles for recurring characters and looks, 7 tone presets (polished, playful, deadpan, cinematic, energetic, app-store, parody), product videos planned around the product in use, a brand kit drafted from your repo or site ( |
Visuals | 16 scene kinds including Claude-written |
Review before render |
|
Audio | System TTS or ElevenLabs, 4 CC0 beds plus locally synthesized scores, beat and downbeat snapping, a synthesized CC0 sound-effect library, sound effects on their peak, motion that reacts to the music, ducking, −14 LUFS |
Footage | Local transcription (~99 languages, speaker turns) with a glossary for names, best-clip |
Captions and languages | Phrase captions clear of platform UI, keyword emphasis, sound-event captions, |
Checks | 30+ lint rules (UI zones, contrast, reading speed, caption sync, insert timing, cuts on the beat, story arc, title length, brand rules, stock phrases, busy crossfades, reveals too fast to read, sound-effect licence and placement, banned effects, acceptance numbers, loop seams, unsafe motion pages), technical QA (loudness, black, frozen, motion density, share of frames moving, loop seam, flashing, A/V sync), |
Export | Per-platform packages, C2PA signing, and an editable timeline (import-tested) for DaVinci Resolve or Final Cut (FCPXML and OTIO) |
Trust and control | Claim |
Generative (prep) | Shot cards compiled into ready-to-paste prompt packs for Seedance, Veo, Kling, Wan, Runway and Hailuo, offline with no spend; provider and publishing keys are optional placeholders until Phase 7/9 |
See what it makes
Everything below was made by the plugin itself: no hand editing.
This README as a reel. docs/media/hero.mp4 is a narrated 1080×1920 final render made from this README. It uses the animated-explainer structure, the technical style and a macOS Premium voice, and every line cites a README line. The project that makes it is examples/readme-hero:
The strips below are frames from preview renders (built-in ffmpeg renderer) made during development.
Built-in scene kinds. The strip shows kinetic text, a stat, a timeline, before/after, a quote, a map and a lower third, from a text-over-music reel with no voiceover:
4 style packs, on the same content. From left: the default, minimal, editorial, technical and energetic:
Languages. The Whisper paper (arXiv 2212.04356) as Hindi (top) and Japanese (bottom) versions made with localize: shaped Devanagari, Japanese line breaking and script-aware captions and covers:
Your own footage. A folder of clips cut to the beat of a bundled music bed, with text over the footage (aesthetic-broll and silent-vlog):
Quick start
Requirements: Claude Code, Node.js 22.13+ and FFmpeg (brew install ffmpeg on macOS). Nothing else is required: the engine is a single bundled file. See Requirements for the full list, and check your machine with /video-studio:doctor.
Inside Claude Code:
/plugin marketplace add harshil-1411/claude_plugin_video_studio
/plugin install video-studio@video-studio-marketplaceThen:
/video-studio:create README.md as a 30-second 9:16 reel for instagram and youtube-shortsClaude reads the source, proposes a hook, a scene plan and a storyboard, and waits for your approval. It then renders a preview, looks at it, then the final, and writes the packages:
dist/
├── reel.mp4 clean-master.mp4 captions.srt captions.vtt cover.jpg
├── video.lock render-manifest.json provenance.json video-spec.json
├── instagram/ video.mp4 cover.jpg captions.* post.json qa.json
└── youtube-shorts/ …Voice: macOS picks your best installed voice (add a Premium voice in System Settings → Accessibility → Spoken Content). An ElevenLabs key (
/plugin→ video-studio → Configure, stored in the OS credential store) is used only when yourpolicy.yamlallows it or you ask for it.HyperFrames renderer (needed for
motionscenes): install Google Chrome, then run/video-studio:doctor: it prints the one-time install command (the pinned@hyperframes/producer, installed into the plugin's data folder) and confirms Chrome starts. Nothing is installed until you run it. Without HyperFrames, every scene kind still renders with ffmpeg, but amotionscene appears as a labelled text stand-in and the render says so.Video URLs:
brew install yt-dlpto ingest YouTube, Vimeo or Loom videos (subtitles become the transcript).Transcription: whisper.cpp (
brew install whisper-cpp); models are downloaded only after you approve.
Nothing needs a key. /plugin → video-studio → Configure shows these fields; secrets are kept in the OS credential store and only reach the engine, never a shell:
Kind | Keys |
In use today | ElevenLabs (voiceover), used only when your |
Placeholders for Phase 7 (AI video) | Runway, HeyGen, fal.ai, Kling, Google Gemini (Veo), BytePlus ModelArk (Seedance), Alibaba DashScope (Wan), MiniMax (Hailuo) |
Placeholders for Phase 9 (publishing) | YouTube client ID and secret, Meta (Instagram) token, LinkedIn token, TikTok client key and secret |
Placeholders can be filled in any time; they have no effect until their integration ships. /video-studio:doctor shows which keys are set, never their values.
Give Claude the reference video with your source:
/video-studio:create my-notes.md as a 20-second 9:16 reel that feels like reference.mp4. The plan skill asks for a reference, a photo and your brand first, and turns "make it feel like this" into acceptance numbers (big changes per second, frozen %, holds)./video-studio:analyze reference.mp4 write_style ref-lookmeasures the reference's motion timing (entrance length, easing, stagger, holds) and saves it as a style pack in your project. Only timing is kept, never its words, frames or audio.After the preview render,
/video-studio:stillsshows each scene on the beats before the final render, and/video-studio:comparescores your render against the reference (frozen %, changes per second, cut rate, loudness).QA fails a render that misses the acceptance numbers, so a slideshow-paced reel cannot pass silently.
git clone https://github.com/harshil-1411/claude_plugin_video_studio.git video-studio
cd video-studio && pnpm install
claude --plugin-dir .Requirements
/video-studio:doctor checks all of this on your machine and says exactly what is missing and how to fix it.
Operating system
OS | Status |
macOS (Apple silicon or Intel) | Fully supported and tested. Everything works, including macOS voices and automatic subject tracking (Apple Vision). |
Linux | Supported. Voice uses |
Windows | Not tested. There is no built-in system voice (use no voice or ElevenLabs), and Chrome is not found automatically (set |
Required software
Tool | Why | Install (macOS) |
Hosts the plugin; Claude writes the plans and pages | see the Claude Code docs | |
Node.js 22.13+ | Runs the engine (it uses Node's built-in SQLite) |
|
FFmpeg and ffprobe, with libx264 and libass | Every render, caption burn-in and QA check. libass (with fribidi) is needed for Arabic, Hebrew and Devanagari captions |
|
Optional software (each unlocks one feature; nothing is installed for you)
Tool | Unlocks |
Google Chrome + the HyperFrames producer |
|
whisper.cpp ( | Transcribing your footage, |
yt-dlp ( | Ingesting YouTube, Vimeo and Loom links |
A macOS Premium or Enhanced voice (System Settings → Accessibility → Spoken Content) | Natural-sounding narration with the free system voice |
An ElevenLabs API key | Premium voiceover (optional and policy-gated; see API keys below) |
FFmpeg with | Tone-mapping HDR phone footage (without it HDR passes through with a warning) |
Hardware
Minimum | Recommended | |
CPU | Any 64-bit CPU | 4+ cores (scenes render in parallel, up to 2 at once) |
Memory | 8 GB | 16 GB. Each parallel scene render wants about 1.5 GB free, and HyperFrames runs Chrome |
Disk | About 25 MB for the plugin | 1–2 GB free for your projects, caches and optional whisper models. A 20 s 1080p reel project is about 50 MB, and the shared cache grows by a few hundred MB over many projects |
GPU | Not needed | Not needed: everything renders on the CPU |
Internet | Only to install the plugin | Only for URL ingest, optional downloads, and paid providers you enable |
How it works
flowchart LR
A["Sources<br/>md · pdf · docx · pptx · url · repo<br/>video · audio · clip folders"] --> B["ingest<br/>ContentIR + evidence refs"]
B --> C["plan<br/>brief · grounded VideoSpec · storyboard"]
C -->|your approval| D["render<br/>voice · scenes · captions · music"]
D --> E["lint + QA<br/>platform contracts · WCAG · loudness"]
E --> F["dist/<platform>/<br/>video · cover · captions · post · qa"]Claude is the creative engine. Skills guide Claude to write the brief and the scene spec.
The engine does the rest. A bundled MCP server validates, renders, runs QA and packages. It is deterministic, cached and has no LLM calls.
Three contracts connect the stages:
ContentIR: the sources, with provenance.VideoSpec: a provider-neutral scene graph.RenderManifest: exactly what happened.
Their JSON Schemas are in
schemas/.Facts are data, not code:
platform-specs/*.yamlrecord each app's limits and UI masks, andprovider-specs/*.yamleach AI video model family's limits and prompt syntax, with source URLs and the date they were checked.
What you can make
Inputs | Markdown, text, PDF, DOCX, PPTX, web pages (a page built by JavaScript can be opened once in an isolated headless Chrome with |
Templates (24) | explain · educational · listicle · faceless-listicle · product-launch · devtool-launch · product-demo · product-ui · case-study · before-after · carousel-story · animated-explainer · text-over-music · talking-head · aesthetic-broll · silent-vlog · oddly-satisfying · ambient-slice-of-life · ui-morph-loop · kinetic-type · ambient-loop · slides-narrated · topic-explainer-9 · product-hero; templates can ask for the inputs they need first (reference video, photo, real UI states) |
Scene kinds (16) | typography · code · chart · stat · diagram · timeline · comparison · split_screen · quote · kinetic_text · lower_third · map · screenshot · cta · end_card, plus |
Voice | macOS |
Audio | 4 bundled CC0 music beds (ducked under speech), locally synthesized scores ( |
Footage | Crop, contain or blurred-pad fits, trim and speed, text overlays, automatic removal of baked-in letterbox bars, and |
Captions | 3–7 word phrases on plates, placed clear of each platform's UI, held long enough to read, broken at speaker changes, with keyword emphasis and sound-event cues like |
Looks | Style packs (minimal, editorial, technical, energetic, or your own in |
Languages |
|
Checks | 30+ lint rules (platform UI zones, contrast, reading speed, caption timing, cues, data inserts on the words that say them, story arc, cutaway rhythm, title length (a heuristic), brand rules, banned effects, acceptance numbers, loop seams, footage quality, stock phrases, busy crossfades, reveal speed, sound effects), technical QA (loudness, black or frozen frames, motion density, share of frames moving, flashing (approximates WCAG 2.3.1; red flashes not measured), A/V sync), |
Trust |
|
Control |
|
Examples
Example | What it shows |
The narrated hero reel above, grounded in a README snapshot | |
A 30 s explainer from Markdown notes (also the golden-frame test) | |
Every Phase 5 scene kind, the | |
A 6 s seamless UI-morph loop written as a | |
A 20 s launch reel for an invented product, made the | |
A tiny web app to try |
Commands
Command | What it does |
| The whole flow, from a source or an idea to packages, with an approval step |
| A short launch reel of something you built, from its repo or URL: the product in use, its own brand, one approval after the preview |
| Brief, grounded spec and storyboard, built on a story arc (hook, open loop, escalation, payoff, CTA); explains every validation issue |
| Local render (preview, then final; cancel anytime), technical QA (including flashing and A/V sync), per-platform packages ( |
| Platform contract checks with a fix loop (UI zones, caption readability and sync, cuts on the beat, story arc); claim coverage against the sources |
| Frames of each scene at chosen times, beats or downbeats, before the full render |
| Prompts for Seedance, Veo, Kling, Wan, Runway and Hailuo compiled from shot cards (offline: nothing generated or spent) |
| Contact sheets, frame strips, transition tiles and crops of a render (lint findings bordered), so Claude looks at the video before handing it over; a before/after page that plays two versions in sync (side by side, stacked or wipe) |
| Golden-frame regression tests; spec, lock and frame diffs between renders |
| Hook × cover A/B sets with an experiment manifest; new aspect, length or platform |
| Language versions from a translation sheet, re-timed for the language |
| Documents, web pages (JavaScript-built ones with your approval), repos, media files and video URLs; local transcription (language detection, speaker turns); standalone clips from a long talk, with shot sheets, subject tracking and cutaways; a reference video's format, motion timing and speech pacing ( |
| Cleans up talking-head footage: shortens pauses (optionally paced like a video you edited), cuts filler words and drops retakes, and checks every join for clipped words (dry run first, new asset on apply) |
| Records a scripted walk through your running app (inputs are blurred) |
| Checks ffmpeg, fonts, Chrome, whisper, HyperFrames and keys |
What it does not do (yet)
No generative video or avatars yet. Adapters are planned (Phase 7, needs keys; the key fields already exist in Configure and do nothing until then). Until then those scenes render as titled placeholder cards, and
prompt_packwrites ready-to-paste prompts for each generator. Sora is intentionally not supported.No posting or analytics. It produces packages and post copy; you upload them. Platform "trending sounds" are added in each app, and
post.jsonreminds you of that.HyperFrames is optional, except for
motionscenes. The built-in ffmpeg renderer covers every other scene kind.motionpages (Claude-written code) need HyperFrames and Google Chrome; without them they render as a reported text stand-in.Some checks are approximations. Flash detection measures average brightness (it follows WCAG 2.3.1 but does not measure red flashes); the title-length band is a rule of thumb, reported as a warning only; motion density counts sudden changes, and
moving_pctcounts frames that move at all (smooth motion and crossfades included), so a very slow drift can sit near its threshold.The editor timeline is a starting point. It places every scene on its exact frame and keeps the audio and captions, but transitions become markers rather than rebuilt dissolves.
Whisper models are downloaded only with your consent (about 148 MB; 488 MB for the speaker-turn model). You can supply SRT/VTT captions instead.
Speaker turns are English-only and label two alternating speakers (S1/S2); rename them if there are more.
Subject tracking is automatic on macOS only. Elsewhere Claude marks the subject from shot sheets.
No colour grading yet. Dark or bright footage gets a warning, not a fix.
Demo capture never starts your app. You start it and give the URL, and every step is approved first.
SVG logos can't be the corner logo (ffmpeg can't read SVG); use a PNG. Project fonts must be TTF or OTF.
The roadmap is in docs/PLAN.md and the current state in docs/HANDOFF.md.
pnpm install
pnpm typecheck # tsc -b
pnpm test # vitest
pnpm schemas # regenerate schemas/*.schema.json
pnpm bundle # build dist/mcp.mjs (single-file ESM, committed)
pnpm smoke # start dist/mcp.mjs over stdio and check its tools
claude plugin validate --strict .claude-plugin/plugin.json # plugin + skills
claude plugin validate --strict . # marketplace
pnpm check # all of the above plus golden frames, stopping at the first failure
pnpm hooks # once: run pnpm check before every git pushAfter changing anything under packages/, rerun pnpm bundle and commit dist/mcp.mjs (pnpm check fails if the committed bundle is stale).
Package | Role |
| zod models → |
| project folders, cache, SQLite ledger, jobs |
| extractors (documents, web, repos, media) |
| ffmpeg, audio mix, captions, QA, ASR, beat detection |
| ffmpeg, footage and HyperFrames renderers; tokens, styles, scripts |
| platform contracts and layout zones |
| TTS backends |
| provider specs and prompt compilers for shot cards |
| the MCP server ( |
Data lives next to the code: skills/, templates/, styles/ (a project can add its own in <project>/styles/), music/, sfx/, fonts/, platform-specs/, provider-specs/ and research-specs/.
Tests that need real Chrome or a whisper model are skipped by default:
VS_TEST_RENDER=1 npx vitest run packages/renderer packages/mcp/src/stills.test.ts # real HyperFrames/Chrome renders
VS_TEST_RENDER=1 VS_UPDATE_GOLDEN=1 npx vitest run tests/golden-frames # record the motion example's goldens
VS_TEST_RENDER=1 npx vitest run packages/mcp/src/render-page.test.ts # render_js isolation on a local JS page
VS_TEST_WHISPER_MODEL=<path to a ggml model> npx vitest run packages/mcp/src/tighten.test.ts
VS_DEBUG_CAPTURE=1 … # trace every Chrome capture stepContributing
Guides for adding archetypes, style packs, platform packs and providers are in docs/contributing/. Ingested content is always treated as untrusted data: the plugin never executes code from sources in its own process. The one exception is opt-in render_js, which opens a page once in an isolated headless Chrome after you approve it, with every request checked by the engine and nothing clicked.
Credits
Some ideas re-expressed from latent-spaces/brag (MIT).
License
Apache-2.0. See LICENSE. The bundled fonts are OFL-1.1 (see fonts/README.md), the bundled music beds are CC0 (see music/README.md), and the bundled sound effects are CC0, synthesized by the plugin (see sfx/README.md).
Related MCP Connectors
Build, run, schedule, and publish AI video pipelines to YouTube and TikTok from any MCP client.
Render video and run AI media tasks from a single declarative JSON request.
- tonpitOAuthcom.tonpit
Video production studio for AI agents: AI media, motion graphics as code, timeline and export.
Transcode, host and caption video from a prompt. Fifteen tools, nine read-only, nothing deletes.
Related MCP Servers
- FlicenseAqualityBmaintenanceEnables AI-powered video editing in IDEs through a pipeline-driven workflow, including proxy ingestion, motion graphics generation, and high-fidelity rendering.4-
- FlicenseNot gradedqualityBmaintenanceEnables AI agents to validate and render branded videos from structured Video Specifications via the Model Context Protocol. Provides tools for health checks, video validation, and deterministic Remotion rendering.-
- AlicenseNot gradedqualityBmaintenanceEnables MCP clients to run project-scoped video editing workflows: propose and approve editing strategies, apply validated plans, review immutable versions, and export final renders via FFmpeg, with durable persistence and approval gates.1MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to turn natural-language creative direction, transcripts, and source media into fully structured, editable video projects, then verify and render delivery files.2MIT