Clip a long video
clip_videoCut one long video into ranked, ready-to-post short clips with burned captions and subject-aware vertical reframing, using transcription to pick self-contained moments.
Instructions
Cut ONE long video into several RANKED, ready-to-post short clips (podcast, webinar, interview, talk, long ad cut -> Reels/Shorts/TikTok). Transcribes with timestamps, picks the strongest SELF-CONTAINED moments, then cuts + reframes each with ffmpeg — no video model renders anything, so it is fast and cheap. THE VERTICAL REFRAME IS SUBJECT-AWARE: one cheap vision call per clip (billed as its own event) picks a SINGLE crop offset held for the whole clip, so a speaker sitting camera-left is not cropped out and the framing never drifts inside a clip; with nothing to discard or no single subject it stays dead centre — read reframedToSubject and each clip's reframeWhy back rather than assuming either way. ACCEPTS: a YouTube link (or Vimeo / Loom / Dailymotion / Streamable / Rumble / Wistia / Twitch / TED), a direct https .mp4/.mov/.webm, or a Hermoso /generated/ URL (upload_file turns a local file into one). NOT supported: TikTok / Instagram / Facebook links, and anything age-restricted, private, members-only, geo-blocked or still LIVE — those fail fast with the real reason and are fully refunded, so ask for a direct file or an upload rather than retrying. Source ~15s to ~600MB; only the first ~40 minutes is analysed (truncated:true says so). Cost: a ~7-credit hold settled to the exact transcription + encode cost, plus the clip-selection model's tokens as their own small event. RETURNS clips[] — each its OWN served mp4 URL, title, hook, ready-to-post caption, 0-100 score and source timecode. SUBTITLES ARE BURNED IN BY DEFAULT (slim white CAPS, thin black outline, bottom safe band, no box) because short-form is watched on mute; captions:false for clean footage. TIMING IS APPROXIMATE, NOT WORD-LEVEL — cues follow the transcript's per-sentence timestamps, split by character count; never promise frame-accurate sync. captionsBurned counts the clips that really carry a burned track and captionNote says why any are bare.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | how many clips to cut, 1-8 (default 4) | |
| video | Yes | the long video to clip — a YouTube/Vimeo/Loom/Dailymotion/Streamable/Rumble/Wistia/Twitch/TED watch URL, a direct https .mp4/.mov/.webm, or a Hermoso /generated/ URL | |
| captions | No | burn subtitles into every clip. DEFAULT TRUE — a clip cut from a podcast or a talk is watched on mute, and the words are the product. Set false for clean footage. A clip whose window carries no readable speech is delivered bare rather than captioned with a guess, and the result says which. | |
| aspectRatio | No | clip shape: '9:16' (default), '1:1', '16:9', any 'W:H', or 'keep' for the source framing |