generate_captions
Generates animated captions from a spoken clip's transcript (the same generator as Frapea's UI): transcript words group into caption text clips over the clip's span, word-timed so the animation follows speech. One undo, and they share a captionGroupId — restyle them all via set_clip_properties with text.wholeCaptionGroup. Needs a transcript (get_transcript tells you). Returns the created clip ids. PLACEMENT: they JOIN the existing caption track whenever they fit in its gaps, and only open a new one when they would overlap (a second language, a second layer). So captioning four narration segments in four calls leaves ONE caption track, not four — no tidying needed afterwards for this. Call manage_tracks {action:'tidy'} at the end of the build for the empty tracks left by everything else. They join the LOWEST caption track, so a picture layer you add ABOVE it after an earlier pass will cover them — captions render behind b-roll rather than looking wrong, which is the hardest kind of mistake to notice. Move the whole lane with manage_tracks {action:'merge'} if you want the captions on top.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| token | Yes | The session token `connect` returned. Pass it on every call — it says which browser to drive. | |
| clipId | Yes | The spoken clip to caption. | |
| preset | No | Style preset: clean (default) | boldPop | karaoke | highlight | minimal | editorial. | |
| position | No | bottom (default) | top. | |
| textCase | No | original (default) | upper | lower. | |
| projectId | Yes | Project id from get_projects. | |
| timelineId | Yes | Timeline containing the spoken clip. | |
| censorProfanity | No | Star out profane words. | |
| maxWordsPerCaption | No | Words per caption before a new one starts (default 4). |