PROJECT-SCOPED: this call acts only on the explicit project_id and returns the project identity with its result. Burned captions. mode='from_transcript' (word-timed from the real transcript, recommended) or mode='off', or items=[{text,start,end,style?}] (source seconds) for text the user dictates. PREMIUM PRESETS (style.preset) are professionally designed looks with real fonts. SAFE DEFAULTS: 'clean' (white Plus Jakarta Sans, complete short phrases, size-only hierarchy), 'documentary' (restrained subtitles on a translucent contrast panel), and 'broadcast' (left-aligned news/explainer lower third). SOCIAL/CREATIVE: 'reels' (FLAGSHIP short-form system: tight two-line hierarchy, warm hero word and multi-stage elastic word landings), 'podcast' (bold white words land on screen as spoken, keywords light up in the accent color, get a highlight box or serif italics, numbers render HUGE), 'beast' (loud MrBeast-style: ALL-CAPS impact font, centered, the spoken word pops in the accent color), 'karaoke' (an accent box follows each spoken word), 'spotlight' (ONE glowing word at a time, centred, uppercase — the modern single-word look for hype/motivation/fast talking; the ONLY preset that belongs mid-frame), 'elegant' (calm lower-third, serif-italic accents — interviews/luxury), 'lyric' (phrase-led music/quote typography), plus stacked/iridescent/chrome/editorial/fashion/luxe/impact/retro/neon composed looks; 'classic' is the plain legacy look. If style is omitted, the tool chooses a coherent preset deterministically from the measured format, speech pace and brief — it does not make every project wear the same caption skin. PLACEMENT: multi-word presets default to the BOTTOM, clear of the face — do not move them to 'middle'; only a single-word-at-a-time look may sit centred. With a preset, sparse emphasis is AUTO-SELECTED from the KEPT transcript when emphasis_words is omitted; measured vocal stress leads on short-form/cached audio, with numbers and semantic outcome words as fallback. pass a verbatim list only when specific words are required, or [] to explicitly disable hierarchy. highlight_color sets the accent (default warm yellow); uppercase overrides the preset's casing; position bottom/top/middle overrides its placement and LOCKS that band for the whole video (no shot-by-shot position changes); anchor_y similarly locks an exact vertical frame fraction. Omit both to let collision-aware placement adapt by shot. Other style fields: color '#RRGGBB', size s|m|l|xl (presets are already big at 'm'), size_scale 0.5-3.0, dynamic:true (legacy karaoke, no preset), animation fade|pop|slide_up|punch|blur_in|whip|flash|rise|drop|elastic|bounce|swing|zoom_blur, or 'none' to turn a preset's animation OFF (instant words), single_line:true to guarantee one rendered row per transcript-caption state regardless of the preset's normal flow/stack layout, max_words_per_caption 1-16. Default modern look is 1-2 words at a time, stacked levels, fade in (preset 'stacked', max_words_per_caption 2) — not a sentence subtitle. Example — modern reel captions: {mode:'from_transcript', style:{preset:'stacked', animation:'fade'}, max_words_per_caption:2}. Example — dictated title card: {items:[{text:'CHAPTER ONE', start:0, end:2.5, style:{preset:'beast'}}]}. Stack presets (stacked/iridescent/chrome/fashion/luxe/editorial/impact) compose 1-2 words across lines of very different SIZES; font picks a bundled family, emphasis 'big' enlarges keywords WITHOUT recolouring them, leading below 1.0 overlaps the lines so a smaller word fades in behind the hero. Production controls include outline_color/outline_width, shadow, background_color/background_opacity, tracking and text_align.