generate_voiceover
Generates an AI voiceover from a script, entirely on-device. The WAV lands in the project's media library as a normal audio asset (with waveform and analysis), the script is stored as asset metadata, and a script-corrected word-level transcript is written so captions work immediately.
CALL THIS BEFORE YOU CUT THE PICTURE. A script does not tell you how long it takes to SAY, so a narrated edit assembled first is built on guessed cut points that the real audio then breaks — every section, graphic and transition has to be retimed. Generate the narration, read its word timings with get_transcript, then lay picture against those frames.
GENERATE IN SEGMENTS, NOT ONE BLOCK — call this 2-4 times (hook / body / close, or one per section) and lay the parts on the timeline. A single long generation comes out FLAT and evenly paced, everything pressed together at one energy; segmenting gives each part its own intent, puts natural air at the joins, and lets you redo one section instead of all of it. Write for the ear: short sentences, one idea each, LINE BREAKS between them — the strongest signal the model has to actually stop. See read_skill {topic:'voiceover'}. Async: returns {jobId, fileName} — fileName is the audio file's final name in the project's media/ folder, decided up front (WAV for local providers, MP3 for ElevenLabs). The job runs on the analysis MASTER tab and its request is persisted on disk: it survives tab closes and browser restarts, and get_job_status answers from disk even in a fresh session. Poll until done; the finished job also carries the asset's mediaRef. Voices/languages come from list_voiceover_models; a local model must be downloaded first (phase 'ready'), otherwise this returns PERMISSION_REQUIRED.
PREMIUM (provider 'elevenlabs' — SPENDS THE USER'S CREDITS): check list_voiceover_models' elevenlabs status object first. Structured refusals you must handle: NO_API_KEY (no key stored — ask the USER to add their key in the Voiceover generator's ElevenLabs onboarding; never ask for, or pass, the key yourself), INSUFFICIENT_CREDITS (carries needed vs remaining — shorten the script or ask the user), CONSENT_REQUIRED (paid features are set to 'ask': a consent dialog with the cost is now open in Frapea and the error carries an actionId — TELL THE USER in your visible reply to confirm it, call await_user_action with that actionId, and on 'granted' retry this call with consentActionId set to it; 'denied' means drop it).
AUDIO TAGS (ElevenLabs): inline bracket tags like [calm], [excited], [whispers] direct the delivery ONLY on models whose list_voiceover_models entry says supportsAudioTags: true (the v3 family). Every other model reads them ALOUD — so only write tags when the chosen model supports them. As a safety net, tags sent to a non-supporting model are stripped before synthesis and the success message says so.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The script to speak (multi-line ok). | |
| speed | No | Speaking rate multiplier (default 1). Local providers accept 0.5–2; ElevenLabs 0.7–1.2. | |
| token | Yes | The session token `connect` returned. Pass it on every call — it says which browser to drive. | |
| voice | No | Voice id from the provider's manifest (default: its first voice for the chosen language). | |
| modelId | No | ElevenLabs only: TTS model id from list_voiceover_models (default 'eleven_multilingual_v2'; turbo/flash cost half). Check the model's supportsAudioTags before writing [calm]-style tags into the text — non-v3 models would read them aloud. | |
| fileName | No | Optional base name for the audio file written into media/. | |
| language | No | Language tag from the provider's manifest (e.g. 'en-us'). Ignored for ElevenLabs — its voices are multilingual. | |
| provider | No | TTS provider id from list_voiceover_models (default 'kokoro'; 'elevenlabs' = premium, needs the user's key + consent). | |
| projectId | Yes | Project id from get_projects. | |
| voiceSettings | No | ElevenLabs only: voice knobs, each optional. stability, similarityBoost and style are 0–1; speakerBoost is a boolean. | |
| consentActionId | No | The actionId from a CONSENT_REQUIRED refusal, after await_user_action returned 'granted'. Good for one job. |