Make an explainer video
make_explainerGenerate a narrated explainer video or AI host episode from a topic or script.
Instructions
AI HOST EPISODE: format 'host_episode' = one presenter talking to camera, the camera changing every piece (frontal, three-quarter, close), the same host and set throughout, native voice, 16:9. Host = creator (saved or preset) or hostImage; words = topic (written for you) or script (verbatim). It returns a 480p DRAFT; HD (720p) is a SEPARATE call with fromDraft (quote it with dryRun, run it only when the user asks). Otherwise: turn a TOPIC into a finished narrated explainer video, in one of TWO LANES (lane). 'blocks' (the default) = 10-second VIDEO blocks, one narrated line per block, hard cuts, a music bed under the voice: an EXPLAINER renders on Gemini Omni at 720p (9:16 by default, or 16:9), a FACELESS CHANNEL video (format:'faceless_channel', or channel history / kids / fairytale) on MiniMax H3 at 2K (16:9 by default) with five cuts per block. 'stills' = a picture film: writes a sectioned script, paints a BURST of pictures per section (about one every 1.5s — most of them one-detail edits of the frame before, so it reads as movement rather than a slideshow), narrates each section with TTS, holds each picture PERFECTLY STILL for its own slice of the narration (the motion is the CUT RATE — a slow move on a still shimmers), then composites the end card (and any on-screen text you asked for) with the Chrome+ffmpeg engine the ads use (text is never model-painted, so it never garbles). BURNED ON-SCREEN TEXT IS OFF BY DEFAULT — the narration carries the point and the pictures carry the story, so the film ships clean unless the user asks otherwise; captions:true adds held key points and subtitles:true adds narration-timed CAPS (see both). The stills lane is an image film WITH motion, not N video-model renders — that's what keeps it affordable. style picks the visual family: the default 'cinematic' is photoreal editorial; every other id is a STYLED, strictly non-photoreal look (illustrated / collage / clay / pixel …) that first renders ONE style-key image and then locks every scene to it, so the whole film holds one look. Cost at the default frame density: a ~130-credit hold for a 60s explainer on the default style, ~100 styled; frameDensity:'lean' roughly halves it and 'minimal' (one picture per section) is ~30. All settle to the exact per-frame image + narration spend (a longer target = more sections = more). Takes SEVERAL minutes — one image render per frame; independent frames are painted concurrently, so it is far faster than the frame count suggests. Needs the writing model and a narration voice engine connected. NOT the tool for a short product ad — use render_ad or generate_video for those, and make_template_ad for the deterministic native formats.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| lane | No | 'blocks' (default) = 10-second video blocks, real motion, one narrated line per block; 'stills' = the picture film (a still about every 1.5s, narrated, no video model; cheaper). Quote either with dryRun. | |
| music | No | music bed under the narration, measured to sit about 14 dB under the voice and sidechain-ducked beneath it. Omit and the KIDS and FAIRYTALE channels get their recommended bed COMPOSED for this film — those two are the only channels a bed is due on unasked, and it costs a small flat fee; every other channel ships dry. 'off' forces silence. 'library' takes a free curated track only, and ships dry when none is on file. NAME A MOOD (upbeat / calm / warm / epic / tense / playful / elegant / hype / chill / dramatic) or DESCRIBE it in words to compose one on ANY channel, at the same fee. hermoso_capabilities reports the exact figure as explainerMusicCredits; quote it before you turn a bed on or pick a mood. | |
| style | No | visual style: 'cinematic' (default, photoreal); styled shortcuts editorial_collage, flat_vector, stickman, whiteboard, ink_marker, silhouette, storybook, paper_diorama, isometric, claymation, pixel_art, watercolor, fluffy_toy, low_poly, stylized_3d, studio_3d (the Kids default), mannequin; or ANY look described in words ('80s anime cel animation'), locked across every frame. Ask rather than pick silently; a styled look costs more. | |
| topic | No | what the explainer should teach or explain — a topic or a short brief (host_episode: or pass `script`) | |
| voice | No | narration voice name — omit for the default warm read | |
| dryRun | No | true = return the exact credits this explainer reserves (its own pricing, stopped at the hold) and render nothing. Quote it before running one; try frameDensity lean or minimal when the balance is short. | |
| format | No | blocks lane: 'explainer' (default; Gemini Omni 720p, 16:9 or 9:16, runs exactly the length asked) or 'faceless_channel' (a YouTube/TikTok faceless channel video; MiniMax H3 at 2K, 16:9 by default, five hard cuts per 10s block, a whole number of blocks). Omit and a history / kids / fairytale channel is a faceless channel video. 'host_episode' = the AI host episode (see the top). | |
| script | No | host_episode: the exact words, said verbatim and split at natural breaks into 4-30s pieces | |
| cameras | No | host_episode: the rotation, ids frontal / three_quarter / close or framings in words; one entry = one fixed camera | |
| channel | No | the CHANNEL TYPE — it sets the pacing, the narration register and the default look, and is orthogonal to `style` (a named style always wins): explainer (casual second-person, fast cuts), history (witty chronological retelling / documentary), kids (fastest, question-first, warm teacher), fairytale (slow, atmospheric myth or folklore). Default 'explainer'. | |
| creator | No | host_episode: the host, a saved creator or a preset by name or id (list_creators) | |
| endCard | No | append the branded end card. DEFAULT FALSE — set true ONLY when the user asks for one | |
| setting | No | host_episode: the set in words (default: written to fit the topic) | |
| upscale | No | optional FINAL upscale — 2 doubles each side, 4 quadruples. Captions and the end card are burned BEFORE it so they upscale with the frame. It is priced BY LENGTH and it is the expensive part — several times the cost of rendering the film itself. hermoso_capabilities reports the exact figures per length as explainerUpscaleCredits. Never turn it on unasked: quote the number and let the user choose. | |
| captions | No | turn ON-SCREEN TEXT on. DEFAULT FALSE, and leave it false unless the user asks — the narration already says the point and the pictures carry it, so the clean film is the better default. `captions:true` on its own burns SUBTITLES (see below), because that is what a caption is for: showing what is being said when the phone is on mute. Slim white CAPS, thin black outline, bottom safe band, no plate, no box. | |
| brandName | No | brand name for the end card — omit to leave it unbranded | |
| fromDraft | No | host_episode: a finished draft's job id, to render it in HD (720p) with the same script, cameras, set and host | |
| hostImage | No | host_episode: a photo URL of the host instead (upload_file for a local file); a real person's face needs a paid plan | |
| subtitles | No | which on-screen text, once `captions` is on. LEAVE IT UNSET (or true) for SUBTITLES — every spoken word, in order, timed to the narration; free, no extra render, no extra credits, and there is NO cue limit, so the whole film is subtitled however long it runs (at most 5 words / 32 characters a line). Set it FALSE only if the user explicitly wants section HEADINGS instead: one short summary label held over each ~7-15s section. That is NOT what is being said — it is a label about it — so it is the wrong answer to "add captions" and to anyone watching on mute. `subtitles:true` also implies `captions:true`. TIMING: each cue is anchored to that section’s REAL measured narration length and distributed inside the section by character count — exact at every section boundary, approximate to a few tenths of a second within one. It is not a word-level speech clock, so never promise frame-accurate sync. | |
| thumbnail | No | host_episode: one thumbnail of the host (default true) | |
| aspectRatio | No | '9:16' default (a faceless channel video and a host episode default to 16:9) | |
| frameDensity | No | how many pictures per second of narration, and therefore what it costs. 'standard' (default) is a frame about every 1.5s — the density a stills film needs to read as a film rather than a slideshow; 'lean' is one about every 2.5s (the longest hold that still reads as a film, ~40% of the frames and ~40% of the cost); 'minimal' is one picture about every 3.5s, the cheapest and the longest any still is ever held, and it reads close to a slideshow. Only drop below the default if the user asked for something cheaper. | |
| durationSeconds | No | target length 20-120s (default 60); drives the section count — ~10s of narration each, 3-8 sections |