Skip to main content
Glama
hermoso-ai

Hermoso

Official

Make an explainer video

make_explainer

Generate a narrated explainer video or AI host episode from a topic or script.

Instructions

AI HOST EPISODE: format 'host_episode' = one presenter talking to camera, the camera changing every piece (frontal, three-quarter, close), the same host and set throughout, native voice, 16:9. Host = creator (saved or preset) or hostImage; words = topic (written for you) or script (verbatim). It returns a 480p DRAFT; HD (720p) is a SEPARATE call with fromDraft (quote it with dryRun, run it only when the user asks). Otherwise: turn a TOPIC into a finished narrated explainer video, in one of TWO LANES (lane). 'blocks' (the default) = 10-second VIDEO blocks, one narrated line per block, hard cuts, a music bed under the voice: an EXPLAINER renders on Gemini Omni at 720p (9:16 by default, or 16:9), a FACELESS CHANNEL video (format:'faceless_channel', or channel history / kids / fairytale) on MiniMax H3 at 2K (16:9 by default) with five cuts per block. 'stills' = a picture film: writes a sectioned script, paints a BURST of pictures per section (about one every 1.5s — most of them one-detail edits of the frame before, so it reads as movement rather than a slideshow), narrates each section with TTS, holds each picture PERFECTLY STILL for its own slice of the narration (the motion is the CUT RATE — a slow move on a still shimmers), then composites the end card (and any on-screen text you asked for) with the Chrome+ffmpeg engine the ads use (text is never model-painted, so it never garbles). BURNED ON-SCREEN TEXT IS OFF BY DEFAULT — the narration carries the point and the pictures carry the story, so the film ships clean unless the user asks otherwise; captions:true adds held key points and subtitles:true adds narration-timed CAPS (see both). The stills lane is an image film WITH motion, not N video-model renders — that's what keeps it affordable. style picks the visual family: the default 'cinematic' is photoreal editorial; every other id is a STYLED, strictly non-photoreal look (illustrated / collage / clay / pixel …) that first renders ONE style-key image and then locks every scene to it, so the whole film holds one look. Cost at the default frame density: a ~130-credit hold for a 60s explainer on the default style, ~100 styled; frameDensity:'lean' roughly halves it and 'minimal' (one picture per section) is ~30. All settle to the exact per-frame image + narration spend (a longer target = more sections = more). Takes SEVERAL minutes — one image render per frame; independent frames are painted concurrently, so it is far faster than the frame count suggests. Needs the writing model and a narration voice engine connected. NOT the tool for a short product ad — use render_ad or generate_video for those, and make_template_ad for the deterministic native formats.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
laneNo'blocks' (default) = 10-second video blocks, real motion, one narrated line per block; 'stills' = the picture film (a still about every 1.5s, narrated, no video model; cheaper). Quote either with dryRun.
musicNomusic bed under the narration, measured to sit about 14 dB under the voice and sidechain-ducked beneath it. Omit and the KIDS and FAIRYTALE channels get their recommended bed COMPOSED for this film — those two are the only channels a bed is due on unasked, and it costs a small flat fee; every other channel ships dry. 'off' forces silence. 'library' takes a free curated track only, and ships dry when none is on file. NAME A MOOD (upbeat / calm / warm / epic / tense / playful / elegant / hype / chill / dramatic) or DESCRIBE it in words to compose one on ANY channel, at the same fee. hermoso_capabilities reports the exact figure as explainerMusicCredits; quote it before you turn a bed on or pick a mood.
styleNovisual style: 'cinematic' (default, photoreal); styled shortcuts editorial_collage, flat_vector, stickman, whiteboard, ink_marker, silhouette, storybook, paper_diorama, isometric, claymation, pixel_art, watercolor, fluffy_toy, low_poly, stylized_3d, studio_3d (the Kids default), mannequin; or ANY look described in words ('80s anime cel animation'), locked across every frame. Ask rather than pick silently; a styled look costs more.
topicNowhat the explainer should teach or explain — a topic or a short brief (host_episode: or pass `script`)
voiceNonarration voice name — omit for the default warm read
dryRunNotrue = return the exact credits this explainer reserves (its own pricing, stopped at the hold) and render nothing. Quote it before running one; try frameDensity lean or minimal when the balance is short.
formatNoblocks lane: 'explainer' (default; Gemini Omni 720p, 16:9 or 9:16, runs exactly the length asked) or 'faceless_channel' (a YouTube/TikTok faceless channel video; MiniMax H3 at 2K, 16:9 by default, five hard cuts per 10s block, a whole number of blocks). Omit and a history / kids / fairytale channel is a faceless channel video. 'host_episode' = the AI host episode (see the top).
scriptNohost_episode: the exact words, said verbatim and split at natural breaks into 4-30s pieces
camerasNohost_episode: the rotation, ids frontal / three_quarter / close or framings in words; one entry = one fixed camera
channelNothe CHANNEL TYPE — it sets the pacing, the narration register and the default look, and is orthogonal to `style` (a named style always wins): explainer (casual second-person, fast cuts), history (witty chronological retelling / documentary), kids (fastest, question-first, warm teacher), fairytale (slow, atmospheric myth or folklore). Default 'explainer'.
creatorNohost_episode: the host, a saved creator or a preset by name or id (list_creators)
endCardNoappend the branded end card. DEFAULT FALSE — set true ONLY when the user asks for one
settingNohost_episode: the set in words (default: written to fit the topic)
upscaleNooptional FINAL upscale — 2 doubles each side, 4 quadruples. Captions and the end card are burned BEFORE it so they upscale with the frame. It is priced BY LENGTH and it is the expensive part — several times the cost of rendering the film itself. hermoso_capabilities reports the exact figures per length as explainerUpscaleCredits. Never turn it on unasked: quote the number and let the user choose.
captionsNoturn ON-SCREEN TEXT on. DEFAULT FALSE, and leave it false unless the user asks — the narration already says the point and the pictures carry it, so the clean film is the better default. `captions:true` on its own burns SUBTITLES (see below), because that is what a caption is for: showing what is being said when the phone is on mute. Slim white CAPS, thin black outline, bottom safe band, no plate, no box.
brandNameNobrand name for the end card — omit to leave it unbranded
fromDraftNohost_episode: a finished draft's job id, to render it in HD (720p) with the same script, cameras, set and host
hostImageNohost_episode: a photo URL of the host instead (upload_file for a local file); a real person's face needs a paid plan
subtitlesNowhich on-screen text, once `captions` is on. LEAVE IT UNSET (or true) for SUBTITLES — every spoken word, in order, timed to the narration; free, no extra render, no extra credits, and there is NO cue limit, so the whole film is subtitled however long it runs (at most 5 words / 32 characters a line). Set it FALSE only if the user explicitly wants section HEADINGS instead: one short summary label held over each ~7-15s section. That is NOT what is being said — it is a label about it — so it is the wrong answer to "add captions" and to anyone watching on mute. `subtitles:true` also implies `captions:true`. TIMING: each cue is anchored to that section’s REAL measured narration length and distributed inside the section by character count — exact at every section boundary, approximate to a few tenths of a second within one. It is not a word-level speech clock, so never promise frame-accurate sync.
thumbnailNohost_episode: one thumbnail of the host (default true)
aspectRatioNo'9:16' default (a faceless channel video and a host episode default to 16:9)
frameDensityNohow many pictures per second of narration, and therefore what it costs. 'standard' (default) is a frame about every 1.5s — the density a stills film needs to read as a film rather than a slideshow; 'lean' is one about every 2.5s (the longest hold that still reads as a film, ~40% of the frames and ~40% of the cost); 'minimal' is one picture about every 3.5s, the cheapest and the longest any still is ever held, and it reads close to a slideshow. Only drop below the default if the user asked for something cheaper.
durationSecondsNotarget length 20-120s (default 60); drives the section count — ~10s of narration each, 3-8 sections

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed14 schema fields changedv0.1.366
    • changedInput schema / properties / aspectRatio / description
      Previous value: -"'9:16' default"New value: +"'9:16' default (a faceless channel video and a host episode default to 16:9)"
    • addedInput schema / properties / cameras
      Added value: +{
      +  "description": "host_episode: the rotation, ids frontal / three_quarter / close or framings in words; one entry = one fixed camera",
      +  "items": {
      +    "type": "string"
      +  },
      +  "type": "array"
      +}
    • addedInput schema / properties / creator
      Added value: +{
      +  "description": "host_episode: the host, a saved creator or a preset by name or id (list_creators)",
      +  "type": "string"
      +}
    • addedInput schema / properties / dryRun
      Added value: +{
      +  "description": "true = return the exact credits this explainer reserves (its own pricing, stopped at the hold) and render nothing. Quote it before running one; try frameDensity lean or minimal when the balance is short.",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / format
      Added value: +{
      +  "description": "blocks lane: 'explainer' (default; Gemini Omni 720p, 16:9 or 9:16, runs exactly the length asked) or 'faceless_channel' (a YouTube/TikTok faceless channel video; MiniMax H3 at 2K, 16:9 by default, five hard cuts per 10s block, a whole number of blocks). Omit and a history / kids / fairytale channel is a faceless channel video. 'host_episode' = the AI host episode (see the top).",
      +  "enum": [
      +    "explainer",
      +    "faceless_channel",
      +    "host_episode"
      +  ],
      +  "type": "string"
      +}
    • changedInput schema / properties / frameDensity / description
      Previous value: -"how many pictures per second of narration, and therefore what it costs. 'standard' (default) is a frame about every 1.5s — the density a stills film needs to read as a film rather than a slideshow; 'lean' is one about every 2.5s (the longest hold that still reads as a film, ~40% of the frames and ~40% of the cost); 'minimal' is ONE picture per narration section, which is cheapest and is frankly a slideshow. Only drop below the default if the user asked for something cheaper."New value: +"how many pictures per second of narration, and therefore what it costs. 'standard' (default) is a frame about every 1.5s — the density a stills film needs to read as a film rather than a slideshow; 'lean' is one about every 2.5s (the longest hold that still reads as a film, ~40% of the frames and ~40% of the cost); 'minimal' is one picture about every 3.5s, the cheapest and the longest any still is ever held, and it reads close to a slideshow. Only drop below the default if the user asked for something cheaper."
    • addedInput schema / properties / fromDraft
      Added value: +{
      +  "description": "host_episode: a finished draft's job id, to render it in HD (720p) with the same script, cameras, set and host",
      +  "type": "string"
      +}
    • addedInput schema / properties / hostImage
      Added value: +{
      +  "description": "host_episode: a photo URL of the host instead (upload_file for a local file); a real person's face needs a paid plan",
      +  "type": "string"
      +}
    • addedInput schema / properties / lane
      Added value: +{
      +  "description": "'blocks' (default) = 10-second video blocks, real motion, one narrated line per block; 'stills' = the picture film (a still about every 1.5s, narrated, no video model; cheaper). Quote either with dryRun.",
      +  "enum": [
      +    "blocks",
      +    "stills"
      +  ],
      +  "type": "string"
      +}
    • addedInput schema / properties / script
      Added value: +{
      +  "description": "host_episode: the exact words, said verbatim and split at natural breaks into 4-30s pieces",
      +  "type": "string"
      +}
    • addedInput schema / properties / setting
      Added value: +{
      +  "description": "host_episode: the set in words (default: written to fit the topic)",
      +  "type": "string"
      +}
    • addedInput schema / properties / thumbnail
      Added value: +{
      +  "description": "host_episode: one thumbnail of the host (default true)",
      +  "type": "boolean"
      +}
    • changedInput schema / properties / topic / description
      Previous value: -"what the explainer should teach or explain — a topic or a short brief"New value: +"what the explainer should teach or explain — a topic or a short brief (host_episode: or pass `script`)"
    • removedInput schema / required
      Removed value: -[
      -  "topic"
      -]
  2. Changed3 schema fields changedv0.1.285
    • changedInput schema / properties / music / description
      Previous value: -"music bed under the narration, measured to sit about 14 dB under the voice and sidechain-ducked beneath it. Omit and the KIDS and FAIRYTALE channels get their recommended bed COMPOSED for this film — those two are the only channels a bed is due on unasked, and it costs a small flat fee; every other channel ships dry. 'off' forces silence. 'library' takes a free curated track only, and ships dry when none is on file. NAME A MOOD — upbeat / calm / warm / epic / tense / playful / elegant / hype / chill / dramatic — to compose one on ANY channel, at the same fee. hermoso_capabilities reports the exact figure as explainerMusicCredits; quote it before you turn a bed on or pick a mood."New value: +"music bed under the narration, measured to sit about 14 dB under the voice and sidechain-ducked beneath it. Omit and the KIDS and FAIRYTALE channels get their recommended bed COMPOSED for this film — those two are the only channels a bed is due on unasked, and it costs a small flat fee; every other channel ships dry. 'off' forces silence. 'library' takes a free curated track only, and ships dry when none is on file. NAME A MOOD (upbeat / calm / warm / epic / tense / playful / elegant / hype / chill / dramatic) or DESCRIBE it in words to compose one on ANY channel, at the same fee. hermoso_capabilities reports the exact figure as explainerMusicCredits; quote it before you turn a bed on or pick a mood."
    • changedInput schema / properties / style / description
      Previous value: -"visual style. 'cinematic' (default) is photoreal; the rest are non-photoreal styled looks — editorial_collage (halftone cutouts + marker accents), flat_vector, stickman, whiteboard, ink_marker, silhouette, storybook (gouache), paper_diorama, isometric, claymation, pixel_art, watercolor, fluffy_toy (felted plush), low_poly, stylized_3d (matte clay render), studio_3d (preschool toy 3D on a white sweep — the Kids default), mannequin (clay-render reenactment figures — a History alternate). Ask the user which they want rather than picking silently; a styled pick costs more (see the cost note)."New value: +"visual style: 'cinematic' (default, photoreal); styled shortcuts editorial_collage, flat_vector, stickman, whiteboard, ink_marker, silhouette, storybook, paper_diorama, isometric, claymation, pixel_art, watercolor, fluffy_toy, low_poly, stylized_3d, studio_3d (the Kids default), mannequin; or ANY look described in words ('80s anime cel animation'), locked across every frame. Ask rather than pick silently; a styled look costs more."
    • removedInput schema / properties / style / enum
      Removed value: -[
      -  "cinematic",
      -  "editorial_collage",
      -  "flat_vector",
      -  "stickman",
      -  "whiteboard",
      -  "ink_marker",
      -  "silhouette",
      -  "storybook",
      -  "paper_diorama",
      -  "isometric",
      -  "claymation",
      -  "pixel_art",
      -  "watercolor",
      -  "fluffy_toy",
      -  "low_poly",
      -  "stylized_3d",
      -  "studio_3d",
      -  "mannequin"
      -]
  3. Changed1 schema field changedv0.1.225
    • changedInput schema / properties / endCard / description
      Previous value: -"append the branded end card (default true)"New value: +"append the branded end card. DEFAULT FALSE — set true ONLY when the user asks for one"
  4. Changed1 schema field changedv0.1.162
    • changedInput schema / properties / frameDensity / description
      Previous value: -"how many pictures per second of narration, and therefore what it costs. 'standard' (default) is a frame about every 1.5s — the density Higgsfield's own stills pipeline enforces; 'lean' is one about every 2.5s (the longest hold that still reads as a film, ~40% of the frames and ~40% of the cost); 'minimal' is ONE picture per narration section, which is cheapest and is frankly a slideshow. Only drop below the default if the user asked for something cheaper."New value: +"how many pictures per second of narration, and therefore what it costs. 'standard' (default) is a frame about every 1.5s — the density a stills film needs to read as a film rather than a slideshow; 'lean' is one about every 2.5s (the longest hold that still reads as a film, ~40% of the frames and ~40% of the cost); 'minimal' is ONE picture per narration section, which is cheapest and is frankly a slideshow. Only drop below the default if the user asked for something cheaper."
  5. Addedv0.1.161

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only flag readOnly=false/openWorld=true/idempotent=false/destructive=false; the description adds far more: it produces a 480p draft with HD as a separate fromDraft call, takes several minutes, concurrency behavior, credit costs with exact figures, the music-bed fee, and that upscale is priced by length and expensive. Defaults for captions/endCard/music shipping dry are all disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The content is valuable but it is a dense wall of text opening on a niche sub-format ('AI HOST EPISODE') before the general purpose, and several ideas (captions/subtitles, upscale cost) are restated in multiple places. It is front-loaded only after the reader wades through the host-episode preamble, which buries the primary use case.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 23-parameter, zero-required, multi-format tool with no output schema, the description covers lane selection, formats, aspect defaults, cost bounds, timing expectations, prerequisites, and sibling routing. Nothing an agent needs to call it correctly is missing, and the draft-vs-HD flow is spelled out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, but the description still adds real meaning: cost and pacing implications of frameDensity, the mood vocabulary for music and when a bed is free vs charged, the subtitle-vs-headings distinction for subtitles, and the exact rule that captions:true alone means subtitles. This is well past the baseline-3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'turn a TOPIC into a finished narrated explainer video,' then enumerates its concrete variants (host_episode, blocks/stills lanes, explainer vs faceless_channel). It explicitly separates itself from adjacent tools ('NOT the tool for a short product ad — use render_ad or generate_video'). An agent can place it against siblings without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when/when-not and named alternatives: host_episode vs blocks vs stills, which channel triggers faceless, and the direct exclusion routing to render_ad / generate_video / make_template_ad. Prerequisites are named too (writing model and narration voice engine must be connected), plus guidance to quote with dryRun before running.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools