Skip to main content
Glama
hermoso-ai

Hermoso

Official

Make an explainer video

make_explainer

Turn any topic into a finished narrated explainer video with synchronized imagery, voiceover, and optional on-screen text.

Instructions

Turn a TOPIC into a finished narrated explainer video. Writes a sectioned script, paints a BURST of pictures per section (about one every 1.5s — most of them one-detail edits of the frame before, so it reads as movement rather than a slideshow), narrates each section with TTS, holds each picture PERFECTLY STILL for its own slice of the narration (the motion is the CUT RATE — a slow move on a still shimmers), then composites the end card (and any on-screen text you asked for) with the Chrome+ffmpeg engine the ads use (text is never model-painted, so it never garbles). BURNED ON-SCREEN TEXT IS OFF BY DEFAULT — the narration carries the point and the pictures carry the story, so the film ships clean unless the user asks otherwise; captions:true adds held key points and subtitles:true adds narration-timed CAPS (see both). It is an image film WITH motion, not N video-model renders — that's what keeps it affordable. style picks the visual family: the default 'cinematic' is photoreal editorial; every other id is a STYLED, strictly non-photoreal look (illustrated / collage / clay / pixel …) that first renders ONE style-key image and then locks every scene to it, so the whole film holds one look. Cost at the default frame density: a ~130-credit hold for a 60s explainer on the default style, ~100 styled; frameDensity:'lean' roughly halves it and 'minimal' (one picture per section) is ~30. All settle to the exact per-frame image + narration spend (a longer target = more sections = more). Takes SEVERAL minutes — one image render per frame; independent frames are painted concurrently, so it is far faster than the frame count suggests. Needs the writing model and a narration voice engine connected. NOT the tool for a short product ad — use render_ad or generate_video for those, and make_template_ad for the deterministic native formats.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
musicNomusic bed under the narration, measured to sit about 14 dB under the voice and sidechain-ducked beneath it. Omit and the KIDS and FAIRYTALE channels get their recommended bed COMPOSED for this film — those two are the only channels a bed is due on unasked, and it costs a small flat fee; every other channel ships dry. 'off' forces silence. 'library' takes a free curated track only, and ships dry when none is on file. NAME A MOOD — upbeat / calm / warm / epic / tense / playful / elegant / hype / chill / dramatic — to compose one on ANY channel, at the same fee. hermoso_capabilities reports the exact figure as explainerMusicCredits; quote it before you turn a bed on or pick a mood.
styleNovisual style. 'cinematic' (default) is photoreal; the rest are non-photoreal styled looks — editorial_collage (halftone cutouts + marker accents), flat_vector, stickman, whiteboard, ink_marker, silhouette, storybook (gouache), paper_diorama, isometric, claymation, pixel_art, watercolor, fluffy_toy (felted plush), low_poly, stylized_3d (matte clay render), studio_3d (preschool toy 3D on a white sweep — the Kids default), mannequin (clay-render reenactment figures — a History alternate). Ask the user which they want rather than picking silently; a styled pick costs more (see the cost note).
topicYeswhat the explainer should teach or explain — a topic or a short brief
voiceNonarration voice name — omit for the default warm read
channelNothe CHANNEL TYPE — it sets the pacing, the narration register and the default look, and is orthogonal to `style` (a named style always wins): explainer (casual second-person, fast cuts), history (witty chronological retelling / documentary), kids (fastest, question-first, warm teacher), fairytale (slow, atmospheric myth or folklore). Default 'explainer'.
endCardNoappend the branded end card (default true)
upscaleNooptional FINAL upscale — 2 doubles each side, 4 quadruples. Captions and the end card are burned BEFORE it so they upscale with the frame. It is priced BY LENGTH and it is the expensive part — several times the cost of rendering the film itself. hermoso_capabilities reports the exact figures per length as explainerUpscaleCredits. Never turn it on unasked: quote the number and let the user choose.
captionsNoturn ON-SCREEN TEXT on. DEFAULT FALSE, and leave it false unless the user asks — the narration already says the point and the pictures carry it, so the clean film is the better default. `captions:true` on its own burns SUBTITLES (see below), because that is what a caption is for: showing what is being said when the phone is on mute. Slim white CAPS, thin black outline, bottom safe band, no plate, no box.
brandNameNobrand name for the end card — omit to leave it unbranded
subtitlesNowhich on-screen text, once `captions` is on. LEAVE IT UNSET (or true) for SUBTITLES — every spoken word, in order, timed to the narration; free, no extra render, no extra credits, and there is NO cue limit, so the whole film is subtitled however long it runs (at most 5 words / 32 characters a line). Set it FALSE only if the user explicitly wants section HEADINGS instead: one short summary label held over each ~7-15s section. That is NOT what is being said — it is a label about it — so it is the wrong answer to "add captions" and to anyone watching on mute. `subtitles:true` also implies `captions:true`. TIMING: each cue is anchored to that section’s REAL measured narration length and distributed inside the section by character count — exact at every section boundary, approximate to a few tenths of a second within one. It is not a word-level speech clock, so never promise frame-accurate sync.
aspectRatioNo'9:16' default
frameDensityNohow many pictures per second of narration, and therefore what it costs. 'standard' (default) is a frame about every 1.5s — the density a stills film needs to read as a film rather than a slideshow; 'lean' is one about every 2.5s (the longest hold that still reads as a film, ~40% of the frames and ~40% of the cost); 'minimal' is ONE picture per narration section, which is cheapest and is frankly a slideshow. Only drop below the default if the user asked for something cheaper.
durationSecondsNotarget length 20-120s (default 60); drives the section count — ~10s of narration each, 3-8 sections
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are minimal (all four hints are false), so the description carries the full burden — and it delivers extensively. It discloses the long-running nature ('Takes SEVERAL minutes — one image render per frame'), the concurrency trait ('independent frames are painted concurrently, so it is far faster than the frame count suggests'), the cost model with concrete credit figures per style and density tier, the motion mechanics (hold-still + cut rate + shimmer), the text pipeline that avoids garbling ('text is never model-painted'), and the on-screen text default. No contradiction with annotations: readOnlyHint=false correctly matches a tool that produces a new artifact.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long (~450 words), but for a 13-parameter video-generation tool with subtle defaults, a cost model, and sibling differentiation, nearly every sentence earns its place. Structure is logical: purpose → motion model → text defaults → style → cost → performance → prerequisites → disambiguation. It drops to a 4 rather than 5 because a few passages (e.g., the 'slow move on a still shimmers' physics and the parenthetical on one-detail edits) add flavor more than agent decision value, and the length makes front-loading slightly harder to absorb at a glance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool of this complexity with no output schema and sparse annotations, the description is remarkably complete: it covers outcome, cost, timing, prerequisites, defaults, and disambiguation. The one genuine gap is the return value — for a multi-minute asynchronous operation, the agent is never told what comes back (a job handle? a URL?) or how to retrieve the result. Given sibling tools like get_job imply an async job pattern, this is a meaningful omission that keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Even with 100% schema coverage (baseline 3), the description adds substantial meaning beyond the schema. It explains the music default behavior (which channels get a composed bed unasked, mood-name composition on any channel, dry shipping, and how to look up exact credits), the captions-vs-subtitles-vs-section-headings distinction with timing approximation caveats, the upscale pricing warning and quote-before-acting rule, the cost implications of each frameDensity tier and what makes a stills film read as a film, and per-style visual descriptions. This materially improves the agent's ability to set parameters correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific, unambiguous statement of purpose: 'Turn a TOPIC into a finished narrated explainer video.' It then walks through the full production pipeline (script → pictures → TTS narration → end-card composite), and explicitly differentiates itself from siblings: 'NOT the tool for a short product ad — use render_ad or generate_video for those, and make_template_ad for the deterministic native formats.' An agent can immediately tell what this tool does and how it differs from the closest alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use and when-not-to-use guidance with named alternatives (render_ad, generate_video, make_template_ad). It also gives strong default-handling instructions: leave captions/subtitles off unless the user asks, never enable upscale unasked and quote the credit figure first, 'Ask the user which [style] they want rather than picking silently,' and only drop frame density below default when the user asks for cheaper. Prerequisites are stated ('Needs the writing model and a narration voice engine connected'). This is unusually actionable routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Install Server

Other Tools

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/hermoso-ai/hermoso'

If you have feedback or need assistance with the MCP directory API, please join our Discord server