Make video thumbnail
make_thumbnailRender click-driving YouTube, Shorts, or Instagram thumbnails via a full concept, casting, render and text pipeline instead of a bare image prompt.
Instructions
Render a click-driving YOUTUBE / Shorts / Instagram THUMBNAIL or video cover through the full production pipeline (concept, casting, scene, render, tweaks, text), not a bare image prompt. Use it for any "thumbnail", "video cover" or MrBeast-style packaging ask INSTEAD of generate_image. About 9 credits per variant; the headline overlay is free.
CONCEPT — open an INFORMATION GAP (the image raises a question the title answers) while staying truthful to the video. Brainstorm ≥5 concepts across the 16 frameworks (ids on framework; combining two is fine) before you pick; hermoso_capabilities has each one's 'realize it with' note and the emotion, overlay, font and rim-colour catalogs.
THREE GATES, all BEFORE you render:
WHO IS IN FRAME — never assume or silently substitute a stranger. A framework with a person and no face photo is refused (nothing charged): ask the user once — themselves (a face photo, identity-locked), a generated person (
castGenericPerson:true), or a people-free framework.TEXT — default is a CLEAN render with the headline TYPESET over it (free, legible, correctly spelled): pass
headline.bakeText:trueonly on an explicit ask for words painted INTO the image. Never infer text intent from the topic.HOW MANY — ask once: one, or a SET (offer 4: one concept at different emotions / camera takes). Default 1;
variantscaps at 16.
emotion is the biggest CTR lever on a face (identity lock is automatic for every face photo). To fix a finished one, re-call with tweak + sourceImage for a surgical edit (emotion / background / background_color / rim_light) — tweaks chain. ALWAYS check the returned postRenderCheck against the image before presenting it.
PROMPT LANGUAGE — write every DESCRIPTIVE field (sceneBrief, keyElements, location, composition, background, topic, each person's describe, every reference) in ENGLISH, translating the user's words: the models render English better. headline, headlineLines and bakedUiText stay verbatim in the user's language.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| font | No | headline font: Anton (default) or any Google Fonts family | |
| logo | No | a brand logo URL or path to place into the composition | |
| split | No | split/panel LAYOUT — only when the user asks for one ("split", "before/after", "versus screen"). "X vs Y" as a SCENE stays one unified frame | |
| takes | No | camera takes per emotion, 1–4: designed framing / low-angle hero / extreme close-up / wide dutch tilt | |
| topic | No | the video's topic — used to pick the hero object when you don't name keyElements | |
| tweak | No | surgical pixel-faithful edit of a FINISHED thumbnail (needs sourceImage): kind emotion / background / background_color / rim_light, or any other kind with the edit in words as value | |
| logo3d | No | first turn the flat logo into a volumetric 3D render (one extra billed image), then composite that | |
| people | No | people described in prose instead of by photo (each still gets the chosen expression) | |
| emotion | No | the expression on the face (default 'shock') — shock · hype · fear · confusion · determination · smug · charisma · disgust · awe · rage · laugh, or your own phrase | |
| bakeText | No | default false. true paints the headline INTO the generation — only on an explicit user ask; it leaks garbled text elsewhere in the frame | |
| emotions | No | render one variant per emotion (variants = emotions × takes, max 16) | |
| headline | No | 2–4 word headline. Typeset OVER the finished render by default (free, always legible); newlines split it into stacked lines | |
| location | No | place, time of day, weather, atmosphere | |
| rimColor | No | colored back+hair light — ONLY when the user names one: 'ice-blue' / 'neon-magenta' / 'toxic-lime' / 'amber-gold' / 'pure-white' | |
| variants | No | how many thumbnails to render (default 1, max 16). Each is its own billed render — offer a set of 4 rather than assuming | |
| framework | No | concept framework id (default 'posed_portrait') — before_after · social_ui · three_step · screenshot · posed_portrait · posed_action · specific_day · graphical · landscape · map_aerial · product · adding_text · repetition · size_difference · news_clip · amplified_reality — or your own concept in words | |
| reference | No | fields YOU extracted by eye from a reference thumbnail. Extract ALL of: brief (one dense sentence on the concept), subject (pose/action generically, NEVER a specific identity), elements, location, composition, background, split (boolean), split_count, person_count (0-3), emotion (one of the 11 presets or 'other'), emotion_detail (one vivid sentence covering eyes, brows, mouth, head angle). emotion + emotion_detail carry the reference's actual facial performance, which is the single biggest CTR lever on a face; split/split_count reproduce its panel structure. The reference image itself is never sent to the model | |
| background | No | override the default bold saturated colour-field background | |
| faceImages | No | up to 3 face photos (URLs or local paths) — each becomes a locked CHARACTER identity, in order | |
| sceneBrief | No | what the thumbnail depicts — the concept in one dense sentence, rendered exactly | |
| aspectRatio | No | '16:9' (YouTube, default) / '9:16' (Shorts) / '4:5' (Instagram) / '4:3' / '1:1' | |
| bakedUiText | No | short label for a text-carrying framework (a chat bubble, a DAY N badge, a news lower-third, a map callout) — needs frameworkRequested:true | |
| composition | No | override the default large-foreground-subject composition | |
| keyElements | No | signature props / effects that make it pop — oversized, flying toward camera | |
| sourceImage | No | the finished thumbnail URL a `tweak` edits; tweaks chain, so feed each accepted output into the next | |
| overlayStyle | No | headline style: beast (default), fire, neon-lime, clean-glass, marker, or your own CSS declarations | |
| forceGenerate | No | render the 'screenshot' framework anyway (it is normally a real video frame, not a generation) | |
| headlineLines | No | explicit headline lines (up to 3) — overrides splitting `headline` on newlines | |
| headlinePlace | No | bottom (default), top, center, or a 0-1 fraction from the top; never over the face | |
| restrainedGrade | No | true for a calm / premium / muted look instead of the default punchy poster grade | |
| castGenericPerson | No | pass true only after the user has explicitly chosen a generated stranger over their own face | |
| frameworkRequested | No | true ONLY when the USER named this framework — it is what authorizes a text-carrying framework (social_ui / news_clip / specific_day / map_aerial) to bake its short UI label |