Skip to main content
Glama
hermoso-ai

Hermoso

Official

Make video thumbnail

make_thumbnail

Render click-driving YouTube, Shorts, or Instagram thumbnails via a full concept, casting, render and text pipeline instead of a bare image prompt.

Instructions

Render a click-driving YOUTUBE / Shorts / Instagram THUMBNAIL or video cover through the full production pipeline (concept, casting, scene, render, tweaks, text), not a bare image prompt. Use it for any "thumbnail", "video cover" or MrBeast-style packaging ask INSTEAD of generate_image. About 9 credits per variant; the headline overlay is free.

CONCEPT — open an INFORMATION GAP (the image raises a question the title answers) while staying truthful to the video. Brainstorm ≥5 concepts across the 16 frameworks (ids on framework; combining two is fine) before you pick; hermoso_capabilities has each one's 'realize it with' note and the emotion, overlay, font and rim-colour catalogs.

THREE GATES, all BEFORE you render:

  1. WHO IS IN FRAME — never assume or silently substitute a stranger. A framework with a person and no face photo is refused (nothing charged): ask the user once — themselves (a face photo, identity-locked), a generated person (castGenericPerson:true), or a people-free framework.

  2. TEXT — default is a CLEAN render with the headline TYPESET over it (free, legible, correctly spelled): pass headline. bakeText:true only on an explicit ask for words painted INTO the image. Never infer text intent from the topic.

  3. HOW MANY — ask once: one, or a SET (offer 4: one concept at different emotions / camera takes). Default 1; variants caps at 16.

emotion is the biggest CTR lever on a face (identity lock is automatic for every face photo). To fix a finished one, re-call with tweak + sourceImage for a surgical edit (emotion / background / background_color / rim_light) — tweaks chain. ALWAYS check the returned postRenderCheck against the image before presenting it.

PROMPT LANGUAGE — write every DESCRIPTIVE field (sceneBrief, keyElements, location, composition, background, topic, each person's describe, every reference) in ENGLISH, translating the user's words: the models render English better. headline, headlineLines and bakedUiText stay verbatim in the user's language.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
fontNoheadline font: Anton (default) or any Google Fonts family
logoNoa brand logo URL or path to place into the composition
splitNosplit/panel LAYOUT — only when the user asks for one ("split", "before/after", "versus screen"). "X vs Y" as a SCENE stays one unified frame
takesNocamera takes per emotion, 1–4: designed framing / low-angle hero / extreme close-up / wide dutch tilt
topicNothe video's topic — used to pick the hero object when you don't name keyElements
tweakNosurgical pixel-faithful edit of a FINISHED thumbnail (needs sourceImage): kind emotion / background / background_color / rim_light, or any other kind with the edit in words as value
logo3dNofirst turn the flat logo into a volumetric 3D render (one extra billed image), then composite that
peopleNopeople described in prose instead of by photo (each still gets the chosen expression)
emotionNothe expression on the face (default 'shock') — shock · hype · fear · confusion · determination · smug · charisma · disgust · awe · rage · laugh, or your own phrase
bakeTextNodefault false. true paints the headline INTO the generation — only on an explicit user ask; it leaks garbled text elsewhere in the frame
emotionsNorender one variant per emotion (variants = emotions × takes, max 16)
headlineNo2–4 word headline. Typeset OVER the finished render by default (free, always legible); newlines split it into stacked lines
locationNoplace, time of day, weather, atmosphere
rimColorNocolored back+hair light — ONLY when the user names one: 'ice-blue' / 'neon-magenta' / 'toxic-lime' / 'amber-gold' / 'pure-white'
variantsNohow many thumbnails to render (default 1, max 16). Each is its own billed render — offer a set of 4 rather than assuming
frameworkNoconcept framework id (default 'posed_portrait') — before_after · social_ui · three_step · screenshot · posed_portrait · posed_action · specific_day · graphical · landscape · map_aerial · product · adding_text · repetition · size_difference · news_clip · amplified_reality — or your own concept in words
referenceNofields YOU extracted by eye from a reference thumbnail. Extract ALL of: brief (one dense sentence on the concept), subject (pose/action generically, NEVER a specific identity), elements, location, composition, background, split (boolean), split_count, person_count (0-3), emotion (one of the 11 presets or 'other'), emotion_detail (one vivid sentence covering eyes, brows, mouth, head angle). emotion + emotion_detail carry the reference's actual facial performance, which is the single biggest CTR lever on a face; split/split_count reproduce its panel structure. The reference image itself is never sent to the model
backgroundNooverride the default bold saturated colour-field background
faceImagesNoup to 3 face photos (URLs or local paths) — each becomes a locked CHARACTER identity, in order
sceneBriefNowhat the thumbnail depicts — the concept in one dense sentence, rendered exactly
aspectRatioNo'16:9' (YouTube, default) / '9:16' (Shorts) / '4:5' (Instagram) / '4:3' / '1:1'
bakedUiTextNoshort label for a text-carrying framework (a chat bubble, a DAY N badge, a news lower-third, a map callout) — needs frameworkRequested:true
compositionNooverride the default large-foreground-subject composition
keyElementsNosignature props / effects that make it pop — oversized, flying toward camera
sourceImageNothe finished thumbnail URL a `tweak` edits; tweaks chain, so feed each accepted output into the next
overlayStyleNoheadline style: beast (default), fire, neon-lime, clean-glass, marker, or your own CSS declarations
forceGenerateNorender the 'screenshot' framework anyway (it is normally a real video frame, not a generation)
headlineLinesNoexplicit headline lines (up to 3) — overrides splitting `headline` on newlines
headlinePlaceNobottom (default), top, center, or a 0-1 fraction from the top; never over the face
restrainedGradeNotrue for a calm / premium / muted look instead of the default punchy poster grade
castGenericPersonNopass true only after the user has explicitly chosen a generated stranger over their own face
frameworkRequestedNotrue ONLY when the USER named this framework — it is what authorizes a text-carrying framework (social_ui / news_clip / specific_day / map_aerial) to bake its short UI label

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changedv0.1.320
    • changedInput schema / properties / emotion / description
      Previous value: -"the expression on the face (default 'shock') — a preset id or your own phrase"New value: +"the expression on the face (default 'shock') — shock · hype · fear · confusion · determination · smug · charisma · disgust · awe · rage · laugh, or your own phrase"
    • changedInput schema / properties / framework / description
      Previous value: -"concept framework id (default 'posed_portrait'), or your own concept in words"New value: +"concept framework id (default 'posed_portrait') — before_after · social_ui · three_step · screenshot · posed_portrait · posed_action · specific_day · graphical · landscape · map_aerial · product · adding_text · repetition · size_difference · news_clip · amplified_reality — or your own concept in words"
  2. Changed7 schema fields changedv0.1.285
    • changedInput schema / properties / font / description
      Previous value: -"headline font (default Anton). Alternatives incl. Bebas Neue, Oswald, Archivo Black, Montserrat, Inter, Playfair Display"New value: +"headline font: Anton (default) or any Google Fonts family"
    • changedInput schema / properties / framework / description
      Previous value: -"concept framework id (default 'posed_portrait'); see the list in this description / hermoso_capabilities"New value: +"concept framework id (default 'posed_portrait'), or your own concept in words"
    • changedInput schema / properties / headlinePlace / description
      Previous value: -"where the headline sits — never over the face (default 'bottom')"New value: +"bottom (default), top, center, or a 0-1 fraction from the top; never over the face"
    • removedInput schema / properties / headlinePlace / enum
      Removed value: -[
      -  "bottom",
      -  "top",
      -  "center"
      -]
    • changedInput schema / properties / overlayStyle / description
      Previous value: -"headline style: 'beast' (default, white + heavy black stroke) / 'fire' / 'neon-lime' / 'clean-glass' / 'marker'"New value: +"headline style: beast (default), fire, neon-lime, clean-glass, marker, or your own CSS declarations"
    • changedInput schema / properties / tweak / description
      Previous value: -"surgical pixel-faithful edit of a FINISHED thumbnail — needs sourceImage"New value: +"surgical pixel-faithful edit of a FINISHED thumbnail (needs sourceImage): kind emotion / background / background_color / rim_light, or any other kind with the edit in words as value"
    • removedInput schema / properties / tweak / properties / kind / enum
      Removed value: -[
      -  "emotion",
      -  "background",
      -  "background_color",
      -  "rim_light"
      -]
  3. Addedv0.1.161

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the generic profile (not read-only, open-world, not idempotent, non-destructive). The description goes well beyond: ~9 credits per variant with the headline overlay free, a people/framework mismatch is refused with nothing charged, postRenderCheck must always be inspected before presenting, identity lock is automatic for face photos, and castGenericPerson requires explicit user consent. This is exactly the kind of cost/auth/refusal context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and sibling differentiation are front-loaded in the first two sentences, then organized under labelled blocks (CONCEPT, THREE GATES, PROMPT LANGUAGE). The sentence count is high, but for a 32-parameter tool most lines carry gating or policy information that would otherwise be lost; only a touch of redundancy (e.g. restating the free headline twice) keeps it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 32 parameters, nested objects, and no output schema, the description covers everything decision-critical: the pre-render gates, the fix-vs-regenerate fork, credit costs, refusal behavior, and the required postRenderCheck verification on the returned result. An agent could drive this tool end-to-end without additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, and the description clearly exceeds it: it flags `emotion` as the biggest CTR lever, explains `variants` cap (16) and the emotion×takes multiplication, constrains `bakeText` to explicit asks only, and gives the prompt-language rule (English for descriptive fields, verbatim for headline/bakedUiText/headlineLines). It stops short of per-parameter coverage for the more mechanical fields (aspectRatio, overlayStyle, headlinePlace), which the schema already handles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (render a thumbnail/video cover) and scopes it precisely: 'through the full production pipeline (concept, casting, scene, render, tweaks, text), not a bare image prompt.' It explicitly names the sibling it is not (generate_image) and the platforms it serves, so an agent can distinguish it without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit routing rule: use this INSTEAD of generate_image for any thumbnail/cover/MrBeast-style packaging ask. It also gives hard preconditions (THREE GATES: who is in frame, text intent, how many variants) and a separate when-to-use path for `tweak` + `sourceImage` on a finished image, including that tweaks chain. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools