Skip to main content
Glama
hermoso-ai

Hermoso

Official

Make video thumbnail

make_thumbnail

Render click-driving video thumbnails for YouTube, Shorts, and Instagram using a full production pipeline: concept frameworks, character casting, scene rendering, and free headline overlay.

Instructions

Render a click-driving YOUTUBE / Shorts / Instagram THUMBNAIL or video cover — the full production pipeline (concept framework → casting → scene → render → surgical tweaks → text), not a bare image prompt. Use this for any "thumbnail", "video cover", "video preview" or MrBeast-style packaging ask INSTEAD of generate_image. About 9 credits per variant; the headline overlay is free.

CONCEPT — every thumbnail must open an INFORMATION GAP (the image raises a question the title answers) while staying truthful to the video. Brainstorm ≥5 concepts across the 16 frameworks before you pick, and feel free to combine two. Frameworks (pass as framework): before_after · social_ui · three_step · screenshot · posed_portrait (the default) · posed_action · specific_day · graphical · landscape · map_aerial · product · adding_text · repetition · size_difference · news_clip · amplified_reality. Call hermoso_capabilities for each one's full 'realize it with' note plus the emotion, overlay-style, font and rim-colour catalogs.

THREE GATES, all BEFORE you render:

  1. WHO IS IN FRAME — never assume and never silently substitute a stranger. If the framework puts a person in frame and no face photo is attached, the tool refuses (nothing rendered, nothing charged) and tells you to ask the user once: themselves (send a face photo → the identity gets locked), a generated person (castGenericPerson:true), or a people-free framework.

  2. TEXT — the default is a CLEAN render with the headline TYPESET OVER THE TOP afterwards (free, always legible, correctly spelled). Just pass headline. Only set bakeText:true if the user explicitly asks for the words painted INTO the image — verified live, that renders the asked-for words correctly but leaks garbled invented text across the rest of the frame. Never infer text intent from the topic or the framework.

  3. HOW MANY — ask once whether they want one thumbnail or a SET (offer 4: the same concept at different emotions and/or camera takes). Default is 1; variants caps at 16.

IDENTITY LOCK is automatic for every attached face photo. emotion is the single biggest CTR lever on a face: shock · hype · fear · confusion · determination · smug · charisma · disgust · awe · rage · laugh (or your own phrase). Finished thumbnail needs a fix? Re-call with tweak + sourceImage for a surgical, pixel-faithful edit (emotion / background / background_color / rim_light) instead of re-rendering — tweaks chain. ALWAYS check the returned postRenderCheck against the image before you present it.

PROMPT LANGUAGE — write every DESCRIPTIVE field in ENGLISH (sceneBrief, keyElements, location, composition, background, topic, each person's describe, and every reference field), translating the user's wording where needed: the image models are trained on English and a non-English scene description renders noticeably worse. Text that gets BAKED OR TYPESET stays verbatim in the user's own language — headline, headlineLines and bakedUiText are never translated.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
fontNoheadline font (default Anton). Alternatives incl. Bebas Neue, Oswald, Archivo Black, Montserrat, Inter, Playfair Display
logoNoa brand logo URL or path to place into the composition
splitNosplit/panel LAYOUT — only when the user asks for one ("split", "before/after", "versus screen"). "X vs Y" as a SCENE stays one unified frame
takesNocamera takes per emotion, 1–4: designed framing / low-angle hero / extreme close-up / wide dutch tilt
topicNothe video's topic — used to pick the hero object when you don't name keyElements
tweakNosurgical pixel-faithful edit of a FINISHED thumbnail — needs sourceImage
logo3dNofirst turn the flat logo into a volumetric 3D render (one extra billed image), then composite that
peopleNopeople described in prose instead of by photo (each still gets the chosen expression)
emotionNothe expression on the face (default 'shock') — a preset id or your own phrase
bakeTextNodefault false. true paints the headline INTO the generation — only on an explicit user ask; it leaks garbled text elsewhere in the frame
emotionsNorender one variant per emotion (variants = emotions × takes, max 16)
headlineNo2–4 word headline. Typeset OVER the finished render by default (free, always legible); newlines split it into stacked lines
locationNoplace, time of day, weather, atmosphere
rimColorNocolored back+hair light — ONLY when the user names one: 'ice-blue' / 'neon-magenta' / 'toxic-lime' / 'amber-gold' / 'pure-white'
variantsNohow many thumbnails to render (default 1, max 16). Each is its own billed render — offer a set of 4 rather than assuming
frameworkNoconcept framework id (default 'posed_portrait'); see the list in this description / hermoso_capabilities
referenceNofields YOU extracted by eye from a reference thumbnail. Extract ALL of: brief (one dense sentence on the concept), subject (pose/action generically, NEVER a specific identity), elements, location, composition, background, split (boolean), split_count, person_count (0-3), emotion (one of the 11 presets or 'other'), emotion_detail (one vivid sentence covering eyes, brows, mouth, head angle). emotion + emotion_detail carry the reference's actual facial performance, which is the single biggest CTR lever on a face; split/split_count reproduce its panel structure. The reference image itself is never sent to the model
backgroundNooverride the default bold saturated colour-field background
faceImagesNoup to 3 face photos (URLs or local paths) — each becomes a locked CHARACTER identity, in order
sceneBriefNowhat the thumbnail depicts — the concept in one dense sentence, rendered exactly
aspectRatioNo'16:9' (YouTube, default) / '9:16' (Shorts) / '4:5' (Instagram) / '4:3' / '1:1'
bakedUiTextNoshort label for a text-carrying framework (a chat bubble, a DAY N badge, a news lower-third, a map callout) — needs frameworkRequested:true
compositionNooverride the default large-foreground-subject composition
keyElementsNosignature props / effects that make it pop — oversized, flying toward camera
sourceImageNothe finished thumbnail URL a `tweak` edits; tweaks chain, so feed each accepted output into the next
overlayStyleNoheadline style: 'beast' (default, white + heavy black stroke) / 'fire' / 'neon-lime' / 'clean-glass' / 'marker'
forceGenerateNorender the 'screenshot' framework anyway (it is normally a real video frame, not a generation)
headlineLinesNoexplicit headline lines (up to 3) — overrides splitting `headline` on newlines
headlinePlaceNowhere the headline sits — never over the face (default 'bottom')
restrainedGradeNotrue for a calm / premium / muted look instead of the default punchy poster grade
castGenericPersonNopass true only after the user has explicitly chosen a generated stranger over their own face
frameworkRequestedNotrue ONLY when the USER named this framework — it is what authorizes a text-carrying framework (social_ui / news_clip / specific_day / map_aerial) to bake its short UI label
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the sparse annotations (readOnlyHint:false, destructiveHint:false), the description discloses rich behavioral traits: refusal behavior when no face photo is attached, credit pricing, free headline overlay, identity lock, bakeText's side effect of garbled text, the default typeset overlay, language translation quirks, tweak chaining, and the requirement to check postRenderCheck. This goes far beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-structured with clear headers (CONCEPT, THREE GATES, IDENTITY LOCK, PROMPT LANGUAGE) and bold keywords. It front-loads the purpose and then explains necessary behavioral nuances. It is appropriately detailed for a tool with 32 parameters, but a few sections could be tightened (e.g., the prompt language paragraph is dense). Overall, it's efficient relative to the complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, no output schema, and the need to guide agents through a multi-step production pipeline, the description covers all essential aspects: concept generation, framework selection, the three gates (identity, text, quantity), identity lock, tweaking, language rules, and the required post-render check. It leaves no significant operational gap for an agent to guess.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema coverage, baseline is 3, but the description substantially enriches parameter meaning. It explains the interplay of framework with hermoso_capabilities, clarifies when castGenericPerson should be true (only after explicit user choice), expands on bakeText, tweak, sourceImage, and reference extraction fields with practical usage context, and grounds each parameter in the production workflow. This adds value well beyond the schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb ('Render') and resource ('YOUTUBE / Shorts / Instagram THUMBNAIL or video cover') and explicitly differentiates itself from generate_image ('INSTEAD of generate_image'). It also conveys the full pipeline nature, making the tool's scope unmistakable even among a huge sibling set.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use this tool ('for any thumbnail, video cover, video preview or MrBeast-style packaging ask') and names the alternative (generate_image). It also gives conditional guidance for related operations (tweak for edits, hermoso_capabilities for framework details) and clarifies the three gate decisions before rendering.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Install Server

Other Tools

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/hermoso-ai/hermoso'

If you have feedback or need assistance with the MCP directory API, please join our Discord server