Skip to main content
Glama

Make video thumbnail

make_thumbnail

Render a click-driving YOUTUBE / Shorts / Instagram THUMBNAIL or video cover — the full production pipeline (concept framework → casting → scene → render → surgical tweaks → text), not a bare image prompt. Use this for any "thumbnail", "video cover", "video preview" or MrBeast-style packaging ask INSTEAD of generate_image. About 9 credits per variant; the headline overlay is free.

CONCEPT — every thumbnail must open an INFORMATION GAP (the image raises a question the title answers) while staying truthful to the video. Brainstorm ≥5 concepts across the 16 frameworks before you pick, and feel free to combine two. Frameworks (pass as framework): before_after · social_ui · three_step · screenshot · posed_portrait (the default) · posed_action · specific_day · graphical · landscape · map_aerial · product · adding_text · repetition · size_difference · news_clip · amplified_reality. Call hermoso_capabilities for each one's full 'realize it with' note plus the emotion, overlay-style, font and rim-colour catalogs.

THREE GATES, all BEFORE you render:

  1. WHO IS IN FRAME — never assume and never silently substitute a stranger. If the framework puts a person in frame and no face photo is attached, the tool refuses (nothing rendered, nothing charged) and tells you to ask the user once: themselves (send a face photo → the identity gets locked), a generated person (castGenericPerson:true), or a people-free framework.

  2. TEXT — the default is a CLEAN render with the headline TYPESET OVER THE TOP afterwards (free, always legible, correctly spelled). Just pass headline. Only set bakeText:true if the user explicitly asks for the words painted INTO the image — verified live, that renders the asked-for words correctly but leaks garbled invented text across the rest of the frame. Never infer text intent from the topic or the framework.

  3. HOW MANY — ask once whether they want one thumbnail or a SET (offer 4: the same concept at different emotions and/or camera takes). Default is 1; variants caps at 16.

IDENTITY LOCK is automatic for every attached face photo. emotion is the single biggest CTR lever on a face: shock · hype · fear · confusion · determination · smug · charisma · disgust · awe · rage · laugh (or your own phrase). Finished thumbnail needs a fix? Re-call with tweak + sourceImage for a surgical, pixel-faithful edit (emotion / background / background_color / rim_light) instead of re-rendering — tweaks chain. ALWAYS check the returned postRenderCheck against the image before you present it.

PROMPT LANGUAGE — write every DESCRIPTIVE field in ENGLISH (sceneBrief, keyElements, location, composition, background, topic, each person's describe, and every reference field), translating the user's wording where needed: the image models are trained on English and a non-English scene description renders noticeably worse. Text that gets BAKED OR TYPESET stays verbatim in the user's own language — headline, headlineLines and bakedUiText are never translated.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
fontNoheadline font (default Anton). Alternatives incl. Bebas Neue, Oswald, Archivo Black, Montserrat, Inter, Playfair Display
logoNoa brand logo URL or path to place into the composition
splitNosplit/panel LAYOUT — only when the user asks for one ("split", "before/after", "versus screen"). "X vs Y" as a SCENE stays one unified frame
takesNocamera takes per emotion, 1–4: designed framing / low-angle hero / extreme close-up / wide dutch tilt
topicNothe video's topic — used to pick the hero object when you don't name keyElements
tweakNosurgical pixel-faithful edit of a FINISHED thumbnail — needs sourceImage
logo3dNofirst turn the flat logo into a volumetric 3D render (one extra billed image), then composite that
peopleNopeople described in prose instead of by photo (each still gets the chosen expression)
emotionNothe expression on the face (default 'shock') — a preset id or your own phrase
bakeTextNodefault false. true paints the headline INTO the generation — only on an explicit user ask; it leaks garbled text elsewhere in the frame
emotionsNorender one variant per emotion (variants = emotions × takes, max 16)
headlineNo2–4 word headline. Typeset OVER the finished render by default (free, always legible); newlines split it into stacked lines
locationNoplace, time of day, weather, atmosphere
rimColorNocolored back+hair light — ONLY when the user names one: 'ice-blue' / 'neon-magenta' / 'toxic-lime' / 'amber-gold' / 'pure-white'
variantsNohow many thumbnails to render (default 1, max 16). Each is its own billed render — offer a set of 4 rather than assuming
frameworkNoconcept framework id (default 'posed_portrait'); see the list in this description / hermoso_capabilities
referenceNofields YOU extracted by eye from a reference thumbnail. Extract ALL of: brief (one dense sentence on the concept), subject (pose/action generically, NEVER a specific identity), elements, location, composition, background, split (boolean), split_count, person_count (0-3), emotion (one of the 11 presets or 'other'), emotion_detail (one vivid sentence covering eyes, brows, mouth, head angle). emotion + emotion_detail carry the reference's actual facial performance, which is the single biggest CTR lever on a face; split/split_count reproduce its panel structure. The reference image itself is never sent to the model
backgroundNooverride the default bold saturated colour-field background
faceImagesNoup to 3 face photos (URLs or local paths) — each becomes a locked CHARACTER identity, in order
sceneBriefNowhat the thumbnail depicts — the concept in one dense sentence, rendered exactly
aspectRatioNo'16:9' (YouTube, default) / '9:16' (Shorts) / '4:5' (Instagram) / '4:3' / '1:1'
bakedUiTextNoshort label for a text-carrying framework (a chat bubble, a DAY N badge, a news lower-third, a map callout) — needs frameworkRequested:true
compositionNooverride the default large-foreground-subject composition
keyElementsNosignature props / effects that make it pop — oversized, flying toward camera
sourceImageNothe finished thumbnail URL a `tweak` edits; tweaks chain, so feed each accepted output into the next
overlayStyleNoheadline style: 'beast' (default, white + heavy black stroke) / 'fire' / 'neon-lime' / 'clean-glass' / 'marker'
forceGenerateNorender the 'screenshot' framework anyway (it is normally a real video frame, not a generation)
headlineLinesNoexplicit headline lines (up to 3) — overrides splitting `headline` on newlines
headlinePlaceNowhere the headline sits — never over the face (default 'bottom')
restrainedGradeNotrue for a calm / premium / muted look instead of the default punchy poster grade
castGenericPersonNopass true only after the user has explicitly chosen a generated stranger over their own face
frameworkRequestedNotrue ONLY when the USER named this framework — it is what authorizes a text-carrying framework (social_ui / news_clip / specific_day / map_aerial) to bake its short UI label

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses numerous behavioral traits beyond the sparse annotations: it refuses to render when a person is in frame without a photo, charges ~9 credits per variant, leaks garbled text if bakeText is used, automatically locks identities, supports chained tweaks, and requires checking postRenderCheck. These are rich, non-obvious behaviors that go well beyond what annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-structured with clear headers (CONCEPT, THREE GATES, IDENTITY LOCK, PROMPT LANGUAGE) and front-loads the core purpose. Each section carries actionable guidance; while some redundancy exists (e.g., the framework list is repeated), the complexity of 32 parameters justifies the length. It earns a 4 rather than 5 due to being somewhat verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and complex generation behavior, the description covers the full workflow: gates, credit costs, language requirements, and post-render verification. It mentions postRenderCheck but doesn't detail the response structure, which is a minor gap. Overall it is highly complete for a generation tool with 32 parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the schema covers all parameters with descriptions (100% coverage), the tool description adds substantial contextual meaning. It explains when to use key parameters (e.g., bakeText only on explicit ask, headline for typesetting over the top, variants cap at 16, emotion as the biggest CTR lever) and clarifies how they interact with the pipeline. This goes far beyond schema help text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: rendering click-driving thumbnails/video covers through a full production pipeline. It explicitly distinguishes itself from generate_image, identifying the exact use case and resource. The verb 'render' and specific platforms (YouTube/Shorts/Instagram) make the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance on when to use this tool vs alternatives: 'Use this for any thumbnail... ask INSTEAD of generate_image.' It also directs the agent to call hermoso_capabilities for framework details, and outlines specific gates (e.g., when a face photo is required) that dictate when to pause and ask the user. This provides clear decision rules.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

A3.7/5.0
Disambiguation2/5

With 293 tools, the surface is enormous and many tools have overlapping purposes—multiple posting tools (post_to_meta, post_to_linkedin, schedule_post, etc.), multiple analytics tools per channel, and several search tools (search_meta_ads, search_instagram, search_reddit...). While each description is detailed, the volume makes it difficult for an agent to reliably distinguish between similar tools without careful reading, leading to frequent misselection.

Naming Consistency4/5

The naming is largely consistent with a verb_noun pattern (post_to_*, list_*, create_*, delete_*, update_*, manage_*). There are clear families for major operations. A few outliers like 'google_business_account', 'hermoso_capabilities', and 'store_get' break the pattern, but the overwhelming majority follow a predictable structure, making navigation somewhat easier.

Tool Count1/5

293 tools is far beyond any reasonable scope for a single MCP server, even for a comprehensive marketing platform. The calibration guide flags 50+ as an extreme mismatch, and this is nearly six times that threshold. Such a large surface overwhelms context windows, increases the probability of misselection, and makes it impractical for agents to learn or use effectively.

Completeness4/5

The tool set covers a vast domain: ad creation and rendering, posting across nine+ social channels, analytics and reporting, file management (Drive/OneDrive), competitor research, brand management, and more. It appears to provide CRUD and lifecycle coverage for most resources. While there may be minor gaps given the immense scope, the overall coverage is impressively comprehensive.