Skip to main content
Glama
hermoso-ai

Hermoso

Official

Product sizzle (music-led)

product_sizzle

Turn a product packshot into a faceless 18–30s music-led video ad: hero clip diced into fast cuts, intercut with spec and CTA cards. Confirm the credit cost before rendering.

Instructions

Render an 18-30s music-led PRODUCT SIZZLE: ONE 15s Seedance 2.0 hero clip of the product, diced into fast cuts and intercut with typeset spec/CTA cards on a brand-coloured grain background, mixed to a music bed. Faceless by design — no people, no voiceover, no spoken lines; the cards carry every word, so nothing is left to a video model's spelling. Pass a real packshot as refImage or the label will not be yours. EXPENSIVE — the hero clip is the only paid leg and it is a full 15s Seedance render: ≈1,040 credits at the DEFAULT 1080p, ≈470 at 720p, ≈220 at 480p, ≈4,130 at 4k (call hermoso_capabilities for the live seedance-2 per-duration numbers; the dicing and the cards are free, and the music bed is already included in the quoted figure). Confirm the spend with the user before calling. For a talking/UGC ad use render_ad or generate_avatar; for a cheap deterministic format use make_template_ad.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
ctaNoclosing CTA line, ≤30 chars
specsNoup to 4 spec lines for the typeset cards, ≤26 chars each
promptYeswhat the sizzle should show — the product, the setting, the look
secondsNofinished length, clamped to 18-30s (default 25). The PAID hero render is always 15s regardless — this only changes how the cuts and cards are packed
refImageNoproduct packshot URL that anchors the real label — strongly recommended
brandNameNobrand name on the cards — defaults to the workspace brand
musicMoodNomusic-bed mood, e.g. driving / cinematic / upbeat
resolutionNohero-clip resolution and therefore the whole cost — DEFAULT '1080p' (≈1,040 credits); '720p' ≈470, '480p' ≈220, '4k' ≈4,130
aspectRatioNo'9:16' default; anything the seedance-2 catalog entry does not list falls back to 9:16

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.1.161

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are sparse (only generic false hints), so the description carries the burden. It discloses the cost in credits per resolution, that the paid hero clip is always 15s regardless of total length, that dicing/cards are free and the music bed is included, and that the user must approve the spend. It also explains the faceless design as a deliberate way to avoid video-model spelling errors. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Long but dense: every sentence carries operational signal, including length, composition, no-people constraint, refImage prerequisite, credit costs, approval requirement, and sibling alternatives. It repeats some schema cost details, but that repetition aids quick decision-making, and the structure is front-loaded and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity, costly 9-parameter tool with no output schema, the description leaves little ambiguity about what is produced, at what cost, under which constraints, and when not to use it. It even points to hermoso_capabilities for live Seedance pricing numbers, closing the remaining factual gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema descriptions already carry cost, resolution, and seconds meanings, so the baseline is 3. The description adds consequential guidance beyond the schema, especially 'Pass a real packshot as refImage or the label will not be yours' and the clarification that seconds only repacks the cuts while the paid leg stays 15s, so it earns a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb and resource: 'Render an 18-30s music-led PRODUCT SIZZLE', then details the exact composition (15s Seedance 2.0 hero clip, diced fast cuts, typeset spec/CTA cards, music bed). It explicitly contrasts with render_ad, generate_avatar, and make_template_ad, so an agent can distinguish it from siblings without inspecting schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use guidance: product sizzles that are faceless, music-led, and text-card-driven. It names alternatives and the conditions that select them ('For a talking/UGC ad use render_ad or generate_avatar; for a cheap deterministic format use make_template_ad'), plus the prerequisite of a real packshot and the requirement to confirm spend with the user before calling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools