Generate AI media (image or animation)
generate_mediaGenerate an AI image or canvas-code-based animation directly into a clip.
kind="image": text-to-image. Pass
prompt. Optional:animation_setting(entry/exit — set it HERE, see below),style_id(from find type='image_gen_style_packs'),reference_image_urlormcp_upload_idfor image-to-image grounding.kind="animation": canvas-code animation rendered from a prompt. Pass
prompt. Optional:voiceover_text(drives timing),base_component_id(reuse a saved animation as the starting point),reference_image_urlormcp_upload_idfor visual grounding.
Generation is asynchronous: the element is created immediately with a stable element_id and rendered in the background. Poll get_clip(select:['busy']) — an EMPTY busy means the render has landed. (This previously said to watch the phantom flag; phantom has never been a key get_clip returns, so there was nothing to poll.)
Set presentation up front. animation_setting is applied to the element as it is created, so the image enters correctly the first time it renders. Doing it afterwards with update_elements means writing to the element that is still generating, which is the write most likely to be refused while the generation holds it.
group is NOT accepted here, unlike add_elements: a generated element is built in the background, and the grouping would be overwritten when the render lands. Add it ungrouped, then call update_elements with group once it appears.
Tip: use this tool whenever the user asks for a "generated", "AI", or "create me a" visual. For uploaded photos / logos / icons / GIFs, use add_elements with element_type='image' and a src or mcp_upload_id instead.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | Top-left X in canvas pixels. | |
| y | Yes | Top-left Y in canvas pixels. | |
| kind | Yes | 'image' = AI text-to-image; 'animation' = canvas-code-based motion graphic. | |
| model | No | Image only. Which image model renders it. Default `gemini-3-pro-image-preview` (Nano Banana Pro), a strong general choice — leave it off unless you want one of the specifics below. `gemini-3.1-flash-image-preview` (Nano Banana 2) and `gemini-2.5-flash-image` are faster and cheaper. `gpt-image-2.5-flare` is OpenAI's fast one, `gpt-image-2.5-sunburst` its most precise editor, and `gpt-image-2` is the one to reach for when the image must carry legible text or follow several reference images. `gpt-image-1` / `gpt-image-1.5` are the older pair and only accept square-ish framing — a wide box is snapped to 4:3 rather than honoured, so do not pick them for a banner. An unlisted value is rejected here rather than silently swapped for the default, which is what the backend does with one. | |
| width | Yes | Width in pixels. | |
| height | Yes | Height in pixels. | |
| prompt | Yes | Generation prompt. For animations, be SPECIFIC: name the UI elements, interaction sequence, timing feel, and visual style. Vague prompts produce bad output. | |
| clip_id | Yes | Clip ID to place the generated element into. | |
| end_time | No | Disappear at (seconds). | |
| style_id | No | Image only. Style preset ID from find(type='image_gen_style_packs'). See resource clueso://docs/generation-styles. | |
| background | No | Image only, and only honoured by `gpt-image-2.5-flare`, `gpt-image-2.5-sunburst`, `gpt-image-1` and `gpt-image-1.5` — every other model is always opaque and ignores this. Pass `false` for a cut-out with no background: icons, logos, stickers, anything meant to sit ON the composition rather than behind it. You MUST set `model` to one of those four in the same call; `gpt-image-2.5-flare` is the one to default to, since the older pair cannot frame wide. Asking for transparency in the PROMPT instead does not work — the model paints a grey-and-white checkerboard as real pixels. | |
| project_id | Yes | Project ID. | |
| start_time | No | Appear at (seconds). | |
| aspect_ratio | No | Image only. Overrides the ratio inferred from width/height. Leave it off and the box decides, which is usually what you want — set it when the box is a placeholder and you know the shape you need. Still clamped to what the model accepts: gpt-image-1/1.5 only do 1:1, 4:3 and 3:4; the gpt-image-2.5 pair adds 16:9 and 9:16; gpt-image-2 has no 3:2 or 2:3, so those become 4:3 and 3:4. The response does not report the ratio used. | |
| mcp_upload_id | No | mcp_upload_id from the upload flow. Resolved server-side to a presigned URL before generation. | |
| voiceover_text | No | Animation only. Paces the motion to the spoken script — and as a side effect sets this clip's voiceover text and triggers speech generation for the clip. | |
| animation_setting | No | Image only. Entry/exit animation, same shape as add_elements' type_data.animation_setting. The design guide asks AI images to enter with `masked_reveal` or a slow fade, so set it here rather than following up with update_elements — that follow-up targets the element while it is still generating. Ignored for kind='animation', which animates through its generated code. | |
| base_component_id | No | Animation only. Reuse a saved animation component as the starting point (from find(type='element_components')). To re-skin its tunable parameters, set parameter_values via update_elements after it renders. | |
| reference_image_url | No | Public URL of a reference image. Mutually exclusive with mcp_upload_id. |