Generate video (Grok Imagine, Wan 2.7, Hailuo 02, Seedance, Kling 2.6, VEO 3.1, Happy Horse)
aetherwave_generate_videoGenerates a short-form video from a text prompt (T2V) or a text prompt + starting image (I2V). Submits, polls, and returns the final video URL. Default model is 'grok-imagine-t2v' (fast, 4-6 cr/s, with built-in KIE -> fal.ai fallback). Use list_video_models for the full lineup with credit cost per second. I2V models (e.g. 'grok-imagine-i2v', 'seedance-pro-i2v') require a public imageUrl. Video generation can take 30s to several minutes; this tool polls with up to an 8-minute budget.
Model selection guide for videos (when the user does not specify a model)
Default: grok-imagine-t2v (4-6 cr/s, fast, has KIE -> fal.ai fallback for redundancy. Best general-purpose).
Pick a different model when the prompt has these signals:
"highest quality" / "premium" / broadcast / commercial ->
veo3.1-qualityorveo3-quality(Google's flagship, fixed 350-560 cr for 8s, 3-5 min)"fast premium" / quick high-quality ->
veo3-fastorveo3.1-fast(84 cr fixed for 8s)Cinematic camera moves / dolly / pan ->
seedance-pro-t2v(3-10 cr/s) orkling-3.0-pro-t2v(26 cr/s)Realistic human motion / faces ->
hailuo-2.3-pro-i2v(I2V, supply imageUrl)Talking head / lip sync ->
kling-avatar-pro(23 cr/s) orinfinitalk(5-17 cr/s)Anime / stylized / fantasy ->
wan-2.7-t2vNSFW / adult ->
wan-22-nsfw-i2v(I2V only; auto-tags adult)Animate this exact image -> any I2V variant (
grok-imagine-i2v,seedance-pro-i2v,hailuo-2.3-pro-i2v)First + last frame interpolation ->
seedance-pro-i2vwith bothimageUrl+endImageUrlCheapest test ->
hailuo-2.0-standard@ 512p (3 cr/s, ~18 cr for 6s) orgrok-imagine-t2v@ 480p (4 cr/s, ~24 cr for 6s)Clip 12-15s ->
grok-imagine-t2v(accepts up to 15s)True 4K ->
kling-3.0-4k-t2v(94 cr/s, expensive but native 4K)
Audio in generated video: grok-imagine-t2v, seedance-pro-t2v, and the VEO 3.x family include audio at base cost (no surcharge). Kling 2.6 and Kling 3.0 are the outliers — they price audio as a +50-100% surcharge (Kling 2.6 doubles the cost, Kling 3.0 Pro adds ~46%). Default to Grok / Seedance / VEO when sound matters and you don't want to think about audio pricing.
Cost framing: resolution and duration drive cost more than model choice. A 6-second 480p Grok generation costs ~24 cr; the same prompt at 1080p Seedance 2 is ~858 cr (35x more). Pick the lowest acceptable resolution + duration first.
For I2V models: imageUrl is required. For first+last-frame models, pass endImageUrl too.
Consistent characters
If the user names a person they already have ("my Amy character"), call aetherwave_list_characters FIRST, append that character's identityBlock verbatim to the prompt, and pass their referenceImages. Do NOT use imageUrl for this - a starting frame switches the engine to first-frame mode and drops the reference images, which is the opposite of what a consistent character needs.
Writing dialogue
duration is authoritative: the engine honours the seconds you ask for to within ~0.1s and then fits the line to that length by changing pace, rather than finishing early. So write the line to the clip length - do NOT estimate the clip length from the line. Measured on a 22-clip production: ~2 words per second of clip. A words-per-minute figure from a voice description describes the character, not the engine, and budgeting by it overran by 35%.
Pass generateAudio: true for any clip with dialogue, and spell hard words the way they should be spoken ("super intelligence" rather than "superintelligence") - the engine reads the text literally and mangles unfamiliar compounds.
⚠️ A take can come back saying the WRONG WORDS and still report success. Re-rendering the identical prompt produced one clean take and one that dropped a whole sentence. For anything that will be assembled unattended, verify the rendered speech against the script before using the clip.
Ask the user only when:
Single generation would cost more than 100 credits and they haven't confirmed
They asked for "the best" with no other signal; surface 2-3 options with cost ranges
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Moderation mode for Grok Imagine. Defaults to 'normal'. | |
| async | No | Submit and return a taskId IMMEDIATELY instead of waiting for the render. Video takes 1-8 minutes and most MCP clients abandon a call at 60s, so a synchronous video call usually fails from the client side even though the render succeeds. Pass true, then poll aetherwave_get_job(taskId). Strongly recommended for video. | |
| model | No | Model ID. Defaults to 'grok-imagine-t2v'. Use list_video_models for the full list. | |
| prompt | Yes | Text description of the video scene. | |
| duration | No | Duration in seconds. Grok Imagine accepts 6-15; other models have their own ranges (see list_video_models). | |
| imageUrl | No | Public URL of starting image. Required for I2V models. | |
| resolution | No | Output resolution. Default depends on model. | |
| aspectRatio | No | Aspect ratio (e.g. '16:9', '9:16', '1:1'). | |
| endImageUrl | No | Public URL of ending image. Supported by some I2V models (first+last frame). | |
| generateAudio | No | Render native speech and sound WITH the video. Off by default, so a clip is SILENT unless you pass true. Free on Seedance 2.x at every resolution (measured: the sounded and silent runs bill identically) and included at base cost on Grok Imagine and VEO 3.x; Kling 2.6/3.0 surcharge for it. Pass true whenever the prompt contains dialogue - a talking head with no audio is not what the caller asked for. | |
| referenceImages | No | Up to 9 image URLs used as identity anchors held consistent ACROSS the whole clip. THIS is the parameter for a consistent character - pass the character's referenceImages from aetherwave_list_characters. Mutually exclusive with imageUrl: a supplied first frame switches the engine to first-frame mode and DROPS these, disabling the only no-drift mechanism there is. Platform https URLs are fetched and encoded for you. |