Skip to main content
Glama

generate_video

Create AI videos from text prompts or input images across 85+ models, then save the finished file to a local folder automatically.

Instructions

Generate a video using kie.ai (85+ models). Downloads to kie/assets/raw/. MODEL GUIDE: Best cinematic→veo-3/text-to-video (50cr/s, audio). Fast+cheap→grok-imagine-video-1-5-preview (1.6-3cr/s, audio, NEW), wan/flash-image-to-video (6-8cr/s measured; alias of wan/2-6-flash). Budget cinematic→hailuo-standard (4cr/s). First→last-frame or anything-from-anything refs→gemini-omni/flash-1-1 (NEW, est. ~63cr per 4s clip). Budget multimodal refs→bytedance/seedance-2-mini (9.5cr/s @480p). 30s single takes→bytedance/seedance-2-5 (NEW). Budget all-rounder w/ audio+templates+extend→pixverse-v6 family (4-9.6cr/s, NEW; I2V is its strength; transition=first/last-frame morph). Multilingual lip-synced dialogue→happyhorse-1-1 T2V/I2V/R2V (NEW). 2K + stereo audio→minimax-h3 (8cr/s @768P, price halved Sept 2026). Per-shot scripted multi-shot→kling-3-omni (14cr/s @720p, NEW; transformation=restyle existing video). Next-gen Wan draft→wan/3-0-video (8cr/s @480P, NEW). Fast Kling→kling/v3-turbo (18cr/s, audio, NEW). Image-to-video→veo-3/image-to-video (include a sound cue like "SFX: room tone" — Veo I2V intermittently fails its audio pass without one; images must be served with their real Content-Type, upload via upload_file), kling/image-to-video. Avatar/talking head→omnihuman-1-5 (premium, NEW), kling/ai-avatar-pro, infinitalk/from-audio. Re-dub existing footage→volcengine/video-to-video-lip-sync (8cr/s, NEW). Motion control→kling/motion-control, wan/animate-move. Extend video→use veo_extend or runway_extend tools. NOTE: Sora 2 family removed (OpenAI API sunset Sept 2026). Use list_models filter="use-case" to explore.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
waitNoSet false to submit and return immediately with the task_id (async mode) — then poll with check_task and fetch with download_result. Recommended for long generations to avoid client-side watchdog timeouts.
modelNoModel ID (e.g. "veo-3/text-to-video", "kling/image-to-video", "wan/3-0-video")veo-3/text-to-video
promptYesVideo description prompt
filenameNoOutput filename (saved to kie/assets/raw/). Auto-generated if omitted.
image_urlsNoInput image URLs for image-to-video models
aspect_ratioNoAspect ratio: 16:9, 9:16, or 1:116:9
download_dirNoAbsolute directory to save the file(s) into (created if missing). Defaults to the server's kie/assets/raw/. Must be absolute — the MCP server's working directory is not the caller's.
model_optionsNoModel-specific options (duration, resolution, mode, etc.)
max_wait_secondsNoOverride the blocking-mode polling budget in seconds (defaults: image 600, video 900, audio 300, speech 300). Ignored when wait=false.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv5.3.0
    • changedInput schema / properties / model / description
      Previous value: -"Model ID (e.g. \"veo-3/text-to-video\", \"sora/text-to-video\", \"kling/image-to-video\")"New value: +"Model ID (e.g. \"veo-3/text-to-video\", \"kling/image-to-video\", \"wan/3-0-video\")"
  2. First observedv4.7.0

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and does well: it discloses per-second credit pricing, which models include audio, a documented failure mode ('Veo I2V intermittently fails its audio pass without one'), and a hard precondition (images must be uploaded via upload_file with real Content-Type). It doesn't cover error handling or what wait=true actually returns, so it stops short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded, but the remainder is a single dense wall of arrow-delimited routing text rather than scannable structure. Most sentences carry model-selection value, though the volume partly overlaps with what list_models exists to provide.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool with nested model_options and no output schema, the description covers the essentials: destination directory, model-selection guidance, upload prerequisite, and failure caveats. It leaves minor gaps around the blocking-mode return value, but nothing that would cause a wrong invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the guide adds real meaning to the `model` parameter that the schema's one-line example cannot (which model fits which use case and cost). It also enriches `image_urls` by mandating upload_file, going beyond the schema's bare 'Input image URLs'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb and resource ('Generate a video'), the backend (kie.ai, 85+ models), and the output location (kie/assets/raw/). It also explicitly routes adjacent work elsewhere ('Extend video→use veo_extend or runway_extend tools'), so an agent can separate it from the extend/upscale siblings without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The MODEL GUIDE is effectively a when-to-use decision tree: cinematic, budget, fast, lip-sync, avatar, motion control, multi-shot, and image-to-video each map to named models. It names alternatives and their selection conditions (e.g. 'Extend video→use veo_extend'), and points to list_models for exploration.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.