Skip to main content
Glama

Create a CraftStory 2.0 talking video (photo + audio)

create_craftstory2_video

Generate a talking-avatar video from a photo and audio clips, with lip-sync, gestures, and natural motion; returns a job ID for tracking progress.

Instructions

Start a craftstory-2 generation: a photo of a person (image_url or image_path, or a custom avatar scene via scene_id) speaks the given audio clips with lip-sync, gestures and natural motion; any length. resolution is WIDTH_HEIGHT (480_832 / 720_1280 portrait, 832_480 / 1280_720 landscape); 1080p is available afterwards via upscale_video. Credits are charged on create (see preview_cost) and refunded if the job fails. Returns the job id and initial status; generation takes 8-15 minutes, so call wait_for_job(model='craftstory-2') repeatedly until it reports done, then get_job_result for the video URL.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
nameNoLabel, used as the download file name
faceswapNoIdentity pass on the result (default true; off for custom avatars)
gesturesNoHow much the avatar moves (default normal)
scene_idNoCustom avatar scene id (from list_avatars) used instead of a photo
avatar_idNoCustom avatar id (from list_avatars); its trained model drives identity
image_urlNoPublic URL of the photo (JPG/PNG)
image_pathNoAbsolute local path of the photo to upload (JPG/PNG/HEIC, <= 20 MB); ~/ is expanded
resolutionYes
user_promptNoOptional motion / scene hint
lipsync_modeNocraftstory (default) / sync_so (alternative engine) / empty (no lip-sync)
audio_clip_idsYesAudio clip ids (from create_audio_clip), played in order

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.2

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, and the description adds important non-obvious behavior: credits are charged on create and refunded on failure, generation takes 8-15 minutes, and the agent must poll wait_for_job(model='craftstory-2') then fetch get_job_result. That workflow guidance is valuable and goes beyond the annotations, though it does not cover rate limits or concurrency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense paragraph that front-loads the core purpose, then enumerates key options and the post-create workflow. It is efficient with little filler, though the length and semicolon-heavy structure make it slightly dense rather than crisply front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent, credit-charging, multi-minute generation tool with no output schema, the description covers the essential lifecycle: input choices, resolution semantics, cost behavior, and the polling/result workflow. It is nearly complete, missing only explicit success/failure payload details, which are partly handled by the named get_job_result sibling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is already high at 91%, so the schema documents most parameters (image_url, image_path, scene_id, avatar_id, gestures, etc.). The description adds meaning by explaining that image_url or image_path provide the photo, scene_id is for a custom avatar scene, resolution is WIDTH_HEIGHT with named portrait/landscape values, and audio_clip_ids are played in order - useful framing that supplements the schema rather than repeating it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ('Start a craftstory-2 generation') and immediately specifies the mechanism: a photo of a person speaks audio clips with lip-sync, gestures, and motion. This clearly distinguishes it from the sibling create_minimax_h3_video and create_audio_clip by naming the craftstory-2 model and the photo+audio inputs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context for when to use this tool (photo + audio clips + lip-sync) and points to alternatives for related steps (scene_id/custom avatars via list_avatars, upscale_video for 1080p, preview_cost for cost preview). However, it does not explicitly state when NOT to use it or directly compare against the sibling video generator create_minimax_h3_video.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.