generate_digital_human
Create a talking video from a portrait image and audio, syncing lip movement and actions to the voice.
Instructions
数字人生成(OmniHuman 1.5):驱动人物图片口型和动作,配合音频生成说话视频。 使用模型:jimeng_realman_avatar_picture_omni_v15
portrait_url: 人物肖像图片 URL(含人物/动漫/宠物)
audio_url: 驱动音频 URL(WAV/MP3,必须小于 60 秒)
resolution: 输出分辨率,720 或 1080(默认 1080)
prompt: 可选提示词,支持中/英/日/韩语,最长 300 字符
注意:此接口使用 CVSubmitTask/CVGetResult(与普通接口不同)。
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | No | ||
| audio_url | Yes | ||
| resolution | No | ||
| portrait_url | Yes |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| result | Yes |