Skip to main content
Glama
vparsatwar-git

higgsfield-mcp-unified

generate_speech_video_tool

Create a talking-head video by combining a face image with WAV audio. Provide image_url and audio_url to generate speech-driven video.

Instructions

Talking-head video from a face image + WAV audio (official backend).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
promptNo
audio_urlYes
image_urlYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
backendYes
model_idYes
job_handleYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it discloses almost nothing: it does not say whether generation is async (siblings like get_status_tool and cancel_job_tool imply it is), whether it costs credits, what the output artifact is, or what constraints apply to inputs. '(official backend)' is the only behavioral hint and it is unexplained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, which is structurally clean. However, at this length for a 3-parameter generation tool it reads as under-specified rather than efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. But with zero annotations, zero schema descriptions, and an undocumented prompt parameter, the definition leaves an agent without the async/job, cost, or input-format context needed to invoke it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It loosely maps 'face image' to image_url and 'WAV audio' to audio_url, but says nothing about the third parameter, prompt, or about URL formats, accepted image types, or audio length limits.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: generating a talking-head video from a face image plus WAV audio. That clearly separates it from generate_video_tool and generate_image_tool. The only gap is the unexplained '(official backend)' qualifier, which hints at an alternative backend without naming it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to choose this tool over generate_video_tool or another speech-video path, and no prerequisites or exclusions. The required image+audio pairing is inferable only from the schema, not the description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.