Skip to main content
Glama

Generate avatar

generate_avatar

Create a talking-head avatar video from an actor and uploaded audio. Provide actorEntityId and audioFileId; set avatarQuality if needed, and expect a few minutes.

Instructions

Generate a talking-head avatar video from an ACTOR entity and an uploaded audio file. Pass actorEntityId and optionally set avatarQuality. Typically takes a few minutes (longer for longer audio). Tell the user that wait up front.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
audioFileIdYesAudio file id for the avatar to lip-sync.
actorEntityIdYesId of a built-in stock actor or an ACTOR entity (vg_enti_...) with an image reference.
avatarQualityNoAvatar generation quality tier. Applies when actorEntityId is provided. Omit to use workspace settings.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
resultsNoGenerated results with download URLs when succeeded.
toolTypeNoTool name (e.g. GENERATE_IMAGE).
attemptIndexNoCurrent or latest attempt index.
toolExecutionIdYesTool execution id (vg_tool_...).
progressPercentageNoCompletion progress 0-100 (present after polling).

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv2.2.1

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the safety profile (readOnly=false, destructive=false, openWorld=false). The description adds genuinely new behavioral context: generation takes 'a few minutes (longer for longer audio)' and the agent should tell the user about the wait up front. That latency/UX disclosure is exactly what annotations do not cover.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action, then optional parameter, then the latency caveat and user-facing instruction. Nothing is wasted, though the final instruction sentence is more UX coaching than tool semantics.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and annotations carry the safety profile. The latency expectation is supplied, leaving sibling differentiation as the only meaningful gap. Adequate for a 3-parameter async generation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents actorEntityId, audioFileId and the avatarQuality enum tiers. The description only restates that actorEntityId is passed and avatarQuality is optional, adding no format or constraint detail beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Generate a talking-head avatar video from an ACTOR entity and an uploaded audio file.' That is clearly distinguishable from generic siblings like generate_video_clip or prompt_to_video_clip, though it never names an alternative to route against.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It notes the two required inputs implicitly but gives no when-to-use or when-not-to-use guidance, and does not distinguish itself from the other video-generation siblings such as generate_video_clip or script_to_video. The only routing signal is the input type (actor + audio).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.