Skip to main content
Glama

generate_video_audio

Generate or replace audio for a video media item. Supports add/replace audio modes and prompt-driven video-to-audio models.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
promptNoPrompt describing the desired audio, music, foley, or ambience
mediaIdYesVideo MediaDocV2 ID to process
modelIdYesVideo-to-audio model ID
videoUrlNoOverride source video URL; otherwise resolves from active media version
audioModeNoWhether generated audio replaces or adds to original audioreplace
extraParamsNoAdditional raw V3 parameters
workspaceIdYesThe workspace ID

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does not disclose whether this is an asynchronous job, how long it takes, whether it costs credits, what happens to the original audio track in "add" vs "replace" mode, or whether the operation is reversible — all material for a generative mutation on existing media.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, with the core action front-loaded before the supporting capability statement. Slightly more could be said about mode semantics, but nothing present is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The input schema is rich and fully documented, which covers parameter-level needs. However, with no output schema and no annotations, the description omits the post-invocation picture (job handle, polling, resulting media version) that an agent needs to chain this generation call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all seven parameters are already documented, including the audioMode enum and the videoUrl override. The description only echoes the mode concept and the model-driven nature of the call, adding no format, default, or interaction detail beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb pair (generate or replace) and a specific resource (audio for a video media item), which is far more precise than the generic siblings. It does not, however, name or distinguish itself from close siblings such as generate_audio, generate_audio_track, or add_audio_track, so an agent must still infer which tool applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Supports add/replace audio modes and prompt-driven video-to-audio models" implies the situations the tool covers, and the video-to-audio framing hints that a video media item is the precondition. There is no explicit when-to-use/when-not statement and no routing guidance against generate_audio, add_audio_track, or lip_sync_video_media, which are all plausible alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources