Skip to main content
Glama

recording_voiceover

Add scripted narration to a completed take: synthesize speech, align segments to timeline events or timestamps, and mux audio over video with optional captions. Starts async; poll recording_status.

Instructions

Attach a scripted narration track to a completed take: the caller supplies the prose, the server synthesizes speech, aligns segments to timeline anchors (or to an explicit at_ms), and muxes the audio over the existing video stream, optionally burning styled captions. Starts an async 'narrating' phase; poll recording_status. GIF cannot carry audio.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
idNoTake id from recording_start; defaults to the latest take.
fitNonatural
voiceNo
engineNoauto
tail_msNo
segmentsYes
offset_msNo
subtitlesNoBurn styled captions from the narration into the video (re-encodes the video). Set false to copy the video without captions.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.1

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the async phase, the required polling via recording_status, and a media-type limitation. It stops short of stating permission requirements, what happens to any existing audio track, or failure/partial-alignment behavior, so it is not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense, front-loaded sentences with no filler; the core action leads and the workflow/limitation follow. Jargon ('muxes', 'narrating') compresses well but borders on under-explaining for a non-expert agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter mutation-style tool with no annotations and no output schema, the description covers the workflow shape and the async contract but leaves several parameters and the return/result behavior unaddressed. Adequate to attempt a correct call, thin on the details needed to call it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25%, so the description has to compensate and only partly does: it clarifies that segments carry prose plus timeline anchors or an explicit at_ms, and that captions can be burned. The enum/general parameters (fit, engine, voice, tail_ms, offset_ms) are left to the schema's bare names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Attach a scripted narration track to a completed take') and then spells out the actual mechanism: synthesize speech, align to anchors, mux over the video stream. This is distinct enough from recording_scenes/timeline that an agent can route to it without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operating context ('completed take', async 'narrating' phase, 'poll recording_status') and one hard exclusion ('GIF cannot carry audio'). It does not name alternatives for cases where narration is not the right call, but the preconditions and follow-up workflow are explicit enough to act on.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.