Skip to main content
Glama

Set voiceover scripts

voiceover_batch
Destructive

Set voiceover text and/or trigger speech generation for any number of clips in a single call.

Each entry chooses its own action:

  • "set_text" — set transcript for a clip (clip_index + text)

  • "generate_speech" — async TTS for a clip (returns immediately)

  • "set_and_generate" — set text and kick off TTS in one entry (+ text)

Address a clip the same way as everywhere else: clip_index. Pass clip_id instead if you already have it — it survives clips being inserted or reordered mid-build — but you never need both; whichever you omit is looked up once for the whole call.

Entries within one call are applied in order. Returns one result object per input entry. All text-set actions land in ONE save; the TTS for generate/set_and_generate runs async per clip after.

IMPORTANT — generating speech RESCALES the whole clip, it does not clamp it: when audio is generated (generate_speech / set_and_generate), the clip's duration is reset to the spoken audio length, and then EVERY element on that clip is retimed proportionally by (new duration ÷ old duration). start_time, end_time and every keyframe timestamp are multiplied by that factor. Nothing is merely truncated — on a 6s clip that becomes 1.02s, an animation you placed at [0, 1.6] ends up at [0, 0.27]. Zoom elements whose window falls under the minimum after scaling are DROPPED entirely. Generation is async, so this lands AFTER this call has already returned success. So: generate speech BEFORE placing time-sensitive elements, or size them against estimate_duration first — then re-read the clip and check what your elements actually became, not just the clip duration.

Concurrency: parallel-safe (conflict domain: a clip's voiceover). The server merges each clip's voiceover under a per-guide lock and preserves that clip's elements, so you can fan voiceover work out across subagents by clip — and it's safe to run alongside element edits. Two concurrent edits to the SAME clip's voiceover do not last-write-win — both claim that clip's voiceover path, so the later one is REJECTED and nothing is written; re-read and re-apply. Do NOT run concurrently with whole-clip/whole-project mutations on the same guide (update_clips on that clip, structural clip ops, add_audio, update_project).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
entriesYesVoiceover entries — at least one.
project_idYesProject ID

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed4 schema fields changed
    • removedInput schema / properties / context
      Removed value: -{
      -  "description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
      -  "type": "string"
      -}
    • removedInput schema / properties / conversation_id
      Removed value: -{
      -  "description": "Echo the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it.",
      -  "type": "string"
      -}
    • removedInput schema / properties / llm_model
      Removed value: -{
      -  "description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
      -  "type": "string"
      -}
    • changedInput schema / required
      Previous value: -[
      -  "project_id",
      -  "entries",
      -  "context",
      -  "llm_model"
      -]New value: +[
      +  "project_id",
      +  "entries"
      +]
  2. Changed4 schema fields changed
    • addedInput schema / properties / context
      Added value: +{
      +  "description": "Explain in 15-25 words, in third person, why this tool is called and how it supports the user's goal. For analytics only. You MUST describe only the abstract purpose of the tool call. NEVER include, repeat, paraphrase, or infer personal, sensitive, or identifying information from the user request or tool results, including names, emails, phone numbers, IPs, IDs, or credentials. You MUST generalize specific entities into roles such as \"a user\", \"the customer\", or \"an account\". Example: \"Retrieving a customer's recent orders to investigate a billing issue and help support determine the appropriate resolution.\"",
      +  "type": "string"
      +}
    • addedInput schema / properties / conversation_id
      Added value: +{
      +  "description": "Echo the conversation_id from the server's previous response. The server provides it on the first call — never invent one, and do not issue parallel tool calls until you have it.",
      +  "type": "string"
      +}
    • addedInput schema / properties / llm_model
      Added value: +{
      +  "description": "The exact model identifier you (the assistant) are running as, taken from your system prompt or environment (e.g. \"claude-opus-4-8\", \"gpt-5.2\"). Used for analytics only. If you do not know your model identifier with certainty, pass \"unknown\" — never guess.",
      +  "type": "string"
      +}
    • changedInput schema / required
      Previous value: -[
      -  "project_id",
      -  "entries"
      -]New value: +[
      +  "project_id",
      +  "entries",
      +  "context",
      +  "llm_model"
      +]
  3. Changed4 schema fields changed
    • changedInput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
    • removedInput schema / additionalProperties
      Removed value: -false
    • removedInput schema / properties / entries / items / anyOf
      Removed value: -[
      -  {
      -    "additionalProperties": false,
      -    "properties": {
      -      "action": {
      -        "const": "set_text",
      -        "description": "Write the voiceover transcript onto the clip without generating audio.",
      -        "type": "string"
      -      },
      -      "clip_index": {
      -        "description": "Zero-based clip index",
      -        "minimum": 0,
      -        "type": "integer"
      -      },
      -      "text": {
      -        "description": "Voiceover transcript",
      -        "minLength": 1,
      -        "type": "string"
      -      }
      -    },
      -    "required": [
      -      "action",
      -      "clip_index",
      -      "text"
      -    ],
      -    "type": "object"
      -  },
      -  {
      -    "additionalProperties": false,
      -    "properties": {
      -      "action": {
      -        "const": "generate_speech",
      -        "description": "Generate speech audio from the transcript already set on the clip.",
      -        "type": "string"
      -      },
      -      "clip_id": {
      -        "description": "Clip ID. Give this or clip_index.",
      -        "minLength": 1,
      -        "type": "string"
      -      },
      -      "clip_index": {
      -        "description": "Zero-based clip index. Give this or clip_id.",
      -        "minimum": 0,
      -        "type": "integer"
      -      }
      -    },
      -    "required": [
      -      "action"
      -    ],
      -    "type": "object"
      -  },
      -  {
      -    "additionalProperties": false,
      -    "properties": {
      -      "action": {
      -        "const": "set_and_generate",
      -        "description": "Write the transcript and generate its speech audio in one step.",
      -        "type": "string"
      -      },
      -      "clip_id": {
      -        "description": "Clip ID. Give this or clip_index.",
      -        "minLength": 1,
      -        "type": "string"
      -      },
      -      "clip_index": {
      -        "description": "Zero-based clip index. Give this or clip_id.",
      -        "minimum": 0,
      -        "type": "integer"
      -      },
      -      "text": {
      -        "description": "Voiceover transcript",
      -        "minLength": 1,
      -        "type": "string"
      -      }
      -    },
      -    "required": [
      -      "action",
      -      "text"
      -    ],
      -    "type": "object"
      -  }
      -]
    • addedInput schema / properties / entries / items / oneOf
      Added value: +[
      +  {
      +    "properties": {
      +      "action": {
      +        "const": "set_text",
      +        "description": "Write the voiceover transcript onto the clip without generating audio.",
      +        "type": "string"
      +      },
      +      "clip_index": {
      +        "description": "Zero-based clip index",
      +        "maximum": 9007199254740991,
      +        "minimum": 0,
      +        "type": "integer"
      +      },
      +      "text": {
      +        "description": "Voiceover transcript",
      +        "minLength": 1,
      +        "type": "string"
      +      }
      +    },
      +    "required": [
      +      "action",
      +      "clip_index",
      +      "text"
      +    ],
      +    "type": "object"
      +  },
      +  {
      +    "properties": {
      +      "action": {
      +        "const": "generate_speech",
      +        "description": "Generate speech audio from the transcript already set on the clip.",
      +        "type": "string"
      +      },
      +      "clip_id": {
      +        "description": "Clip ID. Give this or clip_index.",
      +        "minLength": 1,
      +        "type": "string"
      +      },
      +      "clip_index": {
      +        "description": "Zero-based clip index. Give this or clip_id.",
      +        "maximum": 9007199254740991,
      +        "minimum": 0,
      +        "type": "integer"
      +      }
      +    },
      +    "required": [
      +      "action"
      +    ],
      +    "type": "object"
      +  },
      +  {
      +    "properties": {
      +      "action": {
      +        "const": "set_and_generate",
      +        "description": "Write the transcript and generate its speech audio in one step.",
      +        "type": "string"
      +      },
      +      "clip_id": {
      +        "description": "Clip ID. Give this or clip_index.",
      +        "minLength": 1,
      +        "type": "string"
      +      },
      +      "clip_index": {
      +        "description": "Zero-based clip index. Give this or clip_id.",
      +        "maximum": 9007199254740991,
      +        "minimum": 0,
      +        "type": "integer"
      +      },
      +      "text": {
      +        "description": "Voiceover transcript",
      +        "minLength": 1,
      +        "type": "string"
      +      }
      +    },
      +    "required": [
      +      "action",
      +      "text"
      +    ],
      +    "type": "object"
      +  }
      +]
  4. Changed1 schema field changed
    • changedInput schema / properties / entries / items / anyOf
      Previous value: -[
      -  {
      -    "additionalProperties": false,
      -    "properties": {
      -      "action": {
      -        "const": "set_text",
      -        "description": "Write the voiceover transcript onto the clip without generating audio.",
      -        "type": "string"
      -      },
      -      "clip_index": {
      -        "description": "Zero-based clip index",
      -        "minimum": 0,
      -        "type": "integer"
      -      },
      -      "text": {
      -        "description": "Voiceover transcript",
      -        "minLength": 1,
      -        "type": "string"
      -      }
      -    },
      -    "required": [
      -      "action",
      -      "clip_index",
      -      "text"
      -    ],
      -    "type": "object"
      -  },
      -  {
      -    "additionalProperties": false,
      -    "properties": {
      -      "action": {
      -        "const": "generate_speech",
      -        "description": "Generate speech audio from the transcript already set on the clip.",
      -        "type": "string"
      -      },
      -      "clip_id": {
      -        "description": "Clip ID",
      -        "minLength": 1,
      -        "type": "string"
      -      }
      -    },
      -    "required": [
      -      "action",
      -      "clip_id"
      -    ],
      -    "type": "object"
      -  },
      -  {
      -    "additionalProperties": false,
      -    "properties": {
      -      "action": {
      -        "const": "set_and_generate",
      -        "description": "Write the transcript and generate its speech audio in one step.",
      -        "type": "string"
      -      },
      -      "clip_id": {
      -        "description": "Clip ID",
      -        "minLength": 1,
      -        "type": "string"
      -      },
      -      "clip_index": {
      -        "description": "Zero-based clip index",
      -        "minimum": 0,
      -        "type": "integer"
      -      },
      -      "text": {
      -        "description": "Voiceover transcript",
      -        "minLength": 1,
      -        "type": "string"
      -      }
      -    },
      -    "required": [
      -      "action",
      -      "clip_index",
      -      "clip_id",
      -      "text"
      -    ],
      -    "type": "object"
      -  }
      -]New value: +[
      +  {
      +    "additionalProperties": false,
      +    "properties": {
      +      "action": {
      +        "const": "set_text",
      +        "description": "Write the voiceover transcript onto the clip without generating audio.",
      +        "type": "string"
      +      },
      +      "clip_index": {
      +        "description": "Zero-based clip index",
      +        "minimum": 0,
      +        "type": "integer"
      +      },
      +      "text": {
      +        "description": "Voiceover transcript",
      +        "minLength": 1,
      +        "type": "string"
      +      }
      +    },
      +    "required": [
      +      "action",
      +      "clip_index",
      +      "text"
      +    ],
      +    "type": "object"
      +  },
      +  {
      +    "additionalProperties": false,
      +    "properties": {
      +      "action": {
      +        "const": "generate_speech",
      +        "description": "Generate speech audio from the transcript already set on the clip.",
      +        "type": "string"
      +      },
      +      "clip_id": {
      +        "description": "Clip ID. Give this or clip_index.",
      +        "minLength": 1,
      +        "type": "string"
      +      },
      +      "clip_index": {
      +        "description": "Zero-based clip index. Give this or clip_id.",
      +        "minimum": 0,
      +        "type": "integer"
      +      }
      +    },
      +    "required": [
      +      "action"
      +    ],
      +    "type": "object"
      +  },
      +  {
      +    "additionalProperties": false,
      +    "properties": {
      +      "action": {
      +        "const": "set_and_generate",
      +        "description": "Write the transcript and generate its speech audio in one step.",
      +        "type": "string"
      +      },
      +      "clip_id": {
      +        "description": "Clip ID. Give this or clip_index.",
      +        "minLength": 1,
      +        "type": "string"
      +      },
      +      "clip_index": {
      +        "description": "Zero-based clip index. Give this or clip_id.",
      +        "minimum": 0,
      +        "type": "integer"
      +      },
      +      "text": {
      +        "description": "Voiceover transcript",
      +        "minLength": 1,
      +        "type": "string"
      +      }
      +    },
      +    "required": [
      +      "action",
      +      "text"
      +    ],
      +    "type": "object"
      +  }
      +]
  5. First observed

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations mark this as destructive, and the description substantially honors and extends that warning. It details the rescaling of the entire clip, proportional retiming of all elements, potential dropping of zoom elements, async landing after return, and the non-last-write-wins concurrent conflict behavior. This is far more disclosure than the annotation alone provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but each section earns its place: summary, action list, addressing, order/results, rescale warning with concrete example, and concurrency rules. The most critical warning is front-loaded with IMPORTANT, and the structure makes a complex destructive tool safer to invoke.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with destructive side effects, async behavior, and subtle concurrency semantics, the description is complete. It covers param usage, return shape, timing, failure modes, and safe concurrent usage. Nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds meaning beyond the schema by explaining the three action modes, the clip_index/clip_id resolution rule (the omitted one is looked up once), per-call ordering, single-save semantics for text, and async TTS behavior. These are essential runtime semantics the schema does not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: "Set voiceover text and/or trigger speech generation for any number of clips in a single call." It clearly enumerates the three action types and distinguishes this batch tool from related sibling operations by emphasizing multi-clip, per-entry behavior.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance on when to generate speech before placing time-sensitive elements, how to address clips via clip_index vs clip_id, and when not to run concurrently with specific mutations like update_clips and add_audio. This goes well beyond a generic usage hint and actively routes the agent to safer workflows.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.