Skip to main content
Glama

Add caption track

add_caption_track

Build a caption track from pre-timed lines (e.g. derived from transcribe_clip's word timings). mode "line-sync" (default) creates one text layer per line, each shown only during its [startFrame, endFrame) window via hold-eased opacity keyframes — the active-line karaoke read; mode "static" makes a single layer with all lines joined. style picks a preset look. Lines default to a lower-third band. The caption layers are always wrapped in a "captions" group so they don't clutter the layers list. Returns the created text element ids plus groupElementId (the captions group).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
xNoCaption band centre x. Default canvas centre.
yNoCaption band centre y. Default lower third (~80% of height).
modeNo"line-sync" (default): one timed layer per line. "static": one layer with all lines.
linesYesCaption lines in order. Each: { text, startFrame, endFrame } — frames are 0-indexed at 30fps.
styleNoCaption look preset. Default "classic".
widthNoBand width. Default ~86% of canvas width.
heightNoBand height. Default ~16% of canvas height.
projectIdYesOpaque project id (a v4 UUID, from list_projects/create_project). Selects which existing project this call mutates.
clip_element_idNoOptional "video.<id>" to WELD the caption lines to (line-sync mode). When set, each line's startFrame/endFrame are treated as its window in the clip's OWN source timeline and the on-timeline position is derived live from the clip's trim — so trimming or sliding the clip retimes/clips the captions, exactly like the clip's welded audio. Omit for fixed project-frame captions.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
okYesWhether the call succeeded.
dataNoThe payload, shaped by the tool.
noteNoWhat to do next when not ready.
errorNoWhy it failed.
statusNoFor cache-backed readers: whether the answer was ready.
editorUrlNoOpens this project in the editor.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changed
    • removedInput schema / properties / anonymousToken
      Removed value: -{
      -  "description": "The token create_anonymous_account returned. This connection has no Morpha key, so every call carries it.",
      -  "type": "string"
      -}
    • changedInput schema / required
      Previous value: -[
      -  "lines",
      -  "projectId",
      -  "anonymousToken"
      -]New value: +[
      +  "lines",
      +  "projectId"
      +]
  2. Changed1 schema field changed
    • changedOutput schema / (root)
      Previous value: -nullNew value: +{
      +  "description": "The result envelope every Morpha tool returns.",
      +  "properties": {
      +    "data": {
      +      "description": "The payload, shaped by the tool.",
      +      "type": [
      +        "object",
      +        "array",
      +        "string",
      +        "number",
      +        "boolean",
      +        "null"
      +      ]
      +    },
      +    "editorUrl": {
      +      "description": "Opens this project in the editor.",
      +      "type": "string"
      +    },
      +    "error": {
      +      "description": "Why it failed.",
      +      "type": "string"
      +    },
      +    "note": {
      +      "description": "What to do next when not ready.",
      +      "type": "string"
      +    },
      +    "ok": {
      +      "description": "Whether the call succeeded.",
      +      "type": "boolean"
      +    },
      +    "status": {
      +      "description": "For cache-backed readers: whether the answer was ready.",
      +      "enum": [
      +        "ready",
      +        "not-ready"
      +      ],
      +      "type": "string"
      +    }
      +  },
      +  "required": [
      +    "ok"
      +  ],
      +  "type": "object"
      +}
  3. Changed2 schema fields changed
    • addedInput schema / properties / anonymousToken
      Added value: +{
      +  "description": "The token create_anonymous_account returned. This connection has no Morpha key, so every call carries it.",
      +  "type": "string"
      +}
    • changedInput schema / required
      Previous value: -[
      -  "lines",
      -  "projectId"
      -]New value: +[
      +  "lines",
      +  "projectId",
      +  "anonymousToken"
      +]
  4. First observed

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only say the tool is not read-only and not destructive; the description carries the full behavioral burden and does so richly. It discloses hold-eased opacity keyframes, active-line karaoke behavior, the always-wrapped 'captions' group, welding semantics with clip trim, and the return of element ids plus groupElementId. All of this goes well beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: mode explanation, karaoke effect, styling, defaults, grouping, return value. It front-loads the core purpose and then increments detail in a logical order. No filler or repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 9 parameters, two modes, nested line objects, and welding behavior, the description covers defaults, frame semantics, group wrapping, return values, and the relationship to other tools like transcribe_clip. The existing output schema is mentioned and the return ids are explained. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so parameters are well documented in the schema itself. The description adds extra semantics on top, including the 30fps frame-index convention, the default lower-third position, and the practical purpose of clip_element_id for spreading lines across split source clips. This is meaningful interpretive value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair ('Build a caption track from pre-timed lines') and immediately clarifies the two modes. It is clearly distinct from sibling tools like add_text_layer, merge_caption_lines, and split_caption_line because it describes building a complete timed caption track with wrapping and styling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives strong usage context, including the provenance of input ('derived from transcribe_clip's word timings'), the default lower-third placement, and the caption-group wrapping behavior. It does not explicitly name alternative tools or say 'use this instead of X', but the detailed behavior makes the intended scenario obvious.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources