Skip to main content
Glama

video_editor_upload_captions

Destructive

Import subtitles (SRT or VTT) onto a video as caption rows that burn in on export. Pass the text inline via srt, or an already-uploaded subtitle asset_id. Cues are parsed into timed rows (absolute seconds). render "static" (default) burns each cue as a line via drawtext; "wordpop" shows one word at a time (TikTok pop); "karaoke" sweeps a highlight across the line — both render via libass, with each cue's words spread evenly across its interval. Replaces existing captions unless append=true.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
srtNoInline SRT or VTT text.
styleNo
y_pctNoVertical position 0..100 (default lower-third; ~65 keeps it out of the UI zone).
appendNo
renderNo
asset_idNoA subtitle asset previously uploaded (kind=subtitle).
video_idNo
project_idYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses the destructive behavior ('Replaces existing captions unless append=true'), which aligns with the destructiveHint annotation. It also explains the rendering differences between 'static', 'wordpop', and 'karaoke', and the parsing of cues into timed rows. This goes beyond what annotations alone provide, offering useful behavioral context without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences long and front-loaded with the core purpose. Each sentence adds meaningful detail: input methods, rendering modes, and replacement behavior. It is structured logically and avoids unnecessary fluff, though it could be slightly more concise in the rendering explanation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 8 parameters, including enums and a destructive flag, and no output schema. The description covers the main functionality, input methods, render modes, and append behavior, but it omits style semantics and the roles of video_id and project_id. While it addresses the most critical aspects, the missing parameter details and lack of return-value guidance leave it incomplete for an agent to fully understand all edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 38%, so the description must compensate. It clarifies srt and asset_id (inline vs. asset), explains render modes in detail, and clarifies append behavior. However, it does not explain the style parameter (which has enums), nor video_id and project_id (though project_id is required). The description adds value for key parameters but leaves others undocumented, so it partially compensates for the low schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action: 'Import subtitles (SRT or VTT) onto a video as caption rows that burn in on export.' It names the resource (subtitles onto a video) and the verb (import), making the tool's purpose unambiguous. It does not explicitly differentiate from sibling tools like video_editor_export_captions, but the purpose is specific enough that an agent can infer when to use it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains two input methods ('Pass the text inline via srt, or an already-uploaded subtitle asset_id') and notes the default behavior ('Replaces existing captions unless append=true'). However, it does not explicitly state when to use this tool over alternatives (e.g., video_editor_upload_file or video_editor_export_captions). The guidance is contextual but not comparative, so it falls short of a strong usage guideline.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.