Skip to main content
Glama

caption_span_add

Idempotent

Add a caption span to change caption style or disable captions over a specific segment of a clip, with start and end defined by word, phrase, event, or duration.

Instructions

Captions off over a stretch of the film, or a different look there.

off=True draws none (over an end card or logo). style changes caption_style fields over the span only; {"max_words": 1, "size": 150, "position": "middle"} over one word draws it alone and large, a beat inside ordinary lines. Addressed like overlay_add: start at a word, phrase or event; end at a word, phrase, event or length. A line never crosses a span's edge; later spans win where two overlap. caption_view draws the result; add_captions writes it.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
offNoDraw no captions over the span. Give this or `style`.
pathNoThe project directory to act on. Omit it — the usual case — when this server is bound to a project (started as `proofcut -C DIR mcp`, or inside a project; `ping` says which): it then resolves to that one bound project, a relative path resolves against it, and a path outside it is refused by name. Unbound, `path` is the whole address and omitting it refuses rather than guessing.
planNoResolve the whole call and report what it would do, writing nothing. Prefer it over doing the thing and undoing it.
afterNoA forward cursor over a phrase's matches: any match at or before this word index is skipped. -1, the default, means from the start.
eventNoStart on this event of clip_id: `name`, or `name#k` when the name repeats.
styleNocaption_style fields (any but `preset`) that differ over the span, on top of the project's look: {"max_words": 1, "size": 150, "position": "middle"} draws each word alone and large. Give this or `off`.
phraseNoStart on this phrase's FIRST word, resolved against clip_id's transcript.
clip_idYesThe clip whose words or events address the span — the transcript the word indices index, or the recording the events belong to.
secondsNoEnd this long after the start.
occurrenceNoDisambiguate a phrase by count when it matches more than once, **1-based** in transcript order among the matches after `after`: 1 is the first, 2 the second. Unset, an ambiguous phrase is refused — listing every candidate's range and text — rather than guessed at.
word_indexNoThe word the span starts on. One of word_index, phrase or event.
until_eventNoEnd on this event of clip_id.
until_phraseNoEnd as this phrase's LAST word ends.
until_word_indexNoEnd as this word ends. One of until_word_index, until_phrase, until_event or seconds.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.43.0

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses important behavior: span-only style changes, line-breaking never crossing span edges, later spans winning on overlap, and where the result is drawn versus written by add_captions. The idempotent and non-destructive hints are not contradicted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but compact, front-loading the core purpose and then adding only high-value operational details: mode semantics, an example, addressing approach, overlap behavior, and downstream tools. No sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity — 14 parameters, multiple start/end addressing modes, and overlap semantics — the description is complete enough for an agent to call it correctly. It covers the mode selection, span boundaries, overlap resolution, and the relationship to caption_view and add_captions, while the schema handles parameter details and the output schema handles return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema description coverage, the schema already documents every parameter. The description adds valuable structural context by summarizing the addressing model (start at word/phrase/event, end at word/phrase/event/length) and by clarifying the off-vs-style relationship. This is an enhancement over the schema rather than a necessary compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific action and resource: applying either no captions or a different caption style over a film stretch. It then names the two modes (off and style) with a concrete example and distinguishes the tool from caption_view and add_captions by clarifying the pipeline roles.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear contextual guidance: off is for end cards/logos, style is for temporary per-span caption_style changes, and the one-word example illustrates a typical use. It also points to overlay_add for addressing semantics. It does not explicitly state when not to use the tool or name an alternative span-editing tool, so it stops short of a perfect 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools