Skip to main content
Glama

render_captions

Burn the same video again from edited words. Send the job's words.json with corrections applied, every word with its startMs and endMs in order (keep confidence if you want it in the new words.json); no speech model runs, what you send is what gets burned. Optionally change the preset or the highlight colour. Free within the job's allowance (3 on a paid account, 1 on the trial) and within 24 hours of completion. Returns the job at once with render.status accepted; poll get_caption_job until it is completed, when the artifact links switch to the new version. A retry is safe only when it carries the same idempotencyKey and the same body; without a key, each accepted call reserves a re-render from the allowance. Completion spends it; failure returns it. After 3 failed re-renders the job is paused for investigation, with rerenderFailureLimitReached true; do not automatically retry. Re-renders count toward the account's concurrent job limit while they run.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
jobIdYesA job id returned by this API.
wordsYesThe edited words.json: every word with its timing, in order.
presetNoOptional. Defaults to the job's preset.
highlightColorNoOptional, highlight preset only. The spoken word's colour as a hex RGB. Defaults to the job's colour.
idempotencyKeyNoOptional. Any string unique to this request; a retry with the same key returns the same result instead of running twice.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply coarse hints (readOnlyHint=false, idempotentHint=false); the description adds the substantive behavior: allowance accounting (3 paid / 1 trial, reserved on accept, spent on completion, returned on failure), the 24-hour window, conditional idempotency via idempotencyKey, the failure-limit pause with rerenderFailureLimitReached, and that re-renders consume the concurrent job limit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then the constraints. It is long, but nearly every sentence carries operational information (allowance, window, idempotency, failure limit); the length is justified by the tool's transactional complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with an output schema, this covers everything an agent needs: inputs, ordering requirement, immediate return shape (render.status accepted), the required polling path, and the retry/failure boundary conditions. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description exceeds it by explaining that the words array must be ordered with startMs/endMs and that confidence is optional, and by spelling out the semantics of idempotencyKey ('a retry with the same key returns the same result instead of running twice'), which the schema describes only briefly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: re-burn the same video from an edited words.json, explicitly distinguishing it from the initial captioning step ('no speech model runs'). An agent can tell it apart from caption_video and get_caption_job without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear triggering context (you have corrected words), names the follow-up tool (poll get_caption_job until completed), and states eligibility windows (free within allowance, within 24 hours of completion) plus an explicit 'do not automatically retry' after the failure limit. It stops short of explicitly framing when to prefer this over caption_video, so 4 rather than 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources