Skip to main content
Glama

caption_video

Start captioning one video. Pass exactly one source: inputUrl, a direct https link to a video file (platform pages such as YouTube or TikTok are refused), or jobId from create_upload after the bytes are uploaded. Choose a preset: highlight (the spoken word takes the accent), clean (white with an outline) or boxed (a plate behind the line). Optionally send a dictionary of names to spell right and a language. Returns a jobId in status probing; poll it with get_caption_job. If the response is lost, replay with the original idempotencyKey and payload or use list_caption_jobs to find recent jobs. The balance is charged by the whole second once the duration is known, and nothing is charged if the job fails.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
jobIdNoThe jobId from create_upload, after the bytes have been uploaded.
presetYesCaption style. highlight, clean or boxed.
inputUrlNoA direct https link to a video file. Platform pages such as YouTube or TikTok are refused.
languageNo"auto" to detect, or a BCP-47 language tag such as "en" or "pt-BR".
dictionaryNoOptional. Up to 1000 names or terms, each at most 6 words, spelled the way they should appear. Used for this job only.
highlightColorNoOptional, highlight preset only. The spoken word's colour as a hex RGB such as #FFE500. Defaults to the CaptionPipe accent.
idempotencyKeyNoOptional. Any string unique to this request; a retry with the same key returns the same result instead of running twice.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
jobIdYes
statusYes
balanceYes
balanceWarningYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial runtime behavior beyond the annotations: the returned jobId is for status probing, replaying with the same idempotencyKey and payload is safe, billing is by the whole second once duration is known, and nothing is charged if the job fails. This does not contradict readOnlyHint=false or idempotentHint=false because idempotency is conditional on the idempotencyKey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action and source constraint, and every sentence earns its place: presets, optional inputs, return behavior, recovery, and billing. It is a dense single paragraph rather than structured bullets, but it is not padded or redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the 7-parameter schema, the oneOf source logic, and important billing/recovery behavior, the description covers all essential guidance: source choice, presets, optional parameters, polling, idempotent recovery, and cost. Since an output schema exists, not detailing the return value is acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema: exactly one of inputUrl/jobId must be provided, highlightColor is for the highlight preset only, dictionary entries are names/terms to spell correctly, and language can be auto-detected. This gives the agent selection logic the schema alone does not fully convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening phrase 'Start captioning one video' states a specific verb and resource, and the body clarifies the two input sources and preset options. It also references adjacent workflow tools (create_upload, get_caption_job, list_caption_jobs), making it easy to distinguish from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit selection rules: pass exactly one of inputUrl (direct https video file, platform pages refused) or jobId from create_upload after upload. It also prescribes follow-up actions: poll with get_caption_job, and on lost response replay with idempotencyKey or use list_caption_jobs. It does not explicitly contrast with render_captions, so there is a small gap in sibling differentiation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources