Skip to main content
Glama

Transcript Kit

Translate subtitles and make a dub script

translate_subtitles
Read-only

Use this when the user wants captions or a spoken script in another language, for example "translate these subtitles to Spanish", "make a French dub script for my talk". Pass the transcript_id from transcribe_media, or the text of an SRT or VTT file the user owns as subtitles with source_lang and confirm_rights true (only after the user says they own the recording or may use it). Always pass target_lang. Limits: about 18,000 characters per call (roughly 20 minutes of speech) and a daily allowance of characters per user and in total. Returns translated SRT and VTT with the original times, and a dub script with one spoken line per segment, its start and end, max_chars for the time it has and a fit of ok, tight or long (trim with include). Translated with Cloudflare Workers AI (m2m100), about $0.0002 per 1,000 characters; over-long dub lines are shortened with a small Llama model. Nothing is stored. Pass library_id only if transcribe_media returned one.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
includeNoWhat to return (default all three): translated SRT captions, translated VTT captions, a dub script with one spoken line per segment.
subtitlesNoThe text of an SRT or WebVTT file the user owns or may use. Give this or transcript_id.
library_idNoThe library_id that transcribe_media returned. Only needed when transcribe_media returned one; leave it out otherwise.
source_langNoLanguage of the text. Needed for subtitles; for a transcript the language it was detected in is used.
target_langYesLanguage to translate into, as a two-letter code such as es, fr, de, ja, ar or hi.
transcript_idNoThe transcript_id from transcribe_media. Give this or subtitles.
confirm_rightsNoSet true only after the user has said they own this recording or have the rights to translate it. Required with subtitles.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
srtNo
vttNo
as_ofYes
modelYes
notesYes
usageYes
sourceYes
statusYes
targetYes
cue_countYes
dub_modelNoPresent when some dub lines were shortened by this model.
dub_scriptNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (readOnlyHint, idempotentHint false) by disclosing hard limits (~18,000 chars/call, ~20 min of speech, daily per-user and total allowances), cost (~$0.0002 per 1,000 chars), the models used (m2m100, small Llama for trimming), the fitting behavior of dub lines (ok/tight/long), and that nothing is stored. These are behavioral traits an agent needs and that structured fields do not carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the trigger condition and examples, then limits, return shape, cost and parameter scoping. It is dense and slightly wall-of-text, with parameter guidance packed into the same paragraph, but nearly every sentence carries information an agent needs, so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even though an output schema exists, the description pre-summarizes the returns (translated SRT/VTT with original times, dub script with max_chars and fit values), covers limits, cost, rights gating and rare parameters. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is 3, but the description adds relational meaning the schema cannot: it ties subtitles to source_lang and confirm_rights true, clarifies target_lang is always required, and scopes library_id to only when transcribe_media returned one. It explains how include interacts with the over-long-line trimming.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (translate subtitles, produce a dub script) and anchors it with concrete user phrasings. It also implicitly distinguishes itself from the sibling transcribe_media by explaining that it consumes a transcript_id or SRT/VTT text rather than producing the transcript.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it (user wants captions or a spoken script in another language) with example utterances. It routes between input modes — pass transcript_id from transcribe_media or owned SRT/VTT text with source_lang — and states the confirm_rights precondition (only after the user says they own the recording).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources