CaptionPipe
Server Details
One call burns captions into your video. Hosted, prepaid, no subscription. MP4 plus SRT and VTT.
- Status
- Healthy
- Uptime
- 96.3% over 22 days
- OAuth
- Works in Glama
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 5 tools
Each tool maps to a distinct stage in the captioning workflow: uploading, starting, polling, listing, and re-rendering. caption_video and render_captions both produce jobs, but one processes a new video while the other burns corrected words into an already-captioned video, so there is no real ambiguity.
All tool names follow the same snake_case verb_noun pattern: create_upload, caption_video, get_caption_job, list_caption_jobs, render_captions. The naming is uniform and predictable.
Five tools cover the full captioning pipeline without excess. The count feels right for the server's purpose: upload, start, poll, list, and re-render are each necessary and none are redundant.
The core lifecycle is well covered: uploading, captioning, status polling, recovering lost jobs, and re-rendering with corrections. The only notable absence is an explicit cancel/delete operation, but expiring artifacts and failure handling make this a minor gap rather than a blocking one.
Available Tools
5 toolscaption_videoAInspect
Start captioning one video. Pass exactly one source: inputUrl, a direct https link to a video file (platform pages such as YouTube or TikTok are refused), or jobId from create_upload after the bytes are uploaded. Choose a preset: highlight (the spoken word takes the accent), clean (white with an outline) or boxed (a plate behind the line). Optionally send a dictionary of names to spell right and a language. Returns a jobId in status probing; poll it with get_caption_job. If the response is lost, replay with the original idempotencyKey and payload or use list_caption_jobs to find recent jobs. The balance is charged by the whole second once the duration is known, and nothing is charged if the job fails.
| Name | Required | Description | Default |
|---|---|---|---|
| jobId | No | The jobId from create_upload, after the bytes have been uploaded. | |
| preset | Yes | Caption style. highlight, clean or boxed. | |
| inputUrl | No | A direct https link to a video file. Platform pages such as YouTube or TikTok are refused. | |
| language | No | "auto" to detect, or a BCP-47 language tag such as "en" or "pt-BR". | |
| dictionary | No | Optional. Up to 1000 names or terms, each at most 6 words, spelled the way they should appear. Used for this job only. | |
| highlightColor | No | Optional, highlight preset only. The spoken word's colour as a hex RGB such as #FFE500. Defaults to the CaptionPipe accent. | |
| idempotencyKey | No | Optional. Any string unique to this request; a retry with the same key returns the same result instead of running twice. |
Output Schema
| Name | Required | Description |
|---|---|---|
| jobId | Yes | |
| status | Yes | |
| balance | Yes | |
| balanceWarning | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial runtime behavior beyond the annotations: the returned jobId is for status probing, replaying with the same idempotencyKey and payload is safe, billing is by the whole second once duration is known, and nothing is charged if the job fails. This does not contradict readOnlyHint=false or idempotentHint=false because idempotency is conditional on the idempotencyKey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and source constraint, and every sentence earns its place: presets, optional inputs, return behavior, recovery, and billing. It is a dense single paragraph rather than structured bullets, but it is not padded or redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the 7-parameter schema, the oneOf source logic, and important billing/recovery behavior, the description covers all essential guidance: source choice, presets, optional parameters, polling, idempotent recovery, and cost. Since an output schema exists, not detailing the return value is acceptable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema: exactly one of inputUrl/jobId must be provided, highlightColor is for the highlight preset only, dictionary entries are names/terms to spell correctly, and language can be auto-detected. This gives the agent selection logic the schema alone does not fully convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening phrase 'Start captioning one video' states a specific verb and resource, and the body clarifies the two input sources and preset options. It also references adjacent workflow tools (create_upload, get_caption_job, list_caption_jobs), making it easy to distinguish from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit selection rules: pass exactly one of inputUrl (direct https video file, platform pages refused) or jobId from create_upload after upload. It also prescribes follow-up actions: poll with get_caption_job, and on lost response replay with idempotencyKey or use list_caption_jobs. It does not explicitly contrast with render_captions, so there is a small gap in sibling differentiation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_uploadAInspect
Only for a video file on the caller's side, at most 2 GB. Returns a PUT-only upload URL for one object, valid for one hour; PUT the raw bytes to it, then call caption_video with the returned jobId. Pass contentLength if you know the file's size in bytes and the URL is bound to it; leave it out if you do not. If the video is already at a direct https URL, skip this tool entirely and pass inputUrl to caption_video. A retry is safe only when it carries the same idempotencyKey; without one, each call creates a new upload.
| Name | Required | Description | Default |
|---|---|---|---|
| sha256 | No | Optional. Hex SHA-256 of the file; verified after upload when provided. | |
| contentLength | No | Optional. Exact size of the file in bytes, at most 2 GB. When given it is bound into the upload URL, so a PUT of any other size is refused; when omitted the size is checked after the upload instead. | |
| idempotencyKey | No | Optional. Any string unique to this request; a retry with the same key returns the same result instead of running twice. |
Output Schema
| Name | Required | Description |
|---|---|---|
| jobId | Yes | |
| balance | Yes | |
| expiresAt | Yes | |
| uploadUrl | Yes | |
| balanceWarning | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations mark this as a non-read-only, non-idempotent mutation but say nothing else. The description adds the operational contract the agent actually needs: the URL is PUT-only, single-object, and expires in one hour, the raw bytes must be uploaded before captioning, and retries are only safe when reusing the idempotencyKey. This is consistent with idempotentHint=false rather than contradicting it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences, front-loaded with the eligibility constraint and the return value, then the PUT workflow, then the idempotency caveat. No filler and nothing repeated from the schema verbatim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, yet the description still tells the agent what the call yields (a one-hour PUT URL and a jobId) and what to do with it, covering the full call sequence for a zero-required-param tool. Nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds decision guidance beyond the schema: pass contentLength only when the size is known because the URL becomes size-bound, otherwise omit it. It does not mention sha256, though the schema fully documents that parameter, so no information is actually lost.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: it returns a PUT-only upload URL for a single object, restricted to a caller-side video file of at most 2 GB. It also cleanly separates itself from caption_video, which is named as the follow-up/alternative tool, so an agent can pick it without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is fully bounded: use it only for a local video file, skip it entirely and pass inputUrl to caption_video when the video already sits at an https URL, and call caption_video with the returned jobId after the PUT. Retry conditions are also spelled out (same idempotencyKey or a new upload is created).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_caption_jobARead-onlyIdempotentInspect
Poll a known job. If the job ID was lost, use list_caption_jobs to find it. While it runs, retryAfterSeconds says how long to wait before calling again. When status is completed, artifacts holds signed links to the captioned MP4, SRT, VTT, ASS and words.json, valid until expiresAt (24 hours after completion), and chargedSeconds is final. After a render_captions call, render says which version was asked for last and whether it is still rendering, done, or failed; the links always belong to the newest finished version, and preset describes the newest request. When status is failed, error says why and nothing was charged. dictionaryCapacityApplied is the dictionary capacity the speech vendor honoured (1000 on the primary vendor, 0 on the fallback), not a count of terms. Every response includes the caller's balance.
| Name | Required | Description | Default |
|---|---|---|---|
| jobId | Yes | A job id returned by this API. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, and non-destructive behavior; the description adds real operational traits: retry timing, signed-link expiry, final chargedSeconds, error/no-charge on failure, dictionaryCapacityApplied meaning, and balance in every response. These are exactly the non-obvious behaviors an agent needs beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Long but dense: every clause covers a distinct operational fact such as polling, completion, versioning, failure, dictionary capacity, and balance. The main action is front-loaded nored there is no filler or repetition of the annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Combined with the full input schema, output schema, and annotations, the description covers polling semantics, state transitions, output links, versioning, billing on failure, dictionary capacity, and the always-present balance. There is no material gap for an agent invoking get_caption_job.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single jobId parameter is already fully documented in the schema as an id returned by this API. The description adds the 'known job' framing and a fallback to list_caption_jobs for the lost-id case, but no new syntactic or format-level detail beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with 'Poll a known job,' naming a specific verb and resource. It immediately differentiates from list_caption_jobs by defining get_caption_job as the tool for a known id, so an agent won't confuse the two.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells when to poll this tool ('while it runs' with retryAfterSeconds) and when the id is lost to call list_caption_jobs instead. It also explains post-render behavior, including versioning semantics after render_captions, which is direct usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_caption_jobsARead-onlyIdempotentInspect
List your caption jobs newest first to recover a lost job ID or review recent results. Returns one page; pass nextCursor with the same status filter to continue. Includes running, failed, and completed jobs, including completed jobs whose files expired. No captioning charge. Use get_caption_job for one job's details and download links. Similar submissions may need user confirmation; a missing row does not make resubmission safe. Retry an interrupted submission with its original idempotencyKey and payload when available.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Number of jobs per page, 1 to 100. Defaults to 25. | |
| cursor | No | The nextCursor from a previous page. Keep the same status filter. | |
| status | No | Filter by one job status; omit for all states. |
Output Schema
| Name | Required | Description |
|---|---|---|
| jobs | Yes | |
| balance | Yes | |
| nextCursor | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint and idempotentHint annotations, the description discloses pagination behavior (one page, nextCursor with same status filter), coverage (running, failed, completed including expired files), billing (no captioning charge), and a caution about user confirmation. It adds substantial behavioral context beyond what annotations alone provide, and does not contradict them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: purpose, pagination, inclusion criteria, cost, alternative tool, safety warning, retry guidance. The most important information is front-loaded, and the description is tight with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists (which likely covers return structure) and annotations cover the read-only/idempotent nature, the description provides all necessary operational details: ordering, pagination, inclusivity, cost, safety, and alternative routing. An agent can call this tool correctly without needing further clarification.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all three parameters have descriptions. The description adds value by explaining the cursor's dependency on the status filter ('pass nextCursor with the same status filter to continue') and the implicit ordering (newest first), which is not in the schema. It doesn't elaborate on the limit param, but the schema already covers that, so the baseline 3 is elevated to 4 due to the cursor/status nuance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: 'List your caption jobs newest first' with clear purposes (recover a lost job ID, review recent results). It also distinguishes itself from the sibling get_caption_job by explicitly saying to use that tool for single-job details and downloads.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance (recover lost ID, review results) and names the alternative (get_caption_job) with a clear condition. It adds a safety warning about resubmission, telling agents that a missing row does not imply safe resubmission, and advises retrying with the original idempotencyKey. This goes beyond basic usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
render_captionsAInspect
Burn the same video again from edited words. Send the job's words.json with corrections applied, every word with its startMs and endMs in order (keep confidence if you want it in the new words.json); no speech model runs, what you send is what gets burned. Optionally change the preset or the highlight colour. Free within the job's allowance (3 on a paid account, 1 on the trial) and within 24 hours of completion. Returns the job at once with render.status accepted; poll get_caption_job until it is completed, when the artifact links switch to the new version. A retry is safe only when it carries the same idempotencyKey and the same body; without a key, each accepted call reserves a re-render from the allowance. Completion spends it; failure returns it. After 3 failed re-renders the job is paused for investigation, with rerenderFailureLimitReached true; do not automatically retry. Re-renders count toward the account's concurrent job limit while they run.
| Name | Required | Description | Default |
|---|---|---|---|
| jobId | Yes | A job id returned by this API. | |
| words | Yes | The edited words.json: every word with its timing, in order. | |
| preset | No | Optional. Defaults to the job's preset. | |
| highlightColor | No | Optional, highlight preset only. The spoken word's colour as a hex RGB. Defaults to the job's colour. | |
| idempotencyKey | No | Optional. Any string unique to this request; a retry with the same key returns the same result instead of running twice. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply coarse hints (readOnlyHint=false, idempotentHint=false); the description adds the substantive behavior: allowance accounting (3 paid / 1 trial, reserved on accept, spent on completion, returned on failure), the 24-hour window, conditional idempotency via idempotencyKey, the failure-limit pause with rerenderFailureLimitReached, and that re-renders consume the concurrent job limit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then the constraints. It is long, but nearly every sentence carries operational information (allowance, window, idempotency, failure limit); the length is justified by the tool's transactional complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with an output schema, this covers everything an agent needs: inputs, ordering requirement, immediate return shape (render.status accepted), the required polling path, and the retry/failure boundary conditions. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description exceeds it by explaining that the words array must be ordered with startMs/endMs and that confidence is optional, and by spelling out the semantics of idempotencyKey ('a retry with the same key returns the same result instead of running twice'), which the schema describes only briefly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: re-burn the same video from an edited words.json, explicitly distinguishing it from the initial captioning step ('no speech model runs'). An agent can tell it apart from caption_video and get_caption_job without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear triggering context (you have corrected words), names the follow-up tool (poll get_caption_job until completed), and states eligibility windows (free within allowance, within 24 hours of completion) plus an explicit 'do not automatically retry' after the failure limit. It stops short of explicitly framing when to prefer this over caption_video, so 4 rather than 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
- Added
list_caption_jobs
4 tool updates
- First observed
caption_video - First observed
create_upload - First observed
get_caption_job - First observed
render_captions
Related MCP Connectors
- caption.shOAuthsh.caption
Burn styled, word-timed karaoke captions into any video - captions API and MCP server for agents.
Narrated, captioned videos from a topic or script, editable scene by scene.
YouTube subtitles and transcripts as SRT, VTT or plain text. Pay per download with credits.
Convert subtitles, transcripts, broadcast captions (SCC/MCC/STL), EDLs, and Premiere files.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceAn MCP server that automatically transcribes video content and burns stylized captions directly into the video file. It leverages the Groq Whisper API for fast transcription and supports multiple visual styles tailored for social media and professional content.-
- AlicenseNot gradedqualityCmaintenanceTurn words, images, and audio into an animated video with MP4 export.MIT

cartcut-mcpofficial
AlicenseNot gradedqualityBmaintenanceEdit video in the CartCut desktop editor: cuts from a transcript, captions, motion, effects.160 npm760MIT
AITuber MCP Serverofficial
AlicenseAqualityBmaintenanceCreate AI-powered videos from any MCP-compatible client. Generate videos with AI narration, visuals, and synced captions for short-form and long-form content.250 npm5MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.