Caption File Check
Server Details
Parse WebVTT, SRT, or TTML for conformance, timing, overlaps, line length, and reading speed.
- Status
- Healthy
- Uptime
- 100.0% over 21 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- powmcp/mcp-server-guide
- GitHub Stars
- 0
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: caption_check handles a single file, while caption_compare handles exactly two files. There is no overlap or ambiguity in which tool to call for a given task.
Both tools follow the same 'caption_' prefix plus a clear verb: caption_check and caption_compare. The naming pattern is perfectly consistent and immediately conveys the operation.
Two tools is on the lean side of the typical range, but the server's scope is narrowly focused on caption file preflight and comparison. Each tool earns its place and together they cover the core workflows without unnecessary extras.
For the stated purpose of caption/subtitle file validation, the surface is complete: one tool for single-file checks and one for side-by-side comparison. There are no obvious missing operations or dead ends within the server's domain.
Available Tools
2 toolscaption_checkARead-onlyInspect
Preflight one caption or subtitle file before delivery. Attach a WebVTT (.vtt), SRT (.srt), or TTML/IMSC (.ttml/.xml) caption file in chat or give a publicly fetchable URL and it downloads the artifact under an SSRF-guarded, 2 MB-capped fetch, detects the format by content and the byte encoding and BOM, then parses it with the real spec engine (the WebVTT spec parser, a dedicated SRT parser, or the imscJS TTML/IMSC profile engine) and returns deterministic evidence: pass/fail, cue count, structural spec and IMSC profile-conformance errors, overlapping / negative / out-of-order / empty cues, per-line length, and per-cue and summary reading speed (CPS) against a named platform profile (netflix-adult-20, netflix-children-17, youtube, ebu-tt). It reports only source-file conformance, timing, encoding, and reading-speed facts against a dated profile. It does not judge translation quality, transcription accuracy, or synchronization against the video. An arbitrary XML or text file that is not a caption track returns an explicit unsupported-format error. Fetching and parsing a full-length track can take up to 60 seconds. WebVTT limitation: the parser drops payload-less cues; cue counts, indices, timing checks and comparisons cover retained cues only, not those omitted cues.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | A publicly fetchable HTTP(S) URL of a WebVTT, SRT, or TTML/IMSC caption file, detected by content not extension. Provide either this or an attached file, never both. Query strings are allowed but never logged. Pasted caption bodies, data URLs, FTP/SFTP, and URLs needing credentials, headers, or cookies are not supported. STRICT RULE: Caption check failed — provide either a caption URL or an attached caption file, not both STRICT RULE: Caption check failed — provide a public caption URL or attach the caption file | |
| file | No | A caption file attached in chat (WebVTT, SRT, or TTML/IMSC), detected by content not extension. Provide either this or url, never both. | |
| profile | No | Named reading-speed profile to judge CPS and line length against. netflix-adult-20 (default, 20 CPS) and netflix-children-17 (17 CPS) are spec-backed; youtube and ebu-tt are advisory ~17 CPS heuristics. |
Output Schema
| Name | Required | Description |
|---|---|---|
| bom | Yes | Byte-order-mark detected: utf-8, utf-16le, utf-16be, or none |
| cps | Yes | Per-cue and summary reading-speed measurements |
| url | Yes | The checked caption source: the submitted URL query- and fragment-stripped, or the attached file's name |
| pass | Yes | True when the file passes all checked format, timing, line-length and reading-speed criteria for the selected profile. False can include an advisory readability threshold breach, not just a specification error; this is not a platform-acceptance verdict |
| timing | Yes | Measured retained-cue timing and empty-text findings; WebVTT indices refer to retained cues, not original source positions |
| profile | Yes | Reading-speed and line-length profile applied to the caption checks |
| cueCount | Yes | Number of caption cues retained for checking; payload-less WebVTT cues are omitted by the parser |
| encoding | Yes | Detected byte encoding of the caption file |
| specUsed | Yes | The dated spec identifier the detected format was checked against |
| checkedAt | Yes | ISO timestamp when the caption result was produced |
| lineLength | Yes | Per-line character counts checked against the profile limit |
| truncation | Yes | Markers identifying checks truncated by configured limits |
| rawCueCount | Yes | Cues parsed before the 20,000-cue cap; excludes payload-less WebVTT cues dropped by the parser |
| ttmlProfile | No | IMSC/TTML profile designator declared in the file, when present |
| detectedFormat | Yes | webvtt, srt, or ttml (covers IMSC), detected by content |
| structuralErrors | Yes | Spec-parser and IMSC/TTML profile-conformance findings |
| structuralErrorCounts | Yes | Counts of structural findings grouped by severity |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses substantial behavioral detail: SSRF-guarded and 2 MB-capped fetching, content-based format detection, the specific parsing engines used, deterministic evidence output, exclusions of quality judgments, explicit unsupported-format error, a 60-second timeout, and a WebVTT limitation about dropped payload-less cues. This far exceeds what annotations alone provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and lengthy, but every sentence contributes meaningful operational detail, from fetch constraints to parser limitations. The main purpose is front-loaded, and the WebVTT caveat is placed at the end where it supplements rather than obscures the core message.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema available, the description does not need to detail return values. It covers input requirements (URL or file), fetch constraints, format detection, parser behavior, timing limits, error handling, scope exclusions, and a known parser limitation. For a tool of this complexity, nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents url, file, and profile. The tool description adds context about format detection and reading-speed profiles, but these are also reflected in the schema parameter descriptions. It does not materially enhance parameter-level understanding beyond the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Preflight one caption or subtitle file before delivery,' a specific verb and resource that clearly identifies the tool's function. It further distinguishes itself from the sibling caption_compare by scoping to a single file and explicitly stating what it does not do (judge translation quality, transcription accuracy, or synchronization).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'before delivery' implies a use case, but there is no explicit when-to-use versus when-not-to-use guidance, nor any mention of the sibling caption_compare as an alternative. An agent must infer that this is for single-file validation rather than comparison.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
caption_compareARead-onlyInspect
Compare exactly two caption or subtitle files under identical bounded checks: an original versus its translated track, before versus after a timing fix, or a format migration. Attach two WebVTT, SRT, or TTML/IMSC caption files in chat or give two publicly fetchable URLs and it runs the same SSRF-guarded, 2 MB-capped preflight as caption_check on each, then aligns them by cue index and reports side-by-side pass status, detected-format and encoding differences, the cue-count delta, timing drift and overlap differences, reading-speed (CPS) regressions, structural-error deltas, and a winner only when the measured checks separate the two. Use caption_check for a single file. It reports only source-file conformance, timing, encoding, and reading-speed facts against a dated profile: never translation quality, transcription accuracy, or sync against the video. Fetching and parsing two live files can take up to 60 seconds. WebVTT limitation: the parser drops payload-less cues; cue counts, indices, timing checks and comparisons cover retained cues only, not those omitted cues.
| Name | Required | Description | Default |
|---|---|---|---|
| urls | No | Exactly two public caption URLs to compare, in the order they should appear side by side (for example original then translated). Provide either this or files, never both. STRICT RULE: Caption compare failed — provide either two caption URLs or two attached caption files, not both STRICT RULE: Caption compare failed — provide two public caption URLs or attach two caption files | |
| files | No | Exactly two caption files attached in chat to compare, in the order they should appear side by side (for example original then translated). Provide either this or urls, never both. | |
| profile | No | Named reading-speed profile applied to both tracks. netflix-adult-20 (default), netflix-children-17, youtube, ebu-tt. |
Output Schema
| Name | Required | Description |
|---|---|---|
| notes | Yes | Notes about incomplete checks or completion under identical limits |
| tracks | Yes | The two tracks in input order |
| winner | No | Track with fewer weighted structural, timing, reading-speed, and line-length problems |
| profile | Yes | Reading-speed and line-length profile applied to the caption checks |
| checkedAt | Yes | ISO timestamp when the caption result was produced |
| comparison | No | Measured differences between the two successfully preflighted tracks |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint/destructiveHint annotations, the description discloses SSRF-guarded and 2 MB-capped fetching, a 60-second execution bound, and a WebVTT parser limitation (payload-less cues are dropped). It also clarifies that only source-file conformance facts are reported, adding significant behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then flows through input modes, behavior, output highlights, exclusions, performance, and limitations. Every sentence carries necessary information with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a comparison tool with two input modes, a profile parameter, an output schema, and meaningful edge cases, the description covers all critical operational details: preflight constraints, alignment method, reported metrics, timing, exclusions, and known parser limitations. An agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already fully documents urls, files, and profile including the strict mutual-exclusion rule. The description adds no parameter-specific meaning beyond what the schema provides, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Compare exactly two caption or subtitle files') and immediately enumerates concrete use cases (original vs. translated, before vs. after timing fix, format migration). It also names the sibling tool caption_check, making the distinction unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly directs agents with 'Use caption_check for a single file,' establishing when this tool applies versus the sibling. It further states what the tool does NOT do ('never translation quality, transcription accuracy, or sync against the video'), which prevents misuse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
- Changed
caption_check4 fields changed- changed
Output schema / properties / cueCount / descriptionPrevious value: -"Number of caption cues retained for checking"New value: +"Number of caption cues retained for checking; payload-less WebVTT cues are omitted by the parser" - changed
Output schema / properties / rawCueCount / descriptionPrevious value: -"Cues parsed before the 20,000-cue cap"New value: +"Cues parsed before the 20,000-cue cap; excludes payload-less WebVTT cues dropped by the parser" - changed
Output schema / properties / timing / descriptionPrevious value: -"Measured cue timing and empty-text findings"New value: +"Measured retained-cue timing and empty-text findings; WebVTT indices refer to retained cues, not original source positions" - changed
Output schema / properties / timing / properties / clean / descriptionPrevious value: -"Checked and no overlapping, negative, out-of-order, or empty cue found, distinguished from not-checked"New value: +"Checked and no overlapping, negative, out-of-order, or empty retained cue found, distinguished from not-checked; omitted payload-less WebVTT cues are not assessed"
1 tool update
- Changed
caption_check4 fields changed- changed
Output schema / properties / cueCount / descriptionPrevious value: -"Number of caption cues retained for checking; payload-less WebVTT cues are omitted by the parser"New value: +"Number of caption cues retained for checking" - changed
Output schema / properties / rawCueCount / descriptionPrevious value: -"Cues parsed before the 20,000-cue cap; excludes payload-less WebVTT cues dropped by the parser"New value: +"Cues parsed before the 20,000-cue cap" - changed
Output schema / properties / timing / descriptionPrevious value: -"Measured retained-cue timing and empty-text findings; WebVTT indices refer to retained cues, not original source positions"New value: +"Measured cue timing and empty-text findings" - changed
Output schema / properties / timing / properties / clean / descriptionPrevious value: -"Checked and no overlapping, negative, out-of-order, or empty retained cue found, distinguished from not-checked; omitted payload-less WebVTT cues are not assessed"New value: +"Checked and no overlapping, negative, out-of-order, or empty cue found, distinguished from not-checked"
1 tool update
- Changed
caption_check4 fields changed- changed
Output schema / properties / cueCount / descriptionPrevious value: -"Number of caption cues retained for checking"New value: +"Number of caption cues retained for checking; payload-less WebVTT cues are omitted by the parser" - changed
Output schema / properties / rawCueCount / descriptionPrevious value: -"Cues parsed before the 20,000-cue cap"New value: +"Cues parsed before the 20,000-cue cap; excludes payload-less WebVTT cues dropped by the parser" - changed
Output schema / properties / timing / descriptionPrevious value: -"Measured cue timing and empty-text findings"New value: +"Measured retained-cue timing and empty-text findings; WebVTT indices refer to retained cues, not original source positions" - changed
Output schema / properties / timing / properties / clean / descriptionPrevious value: -"Checked and no overlapping, negative, out-of-order, or empty cue found, distinguished from not-checked"New value: +"Checked and no overlapping, negative, out-of-order, or empty retained cue found, distinguished from not-checked; omitted payload-less WebVTT cues are not assessed"
2 tool updates
- First observed
caption_check - First observed
caption_compare
Related MCP Connectors
Convert subtitles, transcripts, broadcast captions (SCC/MCC/STL), EDLs, and Premiere files.
Validate video cut plans and parse timed subtitles for Laqta’s local browser video editor.
Convert, clean, retime and validate SRT, WebVTT, ASS/SSA, SBV and TTML subtitles. 8 tools.
Validate JSON, YAML, XML and CSV with exact line/column errors and silent-corruption warnings.
Related MCP Servers
- FlicenseAqualityDmaintenanceEnables processing and translating SRT subtitle files with intelligent conversation detection and context preservation. Supports parsing, validation, chunking of large files, and translation while maintaining precise timing and HTML formatting.6-
- AlicenseBqualityBmaintenanceEnables MCP clients to open, read, edit, verify, and save Aegisub/ASS subtitle documents, covering event lines, styles, timing and QC, karaoke, override tags, vector drawings, clipping, and font metrics. Includes a libass-backed verification layer and byte-faithful round-trips, served over stdio or streamable HTTP.121MIT
- AlicenseAqualityCmaintenanceParses NACHA/ACH files into structured JSON and summaries, performing structural and arithmetic validation with IAT support, enabling automated file inspection and validation through MCP clients.2MIT
- AlicenseAqualityCmaintenanceLoudness compliance verdicts against formal broadcast standards (EBU R128, ATSC A/85). Measures audio/video with ffmpeg and returns a pass/fail verdict with exact deltas and remediation parameters.232 PyPI1MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.