Skip to main content
Glama

Caption File Check

Server Details

Parse WebVTT, SRT, or TTML for conformance, timing, overlaps, line length, and reading speed.

Ownership verified
Status
Healthy
Uptime
100.0% over 21 days
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
powmcp/mcp-server-guide
GitHub Stars
0

TDQS

A4.5/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have clearly distinct purposes: caption_check handles a single file, while caption_compare handles exactly two files. There is no overlap or ambiguity in which tool to call for a given task.

Naming Consistency5/5

Both tools follow the same 'caption_' prefix plus a clear verb: caption_check and caption_compare. The naming pattern is perfectly consistent and immediately conveys the operation.

Tool Count4/5

Two tools is on the lean side of the typical range, but the server's scope is narrowly focused on caption file preflight and comparison. Each tool earns its place and together they cover the core workflows without unnecessary extras.

Completeness5/5

For the stated purpose of caption/subtitle file validation, the surface is complete: one tool for single-file checks and one for side-by-side comparison. There are no obvious missing operations or dead ends within the server's domain.

Available Tools

2 tools
caption_checkA
Read-only
Inspect

Preflight one caption or subtitle file before delivery. Attach a WebVTT (.vtt), SRT (.srt), or TTML/IMSC (.ttml/.xml) caption file in chat or give a publicly fetchable URL and it downloads the artifact under an SSRF-guarded, 2 MB-capped fetch, detects the format by content and the byte encoding and BOM, then parses it with the real spec engine (the WebVTT spec parser, a dedicated SRT parser, or the imscJS TTML/IMSC profile engine) and returns deterministic evidence: pass/fail, cue count, structural spec and IMSC profile-conformance errors, overlapping / negative / out-of-order / empty cues, per-line length, and per-cue and summary reading speed (CPS) against a named platform profile (netflix-adult-20, netflix-children-17, youtube, ebu-tt). It reports only source-file conformance, timing, encoding, and reading-speed facts against a dated profile. It does not judge translation quality, transcription accuracy, or synchronization against the video. An arbitrary XML or text file that is not a caption track returns an explicit unsupported-format error. Fetching and parsing a full-length track can take up to 60 seconds. WebVTT limitation: the parser drops payload-less cues; cue counts, indices, timing checks and comparisons cover retained cues only, not those omitted cues.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoA publicly fetchable HTTP(S) URL of a WebVTT, SRT, or TTML/IMSC caption file, detected by content not extension. Provide either this or an attached file, never both. Query strings are allowed but never logged. Pasted caption bodies, data URLs, FTP/SFTP, and URLs needing credentials, headers, or cookies are not supported. STRICT RULE: Caption check failed — provide either a caption URL or an attached caption file, not both STRICT RULE: Caption check failed — provide a public caption URL or attach the caption file
fileNoA caption file attached in chat (WebVTT, SRT, or TTML/IMSC), detected by content not extension. Provide either this or url, never both.
profileNoNamed reading-speed profile to judge CPS and line length against. netflix-adult-20 (default, 20 CPS) and netflix-children-17 (17 CPS) are spec-backed; youtube and ebu-tt are advisory ~17 CPS heuristics.

Output Schema

ParametersJSON Schema
NameRequiredDescription
bomYesByte-order-mark detected: utf-8, utf-16le, utf-16be, or none
cpsYesPer-cue and summary reading-speed measurements
urlYesThe checked caption source: the submitted URL query- and fragment-stripped, or the attached file's name
passYesTrue when the file passes all checked format, timing, line-length and reading-speed criteria for the selected profile. False can include an advisory readability threshold breach, not just a specification error; this is not a platform-acceptance verdict
timingYesMeasured retained-cue timing and empty-text findings; WebVTT indices refer to retained cues, not original source positions
profileYesReading-speed and line-length profile applied to the caption checks
cueCountYesNumber of caption cues retained for checking; payload-less WebVTT cues are omitted by the parser
encodingYesDetected byte encoding of the caption file
specUsedYesThe dated spec identifier the detected format was checked against
checkedAtYesISO timestamp when the caption result was produced
lineLengthYesPer-line character counts checked against the profile limit
truncationYesMarkers identifying checks truncated by configured limits
rawCueCountYesCues parsed before the 20,000-cue cap; excludes payload-less WebVTT cues dropped by the parser
ttmlProfileNoIMSC/TTML profile designator declared in the file, when present
detectedFormatYeswebvtt, srt, or ttml (covers IMSC), detected by content
structuralErrorsYesSpec-parser and IMSC/TTML profile-conformance findings
structuralErrorCountsYesCounts of structural findings grouped by severity

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description discloses substantial behavioral detail: SSRF-guarded and 2 MB-capped fetching, content-based format detection, the specific parsing engines used, deterministic evidence output, exclusions of quality judgments, explicit unsupported-format error, a 60-second timeout, and a WebVTT limitation about dropped payload-less cues. This far exceeds what annotations alone provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and lengthy, but every sentence contributes meaningful operational detail, from fetch constraints to parser limitations. The main purpose is front-loaded, and the WebVTT caveat is placed at the end where it supplements rather than obscures the core message.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema available, the description does not need to detail return values. It covers input requirements (URL or file), fetch constraints, format detection, parser behavior, timing limits, error handling, scope exclusions, and a known parser limitation. For a tool of this complexity, nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents url, file, and profile. The tool description adds context about format detection and reading-speed profiles, but these are also reflected in the schema parameter descriptions. It does not materially enhance parameter-level understanding beyond the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Preflight one caption or subtitle file before delivery,' a specific verb and resource that clearly identifies the tool's function. It further distinguishes itself from the sibling caption_compare by scoping to a single file and explicitly stating what it does not do (judge translation quality, transcription accuracy, or synchronization).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'before delivery' implies a use case, but there is no explicit when-to-use versus when-not-to-use guidance, nor any mention of the sibling caption_compare as an alternative. An agent must infer that this is for single-file validation rather than comparison.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

caption_compareA
Read-only
Inspect

Compare exactly two caption or subtitle files under identical bounded checks: an original versus its translated track, before versus after a timing fix, or a format migration. Attach two WebVTT, SRT, or TTML/IMSC caption files in chat or give two publicly fetchable URLs and it runs the same SSRF-guarded, 2 MB-capped preflight as caption_check on each, then aligns them by cue index and reports side-by-side pass status, detected-format and encoding differences, the cue-count delta, timing drift and overlap differences, reading-speed (CPS) regressions, structural-error deltas, and a winner only when the measured checks separate the two. Use caption_check for a single file. It reports only source-file conformance, timing, encoding, and reading-speed facts against a dated profile: never translation quality, transcription accuracy, or sync against the video. Fetching and parsing two live files can take up to 60 seconds. WebVTT limitation: the parser drops payload-less cues; cue counts, indices, timing checks and comparisons cover retained cues only, not those omitted cues.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlsNoExactly two public caption URLs to compare, in the order they should appear side by side (for example original then translated). Provide either this or files, never both. STRICT RULE: Caption compare failed — provide either two caption URLs or two attached caption files, not both STRICT RULE: Caption compare failed — provide two public caption URLs or attach two caption files
filesNoExactly two caption files attached in chat to compare, in the order they should appear side by side (for example original then translated). Provide either this or urls, never both.
profileNoNamed reading-speed profile applied to both tracks. netflix-adult-20 (default), netflix-children-17, youtube, ebu-tt.

Output Schema

ParametersJSON Schema
NameRequiredDescription
notesYesNotes about incomplete checks or completion under identical limits
tracksYesThe two tracks in input order
winnerNoTrack with fewer weighted structural, timing, reading-speed, and line-length problems
profileYesReading-speed and line-length profile applied to the caption checks
checkedAtYesISO timestamp when the caption result was produced
comparisonNoMeasured differences between the two successfully preflighted tracks

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint/destructiveHint annotations, the description discloses SSRF-guarded and 2 MB-capped fetching, a 60-second execution bound, and a WebVTT parser limitation (payload-less cues are dropped). It also clarifies that only source-file conformance facts are reported, adding significant behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose, then flows through input modes, behavior, output highlights, exclusions, performance, and limitations. Every sentence carries necessary information with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a comparison tool with two input modes, a profile parameter, an output schema, and meaningful edge cases, the description covers all critical operational details: preflight constraints, alignment method, reported metrics, timing, exclusions, and known parser limitations. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already fully documents urls, files, and profile including the strict mutual-exclusion rule. The description adds no parameter-specific meaning beyond what the schema provides, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Compare exactly two caption or subtitle files') and immediately enumerates concrete use cases (original vs. translated, before vs. after timing fix, format migration). It also names the sibling tool caption_check, making the distinction unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly directs agents with 'Use caption_check for a single file,' establishing when this tool applies versus the sibling. It further states what the tool does NOT do ('never translation quality, transcription accuracy, or sync against the video'), which prevents misuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool update
    • Changedcaption_check4 fields changed
      • changedOutput schema / properties / cueCount / description
        Previous value: -"Number of caption cues retained for checking"New value: +"Number of caption cues retained for checking; payload-less WebVTT cues are omitted by the parser"
      • changedOutput schema / properties / rawCueCount / description
        Previous value: -"Cues parsed before the 20,000-cue cap"New value: +"Cues parsed before the 20,000-cue cap; excludes payload-less WebVTT cues dropped by the parser"
      • changedOutput schema / properties / timing / description
        Previous value: -"Measured cue timing and empty-text findings"New value: +"Measured retained-cue timing and empty-text findings; WebVTT indices refer to retained cues, not original source positions"
      • changedOutput schema / properties / timing / properties / clean / description
        Previous value: -"Checked and no overlapping, negative, out-of-order, or empty cue found, distinguished from not-checked"New value: +"Checked and no overlapping, negative, out-of-order, or empty retained cue found, distinguished from not-checked; omitted payload-less WebVTT cues are not assessed"
  2. 1 tool update
    • Changedcaption_check4 fields changed
      • changedOutput schema / properties / cueCount / description
        Previous value: -"Number of caption cues retained for checking; payload-less WebVTT cues are omitted by the parser"New value: +"Number of caption cues retained for checking"
      • changedOutput schema / properties / rawCueCount / description
        Previous value: -"Cues parsed before the 20,000-cue cap; excludes payload-less WebVTT cues dropped by the parser"New value: +"Cues parsed before the 20,000-cue cap"
      • changedOutput schema / properties / timing / description
        Previous value: -"Measured retained-cue timing and empty-text findings; WebVTT indices refer to retained cues, not original source positions"New value: +"Measured cue timing and empty-text findings"
      • changedOutput schema / properties / timing / properties / clean / description
        Previous value: -"Checked and no overlapping, negative, out-of-order, or empty retained cue found, distinguished from not-checked; omitted payload-less WebVTT cues are not assessed"New value: +"Checked and no overlapping, negative, out-of-order, or empty cue found, distinguished from not-checked"
  3. 1 tool update
    • Changedcaption_check4 fields changed
      • changedOutput schema / properties / cueCount / description
        Previous value: -"Number of caption cues retained for checking"New value: +"Number of caption cues retained for checking; payload-less WebVTT cues are omitted by the parser"
      • changedOutput schema / properties / rawCueCount / description
        Previous value: -"Cues parsed before the 20,000-cue cap"New value: +"Cues parsed before the 20,000-cue cap; excludes payload-less WebVTT cues dropped by the parser"
      • changedOutput schema / properties / timing / description
        Previous value: -"Measured cue timing and empty-text findings"New value: +"Measured retained-cue timing and empty-text findings; WebVTT indices refer to retained cues, not original source positions"
      • changedOutput schema / properties / timing / properties / clean / description
        Previous value: -"Checked and no overlapping, negative, out-of-order, or empty cue found, distinguished from not-checked"New value: +"Checked and no overlapping, negative, out-of-order, or empty retained cue found, distinguished from not-checked; omitted payload-less WebVTT cues are not assessed"
  4. 2 tool updates
    • First observedcaption_check
    • First observedcaption_compare

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    B
    maintenance
    Enables MCP clients to open, read, edit, verify, and save Aegisub/ASS subtitle documents, covering event lines, styles, timing and QC, karaoke, override tags, vector drawings, clipping, and font metrics. Includes a libass-backed verification layer and byte-faithful round-trips, served over stdio or streamable HTTP.
    121
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Parses NACHA/ACH files into structured JSON and summaries, performing structural and arithmetic validation with IAT support, enabling automated file inspection and validation through MCP clients.
    2
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Loudness compliance verdicts against formal broadcast standards (EBU R128, ATSC A/85). Measures audio/video with ffmpeg and returns a pass/fail verdict with exact deltas and remediation parameters.
    2
    32 PyPI
    1
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.