Skip to main content
Glama

Caption File Check

caption_check

Read-only

Preflight one caption or subtitle file before delivery. Attach a WebVTT (.vtt), SRT (.srt), or TTML/IMSC (.ttml/.xml) caption file in chat or give a publicly fetchable URL and it downloads the artifact under an SSRF-guarded, 2 MB-capped fetch, detects the format by content and the byte encoding and BOM, then parses it with the real spec engine (the WebVTT spec parser, a dedicated SRT parser, or the imscJS TTML/IMSC profile engine) and returns deterministic evidence: pass/fail, cue count, structural spec and IMSC profile-conformance errors, overlapping / negative / out-of-order / empty cues, per-line length, and per-cue and summary reading speed (CPS) against a named platform profile (netflix-adult-20, netflix-children-17, youtube, ebu-tt). It reports only source-file conformance, timing, encoding, and reading-speed facts against a dated profile. It does not judge translation quality, transcription accuracy, or synchronization against the video. An arbitrary XML or text file that is not a caption track returns an explicit unsupported-format error. Fetching and parsing a full-length track can take up to 60 seconds. WebVTT limitation: the parser drops payload-less cues; cue counts, indices, timing checks and comparisons cover retained cues only, not those omitted cues.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlNoA publicly fetchable HTTP(S) URL of a WebVTT, SRT, or TTML/IMSC caption file, detected by content not extension. Provide either this or an attached file, never both. Query strings are allowed but never logged. Pasted caption bodies, data URLs, FTP/SFTP, and URLs needing credentials, headers, or cookies are not supported. STRICT RULE: Caption check failed — provide either a caption URL or an attached caption file, not both STRICT RULE: Caption check failed — provide a public caption URL or attach the caption file
fileNoA caption file attached in chat (WebVTT, SRT, or TTML/IMSC), detected by content not extension. Provide either this or url, never both.
profileNoNamed reading-speed profile to judge CPS and line length against. netflix-adult-20 (default, 20 CPS) and netflix-children-17 (17 CPS) are spec-backed; youtube and ebu-tt are advisory ~17 CPS heuristics.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
bomYesByte-order-mark detected: utf-8, utf-16le, utf-16be, or none
cpsYesPer-cue and summary reading-speed measurements
urlYesThe checked caption source: the submitted URL query- and fragment-stripped, or the attached file's name
passYesTrue when the file passes all checked format, timing, line-length and reading-speed criteria for the selected profile. False can include an advisory readability threshold breach, not just a specification error; this is not a platform-acceptance verdict
timingYesMeasured retained-cue timing and empty-text findings; WebVTT indices refer to retained cues, not original source positions
profileYesReading-speed and line-length profile applied to the caption checks
cueCountYesNumber of caption cues retained for checking; payload-less WebVTT cues are omitted by the parser
encodingYesDetected byte encoding of the caption file
specUsedYesThe dated spec identifier the detected format was checked against
checkedAtYesISO timestamp when the caption result was produced
lineLengthYesPer-line character counts checked against the profile limit
truncationYesMarkers identifying checks truncated by configured limits
rawCueCountYesCues parsed before the 20,000-cue cap; excludes payload-less WebVTT cues dropped by the parser
ttmlProfileNoIMSC/TTML profile designator declared in the file, when present
detectedFormatYeswebvtt, srt, or ttml (covers IMSC), detected by content
structuralErrorsYesSpec-parser and IMSC/TTML profile-conformance findings
structuralErrorCountsYesCounts of structural findings grouped by severity

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed4 schema fields changed
    • changedOutput schema / properties / cueCount / description
      Previous value: -"Number of caption cues retained for checking"New value: +"Number of caption cues retained for checking; payload-less WebVTT cues are omitted by the parser"
    • changedOutput schema / properties / rawCueCount / description
      Previous value: -"Cues parsed before the 20,000-cue cap"New value: +"Cues parsed before the 20,000-cue cap; excludes payload-less WebVTT cues dropped by the parser"
    • changedOutput schema / properties / timing / description
      Previous value: -"Measured cue timing and empty-text findings"New value: +"Measured retained-cue timing and empty-text findings; WebVTT indices refer to retained cues, not original source positions"
    • changedOutput schema / properties / timing / properties / clean / description
      Previous value: -"Checked and no overlapping, negative, out-of-order, or empty cue found, distinguished from not-checked"New value: +"Checked and no overlapping, negative, out-of-order, or empty retained cue found, distinguished from not-checked; omitted payload-less WebVTT cues are not assessed"
  2. Changed4 schema fields changed
    • changedOutput schema / properties / cueCount / description
      Previous value: -"Number of caption cues retained for checking; payload-less WebVTT cues are omitted by the parser"New value: +"Number of caption cues retained for checking"
    • changedOutput schema / properties / rawCueCount / description
      Previous value: -"Cues parsed before the 20,000-cue cap; excludes payload-less WebVTT cues dropped by the parser"New value: +"Cues parsed before the 20,000-cue cap"
    • changedOutput schema / properties / timing / description
      Previous value: -"Measured retained-cue timing and empty-text findings; WebVTT indices refer to retained cues, not original source positions"New value: +"Measured cue timing and empty-text findings"
    • changedOutput schema / properties / timing / properties / clean / description
      Previous value: -"Checked and no overlapping, negative, out-of-order, or empty retained cue found, distinguished from not-checked; omitted payload-less WebVTT cues are not assessed"New value: +"Checked and no overlapping, negative, out-of-order, or empty cue found, distinguished from not-checked"
  3. Changed4 schema fields changed
    • changedOutput schema / properties / cueCount / description
      Previous value: -"Number of caption cues retained for checking"New value: +"Number of caption cues retained for checking; payload-less WebVTT cues are omitted by the parser"
    • changedOutput schema / properties / rawCueCount / description
      Previous value: -"Cues parsed before the 20,000-cue cap"New value: +"Cues parsed before the 20,000-cue cap; excludes payload-less WebVTT cues dropped by the parser"
    • changedOutput schema / properties / timing / description
      Previous value: -"Measured cue timing and empty-text findings"New value: +"Measured retained-cue timing and empty-text findings; WebVTT indices refer to retained cues, not original source positions"
    • changedOutput schema / properties / timing / properties / clean / description
      Previous value: -"Checked and no overlapping, negative, out-of-order, or empty cue found, distinguished from not-checked"New value: +"Checked and no overlapping, negative, out-of-order, or empty retained cue found, distinguished from not-checked; omitted payload-less WebVTT cues are not assessed"
  4. First observed

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description discloses substantial behavioral detail: SSRF-guarded and 2 MB-capped fetching, content-based format detection, the specific parsing engines used, deterministic evidence output, exclusions of quality judgments, explicit unsupported-format error, a 60-second timeout, and a WebVTT limitation about dropped payload-less cues. This far exceeds what annotations alone provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and lengthy, but every sentence contributes meaningful operational detail, from fetch constraints to parser limitations. The main purpose is front-loaded, and the WebVTT caveat is placed at the end where it supplements rather than obscures the core message.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema available, the description does not need to detail return values. It covers input requirements (URL or file), fetch constraints, format detection, parser behavior, timing limits, error handling, scope exclusions, and a known parser limitation. For a tool of this complexity, nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents url, file, and profile. The tool description adds context about format detection and reading-speed profiles, but these are also reflected in the schema parameter descriptions. It does not materially enhance parameter-level understanding beyond the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Preflight one caption or subtitle file before delivery,' a specific verb and resource that clearly identifies the tool's function. It further distinguishes itself from the sibling caption_compare by scoping to a single file and explicitly stating what it does not do (judge translation quality, transcription accuracy, or synchronization).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'before delivery' implies a use case, but there is no explicit when-to-use versus when-not-to-use guidance, nor any mention of the sibling caption_compare as an alternative. An agent must infer that this is for single-file validation rather than comparison.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.