Skip to main content
Glama

verify

Read-onlyIdempotent

Transcribe a finished render and diff it against the timeline to catch surviving retakes and dropped words.

Instructions

Transcribe a finished render and diff it against what the timeline says.

Run this after rendering, before calling an edit done. It transcribes the render with whisper and compares that word sequence to the one the timeline should play, which is the only check that catches a retake still in the picture: whisper collapses an immediate repeat into a single utterance, so a doubled phrase can be invisible in the source transcript and still be in the render.

Read repeated first — an entry there is a phrase the render plays more times than the timeline expects, i.e. a surviving retake, with the heard word index to look at. dropped is the opposite: words the timeline expects that the render never says, usually a cut that reached too far.

A clean single-pass result is not proof. This check has a known blind spot: the render's transcript is itself one whisper pass, which collapses a repeat the same way the source transcript did — three retakes survived a correct run of it on a real video. Set windowed=True to transcribe in short overlapping windows instead, which is what found them. It costs one whisper run over 2x the audio and uses a deliberately smaller model, so run the default first and escalate to it before calling an edit finished.

loud_gaps comes back either way and trusts no transcript: it measures the render's own energy and reports holes in the heard word map that hold sound anyway. An entry is a place to listen, not a verdict — a music bed or an attenuated noise can produce one. Read speech_db/threshold_db beside it.

similarity around 0.97 is normal on a clean render — whisper spells its own output differently on a second pass ("whodunit" / "who done it", "4" / "four"). Treat it as triage; diff is the artifact. Transcription takes minutes on a long render, and the result is cached under cache/verify/ and reported as heard_transcript — pass it back as transcript_path to re-diff without re-transcribing.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
pathNoThe project directory to act on. Omit it — the usual case — when this server is bound to a project (started as `proofcut -C DIR mcp`, or inside a project; `ping` says which): it then resolves to that one bound project, a relative path resolves against it, and a path outside it is refused by name. Unbound, `path` is the whole address and omitting it refuses rather than guessing.
modelNoThe whisper model for the single-pass transcription.
renderYesThe finished render to transcribe and diff against the timeline. The expected words include a sound's or an inset's own when its clip has a transcript (`placed_audio`) — a narrator take placed as a sound is the film's voice. A voice sound (`ducks`) with no transcript is listed in `voice_sounds_untranscribed` and not checked: transcribe it first.
windowNoLength of each window in the windowed pass, in seconds.
clip_idNoDiff against one transcript's expected words rather than all of them.
overlapNoHow far each window overlaps the one before, in seconds.
languageNoForce a language code for it.
windowedNoTranscribe in short overlapping windows instead of one pass. **A clean single-pass result is not proof** — one pass collapses an immediate repeat the same way the source transcript did, and three surviving retakes passed a correct single-pass run on a real video. It costs a run over twice the audio and a smaller model.
transcript_pathNoAn existing transcription of `render` — what a previous run cached and reported as `heard_transcript`. Pass it back to re-diff without spending the minutes again.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv0.38.0
    • changedInput schema / properties / render / description
      Previous value: -"The finished render to transcribe and diff against the timeline."New value: +"The finished render to transcribe and diff against the timeline. The expected words include a sound's or an inset's own when its clip has a transcript (`placed_audio`) — a narrator take placed as a sound is the film's voice. A voice sound (`ducks`) with no transcript is listed in `voice_sounds_untranscribed` and not checked: transcribe it first."
  2. Changed29 schema fields changedv0.25.0
    • removedInput schema / properties / clip_id / anyOf
      Removed value: -[
      -  {
      -    "type": "string"
      -  },
      -  {
      -    "type": "null"
      -  }
      -]
    • addedInput schema / properties / clip_id / description
      Added value: +"Diff against one transcript's expected words rather than all of them."
    • removedInput schema / properties / clip_id / title
      Removed value: -"Clip Id"
    • addedInput schema / properties / clip_id / type
      Added value: +[
      +  "string",
      +  "null"
      +]
    • removedInput schema / properties / language / anyOf
      Removed value: -[
      -  {
      -    "type": "string"
      -  },
      -  {
      -    "type": "null"
      -  }
      -]
    • addedInput schema / properties / language / description
      Added value: +"Force a language code for it."
    • removedInput schema / properties / language / title
      Removed value: -"Language"
    • addedInput schema / properties / language / type
      Added value: +[
      +  "string",
      +  "null"
      +]
    • removedInput schema / properties / model / anyOf
      Removed value: -[
      -  {
      -    "type": "string"
      -  },
      -  {
      -    "type": "null"
      -  }
      -]
    • addedInput schema / properties / model / description
      Added value: +"The whisper model for the single-pass transcription."
    • removedInput schema / properties / model / title
      Removed value: -"Model"
    • addedInput schema / properties / model / type
      Added value: +[
      +  "string",
      +  "null"
      +]
    • addedInput schema / properties / overlap / description
      Added value: +"How far each window overlaps the one before, in seconds."
    • removedInput schema / properties / overlap / title
      Removed value: -"Overlap"
    • removedInput schema / properties / path / anyOf
      Removed value: -[
      -  {
      -    "type": "string"
      -  },
      -  {
      -    "type": "null"
      -  }
      -]
    • addedInput schema / properties / path / description
      Added value: +"The project directory to act on. Omit it — the usual case — when this server is bound to a project (started as `proofcut -C DIR mcp`, or inside a project; `ping` says which): it then resolves to that one bound project, a relative path resolves against it, and a path outside it is refused by name. Unbound, `path` is the whole address and omitting it refuses rather than guessing."
    • removedInput schema / properties / path / title
      Removed value: -"Path"
    • addedInput schema / properties / path / type
      Added value: +[
      +  "string",
      +  "null"
      +]
    • addedInput schema / properties / render / description
      Added value: +"The finished render to transcribe and diff against the timeline."
    • removedInput schema / properties / render / title
      Removed value: -"Render"
    • removedInput schema / properties / transcript_path / anyOf
      Removed value: -[
      -  {
      -    "type": "string"
      -  },
      -  {
      -    "type": "null"
      -  }
      -]
    • addedInput schema / properties / transcript_path / description
      Added value: +"An existing transcription of `render` — what a previous run cached and reported as `heard_transcript`. Pass it back to re-diff without spending the minutes again."
    • removedInput schema / properties / transcript_path / title
      Removed value: -"Transcript Path"
    • addedInput schema / properties / transcript_path / type
      Added value: +[
      +  "string",
      +  "null"
      +]
    • addedInput schema / properties / window / description
      Added value: +"Length of each window in the windowed pass, in seconds."
    • removedInput schema / properties / window / title
      Removed value: -"Window"
    • addedInput schema / properties / windowed / description
      Added value: +"Transcribe in short overlapping windows instead of one pass. **A clean single-pass result is not proof** — one pass collapses an immediate repeat the same way the source transcript did, and three surviving retakes passed a correct single-pass run on a real video. It costs a run over twice the audio and a smaller model."
    • removedInput schema / properties / windowed / title
      Removed value: -"Windowed"
    • removedInput schema / title
      Removed value: -"verifyArguments"
  3. First observedv0.24.0

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and idempotentHint=true, and the description richly complements them: it discloses the caching behavior (cache/verify/, reported as heard_transcript), the cost profile (minutes per run, windowed costs 2x audio with a smaller model), the known blind spot (single-pass collapses repeats, three retakes survived a correct run), and interpretive context (similarity ~0.97 is normal; loud_gaps is a listen-point, not a verdict). This goes well beyond what annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with front-loaded purpose, bolded warnings, and clearly separated paragraphs, and nearly every sentence carries behavioral value. However, it is long for an MCP description and carries some redundancy — the windowed blind-spot explanation appears in both the description and the windowed parameter's schema text, and the loud_gaps triage guidance repeats the similarity triage point. Tightening would help.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, complex workflow tool, the description is essentially complete: it covers the triggering workflow, escalation strategy, output interpretation (repeated, dropped, loud_gaps, similarity, diff), the blind-spot caveat, and the caching/transcript_path round-trip. An output schema exists, so return values need not be spelled out, and the description still explains the key result fields an agent must act on.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds real value on top: it explains the purpose of transcript_path via the caching mechanism, clarifies that render includes placed_audio transcripts and untranscribed voice sounds, and reinforces windowed's escalation semantics. Not every parameter gets prose (model, window, overlap, language are left to the schema), but the schema itself documents them fully.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a precise verb+resource+scope statement: 'Transcribe a finished render and diff it against what the timeline says.' It states the exact trigger condition (after rendering, before edit done) and the specific failure it catches (a surviving retake), which clearly distinguishes it from siblings like transcribe and transcript_checks without needing to open their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to run it ('Run this after rendering, before calling an edit done') and gives a clear escalation path ('run the default first and escalate to it before calling an edit finished'). It also warns against false confidence from a clean result. It does not name the sibling tools (transcribe, hear, transcript_checks) that an agent might confuse it with, but the workflow positioning is clear enough to route the agent correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools