casefile-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@casefile-mcpMake a walkthrough from recordings/login-test.mp4, highlight errors, blur hostname"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Casefile
Turn a screen recording of a test session into documentation: a narrated, captioned, highlighted walkthrough video plus a subtitle file, driven by any AI assistant through MCP.
Casefile runs locally, works offline once the models are downloaded, and is free and open source (MIT).
The problem
People record a test run, a bug reproduction or a configuration session, and then the recording sits in a folder. Nobody writes it up. A raw 3-minute capture is hard to review: nothing is labelled, the important value is on screen for two seconds, and there may be a hostname you shouldn't share.
Casefile lets an AI assistant turn that recording into something a colleague can follow. The result has chapters, narration, captions, highlight boxes on the fields that matter, blurred secrets and bad frames removed. The assistant does not watch the video. It reads the video with cheap text tools (OCR, chart extraction) and writes one JSON edit spec. A local renderer does the rest.
Related MCP server: screencast
What works today vs. roadmap
Capability | Status |
Narrated, captioned 1080p walkthrough video (H.264/AAC) + | ✅ works (v0.1) |
Highlight boxes, spotlight, blur, camera zoom/pan, chart annotations, intro/outro cards | ✅ works |
Offline OCR text location ( | ✅ works |
Offline TTS narration (Kokoro), cached per line | ✅ works |
Bad-frame repair (black frames, glitches, popups) | ✅ works |
MCP server (9 tools) + identical CLI | ✅ works |
Configurable brand (logo, accent, ink colours) | ✅ works |
Written step-by-step docs (Markdown with screenshots, Confluence export) | 🗺️ planned (v0.2) |
Transcribing the narrator's own voice from the recording (offline STT) | 🗺️ planned (v0.3) |
Jira integration (read test-case steps, attach outputs to the issue) | 🗺️ planned (v0.4) |
Hosted / team tier | 🗺️ planned, not built. The core stays MIT and free |
See ROADMAP.md.
60-second quickstart
1. System dependencies
ffmpeg 6+ (
ffmpegandffprobeon PATH). Windows:winget install Gyan.FFmpeg. Debian/Ubuntu:apt install ffmpeg. macOS:brew install ffmpeg.Tesseract 5. Windows:
winget install UB-Mannheim.TesseractOCR(auto-detected inC:\Program Files\Tesseract-OCR, so PATH is not needed). Debian/Ubuntu:apt install tesseract-ocr. macOS:brew install tesseract.
2. Install Casefile
Casefile is not on PyPI yet. Install it from a checkout. The distribution name will be casefile-mcp.
git clone https://github.com/asrgrimmi-cyber/casefile && cd casefile
uv tool install . # puts `casefile` and `casefile-mcp` on PATH
casefile doctor # checks ffmpeg/tesseract/fonts, downloads the Kokoro TTS model + voices (~340 MB, once)Models and caches live in ~/.casefile/. Set CASEFILE_HOME to move them.
3. Connect it to your assistant
Claude Code
claude mcp add casefile -- casefile-mcp
# optional sandbox: only files under this folder can be read or written
claude mcp add casefile -e CASEFILE_WORKDIR="$PWD" -- casefile-mcpClaude Desktop: edit claude_desktop_config.json (Windows %APPDATA%\Claude\, macOS
~/Library/Application Support/Claude/) and restart Claude Desktop.
{
"mcpServers": {
"casefile": {
"command": "casefile-mcp",
"env": {"CASEFILE_WORKDIR": "C:\\Users\\me\\Videos\\tests"}
}
}
}If casefile-mcp is not on PATH, use the full path, for example C:\\path\\to\\casefile\\.venv\\Scripts\\casefile-mcp.exe.
Hermes Agent: add this to config.yaml in your Hermes home, then run hermes mcp test casefile.
mcp_servers:
casefile:
command: casefile-mcp
env:
CASEFILE_WORKDIR: C:/Users/me/Videos/testsAny other MCP client: run casefile-mcp over stdio, or casefile-mcp --http 127.0.0.1:8765 for streamable HTTP at /mcp.
4. Give the assistant the skill (recommended)
skill/casefile/SKILL.md holds the workflow, tool budgets, a spec cheat-sheet and a review
checklist. Install it as a Claude/Hermes skill, or paste it into your project instructions.
Example prompt
Use the casefile tools.
recordings/login-test.mp4is a recording of test case TC-042 (login with an expired password). Make a walkthrough of about a minute: intro card, one chapter per screen, highlight the error message and the "password expired" status, blur the hostname in the address bar, and end with a summary card. Render a preview first and show me the plan.
The assistant will probe the video, locate the fields with OCR, write edit.json, plan, render a preview, check it
against the checklist and then render the final MP4 and .srt.
How it works
recording.mp4
│
▼
┌──────────┐ scenes, bad frames, ┌──────────────────────────────┐
│ probe │── popups, layout changes ▶│ AI assistant (any MCP client)│
└──────────┘ │ reads text, never watches │
┌──────────────────────────────┐ │ the video │
│ find_text / ocr_region / │◀──────│ │
│ track_text / chart_extract / │ boxes │ writes ONE edit spec (JSON) │
│ contact_sheet (rare, images) │──────▶│ │
└──────────────────────────────┘ └──────────────┬───────────────┘
│ edit.json
▼
┌──────────────┐ ┌────────────────────┐ ┌─────────────────────┐
│ plan │────▶│ render preview │────▶│ render final │
│ (TTS timing, │ │ 960x540 + 4 thumbs │ │ 1920x1080 MP4 + SRT │
│ validation) │ └────────────────────┘ └─────────────────────┘
└──────────────┘One core library, with the MCP server and the CLI as thin wrappers that return the same compact JSON. Details are in docs/HOW_IT_WORKS.md. The spec format is in docs/spec.md, and the tools are in docs/tools.md.
CLI
Every MCP tool has a CLI twin that prints the same JSON:
casefile probe demo.mp4
casefile find-text demo.mp4 3.0 "Band" --row
casefile plan examples/fixture.json --no-synth
casefile render examples/fixture.json --preview
casefile render examples/fixture.jsonexamples/fixture.json renders against the synthetic test video, which python tests/fixtures/make_fixture.py builds.
Branding
Cards and highlights use spec.brand: {"logo": "path/to/logo.png", "accent": "#4F46E5", "ink": "#111827"}.
With no logo, cards show a small "Casefile" text wordmark. The defaults are indigo #4F46E5 and ink #111827, and the
font is Poppins (bundled).
How well it works
Measured numbers, including the weak spots, are in docs/RESULTS.md. In short, the synthetic-fixture gates are exact or within 1–2 px. One real 2:35 1440p recording became a 1:17 walkthrough. On that recording the first review pass found 3 misplaced highlights, which were then fixed.
Limitations
Output today is video + subtitles only. Written docs, voice transcription and Jira are on the roadmap, not built.
Tested on Windows 11 so far. CI runs the fast tests on Ubuntu and Windows. macOS is untested.
Not validated at scale: a synthetic fixture plus a small number of real recordings.
The first
probeof a 1440p recording takes minutes (frame extraction), andtrack_textcosts ~2 s per frame at 1440p.OCR is noisy on full 1440p frames. Pass a
region.Scene detection over-reports on live-updating charts.
Narration is English (Kokoro voices). Other languages are not tested.
The assistant still needs to review the preview. Terminals that scroll mid-shot can move text out from under a highlight.
Contributing
See CONTRIBUTING.md. How the project was built with an orchestrator agent and parallel sub-agents is in docs/BUILD_STORY.md.
License
MIT, see LICENSE. The bundled fonts keep their own licenses: Poppins (SIL OFL 1.1) and DejaVu. The Kokoro
model files are downloaded by casefile doctor from their upstream release and are not included in this repository.
Available Tools
9 toolschart_extractB
Read coloured curves in axis units -> {series:{name:{points,min,max,zero_crossings,coverage}}}.
| Name | Required | Description | Default |
|---|---|---|---|
| t | Yes | time in s | |
| video | Yes | video path (relative to workdir ok) | |
| series | Yes | [{name,rgb:[r,g,b],tol?:40}] | |
| samples | No | ||
| x_ticks | Yes | [[px,value],..] >=2 | |
| y_ticks | Yes | [[py,value],..] >=2 | |
| plot_region | Yes | [x0,y0,x1,y1] plot area px |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it partially discharges it by disclosing the return structure ({series:{name:{points,min,max,zero_crossings,coverage}}}) which is not in any output schema. However, it says nothing about failure modes (e.g. too few ticks, unmatched colour), what 'coverage' or 'tol' mean operationally, or any resource/permission constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence that front-loads the action and the output shape. The arrow-and-brace notation is dense but information-bearing; nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool doing pixel-to-axis calibration, the description is thin: it covers the return shape but not the coordinate/calibration model or the colour-matching semantics. It is minimally adequate given the 86% schema coverage, but an agent would likely need trial and error.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 86%, so the schema already documents most parameters including the terse [px,value] tick format. The description adds nothing about the tricky parameters (plot_region, series rgb/tol, samples default) beyond implying series names surface in the output, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read coloured curves in axis units') and even sketches the return shape. An agent can tell it apart from siblings like ocr_region or track_text, though it does not explicitly name a sibling or state the conversion pipeline.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites (e.g. video must already be loaded/known path), and no mention of alternatives among the sibling tools. The agent must infer applicability purely from the name and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contact_sheetC
Look at frames: one PNG grid with burned-in timestamps (+ text with times).
| Name | Required | Description | Default |
|---|---|---|---|
| n | No | ||
| t0 | No | ||
| t1 | No | ||
| crop | No | [x0,y0,x1,y1] source px; null=full frame | |
| times | No | exact times; else t0/t1/n | |
| video | Yes | video path (relative to workdir ok) | |
| max_width | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It does disclose the output shape (one PNG grid with burned-in timestamps, plus text with times), which is genuinely useful behavioral context. However, it says nothing about permissions, cost of processing a video, or how frames are chosen, leaving significant gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single short, front-loaded sentence with no wasted padding. The cryptic parenthetical '(+ text with times)' is slightly awkward but the overall size is appropriate for the content given.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter video-processing tool with no annotations, no output schema, and 43% parameter coverage, this description is far too sparse. It gives a hint of the return artifact but omits the meaning of most parameters and any usage context, so an agent cannot reliably invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 43%: n, t0, t1 and max_width have no schema descriptions, and the description does not compensate by explaining what they mean (frame count, start/end time, output width). The mention of 'timestamps'/'times' loosely touches the times param but adds no syntax or format detail, so the undocumented parameters remain ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description hints at the resource (frames from a video) and the output artifact (a single PNG grid with burned-in timestamps plus accompanying time text), which is more than a tautology. But the verb 'Look at' is vague and doesn't state that the tool generates/extracts a contact sheet, nor does it differentiate from siblings like probe, find_text, or chart_extract. Purpose is only partially conveyed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as probe, ocr_region, or track_text, and no mention of prerequisites or exclusions. The agent must infer that this is the frame-sampling tool purely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_textB
Locate text on the frame at t -> {t,matches:[{text,box,conf,row_box?}]}. Use boxes for overlays.
| Name | Required | Description | Default |
|---|---|---|---|
| t | Yes | time in s | |
| row | No | also return row_box of the form/table row | |
| fuzzy | No | ||
| limit | No | ||
| query | Yes | text (multi-word ok) or 're:<regex>' | |
| video | Yes | video path (relative to workdir ok) | |
| region | No | [x0,y0,x1,y1] source px; null=full frame |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries full behavioral burden. It usefully discloses the return structure (t, matches with text/box/conf/row_box), but says nothing about fuzzy-match behavior, default limit, region semantics, or performance/cost. Partial disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the action and return type, followed by the one practical usage tip. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, but the description compensates with an inline return shape, which is genuinely helpful for a 7-param tool. However, with no annotations and thin usage/routing guidance, the definition is not fully complete for correct selection.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 71%, so most parameters are already documented there. The description adds nothing about fuzzy, limit, or region beyond the schema, though it does show row_box in the return when row=true. Baseline 3 for moderate schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (locate text on the frame at time t) and even sketches the return shape. It does not distinguish itself from the sibling ocr_region, which an agent could reasonably pick instead, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus ocr_region or track_text, and no prerequisites (e.g. video must exist, region vs full frame). The 'Use boxes for overlays' note hints at a downstream use but not tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ocr_regionB
Read all text lines in region at t -> {t,lines:[{text,box}]} in reading order.
| Name | Required | Description | Default |
|---|---|---|---|
| t | Yes | time in s | |
| video | Yes | video path (relative to workdir ok) | |
| region | No | [x0,y0,x1,y1] source px; null=full frame |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses the return structure and reading order, which is useful, but does not explicitly state that it is read-only, has no side effects, or describe any error conditions or performance characteristics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with zero waste. It is front-loaded with the core action and compactly includes the output shape and ordering constraint, which is valuable given there is no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple OCR tool with a rich input schema and no output schema, the description provides the essential return format inline. It covers core invocation needs but leaves some behavioral aspects (e.g., error handling, read-only confirmation) implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters. The description only restates 'region at t' without adding any syntax or format details beyond what the schema provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: 'Read all text lines in region at t'. It also defines the output format. However, it does not distinguish this tool from siblings like find_text or track_text, which likely also deal with text extraction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as find_text or track_text. The description only implies usage through the 'at t' phrasing, but does not explicitly state when or when not to choose this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
planA
Validate + time a spec -> {total,frames,shots:[{i,chapter,T0,dur,lines}],warnings}. Run before render.
| Name | Required | Description | Default |
|---|---|---|---|
| spec | Yes | edit spec object, JSON text, or .json path | |
| synth | No | real TTS durations (else estimate) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It discloses the return shape but says nothing about side effects, permissions, idempotency, or whether 'synth=true' triggers real TTS calls with associated cost or latency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences with zero waste; purpose and output shape are front-loaded, and the ordering instruction follows immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter planning tool with no output schema, the description does supply the return structure and the sequencing rule. However, it leaves gaps around the meaning of returned fields, potential side effects, and whether repeated calls are safe.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters fully. The description adds no extra meaning beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States specific verbs (Validate, time) and the resource (spec), and explicitly positions itself relative to a sibling ('Run before render'), so an agent can tell it apart from render and tts without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear when-to-use cue ('Run before render'), which is more than implied usage. It does not state when not to use the tool or name a specific alternative beyond the implicit sibling reference, so it falls short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
probeC
Video facts: duration,fps,size,frames,audio_db,scenes[t],bad_frames[{frame,t,kind}],popups[t],layout_changes[t].
| Name | Required | Description | Default |
|---|---|---|---|
| video | Yes | video path (relative to workdir ok) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It lists returned fields, which suggests a read-only probe, but says nothing about permissions, side effects, expected failure modes, or whether the operation is non-mutating. Return fields alone are not behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single dense line with no filler and the resource is front-loaded before the field list. However, the unspaced comma-run makes the output fields harder to scan than a short structured list would be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must carry return-value meaning. It enumerates top-level fields but leaves key semantics unexplained, such as what the '[t]' notation means, what values 'kind' in bad_frames can take, and how scenes or popups are structured.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter is fully documented in the schema as 'video path (relative to workdir ok)'. The description adds no syntax, format, or constraint details beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (video) and enumerates returned facts, but supplies no verb like 'extract' or 'inspect', so the action is only implied. It is enough to distinguish this from text-oriented siblings such as find_text and ocr_region, but does not crisply state what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no condition under which to prefer this over siblings like contact_sheet or chart_extract, and no prerequisites. The agent must infer the triggering scenario entirely on its own.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
renderB
Render spec to MP4 (+SRT) -> {path,duration,frames,size_mb,srt,thumbs}.
| Name | Required | Description | Default |
|---|---|---|---|
| out | No | output .mp4 path | |
| spec | Yes | edit spec object, JSON text, or .json path | |
| thumbs | No | also return one image tiling 4 thumbnails | |
| preview | No | 960x540@15fps fast draft |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the full behavioral burden. It discloses the return object shape and mentions optional SRT generation, but it does not explain side effects, file writes, overwrite behavior, permissions, or runtime characteristics of a render operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact line with no wasted words. It front-loads the verb and artifact, then lists the return keys, making it highly efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter render tool with no annotations and no output schema, the description is only minimally complete. It partially compensates for the missing output schema by listing return fields, but it omits critical behavioral and usage context needed for confident invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents every parameter. The description adds little beyond restating that it renders a spec and outputs MP4/SRT, which is baseline adequacy rather than added semantic value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: render a spec to MP4 and optionally SRT. The output shape is also given, so an agent can tell what the tool produces. However, it does not explicitly differentiate this tool from any sibling, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance about when to use this tool, when not to use it, or which alternatives exist among the sibling tools. It simply states what the tool does, leaving usage context entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
track_textC
When/where text is visible in [t0,t1] -> {query,spans:[{t0,t1,box_start,box_end,moved}]}.
| Name | Required | Description | Default |
|---|---|---|---|
| t0 | Yes | ||
| t1 | Yes | ||
| step | No | sample interval s | |
| fuzzy | No | ||
| query | Yes | ||
| video | Yes | video path (relative to workdir ok) | |
| region | No | [x0,y0,x1,y1] source px; null=full frame |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does usefully sketch the return payload ({query, spans:[{t0,t1,box_start,box_end,moved}]}), which is real information given there is no output schema. But it says nothing about cost, sampling behavior, permissions, or what 'moved' means.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single line with zero padding, so nothing is wasted and the time-range constraint is front-loaded. But the compression is excessive enough that clarity is lost, which is under-specification rather than true conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 7 parameters, no annotations, no output schema, and 43% schema coverage, the description is the only source of behavioral and parameter context and it delivers only a partial return sketch. An agent cannot reliably call this tool correctly without guessing at step, fuzzy, region, and the meaning of the returned fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 43%, so the description must compensate for step, fuzzy, region, video, and query. It only reinforces t0/t1 and query, adding no format, range, or default details for the undocumented parameters. The gap is left unfilled.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The notation 'When/where text is visible in [t0,t1]' conveys a temporal+spatial text tracking operation, which is distinguishable from find_text and ocr_region. However it is written in cryptic pseudo-notation rather than a plain verb+resource statement, so an agent must decode it. The '->' return sketch hints at the resource but never names it explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose this over find_text, ocr_region, or chart_extract, all of which involve text detection. No exclusions, prerequisites, or context of use are given. The agent must infer selection from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ttsC
Synthesise narration (cached WAVs) -> {durations:{id:s},cached:[ids]}.
| Name | Required | Description | Default |
|---|---|---|---|
| lines | Yes | [{id,text}] | |
| speed | No | ||
| voice | No | af_heart | |
| lexicon | No | word->spoken form |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose one real behavioral trait — outputs are cached WAVs — and sketches the return shape, which is useful. However, it says nothing about permissions, whether files are written to disk, latency, or how the cache is invalidated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single short line with no wasted words and the core action is front-loaded. However, the telegraphic arrow-and-brace notation is cryptic and forces the agent to decode rather than read, trading brevity for clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully encodes the return shape ({durations, cached}), which offsets some of the burden. But for a four-parameter tool with half its parameters undocumented, the definition leaves too much unspecified to be considered complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 50% schema coverage, the description should compensate for undocumented parameters, but it mentions none of the four (lines, speed, voice, lexicon). The trailing object notation describes the response, not the inputs, so it adds zero meaning beyond the schema fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Synthesise narration') and notes the artifact produced (cached WAVs), so an agent can tell this is a text-to-speech tool. It does not differentiate from siblings, though the sibling list (probe, ocr_region, chart_extract, etc.) makes it reasonably distinct by domain.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of prerequisites, and no alternatives named. The parenthetical '(cached WAVs)' hints that returns may be reused from cache but gives no condition under which that matters.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.1.0- First observed
chart_extract - First observed
contact_sheet - First observed
find_text - First observed
ocr_region - First observed
plan - First observed
probe - First observed
render - First observed
track_text - First observed
tts
TDQS
Scored across 9 tools
Most tools are clearly distinct: probe, chart_extract, contact_sheet, tts, plan, and render each serve different pipeline stages. find_text, ocr_region, and track_text overlap around OCR but differ by output and granularity, so minor confusion is possible.
All names use lowercase snake_case, but the set mixes bare verbs (probe, plan, render), verb_noun names (find_text, track_text), and noun/acronym names (contact_sheet, tts). The pattern is readable but not fully predictable.
9 tools is well-scoped for a video analysis and rendering pipeline. Each tool maps to a meaningful operation rather than feeling redundant or excessive.
The surface covers analysis, text/chart extraction, visual inspection, narration, validation, and rendering, which is a solid lifecycle. Minor gaps exist around audio transcription or direct spec-building/editing, but agents can work around them.
Maintenance
Related MCP Connectors
Screen recording & video platform: search, share, transcribe, translate videos & AI meeting notes
Screen recording, meeting notes, and voice dictation - all with AI
The AI video studio: records your web app, writes and voices the script, renders a tutorial MP4.
Narrated demo and tutorial videos of your web app, from a guide or your coding agent.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables AI coding agents to capture screen and voice recordings, extract timestamped frames, and receive structured Markdown reports with context for bug fixing and UI feedback.29 npm18MIT
- AlicenseNot gradedqualityAmaintenanceEnables an AI agent to operate Windows applications through vision-driven UI Automation and record polished demo videos with pre-click camera zoom, narration, and cinematic effects.MIT
- AlicenseNot gradedqualityBmaintenanceEnables an AI agent to operate a web app, record a deterministic session, and compile an editable project into a polished MP4 with zooms, cuts, captions, callouts, and reproducible renders.MIT
- AlicenseNot gradedqualityAmaintenanceProvides local, offline transcription, keyframe extraction, OCR, and pre-publish review of audio, video, and image files, enabling AI agents to see and hear media without cloud or API keys.36 npmApache 2.0