babelscribe
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@babelscribeTranscribe the newest video in my Downloads and save SRT subtitles"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
babelscribe — free, offline speech-to-text and subtitles on any GPU (AMD, NVIDIA, Intel, Apple)
babelscribe turns any audio or video file into subtitles (SRT, VTT) and text, in any of Whisper's 99 languages, on your own graphics card — AMD Radeon, NVIDIA, Intel Arc or Apple Silicon — or the CPU. No CUDA needed, no cloud, free (MIT). It runs OpenAI Whisper through whisper.cpp with Vulkan, CUDA or Metal, and works from the command line, from Python, or from AI apps (Claude, Codex, Antigravity, Gemini CLI, Cursor) as an MCP server.
pip install babelscribe # Python 3.9+
babelscribe interview.mp4 # auto-detects language, picks your best GPU -> interview.srt + interview.jsonbabelscribe is a thin, friendly layer on top of whisper.cpp. It exists because
getting fast Whisper on a non-NVIDIA card (e.g. an AMD Radeon on Windows) still means building whisper.cpp with the
Vulkan SDK yourself. babelscribe downloads a prebuilt whisper-cli for your system and handles the rest.
Your hardware | Backend used | Prebuilt asset |
AMD / NVIDIA / Intel GPU on Windows | Vulkan |
|
NVIDIA on Windows | Vulkan (same build) |
|
NVIDIA on Linux (alternative) | CUDA |
|
AMD / NVIDIA / Intel GPU on Linux | Vulkan |
|
Apple Silicon | Metal |
|
No usable GPU | CPU |
|
A 19-minute English TED talk transcribes in 48 seconds on an AMD Radeon RX 9070 XT with 1.5% word error — see the benchmark below.
Use it from an AI app (MCP) — no terminal needed
babelscribe is also an MCP server, so any AI app that supports local MCP servers can
transcribe files on your computer with your GPU. Ask in plain words — "make Thai subtitles for the newest video in
my Downloads" — and the app finds the file, transcribes it, writes .srt / .txt next to it, and can then proofread names,
translate or summarise the text for you.
Tools: transcribe (file → subtitles + text), find_media (newest audio/video in Downloads, Videos, Desktop, …),
list_devices, list_languages_and_models.
Claude Desktop (one click): download babelscribe.mcpb from the latest release
and double-click it (or Settings → Extensions → Install extension).
Everything else runs the same command — uv fetches babelscribe for you:
uvx babelscribe mcp
App | How to add it |
Claude Code |
|
OpenAI Codex CLI |
|
Google Antigravity | agent panel → ⋯ → MCP Servers → Manage MCP Servers → View raw config, add the JSON below |
Gemini CLI | add the JSON below to |
Cursor | add the JSON below to |
VS Code (Copilot) |
|
{
"mcpServers": {
"babelscribe": { "command": "uvx", "args": ["babelscribe", "mcp"], "timeout": 3600000 }
}
}(timeout is in milliseconds and only some apps read it.) Already installed with pip? Use "command": "babelscribe", "args": ["mcp"].
Web-only chat apps (e.g. grok.com, chatgpt.com) can only reach servers on the internet, not your computer, so they can't use your GPU or files.
Related MCP server: video-transcriber-mcp
Benchmark: real talks, human captions as the answer key
Four TED / TEDx talks, scored against the human-made captions in the spoken language (bench/bench.py, reproducible).
Error = word error rate (WER) for space-separated languages, character error rate (CER) for Japanese and Thai.
GPU: AMD Radeon RX 9070 XT via Vulkan. default = large-v3-turbo; accurate = --accurate (see the FLEURS section).
Language | Talk | Length | default: time / error |
|
English | 19.3 min | 48 s (24x) · WER 1.5% | 107 s (11x) · WER 1.5% | |
Japanese | 16.2 min | 47 s (20x) · CER 4.7% | 115 s (8x) · CER 4.2% | |
Spanish | 19.9 min | 56 s (21x) · WER 13.8% | 134 s (9x) · WER 13.5% | |
Thai | นิติ ชัยชิตาทร — โปรดเรียกฉันด้วยนามอันแท้จริง (TEDxBangkok) | 14.1 min | 74 s (12x) · CER 22.7% | 565 s (2x) · CER 16.4% |
How to read it: TED captions are edited for reading (fillers dropped, light rewording), so these numbers are an upper bound —
most of the Spanish "errors" are the speaker's actual words versus the tidied caption. Thai is genuinely harder: fast,
casual speech with slang; --accurate (Pathumma Whisper text + turbo timing) cuts its error by more than a quarter.
Talks are used only to measure accuracy; their transcripts are not redistributed (TED content is CC BY-NC-ND).
Benchmark: 16 languages on FLEURS
Google FLEURS test set, 50 utterances per language, human-verified
verbatim transcripts (bench/fleurs.py, reproducible). Same normaliser for every language (Whisper's rule: lower-case,
drop punctuation and non-spacing marks, numbers spelled out). WER for space-separated languages, CER for ja / zh / ko / th.
Language | default (turbo) |
| model |
|
English | 6.1% | 5.7% | large-v3, beam 5 | 5.6% |
Spanish | 4.1% | 4.2% | large-v3, beam 5 | 2.6% |
French | 7.4% | 7.3% | large-v3, beam 5 | 7.2% |
German | 4.3% | 4.0% | large-v3, beam 5 | 3.3% |
Portuguese | 8.5% | 7.7% | large-v3, beam 5 | 5.3% |
Italian | 6.7% | 6.5% | large-v3, beam 5 | 3.8% |
Russian | 6.8% | 5.8% | large-v3, beam 5 | 5.6% |
Arabic | 11.0% | 10.6% | large-v3, beam 5 | 9.7% |
Hindi | 28.4% | 12.6% | vasista22/whisper-hindi-large-v2 + turbo timing | 11.1% |
Indonesian | 9.8% | 7.8% | large-v3, beam 5 | 5.6% |
Vietnamese | 10.8% | 9.1% | large-v3, beam 5 | 9.1% |
Turkish | 6.2% | 6.7% | large-v3, beam 5 | 5.8% |
Japanese (CER) | 6.4% | 5.7% | large-v3, beam 5 | 4.6% |
Chinese (CER) | 6.3% | 5.3% | large-v3, beam 5 | 4.8% |
Korean (CER) | 4.1% | 3.9% | large-v3, beam 5 | 3.1% |
Thai (CER) | 15.9% | 8.9% | Pathumma Whisper (NECTEC) + turbo timing | 8.8% |
* Some FLEURS clips contain more speech than their reference transcript, so a correct model is charged for words the
reference leaves out. Edge-trimmed ignores extra words before the first / after the last reference word; both numbers
are stored by bench/fleurs.py. 50 utterances per language means differences under ~0.5 points are noise.
Also measured and not used, because large-v3 was as good or better: large-v2 (all languages), PhoWhisper-large (vi 19.0%), whisper-large-v3 dialectal / code-switching Arabic fine-tunes (12.2% / 15.2%), Typhoon Whisper (th 11.9%), Thonburian Whisper (th 9.1% — kept as an option), Vaani Hindi (16.5%).
Why babelscribe (vs. what already exists)
GPU on AMD / Intel | Windows, no build step | Video in, subtitles out | Long files don't loop | Better text for your language | |
babelscribe | ✅ Vulkan | ✅ prebuilt | ✅ ffmpeg bundled | ✅ | ✅ hybrid fine-tune text + turbo timing |
whisper.cpp (raw) | ✅ Vulkan — if you compile it | ❌ official releases ship no Windows Vulkan build | ❌ WAV 16 kHz only | ⚠️ you must know the flag | ❌ |
faster-whisper / WhisperX | ❌ GPU = NVIDIA CUDA only | ✅ pip | ✅ | ⚠️ | ⚠️ manual |
Cloud APIs | n/a (cloud) | ✅ | ✅ | ✅ | ❌ — and your audio leaves your machine, paid per minute |
babelscribe does not replace those projects — it stands on whisper.cpp and simply removes the hard parts: compiling for your GPU, converting media, picking the right device, avoiding the long-file repeat bug, and combining a language-specific fine-tune with accurate timestamps.
FAQ
How do I run Whisper on an AMD GPU on Windows?
pip install babelscribe, then babelscribe video.mp4. It downloads a Vulkan build of whisper.cpp that runs on AMD Radeon
(and Intel Arc / NVIDIA) cards on Windows and Linux — no ROCm, no CUDA, no compiling.
How do I make subtitles (SRT) from a video for free, offline?
babelscribe video.mp4 -f srt writes video.srt next to the video. Nothing is uploaded; it runs on your own computer.
What is the most accurate free transcription for Thai?
babelscribe video.mp4 -l th --accurate — Pathumma Whisper (NECTEC) for the text plus Whisper turbo for timing:
CER 8.9% on Google FLEURS vs 15.9% for plain Whisper turbo. Thonburian Whisper is available too (--text-model thai-thonburian).
Can Claude / ChatGPT Codex / Gemini transcribe a video on my computer?
Yes — add babelscribe as an MCP server (see Use it from an AI app). Claude Desktop installs it with one click from
babelscribe.mcpb. The AI app can then find a file, transcribe it on your GPU, and proofread, translate or summarise the text.
How fast is it? About 20x real time with Whisper large-v3-turbo on an AMD Radeon RX 9070 XT: a 19-minute talk in 48 seconds.
Which languages are supported? All 99 Whisper languages, with automatic language detection. FLEURS error rates for 16 of them are in the benchmark table.
Is it better than faster-whisper or WhisperX? Those are excellent on NVIDIA GPUs; on AMD / Intel GPUs they fall back to the CPU. babelscribe's niche is any GPU, zero setup, and better Thai / Hindi through community fine-tunes. If you have an NVIDIA card and like Python, faster-whisper is a fine choice.
Languages
All 99 languages Whisper was trained on, auto-detected or forced with -l:
af am ar as az ba be bg bn bo br bs ca cs cy da de el en es et eu fa fi fo fr gl gu ha haw he hi hr ht hu hy id is it ja jw ka kk km kn ko la lb ln lo lt lv mg mi mk ml mn mr ms mt my ne nl nn no oc pa pl ps pt ro ru sa sd si sk sl sn so sq sr su sv sw ta te tg th tk tl tr tt uk ur uz vi yi yo yue zh
Accuracy follows Whisper's own training data: excellent for high-resource languages (English, Spanish, Japanese, …),
weaker for low-resource ones. That is what hybrid mode is for — a community fine-tune for one language can be plugged in
with one line in babelscribe/models.py (Thai and Hindi so far). PRs adding fine-tunes for other
languages are the most valuable contribution.
Limitations (honest)
Tested end to end so far on an AMD Radeon RX 9070 XT (Windows, Vulkan). CUDA, Linux and macOS builds are produced by CI; reports from those machines are welcome.
Hybrid mode runs two models, so it is slower (≈1 min per minute of audio with a large fine-tune on that GPU).
Proper nouns can still be misspelled — check names before publishing subtitles.
Features
Any input — mp4, mkv, mov, mp3, wav, m4a… (ffmpeg is bundled through
imageio-ffmpeg).Any language —
-l autodetects it; or pass-l th,-l ja,-l es…Picks the right GPU — prefers a discrete card over an integrated one;
babelscribe deviceslists them,--device Noverrides.Long files that don't loop — runs whisper with
--max-context 0, which stops the classic "same sentence repeated forever" hallucination on long recordings.Hybrid mode for better spelling in your language — community fine-tunes (e.g. Thai Thonburian Whisper) spell far better but often lose timestamps.
--text-modeltakes the text from the fine-tune and the timing fromlarge-v3-turbo, aligned character by character (works for languages without spaces), cut only at word boundaries, with the timing model filling any words the fine-tune skipped.--accurate— slower, fewest errors: large-v3 with beam search, or for Thai and Hindi the best community fine-tune (hybrid). Chosen per language from the FLEURS benchmark above.Outputs —
srt,vtt,txt,json(segments with token timings).
Usage
babelscribe talk.mp4 -f srt,vtt,txt,json # all formats
babelscribe podcast.mp3 -l en -m large-v3 # pick language and model
babelscribe talk.mp4 -l hi --accurate # slower, fewest errors (best model per language)
babelscribe vo.wav -l th --text-model thai-thonburian # hybrid: pick a Thai fine-tune yourself
babelscribe devices # GPUs whisper.cpp can see
babelscribe models # models and fine-tunes
babelscribe talk.mp4 --bin /path/to/whisper-cli # use your own whisper.cpp buildFrom Python:
from babelscribe.api import transcribe_file
r = transcribe_file("talk.mp4", lang="auto", formats=["srt", "txt"]) # or accurate=True
print(r["lang"], r["device"], r["files"]); print(r["text"][:200])Models download on first use to ~/.babelscribe/models (BABELSCRIBE_MODELS to change). Hybrid fine-tunes are converted on your
machine from their original Hugging Face repo — install pip install "babelscribe[finetune]" once; converted weights are never
redistributed, so each fine-tune keeps its own licence. Thai word boundaries: pip install "babelscribe[thai]".
Add a language fine-tune
Add one entry to FINETUNES in babelscribe/models.py (Hugging Face repo, language code, whether it needs beam search) and open a PR.
Prebuilt binaries
.github/workflows/build-binaries.yml builds whisper-cli + whisper-quantize for every row of the table above and attaches
them to each v* release. Point BABELSCRIBE_RELEASES at another URL to self-host.
ภาษาไทย
ใช้ผ่าน Claude Desktop ได้โดยไม่ต้องพิมพ์คำสั่ง: โหลด babelscribe.mcpb จากหน้า Releases แล้วดับเบิลคลิก จากนั้นพิมพ์ในแชทว่า "ทำซับไทยให้คลิปล่าสุดในโฟลเดอร์ดาวน์โหลด"
ถอดเสียงจากไฟล์เสียงหรือวิดีโอได้ทุกภาษา บนการ์ดจอทุกยี่ห้อ (AMD / NVIDIA / Intel ผ่าน Vulkan, Apple ผ่าน Metal) หรือ CPU
ภาษาไทยแนะนำ babelscribe ไฟล์.mp4 -l th --accurate — ข้อความจาก Pathumma Whisper (NECTEC) ที่ผิดน้อยที่สุดใน FLEURS (CER 8.9% เทียบ turbo 15.9%) + เวลาจาก large-v3-turbo
หรือเลือก Thonburian Whisper เอง: --text-model thai-thonburian
Privacy Policy
babelscribe runs entirely on your own computer. Last updated 2026-10-06.
Data collection: none. babelscribe has no telemetry, analytics, accounts or crash reporting, and never uploads your audio, video, transcripts or file names anywhere.
What it processes and where: the media file you choose is converted and transcribed locally; subtitles and text are written to your disk (next to the file, or the folder you choose). Through MCP, the transcript text is returned to the AI app that called the tool — what that app does with it is governed by that app's own privacy policy.
Network access (downloads only): on first use it downloads the
whisper-cliprogram from this project's GitHub releases, speech models from Hugging Face (huggingface.co/ggerganov/whisper.cpp, and for--accurateThai/Hindi the fine-tune's own repository), and Python packages from PyPI when installed with pip/uv. These are plain downloads; no personal data is sent. GitHub, Hugging Face and PyPI see a normal download request (IP address, user agent) under their own privacy policies.Storage and retention: downloaded programs and models are cached in
~/.babelscribe(orBABELSCRIBE_HOME/BABELSCRIBE_MODELS) until you delete that folder. Outputs stay wherever they were written until you delete them. babelscribe keeps no other data.Third-party sharing: none.
Contact: open an issue at https://github.com/phonology024/babelscribe/issues
Credits & licence
MIT. Built on whisper.cpp (MIT) and OpenAI Whisper models (MIT). Fine-tunes belong to their authors: Pathumma Whisper by NECTEC, Thonburian Whisper by biodatlab, whisper-hindi-large-v2 by vasista22 (Speech Lab, IIT Madras).
Available Tools
4 toolsfind_mediaFind audio/video filesARead-only
Find audio/video files on this computer, newest first. Searches Downloads, Videos, Desktop, Music and Documents (two levels deep) unless folder is given. Use it when the user names a file without a full path.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| folder | No | ||
| name_contains | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint and openWorldHint, so the description carries the rest: it discloses the default search set (Downloads, Videos, Desktop, Music, Documents), the two-level depth limit, the newest-first ordering, and that folder overrides the defaults. It omits the default result cap (limit=15) and what happens on an empty result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and scope, then the selection rule. No filler and every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need no explanation, and safety is covered by annotations. The description supplies scope, depth, ordering, and usage context; only the limit parameter and empty-result behavior are left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 3 parameters, so the description must compensate. It explains folder's override behavior well and the "names a file" phrasing implies name_contains, but limit (default 15, and any upper bound) is never explained in either place.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Find) and resource (audio/video files) plus the machine scope and sort order, which is far more than a restatement of the name. It does not explicitly contrast itself with siblings like transcribe, but the purpose is unambiguous on its own.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use it when the user names a file without a full path" gives a concrete triggering condition. There is no explicit when-not guidance or named alternative, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_devicesList GPUsBRead-only
GPUs whisper.cpp can use on this computer and which one babelscribe picks automatically.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so safety and open-world nature are covered. The description adds useful context that the result includes the automatically chosen device, but says nothing about ordering, stability, or what happens when no GPU is usable, which a discovery tool with no output schema would benefit from.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no padding and the key qualifier ('babelscribe picks automatically') up front. It loses a point only because the sentence is an ungrammatical fragment that takes a second pass to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and no output schema, the description should carry the full burden of describing what comes back. It hints at the content (list of GPUs plus the auto-selected one) but does not specify the return shape, device naming, or the empty/no-GPU case, leaving a real gap for a tool whose only purpose is discovery.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema is trivially complete and there is nothing for the description to compensate for. Baseline 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (GPUs usable by whisper.cpp on this machine) and adds a distinguishing detail about automatic selection, but it is a noun-phrase fragment with no explicit verb and largely restates the title 'List GPUs'. An agent can infer it is a listing/discovery tool, but the purpose is implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus siblings like list_languages_and_models, nor any note on whether it should be checked before transcribe to confirm hardware support. Usage is left entirely to inference from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_languages_and_modelsList languages and modelsCRead-only
Models available, and which model --accurate uses per language.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds no behavioral context beyond that - no indication of return format, cost, or whether the listing is static or environment-dependent, despite the openWorldHint=false implying a fixed local catalog.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short fragment with no wasted words, and the key information (models + per-language --accurate mapping) is front-loaded. It is terse to the point of being a label rather than a sentence, which slightly limits clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter lookup with no output schema, the description conveys what kind of data comes back but not its shape (e.g., list vs. language-keyed map) or how it relates to other tools. Adequate but leaves the agent guessing about the return structure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. The schema is trivially complete and the description correctly implies the output is a static mapping rather than something parameterized.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (available models) and adds a specific mapping detail ('which model --accurate uses per language'), which helps distinguish it from siblings like transcribe. However, it is a noun fragment with no verb, so the agent must infer that the tool lists/enumerates rather than computes or modifies anything.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit statement of when to call this tool or what it replaces. The reference to the '--accurate' flag hints at a pre-transcription lookup use case, but the agent must infer that this is a discovery step before calling transcribe, and there are no exclusions or alternatives named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transcribeTranscribe audio/video to subtitlesAIdempotent
Transcribe a local audio or video file into subtitles / text.
file_path: absolute path to the media file (mp4, mkv, mov, mp3, wav, m4a, ...). language: ISO code such as en, th, ja, es, hi — or "auto" to detect. accurate: slower, fewest errors (large-v3 with beam search, or the best fine-tune for Thai / Hindi). formats: comma list from srt, vtt, txt, json. output_dir: folder for the output files; empty = next to the media file. Returns the files written, detected language, device used, and the transcript text.
| Name | Required | Description | Default |
|---|---|---|---|
| formats | No | srt,txt | |
| accurate | No | ||
| language | No | auto | |
| file_path | Yes | ||
| output_dir | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare non-readonly, non-destructive, idempotent, open-world, so the safety profile is partially covered; the description adds real value beyond that by disclosing that output files are written, where they land ('empty = next to the media file'), the accurate-mode performance cost, and what is returned. It stops short of covering failure modes or overwrite behavior on re-runs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded one-line purpose followed by terse per-parameter lines and a return summary; every sentence adds information and nothing is repeated from the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming the return payload (files written, detected language, device, transcript text), and all five parameters are explained. Nothing an agent needs to invoke this correctly appears to be missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden and does so for all five parameters: path form and accepted containers, ISO language codes plus 'auto', the accuracy/speed tradeoff, the valid format set (srt, vtt, txt, json) which is not enumerated in the schema, and the output_dir default behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Transcribe a local audio or video file into subtitles / text') and names the output artifacts, which clearly separates it from siblings like find_media and list_devices. An agent can identify the tool's job without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the emphasis on a 'local' file with an 'absolute path' hints this is not for discovery (find_media) and the accurate-mode tradeoff ('slower, fewest errors') is a useful when-to-pick hint. However, no alternative tools are named and there is no explicit when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.3.2- First observed
find_media - First observed
list_devices - First observed
list_languages_and_models - First observed
transcribe
TDQS
Scored across 4 tools
Each tool targets a clearly distinct action: list_devices (GPU discovery), list_languages_and_models (model/language info), transcribe (the core operation), and find_media (file discovery). No two tools overlap in purpose, so an agent can select unambiguously.
Consistently snake_case with verb-led names (list_devices, list_languages_and_models, find_media), which is predictable. The lone deviation is 'transcribe', which is verb-only with no noun object, though its meaning is still clear.
Four tools is well-scoped for a transcription server: two read-only discovery helpers, one file-locator helper, and one core action. Nothing feels redundant or missing from the count itself.
The surface covers the full workflow: locate media, inspect devices/models, and transcribe into multiple formats. Minor gaps exist (no batch processing, cancellation, or explicit model-download tool), but agents can work around these.
Maintenance
Related MCP Connectors
Transcribe audio and video files to text, SRT and VTT subtitles in 99 languages with Whisper.
AI transcription from URLs or files. 119 languages, diarization, SRT/VTT/text export.
Transcribe audio & video to text for AI agents: 100+ languages, speaker labels, webhooks.
- RendobarOAuthcom.rendobar
Transform video, audio and images, and generate media from prompts. FFmpeg, captions, models.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceEnables high-performance audio transcription using Faster Whisper with CUDA acceleration, supporting single and batch audio file processing with multiple output formats (VTT, SRT, JSON).-
- AlicenseNot gradedqualityAmaintenanceEnables high-performance, offline transcription of videos from 1000+ platforms and local files using whisper.cpp, with support for multiple model sizes, languages, and output formats over stdio or HTTP.18Apache 2.0
- AlicenseAqualityCmaintenanceEnables MCP clients to transcribe audio/video files locally, generate SRT subtitles, and burn captions into videos via tool calls, without a cloud API.3MIT
- AlicenseAqualityAmaintenanceEnables AI clients to watch local video files or YouTube/Bilibili and other supported URLs, receiving timestamped transcripts, subtitles, searchable text and keyframe contact sheets. Runs fully offline with local speech recognition and bundled ffmpeg, so no API key or cloud upload is needed.673 PyPI9MIT