babelscribe
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| transcribeA | Transcribe a local audio or video file into subtitles / text. file_path: absolute path to the media file (mp4, mkv, mov, mp3, wav, m4a, ...). language: ISO code such as en, th, ja, es, hi — or "auto" to detect. accurate: slower, fewest errors (large-v3 with beam search, or the best fine-tune for Thai / Hindi). formats: comma list from srt, vtt, txt, json. output_dir: folder for the output files; empty = next to the media file. Returns the files written, detected language, device used, and the transcript text. |
| find_mediaA | Find audio/video files on this computer, newest first. Searches Downloads, Videos, Desktop, Music and Documents (two levels deep) unless folder is given. Use it when the user names a file without a full path. |
| list_devicesB | GPUs whisper.cpp can use on this computer and which one babelscribe picks automatically. |
| list_languages_and_modelsC | Models available, and which model --accurate uses per language. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 4 tools
Each tool targets a clearly distinct action: list_devices (GPU discovery), list_languages_and_models (model/language info), transcribe (the core operation), and find_media (file discovery). No two tools overlap in purpose, so an agent can select unambiguously.
Consistently snake_case with verb-led names (list_devices, list_languages_and_models, find_media), which is predictable. The lone deviation is 'transcribe', which is verb-only with no noun object, though its meaning is still clear.
Four tools is well-scoped for a transcription server: two read-only discovery helpers, one file-locator helper, and one core action. Nothing feels redundant or missing from the count itself.
The surface covers the full workflow: locate media, inspect devices/models, and transcribe into multiple formats. Minor gaps exist (no batch processing, cancellation, or explicit model-download tool), but agents can work around these.