tts-audio-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| TTS_PYTHON_BIN | No | Python binary with librosa installed | .venv/bin/python3 |
| WHISPER_BINARY | No | Path to whisper.cpp binary | whisper-cli |
| WHISPER_MODEL_PATH | No | Path to Whisper model file | ~/Documents/_dev/tts-audio-mcp/models/ggml-large-v3-turbo.bin |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| transcribeA | Transcribe an audio file to text with word-level timestamps using Whisper |
| quality_scoreA | Analyze speech quality metrics of an audio file — pitch variation, energy, pacing, silence ratio. Detects robotic tone, monotone speech, and audio issues. |
| compare_ttsA | Compare TTS audio output against expected text — identifies mispronunciations, inserted/deleted words, and Word Error Rate (WER) |
| analyze_ttsA | Full TTS audio analysis — transcription, quality scores, pacing analysis, and optional mispronunciation detection. Returns a comprehensive report for debugging TTS issues. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 4 tools
Each tool has a clearly distinct purpose: transcribe focuses solely on speech-to-text, quality_score on speech metrics, compare_tts on TTS-to-text alignment, and analyze_tts explicitly as a comprehensive bundle. The descriptions make the boundaries clear despite analyze_tts overlapping with the others.
Tool names mostly follow snake_case verb-noun pattern (transcribe, compare_tts, analyze_tts), but quality_score deviates by using a noun-noun form rather than a verb. Overall still readable and predictable, with only minor inconsistency.
Four tools is a well-scoped count for an audio analysis server. Each tool earns its place, offering both specialized operations and a comprehensive aggregate, without unnecessary bloat.
The domain of TTS audio analysis is fully covered: transcription, quality scoring, comparison against expected text, and a full-analysis option. There are no obvious gaps for a debugging workflow, as analyze_tts bundles all core features.