mcp-agent-chatterbox
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| HF_HOME | No | Override the Hugging Face cache location | |
| CHATTERBOX_MODEL | No | Model used when speak names none | turbo |
| CHATTERBOX_DEVICE | No | auto | cpu | cuda | cuda:N | auto |
| CHATTERBOX_VRAM_MB | No | Free-VRAM floor for a load attempt | 4096 |
| CHATTERBOX_AUTOPLAY | No | Play audio after writing | 1 |
| CHATTERBOX_LOGS_DIR | No | Rotating log file location | logs |
| CHATTERBOX_GPU_INDEX | No | Pin a CUDA ordinal. Overrides the auto heuristic | |
| CHATTERBOX_MAX_CHARS | No | Reject longer single requests | 4000 |
| CHATTERBOX_OUTPUT_DIR | No | Where WAVs are written | tts_output |
| CHATTERBOX_VOICES_DIR | No | Where reference clips live | voices |
| CHATTERBOX_STRICT_VRAM | No | 1 refuses a load below the floor instead of warning | 0 |
| CHATTERBOX_CHUNK_PAUSE_MS | No | Silence inserted between concatenated chunks | 250 |
| CHATTERBOX_MAX_CHUNK_CHARS | No | Long text auto-splits into sentence-aligned chunks of up to this many chars | 500 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| tts_statusA | Report Chatterbox TTS readiness: installed packages, resolved device, per-GPU free/total VRAM, the resident model, the models available, the languages the multilingual model speaks, and the paralinguistic tags turbo understands. Loads no weights — safe as a preflight before a speak call. On a machine where a large model already occupies a GPU, check free_mb here first: a card with less than ~4096 MB free will fail to load. |
| tts_unloadA | Release the currently loaded Chatterbox model and return its VRAM to the GPU. Use this when another process needs the card, or to drop a 500M model before a lighter turbo load. Safe to call when nothing is loaded. |
| speakA | Speak text aloud with Chatterbox on the local GPU, save it as a WAV, and by default play it through the system audio device. model: "turbo" (default, 350M, English, no reference clip needed, supports inline [laugh]/[chuckle] tags), "multilingual" (500M, 23 languages, requires a reference clip) or "original" (500M, English, requires a reference clip). voice: name of a reference clip in the voices directory (its filename without extension). Required for multilingual/original. reference_clip: path to a wav/mp3/flac clip, as an alternative to voice=. Supply as much clean continuous speech as you have. Short references measurably increase fine clicks and crackle: a 40 s reference beat an 11.8 s excerpt of the same recording by a wide margin, several times the model's own run-to-run variance. Under ~5 s is genuinely too little. Do not normalise or limit the clip -- that shifts output level and adds artefacts. A mastered or compressed recording is fine. Sample rate and channel count need not match anything. language: ISO 639-1 code, multilingual only (e.g. "de", "en", "fr"). t3_model: "v2" (default) or "v3" — multilingual checkpoint. temperature: sampling temperature (0.05-5.0). Honored by all models. top_p: nucleus-sampling cutoff (0.0-1.0). Honored by all models. top_k: top-k sampling size (0-1000). Honored by turbo only — the 500M models have no such parameter. repetition_penalty: penalise repeated tokens (1.0-2.0). Honored by all models. norm_loudness: normalize output to -27 LUFS. Honored by turbo only — the 500M models have no such parameter. exaggeration/cfg_weight: style controls (0.0-2.0 / 0.0-1.0). Honored ONLY by the 500M models (multilingual/original); turbo ignores both, so setting them with model="turbo" is dropped with a warning, not an error. cfg_weight>0 doubles the text tokens for CFG guidance. seed: reseed torch (CPU + CUDA) so re-renders are reproducible within the resident session; 0 (the upstream convention) or unset keeps random sampling. Byte-identical output across a server restart is not guaranteed — Chatterbox is nondeterministic across CUDA kernel choices. Left unset, every knob keeps the model's own tuned default (e.g. turbo runs cfg_weight 0.0 with top_k 1000; the 500M models run cfg_weight 0.5). play_audio: set false to only write the file. Defaults to the server setting (on). wait: block until playback finishes instead of returning immediately. filename: output basename; a safe name is generated when omitted. The first call downloads the model from Hugging Face (hundreds of MB to a few GB) and can take a while; later calls reuse the cached weights. Returns JSON with the output filename, duration, sample rate, model, device and whether it played. |
| list_voicesA | List the reference clips available for voice cloning — the files in the voices directory. Returns each clip's name (what you pass as voice=), filename, size, duration, sample rate and channel count, plus which models can use it. The multilingual and original models require one of these; turbo does not. For good cloning use as much clean single-speaker audio as you have, not a short excerpt — clipping a good recording down measurably adds fine clicks and crackle. A mastered or compressed source is fine; do not normalise it. Sample rate and channel count need not match anything. Clips too short to be safe carry a quality_note. |
| stop_speechA | Stop audio that is currently playing. Useful for cutting off a long utterance started with wait=false. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 5 tools
Each tool has a clearly distinct purpose: tts_status reports readiness, tts_unload releases VRAM, speak synthesizes audio, list_voices enumerates reference clips, and stop_speech halts playback. There is no overlap or boundary confusion between any pair.
Naming is mixed: two tools use a 'tts_' prefix (tts_status, tts_unload) while the others do not, and 'speak' is a bare verb while list_voices/stop_speech follow verb_noun. It is still readable snake_case, but the conventions are not predictable as a set.
Five tools is well-scoped for a local TTS server: status, unload, speak, list_voices, and stop_speech each cover a necessary lifecycle operation without redundancy. Nothing feels thin or bloated.
The surface covers the core TTS lifecycle (preflight, synthesize/play, list voices, stop, unload) and the speak tool exposes rich model controls. Minor gaps exist, such as no explicit preload of a model or enumeration of previously generated output files, but agents can work around these.