Skip to main content
Glama
ChristofMilius

mcp-agent-chatterbox

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
HF_HOMENoOverride the Hugging Face cache location
CHATTERBOX_MODELNoModel used when speak names noneturbo
CHATTERBOX_DEVICENoauto | cpu | cuda | cuda:Nauto
CHATTERBOX_VRAM_MBNoFree-VRAM floor for a load attempt4096
CHATTERBOX_AUTOPLAYNoPlay audio after writing1
CHATTERBOX_LOGS_DIRNoRotating log file locationlogs
CHATTERBOX_GPU_INDEXNoPin a CUDA ordinal. Overrides the auto heuristic
CHATTERBOX_MAX_CHARSNoReject longer single requests4000
CHATTERBOX_OUTPUT_DIRNoWhere WAVs are writtentts_output
CHATTERBOX_VOICES_DIRNoWhere reference clips livevoices
CHATTERBOX_STRICT_VRAMNo1 refuses a load below the floor instead of warning0
CHATTERBOX_CHUNK_PAUSE_MSNoSilence inserted between concatenated chunks250
CHATTERBOX_MAX_CHUNK_CHARSNoLong text auto-splits into sentence-aligned chunks of up to this many chars500

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
tts_statusA

Report Chatterbox TTS readiness: installed packages, resolved device, per-GPU free/total VRAM, the resident model, the models available, the languages the multilingual model speaks, and the paralinguistic tags turbo understands. Loads no weights — safe as a preflight before a speak call. On a machine where a large model already occupies a GPU, check free_mb here first: a card with less than ~4096 MB free will fail to load.

tts_unloadA

Release the currently loaded Chatterbox model and return its VRAM to the GPU. Use this when another process needs the card, or to drop a 500M model before a lighter turbo load. Safe to call when nothing is loaded.

speakA

Speak text aloud with Chatterbox on the local GPU, save it as a WAV, and by default play it through the system audio device.

model: "turbo" (default, 350M, English, no reference clip needed, supports inline [laugh]/[chuckle] tags), "multilingual" (500M, 23 languages, requires a reference clip) or "original" (500M, English, requires a reference clip). voice: name of a reference clip in the voices directory (its filename without extension). Required for multilingual/original. reference_clip: path to a wav/mp3/flac clip, as an alternative to voice=. Supply as much clean continuous speech as you have. Short references measurably increase fine clicks and crackle: a 40 s reference beat an 11.8 s excerpt of the same recording by a wide margin, several times the model's own run-to-run variance. Under ~5 s is genuinely too little. Do not normalise or limit the clip -- that shifts output level and adds artefacts. A mastered or compressed recording is fine. Sample rate and channel count need not match anything. language: ISO 639-1 code, multilingual only (e.g. "de", "en", "fr"). t3_model: "v2" (default) or "v3" — multilingual checkpoint. temperature: sampling temperature (0.05-5.0). Honored by all models. top_p: nucleus-sampling cutoff (0.0-1.0). Honored by all models. top_k: top-k sampling size (0-1000). Honored by turbo only — the 500M models have no such parameter. repetition_penalty: penalise repeated tokens (1.0-2.0). Honored by all models. norm_loudness: normalize output to -27 LUFS. Honored by turbo only — the 500M models have no such parameter. exaggeration/cfg_weight: style controls (0.0-2.0 / 0.0-1.0). Honored ONLY by the 500M models (multilingual/original); turbo ignores both, so setting them with model="turbo" is dropped with a warning, not an error. cfg_weight>0 doubles the text tokens for CFG guidance. seed: reseed torch (CPU + CUDA) so re-renders are reproducible within the resident session; 0 (the upstream convention) or unset keeps random sampling. Byte-identical output across a server restart is not guaranteed — Chatterbox is nondeterministic across CUDA kernel choices. Left unset, every knob keeps the model's own tuned default (e.g. turbo runs cfg_weight 0.0 with top_k 1000; the 500M models run cfg_weight 0.5). play_audio: set false to only write the file. Defaults to the server setting (on). wait: block until playback finishes instead of returning immediately. filename: output basename; a safe name is generated when omitted.

The first call downloads the model from Hugging Face (hundreds of MB to a few GB) and can take a while; later calls reuse the cached weights. Returns JSON with the output filename, duration, sample rate, model, device and whether it played.

list_voicesA

List the reference clips available for voice cloning — the files in the voices directory. Returns each clip's name (what you pass as voice=), filename, size, duration, sample rate and channel count, plus which models can use it. The multilingual and original models require one of these; turbo does not.

For good cloning use as much clean single-speaker audio as you have, not a short excerpt — clipping a good recording down measurably adds fine clicks and crackle. A mastered or compressed source is fine; do not normalise it. Sample rate and channel count need not match anything. Clips too short to be safe carry a quality_note.

stop_speechA

Stop audio that is currently playing. Useful for cutting off a long utterance started with wait=false.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A4.2/5.0

Scored across 5 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: tts_status reports readiness, tts_unload releases VRAM, speak synthesizes audio, list_voices enumerates reference clips, and stop_speech halts playback. There is no overlap or boundary confusion between any pair.

Naming Consistency3/5

Naming is mixed: two tools use a 'tts_' prefix (tts_status, tts_unload) while the others do not, and 'speak' is a bare verb while list_voices/stop_speech follow verb_noun. It is still readable snake_case, but the conventions are not predictable as a set.

Tool Count5/5

Five tools is well-scoped for a local TTS server: status, unload, speak, list_voices, and stop_speech each cover a necessary lifecycle operation without redundancy. Nothing feels thin or bloated.

Completeness4/5

The surface covers the core TTS lifecycle (preflight, synthesize/play, list voices, stop, unload) and the speak tool exposes rich model controls. Minor gaps exist, such as no explicit preload of a model or enumeration of previously generated output files, but agents can work around these.

Maintenance

ActivityMaintained
ResponsivenessNo issues