speak
Generate speech audio from text with Chatterbox on local GPU, save as WAV, and play through system audio.
Instructions
Speak text aloud with Chatterbox on the local GPU, save it as a WAV, and by default play it through the system audio device.
model: "turbo" (default, 350M, English, no reference clip needed, supports inline [laugh]/[chuckle] tags), "multilingual" (500M, 23 languages, requires a reference clip) or "original" (500M, English, requires a reference clip). voice: name of a reference clip in the voices directory (its filename without extension). Required for multilingual/original. reference_clip: path to a wav/mp3/flac clip, as an alternative to voice=. Supply as much clean continuous speech as you have. Short references measurably increase fine clicks and crackle: a 40 s reference beat an 11.8 s excerpt of the same recording by a wide margin, several times the model's own run-to-run variance. Under ~5 s is genuinely too little. Do not normalise or limit the clip -- that shifts output level and adds artefacts. A mastered or compressed recording is fine. Sample rate and channel count need not match anything. language: ISO 639-1 code, multilingual only (e.g. "de", "en", "fr"). t3_model: "v2" (default) or "v3" — multilingual checkpoint. temperature: sampling temperature (0.05-5.0). Honored by all models. top_p: nucleus-sampling cutoff (0.0-1.0). Honored by all models. top_k: top-k sampling size (0-1000). Honored by turbo only — the 500M models have no such parameter. repetition_penalty: penalise repeated tokens (1.0-2.0). Honored by all models. norm_loudness: normalize output to -27 LUFS. Honored by turbo only — the 500M models have no such parameter. exaggeration/cfg_weight: style controls (0.0-2.0 / 0.0-1.0). Honored ONLY by the 500M models (multilingual/original); turbo ignores both, so setting them with model="turbo" is dropped with a warning, not an error. cfg_weight>0 doubles the text tokens for CFG guidance. seed: reseed torch (CPU + CUDA) so re-renders are reproducible within the resident session; 0 (the upstream convention) or unset keeps random sampling. Byte-identical output across a server restart is not guaranteed — Chatterbox is nondeterministic across CUDA kernel choices. Left unset, every knob keeps the model's own tuned default (e.g. turbo runs cfg_weight 0.0 with top_k 1000; the 500M models run cfg_weight 0.5). play_audio: set false to only write the file. Defaults to the server setting (on). wait: block until playback finishes instead of returning immediately. filename: output basename; a safe name is generated when omitted.
The first call downloads the model from Hugging Face (hundreds of MB to a few GB) and can take a while; later calls reuse the cached weights. Returns JSON with the output filename, duration, sample rate, model, device and whether it played.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | ||
| text | Yes | ||
| wait | No | ||
| model | No | ||
| top_k | No | ||
| top_p | No | ||
| voice | No | ||
| filename | No | ||
| language | No | ||
| t3_model | No | ||
| cfg_weight | No | ||
| play_audio | No | ||
| temperature | No | ||
| exaggeration | No | ||
| norm_loudness | No | ||
| reference_clip | No | ||
| repetition_penalty | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| result | Yes |