Skip to main content
Glama
ChristofMilius

mcp-agent-chatterbox

mcp-agent-chatterbox

MCP server for local, GPU-accelerated text-to-speech with Chatterbox (Resemble AI, MIT). Text goes in, spoken audio comes out — no cloud service, no API key, no per-character cost.

Ships two things:

  • An MCP tool surface (speak, tts_status, tts_unload, list_voices, stop_speech) that any MCP client can call.

  • An opencode plugin (.opencode/plugins/tts.js) that adds a speak tool and a /speak command to opencode itself, and installs itself globally so it works in every project, not just this repo.

The three models

Key

Params

Languages

Reference clip

Notes

turbo

350M

English

not needed

Default. Fastest, lowest VRAM. Built-in voice. Understands paralinguistic tags inline: [laugh], [chuckle], [cough], [sigh], [whisper].

multilingual

500M

23 (incl. German)

required

language="de", en, fr, … t3_model="v2" (default) or "v3".

original

500M

English

required

The CFG / exaggeration-tuning model, for delivery style control.

Only turbo ships a stock voice. The other two raise AssertionError: Please prepare_conditionals first or specify audio_prompt_path without a reference clip, so the tool checks this up front and returns a reference_clip_required payload naming the fix instead of a stack trace.

Related MCP server: voice-mcp-server

Requirements

  • Python 3.13 (see Why 3.13 only)

  • An NVIDIA GPU with a driver of 525+ for the CUDA 12.4 wheels. CPU works but is slow.

  • ~6 GB of free VRAM for a comfortable margin (turbo needs ~2-3 GB, the 500M models ~3-4 GB)

  • Windows for audio playback. On other platforms the WAV is still written; the played field comes back false with a reason.

Installation

uv sync --extra dev

uv sync installs the CUDA build of PyTorch automatically. pyproject.toml pins torch to PyTorch's official CUDA index via [tool.uv.sources], so GPU acceleration is a property of the environment rather than an accident of the platform.

Verify

uv run mcp-agent-chatterbox doctor

Prints the resolved device, per-GPU free/total VRAM, the models, and any reference clips found. On a multi-GPU box this is also how you find out which card is free.

Quick CUDA check:

uv run python -c "import torch; print(torch.__version__, torch.cuda.is_available())"

Expect 2.6.0+cu124 True.

Why setuptools<81 is pinned

This is the one dependency in pyproject.toml that looks arbitrary, so here is the whole story.

The turbo model embeds a copy of Perth, Resemble AI's audio watermarking library. That copy is vendored inside the chatterbox-tts sdist — it ships as loose files with no package metadata, and there is no perth-watermarker distribution on PyPI to replace it with.

perth/perth_net/__init__.py starts with:

from pkg_resources import resource_filename

pkg_resources was part of setuptools until setuptools 81 dropped it. So on a current setuptools the import raises ImportError, which perth/__init__.py swallows and turns into PerthImplicitWatermarker = None — and ChatterboxTurboTTS.__init__ then calls that None:

TypeError: 'NoneType' object is not callable

The failure is silent by design upstream (the try/except ImportError is intentional, so Perth works without the neural watermarker), which is exactly why it surfaces as a confusing TypeError at load time rather than a clear message. Pinning setuptools<81 restores pkg_resources and the turbo model loads.

Consequences worth knowing:

  • The pin is a runtime requirement, not a build-time one. The dependency graph still works without it; turbo just fails to load.

  • Only the turbo model is affected. multilingual and original never touch Perth and load fine on any setuptools.

  • If a future chatterbox-tts release fixes the vendored import, drop the pin. The symptom to watch for is the TypeError above returning.

Model weights

Weights are not bundled. They download from Hugging Face on the first speak() call and are cached in the standard location (%USERPROFILE%\.cache\huggingface\hub on Windows, override with HF_HOME). Subsequent calls reuse the cache.

  • turbo → ResembleAI/chatterbox-turbo

  • multilingual / original → ResembleAI/chatterbox

Nothing is downloaded by doctor or tts_status, so those stay fast and work offline.

Registering the MCP server

Add to ~/.config/opencode/opencode.jsonc:

{
  "mcp": {
    "mcp-agent-chatterbox": {
      "type": "local",
      "command": [
        "uv", "run", "--directory", "<absolute path to this repo>",
        "mcp-agent-chatterbox", "serve"
      ],
      "environment": {
        "CHATTERBOX_MODEL": "turbo"
      },
      "enabled": true
    }
  }
}

Restart opencode. The server speaks stdio, loads no weights at startup, and stays responsive while the model is idle.

The opencode plugin

.opencode/plugins/tts.js adds a speak tool and a /speak command to opencode. It self-replicates into ~/.config/opencode/plugins/ the first time opencode loads it from this repo, so the tool is available in every project rather than only sessions started inside this folder.

The plugin deliberately does not auto-speak assistant replies: loading a model costs seconds and VRAM, and unsolicited audio on every turn would be awful. Speaking stays an explicit choice, by tool call or by /speak.

Tools

speak

The main call. Renders text, writes a WAV to the output dir, plays it.

Parameter

Default

Meaning

text

—

What to say. Required.

model

turbo

turbo | multilingual | original

voice

—

Reference clip name in the voices dir (filename without extension)

reference_clip

—

Path to a wav/mp3/flac, as an alternative to voice

language

—

ISO 639-1 code, multilingual only

t3_model

v2

v2 | v3, multilingual only

exaggeration

per model

Emotional range. 500M models only — turbo ignores it

cfg_weight

per model

Guidance strength. 500M models only, turbo 0.0 / 500M 0.5

temperature

per model

Sampling temperature (0.05–5.0)

top_p

per model

Nucleus-sampling cutoff (0.0–1.0)

top_k

per model

Top-k sampling size (0–1000). Turbo only — the 500M models have no such parameter

repetition_penalty

per model

Penalise repeated tokens (1.0–2.0)

norm_loudness

per model

Normalize to −27 LUFS. Turbo only — the 500M models have no such parameter

seed

random

Reseed torch for a reproducible re-render within the resident session; 0 means random. Not guaranteed byte-identical across a restart

play_audio

server setting

false writes the file only

wait

false

Block until playback finishes

filename

generated

Output basename

Style parameters left unset keep each model's own tuned values rather than a single global default — turbo is tuned for latency at cfg_weight=0.0, the 500M models for fidelity at 0.5, and forcing one number onto all three audibly degrades two of them. A knob a model does not support (turbo ignores exaggeration/cfg_weight; the 500M models take no top_k/norm_loudness) is dropped with a warning rather than crashing.

Text longer than CHATTERBOX_MAX_CHUNK_CHARS is spoken as one file but rendered in sentence-aligned chunks: Chatterbox's own generate() truncates long inputs (turbo degrades past roughly 600 characters and comes back shorter than a shorter prompt), so a single oversized render would be garbled. The chunks share the same voice, style knobs and reference clip, and are stitched together with CHATTERBOX_CHUNK_PAUSE_MS of silence between them.

// a quick aside in the stock voice
{ "text": "Build finished. All tests pass.", "model": "turbo" }

// German, cloned voice
{ "text": "Der Build ist fertig.", "model": "multilingual", "language": "de" }

// expressive narration — exaggeration tunes the 500M models; turbo ignores it
{ "text": "And then [chuckle] it compiled on the first try.", "model": "original",
  "exaggeration": 0.4 }

tts_status

Runtime readiness without loading anything: packages, torch/CUDA versions, resolved device, free VRAM per GPU, resident model, available models, the multilingual language list, and the tags turbo understands. Use it as a preflight.

tts_unload

Releases the resident model and returns its VRAM. Needed when another process wants the card. Only one model is resident at a time — asking for a different one evicts the previous automatically — but this frees it entirely.

list_voices

The reference clips available for cloning, with names to pass as voice=.

stop_speech

Cuts off playback started with wait=false.

Voice cloning

Put reference clips in voices/ (create it; it is not tracked by git). Then refer to a clip by its filename stem:

voices/seven.wav   →   speak(text, voice="seven")

Names resolve case-insensitively, and a unique prefix works (voice="sev"). An ambiguous prefix returns the candidates rather than picking one.

How much reference audio to use

Give it as much clean single-speaker speech as you have. Do not trim a good recording down to a short excerpt.

This tool's own documentation used to say "5-15 seconds". That is wrong, and it was wrong in the direction that costs quality, so it was replaced with what an A/B test actually found. Comparing a 40.1 s reference against an 11.8 s excerpt of that same recording — five texts per arm, plus a same-reference re-render arm as the noise floor — gave this for impulsive discontinuities per second (adaptive Laplacian impulse detector):

Threshold

40.1 s reference

11.8 s excerpt

noise floor

k=6 fine

283.6

459.0

5.1

k=10 moderate

108.8

173.6

7.7

k=20 large

72.5

80.1

5.2

The short reference produced ~36 % more moderate and ~44 % more fine impulses, several times the model's own sampling variance, and the effect survived controlling for gain. Worst-spike magnitude, HF energy and zero-crossing irregularity were indistinguishable, so what degrades is fine crackle rather than loud pops. The failure mode is too little audio, not too much.

We did not test 20 s, 60 s or longer, so this does not establish an optimum — only that truncating a good recording to "5-15 s" measurably degrades it. Under about 5 s is genuinely too little.

Three more things that measurement contradicted:

  • A processed or mastered source is fine. The 40 s reference above was an Audacity-mastered file and was the better of the two. Compression is not the problem; shortness is.

  • Do not normalise, gain-match or limit the reference. Applying +23 dB of peak normalisation to the excerpt moved rendered output level by 12.1 dB (RMS −23.1 vs −35.2 dBFS, against a 0.35 dB level noise floor), and the hotter output was measurably grainier. Leave the source level alone.

  • Sample rate and channel count need not match anything. The 40 s reference was 48 kHz stereo and needed no preparation; Chatterbox resamples and downmixes internally.

What you still control: no music, no background noise, no second speaker, and prefer continuous speech over silence-padded audio.

Caveat on all of the above: n=5 per arm on a single speaker, and the detector measures discontinuities, not timbre, prosody or speaker similarity. It quantifies one artefact class, not overall quality — your own listening is still the arbiter.

list_voices reports each clip's duration, sample rate and channel count, and flags clips short enough to be worth warning about, precisely because duration is what predicted quality here.

CLI

# diagnostics: device, per-GPU VRAM, models, voices
uv run mcp-agent-chatterbox doctor

# synthesize without going through MCP
uv run mcp-agent-chatterbox speak "Build finished." --model turbo
uv run mcp-agent-chatterbox speak "Guten Morgen." --model multilingual --voice seven --language de
uv run mcp-agent-chatterbox speak "..." --no-play --wait

# run the server directly
uv run mcp-agent-chatterbox serve

The CLI speak subcommand shares its code path with the MCP tool (mcp_agent_chatterbox.speak.speak_once), so it is the fastest way to check a fresh install.

Web surface (HTTP transport)

Yes, there is one, and it is the least-comfortable part of this server. Read this before binding it to anything but loopback.

# streamable-http on loopback (the default host)
uv run mcp-agent-chatterbox serve --http --port 8123
# endpoint: http://127.0.0.1:8123/mcp

serve takes --http (a flag), not --transport <name>. The underlying server.run() also accepts "sse", but no CLI path reaches it — it is library-only and deprecated in the MCP spec. Prefer streamable-http.

What was verified. A real MCP client over real uvicorn: 5 tools advertised, tts_status, speak and tts_unload all correct, plus a 4-client concurrent burst on a cold engine that produced exactly one weight load. That last one is the load-bearing test — see the concurrency note below.

No authentication. None. There is no token, no TLS, no per-client identity. --host defaults to 127.0.0.1, which is the only thing keeping this off your network. --host 0.0.0.0 publishes an unauthenticated text-to-speech endpoint to every host that can route to you. Do not do that on an untrusted network.

One process, one engine, all clients share it. The HTTP server builds a single AppContext, so:

  • stop_speech() stops playback for every connected client, not just yours.

  • tts_unload() frees the GPU out from under any in-flight speak.

  • There is no per-client output directory; all sessions write to the same output_dir.

Concurrent clients are a queue, not a parallel server. Measured with 4 simultaneous clients: per-call wall times of 1.63s / 3.34s / 4.72s / 6.22s — a staircase, because synthesize() holds an engine-wide lock (see below). Four clients took 6.4s where a parallel server would take ~1.6s. That is the correct trade for one GPU and one resident model: correctness over throughput. tts_status deliberately takes no lock and stays instant even mid-queue.

The concurrency bug this surface would have had. The engine was written for stdio, where the transport serialises one client's calls, so load() was an unlocked check-then-act over a process-global model cache. Under HTTP that is a live race: two threads both miss the cache check and both call from_pretrained() — 4 copies of turbo's weights on a 12 GB card that the VRAM preflight had already cleared — and a request for a different model could gc.collect() + torch.cuda.empty_cache() while another request was still generating. load(), unload() and synthesize() are now serialised behind a re-entrant lock, and tests/test_concurrency.py fails if that lock is removed.

Treat it as single-operator. It is fine for a local dashboard, a second harness on the same box, or testing. It is not an authenticated service.

Configuration

All settings are environment variables, read at startup. No secrets.

Variable

Default

Meaning

CHATTERBOX_MODEL

turbo

Model used when speak names none

CHATTERBOX_DEVICE

auto

auto | cpu | cuda | cuda:N

CHATTERBOX_GPU_INDEX

unset

Pin a CUDA ordinal. Overrides the auto heuristic

CHATTERBOX_OUTPUT_DIR

tts_output

Where WAVs are written

CHATTERBOX_VOICES_DIR

voices

Where reference clips live

CHATTERBOX_LOGS_DIR

logs

Rotating log file location

CHATTERBOX_AUTOPLAY

1

Play audio after writing

CHATTERBOX_MAX_CHARS

4000

Reject longer single requests

CHATTERBOX_MAX_CHUNK_CHARS

500

Long text auto-splits into sentence-aligned chunks of up to this many chars

CHATTERBOX_CHUNK_PAUSE_MS

250

Silence inserted between concatenated chunks

CHATTERBOX_VRAM_MB

4096

Free-VRAM floor for a load attempt

CHATTERBOX_STRICT_VRAM

0

1 refuses a load below the floor instead of warning

HF_HOME

—

Override the Hugging Face cache location

Relative paths resolve against the project root, never the process CWD, since MCP harnesses spawn servers with an unpredictable working directory.

Choosing a GPU

auto picks the CUDA device with the most free memory, which matters when one card is occupied by something else — a loaded LLM, a rendering job. On a box where LM Studio holds both GPUs, doctor shows the headroom per card and CHATTERBOX_GPU_INDEX pins a specific one.

Design notes

A card that cannot be queried is skipped, not fatal. If torch.cuda.mem_get_info() throws for every ordinal — a driver that has not finished initialising, a card held by another process — auto falls back to CPU instead of raising. A gpu_index pin outside the visible range is the one case that errors, because there the user asked for a specific card and silently using another would be worse than failing.

One resident model. MAX_RESIDENT = 1. A 350M turbo model and a 500M multilingual model do not both fit on a 12 GB card that already has an LLM resident. Asking for a different model evicts the previous one and empties the CUDA allocator cache.

The engine is locked; the status reads are not. load(), unload() and synthesize() serialise behind one re-entrant lock, because the model cache is process-global and a single model instance is not safe to call generate() on from two threads. tts_status, gpu_report and free_vram_mb() deliberately take no lock, so a ten-second cold weight load never stalls a diagnostic call — which is the call you need while staring at a slow first request. The lock is invisible over stdio and is the difference between working and OOM-ing over HTTP; the reasoning and the measured numbers are in the web-surface section.

The device string goes straight to from_pretrained(). Chatterbox's own .to(device) is incomplete upstream: in ChatterboxTTS and ChatterboxMultilingualTTS it moves only t3 and gen, leaving ve, s3gen and conds on the old device and never updating self.device — so generate() would then move tensors to the wrong place. In ChatterboxTurboTTS it iterates a gen attribute that __init__ never sets, so it raises AttributeError. There is therefore no load-on-CPU-then-move path, which is why free VRAM is checked before a load instead of after a failure.

VRAM is checked before, not during. Because weights land directly on the device, a load started with too little headroom dies part way through reading the checkpoint with a bare torch.cuda.OutOfMemoryError, after the download. The engine queries free VRAM first, warns (or refuses under CHATTERBOX_STRICT_VRAM), and maps a genuine OOM onto an actionable cuda_out_of_memory payload.

Weights are never loaded at startup. A status or voice-listing call must not pay the multi-second torch import plus a multi-gigabyte fetch.

Troubleshooting

cuda_out_of_memory / "not enough free VRAM". Something else is holding the card. Check doctor for per-GPU free memory, unload the model in LM Studio, or set CHATTERBOX_GPU_INDEX to a freer card.

reference_clip_required. multilingual and original have no stock voice. Add a clip to voices/ and pass voice=, or use model="turbo".

voice_not_found. Run list_voices for the exact names.

unsupported_language. language takes an ISO 639-1 code and only applies to multilingual. tts_status lists all 23.

First call is very slow. That is the Hugging Face download. Later calls reuse the cache.

Audio does not come out. Check played and playback_note in the response. Playback needs Windows; on other platforms the WAV is still written. stop_speech clears anything queued.

turbo and the watermarker. Turbo's output carries Resemble AI's watermarking mark, added by the vendored Perth library. Turning it off would be a licence question, not a technical one, so it stays on.

Logs. logs/mcp_agent_chatterbox.log (rotating, 5 MB × 5). Unhandled tracebacks go there, never into the model's context. The console handler writes to stderr with errors="replace", so a character the Windows console code page cannot represent degrades to ? instead of raising or garbling the line — Chatterbox logs a ✅ of its own, so this is not hypothetical.

Development

uv sync --extra dev
uv run pytest              # unit tests, no model download
uv run ruff check .
uv run ruff format .

Tests stub torch and the Chatterbox classes, so the suite runs in seconds and does not touch the network or the GPU. For a real end-to-end check use mcp-agent-chatterbox speak.

License

MIT — see LICENSE. Chatterbox itself is MIT, © Resemble AI.

Available Tools

5 tools
list_voicesA

List the reference clips available for voice cloning — the files in the voices directory. Returns each clip's name (what you pass as voice=), filename, size, duration, sample rate and channel count, plus which models can use it. The multilingual and original models require one of these; turbo does not.

For good cloning use as much clean single-speaker audio as you have, not a short excerpt — clipping a good recording down measurably adds fine clicks and crackle. A mastered or compressed source is fine; do not normalise it. Sample rate and channel count need not match anything. Clips too short to be safe carry a quality_note.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and does well: it describes the returned fields (name, filename, size, duration, sample rate, channel count, model compatibility) and notes that too-short clips carry a quality_note. It does not explicitly state that the operation is read-only or mention any auth/rate-limit behavior, but those are largely implied by 'List' and not critical here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first paragraph is front-loaded and efficient. The second paragraph, while relevant to voice cloning, contains detailed audio-quality advice ('clipping... adds fine clicks and crackle', 'do not normalise it') that is not necessary for selecting or invoking list_voices and adds tangential length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter listing tool with an output schema, the description is complete enough: it explains what is returned, notes the quality_note field, and clarifies which models depend on these clips. No critical information for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. The description mentions voice= only as context for other tools, not as a parameter of this one, and the schema is empty.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'List the reference clips available for voice cloning — the files in the voices directory.' It is immediately distinguishable from siblings like speak, stop_speech, tts_status, and tts_unload, all of which perform different actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides clear context for when the listed clips are needed: 'The multilingual and original models require one of these; turbo does not.' There are no alternative listing tools to compare against, but the description does not explicitly state exclusions or prerequisites beyond model compatibility.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

speakA

Speak text aloud with Chatterbox on the local GPU, save it as a WAV, and by default play it through the system audio device.

model: "turbo" (default, 350M, English, no reference clip needed, supports inline [laugh]/[chuckle] tags), "multilingual" (500M, 23 languages, requires a reference clip) or "original" (500M, English, requires a reference clip). voice: name of a reference clip in the voices directory (its filename without extension). Required for multilingual/original. reference_clip: path to a wav/mp3/flac clip, as an alternative to voice=. Supply as much clean continuous speech as you have. Short references measurably increase fine clicks and crackle: a 40 s reference beat an 11.8 s excerpt of the same recording by a wide margin, several times the model's own run-to-run variance. Under ~5 s is genuinely too little. Do not normalise or limit the clip -- that shifts output level and adds artefacts. A mastered or compressed recording is fine. Sample rate and channel count need not match anything. language: ISO 639-1 code, multilingual only (e.g. "de", "en", "fr"). t3_model: "v2" (default) or "v3" — multilingual checkpoint. temperature: sampling temperature (0.05-5.0). Honored by all models. top_p: nucleus-sampling cutoff (0.0-1.0). Honored by all models. top_k: top-k sampling size (0-1000). Honored by turbo only — the 500M models have no such parameter. repetition_penalty: penalise repeated tokens (1.0-2.0). Honored by all models. norm_loudness: normalize output to -27 LUFS. Honored by turbo only — the 500M models have no such parameter. exaggeration/cfg_weight: style controls (0.0-2.0 / 0.0-1.0). Honored ONLY by the 500M models (multilingual/original); turbo ignores both, so setting them with model="turbo" is dropped with a warning, not an error. cfg_weight>0 doubles the text tokens for CFG guidance. seed: reseed torch (CPU + CUDA) so re-renders are reproducible within the resident session; 0 (the upstream convention) or unset keeps random sampling. Byte-identical output across a server restart is not guaranteed — Chatterbox is nondeterministic across CUDA kernel choices. Left unset, every knob keeps the model's own tuned default (e.g. turbo runs cfg_weight 0.0 with top_k 1000; the 500M models run cfg_weight 0.5). play_audio: set false to only write the file. Defaults to the server setting (on). wait: block until playback finishes instead of returning immediately. filename: output basename; a safe name is generated when omitted.

The first call downloads the model from Hugging Face (hundreds of MB to a few GB) and can take a while; later calls reuse the cached weights. Returns JSON with the output filename, duration, sample rate, model, device and whether it played.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNo
textYes
waitNo
modelNo
top_kNo
top_pNo
voiceNo
filenameNo
languageNo
t3_modelNo
cfg_weightNo
play_audioNo
temperatureNo
exaggerationNo
norm_loudnessNo
reference_clipNo
repetition_penaltyNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With zero annotations, the description carries the full burden and does so well: it discloses the first-call model download (hundreds of MB to a few GB, slow) and subsequent cache reuse, the file-write and audio-play side effects, that play_audio defaults to a server setting, that wait blocks, and that a safe filename is auto-generated. It also surfaces edge behavior an agent could not guess — ignored parameters produce a warning not an error, cfg_weight doubles token count, and byte-identical output is not guaranteed across a restart due to CUDA nondeterminism.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The one-line purpose is front-loaded, followed by a consistent per-parameter reference block, which is an efficient structure for 17 parameters. It runs long and the reference-clip anecdote (40 s vs 11.8 s excerpt) is more detailed than strictly necessary, but overall each entry earns roughly its place given the parameter count.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 17-parameter, annotation-free synthesis tool, the description is complete: it covers defaults, model/param interactions, side effects, download latency, and even summarizes the JSON return. With an output schema present it correctly does not need to detail return structure, so nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must supply all parameter meaning, and it does: every knob (model, voice, reference_clip, language, t3_model, temperature, top_p, top_k, repetition_penalty, norm_loudness, exaggeration/cfg_weight, seed, play_audio, wait, filename) gets a definition, valid range, and the model-compatibility rule (turbo-only vs 500M-only params). It even adds non-obvious practical guidance on reference-clip length and loudness normalization that the schema could never convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb and resource: synthesize text with a named engine, write a WAV, and by default play it. It clearly reads as the 'do the speech synthesis' action versus the lifecycle siblings (tts_status, tts_unload, stop_speech, list_voices). However, it never names or contrasts those siblings explicitly, so distinctness is inferred rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains the conditions that select a model (turbo needs no reference clip; multilingual/original require one) and the meaning of defaults, which is useful context. But there is no explicit when-to-use/when-not guidance and no routing to siblings — e.g. it never says to use list_voices to discover a valid voice name or stop_speech to interrupt playback.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stop_speechA

Stop audio that is currently playing. Useful for cutting off a long utterance started with wait=false.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It conveys the core effect (stopping playing audio) and a usage trigger, but omits details such as whether it stops all playback or only the current utterance, and what happens when nothing is playing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences, front-loaded with the core action and followed by a relevant usage note. Every sentence earns its place with no wasted wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (no parameters, output schema present), the description covers the essential purpose and one usage context. It falls short only by not distinguishing this tool from related siblings like tts_unload, which an agent may need to choose correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is no parameter semantics to explain. The empty input schema is self-explanatory, and the description appropriately does not add unnecessary parameter details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Stop audio that is currently playing.' This clearly identifies the operation, but it does not differentiate the tool from siblings such as tts_unload, which may also stop audio, so the sibling-routing aspect is missing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides a clear usage scenario: 'Useful for cutting off a long utterance started with wait=false.' However, it does not name alternative tools or state when not to use this tool, leaving the agent to infer the boundary with tts_unload.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tts_statusA

Report Chatterbox TTS readiness: installed packages, resolved device, per-GPU free/total VRAM, the resident model, the models available, the languages the multilingual model speaks, and the paralinguistic tags turbo understands. Loads no weights — safe as a preflight before a speak call. On a machine where a large model already occupies a GPU, check free_mb here first: a card with less than ~4096 MB free will fail to load.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full safety burden and does so: 'Loads no weights — safe as a preflight' discloses the non-mutating behavior, and it warns that a card with under ~4096 MB free will fail to load. It omits any mention of latency or whether the check itself touches the GPU.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the primary purpose and then the safety note and the failure heuristic. The opening enumeration is long but each item is genuinely informative, so it earns its length without much waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter status tool with an output schema, the description need not explain return values, and it covers the remaining critical context: that the call is weight-free and what VRAM threshold causes downstream load failures. Nothing an agent needs before calling is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline for a parameterless tool is 4. The description adds no misleading parameter-like details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Report') and resource ('Chatterbox TTS readiness'), then enumerates exactly what is reported: packages, device, VRAM, resident/available models, languages, and tags. It also positions itself against siblings by naming its role as a preflight before a speak call.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly frames the tool as a preflight before a speak call and tells the agent to check free_mb here first when a GPU is occupied. It does not, however, state when not to use it or contrast it with list_voices or other siblings, so guidance is clear but incomplete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tts_unloadA

Release the currently loaded Chatterbox model and return its VRAM to the GPU. Use this when another process needs the card, or to drop a 500M model before a lighter turbo load. Safe to call when nothing is loaded.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does reasonably well: it discloses the side effect (model unloaded, VRAM freed) and idempotency ('safe to call when nothing is loaded'). It stops short of saying whether in-flight speech is interrupted or whether a reload is required afterward.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action, then the trigger conditions, then the safety reassurance. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary, and a zero-parameter unload tool needs little else. The only meaningful gap is the absence of explicit routing versus the sibling stop_speech.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the baseline there is nothing for the description to document. No parameter-level gaps exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: release the loaded Chatterbox model and return its VRAM to the GPU. An agent can tell this is an unload/teardown operation. It does not, however, explicitly differentiate itself from the sibling stop_speech, which arguably also halts model activity, so it falls short of the 5 bar.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete triggering scenarios: another process needs the card, or dropping a 500M model before a lighter turbo load. It also preempts the caller's main worry with 'Safe to call when nothing is loaded.' No explicit when-not-to-use or naming of stop_speech as the lighter-weight alternative, so not a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.1.0
    • First observedlist_voices
    • First observedspeak
    • First observedstop_speech
    • First observedtts_status
    • First observedtts_unload

TDQS

A4.2/5.0

Scored across 5 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: tts_status reports readiness, tts_unload releases VRAM, speak synthesizes audio, list_voices enumerates reference clips, and stop_speech halts playback. There is no overlap or boundary confusion between any pair.

Naming Consistency3/5

Naming is mixed: two tools use a 'tts_' prefix (tts_status, tts_unload) while the others do not, and 'speak' is a bare verb while list_voices/stop_speech follow verb_noun. It is still readable snake_case, but the conventions are not predictable as a set.

Tool Count5/5

Five tools is well-scoped for a local TTS server: status, unload, speak, list_voices, and stop_speech each cover a necessary lifecycle operation without redundancy. Nothing feels thin or bloated.

Completeness4/5

The surface covers the core TTS lifecycle (preflight, synthesize/play, list voices, stop, unload) and the speak tool exposes rich model controls. Minor gaps exist, such as no explicit preload of a model or enumeration of previously generated output files, but agents can work around these.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers