mcp-agent-chatterbox
Related Servers
Alternatives to mcp-agent-chatterbox
No user-submitted related servers found.
Related Servers
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to speak and listen in real-time with interruption handling, using local ML models and hot-swappable adapters.32 npmMIT
- AlicenseAqualityCmaintenanceEnables AI agents to perform local audio tasks such as speech synthesis, voice cloning, music and sound effect generation, and audio editing through MCP, with GPU models loaded on demand and released after idle.6MIT
- AlicenseNot gradedqualityDmaintenanceEnables MCP clients to speak by running local voice models using Chatterbox Turbo TTS or Kokoro TTS, with support for voice cloning, paralinguistic tags, and multiple voices.40 npm13MIT
- FlicenseNot gradedqualityDmaintenanceEnables Claude to speak text with an embedded audio player, supporting 54 voices, voice cloning, and playback controls, all running locally.3-
- FlicenseNot gradedqualityBmaintenanceEnables AI agents to generate narration in a single fixed, pre-approved voice, returning audio with per-line timing, hashes, and automatic quality checks on content, speaker identity, and pace, all locally and reproducibly.-
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to generate and play high-quality text-to-speech audio using the Kokoro model, with support for multiple voices, adjustable speaking speed, and audio caching.-
TDQS
Scored across 5 tools
Each tool has a clearly distinct purpose: tts_status reports readiness, tts_unload releases VRAM, speak synthesizes audio, list_voices enumerates reference clips, and stop_speech halts playback. There is no overlap or boundary confusion between any pair.
Naming is mixed: two tools use a 'tts_' prefix (tts_status, tts_unload) while the others do not, and 'speak' is a bare verb while list_voices/stop_speech follow verb_noun. It is still readable snake_case, but the conventions are not predictable as a set.
Five tools is well-scoped for a local TTS server: status, unload, speak, list_voices, and stop_speech each cover a necessary lifecycle operation without redundancy. Nothing feels thin or bloated.
The surface covers the core TTS lifecycle (preflight, synthesize/play, list voices, stop, unload) and the speak tool exposes rich model controls. Minor gaps exist, such as no explicit preload of a model or enumeration of previously generated output files, but agents can work around these.