voxcpm-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| VOXCPM_PYTHON | No | Python executable with VoxCPM2 + CUDA (e.g., path to venv python) | |
| VOXCPM_OUTPUT_DIR | No | Directory where WAV files are saved | ./voxcpm_output |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| synthesizeA | Synthesize speech from text using VoxCPM2 (2B diffusion TTS, 48 kHz). Returns the path to the output WAV file and its duration. |
| synthesize_with_cloneA | Synthesize speech cloning a voice from a reference WAV. The reference WAV sets the speaker identity, prosody, and style. Both reference and output are 48 kHz mono WAV. |
| preload_modelA | Load VoxCPM2 into VRAM now (takes ~10 s on RTX 4060 Laptop). Call this before synthesize if you want the first synthesis to be fast. |
| pingA | Check that the VoxCPM2 worker subprocess is alive. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 4 tools
Each tool targets a distinct operation: health check, model preloading, standard synthesis, and cloned synthesis. There is no overlap in functionality.
All tool names use a consistent snake_case pattern with clear, imperative verbs (ping, preload, synthesize, synthesize_with_clone).
Four tools covers the essential workflow of a TTS server: ensure service is alive, preload model for fast synthesis, generate speech, and clone voice. No extraneous or missing tools.
Core synthesis and cloning are covered, but missing advanced features like listing available voices or managing models. Minor gap for production use.