VoiceStudio MCP
OfficialServer Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| VOICESTUDIO_URL | Yes | The URL of your VoiceStudio backend. For local stdio, set this to your backend before running the server. |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| voicestudio_statusB | Check backend health, native MCP discovery, HTTP API operation count and file-transfer configuration. |
| voicestudio_search_apiB | Search every live VoiceStudio OpenAPI HTTP operation, including dubbing, audiobooks, conversion, models, batch jobs, profiles, pronunciation, projects, watermarking, settings and workers. Offline source catalog is informational only. |
| voicestudio_get_operationA | Get a live operation's exact request parameters, requestBody, responses and referenced component schemas. Inspect this before calling it. |
| voicestudio_call_apiA | Execute one discovered HTTP operation against the configured backend. JSON, forms, repeated multipart files and finite SSE responses are supported. Binary outputs become timed file links. Mutations may delete data, change settings, download models, contact providers or place calls; use only within the user's request. Use start_job for long renders. |
| voicestudio_call_nativeB | Call a discovered native MCP tool with its original arguments. For uploaded reference/audio files use file_arguments={ref_audio_base64:file_id} for clone_voice or {audio_base64:file_id} for transcribe. Large audio uses multipart HTTP API instead. |
| voicestudio_start_jobA | Start a background operation and immediately return a gateway job ID. For kind=api use the operation ID as name and call_api arguments without operation_id. For kind=native use a native tool name and arguments={arguments:{...},file_arguments:{...}}. Gateway status persists; work is not automatically retried after restart. |
| voicestudio_job_statusB | Poll a gateway background job. Backend batch/dub job IDs are separate; inspect their HTTP operations for progress and cancellation. |
| voicestudio_cancel_jobA | Cancel the gateway's wait. Backend synthesis may continue. To abort backend work use its discovered cancel/abort HTTP operation. |
| voicestudio_create_uploadA | Create a file slot and a one-use signed PUT URL for direct audio, video, manuscript or subtitle upload. Send raw file bytes to that URL; then use its file_id in API multipart fields. Requires PUBLIC_BASE_URL for remote clients. |
| voicestudio_upload_base64A | Stage a small file (default max 8 MiB) from base64. Prefer direct signed PUT upload for large files; never ask the user to paste base64 into chat. |
| voicestudio_file_infoB | Inspect a staged/generated file and renew its timed download link while retained. Links provide access to anyone holding them; do not publish private recordings. |
| voicestudio_read_resourceA | Read a native MCP resource such as voice:// or history://recent. Only the backend's advertised resource namespaces are accepted. |
| generate_speechA | Generate speech audio from text. |
| list_voicesA | List all saved voice profiles. Returns a JSON array of voice profiles with id, name, type (clone/design), and personality. |
| list_personalitiesA | List available voice personality presets. |
| list_languagesA | List a sample of supported TTS languages. VoiceStudio supports 646 languages. This returns the most popular ones plus a note about the full count. |
| transcribeA | Transcribe spoken audio to text. |
| check_healthA | Check if the VoiceStudio backend is running and what GPU device is active. |
| clone_voiceA | Clone a new voice profile from a reference audio sample. |
| describe_voiceA | Preview how a voice description maps onto voice-design attributes. |
| design_voiceA | Design and save a new voice profile from a text description. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
| Capability coverage | |
| Desk2Quant voice workflows |
TDQS
Scored across 21 tools
Most native tools (generate_speech, transcribe, clone_voice, design_voice, describe_voice) have distinct purposes, but check_health and voicestudio_status clearly overlap (both report backend health/GPU), and the generic API layer (search_api/get_operation/call_api/call_native) creates unclear boundaries against the native tools—agents must decide which interface to use. The dual upload paths (create_upload vs upload_base64) are clarified by descriptions but still close.
Two conventions coexist: a large voicestudio_* prefixed family (status, search_api, call_api, start_job, upload_base64, etc.) and an unprefixed native family (generate_speech, clone_voice, transcribe, list_voices, check_health). Both are snake_case and readable, but the mix—and especially check_health being unprefixed while the overlapping voicestudio_status is prefixed—makes the set feel inconsistent.
21 tools sits in the heavy/borderline range for a voice platform. The native voice tools (7) and job/file helpers are justified, but the sizeable generic gateway meta-layer (search/get/call API, call_native, start/status/cancel job, upload, file_info, read_resource) inflates the count and adds surface that largely mirrors what the native tools already do.
The surface covers the core lifecycle: synthesis, transcription, cloning, design+preview, voice/personality/language listing, file upload, and background job control, plus a dynamic API passthrough for anything unlisted. Minor gaps exist (no explicit delete/update/rename of profiles, history only reachable via read_resource), but these are workable.