Skip to main content
Glama
AIM-IT4
by AIM-IT4
README.md
# VoiceStudio MCP for Desk2Quant

A deployable MCP gateway and private plugin source package for
[debpalash/VoiceStudio](https://github.com/debpalash/VoiceStudio).
It keeps inference in VoiceStudio and exposes its existing native tools plus
its broader HTTP API over Streamable HTTP or local stdio.

**Start here: [DEPLOY.md](DEPLOY.md).** The ZIP is source, not a running server.
After deployment your private MCP endpoint is `https://<your-gateway>/<secret>/mcp`.
No live domain or credentials are embedded in this package.

## What is included

| Capability | How it is accessed |
|---|---|
| Speech / narration | Native `generate_speech`; advanced parameters via HTTP API |
| Voice cloning / voice design | Native `clone_voice`, `describe_voice`, `design_voice` |
| Transcription | Native `transcribe`; multipart API for larger recordings |
| Voice / personality / language discovery | Native `list_voices`, `list_personalities`, `list_languages` |
| Video dubbing, translation and subtitles | Discover, inspect and call the live `/dub/*` and related operations |
| Audiobooks and long-form production | Live audiobook / longform operations, including finite SSE rendering |
| Batch jobs, projects, conversion and pronunciation | Live API discovery and execution |
| Models, engines, watermarking, settings and workers | Live API discovery; upstream permission and hardware rules apply |
| Audio, video, subtitles and manuscript files | Signed uploads, repeated multipart file fields and timed downloads |
| Long operations | Bounded background jobs and persistent status without automatic retry |
| MCP voice/history resources | `voicestudio_read_resource` |

There are **9 native tools + 12 gateway tools** in the bundled schema. Native
discovery also picks up additional tools exposed by your backend. The source
catalog contains **337 HTTP routes** inspected at the commit in `UPSTREAM.json`;
actual operation IDs, schemas and availability always come from your running
backend's OpenAPI document. This is API coverage, not a claim that every desktop
feature or every engine is usable on every host.

## Package files

- `Dockerfile`, `railway.json`: deploy the small gateway.
- `docker-compose.yml`, `docker-compose.gpu.yml`: gateway + official backend image.
- `.env.example`: configuration; `scripts/init_env.py` generates fresh secrets.
- `plugin.json`, `mcp.json`: portable **local** plugin with a real stdio entrypoint.
- `scripts/export_remote_plugin.py`: verify your deployed endpoint and make a
  separate remote plugin ZIP containing its actual URL.
- `scripts/smoke_test.py`: safe deployed check of discovery and status.
- `scripts/upload_file.py`: upload a local file without putting base64 in a chat.
- `CAPABILITIES.md`: coverage and limitations.
- `VERIFICATION.md`: what was actually tested.
- `uv.lock`: frozen Python dependency resolution.

## Important limits

1. A running VoiceStudio backend and installed models are required. The gateway
   does not contain model weights and cannot synthesize audio by itself.
2. A small Railway/Render instance can host the gateway. Running the complete
   VoiceStudio inference workload there is a separate resource decision; do not
   assume a 1 GB/free instance will support large speech models or video dubbing.
3. Native desktop controls (microphone widget, OS hotkeys, native file pickers,
   reveal-in-folder, desktop runtime management) and continuous WebSocket
   sessions remain in the original app. Finite HTTP/SSE workflows are supported.
4. The initial `mcp.json` is local. A ChatGPT web/mobile connection needs the
   deployed HTTPS endpoint or the remote plugin exported after verification.
   Installing a ZIP does not start hosting. Availability depends on the host's
   current custom-MCP/plugin support; this package does not promise mobile-only
   installation.
5. Secret URLs are a private single-user access method. They are not OAuth.
   Treat the full MCP URL as a password. Use an OAuth gateway for shared/public
   distribution. Optional bearer authentication is supported for clients that
   send headers; this package does not implement an OAuth authorization server.
6. Hosting, GPU resources, optional paid providers, telephony and some models
   can incur costs. Open source does not make these services free.
7. Upstream is AGPL-3.0; model licenses differ. This gateway source is provided
   under the included AGPL-3.0 license. Only clone a speaker's voice with permission.

No changes to Desk2Quant, its repository, payments or existing services are
needed to use this separate package.

## Development

```bash
uv sync --frozen
uv run pytest
```

For local stdio, set `VOICESTUDIO_URL` to your backend, then run
`uv run voicestudio-mcp --transport stdio`. Logs go to stderr.

## Sources

- [Upstream MCP documentation](https://github.com/debpalash/VoiceStudio/blob/main/docs/mcp.md)
- [Upstream Docker instructions](https://github.com/debpalash/VoiceStudio/blob/main/docs/install/docker.md)
- [Upstream authentication](https://github.com/debpalash/VoiceStudio/blob/main/docs/api-auth.md)
- [Official Python MCP SDK](https://github.com/modelcontextprotocol/python-sdk)

The package pins MCP SDK 1.28.1 to match the inspected upstream integration;
it does not depend on the current SDK main branch's v2 interface.

TDQS

B3.4/5.0

Scored across 21 tools

Disambiguation3/5

Most native tools (generate_speech, transcribe, clone_voice, design_voice, describe_voice) have distinct purposes, but check_health and voicestudio_status clearly overlap (both report backend health/GPU), and the generic API layer (search_api/get_operation/call_api/call_native) creates unclear boundaries against the native tools—agents must decide which interface to use. The dual upload paths (create_upload vs upload_base64) are clarified by descriptions but still close.

Naming Consistency3/5

Two conventions coexist: a large voicestudio_* prefixed family (status, search_api, call_api, start_job, upload_base64, etc.) and an unprefixed native family (generate_speech, clone_voice, transcribe, list_voices, check_health). Both are snake_case and readable, but the mix—and especially check_health being unprefixed while the overlapping voicestudio_status is prefixed—makes the set feel inconsistent.

Tool Count3/5

21 tools sits in the heavy/borderline range for a voice platform. The native voice tools (7) and job/file helpers are justified, but the sizeable generic gateway meta-layer (search/get/call API, call_native, start/status/cancel job, upload, file_info, read_resource) inflates the count and adds surface that largely mirrors what the native tools already do.

Completeness4/5

The surface covers the core lifecycle: synthesis, transcription, cloning, design+preview, voice/personality/language listing, file upload, and background job control, plus a dynamic API passthrough for anything unlisted. Minor gaps exist (no explicit delete/update/rename of profiles, history only reachable via read_resource), but these are workable.

Maintenance

ActivityMaintained
ResponsivenessNo issues