Skip to main content
Glama
AIM-IT4
by AIM-IT4

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
VOICESTUDIO_URLYesThe URL of your VoiceStudio backend. For local stdio, set this to your backend before running the server.

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
voicestudio_statusB

Check backend health, native MCP discovery, HTTP API operation count and file-transfer configuration.

voicestudio_search_apiB

Search every live VoiceStudio OpenAPI HTTP operation, including dubbing, audiobooks, conversion, models, batch jobs, profiles, pronunciation, projects, watermarking, settings and workers. Offline source catalog is informational only.

voicestudio_get_operationA

Get a live operation's exact request parameters, requestBody, responses and referenced component schemas. Inspect this before calling it.

voicestudio_call_apiA

Execute one discovered HTTP operation against the configured backend. JSON, forms, repeated multipart files and finite SSE responses are supported. Binary outputs become timed file links. Mutations may delete data, change settings, download models, contact providers or place calls; use only within the user's request. Use start_job for long renders.

voicestudio_call_nativeB

Call a discovered native MCP tool with its original arguments. For uploaded reference/audio files use file_arguments={ref_audio_base64:file_id} for clone_voice or {audio_base64:file_id} for transcribe. Large audio uses multipart HTTP API instead.

voicestudio_start_jobA

Start a background operation and immediately return a gateway job ID. For kind=api use the operation ID as name and call_api arguments without operation_id. For kind=native use a native tool name and arguments={arguments:{...},file_arguments:{...}}. Gateway status persists; work is not automatically retried after restart.

voicestudio_job_statusB

Poll a gateway background job. Backend batch/dub job IDs are separate; inspect their HTTP operations for progress and cancellation.

voicestudio_cancel_jobA

Cancel the gateway's wait. Backend synthesis may continue. To abort backend work use its discovered cancel/abort HTTP operation.

voicestudio_create_uploadA

Create a file slot and a one-use signed PUT URL for direct audio, video, manuscript or subtitle upload. Send raw file bytes to that URL; then use its file_id in API multipart fields. Requires PUBLIC_BASE_URL for remote clients.

voicestudio_upload_base64A

Stage a small file (default max 8 MiB) from base64. Prefer direct signed PUT upload for large files; never ask the user to paste base64 into chat.

voicestudio_file_infoB

Inspect a staged/generated file and renew its timed download link while retained. Links provide access to anyone holding them; do not publish private recordings.

voicestudio_read_resourceA

Read a native MCP resource such as voice:// or history://recent. Only the backend's advertised resource namespaces are accepted.

generate_speechA

Generate speech audio from text.

    Args:
        text: The text to synthesize into speech.
        language: Target language (ISO code or 'Auto'). 646 languages
            supported. Omit to use the voice profile's saved language;
            an explicit 'Auto' overrides it.
        profile_id: ID of a saved voice profile to clone. Omit to use this
            agent's bound voice (Settings → MCP), else the default voice.
        instruct: Style instruction (e.g. 'whisper', 'excited', 'narrator').
        speed: Speech speed multiplier (0.5–2.0, default 1.0).
        steps: Diffusion steps (8=fast/draft, 16=balanced, 32=quality).
        format: File and URL format: wav (default), ogg or opus. Both
            ogg and opus carry Opus in Ogg; requires files/both mode and ffmpeg.

    Returns:
        JSON with audio_id, generation_time_s, audio_duration_s and the
        audio shaped by OMNIVOICE_MCP_OUTPUT_MODE: base64 WAV data
        ('resources', the default), a URL plus an optional file ('files'),
        or both ('both'). Prefer 'files' for LLM agents.
     For long operations use voicestudio_start_job.
list_voicesA

List all saved voice profiles.

Returns a JSON array of voice profiles with id, name, type (clone/design), and personality.

list_personalitiesA

List available voice personality presets.

    Returns presets like Narrator, Casual, News Anchor, etc. with their
    instruct text. Use the instruct text with generate_speech.
    
list_languagesA

List a sample of supported TTS languages.

VoiceStudio supports 646 languages. This returns the most popular ones plus a note about the full count.

transcribeA

Transcribe spoken audio to text.

    Pass exactly one of audio_base64 or audio_path.

    Args:
        audio_base64: Base64-encoded audio bytes (wav/mp3/webm/m4a).
        audio_path: Path to an audio file under OMNIVOICE_MCP_BASE_PATH
            (relative to it, or absolute inside it). The base path is the
            security boundary: with none configured, paths are refused.
            Prefer this lane for LLM agents - the audio never enters the
            agent's context.
        language: Optional language hint; omit for auto-detect.

    Returns:
        JSON with the recognized text, language, and duration.
     For long operations use voicestudio_start_job.
check_healthA

Check if the VoiceStudio backend is running and what GPU device is active.

clone_voiceA

Clone a new voice profile from a reference audio sample.

    The new voice is immediately available for use with generate_speech
    (pass the returned profile_id as the profile_id argument). Pass
    exactly one of ref_audio_base64 or ref_audio_path.

    Args:
        name: A human-friendly name for the cloned voice.
        ref_audio_base64: Base64-encoded audio (WAV, MP3, FLAC, etc.) of
            the reference voice — 5-30 seconds of clean single-speaker
            speech.
        ref_text: Optional transcript of the reference audio (improves
            quality for some engines).
        instruct: Optional style instruction (e.g. 'whisper', 'excited').
        language: Language of the reference audio (ISO code or 'Auto').
        ref_audio_path: Path to the reference audio under
            OMNIVOICE_MCP_BASE_PATH (relative to it, or absolute inside
            it); refused when no base path is configured. Prefer this
            lane for LLM agents - the clip never enters the context.

    Returns:
        JSON with the new profile's id, name, and kind.
     For long operations use voicestudio_start_job.
describe_voiceA

Preview how a voice description maps onto voice-design attributes.

    Nothing is saved. Use this before design_voice to see what the
    description will produce. The design space is small and fixed; only
    these tokens (and close synonyms) are understood:
      Gender: male, female
      Age: child, teenager, young adult, middle-aged, elderly
      Pitch: very low / low / moderate / high / very high pitch
      Style: whisper
      EnglishAccent: american, british, australian, canadian, indian,
        japanese, korean, chinese, russian, portuguese accent
      ChineseDialect (Chinese speech; overrides an English accent):
        sichuan, dongbei / northeastern chinese, henan, shaanxi, gansu,
        guilin, guizhou, jinan, ningxia, qingdao, shijiazhuang, yunnan
        (e.g. "sichuan dialect")
    Timbre words ("gravelly", "raspy") and other accents are ignored and
    reported in `unmatched`. For a voice outside this space, use
    clone_voice with reference audio instead.

    Args:
        description: Free-text description, e.g. "an elderly man with a
            deep voice and a british accent".

    Returns:
        JSON with attrs (category → token or "Auto"), instruct, matched
        and unmatched.
    
design_voiceA

Design and save a new voice profile from a text description.

    The description is mapped onto the same attributes describe_voice
    previews (see its docstring for the vocabulary). The backend tries to
    render a fixed-seed identity sample at save time; if the voice engine
    isn't ready, the profile is still saved and the same sample is
    rendered on first use, so the voice stays stable across
    generate_speech calls either way. Pass the returned profile_id to
    generate_speech. Refuses a description that matches no attribute.

    Args:
        name: A human-friendly name for the new voice.
        description: Free-text description of the voice.
        language: The voice's saved language (ISO code or 'Auto'); used
            for its sample and by generate_speech calls that omit one.

    Returns:
        JSON with the new profile's id, name, kind, the attrs used, and
        any unmatched description fragments.
     For long operations use voicestudio_start_job.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription
Capability coverage
Desk2Quant voice workflows

TDQS

B3.4/5.0

Scored across 21 tools

Disambiguation3/5

Most native tools (generate_speech, transcribe, clone_voice, design_voice, describe_voice) have distinct purposes, but check_health and voicestudio_status clearly overlap (both report backend health/GPU), and the generic API layer (search_api/get_operation/call_api/call_native) creates unclear boundaries against the native tools—agents must decide which interface to use. The dual upload paths (create_upload vs upload_base64) are clarified by descriptions but still close.

Naming Consistency3/5

Two conventions coexist: a large voicestudio_* prefixed family (status, search_api, call_api, start_job, upload_base64, etc.) and an unprefixed native family (generate_speech, clone_voice, transcribe, list_voices, check_health). Both are snake_case and readable, but the mix—and especially check_health being unprefixed while the overlapping voicestudio_status is prefixed—makes the set feel inconsistent.

Tool Count3/5

21 tools sits in the heavy/borderline range for a voice platform. The native voice tools (7) and job/file helpers are justified, but the sizeable generic gateway meta-layer (search/get/call API, call_native, start/status/cancel job, upload, file_info, read_resource) inflates the count and adds surface that largely mirrors what the native tools already do.

Completeness4/5

The surface covers the core lifecycle: synthesis, transcription, cloning, design+preview, voice/personality/language listing, file upload, and background job control, plus a dynamic API passthrough for anything unlisted. Minor gaps exist (no explicit delete/update/rename of profiles, history only reachable via read_resource), but these are workable.

Maintenance

ActivityMaintained
ResponsivenessNo issues