Skip to main content
Glama
AIM-IT4
by AIM-IT4

clone_voice

Create a new voice profile from a reference audio sample, then use the returned profile_id to generate speech.

Instructions

Clone a new voice profile from a reference audio sample.

    The new voice is immediately available for use with generate_speech
    (pass the returned profile_id as the profile_id argument). Pass
    exactly one of ref_audio_base64 or ref_audio_path.

    Args:
        name: A human-friendly name for the cloned voice.
        ref_audio_base64: Base64-encoded audio (WAV, MP3, FLAC, etc.) of
            the reference voice — 5-30 seconds of clean single-speaker
            speech.
        ref_text: Optional transcript of the reference audio (improves
            quality for some engines).
        instruct: Optional style instruction (e.g. 'whisper', 'excited').
        language: Language of the reference audio (ISO code or 'Auto').
        ref_audio_path: Path to the reference audio under
            OMNIVOICE_MCP_BASE_PATH (relative to it, or absolute inside
            it); refused when no base path is configured. Prefer this
            lane for LLM agents - the clip never enters the context.

    Returns:
        JSON with the new profile's id, name, and kind.
     For long operations use voicestudio_start_job.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
nameYes
instructNo
languageNoAuto
ref_textNo
ref_audio_pathNo
ref_audio_base64No

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose several non-obvious traits: immediate availability of the new voice, the base-path refusal condition, and the mutually exclusive audio-input rule. It omits auth/permission requirements, duplicate-name behavior, and any cost or rate-limit context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first sentence front-loads the purpose and is followed by a well-organized Args/Returns block, so scanning is easy. It is slightly verbose for a six-parameter tool, but each line adds meaning rather than restating the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a creation tool with no annotations and a documented output schema, the description covers the creation semantics, the input constraint, the fallback routing, and the return shape. Remaining gaps are error/failure behavior beyond the base-path refusal and any permission requirements.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry parameter meaning, and it documents all six parameters with real semantic value: duration and cleanliness guidance (5-30 seconds, single speaker) for the audio, ISO code or 'Auto' for language, and example style instructions for instruct. This fully compensates for the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb and resource: cloning a voice profile from a reference audio sample. It clearly separates this from generate_speech by framing the output as an input to it. It never names the closest sibling (design_voice), so an agent must infer the difference between cloning from a sample versus designing a voice.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the core constraint (pass exactly one of ref_audio_base64 or ref_audio_path), recommends a preferred lane for LLM agents (ref_audio_path, so the clip never enters context), and routes long operations to voicestudio_start_job. It also tells the agent how to consume the result via generate_speech.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.