Skip to main content
Glama
AIM-IT4
by AIM-IT4

design_voice

Create a voice profile from a text description. Saves the profile and returns an ID for use with generate_speech.

Instructions

Design and save a new voice profile from a text description.

    The description is mapped onto the same attributes describe_voice
    previews (see its docstring for the vocabulary). The backend tries to
    render a fixed-seed identity sample at save time; if the voice engine
    isn't ready, the profile is still saved and the same sample is
    rendered on first use, so the voice stays stable across
    generate_speech calls either way. Pass the returned profile_id to
    generate_speech. Refuses a description that matches no attribute.

    Args:
        name: A human-friendly name for the new voice.
        description: Free-text description of the voice.
        language: The voice's saved language (ISO code or 'Auto'); used
            for its sample and by generate_speech calls that omit one.

    Returns:
        JSON with the new profile's id, name, kind, the attrs used, and
        any unmatched description fragments.
     For long operations use voicestudio_start_job.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
nameYes
languageNoAuto
descriptionYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: it discloses that a fixed-seed sample is rendered at save time or on first use, that the profile is saved even if the engine isn't ready, that the voice stays stable across generate_speech calls, and that the tool refuses descriptions matching no attribute. These are non-obvious behaviors an agent could not infer from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then behavior, then a structured Args/Returns block, with no filler sentences. The trailing 'For long operations use voicestudio_start_job' is slightly abrupt but useful; the Args/Returns restatement is mild redundancy given an output schema exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter creation tool with no annotations, the definition covers behavior, parameters, failure mode, and downstream usage. An output schema exists, so the Returns section is somewhat redundant, and the cross-tool routing to start_job is a single dangling sentence rather than integrated guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it documents all three parameters: name (human-friendly label), description (free text mapped onto attribute vocabulary), and language (ISO code or 'Auto', used for the sample and by generate_speech calls that omit one). It is nearly complete, though it doesn't enumerate example ISO codes or validate the 'Auto' fallback behavior deeply.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence gives a specific verb and resource: design and save a new voice profile from a text description. It also ties the input vocabulary to a named sibling (describe_voice) and tells the agent to pass profile_id to generate_speech, so its role in the workflow is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear operational context: the description must match a known attribute vocabulary, the result's profile_id feeds generate_speech, and long operations should use voicestudio_start_job. It does not explicitly say when to prefer this over clone_voice, which is the most plausible alternative, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.