Skip to main content
Glama
AIM-IT4
by AIM-IT4

VoiceStudio MCP for Desk2Quant

A deployable MCP gateway and private plugin source package for debpalash/VoiceStudio. It keeps inference in VoiceStudio and exposes its existing native tools plus its broader HTTP API over Streamable HTTP or local stdio.

Start here: DEPLOY.md. The ZIP is source, not a running server. After deployment your private MCP endpoint is https://<your-gateway>/<secret>/mcp. No live domain or credentials are embedded in this package.

What is included

Capability

How it is accessed

Speech / narration

Native generate_speech; advanced parameters via HTTP API

Voice cloning / voice design

Native clone_voice, describe_voice, design_voice

Transcription

Native transcribe; multipart API for larger recordings

Voice / personality / language discovery

Native list_voices, list_personalities, list_languages

Video dubbing, translation and subtitles

Discover, inspect and call the live /dub/* and related operations

Audiobooks and long-form production

Live audiobook / longform operations, including finite SSE rendering

Batch jobs, projects, conversion and pronunciation

Live API discovery and execution

Models, engines, watermarking, settings and workers

Live API discovery; upstream permission and hardware rules apply

Audio, video, subtitles and manuscript files

Signed uploads, repeated multipart file fields and timed downloads

Long operations

Bounded background jobs and persistent status without automatic retry

MCP voice/history resources

voicestudio_read_resource

There are 9 native tools + 12 gateway tools in the bundled schema. Native discovery also picks up additional tools exposed by your backend. The source catalog contains 337 HTTP routes inspected at the commit in UPSTREAM.json; actual operation IDs, schemas and availability always come from your running backend's OpenAPI document. This is API coverage, not a claim that every desktop feature or every engine is usable on every host.

Related MCP server: ChatATP Studio MCP Server

Package files

  • Dockerfile, railway.json: deploy the small gateway.

  • docker-compose.yml, docker-compose.gpu.yml: gateway + official backend image.

  • .env.example: configuration; scripts/init_env.py generates fresh secrets.

  • plugin.json, mcp.json: portable local plugin with a real stdio entrypoint.

  • scripts/export_remote_plugin.py: verify your deployed endpoint and make a separate remote plugin ZIP containing its actual URL.

  • scripts/smoke_test.py: safe deployed check of discovery and status.

  • scripts/upload_file.py: upload a local file without putting base64 in a chat.

  • CAPABILITIES.md: coverage and limitations.

  • VERIFICATION.md: what was actually tested.

  • uv.lock: frozen Python dependency resolution.

Important limits

  1. A running VoiceStudio backend and installed models are required. The gateway does not contain model weights and cannot synthesize audio by itself.

  2. A small Railway/Render instance can host the gateway. Running the complete VoiceStudio inference workload there is a separate resource decision; do not assume a 1 GB/free instance will support large speech models or video dubbing.

  3. Native desktop controls (microphone widget, OS hotkeys, native file pickers, reveal-in-folder, desktop runtime management) and continuous WebSocket sessions remain in the original app. Finite HTTP/SSE workflows are supported.

  4. The initial mcp.json is local. A ChatGPT web/mobile connection needs the deployed HTTPS endpoint or the remote plugin exported after verification. Installing a ZIP does not start hosting. Availability depends on the host's current custom-MCP/plugin support; this package does not promise mobile-only installation.

  5. Secret URLs are a private single-user access method. They are not OAuth. Treat the full MCP URL as a password. Use an OAuth gateway for shared/public distribution. Optional bearer authentication is supported for clients that send headers; this package does not implement an OAuth authorization server.

  6. Hosting, GPU resources, optional paid providers, telephony and some models can incur costs. Open source does not make these services free.

  7. Upstream is AGPL-3.0; model licenses differ. This gateway source is provided under the included AGPL-3.0 license. Only clone a speaker's voice with permission.

No changes to Desk2Quant, its repository, payments or existing services are needed to use this separate package.

Development

uv sync --frozen
uv run pytest

For local stdio, set VOICESTUDIO_URL to your backend, then run uv run voicestudio-mcp --transport stdio. Logs go to stderr.

Sources

The package pins MCP SDK 1.28.1 to match the inspected upstream integration; it does not depend on the current SDK main branch's v2 interface.

Available Tools

21 tools
check_healthA

Check if the VoiceStudio backend is running and what GPU device is active.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral burden, but this is a zero-parameter read-only probe where the risk surface is inherently small. It does disclose the two things the response covers (running state, active GPU), yet never states it is non-mutating or whether it is safe to poll repeatedly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence naming both outputs with zero filler. Perfectly sized for a trivial probe.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description need not explain return values, and it correctly names the two pieces of information returned. The only gap is the missing relationship to the overlapping voicestudio_status sibling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema declares zero parameters, so there is nothing for the description to clarify. Baseline 4 applies; no parameter confusion is possible.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: checks backend liveness and reports the active GPU device. That is more informative than a bare 'health check'. However, it does not distinguish itself from the sibling voicestudio_status, which plausibly overlaps with this tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance at all: nothing says whether to call this before a job, on failure, or how it relates to voicestudio_status. The only usage signal is implied by the word 'check'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clone_voiceA

Clone a new voice profile from a reference audio sample.

    The new voice is immediately available for use with generate_speech
    (pass the returned profile_id as the profile_id argument). Pass
    exactly one of ref_audio_base64 or ref_audio_path.

    Args:
        name: A human-friendly name for the cloned voice.
        ref_audio_base64: Base64-encoded audio (WAV, MP3, FLAC, etc.) of
            the reference voice — 5-30 seconds of clean single-speaker
            speech.
        ref_text: Optional transcript of the reference audio (improves
            quality for some engines).
        instruct: Optional style instruction (e.g. 'whisper', 'excited').
        language: Language of the reference audio (ISO code or 'Auto').
        ref_audio_path: Path to the reference audio under
            OMNIVOICE_MCP_BASE_PATH (relative to it, or absolute inside
            it); refused when no base path is configured. Prefer this
            lane for LLM agents - the clip never enters the context.

    Returns:
        JSON with the new profile's id, name, and kind.
     For long operations use voicestudio_start_job.
ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
instructNo
languageNoAuto
ref_textNo
ref_audio_pathNo
ref_audio_base64No

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose several non-obvious traits: immediate availability of the new voice, the base-path refusal condition, and the mutually exclusive audio-input rule. It omits auth/permission requirements, duplicate-name behavior, and any cost or rate-limit context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first sentence front-loads the purpose and is followed by a well-organized Args/Returns block, so scanning is easy. It is slightly verbose for a six-parameter tool, but each line adds meaning rather than restating the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a creation tool with no annotations and a documented output schema, the description covers the creation semantics, the input constraint, the fallback routing, and the return shape. Remaining gaps are error/failure behavior beyond the base-path refusal and any permission requirements.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry parameter meaning, and it documents all six parameters with real semantic value: duration and cleanliness guidance (5-30 seconds, single speaker) for the audio, ISO code or 'Auto' for language, and example style instructions for instruct. This fully compensates for the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb and resource: cloning a voice profile from a reference audio sample. It clearly separates this from generate_speech by framing the output as an input to it. It never names the closest sibling (design_voice), so an agent must infer the difference between cloning from a sample versus designing a voice.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the core constraint (pass exactly one of ref_audio_base64 or ref_audio_path), recommends a preferred lane for LLM agents (ref_audio_path, so the clip never enters context), and routes long operations to voicestudio_start_job. It also tells the agent how to consume the result via generate_speech.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

describe_voiceA

Preview how a voice description maps onto voice-design attributes.

    Nothing is saved. Use this before design_voice to see what the
    description will produce. The design space is small and fixed; only
    these tokens (and close synonyms) are understood:
      Gender: male, female
      Age: child, teenager, young adult, middle-aged, elderly
      Pitch: very low / low / moderate / high / very high pitch
      Style: whisper
      EnglishAccent: american, british, australian, canadian, indian,
        japanese, korean, chinese, russian, portuguese accent
      ChineseDialect (Chinese speech; overrides an English accent):
        sichuan, dongbei / northeastern chinese, henan, shaanxi, gansu,
        guilin, guizhou, jinan, ningxia, qingdao, shijiazhuang, yunnan
        (e.g. "sichuan dialect")
    Timbre words ("gravelly", "raspy") and other accents are ignored and
    reported in `unmatched`. For a voice outside this space, use
    clone_voice with reference audio instead.

    Args:
        description: Free-text description, e.g. "an elderly man with a
            deep voice and a british accent".

    Returns:
        JSON with attrs (category → token or "Auto"), instruct, matched
        and unmatched.
    
ParametersJSON Schema
NameRequiredDescriptionDefault
descriptionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden, and it does: 'Nothing is saved' declares the non-mutating nature, the fixed token vocabulary declares the input constraints, and the ignored-timbre behavior is disclosed along with where it surfaces ('reported in `unmatched`'). That is more side-effect and constraint disclosure than most annotated tools provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose and the safety fact before the long but necessary token enumeration. Slight waste in the boilerplate 'Args:'/'Returns:' blocks, the latter of which restates an output schema that already exists, but the token list is essential for correct invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter preview tool with no annotations, the description supplies the purpose, the safety profile, the accepted vocabulary, the fallback tool, and the output shape. Nothing an agent needs to invoke it correctly or interpret the result is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single required parameter, so the description must compensate, and it does with a concrete format example ('an elderly man with a deep voice and a british accent') plus the full vocabulary of understood tokens and the fact that out-of-vocabulary words are ignored rather than rejected.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (preview) plus the transformation it performs — mapping a free-text voice description onto design attributes — and implicitly distinguishes itself from design_voice (the committing sibling) and clone_voice (the out-of-space fallback). An agent can tell exactly what this tool produces without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit sequencing: 'Use this before design_voice to see what the description will produce.' Explicit exclusion with a named alternative: 'For a voice outside this space, use clone_voice with reference audio instead.' Both the when-to-use and when-not-to-use paths are covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

design_voiceA

Design and save a new voice profile from a text description.

    The description is mapped onto the same attributes describe_voice
    previews (see its docstring for the vocabulary). The backend tries to
    render a fixed-seed identity sample at save time; if the voice engine
    isn't ready, the profile is still saved and the same sample is
    rendered on first use, so the voice stays stable across
    generate_speech calls either way. Pass the returned profile_id to
    generate_speech. Refuses a description that matches no attribute.

    Args:
        name: A human-friendly name for the new voice.
        description: Free-text description of the voice.
        language: The voice's saved language (ISO code or 'Auto'); used
            for its sample and by generate_speech calls that omit one.

    Returns:
        JSON with the new profile's id, name, kind, the attrs used, and
        any unmatched description fragments.
     For long operations use voicestudio_start_job.
ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
languageNoAuto
descriptionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: it discloses that a fixed-seed sample is rendered at save time or on first use, that the profile is saved even if the engine isn't ready, that the voice stays stable across generate_speech calls, and that the tool refuses descriptions matching no attribute. These are non-obvious behaviors an agent could not infer from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then behavior, then a structured Args/Returns block, with no filler sentences. The trailing 'For long operations use voicestudio_start_job' is slightly abrupt but useful; the Args/Returns restatement is mild redundancy given an output schema exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter creation tool with no annotations, the definition covers behavior, parameters, failure mode, and downstream usage. An output schema exists, so the Returns section is somewhat redundant, and the cross-tool routing to start_job is a single dangling sentence rather than integrated guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it documents all three parameters: name (human-friendly label), description (free text mapped onto attribute vocabulary), and language (ISO code or 'Auto', used for the sample and by generate_speech calls that omit one). It is nearly complete, though it doesn't enumerate example ISO codes or validate the 'Auto' fallback behavior deeply.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence gives a specific verb and resource: design and save a new voice profile from a text description. It also ties the input vocabulary to a named sibling (describe_voice) and tells the agent to pass profile_id to generate_speech, so its role in the workflow is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear operational context: the description must match a known attribute vocabulary, the result's profile_id feeds generate_speech, and long operations should use voicestudio_start_job. It does not explicitly say when to prefer this over clone_voice, which is the most plausible alternative, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_speechA

Generate speech audio from text.

    Args:
        text: The text to synthesize into speech.
        language: Target language (ISO code or 'Auto'). 646 languages
            supported. Omit to use the voice profile's saved language;
            an explicit 'Auto' overrides it.
        profile_id: ID of a saved voice profile to clone. Omit to use this
            agent's bound voice (Settings → MCP), else the default voice.
        instruct: Style instruction (e.g. 'whisper', 'excited', 'narrator').
        speed: Speech speed multiplier (0.5–2.0, default 1.0).
        steps: Diffusion steps (8=fast/draft, 16=balanced, 32=quality).
        format: File and URL format: wav (default), ogg or opus. Both
            ogg and opus carry Opus in Ogg; requires files/both mode and ffmpeg.

    Returns:
        JSON with audio_id, generation_time_s, audio_duration_s and the
        audio shaped by OMNIVOICE_MCP_OUTPUT_MODE: base64 WAV data
        ('resources', the default), a URL plus an optional file ('files'),
        or both ('both'). Prefer 'files' for LLM agents.
     For long operations use voicestudio_start_job.
ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
speedNo
stepsNo
formatNowav
instructNo
languageNo
profile_idNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the output-mode behavior (resources/files/both via OMNIVOICE_MCP_OUTPUT_MODE), the dependency that ogg/opus 'requires files/both mode and ffmpeg', and the fallback chain for voice selection. It does not mention auth requirements, rate limits, or cost, which is a minor gap for a generation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action, then a tidy Args list where each line earns its place, followed by Returns and the sibling pointer. Slightly verbose and the dangling 'For long operations' sentence sits awkwardly after the Returns block, but nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a seven-parameter generation tool with output modes, the description covers every decision point: parameter meanings, defaults, format constraints with their prerequisite (ffmpeg, files/both mode), output-mode selection, and long-operation escalation. An output schema exists, so the Returns detail is bonus rather than a necessity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the prose must compensate, and it does: all seven parameters get semantics beyond their names — language as ISO code or 'Auto' with 646 languages, steps as 8=draft/16=balanced/32=quality, speed range 0.5–2.0, and valid format values. This is far richer than the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Generate speech audio from text'), making it unmistakable against the reverse-direction sibling transcribe and the voice-creation siblings (clone_voice, design_voice). No ambiguity about what the tool produces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit alternative routing: 'For long operations use voicestudio_start_job' gives the when-not condition, and 'Prefer files for LLM agents' guides mode selection. Default-resolution rules for language and profile_id (omit to inherit binding, explicit 'Auto' overrides) tell the agent exactly when to pass or omit each optional argument.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_languagesA

List a sample of supported TTS languages.

VoiceStudio supports 646 languages. This returns the most popular ones plus a note about the full count.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does disclose a genuinely important behavioral trait: the result is a curated sample, not the complete set of 646 languages, and it includes a count note. It does not need to describe return shape since an output schema exists. Read-only listing needs no auth/rate disclosure, so this is near-complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Short and front-loaded, with the sampling caveat directly after the purpose. The second and third sentences overlap somewhat ("sample" vs "most popular ones plus a note about the full count"), which is mild redundancy rather than a structural flaw.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless read tool with an output schema, the description supplies everything an agent needs: what it returns, that the list is partial, and that the total count is surfaced. No open questions remain that would affect invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4; there are no parameter semantics to add meaning to. Nothing in the description misrepresents the no-argument interface.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("List ... supported TTS languages") and immediately qualifies the scope as a sample rather than the full 646-entry catalog. The resource is unambiguous against siblings like list_voices and list_personalities, so an agent can select it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the agent can infer this is the discovery call for language codes, but there is no explicit when-to-use guidance, no mention of prerequisites, and no routing to alternatives. Adequate but with a clear gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_personalitiesA

List available voice personality presets.

    Returns presets like Narrator, Casual, News Anchor, etc. with their
    instruct text. Use the instruct text with generate_speech.
    
ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It discloses the return contents (named presets with instruct text) and how to consume them, but says nothing about auth needs, caching, or whether the preset list is static. For a zero-param read tool this is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the purpose and followed by the return shape and the follow-up action. No padding, though the sentence fragment structure is slightly loose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-argument list tool with an output schema already describing return values, the description covers purpose, example values, and downstream usage. The only omission is the sibling distinction against list_voices, which is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes no parameters, so there is nothing to disambiguate and the baseline of 4 applies. The description adds no parameter detail because none is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List available voice personality presets'), and the resource is distinct from siblings like list_voices and list_languages. It does not explicitly contrast itself against those siblings, but the noun 'personality presets' is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear downstream context: 'Use the instruct text with generate_speech,' which tells the agent why to call this and what to do with the result. It stops short of saying when to prefer this over list_voices, so no explicit alternative routing is offered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_voicesA

List all saved voice profiles.

Returns a JSON array of voice profiles with id, name, type (clone/design), and personality.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden; it discloses that this is a read-style enumeration and names the returned fields. However, it says nothing about permissions/auth, whether results are paginated or bounded, or whether the list is filtered by user or workspace — relevant gaps for an unannotated tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, with the core action front-loaded and the return shape following immediately. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not explain return values, and with zero parameters there is no input complexity to cover. It is nearly complete, missing only the access/prerequisite context that annotations would otherwise have supplied.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate and the baseline of 4 applies. Schema coverage is also 100%.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List all saved voice profiles') with scope ('all saved'), so the agent knows this is an enumeration tool. It does not distinguish itself from nearby siblings like describe_voice or list_personalities, but the verb+resource pairing is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of alternatives (e.g. describe_voice for a single profile), and no stated prerequisites. The only implied usage cue is that it takes no parameters, which the schema already conveys.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

transcribeA

Transcribe spoken audio to text.

    Pass exactly one of audio_base64 or audio_path.

    Args:
        audio_base64: Base64-encoded audio bytes (wav/mp3/webm/m4a).
        audio_path: Path to an audio file under OMNIVOICE_MCP_BASE_PATH
            (relative to it, or absolute inside it). The base path is the
            security boundary: with none configured, paths are refused.
            Prefer this lane for LLM agents - the audio never enters the
            agent's context.
        language: Optional language hint; omit for auto-detect.

    Returns:
        JSON with the recognized text, language, and duration.
     For long operations use voicestudio_start_job.
ParametersJSON Schema
NameRequiredDescriptionDefault
languageNo
audio_pathNo
audio_base64No

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the OMNIVOICE_MCP_BASE_PATH security boundary, that paths are refused when no base path is configured, and that the path lane keeps audio out of the agent's context. It omits auth/permission requirements and size or rate limits, so it falls short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose and the one-of constraint, then organized into Args/Returns sections with no filler sentences. Slight redundancy in restating the return payload despite an output schema existing, and the docstring formatting is heavier than strictly necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter tool with an output schema already covering return values, the description supplies everything else an agent needs: purpose, parameter constraints, the security boundary for file paths, and the escalation path to a job-based tool. No meaningful gap remains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: each of the three params is explained with accepted formats (wav/mp3/webm/m4a), path resolution semantics relative to the base path, and the mutual-exclusivity constraint that the schema does not encode. Language is documented as an optional auto-detect hint.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Transcribe spoken audio to text') that an agent can distinguish from generate_speech and the voicestudio_* job tools. It also names voicestudio_start_job as the alternative for long operations, so the tool's niche is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit selection rule ('Pass exactly one of audio_base64 or audio_path'), a preference recommendation for LLM agents (path lane, because audio never enters context), and a routing rule to voicestudio_start_job for long operations. This is when-to-use plus alternatives, not inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

voicestudio_call_apiA
Destructive

Execute one discovered HTTP operation against the configured backend. JSON, forms, repeated multipart files and finite SSE responses are supported. Binary outputs become timed file links. Mutations may delete data, change settings, download models, contact providers or place calls; use only within the user's request. Use start_job for long renders.

ParametersJSON Schema
NameRequiredDescriptionDefault
formNo
filesNo
queryNo
headersNo
json_bodyNo
path_paramsNo
operation_idYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations already declaring destructive=false/idempotent=false (here destructiveness is disclosed), the description still adds real context: supported payload shapes (JSON, forms, multipart files, finite SSE) and the non-obvious behavior that binary outputs become timed file links. The mutation enumeration adds specificity beyond destructiveHint=true, though it does not cover auth or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Sentences are dense but each carries a distinct fact (capabilities, output handling, mutation risk, alternative tool), and the core purpose leads. Slightly list-heavy in the mutation clause but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an open-world, destructive, parameterless-described API caller with no output schema, the description covers capability and safety reasonably well. However, the absence of any operation_id provenance or payload-field mapping leaves an agent without enough to construct a correct call from the description alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 7 parameters, so the description must compensate and largely does not: it never mentions operation_id, json_body, path_params, query, or headers. The loose references to 'JSON, forms, repeated multipart files' hint at json_body/form/files but do not map to any named parameter, leaving an agent to guess field usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Execute') and resource ('one discovered HTTP operation against the configured backend'), so the agent knows it is a generic API caller rather than a domain tool. It does not explicitly distinguish itself from the sibling voicestudio_call_native, which is the nearest alternative, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear constraint ('use only within the user's request') and routes long-running work to a named alternative ('Use start_job for long renders'). It lacks guidance on when to prefer this over call_native or how operation_id is discovered, so it is strong but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

voicestudio_call_nativeB
Destructive

Call a discovered native MCP tool with its original arguments. For uploaded reference/audio files use file_arguments={ref_audio_base64:file_id} for clone_voice or {audio_base64:file_id} for transcribe. Large audio uses multipart HTTP API instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
argumentsNo
tool_nameYes
file_argumentsNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safety profile (destructive, open-world, non-idempotent), so the description is not the sole carrier of that information. It adds the file-passing behavior for reference/audio files, which is genuinely useful, but does not disclose that behavior is entirely delegated to the target tool, nor what happens on an unknown tool_name or a failed downstream call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action before the file-handling detail. The compact notation ({ref_audio_base64:file_id}) is dense but parses quickly; nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no parameter descriptions, the description should do more: it omits how native tools are discovered, what the response looks like, and how errors surface for a destructive dispatcher. The file_arguments coverage is a real contribution, but the discovery/response side of this meta-tool is left unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden, and it does clarify file_arguments with concrete key/value examples tied to specific target tools. It says nothing about the `arguments` passthrough object or how tool_name is expected to be resolved (the 'discovered' mechanism), leaving a meaningful gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it calls a discovered native MCP tool, passing through the original arguments. The word 'native' implicitly separates it from the sibling voicestudio_call_api, but the description never states that contrast explicitly, so an agent must infer the boundary from the names alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete when-to-use guidance for the file_arguments mechanism (clone_voice vs transcribe key names) and one explicit when-not rule ('Large audio uses multipart HTTP API instead'). However, it never says when to pick this dispatcher over voicestudio_call_api, voicestudio_start_job, or calling generate_speech/clone_voice directly, which is the main selection decision for an agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

voicestudio_cancel_jobA
Destructive

Cancel the gateway's wait. Backend synthesis may continue. To abort backend work use its discovered cancel/abort HTTP operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true, idempotentHint=false, readOnlyHint=false, so the mutation profile is covered. The description adds genuinely useful context beyond that: canceling only stops the gateway wait while backend synthesis may continue, which warns the agent the operation is partial and not a true abort.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place: the action, the crucial caveat, and the routing to the alternative. The key scope limitation is front-loaded immediately after the verb.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so return behavior needn't be explained, and the destructive/partial nature is well covered. The gap is the required job_id, which an agent cannot confidently supply from either the schema or the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The sole parameter job_id has 0% schema description coverage, so the description carries the full burden, yet it says nothing about what job_id is, where it comes from, or its format. The one sentence of context is entirely behavioral, leaving the required parameter undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Cancel) and a precise, non-obvious scope: "the gateway's wait" rather than the job itself. This distinction is essential given the name 'cancel_job' would otherwise mislead, and it implicitly separates this from backend abort operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: use this to stop waiting, and use a discovered cancel/abort HTTP operation to abort backend work. That is a clear when-to-use-this-vs-alternative statement, though the alternative is described generically rather than by naming the sibling (e.g. voicestudio_call_api).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

voicestudio_create_uploadA
Destructive

Create a file slot and a one-use signed PUT URL for direct audio, video, manuscript or subtitle upload. Send raw file bytes to that URL; then use its file_id in API multipart fields. Requires PUBLIC_BASE_URL for remote clients.

ParametersJSON Schema
NameRequiredDescriptionDefault
filenameYes
mime_typeNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With destructiveHint=true already declared, the description adds real context: the URL is one-use and signed, a persistent file slot is created, and remote clients require PUBLIC_BASE_URL. It never explains why the operation is flagged destructive (slot creation/consumption), so a 5 is not warranted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the operation and its return artifact (file_id), then the follow-up step, then the environment prerequisite. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter creation tool with no output schema, the description covers the workflow, the signing/reuse constraint, and the env requirement, and it names file_id as the handoff value. The remaining gap is that neither input parameter's semantics are addressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and neither 'filename' nor 'mime_type' is explained in the description. The only hint is the enumeration of media kinds, which loosely maps to mime_type; the required filename parameter and any naming/format constraints are undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Create a file slot and a one-use signed PUT URL', plus the accepted media kinds (audio, video, manuscript, subtitle). It does not explicitly name voicestudio_upload_base64 as the contrasting sibling, but 'direct ... Send raw file bytes' implicitly distinguishes it from the base64 path.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear two-step usage recipe: PUT raw bytes to the returned URL, then reference file_id in API multipart fields, plus a deployment prerequisite (PUBLIC_BASE_URL for remote clients). It lacks an explicit when-to-use-this-vs-upload_base64 statement or any exclusions, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

voicestudio_file_infoB
Read-onlyIdempotent

Inspect a staged/generated file and renew its timed download link while retained. Links provide access to anyone holding them; do not publish private recordings.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_idYes

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds real context beyond the annotations: the link is timed and gets renewed on this call, and the link is a bearer-style capability usable by anyone holding it, with a warning not to publish private recordings. The only gap is the mild tension with readOnlyHint=true, since renewing an expiry is a side effect (though idempotent and non-destructive, matching the other hints).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core action, with the security caveat second. Nothing is redundant, though the security sentence is phrased as an admonition rather than operational detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should say what 'inspect' returns (metadata fields, expiry window, link format) and it does not. For a simple one-parameter lookup this is adequate but leaves the agent guessing at the response shape and link lifetime.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single parameter file_id is undocumented. The description partially compensates by implying the id refers to a staged or generated file, but adds no format, prefix, or lookup semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: inspecting a 'staged/generated file' and renewing its 'timed download link'. That distinguishes it from the job/operation/upload siblings, though it never names an alternative directly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives no explicit when-to-use or when-not-to-use guidance. The only implied trigger is having a file already staged or generated, and no alternative (e.g. read_resource) is mentioned for retrieving file content instead of metadata.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

voicestudio_get_operationA
Read-onlyIdempotent

Get a live operation's exact request parameters, requestBody, responses and referenced component schemas. Inspect this before calling it.

ParametersJSON Schema
NameRequiredDescriptionDefault
operation_idYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, idempotent, open-world, and non-destructive behavior. The description adds useful behavioral context by specifying exactly what data is retrieved: request parameters, requestBody, responses, and referenced component schemas.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the retrieval purpose and followed by a clear usage directive. No filler or redundant restatement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Annotations cover the safety profile and there is no output schema, so the description need not explain return values. However, for a 1-param tool with 0% schema description coverage, it should say more about operation_id, such as where it comes from or what format it uses.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the single operation_id parameter is undocumented in the schema. The description implies it identifies an operation but does not explain its format, source, or how to obtain it, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb (Get) and resource (a live operation's exact request parameters, requestBody, responses, and referenced component schemas). This clearly distinguishes the tool from execution-oriented siblings like voicestudio_call_api and voicestudio_call_native.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Inspect this before calling it' gives a clear usage context: retrieve operation metadata prior to invoking the operation. It does not name alternatives explicitly or state when not to use the tool, but the intended workflow is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

voicestudio_job_statusB
Read-onlyIdempotent

Poll a gateway background job. Backend batch/dub job IDs are separate; inspect their HTTP operations for progress and cancellation.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, non-destructive and openWorld, so the safety profile is covered. The description adds that this is a polling operation and that job-ID namespaces differ, but says nothing about terminal states, return content, or whether polling must be repeated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action. The second sentence is slightly cryptic but every clause carries information; little waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and one undocumented parameter, the description should ideally describe the status payload or polling semantics. It covers the namespace distinction but leaves the agent without a clear picture of what a poll returns or when to stop.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single job_id parameter is undocumented. The description partially compensates by clarifying that gateway job IDs are distinct from backend batch/dub IDs, which is real semantic value, but it never says where the ID comes from or its format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Poll a gateway background job') and scopes it against backend batch/dub jobs, which are declared separate. This differentiates it from related siblings, though it never names them directly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implicitly tells you to use this for gateway jobs and warns that backend batch/dub job IDs are handled elsewhere via their HTTP operations. However, the routing advice ('inspect their HTTP operations') is vague and does not name the alternative sibling (e.g. voicestudio_get_operation), leaving the agent to infer the correct call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

voicestudio_read_resourceA
Read-onlyIdempotent

Read a native MCP resource such as voice:// or history://recent. Only the backend's advertised resource namespaces are accepted.

ParametersJSON Schema
NameRequiredDescriptionDefault
uriYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is fully covered by structured data. The description adds one behavioral constraint (rejection of unadvertised namespaces), but says nothing about the returned content shape or error format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and followed by the key constraint. No filler, no restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description would ideally hint at what a resource read returns, but it omits the return shape entirely. The namespace-acceptance rule covers the main failure mode, and annotations cover safety, leaving the definition adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single uri parameter has 0% schema description coverage, so the description carries the burden and does so with concrete format examples (voice://<profile_id>, history://recent). It stops short of enumerating the full set of accepted namespaces, so an agent still has to guess at edge-case URIs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (Read) and resource (native MCP resource), plus concrete URI examples (voice://<profile_id>, history://recent) that make the target unambiguous. It distinguishes the tool implicitly from API/job-oriented siblings like voicestudio_call_api or voicestudio_start_job, but never names an alternative to route against explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the URI-based examples and the constraint that only the backend's advertised namespaces are accepted, which steers the agent toward valid inputs. However, it never states when to prefer this over voicestudio_call_api or voicestudio_call_native, nor any prerequisite for discovering valid namespaces.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

voicestudio_search_apiB
Read-onlyIdempotent

Search every live VoiceStudio OpenAPI HTTP operation, including dubbing, audiobooks, conversion, models, batch jobs, profiles, pronunciation, projects, watermarking, settings and workers. Offline source catalog is informational only.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNo
limitNo
queryNo
methodNo
offsetNo
refreshNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, covering the safety profile. The description adds the live-vs-offline distinction, but says nothing about pagination (offset/limit behavior), result format, or freshness (the 'refresh' flag).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and scope, with no filler. The domain enumeration is long but serves the purpose of defining the search surface.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter search tool with no output schema and no schema descriptions, the description establishes what is searched but omits how the filters behave and what the response contains, leaving meaningful gaps for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Six parameters with 0% schema description coverage, and the description supplies no parameter meaning at all — it never explains query, tags, method, limit, offset, or refresh. The enumerated domains loosely hint at what 'tags' might contain but stop short of mapping to any parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Search) and resource (every live VoiceStudio OpenAPI HTTP operation), then enumerates the covered domains (dubbing, audiobooks, conversion, models, etc.), so an agent knows exactly what surface is being searched. It distinguishes the live operation surface from the 'offline source catalog' but does not explicitly separate itself from siblings like voicestudio_get_operation or voicestudio_call_api.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The note that the offline source catalog 'is informational only' implies this tool is the right entry point for live search, but there is no explicit when-to-use/when-not or named alternative. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

voicestudio_start_jobA
Destructive

Start a background operation and immediately return a gateway job ID. For kind=api use the operation ID as name and call_api arguments without operation_id. For kind=native use a native tool name and arguments={arguments:{...},file_arguments:{...}}. Gateway status persists; work is not automatically retried after restart.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYes
nameYes
argumentsNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds real behavioral facts beyond the annotations: the call returns immediately with a job ID, gateway status persists, and work is NOT automatically retried after restart. That last point is important operational context not conveyed by destructiveHint/openWorldHint. It still doesn't say how to subsequently poll or cancel (the job_status/cancel_job siblings exist).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose in the first sentence, then dispenses the two mode-specific usage rules compactly. Dense but every clause carries information; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a fire-and-forget job launcher with no output schema, the description covers the return (job ID) and the persistence/retry semantics. It would be stronger if it pointed at voicestudio_job_status and voicestudio_cancel_job for follow-up, which is the natural next step for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden and largely does: it explains that name holds the operation ID for kind=api, and that arguments for kind=native nests arguments and file_arguments. The internal shape of arguments for kind=api is not spelled out, leaving some ambiguity in a nested free-form object.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Start a background operation') plus the immediate return value ('gateway job ID'), which distinguishes it from the synchronous call_api/call_native siblings. It does not explicitly name the sync alternatives, so an agent must infer that this is the async variant.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete routing guidance for both modes: for kind=api use the operation ID as name and call_api arguments without operation_id; for kind=native use a native tool name with the arguments/file_arguments shape. It lacks an explicit 'use this instead of call_api when you want non-blocking execution' statement, so the when-not is left implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

voicestudio_statusB
Read-onlyIdempotent

Check backend health, native MCP discovery, HTTP API operation count and file-transfer configuration.

ParametersJSON Schema
NameRequiredDescriptionDefault
refreshNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint, so the safe-read profile is covered by structured data. The description adds value by disclosing what categories of information are probed, but says nothing about auth requirements, rate limits, or the effect of the refresh parameter, so it only modestly exceeds the annotation baseline.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence that front-loads the verb and lists the checked subsystems with no filler. It is appropriately sized, though the list format is dense rather than explanatory.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the burden of explaining what comes back; it does enumerate the probed areas but gives no sense of response shape or how to interpret results. The unaddressed refresh parameter is also a completeness gap for a status/diagnostic tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter, refresh, has 0% schema description coverage and is never mentioned in the description despite its name implying a forced re-check versus a cached read. With one undocumented parameter, the description should compensate, and it does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ("Check") and enumerates the concrete things inspected: backend health, native MCP discovery, HTTP API operation count, and file-transfer configuration. It clearly states what the tool does, but it does not differentiate itself from the sibling check_health, whose purpose sounds overlapping.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus check_health or any other diagnostic sibling, and no prerequisites or exclusions are stated. The agent must infer that this is a broader diagnostic aggregation, but nothing in the text confirms it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

voicestudio_upload_base64A
Destructive

Stage a small file (default max 8 MiB) from base64. Prefer direct signed PUT upload for large files; never ask the user to paste base64 into chat.

ParametersJSON Schema
NameRequiredDescriptionDefault
filenameYes
mime_typeNo
content_base64Yes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the safety profile (readOnly=false, destructive=true, openWorld=true, non-idempotent), so the bar is lower. The description adds useful behavioral context beyond them: an 8 MiB default size limit, the staging nature of the operation, and a workflow anti-pattern to avoid. It does not explain what the destructiveHint implies (e.g. overwriting an existing staged file), leaving one gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences with no filler. The primary action and size cap are front-loaded, followed by the alternative and the anti-pattern guidance, in decreasing order of importance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter mutation tool with no output schema, the description covers the workflow and the size constraint, and annotations cover safety. It stops short of the parameter-level detail (filename/mime_type behavior, overwrite semantics) and does not indicate what a successful stage returns (a handle or ID) that a caller probably needs to proceed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and none of the three parameters have schema descriptions, so the description must carry the full load. It only clarifies content_base64 (base64 source, size capped) and leaves filename (does it overwrite an existing entry? is the extension meaningful?) and mime_type entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ("Stage") and resource ("a small file ... from base64"), so an agent immediately knows what operation this performs. It also carves out the large-file case against the signed-PUT path, though it does not name the sibling tool (e.g. voicestudio_create_upload) explicitly. Clear purpose with slightly implicit sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the condition for use (small files, default 8 MiB cap), the alternative for large files (direct signed PUT upload), and a hard behavioral rule (never ask the user to paste base64 into chat). The when/when-not/alternative triad is all present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 21 tool updatesv0.1.0
    • First observedcheck_health
    • First observedclone_voice
    • First observeddescribe_voice
    • First observeddesign_voice
    • First observedgenerate_speech
    • First observedlist_languages
    • First observedlist_personalities
    • First observedlist_voices
    • First observedtranscribe
    • First observedvoicestudio_call_api
    • First observedvoicestudio_call_native
    • First observedvoicestudio_cancel_job
    • First observedvoicestudio_create_upload
    • First observedvoicestudio_file_info
    • First observedvoicestudio_get_operation
    • First observedvoicestudio_job_status
    • First observedvoicestudio_read_resource
    • First observedvoicestudio_search_api
    • First observedvoicestudio_start_job
    • First observedvoicestudio_status
    • First observedvoicestudio_upload_base64

TDQS

B3.4/5.0

Scored across 21 tools

Disambiguation3/5

Most native tools (generate_speech, transcribe, clone_voice, design_voice, describe_voice) have distinct purposes, but check_health and voicestudio_status clearly overlap (both report backend health/GPU), and the generic API layer (search_api/get_operation/call_api/call_native) creates unclear boundaries against the native tools—agents must decide which interface to use. The dual upload paths (create_upload vs upload_base64) are clarified by descriptions but still close.

Naming Consistency3/5

Two conventions coexist: a large voicestudio_* prefixed family (status, search_api, call_api, start_job, upload_base64, etc.) and an unprefixed native family (generate_speech, clone_voice, transcribe, list_voices, check_health). Both are snake_case and readable, but the mix—and especially check_health being unprefixed while the overlapping voicestudio_status is prefixed—makes the set feel inconsistent.

Tool Count3/5

21 tools sits in the heavy/borderline range for a voice platform. The native voice tools (7) and job/file helpers are justified, but the sizeable generic gateway meta-layer (search/get/call API, call_native, start/status/cancel job, upload, file_info, read_resource) inflates the count and adds surface that largely mirrors what the native tools already do.

Completeness4/5

The surface covers the core lifecycle: synthesis, transcription, cloning, design+preview, voice/personality/language listing, file upload, and background job control, plus a dynamic API passthrough for anything unlisted. Minor gaps exist (no explicit delete/update/rename of profiles, history only reachable via read_resource), but these are workable.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers