VoiceLabs
Server Details
AI voice generation: text-to-speech and voice cloning from any MCP client.
- Status
- Healthy
- Uptime
- 15.0% over 48 days
- OAuth
- Works in Glama
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 9 tools
Most tools target distinct actions (clone vs ensure vs speak vs transcribe vs list), and the descriptions explicitly contrast clone_voice_profile with ensure_voice_profile. The only mild friction is the two profile-creation tools and the meta-pair search_tools/execute_typescript, but descriptions make boundaries clear.
All names are snake_case with a verb_noun pattern (clone_voice_profile, get_generation, list_captures, search_tools). Two bare verbs (speak, transcribe) deviate slightly from the noun-suffixed pattern but remain readable and conventional for TTS/STT.
Nine tools is well-scoped for a voice platform, covering voice creation, listing, TTS, STT, generation polling, capture listing, and a meta-orchestration pair. Each tool earns its place with no redundancy.
The surface covers creation and read paths plus TTS/STT, but has no update or delete for voice profiles or captures, and no get-by-id for profiles/captures beyond generation polling. Agents can work around this but lifecycle coverage is partial.
Available Tools
9 toolsclone_voice_profileClone a voice from a recordingAInspect
Create a cloned voice on the user's VoiceLabs account from a recording the caller provides: a name, the language, the exact text spoken in the recording, and the audio either inline as base64 (audioBase64) or as an https:// URL (audioUrl). The recording must be 2–30 seconds of clear speech. Returns the new profile's id, name, voice type and engine; pass the id to speak. Not idempotent: a second call with the same name is refused, so check list_voice_profiles first. Only clone a voice whose owner has agreed to it.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The name for the new cloned voice. Must not already exist on the account — call list_voice_profiles first if unsure; a duplicate is refused with INVALID_REQUEST. | |
| audioUrl | No | An https:// URL the recording can be downloaded from (public host, no redirects, up to 10 MiB). Give exactly one of audioBase64 or audioUrl. | |
| language | No | ISO language code of the recording and the voice. Defaults to "en". | |
| audioBase64 | No | The recording inline as base64 (WAV, or any of mp3/m4a/ogg/flac/aac/webm/opus; up to 10 MiB decoded). 2–30 seconds of clear speech. Give exactly one of audioBase64 or audioUrl. | |
| referenceText | Yes | The exact words spoken in the recording (up to 1000 characters). The engine aligns the clone against this transcript, so it must match what is said. |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | |
| name | Yes | |
| engine | Yes | |
| language | Yes | |
| sampleId | Yes | |
| voiceType | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=false, idempotentHint=false and openWorldHint=true, and the description reinforces the non-idempotency with a concrete consequence ('a second call with the same name is refused'). It also adds a real constraint the annotations don't convey — the 2–30 second recording requirement — and an ethical precondition (owner consent). Not a full 5 because the duplicate-refusal mechanics are largely repeated from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense paragraph with the core purpose front-loaded, followed by inputs, return shape, the non-idempotency warning, and the consent caveat. Each sentence carries information; the parameter enumeration is slightly redundant against the 100%-covered schema but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described — yet the description still names the key returned fields (id, name, voice type, engine) and how to use the id. With annotations covering the safety profile and the schema covering every parameter, the remaining behavioral gaps (non-idempotency, consent, length constraint) are all addressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all five parameters in detail (formats, size limits, enum languages, transcript matching). The description restates the inputs at a high level — name, language, spoken text, and audio via base64 or URL — but adds no syntax or format meaning beyond what the schema provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a cloned voice ... from a recording') and enumerates the required inputs, so the agent knows exactly what the tool produces. It routes to list_voice_profiles and speak, but never distinguishes itself explicitly from the sibling ensure_voice_profile, which appears to be the idempotent variant — a one-line contrast would have earned a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational context: check list_voice_profiles first, pass the returned id to speak, and only clone with the owner's consent. It names a concrete prerequisite tool, though it stops short of stating when to prefer this over ensure_voice_profile.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ensure_voice_profileEnsure a voice profile existsAIdempotentInspect
Make sure a voice profile with the given name exists on the user's VoiceLabs account, creating it from one of VoiceLabs' built-in (preset) voices if it does not. Idempotent: calling it again with the same name returns the same profile with created: false and creates nothing. Use this when list_voice_profiles comes back empty, or when speak reports NOT_FOUND for a voice, then pass the returned id to speak. It cannot clone a real person's voice — that is clone_voice_profile, which needs the voice:clone scope and a recording.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The profile name to ensure exists. Matched EXACTLY against the account's existing profiles — the same uniqueness key the engine enforces — so a differently-cased name is a different profile. | |
| engine | No | Which built-in voice catalogue to draw from. Defaults to "kokoro", VoiceLabs' fast general-purpose preset engine. | |
| language | No | ISO language code for the profile. Defaults to "en". | |
| description | No | Optional human note stored on the profile. Ignored when the profile already exists. | |
| builtinVoiceId | No | The built-in voice to use. Defaults to "af_heart", the engine's own default Kokoro voice. The ids VoiceLabs publishes (case-sensitive) — kokoro: af_alloy, af_aoede, af_bella, af_heart, af_jessica, af_kore, af_nicole, af_nova, af_river, af_sarah, af_sky, am_adam, am_echo, am_eric, am_fenrir, am_liam, am_michael, am_onyx, am_puck, am_santa, bf_alice, bf_emma, bf_isabella, bf_lily, bm_daniel, bm_fable, bm_george, bm_lewis, ef_dora, em_alex, em_santa, ff_siwis, hf_alpha, hf_beta, hm_omega, hm_psi, if_sara, im_nicola, jf_alpha, jf_gongitsune, jf_nezumi, jf_tebukuro, jm_kumo, pf_dora, pm_alex, pm_santa, zf_xiaobei, zf_xiaoni, zf_xiaoxiao, zf_xiaoyi. qwen_custom_voice: Vivian, Serena, Uncle_Fu, Dylan, Eric, Ryan, Aiden, Ono_Anna, Sohee. Any other id is refused with INVALID_REQUEST. |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | |
| name | Yes | |
| engine | Yes | |
| created | Yes | |
| language | Yes | |
| voiceType | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and destructiveHint=false, so the safety profile is covered, but the description adds real value on top: the exact idempotent contract ('returns the same profile with `created: false` and creates nothing') and the boundary that it cannot clone a real person's voice. It stops short of describing error behavior beyond that, but this is well above the annotation baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences with no filler; the creation semantics come first, then the usage triggers, then the sibling exclusion. It is dense but every clause carries information, so it is efficient rather than padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 5 parameters, an output schema, and full annotation coverage, the description supplies exactly what the structured fields cannot: the idempotency contract, the failure triggers that prompt calling it, and the hand-off to speak. An agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with enum values, defaults, and the exact-match/case-sensitivity rule all documented in the schema itself. The description adds only the informal notion of 'the given name' and 'the returned id', so the baseline 3 for a fully documented schema is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a precise verb and resource ('Make sure a voice profile with the given name exists on the user's VoiceLabs account, creating it from one of VoiceLabs' built-in (preset) voices if it does not') and explicitly distinguishes itself from clone_voice_profile, which a sibling tool. An agent can route to it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the trigger conditions ('when list_voice_profiles comes back empty, or when speak reports NOT_FOUND for a voice'), tells the agent what to do next ('pass the returned id to speak'), and identifies the alternative it is not (clone_voice_profile, which needs the voice:clone scope and a recording). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
execute_typescriptExecute TypeScriptBDestructiveInspect
Run a short TypeScript program that orchestrates VoiceLabs MCP tools (found via search_tools) in an isolated sandbox. See the server instructions for the full contract.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | A TypeScript program body run in an isolated sandbox. Call the external_* functions search_tools declared; the program must `return` its result. |
Output Schema
| Name | Required | Description |
|---|---|---|
| logs | No | |
| error | No | |
| result | No | |
| success | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, openWorldHint=true, non-idempotent, and not read-only, so the risk profile is covered structurally. The description adds the meaningful detail that execution happens in an 'isolated sandbox' and that the program must return its result, but it does not say what constraints 'short' imposes (time, memory, tool-call budget) or what happens on sandbox failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no padding, with the core operation front-loaded. The closing pointer to server instructions is a necessary redirect for a tool whose full contract is large, though it does consume space without adding local information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value description is not required, and the sandbox/return contract is sketched. However, for a destructive, open-world code-execution tool the behavioral contract is deferred to 'server instructions', leaving resource limits and failure semantics unspecified in the definition itself.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter with 100% schema description coverage, and the schema already explains that 'code' is a TypeScript body calling declared external_* functions that must return its result. The description only adds the vague 'short' qualifier, so the baseline of 3 for high coverage is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Run/Execute) and resource (a TypeScript program) and pins the domain: orchestrating VoiceLabs MCP tools in an isolated sandbox. It also names the sibling 'search_tools' as the discovery path, so an agent can place it relative to the other tools without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use for multi-tool orchestration ('orchestrates VoiceLabs MCP tools') but never states when to prefer this over calling a sibling tool directly, nor any exclusions or limits. The explicit instruction is to consult server instructions, which is an out-of-band dependency rather than guidance in the definition itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_generationCheck a generationARead-onlyIdempotentInspect
Look up a VoiceLabs generation by id to poll its status. Returns status (generating | completed | failed) and, once completed, the audio URL; a failed one carries a short human-readable error and a failureKind. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| generationId | Yes | The id returned by the speak tool. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | Yes | |
| status | Yes | |
| profile | Yes | |
| audioUrl | Yes | |
| failureKind | Yes | |
| generationId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/non-destructive, so the safety burden is covered; the description still adds the state machine (generating | completed | failed) and the failure payload (human-readable error plus failureKind). That is genuine behavioral context beyond the annotations, though the return details overlap the existing output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, front-loaded with the action and purpose before the return-value detail. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, read-only polling tool with full annotation coverage and an output schema, nothing needed to call it correctly is missing. The added status/failure semantics are a bonus rather than a gap-filler.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single parameter, and the description's framing ('by id', id from the speak tool) matches rather than extends the schema's own description. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (look up), resource (a VoiceLabs generation), and scope (by id, to poll status), and explicitly ties the id's provenance to the sibling `speak` tool. An agent can distinguish this from `speak` or `transcribe` without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly frames the context ('to poll its status') and names where the required id comes from ('returned by the speak tool'). It stops short of explicit when-not-to-use or naming an alternative polling/listing tool, so it is clear context without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_capturesList capturesARead-onlyIdempotentInspect
List the user's recent VoiceLabs captures (dictations, recordings, uploads) with their transcripts, most-recent first. Paginated via limit (1-200) and offset. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| limit | Yes | |
| total | Yes | |
| offset | Yes | |
| captures | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so the 'Read-only' text is redundant. The description adds useful behavioral context: results are ordered most-recent first and pagination uses limit/offset. No auth, rate-limit, or return-shape details beyond transcripts are given.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose and scope, then pagination and read-only. Every element serves the agent, with no repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read-only list tool with an output schema and full annotations, the description covers purpose, ordering, and pagination. It omits explicit usage routing, but otherwise provides enough to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It identifies limit and offset as pagination controls and repeats the limit range (1-200), but does not explain offset semantics or the default values already present in the schema. This partially compensates but leaves gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List), resource (captures), scope (user's recent), and categories (dictations, recordings, uploads) with transcripts and ordering. This distinguishes it from sibling list_voice_profiles and other tools without needing to reference them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides implied usage by specifying the tool lists captures, but does not state when to use it versus alternatives like get_generation or search_tools, nor does it give exclusions or prerequisites. An agent can infer it is the general capture-list endpoint, but explicit guidance is absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_voice_profilesList voice profilesARead-onlyIdempotentInspect
List the user's VoiceLabs voice profiles — both cloned voices and presets — with language, engine, and usage counts. Read-only. A brand-new account has none yet: an empty list comes back with an emptyState explaining how the user creates their first voice, which is what to relay instead of reporting that the tool failed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| total | Yes | |
| profiles | Yes | |
| emptyState | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already establish readOnlyHint, idempotentHint, and non-destructive semantics, so the description's 'Read-only' line is largely redundant. Its real added value is disclosing the empty-account behavior and the `emptyState` payload, a trait not derivable from the annotations or schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the operation and scope before the empty-state caveat. The standalone 'Read-only' sentence is redundant with readOnlyHint=true and is the only real waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return fields need not be re-explained, yet the description still covers the one non-obvious outcome (emptyState on a fresh account). Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate — baseline 4 applies. The schema already closes the object with additionalProperties: false.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (the user's VoiceLabs voice profiles), and enumerates what is returned — cloned voices, presets, language, engine, usage counts. It is clearly distinguishable from write-oriented siblings like clone_voice_profile, but it never explicitly names or contrasts an alternative tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete usage directive for the empty case: an empty list is normal on a new account, and the `emptyState` should be relayed rather than reported as a failure. It does not, however, say when to prefer this over ensure_voice_profile or clone_voice_profile.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_toolsSearch ToolsARead-onlyIdempotentInspect
Find the VoiceLabs MCP tools this connection can reach, returned as TypeScript declarations to call from execute_typescript. See the server instructions for the full contract.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Case-insensitive substring matched against each tool's name, title and description. Omit to return every tool this connection can reach. |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | Yes | |
| declarations | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds the return format (TypeScript declarations) and the fact that results are limited to this connection's reachable tools, but it pushes the 'full contract' off to server instructions rather than disclosing behavior directly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the discover-tools purpose, then the return format. No filler and no repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and 100% parameter coverage, the description needn't explain return values, yet it usefully states the TypeScript-declaration form and the connection-scoped result set. The only weakness is leaning on external server instructions for the calling contract.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% — the single 'query' parameter is fully documented as a case-insensitive substring match with omit-to-return-all semantics. The description adds nothing about the parameter, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Find the VoiceLabs MCP tools this connection can reach.' It also says what comes back (TypeScript declarations callable from execute_typescript), which distinguishes it from the action-oriented siblings. It stops short of explicitly contrasting itself with those siblings, so not a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied via 'to call from execute_typescript' and the connection-scoping phrase; there is no explicit 'use this when you don't know which tool to call' or when-not guidance. It defers the contract to server instructions rather than stating the selection context itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
speakGenerate speechAInspect
Generate speech from text in one of the user's VoiceLabs voice profiles. Returns a generationId with status "generating"; poll get_generation with that id to retrieve the audio URL once it completes. Pass profileId or profileName to choose the voice.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The text to speak. | |
| language | No | ISO language code (defaults to the profile/engine default). | |
| profileId | No | Voice profile id (takes precedence). | |
| profileName | No | Voice profile name. |
Output Schema
| Name | Required | Description |
|---|---|---|
| status | Yes | |
| profile | Yes | |
| audioUrl | Yes | |
| generationId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare a non-read-only, non-idempotent mutation, but the description adds genuinely useful behavior the annotations cannot convey: the call is asynchronous and returns status 'generating' rather than final audio, requiring a poll. It does not mention auth requirements, quota limits, or whether repeated calls create duplicate generations (relevant given idempotentHint=false).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences: the async contract and the retrieval path come first, then the voice-selection note. No filler, every clause carries actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an async generation tool with an output schema and fully documented parameters, the description supplies the one thing the schema cannot: that the response is a queued generationId needing a get_generation poll. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented, including the profileId precedence rule. The description only restates the profileId/profileName selection without adding format or constraint detail beyond the schema, which matches the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (generate) and resource (speech) scoped to the user's VoiceLabs voice profiles, which cleanly separates it from transcribe (the inverse operation) and the voice-profile management siblings. An agent can identify the tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit follow-up workflow: poll get_generation with the returned generationId to retrieve the audio URL, and explains how to choose the voice via profileId or profileName. It does not state when *not* to use this tool or contrast it against transcribe, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transcribeTranscribe audioAInspect
Transcribe an audio file to text using VoiceLabs' Whisper. Give the audio exactly one way: audioUrl, a public https:// URL serving the file itself (wav, mp3, m4a, aac, ogg, flac, aiff or webm; up to 10 MiB), or audioBase64, the file's real bytes (≤10 MiB of base64 text). A payload that is not real audio — under 1 KiB, a bare header, or made-up placeholder bytes — is refused without being transcribed. If the user has an audio file you cannot read the bytes of, ask them for a public link, or to upload it on the Captures page of the VoiceLabs studio (https://app.voicelabs.now/studio). Returns the transcript and the created capture id.
| Name | Required | Description | Default |
|---|---|---|---|
| audioUrl | No | A public https:// URL that serves the audio file itself (wav, mp3, m4a, aac, ogg, flac, aiff or webm; no redirects, up to 10 MiB). Give exactly one of audioBase64 or audioUrl. | |
| language | No | ISO language code (auto-detected when omitted). | |
| audioBase64 | No | The audio file's real bytes as base64 (≤10 MiB of base64 text; at least 1 KiB of audio). Only use this when you actually hold the file's bytes — a made-up or placeholder payload is refused. Give exactly one of audioBase64 or audioUrl. |
Output Schema
| Name | Required | Description |
|---|---|---|
| language | Yes | |
| captureId | Yes | |
| durationMs | Yes | |
| transcript | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare the write profile (readOnlyHint=false, openWorldHint=true, non-idempotent), and the description adds non-obvious behavior on top: fake/placeholder payloads under 1 KiB or bare headers are refused without transcription. It also discloses the return shape (transcript + capture id), which the annotations do not cover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then limits, then validation, then the fallback. Efficient overall, though the final sentence about the Captures page is somewhat long and could be tighter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers input constraints, validation failure modes, and the fallback path, and an output schema exists so return values need not be spelled out. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the mutual-exclusivity rule, format list, and size limits are already documented in the schema. The description restates these (10 MiB, format list, 'real bytes') but adds little parameter-level meaning beyond the schema baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('Transcribe an audio file to text') and names the engine (VoiceLabs' Whisper), which cleanly distinguishes it from siblings like speak and list_captures. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the exactly-one-input rule (audioUrl or audioBase64) and provides a concrete fallback workflow when the agent cannot read the bytes: ask for a public link or direct the user to the Captures page. Conditions for each path are stated rather than implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
- First observed
clone_voice_profile - First observed
ensure_voice_profile - First observed
execute_typescript - First observed
get_generation - First observed
list_captures - First observed
list_voice_profiles - First observed
search_tools - First observed
speak - First observed
transcribe
Related MCP Connectors
Generate AI images, video, music, and sound effects, and upscale them, from any MCP client.
Generate AI images and videos from any compatible MCP client.
Generate game-ready 3D models, textures, and audio from natural language, over MCP.
MCP server for AI dialogue using various LLM models via AceDataCloud
Related MCP Servers
- FlicenseAqualityCmaintenanceAn MCP server that enables voice-to-voice AI conversations using ElevenLabs for speech synthesis and recognition, with tools for voice management, text-to-speech, and speech-to-text.7-
- AlicenseNot gradedqualityCmaintenanceAn MCP server for AI voice synthesis with an inline audio player, allowing users to give their AI assistant a custom cloned voice using DashScope or ElevenLabs.MIT
- AlicenseAqualityDmaintenanceAn MCP server that enables LLMs to generate spoken audio from text using OpenAI's Text-to-Speech API, supporting various voices, models, and audio formats.113 npm1MIT
- FlicenseNot gradedqualityCmaintenanceAn MCP server that provides text-to-speech, speech-to-text, and voice management via ElevenLabs API.1-
Glama MCP Gateway
Add one secure layer between your agents and this server.