Skip to main content
Glama

Generate Sequencer audio

generate_audio
Destructive

Generate AI audio through Sequencer. Use this whenever the user asks to make, create, generate, or render speech, voiceover, narration, music, a song, or a sound effect with Sequencer or names an audio/music model/provider Sequencer supports: Cartesia, ElevenLabs, Fish Audio, Lyria, Suno, or another audio model. If no speech model is specified, Sequencer uses Cartesia Sonic 3.5 (cartesia-sonic-3-5). An active skill's explicit model still takes priority. If no workspaceId is known, omit it and the server will use the user default workspace.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modelNoModel ID, for example cartesia-sonic-3-5 for Cartesia Sonic 3.5, elevenlabs-tts-v3 for speech, lyria-002 for music, or another ID returned by get_model_catalog. Uses cartesia-sonic-3-5 if not specified.
promptYesText to speak (for TTS) or music prompt
voiceIdNoVoice ID for TTS models
maxCostUsdNoMaximum charge for each output. Checked against authoritative pricing before generation.
sparkTaskIdNoOriginating Spark task for library history and recovery.
stylePromptNoDelivery and accent direction for supported speech models, for example: warm Argentine Spanish with natural Rioplatense pronunciation.
workspaceIdNoOptional workspace ID. If omitted, Sequencer uses the user default/personal workspace.
idempotencyKeyNoStable request key. Retries with the same key reuse the existing output and do not start another paid generation.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
dataNo
errorNo
successYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=true, so the paid/generative nature is covered structurally. The description adds the model-resolution precedence ('An active skill's explicit model still takes priority') and the workspace-defaulting behavior, but those largely restate the model and workspaceId schema descriptions, and it never states outright that this starts a billable generation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and trigger conditions are front-loaded in the first two sentences, which is the right ordering for a router-style tool. The trailing sentences about default models and workspaceId partly duplicate the schema, adding mild redundancy, but the text stays free of filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and the description covers provider coverage, default model fallback, skill precedence, and workspace defaulting. It leaves cost behavior and the interaction between voiceId/stylePrompt and specific model families to the schema, which is acceptable but not exhaustive for a paid, destructive generation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all eight parameters are already documented in the schema, making 3 the baseline. The description reinforces model defaults and workspaceId omission, and adds the skill-priority rule, but contributes no syntax or format detail beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('Generate AI audio through Sequencer') and enumerates the modalities it covers: speech, voiceover, narration, music, song, sound effects. It does not, however, differentiate itself from close siblings such as generate_audio_track or generate_video_audio, so an agent still has to infer which generator to pick for timeline- or video-scoped audio.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives strong positive routing: 'Use this whenever the user asks to make, create, generate, or render speech, voiceover, narration, music, a song, or a sound effect with Sequencer or names an audio/music model/provider Sequencer supports.' It also names the supported providers and the default model resolution rule. What is missing is any when-not guidance relative to the adjacent audio tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources