Generate Sequencer audio
generate_audioGenerate AI audio through Sequencer. Use this whenever the user asks to make, create, generate, or render speech, voiceover, narration, music, a song, or a sound effect with Sequencer or names an audio/music model/provider Sequencer supports: Cartesia, ElevenLabs, Fish Audio, Lyria, Suno, or another audio model. If no speech model is specified, Sequencer uses Cartesia Sonic 3.5 (cartesia-sonic-3-5). An active skill's explicit model still takes priority. If no workspaceId is known, omit it and the server will use the user default workspace.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Model ID, for example cartesia-sonic-3-5 for Cartesia Sonic 3.5, elevenlabs-tts-v3 for speech, lyria-002 for music, or another ID returned by get_model_catalog. Uses cartesia-sonic-3-5 if not specified. | |
| prompt | Yes | Text to speak (for TTS) or music prompt | |
| voiceId | No | Voice ID for TTS models | |
| maxCostUsd | No | Maximum charge for each output. Checked against authoritative pricing before generation. | |
| sparkTaskId | No | Originating Spark task for library history and recovery. | |
| stylePrompt | No | Delivery and accent direction for supported speech models, for example: warm Argentine Spanish with natural Rioplatense pronunciation. | |
| workspaceId | No | Optional workspace ID. If omitted, Sequencer uses the user default/personal workspace. | |
| idempotencyKey | No | Stable request key. Retries with the same key reuse the existing output and do not start another paid generation. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| data | No | ||
| error | No | ||
| success | Yes |