step_text_to_speech
Transform text, scripts, or narration into lifelike speech audio. Specify voice, speed, volume, and style for tailored spoken output.
Instructions
Generate speech audio with Step Plan TTS. Use this when the user asks to turn text, scripts, narration, ads, reports, or dialogue into an audio file. Before calling, pass the exact text to be spoken, not a vague reference to prior conversation.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | Exact text to synthesize. Include the full text; do not pass vague references. | |
| speed | No | Optional speech speed multiplier. | |
| voice | No | Step TTS voice id. Default linjiajiejie, a warm natural female voice. | linjiajiejie |
| volume | No | Optional volume multiplier. | |
| instruction | No | Optional speaking style, emotion, pacing, or performance instruction. | |
| response_format | No | Audio file format. Default mp3. | mp3 |