clone_voice
Clone a voice you own or have permission to use, from a single audio sample. Returns a reusable voice_id for text_to_speech. Requires consent=true; cloning a voice to impersonate someone is prohibited. High-fidelity reproduction capturing tone, cadence, and accent. SAMPLE REQUIREMENTS: MP3, M4A or WAV, 10 seconds to 5 minutes, 20 MB max, one speaker, no background music — a sample outside these limits is rejected upstream and the clone fails. Turbo (faster) or HD (higher quality) modes. 7,500 sats per clone. Pay per request with Bitcoin Lightning — no API key or signup needed. Requires create_payment with toolName='clone_voice'.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Voice model: turbo (faster) or hd (higher quality) | speech-02-turbo |
| consent | Yes | Required. Set true to attest that you own this voice or have the speaker's permission to clone it. Cloning a voice to impersonate someone is prohibited (sats4ai.com/terms). | |
| accuracy | No | Text validation accuracy 0-1 (default 0.7) | |
| paymentId | Yes | Valid payment ID (must be paid) | |
| voiceFileUrl | Yes | Public URL to audio file of the voice to clone. MP3, M4A or WAV, 10 seconds to 5 minutes, 20 MB max. One speaker, no background music. |