describe_voice
Preview how free-text voice descriptions map to supported design attributes (gender, age, pitch, accent) before generating a voice; unmatched terms are reported.
Instructions
Preview how a voice description maps onto voice-design attributes.
Nothing is saved. Use this before design_voice to see what the
description will produce. The design space is small and fixed; only
these tokens (and close synonyms) are understood:
Gender: male, female
Age: child, teenager, young adult, middle-aged, elderly
Pitch: very low / low / moderate / high / very high pitch
Style: whisper
EnglishAccent: american, british, australian, canadian, indian,
japanese, korean, chinese, russian, portuguese accent
ChineseDialect (Chinese speech; overrides an English accent):
sichuan, dongbei / northeastern chinese, henan, shaanxi, gansu,
guilin, guizhou, jinan, ningxia, qingdao, shijiazhuang, yunnan
(e.g. "sichuan dialect")
Timbre words ("gravelly", "raspy") and other accents are ignored and
reported in `unmatched`. For a voice outside this space, use
clone_voice with reference audio instead.
Args:
description: Free-text description, e.g. "an elderly man with a
deep voice and a british accent".
Returns:
JSON with attrs (category → token or "Auto"), instruct, matched
and unmatched.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| description | Yes |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| result | Yes |