Why choose S1?
The platform guide recommends S1 as a choice focused on long-text stability, suitable for content where clear narration is the priority. Actual performance depends on the script, voice reference, and generation settings, so you should first listen to representative passages.
Where should I specify the model?
Specify s1 through the HTTP request header model. If unspecified, the default is s2-pro; the voice reference_id is used for voice selection and is configured separately from the synthesis model.
How do I choose between saving a voice and one-time cloning?
For fixed characters or serialized content, use reference_id to reuse a voice; for one-time projects, use references to provide recordings and transcripts. The two are mutually exclusive, and one-time cloning accepts only one reference sample.
How do I control speech speed and output format?
prosody.speed=1.0 indicates the original speed, and volume uses dB, with 0 meaning no change in volume. You can choose mp3, wav, or pcm; the latter two both use the WAV container, and MP3 bitrates can be 64, 128, or 192.
How do I track completion for long scripts?
After setting callback_url, first save task_id and started_at, then wait for the completion callback; you can also query by task ID. Only after obtaining audio_url can you proceed to playback or editing.