Enhance natural voiceover workflows with next-generation speech capabilities
Fish Audio S2.1 Pro is Fish Audio's recommended production speech model for text-to-speech and voice cloning. Official materials describe improvements in quality, latency, and throughput compared with S2 Pro, making it suitable for applications that require natural voices and ongoing voiceover production. This platform explicitly selects this model through model: s2.1-pro, retaining the interfaces for text, voice references, output formats, and task callbacks.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and interface features
Clarify the model, inputs and outputs, and invocation method before selecting a solution.
Model selection
HTTP request header model: s2.1-pro
Platform interface
POST /fish/tts; text is a required input
Audio delivery
Returns audio_url synchronously; supports asynchronous callbacks via callback_url
Output formats
mp3, wav, pcm; both wav and pcm return a WAV container
Voice reference
reference_id or a single references sample; the two are mutually exclusive
Prosody control
prosody.speed controls speech rate, and prosody.volume controls volume gain
Capability descriptions combine public model materials and this platform's documentation; refer to the relevant API and pricing sections for actual parameters, outputs, and billing rules.
Core capabilities
Learn about the voice tasks that s2.1-pro is suited to handle.
Production-oriented speech upgrade
Officially, S2.1 Pro is positioned as the recommended production model, with stated improvements in quality, latency, and throughput compared with S2 Pro. During selection, verify these changes using your own scripts and voice samples; do not turn official directional descriptions into fixed latency or concurrency commitments.
Reuse existing voice assets
You can use an existing voice through reference_id, or provide a reference recording and accurate verbatim transcript in a single request. Reusing voice materials helps compare new and old models, but generated results still need to be checked for tone, pronunciation, and character consistency.
Organize production tasks in one interface
Submit synthesis using text and format settings, then retrieve audio_url when complete. Long scripts can be configured with callback_url: save the task ID first, then receive the completed result; applications do not need to present a task still waiting as delivered audio.
Applicable Scenarios
Choose based on specific content and delivery method.
Voiceovers for Content at Scale
For frequently updated knowledge content, product explanations, or product tutorials, first select representative scripts for evaluation, then gradually migrate the production workflow. Save model and voice information to make it easier to identify settings when differences in delivery occur.
Application Interactions with Natural Voices
Convert navigation, explanations, and help text into playable audio, focusing on the clarity and waiting experience that real users hear. This endpoint delivers audio links; real-time conversational applications require separate design for playback and buffering.
Upgrade Evaluation for Existing Voiceover Workflows
Compare S2 Pro and S2.1 Pro using the same voice and text, checking proper nouns, long sentences, pauses, and the amount of post-processing required. Use small-batch listening tests to decide the scope of migration, avoiding direct bulk replacement of existing finished content.
How to Choose This Model
Compare based on scripts, voices, and production costs.
Explicitly Use the New-Generation Model
This endpoint still defaults to s2-pro; to use S2.1 Pro, set model: s2.1-pro. Do not assume that a request without a specified model already uses it solely based on its official positioning as the “recommended production model.”
Measure Upgrade Benefits with Your Own Samples
Official materials do not provide a fixed percentage improvement that can be promised for this page. Compare audio quality, task wait times, review, and revision costs before deciding whether to replace an existing workflow; S1 can also remain a candidate for stable long-form narration.
Getting Started
Listen to samples first, then expand to production content.
Prepare Scripts and Voices
Clarify the purpose of the narration, proofread the script, and choose a saved voice or prepare a reference recording and transcript.
Generate Short Samples and Review Them
Explicitly specify s2.1-pro and use representative passages to confirm pronunciation, pacing, and output format.
Integrate into the Production Workflow
Save settings and task IDs, retrieve the audio link after completion, then arrange playback, editing, or content delivery.
Usage Boundaries
Understand the synthesis method and delivery scope.
This endpoint is for text-to-speech and is not equivalent to speech recognition, music generation, or native real-time audio streaming. Long-form content should be produced in segments and reviewed, checking numbers, abbreviations, and proper nouns.
One-time cloning accepts only one publicly accessible HTTPS MP3/WAV recording sample and its accurate transcript; it does not accept Base64, data URIs, or URLs with credentials, and it does not automatically save a long-term voice.
format=pcm returns a WAV container and cannot be processed directly as raw PCM bytes; opus is not supported. For the supported scope of generation parameters and actual billing, see this platform's API documentation and pricing.
Frequently Asked Questions
Answers to common questions when using s2.1-pro.
What has changed in S2.1 Pro compared with S2 Pro?
Official information states that S2.1 Pro improves quality, latency, and throughput, and lists it as the recommended production model. No specific degree of improvement is promised here; migration should be evaluated using your own scripts, voices, and actual call results.
Where should the model name be specified?
Specify s2.1-pro through the HTTP request header model. If not specified, the default is s2-pro; the voice reference_id belongs to voice selection and is configured separately from the synthesis model.
How do I choose between saving a voice and one-time cloning?
For recurring characters or serialized content, you can reuse a voice with reference_id; for one-off projects, you can provide recordings and transcripts with references. The two are mutually exclusive, and one-time cloning accepts only one reference sample.
How do I control speech speed and output format?
prosody.speed=1.0 indicates the original speed, volume uses dB, and 0 means no change to volume. You can choose mp3, wav, or pcm; the latter two both use the WAV container, and MP3 bitrates can be 64, 128, or 192.
How do I track completion results for long scripts?
After setting callback_url, first save task_id and started_at, then wait for the completion callback; you can also query by task ID. Only after obtaining audio_url can you proceed to playback or editing.
Model information · Updated: 2026-10-01. For call parameters and billing rules, see the API and pricing sections.