Ace Data Cloud · AI Speech Synthesis

Fish Audio Speech Synthesis and Voice Cloning

High-expressiveness text-to-speech (TTS) powered by Fish Audio. A single piece of text can synthesize natural human speech, supporting three models: s1 / s2-pro / s2.1-pro, voice cloning, and searchable access to a library of 1.8M+ public voices. The request body is consistent with the official version—just replace it with this platform's Token.

🐟 s1 / s2-pro / s2.1-pro 🎙️ Voice Cloning 🌍 Multilingual Support
1.8M+
Public Voice Library
3
TTS Models Available
$0.011
Per Thousand Characters, Starting At
Seconds
Short Text Synthesis Response

Core Features

From text-to-speech to voice cloning, one API covers the complete speech synthesis workflow.

🗣️

High-Expressiveness TTS

Submit the text to synthesize with text. s2-pro offers stronger emotion and expressiveness, while s1 is more stable and less likely to drift on long text. Switch via the model request header.

🎙️

Voice Cloning

Pass reference_id to reuse an existing voice, or use references[].audio + text to instantly replicate a voice in a single request—making the synthesized voice closely match the specified speaker.

📚

1.8M+ Voice Library

Search public voices via GET /fish/model, filter by language, tag, and title, then use the obtained _id as reference_id.

⚡

Both Synchronous and Asynchronous

For short text, synchronous requests directly return audio_url and the current cost; for long text, pass callback_url to immediately receive a task_id. Upon completion, both callbacks and final Tasks results include the actual billed quota.

One request, get voice audio

Simply provide text and format to synthesize, and synchronously receive an audio_url that can be played or downloaded directly. Field names are consistent with the official Fish API, enabling zero-cost migration.

cURL · POST /fish/tts
curl -X POST https://api.acedata.cloud/fish/tts \
  -H "authorization: Bearer YOUR_API_KEY" \
  -H "content-type: application/json" \
  -d '{
    "text": "The weather is so nice today, let's go for a walk together.",
    "format": "mp3"
  }'
# Synchronously returns a direct link to playable audio (verified):
# {
#   "audio_url": "https://cdn.acedata.cloud/assets/examples/fish/5ade0339-5f11-487e-aacc-06a908271706-8e3fcb0e5547.mp3"
# }

Get started quickly in 3 steps

From obtaining a Token to getting audio, it only takes a few minutes.

1

Get a Token

Apply for an API Token in the console. One Token lets you call all platform services, and free credits are provided on first use.

2

Submit text

Call POST /fish/tts and pass in text and format (mp3 or pcm).

3

Get audio

Synchronously receive audio_url and cost; play directly or bill based on actual Credits. For long text, read cost from the callback and Tasks final result.

What is Fish Audio suitable for?

From audio content to in-app voice, covering most speech synthesis needs.

📖

Audiobooks / Podcasts

Batch-convert articles and manuscripts into natural voices, using s1 to reliably process long-form content.

🎬

Video Dubbing / Voiceovers

Generate narration for short videos, tutorials, and explainers, using prosody to adjust speaking speed and volume.

☎️

IVR / Customer Service Voice

Dynamically generate prompts and scripts, using a fixed voice to maintain consistent brand voice.

♿

Accessibility Read-Aloud

Provide read-aloud capabilities for web pages and documents, outputting pcm for easy real-time client-side concatenation and playback.

🎮

Games / Character Voice Acting

Use voice cloning to create custom voices for characters and quickly cover large amounts of dialogue.

🧑‍💻

Digital Humans / Virtual Hosts

Synthesize scripts into speech in real time to power digital-human voiceovers and virtual-host livestreaming scenarios.

Two Core Capabilities, One API

Text-to-speech and voice management, with field names fully consistent with Fish's official API

🗣️ Text-to-Speech TTS

Submit text to synthesize speech, with support for format(mp3/pcm), prosody speaking speed and volume, mp3_bitrate bitrate, and sample_rate sample rate.

POST /fish/tts model: s1 / s2-pro

🎙️ Voice Management Model

Clone new voices using reference audio, search a public library of 1.8M+ voices, or query details for a single voice by _id—the returned _id can be used directly as reference_id.

POST /fish/model GET /fish/model GET /fish/model/{id}

Fish Audio Pricing

Pay as you go, no subscription fees, no hidden charges.

Voice search (GET /fish/model) is free

Pay as you go
Pay by usage
$0.011 / per 1,000 characters

TTS is billed by synthesized text length; voice creation and search are free

  • ✓ TTS synthesis approx. $0.011 / per 1,000 characters (30% off official price)
  • ✓ Voice creation and search—free
  • ✓ Voice search GET—free
  • ✓ Free credits upon first registration
View pricing details
Enterprise
Custom

Dedicated plans for high-usage teams

  • ✓ Tiered discounts based on usage
  • ✓ Priority support and account manager
  • ✓ Higher concurrency and capacity
  • ✓ SLA guarantee
  • ✓ Private deployment options
Contact sales

Frequently Asked Questions

Common questions about using the Fish Audio API

How does Fish Audio charge?›

TTS is billed by the UTF-8 byte length of synthesized text, approximately $0.011 / thousand characters (that is, $10.5 / million bytes, 30% off Fish's official $15); the s1, s2-pro, and s2.1-pro models are priced the same. Voice creation (POST /fish/model) and voice lookup (GET /fish/model) are free. New registrations receive free credits to try it directly.

Which audio formats are supported?›

The format field in the request body must be explicitly specified; available options are mp3 or pcm. mp3 can use mp3_bitrate (64/128/192) to control the bitrate; pcm is suitable for real-time browser concatenation or client-side post-processing (mixing, speed adjustment).

How do I clone or specify a voice?›

There are two ways: pass reference_id to reuse an existing voice (you can use GET /fish/model to search the public library and obtain the _id, or use POST /fish/model to clone your own voice); or use references[].audio + text for instant cloning in a single request. Choose one of the two.

What is the difference between s1 and s2-pro?›

Switch via the model request header; the default is s2-pro. s2-pro has stronger expressiveness and emotion, making it suitable for short sentences and voiceovers; s1 is more stable and less likely to drift on long text, making it suitable for long-form content such as audiobooks.

What should I do if long-text synthesis is very slow?›

Synthesizing long text in one request may take from over ten seconds to tens of seconds. Pass callback_url in the request body, and the API will immediately return {task_id, started_at}; when complete, it will POST the result back to that URL. You can also use the Fish Tasks API to actively retrieve it by task_id.

How long is the returned audio link valid?›

audio_url points to the platform CDN and can be played or downloaded directly. The extension follows the output format; it is still recommended to save important results to your own storage.

Explore More AI Audio Services

Ace Data Cloud offers a variety of AI audio and music generation APIs

Start Synthesizing Speech with Fish Audio Now

Get natural-sounding voices from a text snippet in seconds. Supports voice cloning and a library of 1.8 million+ voices.