Frequently Asked Questions

Answers to common questions about Soundwaver. If you cannot find what you are looking for, feel free to contact us.

Product

Soundwaver offers three core capabilities: Text-to-Speech (TTS), Voice Cloning, and Voice Design. Type text to generate speech, upload audio to replicate any voice, or describe a voice in words to create a brand-new one from scratch.

Text-to-Speech uses built-in high-quality voices for quick voiceovers; Voice Cloning precisely recreates a voice from your uploaded sample; Voice Design generates an entirely new voice from a text description, with no audio sample required.

Chinese and English are currently supported. Built-in voices each map to a language — for example Qingqing and Rourou for Chinese, and Aria and Luna for English. Voice Cloning and Voice Design can work across languages.

Output supports WAV, MP3, and PCM16, with optional real-time streaming. Input audio supports common formats such as MP3, WAV, OGG, M4A, and FLAC.

We support low-latency real-time streaming output. Non-streaming requests are typically completed within seconds, depending on text length and plan.

Soundwaver runs on our proprietary WaveMind™ voice engine, supporting Chinese, English, and multiple Chinese dialects, with Director Mode for nuanced, actor-like performance control.

Voices

Chinese built-in voices include Qingqing, Rourou, Leilei, and Chenfeng. English built-in voices include Aria, Luna, Rex, and Cole. Pick whichever fits your content style.

We recommend a clear 5–30 second recording with minimal background noise. Common formats such as WAV, MP3, and FLAC are supported.

Voice cloning reproduces timbre, intonation, and speaking style with high fidelity. Using a clear, representative sample yields the best results.

Yes. Describe the desired pace, emotion, role-play, or even dialect in natural language — for example "announce good news in a bright, upbeat, slightly excited tone" — and the matching style will be generated.

Voice Design lets you create a brand-new voice from a text description without any audio sample. Describe timbre, tone, age, gender, and more, and AI will generate a unique voice accordingly.

Pricing & Account

The Free plan offers a limited monthly TTS credits quota and voice clone count, with standard-quality output. Upgrading unlocks higher quotas, HD quality, and more features.

Sign in and go to the "Subscription" tab in your Dashboard to upgrade or manage your subscription. After cancelling, your current cycle remains active until it ends.

Each generation is billed by the audio's actual duration in seconds (1 credit = 1 second). Voice cloning and voice design cost nothing extra. View your Credits balance in real time from the "Usage" tab in your Dashboard.

Privacy & Security

We protect your data with encrypted transport and never share your audio or content with third parties. You can also delete your voices and history from your account at any time.

Yes. You must obtain the voice owner's explicit consent before cloning. Unauthorized cloning or use for any illegal or fraudulent purpose is strictly prohibited.

Paid plans support commercial use for videos, podcasts, ads, and more. The Free plan is limited to personal and non-commercial use only.

Audio files are managed and downloaded by you. Our servers do not permanently store your generated results, and you can delete history and custom voices at any time.