The complete quest handbook from newbie village to voiceover master — six levels that unlock the superpower of voice, step by step. Not just how to use it, but how to use it brilliantly!
Level 1 · Quick Start
Generate your first voice in 3 steps
First time in Soundwaver? Follow these three steps and hear AI speak within a minute.
Click "Studio" in the top-right corner to see the three feature tabs. Beginners should stay on "Text to Speech" for now.
No installation needed — open your browser and start creating.
Type any sentence, then choose a voice from the list. For English try Aria, Luna, Rex, or Cole.
Sliders adjust speed (10-200%) and pitch — start with defaults first.
Hit generate and listen within seconds. Happy with it? Download as WAV or MP3 for your videos, slides, or podcast.
Every generation is saved to History automatically — no panic if you close the tab.
Level 2 · Text to Speech
Master tags, styles & sing mode
Text to Speech is more than a "script reader". Learn the panel and the tag spells, and it can perform an entire play for you.
Built-in voices each have a language and personality — Aria, Luna, Rex, Cole for English; Qingqing, Rourou, Leilei, Chenfeng for Chinese.
Free adjustment from 10% to 200%. Narration works best at 90-110%; fast-paced ads can go up to 130%.
Lower it for a deep, magnetic tone; raise it for a cute, childlike sound. Fine-tune together with speed for the most natural result.
Once enabled, AI performs with melody — great for jingles, birthday wishes, or fun short videos.
Insert tags into your text to direct the AI's emotion and breath. Remember just two formats:
Wrap in parentheses and place at the very beginning. Stack multiple tags to set the acting tone of the whole line.
(Excited)We won! What an incredible match!
Wrap in square brackets and append to the end of a sentence to control pauses, breaths, and realistic sounds.
This news is truly unbelievable[Sigh] give me a moment.
Tap a tag to add it to the line — feel the AI acting change!
Welcome to Soundwaver — the playground of voices. ← Tap tags to try
( ) (Style/Language tags)
[ ] [Audio tags]
Tap a tag and it synthesizes and plays right away — hear it first!
Pro Tip:Generate the same line with different tags and compare — the fastest way to build "tag intuition".
Level 3 · Voice Cloning
Clone any voice from one clip
Want the AI to sound like you (or someone who has authorized you)? Voice Cloning needs only two things: a clean audio clip and its transcript.
A clear 5-30 second recording is ideal: quiet room, stable distance from the mic, natural pace. MP3, WAV, FLAC and more are supported.
Type out what is said in the clip and submit both together. The closer the transcript matches the audio, the more accurate the cloned timbre and rhythm.
After a short wait, the new voice appears in your voice list. Give it a memorable name so you can call it anytime in Text to Speech.
Pro Tip:Clone quality is 80% about the material. Spend five minutes re-recording a clean clip instead of forcing a noisy meeting recording.
Level 4 · Voice Design
Craft a unique voice from words
No audio on hand? No problem — Voice Design lets you "sculpt" a never-before-heard voice from words. The CharacterWizard walks you through five steps.
Decide who this voice is: age, gender, personality, and use case. The more specific the description, the better AI captures the feel.
Choose to design a brand-new voice from scratch, or blend traits from existing voices as the base.
Describe the timbre in words, e.g. "a 30-year-old man, deep and raspy, like a late-night radio host". Not sure what to write? Apply a built-in description template or let AI polish expand it.
AI generates candidate voices from your description. Not satisfied? Adjust the words and regenerate until it clicks.
Name and save the voice, then pick it directly from the voice list in Text to Speech.
Good"A 60-year-old grandpa, raspy with a hint of laughter, slow-paced, like a grandfather telling stories by the fireplace."
Bad"an old voice." — too vague; the AI can only guess.
Pro Tip:Describing "what it sounds like" beats listing parameters. Use professions, scenes, and metaphors — "like a late-night radio host" works better than "low frequency, slow speed".
Level 5 · Voices & History
Manage your voice library
The longer you use it, the more voices and files you accumulate. This level teaches you to keep your voice assets in perfect order.
Cloned and designed voices live in the voice list alongside built-in ones. Click to preview or use them for synthesis directly.
Every generation is saved automatically — replay, download, or delete anytime. Searching by content keywords is the fastest way to find a specific file.
The "Usage" tab in the Dashboard shows your Credits balance in real time. Upgrade before you run out.
Pro Tip:Build the habit of downloading right after generating. History is convenient, but keeping the original file yourself is the safest.
Level 6 · Pro Tactics
Hidden tricks of director-level dubbing
Cleared all the basics? Here are the hidden tricks that take your voiceover from "usable" to "professional" — battle-tested by power users.
Style tags stack — "(Cheerful)(Whisper)" produces a mysteriously lowered yet happy delivery. Experiment with combos to unlock brand-new expressions.
Give directions like a director: "(Excited) announce it like a sports commentator", then the line. AI understands role settings and stays in character.
Don't generate long scripts in one go. Split by paragraph, tune speed and emotion per chunk, then stitch together for far better pacing control.
Ellipses "..." create hesitation, exclamation marks raise the volume, question marks lift the intonation. Change punctuation, not words, and the emotion shifts instantly.
Open with a jingle in Sing Mode, switch to narration for the body, end with [Laugh] — a memorable show intro in three minutes.
Cloned voices eat tags too! Add (Sad) or raise the speed, and one voice instantly gains multiple acting ranges.
(Excited) announce like a sports commentator — Ladies and gentlemen, ten seconds left... and it's in! (Deep) switch to a late-night documentary narration... this city never truly sleeps. (Gentle) like telling a bedtime story, slow it down: and so, the little fox wandered into the forest.
Pro Tip:Write down great parameter combos (voice + speed + pitch + tags) in a note. Reuse them next time for stable quality and saved time.