From your first line of speech to director-level performance — an 8-level quest handbook with copy-paste scripts and pitfall lists at every step, turning "it makes sound" into "it acts".
EN & ZH
model & interface
8+
built-in voices
30 sec
to clone a voice
Seconds
text to speech
The free plan includes about 30 minutes of voice per month. No credit card required.
Level 1 · Quick Start
Generate your first voice in 3 steps
After this level you can: generate your first piece of speech within a minute.
First time in Soundwaver? Follow the walkthrough animation and three simple steps, and you will hear AI speak for you within a minute.
Line
STEP 1 · Type what you want to say
The animation loops automatically — pause anytime
Click "Studio" in the top-right corner to see the three feature tabs. Beginners should stay on "Text to Speech" for now — later levels unlock the rest.
No installation needed — open your browser and start creating.
Type any sentence, then choose a voice from the list. For English try Aria, Luna, Rex, or Cole.
Sliders adjust speed (10-200%) and pitch — start with defaults first.
Hit generate and listen within seconds. Happy with it? Download as WAV or MP3 for your videos, slides, or podcast.
Every generation is saved to History automatically — no panic if you close the tab.
Check punctuation before you generate — "…" creates hesitation, "!" raises the volume, "?" adds a rising tone. Change only the punctuation, and the emotion changes instantly.
Got it? Take your first line of speech into the Studio now.
The free plan includes about 30 minutes of voice per month. No credit card required.
Pro Tip:For important lines, generate several takes and pick the best — the same input performs differently every time (see Lv.7 Engineering & Quality).
Level 2 · Text to Speech
Make AI act, not just read
After this level you can: use tags and punctuation to make the same sentence perform completely different emotions.
Text to Speech is more than a "script reader". Learn the control panel, the tag spells, and punctuation-as-acting, and it can perform an entire play for you.
Built-in voices each have a language and personality — Aria, Luna, Rex, Cole for English; Qingqing, Rourou, Leilei, Chenfeng for Chinese.
Free adjustment from 10% to 200%. Narration works best at 90-110%; fast-paced ads can go up to 130%.
Lower it for a deep, magnetic tone; raise it for a cute, childlike sound. Fine-tune together with speed for the most natural result.
Once enabled, AI performs with melody — great for jingles, birthday wishes, or fun short videos.
Insert tags into your text to direct the AI's emotion and breath. Remember just two formats:
Wrap in parentheses and place at the very beginning. Stack multiple tags to set the acting tone of the whole line.
(Excited)We won! What an incredible match!
Wrap in square brackets and append to the end of a sentence to control pauses, breaths, and realistic sounds.
This news is truly unbelievable[Sigh] give me a moment.
Line
Type the line first
The animation loops automatically — pause anytime
Tap tags into the line, hit "Generate" and hear the result instantly — this is the core studio experience.
Tap a tag to add it to the line — feel the AI acting change!
Welcome to Soundwaver — the playground of voices. ← Tap tags to try
( ) (Style/Language tags)
[ ] [Audio tags]
Tap a tag and it synthesizes and plays right away — hear it first!
Change only the punctuation, not a single word, and AI performs a completely different tone — ellipses linger with hesitation, exclamation marks burst open. Listen to this pair:
Like this effect? Use the same voices in the Studio →
Advanced: Director Mode — tell AI "who is speaking, where, and how" →
Bring this tag script into the Studio and hear AI act it out.
The free plan includes about 30 minutes of voice per month. No credit card required.
Pro Tip:Generate the same line with different tags and compare — the fastest way to build "tag intuition".
Level 3 · Director Mode
Three lines that keep AI in character
After this level you can: write three-line director instructions that keep AI in character from start to finish.
Emotion tags control "how this line feels"; Director Mode controls "who this person is, where they are, and how they speak". Put the three-line instruction in the "Style Notes" field (never read aloud), and the line itself goes into the input box as usual.
Hear Director Mode in action (Baihua · Late-night Radio)
Character: A late-night radio host Scene: A live studio at 2 a.m. Direction: Keep it low and unhurried.
...The city is asleep. I am still here.
Like this effect? Use the same voices in the Studio →
Who is speaking — identity sets the vocal texture and word choice
Character: A late-night radio host
Where they speak — space and time both shape the performance
Scene: A live studio at 2 a.m.
How to speak — concrete performance notes on pace, resonance, breath, and more
Direction: Keep it low and unhurried.
Style Notes field (three-line instruction)
Character: (Who is speaking, e.g. a late-night radio host) Scene: (Where, e.g. a live studio at 2 a.m.) Direction: (How, e.g. keep it low and unhurried)
Line box (only the words to be spoken)
...The city is asleep. I am still here.
The "Style Notes" field is the director's chair — instructions go here and are never read aloud; every character in the line box is read aloud.
Character: (Who is speaking, e.g. a late-night radio host) Scene: (Where, e.g. a live studio at 2 a.m.) Direction: (How, e.g. keep it low and unhurried)
Line
STEP 1 · Write the three lines in the Style Notes field
The animation loops automatically — pause anytime
Director Mode works on the free plan — templates already written for you.
The free plan includes about 30 minutes of voice per month. No credit card required.
Pro Tip:Director instructions describe "how to perform" — never put the spoken line inside them; for line-level emotion, layer (style) tags on top.
Level 4 · Voice Design
Sculpt a unique voice from one description
After this level you can: describe a voice in 1-4 sentences and let the CharacterWizard sculpt a unique timbre.
No audio on hand? No problem — Voice Design lets you "sculpt" a never-before-heard voice from words. The CharacterWizard walks you through five steps — watch the animation here:
…
Character setup: who is this voice, where is it used
The animation loops automatically — pause anytime
Decide who this voice is: age, gender, personality, and use case. Keep the description to 1-4 sentences — too short and the AI guesses wildly; too long and the traits fight each other.
Choose to design a brand-new voice from scratch, or blend traits from existing voices as the base.
Describe the timbre in words, e.g. "a 30-year-old man, deep and raspy, like a late-night radio host". Age and timbre must be specific. Not sure what to write? Apply a built-in description template or let AI polish expand it.
AI generates candidate voices from your description. Not satisfied? Adjust the words and regenerate until it clicks.
Name and save the voice, then pick it directly from the voice list in Text to Speech.
Good"A 60-year-old grandpa, raspy with a hint of laughter, slow-paced, like a grandfather telling stories by the fireplace."
Bad"an old voice." — too vague; the AI can only guess.
The description formula: profession + scene + metaphor. "Like a late-night radio host" beats listing parameters like "low frequency, slow speech".
Want to dig deeper into the WaveMind™ voice engine behind this? →
Bring your description to the Design Wizard — your voice in two minutes.
The free plan includes about 30 minutes of voice per month. No credit card required.
Pro Tip:Not satisfied? Change the description and regenerate — iterating over multiple rounds is the norm. Change only one variable per round (age / timbre / tone) and you will land the fastest.
Level 5 · Voice Cloning
Clone any voice from a 5-30s clip
After this level you can: clone a voice from a 5-30 second clip, and make it perform with tags.
Want the AI to sound like you (or someone who has authorized you)? Voice Cloning needs only two things: a clean audio clip and its transcript.
A clear 5-30 second recording is ideal: quiet room, stable distance from the mic, natural pace. MP3, WAV, FLAC and more are supported.
Type out what is said in the clip and submit both together. The closer the transcript matches the audio, the more accurate the cloned timbre and rhythm.
After a short wait, the new voice appears in your voice list. Give it a memorable name so you can call it anytime in Text to Speech.
STEP 1 · Upload a clean 5-30s clip
The animation loops automatically — pause anytime
A cloned voice takes tags too! Add (sad) or raise the speed, and one voice gains many acting styles — see the combo tricks in Lv.8.
Clip ready? Upload it in the Studio to start cloning.
The free plan includes about 30 minutes of voice per month. No credit card required.
Pro Tip:Clone quality is 80% about the material. Spend five minutes re-recording a clean clip instead of forcing a noisy meeting recording.
Level 6 · Sing Mode
(Singing) at the start, and it sings
After this level you can: one (Singing) tag, and AI opens its mouth already singing.
Sing Mode needs no switch to find — add the (Singing) tag at the very front of the line, and AI performs with a sense of melody.
(Singing) Happy birthday to you~ Happy birthday to you~
The (Singing) tag must sit at the very start of the text and applies to the whole sentence. Great for birthday wishes, show intros, and playful greetings.
Sing one jingle line in Sing Mode for the intro, switch back to narration for the body, and end with [Laugh] — a memorable show intro in three minutes. Since (Singing) applies to the whole sentence, the mix trick is to generate the sung part and the spoken part as two separate clips, then stitch them in your editor.
Listen: a full (Singing) line
Like this effect? Use the same voices in the Studio →
Take the jingle recipe into the Studio and sing it out.
The free plan includes about 30 minutes of voice per month. No credit card required.
Pro Tip:Layer style tags like (Cheerful) on top to add emotion; Sing Mode also interacts with the speed slider — listen to the default first, then fine-tune.
Level 7 · Engineering & Quality
Multi-take iteration & long-form dubbing
After this level you can: reliably produce deliverable work — multi-take iteration, long-form chunking, and asset management.
The distance between "usable" and "a good piece of work" is a set of engineering habits. This level upgrades random generation into a controllable production line.
With the same prompt, AI performs differently every time. For important lines, never settle for the first take — hit generate three times in a row and pick the most natural one. Listen to how these three takes differ (same prompt, three independent syntheses):
Same prompt · three takes compared
Split by paragraph or emotion, and generate each chunk on its own
Fine-tune speed, pitch, and tags per chunk
Stitch in your editor — pacing control is far more precise
Cloned and designed voices are collected in the voice list alongside built-in ones. Click to preview, or use them for synthesis directly.
Every generation is saved automatically — replay, download, or delete anytime. Searching by content keywords is the fastest way to find a specific file.
The "Usage" tab in the Dashboard shows your Credits balance in real time. Upgrade before you run out.
No voice of your own yet? See the 30-second voice cloning tutorial first →
Keep every take in History and pick the best one.
The free plan includes about 30 minutes of voice per month. No credit card required.
Pro Tip:Build the habit of downloading right after generating. History is convenient for review, but keeping the original file yourself is the safest.
Level 8 · Pro Tactics
Combos the veterans actually use
After this level you can: combine every skill from the previous levels into a director-grade performance.
Cleared all the basics? Here are the combo tricks power users actually rely on — collected from real-world sessions.
Style tags stack — "(Cheerful)(Whisper)" produces a mysteriously lowered yet happy delivery. Experiment with combos to unlock brand-new expressions.
Not enough flavor in your director instructions? Add one dimension to the "Direction" line (resonance / pacing & pauses / breath-voice shifts / articulation) — precise like a mixing console. Full tutorial in the Director Mode guide.
Ellipses "..." create hesitation, exclamation marks raise the volume, question marks lift the intonation. Change punctuation, not words, and the emotion shifts instantly (demo clips in Lv.2).
Cloned voices eat tags too! Add (Sad) or raise the speed, and one voice instantly gains multiple acting ranges.
Open with a jingle in Sing Mode, switch to narration for the body, end with [Laugh] — a memorable show intro in three minutes.
Save successful three-line instructions in a note, then copy and tweak next time. A veteran's speed comes from templates, not from inventing from scratch every time.
(Excited) announce like a sports commentator — Ladies and gentlemen, ten seconds left... and it's in! (Deep) switch to a late-night documentary narration... this city never truly sleeps. (Gentle) like telling a bedtime story, slow it down: and so, the little fox wandered into the forest.
The tactics are yours — the Studio is your dojo.
The free plan includes about 30 minutes of voice per month. No credit card required.
Pro Tip:Write down great parameter combos (voice + speed + pitch + tags) in a note. Reuse them next time for stable quality and saved time.
Line-by-line anatomy of the three-line prompt, four-dimension direction with audio demos, 4 copy-ready scene examples, failure fixes and the iteration workflow.
The free plan includes about 30 minutes of voice per month. No credit card required.