Official Player Handbook

Soundwaver Guide

The complete quest handbook from newbie village to voiceover master — six levels that unlock the superpower of voice, step by step. Not just how to use it, but how to use it brilliantly!

Estimated clear time: About 10 min
1

Level 1 · Quick Start

Quick Start

Generate your first voice in 3 steps

First time in Soundwaver? Follow these three steps and hear AI speak within a minute.

1

Enter the Studio

Click "Studio" in the top-right corner to see the three feature tabs. Beginners should stay on "Text to Speech" for now.

No installation needed — open your browser and start creating.

2

Type Text, Pick a Voice

Type any sentence, then choose a voice from the list. For English try Aria, Luna, Rex, or Cole.

Sliders adjust speed (10-200%) and pitch — start with defaults first.

3

Generate, Listen & Download

Hit generate and listen within seconds. Happy with it? Download as WAV or MP3 for your videos, slides, or podcast.

Every generation is saved to History automatically — no panic if you close the tab.

2

Level 2 · Text to Speech

Text to Speech

Master tags, styles & sing mode

Text to Speech is more than a "script reader". Learn the panel and the tag spells, and it can perform an entire play for you.

Know the Control Panel

Voice Picker

Built-in voices each have a language and personality — Aria, Luna, Rex, Cole for English; Qingqing, Rourou, Leilei, Chenfeng for Chinese.

Speed Slider

Free adjustment from 10% to 200%. Narration works best at 90-110%; fast-paced ads can go up to 130%.

Pitch Slider

Lower it for a deep, magnetic tone; raise it for a cute, childlike sound. Fine-tune together with speed for the most natural result.

Sing Mode

Once enabled, AI performs with melody — great for jingles, birthday wishes, or fun short videos.

Tag System: Magic Spells for Voice Acting

Insert tags into your text to direct the AI's emotion and breath. Remember just two formats:

(Style/Language tags) → at the start

Wrap in parentheses and place at the very beginning. Stack multiple tags to set the acting tone of the whole line.

(Cheerful)(Whisper)(Sad)(Angry)(中文)

(Excited)We won! What an incredible match!

[Audio tags] → at the end

Wrap in square brackets and append to the end of a sentence to control pauses, breaths, and realistic sounds.

[Pause][Sigh][Laugh][Breath]

This news is truly unbelievable[Sigh] give me a moment.

Interactive Dojo

Tap a tag to add it to the line — feel the AI acting change!

Welcome to Soundwaver — the playground of voices. ← Tap tags to try

( ) (Style/Language tags)

[ ] [Audio tags]

Tap a tag and it synthesizes and plays right away — hear it first!

Pro Tip:Generate the same line with different tags and compare — the fastest way to build "tag intuition".

3

Level 3 · Voice Cloning

Voice Cloning

Clone any voice from one clip

Want the AI to sound like you (or someone who has authorized you)? Voice Cloning needs only two things: a clean audio clip and its transcript.

1

Prepare a Clean Clip

A clear 5-30 second recording is ideal: quiet room, stable distance from the mic, natural pace. MP3, WAV, FLAC and more are supported.

2

Upload Audio + Transcript

Type out what is said in the clip and submit both together. The closer the transcript matches the audio, the more accurate the cloned timbre and rhythm.

3

Generate & Name Your Voice

After a short wait, the new voice appears in your voice list. Give it a memorable name so you can call it anytime in Text to Speech.

Do This

  • Record in a quiet space with no echo or noise
  • Keep the clip clear and the pace natural
  • Get the voice owner's explicit consent first
  • Make the transcript match the audio content

Avoid This

  • Background music, fans, or keyboard sounds
  • Clips too short (under 5s) or heavily chopped
  • Cloning someone's voice without consent
  • Audio processed by a voice changer

Pro Tip:Clone quality is 80% about the material. Spend five minutes re-recording a clean clip instead of forcing a noisy meeting recording.

4

Level 4 · Voice Design

Voice Design

Craft a unique voice from words

No audio on hand? No problem — Voice Design lets you "sculpt" a never-before-heard voice from words. The CharacterWizard walks you through five steps.

1

Character Setup

Decide who this voice is: age, gender, personality, and use case. The more specific the description, the better AI captures the feel.

2

Voice Source

Choose to design a brand-new voice from scratch, or blend traits from existing voices as the base.

3

Persona Details

Describe the timbre in words, e.g. "a 30-year-old man, deep and raspy, like a late-night radio host". Not sure what to write? Apply a built-in description template or let AI polish expand it.

4

Generate & Preview

AI generates candidate voices from your description. Not satisfied? Adjust the words and regenerate until it clicks.

5

Save the Character

Name and save the voice, then pick it directly from the voice list in Text to Speech.

Good"A 60-year-old grandpa, raspy with a hint of laughter, slow-paced, like a grandfather telling stories by the fireplace."

Bad"an old voice." — too vague; the AI can only guess.

Three Common Pitfalls
  • • Gender without age or timbre (e.g. "a woman") — the AI can only guess randomly
  • • Stacking conflicting adjectives (deep AND bright AND lively) — they cancel each other out
  • • Writing stage directions into the description (e.g. "please sound excited") — describe the voice itself instead

Pro Tip:Describing "what it sounds like" beats listing parameters. Use professions, scenes, and metaphors — "like a late-night radio host" works better than "low frequency, slow speed".

5

Level 5 · Voices & History

Voices & History

Manage your voice library

The longer you use it, the more voices and files you accumulate. This level teaches you to keep your voice assets in perfect order.

Custom Voice Library

Cloned and designed voices live in the voice list alongside built-in ones. Click to preview or use them for synthesis directly.

History Records

Every generation is saved automatically — replay, download, or delete anytime. Searching by content keywords is the fastest way to find a specific file.

Usage Monitoring

The "Usage" tab in the Dashboard shows your Credits balance in real time. Upgrade before you run out.

Pro Tip:Build the habit of downloading right after generating. History is convenient, but keeping the original file yourself is the safest.

6

Level 6 · Pro Tactics

Pro Tactics

Hidden tricks of director-level dubbing

Cleared all the basics? Here are the hidden tricks that take your voiceover from "usable" to "professional" — battle-tested by power users.

1

Stacked-Tag Acting

Style tags stack — "(Cheerful)(Whisper)" produces a mysteriously lowered yet happy delivery. Experiment with combos to unlock brand-new expressions.

2

Director's Script Style

Give directions like a director: "(Excited) announce it like a sports commentator", then the line. AI understands role settings and stays in character.

3

Chunk Long Scripts

Don't generate long scripts in one go. Split by paragraph, tune speed and emotion per chunk, then stitch together for far better pacing control.

4

Punctuation Is Acting

Ellipses "..." create hesitation, exclamation marks raise the volume, question marks lift the intonation. Change punctuation, not words, and the emotion shifts instantly.

5

Mix Singing & SFX

Open with a jingle in Sing Mode, switch to narration for the body, end with [Laugh] — a memorable show intro in three minutes.

6

Cloned Voice × Tags = Your Own Voice Actor

Cloned voices eat tags too! Add (Sad) or raise the speed, and one voice instantly gains multiple acting ranges.

Full Director-Mode Example

(Excited) announce like a sports commentator — Ladies and gentlemen, ten seconds left... and it's in!
(Deep) switch to a late-night documentary narration... this city never truly sleeps.
(Gentle) like telling a bedtime story, slow it down: and so, the little fox wandered into the forest.

Completion Checklist0/3

Pro Tip:Write down great parameter combos (voice + speed + pitch + tags) in a note. Reuse them next time for stable quality and saved time.

Practice in the Studio