Emotion tags tell AI "how this line feels"; Director Mode tells it "who you are, where you are, and how to speak". This guide breaks the three-line instruction down word by word — everything is copy-ready.
Hear Director Mode in action (Baihua · Late-night Radio)
Style Notes field (instructions go here)
Character: A late-night radio host Scene: A live studio at 2 a.m. Direction: Keep it low and unhurried.
Line box (the content to be spoken)
...The city is asleep. I am still here.
Director instructions go into the Studio's "Style Notes" field — not a single word is read aloud; every character in the line box is read aloud.
Line 1 · Character (who is speaking)
Identity sets vocal texture and word choice — the more specific the better: "a late-night radio host" beats "a host"; "a friend who works the late shift" paints a clearer picture than "a friend".
Character: A late-night radio host
Line 2 · Scene (where they speak)
Space and time give AI the environment of the performance. The same line sounds completely different in a 2 a.m. studio versus a busy shopping mall.
Scene: A live studio at 2 a.m.
Line 3 · Direction (how to speak)
Concrete performance notes: pace, resonance, breath, pauses. This is the most "tunable" line — see the Four-Dimension Direction below.
Direction: Keep it low and unhurried.
Copy-ready template
Character: (who is speaking) Scene: (where) Direction: (how)
The more specific the "Direction" line, the more controllable the performance. The four dimensions work alone or in combination — same line, change one dimension at a time, and compare:
First two lines shared
Character: A gentle storyteller Scene: Inside a car on a winter night
Line shared by all four demos
The night wind is cold. We parked by the roadside, and neither of us said a word.
1Resonance
Where the voice is supported — chest resonance is deep and solid; head resonance is bright and open.
Direction: More chest resonance, keep the voice deep and solid
2Pacing & Pauses
The rhythm of speed and pauses — slowing half a beat and leaving space often grips more than rushing.
Direction: Slow the pace, pause half a beat at commas, let the ending breathe
3Breath–Voice Shifts
The balance and transitions between breathy and full voice — breathy is intimate, full voice is assured; gradual shifts sound most natural.
Direction: Start with breathy onset, shift to full voice mid-line, fade out on a soft breath
4Articulation
The clarity of consonants and endings — crisp articulation reads professional; soft, lazy edges read familiar.
Direction: Crisp articulation, land every onset and ending clearly, no mumbling
Pro Tip:Tap to listen — hear what changing a single dimension does to the same line
Each template is a full script of "three-line instruction + line" — paste it into the Studio and it works.
Late-night Radio Opener
Low-temperature storytelling — slower pace, resonance sinking down.
Character: A late-night radio host Scene: A live studio at 2 a.m. Direction: Keep it low and unhurried.
...The city is asleep. I am still here.
Ad Read
Energy and urgency — crisp articulation, every line pushing forward.
Character: A high-energy ad voice actor Scene: A prime-time TV commercial Direction: Crisp articulation, slightly faster pace, rising cheerful endings
Still struggling with voiceovers? In seconds, turn text into pro narration!
Audiobook Narration
Built for long listening — breathy foundation, steady pace.
Character: An audiobook narrator Scene: A quiet recording booth Direction: Steady pace, breathy line openings, fade out gently at line ends
That night, the first snow fell on the little town.
Short-video Narration
Hook fast — the first two seconds must grab people.
Character: A short-video narrator Scene: A bedroom studio late at night Direction: Open with high energy to grab attention, then lower and slow down for the reveal
Can you believe it? This beef noodle shop has been hidden in this alley for thirty years.
The emotion lands flat — it still sounds "neutral"
Add one concrete action to the "Direction" line (e.g. "with breath, slow the pace"). Vague commands like "more emotional" do nothing; concrete actions work.
Overacted — AI tries too hard
Delete stacked adjectives and keep one direction only. When instructions fight each other (excited AND gentle), the AI can only flail.
Tags or instructions get read aloud
Check punctuation width and placement: (style) at line start, [audio] at line end; the three-line instruction only goes in the "Style Notes" field — put it in the line box and it will be read aloud.
The first generation misses — don't rush to rewrite
Generate several takes first (same input ≠ same output), pick the closest one, then fine-tune. Out of ten takes, one is always right.
Director Mode is a probability game — the same prompt performs differently every time. The mature workflow:
Generate 3 takes with the same prompt in a row
Listen and compare, pick the take closest to your imagination
Fine-tune the instruction (add or remove one dimension) and regenerate
Listen to how these three takes differ — identical text and instruction, three independent syntheses:
Same prompt · three takes compared
No. The three-line instruction lives in the "Style Notes" field and is only sent to the model as performance guidance; only the content of the line box is spoken.
No hard limit, but 1-3 lines of concrete instruction work best. Piling up instructions makes them fight each other and dilutes the point.
Yes. Director Mode is independent of the voice source — cloned, designed, and built-in voices all take the three-line instruction, fully compatible with (style) tags.
Emotion tags are a per-line "emotion filter"; Director Mode shapes the whole passage's "character and space". Common combo: three-line instruction in the Style Notes field, plus a (Cheerful) tag at the start of the line for per-line tweaks.
Yes. Director Mode is built into the Studio. The free plan includes about 30 minutes of voice per month, no credit card required.
Take the three-line instruction above into the Studio, swap in your line — and your next "actable" voice appears in two minutes.
The free plan includes about 30 minutes of voice per month. No credit card required.