The Complete Guide to Director Mode

Director Mode: three lines of instruction keep AI in character

Emotion tags tell AI "how this line feels"; Director Mode tells it "who you are, where you are, and how to speak". This guide breaks the three-line instruction down word by word — everything is copy-ready.

The free plan includes about 30 minutes of voice per month. No credit card required.

Hear Director Mode in action (Baihua · Late-night Radio)

Style Notes field (instructions go here)

Character: A late-night radio host
Scene: A live studio at 2 a.m.
Direction: Keep it low and unhurried.

Line box (the content to be spoken)

...The city is asleep. I am still here.

Anatomy of the three-line instruction

Director instructions go into the Studio's "Style Notes" field — not a single word is read aloud; every character in the line box is read aloud.

1

Line 1 · Character (who is speaking)

Identity sets vocal texture and word choice — the more specific the better: "a late-night radio host" beats "a host"; "a friend who works the late shift" paints a clearer picture than "a friend".

Character: A late-night radio host
2

Line 2 · Scene (where they speak)

Space and time give AI the environment of the performance. The same line sounds completely different in a 2 a.m. studio versus a busy shopping mall.

Scene: A live studio at 2 a.m.
3

Line 3 · Direction (how to speak)

Concrete performance notes: pace, resonance, breath, pauses. This is the most "tunable" line — see the Four-Dimension Direction below.

Direction: Keep it low and unhurried.

Copy-ready template

Character: (who is speaking)
Scene: (where)
Direction: (how)

Four-Dimension Direction: four knobs for tuning the "Direction" line

The more specific the "Direction" line, the more controllable the performance. The four dimensions work alone or in combination — same line, change one dimension at a time, and compare:

First two lines shared

Character: A gentle storyteller
Scene: Inside a car on a winter night

Line shared by all four demos

The night wind is cold. We parked by the roadside, and neither of us said a word.

1Resonance

Where the voice is supported — chest resonance is deep and solid; head resonance is bright and open.

Direction: More chest resonance, keep the voice deep and solid

2Pacing & Pauses

The rhythm of speed and pauses — slowing half a beat and leaving space often grips more than rushing.

Direction: Slow the pace, pause half a beat at commas, let the ending breathe

3Breath–Voice Shifts

The balance and transitions between breathy and full voice — breathy is intimate, full voice is assured; gradual shifts sound most natural.

Direction: Start with breathy onset, shift to full voice mid-line, fade out on a soft breath

4Articulation

The clarity of consonants and endings — crisp articulation reads professional; soft, lazy edges read familiar.

Direction: Crisp articulation, land every onset and ending clearly, no mumbling

Pro Tip:Tap to listen — hear what changing a single dimension does to the same line

4 complete scene templates (copy and use)

Each template is a full script of "three-line instruction + line" — paste it into the Studio and it works.

Late-night Radio Opener

Low-temperature storytelling — slower pace, resonance sinking down.

Character: A late-night radio host
Scene: A live studio at 2 a.m.
Direction: Keep it low and unhurried.
...The city is asleep. I am still here.

Ad Read

Energy and urgency — crisp articulation, every line pushing forward.

Character: A high-energy ad voice actor
Scene: A prime-time TV commercial
Direction: Crisp articulation, slightly faster pace, rising cheerful endings
Still struggling with voiceovers? In seconds, turn text into pro narration!

Audiobook Narration

Built for long listening — breathy foundation, steady pace.

Character: An audiobook narrator
Scene: A quiet recording booth
Direction: Steady pace, breathy line openings, fade out gently at line ends
That night, the first snow fell on the little town.

Short-video Narration

Hook fast — the first two seconds must grab people.

Character: A short-video narrator
Scene: A bedroom studio late at night
Direction: Open with high energy to grab attention, then lower and slow down for the reveal
Can you believe it? This beef noodle shop has been hidden in this alley for thirty years.

Common failures and fixes

The emotion lands flat — it still sounds "neutral"

Add one concrete action to the "Direction" line (e.g. "with breath, slow the pace"). Vague commands like "more emotional" do nothing; concrete actions work.

Overacted — AI tries too hard

Delete stacked adjectives and keep one direction only. When instructions fight each other (excited AND gentle), the AI can only flail.

Tags or instructions get read aloud

Check punctuation width and placement: (style) at line start, [audio] at line end; the three-line instruction only goes in the "Style Notes" field — put it in the line box and it will be read aloud.

The first generation misses — don't rush to rewrite

Generate several takes first (same input ≠ same output), pick the closest one, then fine-tune. Out of ten takes, one is always right.

The iteration workflow: three takes, pick one

Director Mode is a probability game — the same prompt performs differently every time. The mature workflow:

  1. 1

    Generate 3 takes with the same prompt in a row

  2. 2

    Listen and compare, pick the take closest to your imagination

  3. 3

    Fine-tune the instruction (add or remove one dimension) and regenerate

Listen to how these three takes differ — identical text and instruction, three independent syntheses:

Same prompt · three takes compared

Director Mode FAQ

No. The three-line instruction lives in the "Style Notes" field and is only sent to the model as performance guidance; only the content of the line box is spoken.

No hard limit, but 1-3 lines of concrete instruction work best. Piling up instructions makes them fight each other and dilutes the point.

Yes. Director Mode is independent of the voice source — cloned, designed, and built-in voices all take the three-line instruction, fully compatible with (style) tags.

Emotion tags are a per-line "emotion filter"; Director Mode shapes the whole passage's "character and space". Common combo: three-line instruction in the Style Notes field, plus a (Cheerful) tag at the start of the line for per-line tweaks.

Yes. Director Mode is built into the Studio. The free plan includes about 30 minutes of voice per month, no credit card required.

The template is already written for you

Take the three-line instruction above into the Studio, swap in your line — and your next "actable" voice appears in two minutes.

The free plan includes about 30 minutes of voice per month. No credit card required.

Keep digging