Same text-to-speech engine. Two scripts. One sounds like a person telling you something. The other sounds like a vending machine reading its own manual.
For a long time I blamed the voice model. Then I ran my own client script and the client’s rewritten version through the identical engine, back to back, and had to admit something annoying: the tool reads what you give it, and what you give it decides whether it sounds human.
Written language and spoken language are different species. You wouldn’t hand a chef a poem and expect dinner. Yet people paste polished blog posts into a voice generator and act betrayed when the result sounds like a hostage reading a statement.
These are the eight rules I use on every voiceover script I write. Each one comes with a before and after, because the difference only makes sense when you can hear it in your head.
Rule 1: One Sentence, One Breath
Humans breathe. If a sentence runs past about twenty words without a stop, the listener starts worrying about you.
Clunky: “Our platform supports multi-device sync and encrypted transfer so you can use it安心ly on any device you own without worrying about your data.” One lump. By the end, nobody knows what mattered.
Fixed: “Our platform syncs across devices. And everything is encrypted — so your data stays yours.” Two sentences. The point surfaces on its own.
Short sentences also give the engine somewhere to vary its pitch. Long sentences flatten out, and flat pitch is the single loudest tell of synthetic speech.
Rule 2: Conclusion First, Reasons Second
Old broadcast advice: give the answer, then the story. Attention is rented, not owned, and you have about thirty seconds of goodwill.
Weak: “After six months of work by the team, through three rounds of architecture changes, we are excited to announce —” They are already gone.
Strong: “It’s three times faster. Here’s how we got there.” Lead with the payoff.
This fixes a second problem for free: engines tend to slide downward in pitch across long wind-ups, so by the time your conclusion arrives, the voice has run out of batteries. Put the conclusion first and the energy starts high.
Rule 3: Cut Half Your Adjectives
Formal writing leans on adjectives for weight. Speech leans on concrete nouns. “An absolutely delicious culinary experience” is empty out loud. “A burger that sprays juice when you bite it” is a picture.
My drafting pass: search for “very,” “really,” “extremely,” “incredibly,” delete them all, read it again. The sentences almost always come out not weaker but faster. Listeners don’t count adjectives; they remember images.
Rule 4: Punctuation Is Your Emotion Remote
If you remember one line from this article, make it this: the AI doesn’t feel your emotion — it executes your punctuation.
- Exclamation point = more energy
- Ellipsis = slow down, leave air (never more than two in a row; the voice dies)
- Em dash = pivot, or the vocal version of biting your tongue
- Period = clean landing
- Comma = keep going, shorter gap
So instead of writing stage directions like “(say this excitedly)” — which engines ignore — write the excitement into the marks. Directing by stage direction loses. Directing by punctuation wins.
Rule 5: Spell Every Number Out
“3.5 million,” “episode 12,” “September 21, 2026” — these are the highest-crash zones in any script, and I covered the why in my blooper article. Here is just the rule: the script serves the ear, and the ear wants spoken words, not displayed ones.
Money: “three point five million,” said the way a human says money. Dates: “September twenty-first.” Years: whatever way you’d actually say it out loud. Thirty seconds of typing saves three minutes of re-recording.
Rule 6: Hook in the First Three Seconds
Course, ad, podcast — the first line decides everything, and text-to-speech is merciless here because the opening sentence is usually where the engine’s pitch is flattest; it hasn’t locked onto an emotional baseline yet.
So the first sentence is always short and always incomplete in information.
Flat: “In this video, we’re going to cover three time-management techniques.” Accurate. Fatal.
Alive: “What was the first thing you procrastinated on today?” The ear leans forward on its own. This isn’t a copywriting trick. It’s a survival trick.
Rule 7: If Your Mouth Stumbles, the Ear Will Too
The oldest check in the book: read it out loud yourself.
Your mouth catches things your eyes have approved a hundred times — overlong proper nouns, three "the"s in a row, that cluster of homophones you thought nobody would notice. Everywhere you trip, the engine will glide instead, and that glide is precisely what manufactured-sounding speech is made of.
My habit: wherever the read goes awkward, rewrite on the spot instead of powering through. Keep going until you can say the whole thing in one breath. There is no shortcut around this rule, but it eliminates more problems than the other seven combined.
Rule 8: Leave the Voice Somewhere to Breathe
Plain text has no room in it. A person recording naturally pauses, inhales, eases off. An engine needs you to mark those spots in words.
Three practical moves:
- Put a period before the punchline or key claim — it buys about a third of a second of air.
- Never butt paragraphs tightly; one extra period between them costs nothing.
- When an important term appears the first time, frame it with commas so it gets a beat on either side.
In other words: write the edit into the script. People who do this are the ones whose output survives the “wait, is that AI?” test.
Get the Script Right and the Tool Follows
All eight rules compress into a single instruction: write with your mouth, not your eyes.
Your eyes will happily read sentences that are miserable to say. Your mouth is an honest critic. Read your script aloud before you generate — that one action outperforms switching tools ten times over.
I have watched people stay stuck in the belief that the problem is the tool: they upgrade the plan, swap to the newest model, and still hand it bookish wall-of-text prose. Three engine upgrades later, the listening experience is unchanged, because the input never moved.
Meanwhile one of my clients changed exactly one thing — rewrote every script to these eight rules. Same engine, same voice, and the comments section asked for the first time whether the voice actor was a real person. The tool didn’t change. The script did.
So before you ask which TTS sounds the most human, ask a blunter question: does the script I’m feeding it sound like something a person would actually say?
In my experience, that question has better follow-through rate than any pricing page on the internet.


