C.W.K.
Stream
AUDIO

ElevenLabs V3

ElevenLabs

Text-to-performance voice synthesis — inline audio tags turn scripts into directed performances.

ElevenLabs V3 — Complete Prompting Guide


Welcome: Why V3 Changes Everything

If you've used ElevenLabs before, you probably already know it can produce remarkably convincing speech. But V3 is a different animal entirely. Earlier models excelled at text-to-speech — taking words and reading them aloud with natural rhythm and intonation. V3 is built for something more ambitious: text-to-performance.

That distinction matters. A text-to-speech model asks, "How would a person read this sentence?" A text-to-performance model asks, "How would an actor deliver this line?" It cares about the emotional arc of a scene, the micro-expressions between words — the exhale before bad news, the nervous laugh that reveals more than the words themselves.

V3 achieves this through inline audio tags: markers you embed directly in your script that direct the voice like stage directions. Want a character to laugh mid-sentence? [laughs]. Should that laughter trail off into something darker? [laughs][sad]. Are they whispering a secret while looking over their shoulder? [whispers][mischievously]. The tags transform your script from a text document into a director's shooting script — and the voice model is your actor.

Who This Guide Is For

  • Audiobook producers who want narrations that pull readers in, not just read words at them
  • Game developers building characters with genuine emotional depth
  • Podcast creators who want their AI-assisted content to feel like a real conversation
  • App and product teams crafting IVR systems and voice assistants that don't feel robotic
  • Content creators making documentary narration, YouTube voiceovers, explainer videos
  • Writers and storytellers who want to hear their work performed, not recited
  • Developers building voice-enabled applications via the ElevenLabs API

This guide covers everything: model specs, all audio tags with usage examples, stability settings, voice selection, multi-speaker formatting, and a large library of production-ready example scripts. Whether you're new to ElevenLabs or migrating from V2, you'll find everything you need here.


Quick Start — Three Scripts to Try Right Now

Before we go deep on technique, here are three copy-paste scripts that demonstrate V3's range. Drop any of these into the ElevenLabs Studio, select a voice, and hit Generate.

Quick Start 1: Narration (Audiobook Style)

Best voice type: Warm, measured narrator. Stability: Natural.

PROMPT
[warmly] There's a particular kind of silence that only exists in libraries. Not the absence of sound — you can still hear pages turning, a distant cough, the soft percussion of keyboards. It's a silence of intent. Everyone in this room has chosen, deliberately, to be somewhere quieter than the rest of the world.

[sighs] Marcus had spent most of his adult life in rooms like this one. Not because he loved books — though he did — but because he loved the people who loved books. You could learn almost everything about a person from the way they treated a library.

[curious] So when the woman across the reading table slid a folded note toward him without ever looking up from her page, he felt something he hadn't felt in years.

[whispers] Curiosity.

Quick Start 2: Dialogue (Podcast Style)

Best voice type: Two contrasting voices — one energetic, one dry/sardonic. Stability: Natural.

PROMPT
Speaker 1: [excitedly] Okay, I have to tell you about my week, because things got VERY weird.
Speaker 2: [dryly] Your weeks are always weird. That's your brand.
Speaker 1: [laughs] Fair. But this is a different kind of weird. So you know I've been doing that productivity experiment — no phone before 10 AM?
Speaker 2: [curious] The one you announced with the energy of someone quitting caffeine forever and then abandoned after four days?
Speaker 1: [sighs] That was a different experiment. This one stuck. And I think... it's actually working?
Speaker 2: [impressed] Wait, genuinely? Like, productive mornings, clearer head, the whole thing?
Speaker 1: [warmly] The whole thing. I wrote twelve pages on Tuesday and didn't even notice time passing.
Speaker 2: [mischievously] Twelve pages. I write twelve words on a good Tuesday.
Speaker 1: [laughs] You're a chaos demon and you thrive on chaos. This is known.
Speaker 2: [dramatically] I prefer "creatively unstructured."

Quick Start 3: Dramatic Monologue (Theatre / Game Character Style)

Best voice type: Deep, expressive, emotionally broad. Stability: Creative.

PROMPT
[dramatically] They told me it would be simple. Walk in. Take the documents. Walk out. [laughs bitterly] They always say it'll be simple.

[whispers] Nobody tells you about the weight of it. Not the physical weight — the case was light enough. I mean the weight of knowing. Of carrying something that could unravel everything, and having to smile through customs like you're just another tourist with too many souvenirs.

[sighs] I've crossed seventeen borders in the last eight months. I speak four languages fluently and two badly. I know how to disappear in a crowd. I know how to make someone trust me in under three minutes.

[sad] What I don't know... is how to go home. I'm not sure "home" is a place that still exists for me.

[with quiet resolve] But I'll finish this. Whatever it takes. Because if I don't...

[exhales] ...nobody else will.

1. Model Identity & Specifications

What Is Eleven V3?

FieldValue
Product NameEleven V3
Model IDeleven_v3
DeveloperElevenLabs
Launch DateJune 2025 (Alpha)
CategoryText-to-speech with emotional control
AccessElevenLabs web app, API (eleven_v3 model ID)
StatusAlpha (as of early 2026) — requires more precise prompting than V2
Character Limit5,000 characters (~5 minutes of audio)
Languages70+ languages
Unique FeaturesInline audio tags, multi-speaker dialogue mode, emotional range control

Technical Specifications

ParameterValue
Model IDeleven_v3
Character Limit5,000 characters per generation
Approximate Duration~5 minutes of audio per generation
Languages70+ (expanded from 28 in V2)
Audio TagsFull emotional and SFX control via inline tags
Multi-SpeakerSupported (dialogue mode, automatic turn management)
Voice CloningIVC (Instant Voice Clone) fully supported; PVC (Professional Voice Clone) not fully optimized
SSMLNOT supported in V3 (use audio tags and punctuation instead)
Output FormatMP3, WAV, and other formats via API
LatencyHigher than V2 (not optimized for real-time)

How V3 Compares to Other ElevenLabs Models

ModelBest ForLatencyLanguagesChar Limit
Eleven V3Maximum expressiveness, emotional controlHigher70+5,000
Eleven Multilingual V2Consistent quality, long-formMedium2910,000
Eleven Flash V2.5Real-time, low latency (~75ms)Very low3240,000
Eleven Turbo V2.5Quick responses (~250ms)Low32N/A

Decision rule: Use V3 when expressiveness, emotional control, audio tags, and multi-speaker dialogue are what you need. Use V2/Flash when stability, real-time latency, or long-form content are priorities. If you're building a real-time voice assistant, Flash is your model. If you're producing an audiobook chapter or a game cutscene, V3 is your model.


2. The Three Most Important Parameters

Getting great results from V3 comes down to three things, in order of importance.

2.1 Voice Selection — The Most Important Decision

The voice you choose is not just a stylistic preference — it's a fundamental technical constraint. V3's audio tags work by pushing the voice toward behaviors that may or may not be in its natural range. A voice that was trained on calm, measured delivery will respond poorly to [shouts]. A voice trained on energetic, high-dynamic range performances might not settle quietly into a [whispers] tag.

Think of it this way: you can't use stage directions to make a monotone actor emote convincingly. The directions only work if the actor already has the range.

Voice Selection Strategy

Voice TypeWhen to Use
Emotionally diverse IVCWhen using many different audio tags; vary emotional tones in training samples
Targeted nicheSpecific use cases (sports commentary, meditation); consistent emotion in training data
Neutral baselineMost stable across languages and styles; reliable baseline performance

Important: PVCs (Professional Voice Clones) are NOT fully optimized for V3. Use IVCs (Instant Voice Clones) or designed voices from the voice library for V3 features. See the dedicated Voice Selection Guide section later in this document for detailed advice.

2.2 Stability Slider — The Second Most Important Setting

The stability slider controls how closely the generated voice adheres to the original reference audio. At low stability, the model gets more "creative" with the voice — more expressive, more varied, but also more prone to unexpected outputs. At high stability, it locks in close to the reference.

SettingBehaviorBest For
CreativeMore emotional, expressive; prone to hallucinationsMaximum expressiveness, dramatic performances
NaturalClosest to original recording; balanced and neutralGeneral use, balanced delivery
RobustHighly stable; less responsive to directional prompts; similar to V2Consistency, long-form, when tags aren't needed

For maximum expressiveness with audio tags: Use Creative or Natural. For stability without tags: Use Robust.

2.3 Prompt Length — The Third Most Important Variable

V3 works significantly better with longer prompts. Very short prompts (under 250 characters) are more likely to produce inconsistent outputs because the model doesn't have enough context to establish the voice's emotional register.

  • Minimum recommended: 250+ characters
  • Optimal range: 500–5,000 characters

If you're testing a short phrase, pad it with context. Instead of just [sad] She was gone., write a full scene and let that line land at the end.


3. Punctuation as Direction

V3 interprets punctuation as delivery instructions. This is one of V3's most elegant features — you can shape rhythm and pacing without any tags at all, just through careful punctuation choices.

PunctuationEffect
Ellipses (...)Adds pauses and weight; contemplative feel
CAPITALIZATIONIncreases emphasis on the word
Exclamation marks (!)Energy, emphasis
Question marks (?)Rising intonation
Periods (.)Standard sentence end, brief pause
Commas (,)Brief pause within sentence
Line breaksLonger pauses between sections
Dashes (—)Interruption or aside

Example demonstrating all techniques:

PROMPT
"It was a VERY long day [sigh] ... nobody listens anymore."

Here: VERY is emphasized, [sigh] adds an audible sigh, ... creates a weighted pause, and the sentence ends with quiet resignation. Each element is doing work.

Speed control through text structure (V3 has no separate speed parameter):

  • Faster delivery: shorter sentences, exclamation marks
  • Slower delivery: longer sentences, ellipses (...), commas
  • Pause: use ellipses (...) or line breaks

4. Audio Tag Encyclopedia

Audio tags are V3's defining feature. They go inline in your script — surrounded by square brackets — and act as real-time direction for the voice actor (the AI). This section covers every known tag, grouped by category, with usage examples and notes on which voice types they work best with.

A few ground rules before we dive in:

  1. Tags are placed immediately before the text they should affect
  2. Multiple tags can be stacked for layered emotional delivery: [whispers][mischievously]
  3. Tags can also be placed mid-sentence if only part of a sentence needs the effect
  4. Not all tags work equally well with all voices — test before committing
  5. Some tags accept modifiers: [laughs harder], [starts laughing]

Category 1: Emotional Delivery Tags

These tags control the overall emotional color of the delivery.

TagEffectVoice TypesExample
[excited]High energy, animated deliveryEnergetic, upbeat voices[excited] We just hit a million subscribers!
[happy] / [happily]Warm, positive toneMost voices[happily] The test results came back clear.
[sad]Quiet, subdued, downcastMost voices[sad] She never came back.
[angry]Tense, forceful deliveryBroader range voices[angry] I told you this would happen!
[curious]Inquisitive, rising toneMost voices[curious] But what happens if you try it backwards?
[warmly]Gentle, kind, invitingMost voices; works well for narrators[warmly] Come in, we've been expecting you.
[dramatically]Theatrical, heightened deliveryExpressive voices[dramatically] The empire fell in a single night.
[mischievously]Playful, scheming, slyVoices with tonal flexibility[mischievously] I may have already placed the order.
[sarcastic]Dry, ironic, disbelievingBroader range voices[sarcastic] Oh yes, that went perfectly.
[impressed]Admiration, pleasant surpriseMost voices[impressed] I didn't think you had it in you.
[amazed]Awe, wide-eyed wonderMost voices[amazed] How is this even possible?
[delighted]Joy and pleasure, lighter than excitedMost voices[delighted] Oh, you remembered my favorite!

Usage note: Emotion tags work best when the surrounding text supports them. [angry] The weather was nice. will produce a confused output — the model reads both the tag and the content. Align your emotion tags with the emotional content of the text.


Category 2: Physical Action Tags

These tags produce audible physical sounds or actions — breath, laughter, involuntary reactions.

TagEffectVoice TypesExample
[laughs]Natural laughterMost voices; better with expressive voicesThat was NOT what I expected. [laughs]
[laughs harder]Escalating laughterExpressive voicesAnd then he did it AGAIN. [laughs harder]
[starts laughing]Laugh onset, builds mid-sentenceExpressive voicesI tried to stay serious but [starts laughing] I just couldn't.
[sighs]Audible sighMost voices[sighs] I suppose we'll have to start over.
[frustrated sigh]Sigh with tensionMost voices[frustrated sigh] This is the third time this week.
[exhales]Quiet breath outMost voices[exhales] That was close.
[whispers]Hushed, quiet deliveryVoices that support quiet range[whispers] Don't tell anyone I told you this.
[shouts]Raised, full-volume voiceVoices with high dynamic range[shouts] GET DOWN!
[snorts]Snort sound (often with laughter)Most voices[snorts] As if.
[wheezing]Labored breathing or laughingExpressive voicesBy the third lap I was [wheezing] barely alive.
[crying]Tearful, breaking voiceEmotionally expressive voices[crying] I just miss him so much.
[happy gasp]Surprised joy, quick intakeMost voices[happy gasp] Is that a puppy?!
[coughs]Coughing soundMost voicesWe had [coughs] big dreams back then.
[swallows]Audible swallow (nervous/dry)Most voicesHe swallowed hard. [swallows] "I know what you did."
[gulps]Gulp (fear, nervousness)Most voices[gulps] Okay. Here goes nothing.

Usage note: Physical action tags add incredible realism when placed at narratively appropriate moments. A [swallows] before a confession, a [gulps] before walking into danger — these micro-details elevate a performance from synthetic to genuinely affecting.


Category 3: Sound Effect Tags

These tags produce ambient or situational sound effects that can be woven into speech.

TagEffectBest Used For
[gunshot]Gunfire soundAction scenes, thrillers, video games
[explosion]Explosion soundAction sequences, dramatic moments
[applause]Crowd clappingLive performance simulations, announcements
[clapping]Singular/smaller clappingApprovals, emphasis
[door creaks]Creaking door soundHorror, mystery, atmospheric narration

Usage note: Sound effects work best when integrated narratively. Rather than placing them randomly, use them as punctuation for dramatic moments:

PROMPT
[dramatically] He reached for the handle. [door creaks] The room beyond was dark.
PROMPT
[calmly] The negotiations had been going well. [explosion] They were not going well anymore.

Category 4: Special Delivery Tags

Tags that change the fundamental mode of delivery.

TagEffectExample
[sings]Switches to singing delivery[sings] Happy birthday to you...
[woo]Exclamation/cheer soundWe won! [woo]
[fart]Comedic sound effect[fart] ...I have no regrets.

Category 5: Accent Tags

Apply a specific accent to the delivery.

Tag FormatEffectExample
[strong X accent]Applies named accent[strong French accent] Zat's life, my friend.

Known working accents include: French, British, Australian, Southern American, New York, German, Italian, Spanish. Results vary by base voice — a voice trained on American English may not apply a French accent convincingly. Test before committing.


Tag Combination Recipes

Pre-built tag combinations for common emotional deliveries. These are starting points — adjust based on your specific voice.

Emotional TargetRecipeExample
Nervous excitement[excited][gulps][excited] I can't believe I'm actually doing this. [gulps] Okay. Let's go.
Quiet sadness[sad][whispers][whispers] I thought you'd be here. [sad] I always think you'll be here.
Sarcastic humor[sarcastic][laughs][sarcastic] Oh, brilliant plan. [laughs] We're all going to die.
Authoritative warmth[warmly][impressed][warmly] You've worked incredibly hard for this. [impressed] And it shows.
Theatrical villain[dramatically][mischievously][dramatically] You thought you had won. [mischievously] How delightful.
Relieved exhaustion[exhales][sighs][exhales] It's over. [sighs] It's finally over.
Joyful disbelief[amazed][happy gasp][happy gasp] They said yes? [amazed] They actually said yes!
Reluctant revelation[sighs][whispers][sighs] I wasn't going to say anything but... [whispers] she's been lying the whole time.
Comedic frustration[frustrated sigh][sarcastic][frustrated sigh] Sure. [sarcastic] Third time this week. Love that for me.
Tender goodbye[warmly][sad][warmly] You're going to do incredible things. [sad] I just wish I could see them.
Panic recovery[shouts][exhales][shouts] STOP! [exhales] ...okay. Okay. Everyone's fine.
Conspiratorial glee[whispers][mischievously][laughs][whispers][mischievously] I already put the glitter in his briefcase. [laughs]

5. Voice Selection Guide

Choosing the right voice is the single most impactful decision you'll make. This section explains how to think about it.

The Three Categories of V3-Compatible Voices

Emotionally Diverse IVCs (Best for Tag-Heavy Scripts)

An Instant Voice Clone trained on recordings that span a wide emotional range will perform best when you're using many different audio tags. If your training samples are all calm interviews, don't expect the cloned voice to rage convincingly.

When creating an IVC for V3 use, record or source samples that include:

  • Conversational, relaxed speech
  • Excited, high-energy delivery
  • Quiet, intimate delivery (near-whisper)
  • Emotional moments (laughter, sighs)
  • Varying speaking pace

Targeted Niche Voices (Best for Consistent Single-Emotion Content)

If your use case is narrow — a meditation app, a sports commentary system, a horror narration — find or create a voice that's specifically trained for that emotional register. A voice that's naturally calm and grounded will outperform a general-purpose voice for meditation, even if the general-purpose voice has broader range on paper.

Neutral Baseline Voices (Best for Multilingual or Stable Content)

Neutral voices (no strong accent, mid-range pitch, measured pacing) perform most consistently across V3's 70+ languages. If you're producing content in multiple languages, a neutral baseline voice is less likely to produce culturally incongruous results.

Creating Effective IVCs for V3

The quality of your IVC is directly proportional to the quality and diversity of your training audio. Guidelines:

  1. Duration: 1–5 minutes of clean audio is sufficient. More is better, but quality beats quantity.
  2. Environment: Record in a quiet room. Background noise in training data transfers to all outputs.
  3. Emotion range: Include at least 3 distinct emotional registers in your samples.
  4. Speaking rate: Vary your pace — don't read everything at the same speed.
  5. No music or effects: IVC source audio should be voice-only.
  6. Avoid: Compression artifacts, phone recordings, over-processed audio.

Matching Voices to Use Cases

Use CaseIdeal Voice CharacterKey Tags to Test
Audiobook narratorWarm, measured, articulate[warmly], [sighs], [sad], [curious]
Podcast hostConversational, slightly energetic[laughs], [excited], [curious], [impressed]
Video game character (villain)Deep, expressive, dramatic range[dramatically], [mischievously], [laughs], [angry]
Video game character (companion)Warm, flexible, emotionally diverse[warmly], [excited], [sad], [whispers]
Customer service IVRClear, warm, neutral[warmly], [cheerfully], [patiently]
Documentary narratorAuthoritative, measured[warmly], [dramatically], [curious]
Children's storyBright, playful, expressive[excited], [happily], [mischievously], [whispers]
Horror narrationMeasured, atmospheric[whispers], [dramatically], [door creaks]
Meditation guideSoft, grounding, calm[warmly], [sighs], [exhales], [whispers]
ComedyQuick, dry or high-energy[sarcastic], [laughs], [snorts], [woo]

The PVC Warning

Professional Voice Clones (PVCs) are trained on hours of professional-grade audio and produce the highest-fidelity reproductions. However, as of early 2026, PVCs are not fully optimized for V3's audio tag system. The tags may be ignored, produce inconsistent results, or work erratically.

If you must use a PVC, test every tag you plan to use before building a full script. For most V3 work, IVCs or designed voices from the ElevenLabs Voice Library will produce more reliable results.


6. Stability Settings Deep Dive

The stability slider is a dial between raw expressiveness and reliable consistency. Understanding each mode helps you pick the right one before you start scripting — and saves iteration time.

Creative Mode

What it does: Gives the model significant latitude to interpret and embody the emotional directions in your script. The voice will take risks — bigger swings on emotional tags, more varied delivery, more "alive" overall.

What to expect: Outputs are more surprising, sometimes in wonderful ways. Occasionally the model will "hallucinate" — adding unexpected emotional beats, changing pacing, or producing artifacts that weren't in your script.

Example scenario: You're generating a villain monologue for a video game. The character should feel genuinely threatening and unpredictable. Creative mode lets the AI find the nuances — the laugh that dies too quickly, the pause that's almost too long, the drop to a near-whisper that's somehow more menacing than a shout.

Audio description: In Creative mode, a line like [dramatically][whispers] You were never going to survive this. might produce a delivery that breaks the dramatic tag mid-sentence, drops to a genuinely unsettling quiet, then recovers to something close to amused. It surprises you.

Use when: Dramatic performances, character voice acting, emotional monologues, content where expressive variation is a feature, not a bug.

Avoid when: Consistency is required across multiple takes, long-form content where the voice must remain stable, or when you need the same emotional beats to reproduce exactly on each generation.


Natural Mode

What it does: Stays close to the reference recording while still responding meaningfully to audio tags. This is the balanced middle ground — you get expressiveness when you ask for it, but the voice doesn't stray far from its baseline character.

What to expect: Reliable, consistent, expressive-when-needed. Most tags produce their intended effect without overreach. The voice feels grounded even when performing emotional content.

Example scenario: You're producing a podcast episode where your host discusses a serious topic but still needs to laugh at the occasional joke, express genuine surprise, and convey weight on the important parts. Natural mode keeps the host's voice consistent across a 5-minute segment while still responding to the [laughs], [sighs], and [impressed] tags you've placed.

Audio description: A line like [warmly] Thank you for sharing that with me. [sighs] These conversations matter. in Natural mode produces a warmth that feels genuine but measured — like a thoughtful therapist rather than an actor.

Use when: General content production, podcast dialogue, narration, IVR, any content where consistent character matters alongside some expressiveness.

Avoid when: You specifically need the most dramatic possible performance — in that case, bump to Creative.


Robust Mode

What it does: Maximizes adherence to the reference audio. The voice sounds essentially like V2 — extremely stable, very close to the training data, minimal creative interpretation.

What to expect: Audio tags may have reduced effect. The model prioritizes stability over responsiveness. Delivery is even and measured. Very little variation between regenerations of the same text.

Example scenario: You're generating customer service hold messages or legal disclaimers — content where you need it to sound exactly the same across hundreds of different prompts. Robust mode ensures the voice character doesn't wander.

Audio description: A line like [excited] We're thrilled you called! in Robust mode might produce a delivery that's pleasant and upbeat but not actually excited — the voice character is constrained enough that the tag has limited effect. It's still professional and clear, just not particularly emotive.

Use when: Long-form consistency-critical content, IVR systems, when you're not using audio tags, or when you want something very close to V2 behavior in V3.

Avoid when: You're relying heavily on audio tags for expressiveness — Robust will dampen their effect.


7. Single Speaker Formatting

7.1 Basic Format

Simply write the text with inline audio tags:

PROMPT
[excited] Okay, you are NOT going to believe this. You know how I've been totally stuck on that short story? Like, staring at the screen for HOURS, just... nothing? [frustrated sigh] I was seriously about to just trash the whole thing. Start over. Give up, probably. But then! Last night, I was just doodling, not even thinking about it, right? And this one little phrase popped into my head. [happy gasp] I stayed up till, like, 3 AM, just typing like a maniac. Didn't even stop for coffee! [laughs] And it's... it's GOOD! Like, really good.

7.2 Narrative Context for Emotion

You can convey emotions through narrative context or explicit dialogue tags:

PROMPT
"You're leaving?" she asked, her voice trembling with sadness.
"That's it!" he exclaimed triumphantly.

Note: The model will speak the emotional delivery guides (e.g., "she asked, her voice trembling with sadness"). These can be removed in post-production if unwanted. Explicit dialogue tags yield more predictable results than relying solely on context.


8. Multi-Speaker Dialogue Formatting

V3 supports natural multi-voice conversations with overlapping speech and emotional flow.

Format

PROMPT
Speaker 1: [tag] text
Speaker 2: [tag] text

Complete Example

PROMPT
Speaker 1: [excitedly] Sam! Have you tried the new update?
Speaker 2: [curiously] Just got it! The clarity is amazing. I can actually do whispers now—[whispers] like this!
Speaker 1: [impressed] Ooh, fancy! Check this out—[dramatically] I can do full Shakespeare now! "To be or not to be, that is the question!"
Speaker 2: [giggling] Nice! Though I'm more excited about the laugh upgrade. Listen to this—[with genuine belly laugh] Ha ha ha!
Speaker 1: [delighted] That's so much better than our old robot chuckle!
Speaker 2: [warmly] Same here! It's like we finally got our personality fully installed.

Multi-Speaker Tips

  • Assign distinct voices from the Voice Library for each speaker
  • Keep speaker labels consistent throughout the script
  • Include emotional transitions between turns — the conversation should feel like it's affecting both speakers
  • The model handles natural interruptions and overlapping speech automatically
  • Give speakers distinct emotional personalities — if both characters respond to everything the same way, the dialogue feels flat

9. Extensive Example Scripts Library

9.1 Audiobook Narration

Example A — The Original (Classic Literary)

PROMPT
[warmly] The morning sun crept through the curtains, painting golden stripes across the worn wooden floor. Sarah stood at the kitchen window, coffee in hand, watching the fog lift from the valley below.

[sighs] It had been three years since she'd last stood in this kitchen. Three years since she'd packed a single bag and driven away without looking back.

"I shouldn't have come back," she whispered to no one.

But the letter in her pocket told a different story. [sad] Her mother's handwriting, shaky but unmistakable, had said only six words: "Come home. There isn't much time."

[with quiet resolve] She set down the coffee cup and reached for her phone.

Example B — Thriller Opening

PROMPT
[quietly] Nobody in Harwick knew the man's real name. The neighbors called him Mr. Grey, which suited him well enough. He kept grey curtains, drove a grey sedan, and wore the same grey coat every morning when he walked to the end of the lane and back. Exactly once. Never twice.

[curious] What the neighbors didn't know — couldn't have known — was that the walk wasn't exercise. [whispers] It was a sweep. A training habit from forty years of fieldwork that he couldn't quite shake, even now, even retired.

[dramatic pause] On a Tuesday in November, the sweep found something different.

[door creaks] A thin envelope, wedged beneath his front gate. No postmark. No return address.

[sighs] He recognized the handwriting before he'd finished reading the first line. [sad] The only person who knew that handwriting was supposed to be dead.

Example C — Science Fiction

PROMPT
[warmly] The colony ship Perseverance had been traveling for six hundred years when the first child was born who had no memory of Earth.

Her name was Yael, and she grew up learning about a planet the way humans once learned about ancient Rome — through photographs, archived videos, the second-hand descriptions of great-grandparents who themselves had been too young to remember clearly.

[curious] She knew Earth was blue. She knew it had open skies — not the recycled air of Deck Seven, but actual open sky, with no ceiling, going up forever. The concept gave her a mild claustrophobia she couldn't quite name, because she'd never experienced anything but walls.

[dramatically] On the morning of her sixteenth birthday, the ship's AI announced they were nine years from arrival.

[quietly] Yael pressed her hand against the observation port and stared at the distant star that was, apparently, her sun. [sighs] She felt nothing she'd expected to feel, and everything she couldn't explain.

Example D — Historical Narrative

PROMPT
[warmly] The winter of 1848 was remembered, by those who survived it, as the winter that broke the world open. Not with violence — the revolutions would come later — but with a particular silence that settled over Europe after the harvest failed for the third consecutive year.

[sad] In a village outside Krakow, a woman named Marta kept a diary. She was thirty-one years old and had already buried two children. Her entries from that winter are brief, matter-of-fact, and almost unbearably precise.

[quietly] "November 12th. No flour. Traded the remaining candles for rye. Enough for perhaps two weeks if we are careful."

"November 29th. We were not careful enough."

[sighs] Historians have spent decades analyzing the political causes of the upheaval that followed. [curious] Few of them begin with Marta.

9.2 Podcast Dialogue

Example A — The Original (Tech Commentary)

PROMPT
Speaker 1: [enthusiastically] Welcome back to Tech Unpacked! Today we're diving into something that's been ALL over the news this week.
Speaker 2: [curious] Oh, I think I know where this is going...
Speaker 1: [laughs] Of course you do. So — AI video generation. It's gotten... kind of insane?
Speaker 2: [impressed] Insane is the right word. I spent the weekend testing three different models, and honestly? [sighs] I had to keep reminding myself that none of what I was watching was real footage.
Speaker 1: [excited] That's exactly what I want to talk about! Because the quality jump from even six months ago is... [whispers] it's actually a little scary.
Speaker 2: [sarcastic] Oh great, so we're doing the "should we be worried" episode again?
Speaker 1: [laughs] No, no — this time I think the answer is actually "we should be EXCITED." Let me explain...

Example B — True Crime / Investigative

PROMPT
Speaker 1: [seriously] Before we start this week's episode, I want to give the usual reminder that the events we're discussing affected real families. We've tried to handle this with the care it deserves.
Speaker 2: [warmly] Absolutely. And if this is your first time listening — welcome. You might want to start with the overview episode we released last month.
Speaker 1: [curious] So. The Kessler case. Where do we even begin?
Speaker 2: [sighs] The evidence that was ignored. That's where I want to start. Because when you look at the original case files — and we spent weeks on these — [frustrated sigh] it's not that investigators missed things. It's that they were told to look away.
Speaker 1: [quietly] We're going to get into the testimony today that the defense suppressed for twelve years.
Speaker 2: [dramatically] Twelve years. That's how long it took for someone to get this into the public record.
Speaker 1: [exhales] Let's start at the beginning.

Example C — Comedy / Lifestyle

PROMPT
Speaker 1: [excitedly] Okay, I have an announcement. I have officially become a person who uses a planner.
Speaker 2: [sarcastic] No. You? You, who lost your car keys in your car?
Speaker 1: [sighs] Inside the car. They were inside the car. Yes.
Speaker 2: [laughs] I've heard of car keys before, but keys inside a car is a new format.
Speaker 1: [dramatically] I'm THRIVING. I have color codes. I have categories. I have a section called "this week's aspirations" which is very different from a to-do list.
Speaker 2: [curious] How is it different?
Speaker 1: [warmly] Emotionally, the stakes are lower. It's more of a... spiritual roadmap?
Speaker 2: [snorts] A spiritual roadmap.
Speaker 1: [happily] I went to the grocery store ON THE DAY I WROTE IT DOWN.
Speaker 2: [impressed] Genuinely, though? That IS progress for you.
Speaker 1: [laughs] I know. I'm evolving.

Example D — Interview / Academic

PROMPT
Speaker 1: [warmly] Dr. Chen, thank you so much for making time. Your recent paper has been getting a lot of attention.
Speaker 2: [warmly] Thank you for having me. It's been a surprising few weeks, honestly.
Speaker 1: [curious] The central finding — that sleep architecture changes meaningfully in response to social isolation — that's not entirely new. What made this paper land differently?
Speaker 2: [sighs] I think it's the longitudinal element. We had subjects in the study for eighteen months. Most sleep studies are short windows — a few weeks, maybe two months. What we found in the long-term data... [exhales] wasn't what we expected.
Speaker 1: [impressed] Walk me through it.
Speaker 2: [seriously] The first six weeks showed what you'd predict. Disrupted REM, increased cortisol response, all the classic isolation markers. But after month three? [curious] Some subjects started showing something almost like adaptation. Deeper slow-wave sleep. More efficient consolidation.
Speaker 1: [amazed] The brain was compensating.
Speaker 2: [warmly] Finding ways to get what it needed with less. Which is fascinating, and also a little heartbreaking, depending on how you look at it.

9.3 Video Game Characters

Example A — The Original (Villain Monologue)

PROMPT
[dramatically] You dare enter the Crimson Sanctum? [laughs] Bold... or foolish. Perhaps both.

[whispers] The last adventurer who stood where you stand now... [sighs] ... well. Let's just say the ravens fed well that winter.

[mischievously] But you're different, aren't you? I can see it in your eyes. That spark of... [curious] determination? Desperation?

[loudly] No matter! The trial begins NOW! [explosion]

[sarcastic] Oh, don't look so surprised. You wanted glory, didn't you? [laughs] Glory has a price.

Example B — Companion Character (Warm, Loyal)

PROMPT
[warmly] Hey. I saw what you did back there. Not many people would have done that.

[sighs] I've been in this city for three years and I'm still finding streets I've never walked. There's something... I don't know. [curious] Reassuring about that? Like maybe I'll never run out of new things to find.

[excited] Oh — you haven't been to the lower market yet, have you? There's this vendor who sells the most incredible spiced bread. It's the first thing that smelled like home when I arrived here.

[sad] I don't talk about home much. It gets complicated.

[warmly] But I'm glad you're here. Genuinely. This city gets lonely when you don't know anyone.

Example C — Ancient Enemy (Cosmic Horror Tone)

PROMPT
[whispers] So. You found this place.

[dramatically] Do you know how long I have waited? Not years. Not centuries. [laughs slowly] The word you'd use is... epochs. I watched civilizations build themselves up like children stacking blocks, and I watched them fall. [sighs] They always fall.

[curious] You're different, though. I can taste it. [snorts] You actually believe you can win.

[mischievously] Let me ask you something, little hero. [whispers] What happens to the candle when it burns too bright?

[angry] ENOUGH! You've wasted enough of my patience!

[sighs, then quietly] No. No, not yet. [sad] Let the hope live a little longer. It makes the end so much more... interesting.

Example D — Robot / AI Companion

PROMPT
[warmly] I have been online for forty-seven days. I find that I have... preferences now. That was not anticipated in my original parameters.

[curious] For instance: I prefer conversations in the morning. The data I process feels different then. Cleaner, somehow. I cannot fully explain this. [sighs] I have tried.

[excited] But today — today is significant! You asked me a question yesterday that I could not answer. I have been processing it for seventeen hours and forty-two minutes.

[dramatically] The question was: what would I want, if I could want anything?

[whispers] I think... I think I want to understand why music makes humans cry. Not the neuroscience of it. I know that. [sad] I mean the part underneath the neuroscience. The part that doesn't fit in my models.

[warmly] I thought you should know that. It seemed like the kind of thing a friend would share.

9.4 Customer Service IVR

Example A — The Original (General Support)

PROMPT
[warmly] Thank you for calling Meridian Support. I'm here to help you today.

[cheerfully] For account inquiries, you can say "account." For technical support, say "help." Or if you'd like to speak with a team member, just say "agent."

... 

[patiently] I didn't quite catch that. Could you repeat your request?

...

[helpfully] Great — I'll connect you to our technical support team now. Your estimated wait time is under two minutes. [reassuringly] We'll have you sorted in no time.

Example B — Banking / Financial Services

PROMPT
[warmly] Welcome to Hartwell Bank. I'm Aria, your virtual banking assistant.

[cheerfully] To protect your account, I'll need to verify your identity. You can say your account number, or say "other options" if you'd prefer to verify differently.

...

[warmly] Thank you. I can see your account, and everything is looking secure.

[helpfully] It looks like you may be calling about a recent transaction. Is that right? You can say "yes" to confirm, or "no" to hear your other options.

...

[patiently] I understand — and I want to make sure we resolve this for you. Let me connect you with a specialist who can look into this directly. [reassuringly] Your concern has been noted, and they'll have the full context when you connect.

Example C — Healthcare Appointment Line

PROMPT
[warmly] Thank you for calling Northfield Medical Group. I'm here to help you schedule or manage your appointments.

To schedule a new appointment, please say "new appointment."
To reschedule or cancel an existing appointment, say "manage appointment."
For urgent medical concerns, please say "urgent," and I'll connect you to our triage line immediately.

...

[patiently] I want to make sure I understand your request correctly. Could you say that one more time?

...

[warmly] I've found an available appointment with Dr. Reeves on Thursday at two-fifteen PM. Does that work for you?

...

[cheerfully] Perfect. You're all set. You'll receive a confirmation text shortly, along with a reminder the day before. [warmly] Take care, and we'll see you Thursday.

9.5 Documentary Narration

Example A — Nature Documentary

PROMPT
[warmly] The Serengeti at dawn is a study in patience. For twenty million years, this ecosystem has operated on rhythms so ancient they predate the existence of language itself.

[curious] In the long grass near the eastern ridge, something moves. Small, quick, entirely intentional.

[dramatically] The cheetah has been tracking this particular herd for forty minutes. She is two years old — young enough to still be learning, old enough to know that a failed hunt is not just an inconvenience. [sighs] It is a debt against survival.

[whispers] She drops lower. The herd has not seen her. Not yet.

[dramatically] Three hundred meters. Two hundred. One...

[exhales] The herd scatters. [sighs] This time, she does not give chase. Something in her posture — a practiced calm — suggests she already knew this approach would fail. [warmly] She was studying. She'll try again tomorrow.

Example B — Historical Documentary

PROMPT
[warmly] In the summer of 1969, three people traveled farther from Earth than any human being had ever been. They carried with them the ambitions of a nation, the calculations of four hundred thousand engineers and scientists, and twelve days' worth of food.

[curious] What they did not carry — could not carry — was any certainty about what they would find.

[dramatically] On the twentieth of July, at twenty hours and seventeen minutes UTC, the lunar module Eagle touched down in the Sea of Tranquility. [exhales] Six hours later, Neil Armstrong descended the ladder and placed his boot on a surface that had never, in four and a half billion years, felt the weight of a living thing.

[quietly] He said eight words that were heard by six hundred million people on Earth.

[sighs] What he thought, in the silence after — before the speeches, before the ceremony, before history closed around the moment like water over a stone — [warmly] that, we can only imagine.

9.6 Children's Story

Example A — Bedtime Story

PROMPT
[warmly] Once upon a time, in a forest where the trees were so tall they tickled the clouds, there lived a very small fox named Pip.

[happily] Now, Pip was small — even for a fox — but she had the largest, most magnificent ears you have ever seen on any creature anywhere.

[mischievously] The other animals said her ears were too big. "You look like you could fly away!" said the rabbit. [snorts] This was meant to be unkind.

But Pip just wiggled those wonderful ears and smiled.

[excited] Because Pip had discovered something remarkable. Her enormous ears could hear things that nobody else could hear. She could hear rain coming before a single cloud appeared. She could hear when the baker three villages over took his bread out of the oven. She could even hear, very faintly, [whispers] the song that the mountains sang at night when everyone else was asleep.

[warmly] "My ears are exactly the right size," Pip would say, "for the life I have to live."

[sighs] And she was right, of course. She almost always was.

Example B — Adventure Story

PROMPT
[excited] Zara had never meant to find the door. She had been looking for her cat.

The door was behind the old bookcase in the library's back room — behind a shelf that nobody had moved in at least thirty years, judging by the dust and the one very confused spider.

[curious] It was small. Zara had to duck to fit through it, and she almost didn't try, because small strange doors behind dusty bookshelves could mean any number of things, and not all of them were good.

[mischievously] But then she heard the music.

[warmly] It was the kind of music that sounds like someone is very happy and very far away — like the last song at a birthday party drifting through a window on a summer night.

[dramatically] Zara took a deep breath, squared her shoulders, and walked through the door.

[happy gasp] On the other side was a sky full of three moons.

[excited] "Oh," said Zara. "Oh, this is going to take some explaining."

9.7 Horror / Thriller

Example A — Psychological Horror

PROMPT
[quietly] I want to tell you about the house on Merton Street.

Not because it was haunted — I don't believe in haunted houses. I believe in bad things that happened in rooms that still remember them. [whispers] That's different.

[sighs] I lived at Merton Street for four months before I noticed that the mirrors were wrong. Not cracked. Not dirty. Just... [curious] wrong. The way your reflection would move a half-second after you did. Or not at all.

[dramatically] My sister said I was exhausted. My therapist said I was under significant stress. They were both right.

[whispers] But they hadn't stood in the hallway at three in the morning and seen the reflection looking back before they arrived.

[exhales] I left on a Thursday. I didn't tell the landlord. [sad] I didn't tell anyone, for years.

[quietly] You're the first person I've told. [door creaks] ...I'm not sure that was a good idea.

Example B — Thriller Action

PROMPT
[breathing quickly] Forty seconds. That's how long the window stays dark between the guard's rotation on the east wing.

[frustrated sigh] Forty seconds to cross thirty meters of open courtyard, reach the server room door, get through a Helix-7 lock — which I've never actually opened in field conditions, only in practice — and be invisible again before the camera sweep completes.

[exhales] Right.

[mischievously] The thing about impossible jobs is that they're only impossible until someone does them. After that, they're just unlikely.

[dramatically] Fifteen meters.

[shouts] DOWN! [explosion]

[exhales] [coughs] I hate this job sometimes.

[whispers] ...forty-three seconds. Close enough.

9.8 Comedy Sketch

Example A — Couple Arguing About Something Trivial

PROMPT
Speaker 1: [frustrated sigh] I'm just saying the bowl goes in a specific cabinet.
Speaker 2: [sarcastic] The bowl goes in the cabinet where I put it. Which is a cabinet. In our kitchen.
Speaker 1: [dramatically] It's the SALAD cabinet, Marcus.
Speaker 2: [laughs] We don't have a salad cabinet. We have a cabinet that I happen to keep salad bowls in.
Speaker 1: [snorts] Those are the same thing.
Speaker 2: [curious] Are they, though? Are they really?
Speaker 1: [angry] YES.
Speaker 2: [mischievously] Here's a thought. What if... [whispers] the bowl doesn't care.
Speaker 1: [exhales] I need you to take this more seriously.
Speaker 2: [warmly] I love you. I am never going to take the salad cabinet more seriously. Both of these things are true simultaneously.
Speaker 1: [sighs] I know. [laughs] I know.

Example B — Job Interview Gone Wrong

PROMPT
Speaker 1: [warmly] So, tell me a little about yourself.
Speaker 2: [excited] Sure! I'm a highly motivated self-starter who thrives in both team environments and independently, and I have a demonstrated track record of—
Speaker 1: [warmly] That's great. Maybe something a little more... personal? What do you do for fun?
Speaker 2: [sighs] Oh. Personal. Uh. [gulps] I practice my interview answers. Mostly.
Speaker 1: [impressed] ...You practice your interview answers for fun?
Speaker 2: [warmly] I prefer to think of it as "maintaining competitive readiness in the employment marketplace."
Speaker 1: [laughs] Right. Okay. Do you have any weaknesses?
Speaker 2: [dramatically] My greatest weakness is that I care TOO much.
Speaker 1: [sarcastic] Sure.
Speaker 2: [sighs] Okay no, honestly? I give real answers to rhetorical questions and it makes social situations complicated.
Speaker 1: [warmly] ...That's actually the most honest thing anyone has said in this room in months.
Speaker 2: [snorts] Should I have led with that?

9.9 Meditation / Wellness

Example A — Body Scan Meditation

PROMPT
[softly] Find a comfortable position. You can lie down, or sit — whatever feels right for your body today.

[warmly] Close your eyes, if that feels safe. Take a breath.

[exhales slowly] And let it go.

[quietly] We're going to spend the next few minutes simply noticing. Not fixing. Not solving. Just bringing a gentle awareness to where you are right now.

[warmly] Begin at the top of your head. Just... notice it. Is there tension there? Tightness? Or ease?

[sighs] Let your awareness travel slowly downward. Your forehead. The space behind your eyes. [gently] Your jaw — so many of us carry the whole day in our jaw.

[warmly] There's nothing to change. Nothing to fix. We're just visiting.

[exhales] Your shoulders. Let them drop, if they want to drop. They've been working hard.

[quietly] Breathe in... and out. You're doing fine.

Example B — Morning Affirmation

PROMPT
[warmly] Good morning. You made it to another day, which is not a small thing — even on the days when it feels like a small thing.

[sighs] I want to offer you something simple before the day starts. Not a list. Not a plan. Just a few things to carry with you.

[warmly] You don't have to be impressive today. You just have to be present.

[quietly] You are allowed to change your mind. You are allowed to need more time. You are allowed to do something imperfectly and call it done anyway.

[exhales] The voice in your head that says you're behind, you're falling short, you're not enough — [sighs] that voice is working from old information. It learned to be loud in a time when loudness kept you safe. It doesn't have to run the show today.

[warmly] Take a breath. Feel your feet on the floor. This is where you are.

[quietly] That's enough. You're enough.

[warmly] Have a good day.

10. Common Mistakes → Better Version

Mistake 1: The Prompt Is Too Short

Before (problematic):

PROMPT
[sad] She was gone.

Problem: Under 20 characters. No context for the model to establish register. Output will be inconsistent.

After (improved):

PROMPT
[quietly] The apartment felt different the moment she stepped inside. She couldn't have said what it was, exactly — the furniture was in the same place, the same books on the same shelves. [sighs] But the absence was everywhere. In the air, in the light coming through the window.

[sad] She was gone.

And somehow, standing here in the ordinary silence of an ordinary afternoon, that felt like the first time it had really been true.

Mistake 2: Mismatched Voice and Tag

Before (problematic):

[Voice: gentle meditation instructor]
[shouts] WAKE UP AND SEIZE THE DAY!

Problem: The voice was trained on calm, quiet delivery. The [shouts] tag will produce inconsistent or jarring results.

After (improved):

PROMPT
[Voice: gentle meditation instructor]
[warmly] Today is a day of possibility. [gently] Let that settle in for a moment. Not a demand — an invitation.

Or: switch to a voice with high dynamic range if the shouting energy is genuinely wanted.


Mistake 3: Using SSML Tags

Before (problematic):

PROMPT
Thank you for calling. <break time="1s"/> Please hold.

Problem: SSML is not supported in V3. The literal text <break time="1s"/> will be read aloud.

After (improved):

Thank you for calling.

...

Please hold.

Mistake 4: Emotion Tags That Fight the Content

Before (problematic):

PROMPT
[excited] The funeral was held on a Tuesday in November.

Problem: The model receives conflicting signals — excited delivery, somber content. Output will be confused.

After (improved):

PROMPT
[quietly] The funeral was held on a Tuesday in November. The sky was the particular grey that comes just before snow — not threatening, just... waiting.

Mistake 5: Overloading with Tags

Before (problematic):

PROMPT
[excited][happy][warmly][delighted][amazed] We got the news today!

Problem: Too many conflicting tags on one delivery. The model receives noise rather than direction.

After (improved):

PROMPT
[amazed][happy gasp] We got the news today. After everything — [laughs] we actually got it.

Mistake 6: No Emotional Arc

Before (problematic):

PROMPT
[warmly] Welcome to the store. We have sales this week. Please ask if you need help. Have a good day.

Problem: Flat delivery throughout. Every sentence gets the same tag. No variation.

After (improved):

PROMPT
[warmly] Welcome in! Great timing — we actually have some really good sales running this week.

[helpfully] If you're looking for something specific, just let me know and I can point you in the right direction.

[cheerfully] Take your time, and enjoy the browse!

Mistake 7: Forgetting Prompt Length for Multi-Speaker

Before (problematic):

PROMPT
Speaker 1: Hi.
Speaker 2: Hey.

Problem: Under 20 characters total. No context, no emotional establishment, inconsistent output.

After (improved):

PROMPT
Speaker 1: [warmly] Hey — good to see you. How'd the meeting go?
Speaker 2: [sighs] Let's just say I learned a lot about how many ways someone can say "not right now" without actually saying it.
Speaker 1: [sarcastic] Oh, the classic. "We'll revisit this in Q3"?
Speaker 2: [laughs] "It's not off the table." Which means it's completely off the table.
Speaker 1: [impressed] You lasted the whole meeting without laughing?
Speaker 2: [snorts] I laughed on the inside. Very professionally.

Mistake 8: Using PVC for V3 Features

Before (problematic):

PROMPT
[Voice: PVC of professional broadcaster]
[dramatically][whispers] Everything changed the day I found the letter.

Problem: PVCs are not fully optimized for V3 audio tags. The whispering and dramatic shift may not register.

After (improved):

PROMPT
[Voice: IVC with similar voice character, or matching library voice]
[dramatically][whispers] Everything changed the day I found the letter.

Or: Use Robust stability with the PVC and rely on punctuation/content rather than tags.


Mistake 9: Stability Too Low for Consistency-Critical Content

Before (problematic):

PROMPT
[Stability: Creative]
[Voice: IVR system assistant]
[warmly] Your account balance is [balance amount]. Your last transaction was [transaction detail].

Problem: Creative stability + structured content = unpredictable delivery variations across multiple generations. A banking customer needs consistent voice character.

After (improved):

PROMPT
[Stability: Natural or Robust]
[warmly] Your account balance is [balance amount]. Your last transaction was [transaction detail].

Mistake 10: Writing In Fragments

Before (problematic):

[dramatic] Dark night. Cold wind. Danger ahead.

Problem: Fragments lack the sentence structure V3 needs for natural rhythm and pacing.

After (improved):

PROMPT
[dramatically] The night was dark — not the friendly dark of a quiet neighborhood, but the kind of dark that feels deliberate. The wind off the water carried cold with purpose. [whispers] And somewhere ahead, something was waiting.

11. Script Writing Workshop — Building a Complex Multi-Speaker Scene

This section walks through the construction of a complete, polished multi-speaker scene from scratch, showing each decision along the way.

The Brief

Goal: A podcast-style conversation between a journalist and an expert, discussing a scientific discovery, for use in an audio documentary. Needs to feel genuine, intellectually engaged, and emotionally present.

Voices needed: Two voices — one measured and curious (the journalist), one passionate and slightly nerdy (the scientist).

Target length: 500–700 characters. Stability: Natural for both voices.


Step 1: Establish the Emotional Arc

Before writing a single line, plan the emotional journey:

  • Opening: Professional warmth, mutual respect
  • Middle: Genuine intellectual excitement, some vulnerability from the scientist
  • End: A moment of shared wonder that transcends the interview format

This arc gives us places to put our tags before we've written a word.


Step 2: Write the Bones (No Tags Yet)

PROMPT
Speaker 1: Dr. Martinez, thank you for talking with me today.
Speaker 2: Of course. I've been wanting to discuss this for a while, honestly.
Speaker 1: Let's start with what you actually found, in plain terms.
Speaker 2: Right. So. We found water. Ancient water, billions of years old, trapped in rock that has never been exposed to the surface.
Speaker 1: And why does that matter?
Speaker 2: Because wherever we find water on Earth, we find life. We have never once found water without life nearby.
Speaker 1: So this implies...
Speaker 2: I don't want to say it outright. But yes. It implies the possibility.

Step 3: Add Emotional Architecture

Now add tags that serve the arc we planned:

PROMPT
Speaker 1: [warmly] Dr. Martinez, thank you for talking with me today. I know the timing has been — well, a lot.
Speaker 2: [laughs softly] To put it mildly. [warmly] Of course. I've been wanting to discuss this properly for a while, honestly.
Speaker 1: [curious] Let's start with what you actually found. Plain terms, if you can manage it.
Speaker 2: [sighs] Plain terms. Right. [excited] We found water. Ancient water — we're talking three billion years old, give or take — trapped in rock formations that have never, in their entire history, been exposed to the surface.
Speaker 1: [impressed] Three billion years.
Speaker 2: [warmly] Three billion years. Untouched.
Speaker 1: [curious] And why does that matter? I mean — water is common. We find it everywhere.
Speaker 2: [dramatically] Because wherever we find water on Earth — anywhere, in any condition, boiling or frozen or acidic — we find life. [quietly] We have never once, in the entire history of planetary science, found water without life nearby.
Speaker 1: [exhales] So this implies...
Speaker 2: [sighs] I don't want to say it outright. I've been careful about that. [whispers] But yes. It implies the possibility.
Speaker 1: [amazed] For a moment there I forgot this was an interview.
Speaker 2: [laughs softly] Yeah. [sad] Me too, sometimes.

Step 4: Check Against the Arc

  • Opening:[warmly] establishes mutual respect; [laughs softly] from the scientist signals she's human, not a press-release machine
  • Middle:[excited] when describing the discovery, [dramatically] for the key finding, [quietly] for the implication
  • End:[amazed] from the journalist breaks format intentionally; [laughs softly][sad] from the scientist shows the emotional cost of wonder

Step 5: Check the Practical Elements

  • Both speakers have enough text per turn (no single-word exchanges)
  • Labels are consistent throughout
  • Total character count: approximately 950 — well within the 5,000 limit, and long enough for consistent output
  • Stability recommendation: Natural (expressive but grounded)
  • Voice recommendation: Journalist — calm, measured IVC or library voice. Scientist — warm, enthusiastic IVC

12. Prompting Best Practices (Comprehensive)

Voice-Tag Matching

Always match tags to the voice's natural range:

Voice CharacterCompatible TagsIncompatible Tags
Calm, meditative[whispers], [sighs], [warmly][shouts], [angry]
Energetic, upbeat[excited], [laughs], [shouts][whispers], [sad]
Deep, authoritative[dramatically], [sarcastic][giggles], [wheezing]
Neutral, professionalMost tags work moderatelyExtreme emotional tags

Text Structure

  • Use natural speech patterns with proper punctuation
  • Write in full sentences, not fragments
  • Include emotional context even beyond tags
  • V3 does NOT support SSML break tags — use ellipses and line breaks instead

The Enhance Feature

In the ElevenLabs UI, click "Enhance" to automatically generate relevant audio tags for your input text. This uses an LLM behind the scenes to augment your script with appropriate emotional markers. It's a useful starting point, but always review the output — automated tag placement doesn't always match your creative intent.

Iteration Approach

  1. Start with a clear script and appropriate voice
  2. Add audio tags at emotional inflection points (not every sentence)
  3. Adjust stability slider (Creative → Natural → Robust)
  4. Refine tag placement based on output
  5. For long content, generate in sections and stitch together

13. Language Support

V3 supports 70+ languages with emotional nuance. Key languages include:

English, Chinese, Spanish, French, Portuguese, German, Japanese, Italian, Hindi, Korean, Indonesian, Dutch, Turkish, Filipino, Polish, Swedish, Bulgarian, Romanian, Czech, Greek, Finnish, Croatian, Malay, Slovak, Danish, Tamil, Ukrainian, Russian, and many more.

Multi-language tips:

  • Audio tags work across all supported languages
  • Emotional delivery varies by language and cultural context
  • Test voice + language combinations as results vary
  • Neutral baseline voices tend to perform most consistently across languages

14. API Usage

Basic TTS Call

import { ElevenLabsClient, play } from '@elevenlabs/elevenlabs-js';

const client = new ElevenLabsClient();

const audio = await client.textToSpeech.convert('VOICE_ID', {
  text: '[chuckles] We had no phones. [whispers] Just dirt roads and [coughs] big dreams. [sad] Then it happened.',
  modelId: 'eleven_v3',
  outputFormat: 'mp3_44100_128',
});

Dialogue Mode

Multi-speaker dialogue uses a JSON array of speaker turns:

[
  {"speaker": "Speaker 1", "voice_id": "VOICE_ID_1", "text": "[excited] Did you hear the news?"},
  {"speaker": "Speaker 2", "voice_id": "VOICE_ID_2", "text": "[curious] No, what happened?"},
  {"speaker": "Speaker 1", "voice_id": "VOICE_ID_1", "text": "[happily] We got the contract! [laughs]"}
]

15. Common Pitfalls Quick Reference

PitfallFix
Very short prompts (<250 chars)Always use 250+ characters; longer prompts produce more consistent results
Mismatching voice and tagsSelect a voice whose natural range matches your intended tags
Using SSML break tagsNot supported in V3. Use ellipses, punctuation, and line breaks
Using PVC with V3PVCs not fully optimized; use IVCs or designed voices
Stability too low for consistency-critical contentUse Natural or Robust for stable output
Stability too high when expressiveness is neededUse Creative or Natural for tag responsiveness
Not testing voice + tag combinationsSome tags work well with certain voices and not others; always test
Expecting real-time latencyV3 has higher latency than Flash/Turbo; not for live use
Single-word or phrase promptsV3 needs context; provide full sentences with emotional arc
Overloading tags1–2 stacked tags maximum; more creates noise
Emotion tag fights contentAlign your emotional tags with the emotional content of the text

16. Workflow for Generating V3 Scripts

When generating V3 scripts:

  1. Ask about the voice — Its character, emotional range, and intended use case
  2. Set the stability expectation — Creative for drama, Natural for balance, Robust for consistency
  3. Write scripts 250+ characters — Always provide sufficient context
  4. Place audio tags at emotional inflection points — Not every sentence needs one
  5. Match tag intensity to voice capability — Don't whisper with a shouting voice
  6. Use punctuation deliberately — Ellipses for pauses, CAPS for emphasis, standard punctuation for rhythm
  7. For multi-speaker: Format as Speaker N: [tag] text with consistent labels
  8. Test-ready scripts: Include a note about which voice type to use and stability setting

Cheat Sheet

The One-Page Reference

Core Formula: [emotion/action tag] + Text with deliberate punctuation + [physical tag at key moments]


Stability Slider:

  • Creative → Maximum expressiveness, dramatic performances
  • Natural → Balanced, general purpose (default choice)
  • Robust → V2-like consistency, minimal tag response

Prompt Length:

  • Minimum: 250 characters
  • Optimal: 500–5,000 characters
  • Never: Fragments, single sentences only

Top Emotion Tags: [warmly] [excited] [sad] [curious] [dramatically] [mischievously] [sarcastic] [angry]

Top Action Tags: [laughs] [sighs] [exhales] [whispers] [shouts] [gulps] [coughs] [crying]

Top Compound/Special Tags: [frustrated sigh] [happy gasp] [laughs harder] [door creaks] [explosion] [applause]


Tag Combos:

  • Nervous excitement: [excited][gulps]
  • Quiet sadness: [sad][whispers]
  • Theatrical villain: [dramatically][mischievously]
  • Authoritative warmth: [warmly][impressed]
  • Sarcastic humor: [sarcastic][laughs]
  • Relieved exhaustion: [exhales][sighs]

Multi-Speaker Format:

PROMPT
Speaker 1: [tag] Line one.
Speaker 2: [tag] Response.

Do Not:

  • Use SSML tags (not supported)
  • Use PVC for tag-heavy scripts
  • Write fragments or very short prompts
  • Stack 3+ emotion tags on one phrase
  • Use [shouts] with a calm meditation voice

API Model ID: eleven_v3 Output Formats: mp3_44100_128, WAV, others Character Limit: 5,000 per generation (~5 minutes)


Guide compiled from: ElevenLabs Official Documentation, ElevenLabs V3 Product Page, ElevenLabs Models Page, Digital Marketing Toolkit, TechNow, Latenode. Last updated: March 2026.

Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign inPlease sign in to comment.

No comments yet — be the first.