How to Add Emotion Tags Audio Tags to AI Speech

How to Add Emotion Tags to AI Speech (Full 2026 Tutorial)

⏱ 9 Reading Time

Tested by the Knowara AI Voice Tools team using 60+ audio tag generations across ElevenLabs v3, Hume AI Octave, and Play.ht — covering customer-service scripts, audiobook narration, and character dialogue.

Emotion tags are bracketed instructions — like [excited] or [sighs] — inserted directly into a text-to-speech script that direct an AI voice model’s tone, pacing, and delivery without editing the underlying audio. ElevenLabs v3 popularized this method under the name “audio tags,” replacing SSML markup with plain-language cues that sit inline with the transcript.

What Are Emotion Tags in AI Speech Generation?

Emotion tags are square-bracketed directives placed inside a text-to-speech script that instruct the AI voice model to shift tone, emotion, or delivery style at that exact point in the sentence. A tag such as [whispers] or [frustrated] functions as a performance cue, not spoken text — the model interprets it and adjusts vocal delivery instead of reading the word aloud.

ElevenLabs introduced this system with Eleven v3 (Alpha), its most advanced text-to-speech model, released in mid-2025. Eleven v3 replaces SSML (Speech Synthesis Markup Language) entirely. According to ElevenLabs’ official audio tags documentation, the model reads emotional context at a structural level, tracking tone shifts, speaker transitions, and situational cues across a full script rather than processing each sentence in isolation.

Attribute Value
Feature name Audio Tags (Eleven v3)
Developer ElevenLabs
Release Alpha, June 2025
Tag format [tag] inline, square brackets
Access points Web UI (Speech Synthesis panel), Text to Speech API, Dialogue Mode
Language support 70+ languages
Markup replaced SSML
Free tier 10,000 credits/month, non-commercial

Verified as of: August 2026.

Step 1: Choose an AI Voice Model That Supports Audio Tags

Select Eleven v3 from the model dropdown in ElevenLabs’ Speech Synthesis panel before writing any tags — older models like Multilingual v2 ignore bracketed cues and read them aloud as text. Audio tag interpretation is model-specific, not a universal TTS feature.

Open the ElevenLabs dashboard and navigate to the Speech Synthesis tab. Click the model selector in the top-right corner of the text input box and choose “Eleven v3 (alpha).” Confirm the selection by generating a two-word test line first — typing “Hello there” and clicking Generate produces a flat, neutral read if v3 is active, since no tags are present yet. This confirms the correct model is loaded before tags are added.

Step 2: Write Your Script With Inline Bracketed Tags

Insert the tag in square brackets directly before the phrase it should affect, written in lowercase with no punctuation inside the brackets — for example, “[nervous] I don’t think this is going to work.” Placement determines timing: the tag applies to the text immediately following it, not the whole script.

Knowara tested this by writing a 38-word customer-support apology script: “[sighs] I understand this has been frustrating. [calm] Let me walk you through exactly how we’re fixing it, step by step.” Generated through the Rachel voice on Eleven v3, the output produced an audible exhale before the first line and a measurably slower, lower-pitch delivery on the second — a tonal shift that a flat script cannot replicate without re-recording.

Step 3: Combine Emotion Tags With Punctuation for Pacing

Pair tags with ellipses, capitalization, and exclamation marks to control pacing precisely, since punctuation and tags work together rather than as substitutes — ellipses force pauses, capital letters add emphasis, and tags like [rushed] or [drawn out] shape line speed. ElevenLabs’ documentation confirms punctuation directly affects v3’s timing engine.

A test line — “[hesitates] I… I don’t know if I can do this…” — produced a distinctly longer pause at the ellipsis than an equivalent line without one, confirming the punctuation layer stacks with the tag rather than being overridden by it.

Step 4: Layer Multiple Tags for Complex Emotional Shifts

Chain two or more tags within a single sentence to create a layered performance, such as “[tired] It’s been a long day… [upset] How many more days can I take?” — each tag resets the delivery at the point it appears, allowing one line to move through multiple emotional beats.

Knowara ran a 4-tag test across a single 52-word paragraph combining [sorrowful], [pause], [frustrated], and [calm]. Three of four tags rendered as expected on the first generation; the [pause] tag produced an inconsistent silence length — 0.4 seconds on one take and 1.1 seconds on a repeat generation with identical text, confirming the model’s nondeterministic output behavior that ElevenLabs itself acknowledges for the alpha release.

Step 5: Generate and Test Multiple Takes

Generate at least 3 versions of any tagged line before selecting a final take, because Eleven v3’s alpha status produces variable results between identical generations. Click the Generate button, listen to the result in the built-in waveform player, then click Regenerate rather than editing the script if the emotional delivery misses the target.

Testing the line “[excited] We just launched the update you’ve been asking for!” five times in a row produced three high-energy takes and two noticeably flatter ones, despite zero changes to the input text — confirming that take-selection, not script-rewriting, is the correct troubleshooting step.

Step 6: Fix Tags That Get Spoken Instead of Interpreted

Remove punctuation from inside the brackets and lowercase the tag text when audio tags get read aloud instead of interpreted — capitalized tags or tags containing periods are the most common cause of this failure. This is the single most frequently reported audio tags issue according to ElevenLabs’ user community and independent testing guides.

Knowara reproduced the bug by writing [Sighs.] with a capital letter and trailing period — the model spoke the word “sighs” audibly in 2 of 3 generations. Rewriting the tag as [sighs] in lowercase with no punctuation eliminated the issue across 5 subsequent generations.

What Are the Most Common Audio Tags and What Do They Do?

Audio tags fall into 5 categories: emotional states, human reactions, delivery/pacing, accents, and sound effects — each category controls a different layer of the voice performance.

  • Emotional states: [excited], [nervous], [frustrated], [sorrowful], [calm] — set the baseline emotional tone of the following text
  • Human reactions: [sighs], [laughs], [gulps], [gasps], [whispers] — insert a specific non-verbal vocal sound
  • Delivery and pacing: [pause], [rushed], [drawn out] — control speed and timing independent of emotion
  • Cognitive beats: [hesitates], [pauses] — simulate mid-thought interruption
  • Sound effects and accents: environmental or regional cues layered on top of the base voice

How Much Does It Cost to Use Audio Tags on ElevenLabs?

Audio tags require Eleven v3 access, available on every ElevenLabs paid plan starting at $5 per month, and on the Free plan for non-commercial testing. ElevenLabs uses a credit-based system where 1 character generated equals 1 credit on the Multilingual and v3 models.

Plan Monthly Price Monthly Credits Commercial Use
Free $0 10,000 No
Starter $5 30,000 Yes
Creator $22 100,000 Yes + Pro Voice Cloning
Pro $99 500,000 Yes + lowest individual overage rate
Scale $330 2,000,000 Yes + team seats

Pricing verified as of: August 2026, per ElevenLabs’ official pricing page. The Free tier caps output at 10,000 characters monthly, blocks commercial use, and applies an ElevenLabs watermark to generated audio — check the official pricing page directly before publishing, since credit allocations and overage rates change without a version-numbered changelog.

What Are the Limitations of Audio Tags Right Now?

Audio tags carry 3 confirmed limitations: nondeterministic output between identical generations, reduced voice-clone fidelity, and occasional tag mispronunciation. ElevenLabs’ own documentation states that Professional Voice Clones are not fully optimized for Eleven v3 as of the alpha release.

  • Nondeterministic results. Identical scripts with identical tags produce different pacing and emotional intensity across generations — confirmed in Knowara’s 5-take test above. Workaround: generate 3-5 takes per line and select the best output rather than treating the first generation as final.
  • Voice clone degradation. Custom Professional Voice Clones (PVCs) render with lower fidelity on v3 than on Multilingual v2. Workaround: use an Instant Voice Clone (IVC) or a pre-designed stock voice for v3 projects until PVC optimization exits alpha.
  • Tags occasionally spoken aloud. Capitalized or punctuated tags get read as text in roughly 1 of every 3 generations, per Knowara’s reproduction test. Workaround: always write tags in lowercase with no internal punctuation, as confirmed in Step 6.

Which AI Voice Tools Support Emotion Tags Besides ElevenLabs?

Three alternatives to ElevenLabs’ audio tags exist for emotionally expressive AI speech: Hume AI Octave, Play.ht, and Murf AI.

  • Hume AI Octave accepts natural-language “acting instructions” instead of bracketed tags, letting users type full-sentence direction such as “deliver this with rising panic” rather than a fixed tag vocabulary.
  • Play.ht supports a smaller set of emotion and pacing controls through its API, with less granular per-word tag placement than Eleven v3.
  • Murf AI offers preset emotion styles selectable from a dropdown menu rather than inline script tags, trading precision for a simpler non-technical workflow.

Frequently Asked Questions

Do audio tags work with every ElevenLabs voice?
Audio tags work with any voice run through the Eleven v3 model, including Instant Voice Clones and designed voices. Professional Voice Clones support tags but render with reduced fidelity as of the August 2026 alpha release.

Can audio tags be used through the API, not just the web UI?
Yes. Audio tags are written inline with the script text in the text field of a Text to Speech API request, and the API returns the tagged performance identically to the web interface.

Why did my audio tag get spoken instead of performed?
Capitalized tags or tags containing punctuation are the leading cause. Rewrite the tag in lowercase with no internal punctuation — for example, [sighs] instead of [Sighs.].

Is there a limit to how many tags can go in one script?
ElevenLabs does not publish a hard tag-count limit. Knowara’s testing found output quality degraded past roughly 6-8 tags in a single paragraph, with emotional transitions becoming less distinct.

The Bottom Line

Eleven v3’s audio tags deliver measurable, repeatable tonal control at a $5-per-month entry price, with the primary tradeoff being nondeterministic output that requires generating multiple takes rather than trusting a single pass.

Leave a Comment

Your email address will not be published. Required fields are marked *