Skip to main content
Audio tags are inline tokens you put in the input text to make speech expressive: laughs, sighs, breaths, pauses, and emotional delivery. No SSML, no extra request fields — the tags travel inside the text and work the same over the HTTP endpoint, the streaming WebSocket, and the playground. There are two tag dialects, and the voice decides which one applies: In the playground, type < on an Atlas voice or [ on Ember to get a picker of the supported tags. Multilingual voices get no picker.
Tags count toward your character quota like any other text. A tagged line is metered on the full input string, tokens included.

Atlas voices — pipe tags

Atlas voices understand a fixed vocabulary of <|…|> tokens. There are two kinds: event tags that insert one sound, and style wraps that color a whole sentence.

Event tags — a single transient sound

An event tag is a bare token followed by a space. It inserts one sound (a laugh, a sigh, a cough…) at that point in the speech.

Supported events

<|uh|> and <|um|> are the most strongly trained events and the cheapest way to make a line sound spontaneous rather than read. Put one where a speaker would genuinely stall — before the word being reached for, not at the start of the sentence — and use at most one or two per sentence:

Placement

Where each event sounds most natural:
  • Lead the sentence: laugh, chuckle, giggle, laugh_harder, cry, hum_tune, tutu_tune, woo, yawn, throat_clear, exhale, sigh, inhale, cough, sniff.
  • Between clauses, mid-sentence: uh, um, breath, long_pause, gulp, snort, lip_smack.
  • Either slot: pause, gasp, sharp_breath.

Write tags bare

The token, a space, then your text — nothing else. Don’t spell the sound out next to the tag (<|laugh|> Hahaha!): the tag already produces it, and the word gets spoken on top.
<|snort|> is a derisive snort (“pfft”), not a nasal sniff — use <|sniff|> for nasal. <|sharp_breath|> is a quick intake of breath, distinct from the slower <|inhale|>.

Style wraps — color a whole sentence

A style wrap sets the emotional delivery of one full sentence. The shape is always <|style_open|>NAME<|style_body|> … <|style_close|>:
To change style, close the current wrap and open a new one back-to-back. Event tags can sit inside a wrap:

Supported styles

Emotions:
Accents:
If the feeling you want isn’t listed (serious, monotone, fast, slow…), leave the sentence untagged — neutral is the correct default. Nearest matches: warm/gentle/soft → tender, happy → joyful, quietly/whisper → whispers, interested → curious, concerned/tense → nervous, impressed → awe, terrified → fearful.

Rules of thumb

  1. Use the exact tokens — no spaces inside <|…|>, and only the tags listed here. Anything else (<|singing|>, <|music|>, <|clap|> …) has no token in the model and is either spoken aloud or dropped.
  2. At most one style per sentence, and a style never carries past <|style_close|>.
  3. Tags describe the voice only — no music, sound effects, or physical actions.
  4. Add emphasis with a CAPITAL word, ? / !, or … ellipses rather than piling on tags.
  5. Pick tags that match the line’s emotion; a contradicting tag sounds worse than none.

Full example

Input text:
My uncle thought the robot vacuum was a cat and fed it milk. I couldn’t stop laughing. Then I got worried it might break.
Tagged:
As an API request:

Ember — bracket tags

The ember voice runs on a different model and takes free-form cues in square brackets anywhere in the text. There is no fixed list: any short description of a sound or a delivery works, so write what you mean.
Cues the model responds to well: Keep cues short (a few words), place them right before the speech they should affect, and don’t put the words you want spoken inside the brackets — bracketed text is a direction, not speech.
Don’t mix dialects. <|laugh|> on Ember and [chuckling] on an Atlas voice are both treated as plain text and may be read aloud.

Streaming

Tags work unchanged in text frames on the streaming endpoint. When the gateway splits a frame into sentences it never cuts inside a <|…|> token, a style wrap, or a [bracket] cue. Keep each tag inside a single text frame — the two halves of a token sent in separate frames are plain text.