> ## Documentation Index
> Fetch the complete documentation index at: https://docs.tts.runatlas.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Audio tags

> Inline tokens in the input text that add laughs, sighs, pauses, and emotional styles to generated speech.

Audio tags are inline tokens you put in the `input` text to make speech expressive: laughs,
sighs, breaths, pauses, and emotional delivery. No SSML, no extra request fields — the tags
travel inside the text and work the same over the HTTP endpoint, the streaming WebSocket,
and the playground.

There are two tag dialects, and the **voice decides which one applies**:

| Voices | Dialect | Shape | Vocabulary |
| - | - | - | - |
| Atlas voices (`lylla` and the rest of the standard catalogue, plus your voice clones) | pipe | `<\|laugh\|>` | Fixed — only the tokens listed on this page render |
| `ember` | bracket | `[chuckling]` | Free-form — any short description in square brackets |
| Multilingual voices (`indic-*`) | — | — | **Not supported** — leave tags out of the `input` |

In the playground, type `<` on an Atlas voice or `[` on Ember to get a picker of the
supported tags. Multilingual voices get no picker.

<Note>
  Tags count toward your character quota like any other text. A tagged line is metered on
  the full `input` string, tokens included.
</Note>

## Atlas voices — pipe tags

Atlas voices understand a fixed vocabulary of `<|…|>` tokens. There are two kinds: **event
tags** that insert one sound, and **style wraps** that color a whole sentence.

### Event tags — a single transient sound

An event tag is a bare token followed by a space. It inserts one sound (a laugh, a sigh, a
cough…) at that point in the speech.

```text theme={null}
<|laugh|> That is hilarious.
She read it twice, <|sigh|> and still missed the clause.
```

#### Supported events

| Category | Tags |
| - | - |
| Hesitation | `<\|uh\|>` `<\|um\|>` |
| Laughter & crying | `<\|laugh\|>` `<\|chuckle\|>` `<\|giggle\|>` `<\|laugh_harder\|>` `<\|cry\|>` |
| Breath | `<\|breath\|>` `<\|sharp_breath\|>` `<\|inhale\|>` `<\|exhale\|>` `<\|sigh\|>` `<\|gasp\|>` |
| Throat & nose | `<\|cough\|>` `<\|throat_clear\|>` `<\|sniff\|>` `<\|yawn\|>` |
| Mouth sounds | `<\|lip_smack\|>` `<\|gulp\|>` |
| Reactions | `<\|snort\|>` `<\|woo\|>` `<\|hum_tune\|>` `<\|tutu_tune\|>` |
| Pacing | `<\|pause\|>` `<\|long_pause\|>` |

`<|uh|>` and `<|um|>` are the most strongly trained events and the cheapest way to make a
line sound spontaneous rather than read. Put one where a speaker would genuinely stall —
before the word being reached for, not at the start of the sentence — and use at most one
or two per sentence:

```text theme={null}
I was, <|um|> honestly not sure what to say.
```

#### Placement

Where each event sounds most natural:

* **Lead the sentence:** laugh, chuckle, giggle, laugh\_harder, cry, hum\_tune, tutu\_tune,
  woo, yawn, throat\_clear, exhale, sigh, inhale, cough, sniff.
* **Between clauses, mid-sentence:** uh, um, breath, long\_pause, gulp, snort, lip\_smack.
* **Either slot:** pause, gasp, sharp\_breath.

#### Write tags bare

The token, a space, then your text — nothing else. Don't spell the sound out next to the
tag (`<|laugh|> Hahaha!`): the tag already produces it, and the word gets spoken on top.

<Note>
  `<|snort|>` is a derisive snort ("pfft"), not a nasal sniff — use `<|sniff|>` for nasal.
  `<|sharp_breath|>` is a quick intake of breath, distinct from the slower `<|inhale|>`.
</Note>

### Style wraps — color a whole sentence

A style wrap sets the emotional delivery of one full sentence. The shape is always
`<|style_open|>NAME<|style_body|> … <|style_close|>`:

```text theme={null}
<|style_open|>excited<|style_body|> I finally got the offer! <|style_close|>
```

To change style, close the current wrap and open a new one back-to-back. Event tags can sit
inside a wrap:

```text theme={null}
<|style_open|>joyful<|style_body|> I finally got the job offer! <|style_close|><|style_open|>nervous<|style_body|> But the commute is <|uh|> rough. <|style_close|>
```

#### Supported styles

**Emotions:**

```text theme={null}
thoughtful  curious   whispers     mischievously  joyful
excited     sad       surprised    angry          annoyed
sarcastic   tender    fearful      flirty         loud
awe         menacing  nervous      smug
```

**Accents:**

```text theme={null}
strong-british-accent  strong-german-accent  strong-french-accent
strong-italian-accent  southern-us-accent
```

If the feeling you want isn't listed (serious, monotone, fast, slow…), leave the sentence
untagged — neutral is the correct default. Nearest matches: warm/gentle/soft → `tender`,
happy → `joyful`, quietly/whisper → `whispers`, interested → `curious`, concerned/tense →
`nervous`, impressed → `awe`, terrified → `fearful`.

### Rules of thumb

1. Use the **exact tokens** — no spaces inside `<|…|>`, and only the tags listed here.
   Anything else (`<|singing|>`, `<|music|>`, `<|clap|>` …) has no token in the model and is
   either spoken aloud or dropped.
2. **At most one style per sentence**, and a style never carries past `<|style_close|>`.
3. Tags describe the **voice only** — no music, sound effects, or physical actions.
4. Add emphasis with a CAPITAL word, `?` / `!`, or `…` ellipses rather than piling on tags.
5. Pick tags that match the line's emotion; a contradicting tag sounds worse than none.

### Full example

Input text:

> My uncle thought the robot vacuum was a cat and fed it milk. I couldn't stop laughing.
> Then I got worried it might break.

Tagged:

```text theme={null}
<|laugh|> My uncle thought the robot vacuum was a cat and fed it milk! I couldn't stop LAUGHING. <|style_open|>nervous<|style_body|> Then I got worried it might break. <|style_close|>
```

As an API request:

```bash theme={null}
curl -X POST "https://api.tts.runatlas.com/v1/audio/speech" \
  -H "Authorization: Bearer $ATLAS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"atlas-tts","voice":"lylla","input":"<|laugh|> My uncle thought the robot vacuum was a cat and fed it milk! <|style_open|>nervous<|style_body|> Then I got worried it might break. <|style_close|>","response_format":"wav"}' \
  --output out.wav
```

## Ember — bracket tags

The `ember` voice runs on a different model and takes **free-form cues in square
brackets** anywhere in the text. There is no fixed list: any short description of a sound or
a delivery works, so write what you mean.

```text theme={null}
[chuckling] You will not believe what happened next.
I told him the truth. [sigh] He did not take it well.
[whisper in small voice] Don't tell anyone, but I think it's a surprise party.
```

Cues the model responds to well:

| Kind | Examples |
| - | - |
| Sounds | `[laughing]` `[chuckle]` `[chuckling]` `[sigh]` `[inhale]` `[exhale]` `[panting]` `[clearing throat]` `[tsk]` `[singing]` `[audience laughter]` |
| Pacing | `[pause]` `[short pause]` `[emphasis]` `[interrupting]` |
| Emotion | `[excited]` `[angry]` `[sad]` `[surprised]` `[delight]` `[shocked]` |
| Delivery | `[whisper in small voice]` `[professional broadcast tone]` `[with strong accent]` `[low voice]` `[loud]` `[shouting]` `[screaming]` |
| Effects | `[pitch up]` `[volume up]` `[volume down]` `[echo]` |

Keep cues short (a few words), place them right before the speech they should affect, and
don't put the words you want spoken inside the brackets — bracketed text is a direction,
not speech.

<Warning>
  Don't mix dialects. `<|laugh|>` on Ember and `[chuckling]` on an Atlas voice are both
  treated as plain text and may be read aloud.
</Warning>

## Streaming

Tags work unchanged in `text` frames on the [streaming endpoint](/api-reference/streaming).
When the gateway splits a frame into sentences it never cuts inside a `<|…|>` token, a style
wrap, or a `[bracket]` cue. Keep each tag inside a single `text` frame — the two halves of a
token sent in separate frames are plain text.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.