Type a sentence into any text-to-speech engine and you'll get a voice reading it back — but "reading" and "performing" are not the same thing. Plain text tells our AI voice engine what to say, but it doesn't tell it how to say it. Should there be a pause before the punchline? Should "Dr." be read as "doctor"? Should the word "never" land harder than the rest of the sentence? That's exactly the gap that Speech Synthesis Markup Language, or SSML, is built to close.
SSML is a markup language, similar in spirit to HTML, that wraps ordinary text in tags to give explicit instructions to a speech engine. Instead of hoping our AI infers the right pacing and tone from context, you can tell it directly. In PlayHT's TTS studio and API, SSML gives you fine-grained control over pauses, emphasis, rate, pitch, and pronunciation — the difference between a voice that sounds fine and one that sounds intentional.
Why plain text isn't enough
Our models are very good at predicting natural-sounding prosody from context, but they're still guessing. A sentence like "That's just great." can be sincere or sarcastic, and nothing in the raw characters tells the engine which one you mean. Numbers are another common failure point: "05/06/2024" could be a date, a fraction, or a code, and a phone number needs to be read digit by digit rather than as one long number. SSML removes the guesswork by letting you annotate the text with your actual intent.
The tags you'll actually use
break — controlling pauses
The <break> tag inserts a pause of a specific length, which is useful for comedic timing, list readability, or just letting a sentence breathe before a key point.
<speak>
Here's the good news<break time="500ms"/>. We shipped it early.
</speak>
You can specify time in milliseconds ("500ms") or seconds ("1s"). A strength attribute ("weak", "medium", "strong", "x-strong") is also supported for cases where you want a natural-feeling pause tied to sentence structure rather than an exact duration.
emphasis — adding stress
The <emphasis> tag tells the engine to stress a word or phrase, shifting pitch, rate, and volume together the way a person would when underlining a point out loud.
<speak>
This offer is available <emphasis level="strong">today only</emphasis>.
</speak>
The level attribute accepts "reduced", "moderate" (the default), and "strong".
prosody — rate and pitch
The <prosody> tag is the most powerful of the group, letting you adjust rate, pitch, and volume over any span of text.
<speak>
<prosody rate="90%" pitch="-2st">
Let's slow down for this next part.
</prosody>
</speak>
rate can be a percentage ("90%" is slightly slower than default, "120%" faster) or a keyword like "slow" or "fast". pitch accepts semitone shifts ("+2st", "-1st") or percentages. Reach for this tag when a narration voice needs more energy in an intro and a calmer tone in a closing line.
say-as — interpreting numbers, dates, and more
The <say-as> tag tells the engine how to interpret a string instead of leaving it to guess the format.
<speak>
Call us at <say-as interpret-as="telephone">1-800-555-0199</say-as>
before <say-as interpret-as="date" format="mdy">12/31/2024</say-as>.
</speak>
Common interpret-as values include "cardinal" and "ordinal" for numbers, "date", "time", "telephone", "currency", and "characters", which spells a string out letter by letter — handy for acronyms and confirmation codes.
Custom pronunciation
For brand names, technical terms, or names our AI voice doesn't pronounce the way you'd like, you can specify the pronunciation directly with a phonetic tag, or override it with a simple text substitution for a quicker fix. This matters most for product names, founder names, or industry jargon that a general-purpose model has rarely, if ever, heard spoken aloud.
Before and after: a few concrete examples
Here's a sentence with no markup at all:
<speak>
I need that report by 5pm. Not 5:30. 5.
</speak>
Read plain, this tends to sound flat — the whole point of the sentence, the deadline, doesn't land. With SSML:
<speak>
I need that report by 5pm.<break time="300ms"/>
Not 5:30.<break time="300ms"/>
<emphasis level="strong">5.</emphasis>
</speak>
Now the pauses create the beat and the emphasis puts weight exactly where it belongs.
Another example — an onboarding message that should feel warm at the start and precise at the end:
<speak>
<prosody rate="105%" pitch="+1st">Welcome aboard!</prosody>
<break time="400ms"/>
Your account code is <say-as interpret-as="characters">A47B9</say-as>.
Please keep it for your records.
</speak>
Without say-as, "A47B9" might get read as a garbled word or a single large number. With it, every character comes through clearly.
Common mistakes to avoid
- Over-tagging. If every third word carries an
<emphasis>tag, nothing sounds emphasized anymore — stress only works in contrast to what surrounds it. Use it sparingly, on the one or two words per sentence that actually carry the meaning. - Stacking too many break tags. A pause after every sentence, everywhere, produces a halting, robotic rhythm instead of natural speech. Save deliberate pauses for the moments that actually need them.
- Invalid nesting. SSML tags must be closed in the order they were opened, just like HTML. Crossing tags so that one closes before another that opened after it will fail to parse correctly.
- Forgetting the wrapping speak tag. All SSML content needs a single root
<speak></speak>element around it to be recognized as SSML rather than plain text. - Extreme prosody values. Pushing rate or pitch too far from the default can make a voice sound strained or artificial. Small, deliberate adjustments read as far more human than dramatic swings.
Where to use SSML in PlayHT
You don't need to write raw markup to benefit from SSML in the PlayHT TTS studio. The studio editor includes an SSML toggle: turn it on and you get inline controls for inserting breaks, emphasis, and pronunciation adjustments directly from the toolbar, with the underlying markup generated for you automatically. This is the fastest path for writers and content teams who want fine control without memorizing tag syntax.
If you're integrating through our developer API, you can pass SSML directly in the raw text field of your request — just wrap your content in a <speak> element and include whatever tags you need. This is the better option when pronunciation and pacing rules need to be generated programmatically, for example when a script is assembled dynamically from a database of names, prices, or dates.
Both paths run through the same underlying engine, so a script you fine-tune in the studio will sound identical when sent through the API, and vice versa.
Start fine-tuning your voice
SSML is the difference between telling our AI voice engine what to say and telling it how to say it. Start small: a handful of <break> and <say-as> tags around your trickiest numbers and dates will fix most of the awkward moments in a script. From there, layer in <emphasis> and <prosody> only where the meaning genuinely calls for it.
If you haven't tried it yet, create a free account and open the studio's SSML toggle on your next script, or check our pricing page if you're ready to move a production workload onto the API.