How do modern text-to-speech systems let you control a voice's tone and pacing?
answer
- describe the delivery, don't mark it up
- free-text style instruction, prompt-like
- inline tags for laughs and pauses
- phonetics fix stubborn names
- unsupported tags may be spoken aloud
basics
~20 sMainly through plain-language style instructions — "warm, unhurried, slight smile" — plus inline tags in the script and phonetic spellings for tricky words. Newer voice models have moved away from the heavy XML-style markup that older engines required.
solid answer
~50 sOlder concatenative and parametric engines were driven by markup: XML-like tags wrapping the text to set rate, pitch, emphasis and pauses. Current neural voice models are steered more like a prompt. You supply a free-text description of the delivery — tone, pace, emotional register, even an implied situation such as "apologetic, speaking to a frustrated customer" — and the model infers prosody across the whole utterance rather than per-tag. Two supporting controls remain: inline tags placed in the script itself to mark a laugh, a whisper or a pause, and phonetic spelling for names and domain terms the model mispronounces. As of mid-2026 some leading voice models explicitly no longer honour the older break-and-prosody markup, so a script written against it will silently render its tags as ignored text — or worse, read them aloud. Always listen to the output rather than assuming a control took effect.
go deeper
Know that current voice models are steered with a plain-language description of the delivery plus inline tags, and that phonetic spelling is the fix for mispronounced names.
Explain why neural models condition on a holistic style description rather than per-token markup, and name the failure modes: silently ignored controls, tags read aloud, and drift over long passages.
Show the production discipline — a pronunciation dictionary, script validation against supported tags, caching fixed assets, sentence-level synthesis for live latency, and auditioning as part of release.
Own delivery as a product surface: consistent voice and register across every touchpoint, and an abstraction over steering so a provider change does not mean rewriting every script.
## Where the control moved Text-to-speech steering used to live in markup. The script was wrapped in XML-style tags that set speaking rate, pitch, volume, emphasis and explicit pause lengths, plus phoneme tags to force a pronunciation. That design fit engines that assembled speech from units or parameters, where each knob mapped to something mechanical inside the synthesizer. Neural voice models do not work that way. They generate prosody holistically from the text and from whatever conditioning they are given, so a global description of *how this should sound* is a better fit than dozens of local tags. As of mid-2026 the dominant interface is a free-text style instruction alongside the script — "warm, unhurried, slight smile", "crisp and factual, no filler", "apologetic but not grovelling" — with providers exposing it as a separate instructions field or as part of the voice configuration. Several leading models have explicitly dropped support for the old break and prosody tags. ## The three controls that matter today **1. Style instruction (global).** Sets register, emotion, pace and persona for the whole utterance. This is prompt-like: concrete sensory direction outperforms adjectives. "Like a pharmacist explaining a dosage to an anxious patient — slow, clear, reassuring" produces a more consistent result than "friendly". **2. Inline tags (local).** Markers placed in the script to trigger a specific behaviour at a point — a laugh, a sigh, a whisper, a beat of hesitation. The available set is model-specific, and unknown tags are a real hazard: depending on the model they are ignored or spoken aloud. **3. Pronunciation control.** For drug names, product names, surnames and acronyms, the reliable levers are a phonetic transcription (IPA is the common notation) or a respelling in ordinary letters that happens to synthesize correctly. A pronunciation dictionary applied across your whole script keeps this consistent instead of fixing terms one at a time. ## Why this matters in a voice product Delivery *is* content in speech. The same sentence read briskly and read slowly communicates different things about urgency and care. In a drive-thru or a support line, an over-cheerful read of an apology is worse than a flat one. Because the style instruction is free text, it can also be varied per situation — a confirmation read briskly, a price correction read more carefully — without rewriting the script. ## Practical gotchas - **Silent no-ops.** If the model ignores a control you assumed worked, nothing errors. The output just sounds ordinary. Audition every change. - **Tags read aloud.** The worst failure is a model speaking a literal tag to a customer. Validate scripts against the model's supported set and strip unknown markup. - **Numbers, dates and units.** "15 mg" and "1/2/26" are ambiguous to a synthesizer. Normalize them in text before synthesis rather than hoping the model guesses your locale. - **Instruction drift over long text.** A style instruction is strongest early; on a long passage the delivery can regress toward neutral. Synthesize in shorter chunks, restating the style, and stitch the audio. - **Non-determinism.** Neural TTS may not produce identical audio for identical input. If you need a fixed asset — a legal disclaimer, a brand line — synthesize it once and cache the audio file rather than regenerating it per request. - **Latency in a live loop.** In a conversational agent, synthesizing sentence by sentence lets the first words play while the rest is still being generated. Style instructions apply per request, so keep them consistent across chunks or the voice will shift mid-reply. ## Interview framing The point to make is not that markup is dead in every product — plenty of deployed systems still use it, and some providers still accept it. It is that the *locus of control* moved from declarative per-token markup to a natural-language description of delivery, because that is what neural voice models condition on well. Say which generation of tooling you are describing, and note that the details are provider-specific and change quickly, so the durable skill is auditioning output and building a pronunciation dictionary — not memorizing one vendor's tag list.
- A customer name is consistently mispronounced by your TTS voice. What do you do?Give the model a phonetic transcription — IPA is the usual notation — or a respelling in ordinary letters that synthesizes correctly, and keep it in a pronunciation dictionary applied across every script so the fix is consistent. Prose instructions like 'pronounce it correctly' do not work, because the model has no reference for what correct means here.
- Why is auditioning the audio non-negotiable when you change a steering control?Unsupported controls usually fail silently: the model ignores the tag or instruction and returns perfectly normal-sounding speech, so nothing in the response signals that your change did nothing. The worse variant is a model reading an unrecognized tag aloud to a customer. Listen to the output, and validate scripts against the model's supported tag set before they reach production.
- How would you keep a long spoken passage from drifting back to a neutral delivery?Synthesize in shorter chunks and restate the style instruction on each request, then stitch the audio, checking the joins for level and pace mismatches. Style conditioning is strongest at the start of an utterance, so a single long request tends to regress. For fixed assets like disclaimers, synthesize once and cache the audio rather than regenerating it.
saying these in an interview costs you the question
- Assuming XML-style markup works on every current voice model
- Expecting prose instructions to fix a mispronounced name
- Not listening to output because the request returned successfully
- Regenerating fixed brand or legal lines on every request
- Treating delivery as cosmetic rather than part of the message