Glossary · Updated Sep 2, 2026

Text-to-speech (TTS)

Synthesising natural spoken audio from written text, often with selectable voices, emotions, and languages.

Definition

Modern text-to-speech uses neural models that produce expressive, human-like voices with control over pace, emphasis, and style, and can stream audio with low latency for conversational agents. Products offer voice libraries, cloning, multilingual output, and APIs. Licensing determines whether generated audio can be used commercially.

Why it matters when choosing a tool

TTS quality and rights matter for narration, accessibility, and voice agents; test voices on your script and confirm commercial terms before publishing.

Where you will meet it

AI Audio & Voice, AI Video Generation, AI Learning & Tutoring

Related terms

Voice cloning · Speech-to-text (transcription) · Latency

Tools where this matters

Reviewed products in the related categories