Glossary · Updated Sep 2, 2026

Speech-to-text (transcription)

Converting spoken audio into written text, the foundation of meeting assistants and captioning tools.

Definition

Speech-to-text models transcribe audio with speaker labels, timestamps, and punctuation, and many now handle dozens of languages and noisy conditions. Accuracy varies with accents, vocabulary, and audio quality. Downstream features such as summaries and action items are only as good as the transcript.

Why it matters when choosing a tool

Transcript accuracy on your own recordings should be tested before trusting summaries or searchable meeting archives.

Where you will meet it

AI Meeting Assistants, AI Audio & Voice, AI Translation & Localization

Related terms

Text-to-speech (TTS) · Multimodal AI · Large language model (LLM)

Tools where this matters

Reviewed products in the related categories