Follow your curiosity

What discovery has been shared with you?

Start with one fact. Explore it, go deeper, then follow whichever branch catches your imagination.

Choose subjects for a surprise

Exploring any topic

Begin your discovery

Your next discovery is one click away.

Choose one or more subjects above, or leave Any Topic selected and let curiosity decide.

Technology

Speech Recognition

Quick fact

Modern speech recognition systems can achieve word error rates below 5% in controlled environments, rivaling human accuracy for many tasks.

Why this is interesting

You talk to your phone and it writes your message—but how does it understand your voice, even when you mumble or talk over background noise?

Read the full explanation

Understanding Speech Recognition

Imagine you're translating a foreign language that has no writing system—you have to listen to sounds and figure out the words. Speech recognition does something similar. First, a microphone captures the sound waves and converts them into a digital signal. This signal is chopped into tiny time slices, and for each slice the system extracts features like the frequencies that matter most (think of a piano keyboard highlighting which keys are pressed). These features are then matched against an acoustic model, a statistical map of how basic speech sounds—called phonemes—behave. A phoneme is like a building block, for example the 'k' sound in 'cat' or the 's' in 'sun'. The system strings phonemes together using a language model, which knows which words are likely to appear next (e.g., 'the white house' is more probable than 'the whyte howse'). Finally, a decoder combines all this information to output the most likely string of words.

A deeper explanation

At its core, speech recognition relies on a statistical inference problem: given a sequence of acoustic observations, find the most probable word sequence. This is formalized using Bayes' theorem, where the posterior probability of a word sequence given the audio is proportional to the product of the acoustic model likelihood and the language model prior. The acoustic model, traditionally built with Hidden Markov Models (HMMs) and Gaussian Mixture Models, is now dominated by deep neural networks (DNNs) that learn hierarchical representations directly from data. These networks are trained on thousands of hours of transcribed speech, learning to ignore irrelevant variations (pitch, accent, background noise) while focusing on distinguishing phonetic content. Language models are typically n-gram or neural network based, capturing syntax and context. The system matters because it enables hands-free interaction, accessibility for the disabled, efficient transcription, and seamless integration into smart environments—transforming how we interface with technology.

Keep FACTREE close

Internet access is required. Updates arrive when you reopen or reload the app. You may need to sign in again in the installed app.