Speech recognition at a glance
- Invented
- Speech recognition was invented in 1952, when Bell Labs' Audrey system recognized the digits 0 to 9 spoken by a single voice.
- First consumer software
- Dragon Dictate, released in 1990 for about $9,000, was the first speech recognition product sold to the public, and it needed a pause after every word.
- First 1,000-word vocabulary
- Carnegie Mellon's Harpy understood 1,011 words in 1976, the target DARPA set for its five-year Speech Understanding Research program.
- Human parity
- In 2017 Microsoft measured a 5.1% word error rate on the Switchboard telephone benchmark, matching the professional transcribers who worked the same calls.
- Open models
- OpenAI's Whisper, released in September 2022, was trained on 680,000 hours of audio and transcribes nearly 100 languages with publicly released weights.
- Today
- Sonix, founded in 2017, transcribes 54+ languages with 99% accuracy and returns a one-hour recording in about 5 to 6 minutes.
The first speech recognition system was invented in 1952
The first speech recognizer was built at Bell Labs in 1952 and it understood exactly ten words: the digits zero through nine.
Audrey, the Automatic Digit Recognizer
K. H. Davis, R. Biddulph and S. Balashek built Audrey from analog circuits that filled a six-foot relay rack. A speaker said a digit into a telephone handset, the machine measured the vowel formants, and a lamp lit beside the digit it heard. Tuned to its designers' voices it scored 97 to 99 percent; with strangers it fell apart. Audrey never left the laboratory, because a human operator dialing numbers was cheaper, but it settled the question of whether a machine could hear at all.
IBM's Shoebox and the 1960s
In 1962 IBM demonstrated Shoebox at the Seattle World's Fair. It recognized 16 spoken words, the ten digits and six arithmetic commands, and drove an adding machine with them. Through the decade laboratories in the United States, Britain, Japan and the Soviet Union built recognizers for vowels, consonants and small word lists, and in 1968 Taras Vintsyuk described dynamic time warping, the alignment method that let a machine compare speech spoken at different speeds. By 1969 the field had ambition but no theory: John Pierce of Bell Labs compared it to schemes for turning water into gasoline, and the lab stopped funding it for years.
DARPA's Speech Understanding Research program took vocabularies past 1,000 words
Between 1971 and 1976 DARPA funded the largest speech recognition effort to date, and Carnegie Mellon's Harpy reached its goal of 1,000 words.
The five-year goal
The Speech Understanding Research program asked for a system that could understand connected speech from several speakers over a 1,000-word vocabulary with under 10 percent error, and it paid Carnegie Mellon, BBN, SRI and others to compete. Three systems reached the finish line: Hearsay-II and Harpy at Carnegie Mellon and HWIM at BBN. Harpy, finished in 1976, understood 1,011 words, roughly a three-year-old's vocabulary, and was the only one to meet the target. Its trick was a search that pruned unlikely word sequences early instead of weighing every possibility, an idea still at the heart of every decoder.
The statistical turn
The same years produced the idea that beat everything else for the next 40 years. At Carnegie Mellon, James Baker's 1975 thesis described the DRAGON system, which treated speech as a hidden Markov model and let probabilities, not hand-written rules, decide which words were said. At IBM, Fred Jelinek's group built the same approach into a research dictation system and added n-gram language models, so the recognizer could guess the next word from the last two. Bell Labs, meanwhile, moved from one voice to many and built recognizers that worked across speakers, which telephone applications required.
Hidden Markov models made speech recognition statistical during the 1980s
The 1980s replaced pattern matching with probability, and vocabularies grew from a few hundred words to 20,000 in a decade.
How a hidden Markov model hears
A hidden Markov model does not try to match a sound to a template. It asks which sequence of speech sounds most probably produced the audio it received, then which sequence of words most probably produced those sounds. Training on recorded speech sets the probabilities, so more data means a better recognizer, a property that would matter enormously once the internet arrived. Combined with n-gram language models, the HMM became the standard architecture in every research lab and stayed there until deep learning displaced it in 2012.
Tangora, Sphinx and the first products
IBM's Tangora recognized a 20,000-word vocabulary by the mid-1980s, though it needed a pause between words and training on each speaker. In 1988 Kai-Fu Lee's Sphinx system at Carnegie Mellon was the first to combine a large vocabulary, continuous speech and speaker independence in one recognizer. Dragon Systems, founded by James and Janet Baker in 1982, began selling recognition software for personal computers. Voice recognition reached consumers for the first time too: Worlds of Wonder's Julie doll answered a handful of spoken phrases in 1987, and telephone companies trialed voice dialing. The toys were crude, but they put the phrase "voice recognition" into ordinary homes.
The 1990s put speech recognition on personal computers
Faster processors turned dictation from a laboratory demonstration into shrink-wrapped software, and a shared benchmark made progress easy to measure.
Dragon Dictate to NaturallySpeaking
Dragon Dictate shipped in 1990 for about $9,000. It was the first speech recognition product a member of the public could buy, and it demanded a pause after every word. In 1997 Dragon NaturallySpeaking removed the pause: it recognized continuous speech at around 100 words a minute on a home PC, and IBM answered with ViaVoice the same year. For the first time a lawyer, a doctor or a writer could talk to a computer and get usable text, provided they trained it and corrected it.
Phone trees and the first benchmarks
In 1996 BellSouth launched VAL, a dial-in voice portal that answered spoken questions, and every automated phone system you have argued with since descends from it. Behind the scenes, DARPA and the National Institute of Standards and Technology ran yearly evaluations on shared recordings such as the Switchboard telephone corpus, and word error rate became the number every lab reported. From then on, the history of speech recognition is largely the history of that number falling.
The 2000s moved speech recognition into the cloud
Accuracy stalled near 80 percent for most of the decade, until Google moved recognition off the device and onto its servers.
The plateau
By 2001 the best systems transcribed dictated speech at roughly 80 percent accuracy, and the following years brought refinements rather than leaps. Desktop dictation stayed a niche, the phone-tree systems stayed frustrating, and research funding thinned. The hidden Markov model was near its ceiling, and the models that would replace it did not yet have the computing power or the data they needed.
Google Voice Search
Google's GOOG-411 directory service, launched in 2007, was a way to collect millions of spoken queries. In November 2008 the company released Google Voice Search as an iPhone app, and speech recognition changed shape: the phone only recorded, and the recognition ran in Google's data centers on models far larger than any device could hold. Google's English system was trained on 230 billion words of search queries, so it could predict what people were likely to say, and every new query fed the next model. Recognition became a service rather than a program, the form it still takes.
