Speech recognition trends and predictions, or, do we need text?
Patti J. Price
Abstract
Patti J. Price
Abstract
Speech is the means of communication used first and foremost by humans. Automatic speech recognition capabilities are still far from those of humans, particularly with respect to degradation in the face of noise, nonnative speakers, and casual speech styles. However, significant advances have been achieved in recent years; automatic speech recognition capabilities now permit us to use speech as an interface for dictation and for information access. In such applications, the interactivity of the system allows the person to adapt to the system (and the system to adapt to the user). The real challenge of the information age is using speech not just as a medium for information access, but as itself a source of information, wherever speech occurs. Examples include: voice annotation for later retrieval, sorting and prioritizing voice messages, access to voice recordings of meetings without the need to take minutes, and access to broadcast media. Speech technologies in support of such applications include not just speech recognition, but also speaker recognition, topic spotting, gisting, and summarization. This talk will survey the state of the art and recent trends in speech as a source of information.
OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Speech is the means of communication used first and foremost by humans. Automatic speech recognition capabilities are still far from those of humans, particularly with respect to degradation in the face of noise, nonnative speakers, and casual speech styles. However, significant advances have been achieved in recent years; automatic speech recognition capabilities now permit us to use speech as an interface for dictation and for information access. In such applications, the interactivity of the system allows the person to adapt to the system (and the system to adapt to the user). The real challenge of the information age is using speech not just as a medium for information access, but as itself a source of information, wherever speech occurs. Examples include: voice annotation for later retrieval, sorting and prioritizing voice messages, access to voice recordings of meetings without the need to take minutes, and access to broadcast media. Speech technologies in support of such applications include not just speech recognition, but also speaker recognition, topic spotting, gisting, and summarization. This talk will survey the state of the art and recent trends in speech as a source of information.
Key concepts: Computer science, Speech recognition, Speech analytics, Automatic summarization, Voice activity detection, Audio mining, Speaker recognition, Speech technology