Intelligent Speech/Audio Processing for Multimedia Applications
Author information unavailable
Abstract
Author information unavailable
Abstract
Intelligent speech and audio processing can provide efficient and smart interfaces for various multimedia applications. Generally, speech is the most natural form of human communication. Audio and music can enhance our emotional impacts and promote interest in multimedia applications. A successful interactive multimedia system must have the capabilities of speech and audio compression, text-to-speech conversion, speech understanding, and music synthesis. The main purpose of speech and audio compression is to provide cost-effective storage or to minimize transmission costs. Text-to-speech converts linguistic information stored as data or text into speech for the applications of talking terminals, alarm systems, and audiotext services. Speech understanding systems make it possible for people to interact with computers using human speech. Its success relies on the integration of a wide variety of speech technologies, including acoustic, lexical, syntactic, semantic, and pragmatic analyses. The applications of music processing for multimedia were mostly realized by means of the combination of music, graphics, video, and other media. Since musical sounds and compositions can be precisely specified and controlled by a computer, we can easily create artificial orchestras, performers, and composers. Nowadays, multimedia systems have become more sophisticated with the advances made in computer and microelectronic technologies. Many applications require efficient processing of speech and audio for interactive presentations and integration with other types of media. The application-specific hardwares are proposed to meet the high-speed, low-cost, lightweight, and low-power requirements. The design example of a speech recognition processor and system for voice-control applications is introduced. The industrial standards and commercial products of speech and audio processing ar e also summarized in this chapter.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Intelligent speech and audio processing can provide efficient and smart interfaces for various multimedia applications. Generally, speech is the most natural form of human communication. Audio and music can enhance our emotional impacts and promote interest in multimedia applications. A successful interactive multimedia system must have the capabilities of speech and audio compression, text-to-speech conversion, speech understanding, and music synthesis. The main purpose of speech and audio compression is to provide cost-effective storage or to minimize transmission costs. Text-to-speech converts linguistic information stored as data or text into speech for the applications of talking terminals, alarm systems, and audiotext services. Speech understanding systems make it possible for people to interact with computers using human speech. Its success relies on the integration of a wide variety of speech technologies, including acoustic, lexical, syntactic, semantic, and pragmatic analyses. The applications of music processing for multimedia were mostly realized by means of the combination of music, graphics, video, and other media. Since musical sounds and compositions can be precisely specified and controlled by a computer, we can easily create artificial orchestras, performers, and composers. Nowadays, multimedia systems have become more sophisticated with the advances made in computer and microelectronic technologies. Many applications require efficient processing of speech and audio for interactive presentations and integration with other types of media. The application-specific hardwares are proposed to meet the high-speed, low-cost, lightweight, and low-power requirements. The design example of a speech recognition processor and system for voice-control applications is introduced. The industrial standards and commercial products of speech and audio processing ar e also summarized in this chapter.
Key concepts: Computer science, Multimedia, Audio mining, Speech processing, Speech recognition, Speech synthesis, Voice activity detection