Speech Recognition Based System to Control Electrical Appliances
Arvinder Singh, Gagandeep Singh
Abstract
Arvinder Singh, Gagandeep Singh
Abstract
Speech is one of the natural forms of communication. Recent development has made it possible to use this in the security system and controlling the devices. In speech recognition, the task is to use a speech sample to select the identity of the person that produced the speech from among a population of speakers. An important pre-processing step in Automatic Speech Recognition systems is to detect the presence of noise. It has been shown that accurate speech endpoint detection improves the isolated word recognition accuracy. Also, proper location of regions of speech reduces the amount of processing. This aspect is also important for mobile telephony. Sensitivity to speech variability, inadequate recognition accuracy, and susceptibility to impersonation are among the main technical hurdles that prevent the widespread adoption of speech-based recognition systems. Speech recognition systems work reasonably well with a quiet background but poorly under noisy conditions or in distorted channels. Such a mismatch in the training and testing has severely limited. The objective of this algorithm is the development of signal processing and analysis techniques that would provide sharply improved speech recognition accuracy in any type of noisy environments. Speech is a natural medium of communication for humans, and in the last decade various speech technologies like automatic speech recognition (ASR), Voice response systems and another similar system have considerably matured. The above systems rely on the clarity of the captured speech but many of the real-world environments include noise and reverberation that mitigate the system performance. The key focus of the project is on the effectiveness of ASR.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Speech is one of the natural forms of communication. Recent development has made it possible to use this in the security system and controlling the devices. In speech recognition, the task is to use a speech sample to select the identity of the person that produced the speech from among a population of speakers. An important pre-processing step in Automatic Speech Recognition systems is to detect the presence of noise. It has been shown that accurate speech endpoint detection improves the isolated word recognition accuracy. Also, proper location of regions of speech reduces the amount of processing. This aspect is also important for mobile telephony. Sensitivity to speech variability, inadequate recognition accuracy, and susceptibility to impersonation are among the main technical hurdles that prevent the widespread adoption of speech-based recognition systems. Speech recognition systems work reasonably well with a quiet background but poorly under noisy conditions or in distorted channels. Such a mismatch in the training and testing has severely limited. The objective of this algorithm is the development of signal processing and analysis techniques that would provide sharply improved speech recognition accuracy in any type of noisy environments. Speech is a natural medium of communication for humans, and in the last decade various speech technologies like automatic speech recognition (ASR), Voice response systems and another similar system have considerably matured. The above systems rely on the clarity of the captured speech but many of the real-world environments include noise and reverberation that mitigate the system performance. The key focus of the project is on the effectiveness of ASR.
Key concepts: Speech recognition, Computer science, Voice activity detection, Speech processing, Background noise, Noise (video), Focus (optics), Speaker recognition