An approach to the classification of American English diphthongs
Michael Gottfried
Abstract
Michael Gottfried
Abstract
Six diphthongs of American English (/au, aɪ, eɪ, oᴜ, ɔɪ, ju/) were produced by four midwestern American speakers (two male, two female) at two tempos (slow, fast) with differing stress (stressed, unstressed) in two contexts ([b_d], [h_d]). Using a plot of the fundamental frequency and the first three formants derived from linear-prediction-coding (LPC) analysis, the onset and offset of each production was determined. The pattern of formants and fundamental frequency at the onset and offset of diphthongs was used to establish a set of parameters that can classify intended productions of the American English diphthongs in varying stress and tempo conditions with an average accuracy of 93%. Results are also presented for diphthong, target-syllable, and sentence durations. The classification results are discussed with respect to hypotheses concerning the perception of diphthongs.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Six diphthongs of American English (/au, aɪ, eɪ, oᴜ, ɔɪ, ju/) were produced by four midwestern American speakers (two male, two female) at two tempos (slow, fast) with differing stress (stressed, unstressed) in two contexts ([b_d], [h_d]). Using a plot of the fundamental frequency and the first three formants derived from linear-prediction-coding (LPC) analysis, the onset and offset of each production was determined. The pattern of formants and fundamental frequency at the onset and offset of diphthongs was used to establish a set of parameters that can classify intended productions of the American English diphthongs in varying stress and tempo conditions with an average accuracy of 93%. Results are also presented for diphthong, target-syllable, and sentence durations. The classification results are discussed with respect to hypotheses concerning the perception of diphthongs.
Key concepts: Diphthong, Formant, American English, Offset (computer science), Mathematics, Sentence, Speech recognition, Computer science