論文

査読有り
2015年10月

Automatic recognition of Japanese vowel length accounting for speaking rate and motivated by perception analysis

SPEECH COMMUNICATION
  • Greg Short
  • ,
  • Keikichi Hirose
  • ,
  • Mariko Kondo
  • ,
  • Nobuaki Minematsu

73
開始ページ
47
終了ページ
63
記述言語
英語
掲載種別
研究論文(学術雑誌)
DOI
10.1016/j.specom.2015.07.001
出版者・発行元
ELSEVIER SCIENCE BV

Automatic recognition of vowel length in Japanese has several applications in speech processing such as for computer assisted language learning (CALL) systems. Standard automatic speech recognition (ASR) systems make use of hidden Markov models (HMMs) to carry out the recognition. However, HMMs are not particularly well-suited for this problem since classification of vowel length is dependent on prosodic information, and since it is a relative feature affected by changes in the durations of surrounding sounds which vary in part due to changes in speaking rates. That being said, it is not obvious how to design an algorithm to account for these contextual dependencies, since there is still not enough known about how humans perceive the contrast. Therefore, in this paper, we conduct perceptual experiments to further understand the mechanism of human vowel length recognition. In our research, we found that the perceptual boundary is largely affected by the vowels two prior, one prior, and following the vowel for which the length is being recognized. Based on these results and the works of others, we propose an algorithm which does post-processing on alignments output by HMMs to automatically recognize vowel length. The method is primarily composed of two levels of processing dealing first with local dependencies and then long-term dependencies. We test several variations of this algorithm. The method we develop is shown to have superior recognition capabilities and be robust against speaking rate differences producing significant improvements. We test this method on three different databases: a speaking rate database, a native database, and a non-native database. (C) 2015 Elsevier B.V. All rights reserved.

リンク情報
DOI
https://doi.org/10.1016/j.specom.2015.07.001
Web of Science
https://gateway.webofknowledge.com/gateway/Gateway.cgi?GWVersion=2&SrcAuth=JSTA_CEL&SrcApp=J_Gate_JST&DestLinkType=FullRecord&KeyUT=WOS:000361576100004&DestApp=WOS_CPL
ID情報
  • DOI : 10.1016/j.specom.2015.07.001
  • ISSN : 0167-6393
  • eISSN : 1872-7182
  • Web of Science ID : WOS:000361576100004

エクスポート
BibTeX RIS