Language-independent pronunciation assessment based on reference-based comparison using self-supervised representations
DOI:
https://doi.org/10.51094/jxiv.5347キーワード:
Computer-assisted language learning、 Pronunciation assessment、 self-supervised models、 dynamic timewarping抄録
This paper proposes a language-independent pronunciation assessment method that utilizes self-supervised learning (SSL) pre-trained models. While most conventional computer-assisted pronunciation training (CAPT) systems rely on automatic speech recognition (ASR) technologies, developing such systems for low-resource languages is difficult. To address this issue, we
investigate an approach that directly compares a learner’s (L2) utterance with a native (L1) reference utterance of the same text to evaluate pronunciation quality without requiring task-specific fine-tuning or additional training. In the proposed method, feature vectors are extracted using the SSL models. The distance between the two utterances is then calculated using dynamic time warping (DTW). To evaluate the performance of the method, we conducted experiments on four distinct corpora of L2 speech in English, Chinese, and Japanese. Experimental results demonstrate that the SSL-based methods significantly outperform the conventional MFCC baseline, showing substantially higher correlations across all corpora. In particular, HuBERT and ContentVec outperform Wav2Vec 2.0, with ContentVec achieving the most stable correlation to human annotations across languages and datasets for fluency evaluation. Conversely, the correlation remains relatively low for prosodic features that heavily depend on pitch, such as Japanese pitch accent, revealing a limitation in current SSL models’ ability to fully capture pitch-related features.
利益相反に関する開示
We have no conflict of interest.ダウンロード *前日までの集計結果を表示します
引用文献
T. Hasumi and M.-S. Chiu, “Technology-enhanced language learning in English language education: Performance analysis, core publications, and emerging trends,” Cogent Education, vol. 11, no. 1, p. 2346044, 2024.
Y. Wang and M. K. Kabilan, “Integrating technology into English learning in higher education: a bibliometric analysis,” Cogent Education, vol. 11, no. 1, p. 2404201, 2024.
M. Amrate and P.-h. Tsai, “Computer-assisted pronunciation training: A systematic review,” ReCALL, vol. 37, no. 1, pp. 22–42, 2025.
C. Blanco, “2025 duolingo language report.”
[Online]. Available: https://blog.duolingo.com/2025duolingo-language-report/
L. Yang, K. Fu, J. Zhang, and T. Shinozaki, “Nonnative acoustic modeling for mispronunciation verification based on language adversarial representation learning,” Neural Networks, vol. 142, pp. 597–607, 2021.
D. Jiang, G. C. Shuang, and A. H. M. Adnan, “Content knowledge base for AI-based Chinese computer-assisted pronunciation training,” Malaysian Journal of Social Sciences and Humanities (MJSSH), vol. 11, no. 3, pp. e003 729–e003 729, 2026.
G. Demenko, A. Wagner, and N. Cylwik, “The use of speech technology in foreign language pronunciation training,” Archives of Acoustics, pp. 309–329, 2010.
S. N. Mehta, A. Roth, C. Munteanu, and S. Chandna, “Ai-based pronunciation assessment and grammaticalerror correction with feedback for the German language,” in International conference on human-computer interaction. Springer, 2025, pp. 388–407.
A. Demin, G. Vorontsov, and D. Chaikovskii, “Human– AI feedback loop for pronunciation training: A mobile application with phoneme-level error highlighting,” Multimodal Technologies and Interaction, vol. 10, no. 1, p. 2, 2025.
V. Martínez-Paricio and J. Koreman, “The Spanish computer-assisted listening and speaking tutor: a multilingual approach to pronunciation training,” Revista Española de Lingüística Aplicada/Spanish Journal of Applied Linguistics, vol. 38, no. 1, pp. 134–161, 2025.
F. J. Yaury and A. Zahra, “Computer-assisted assessment of phonetic and prosodic features in spanish pronunciation practice,” in 2026 International Seminar on Intelligent Business and Edge-Computing Research (ISIBER). IEEE, 2026, pp. 167–171.
E. Pyshkin, A. Kusakari, J. Blake, N. B. Pham, and N. Bogach, “Multimodal modeling of the mora-timed rhythm of Japanese and its application to computerassisted pronunciation training,” in 2023 14th IIAI International Congress on Advanced Applied Informatics (IIAI-AAI). IEEE, 2023, pp. 174–179.
L. Campbell and A. Belew, Cataloguing the world’s endangered languages. Routledge London and New York, 2018, vol. 711.
L. Bromham, X. Hua, C. Algy, and F. Meakins, “Language endangerment: A multidimensional analysis of risk factors,” Journal of Language Evolution, vol. 5, no. 1, pp. 75–91, 2020.
G. F. Simons and M. P. Lewis, “The world’s languages in crisis: A 20-year update,” in 26th Linguistics Symposium: Language Death, Endangerment, Documentation, and Revitalization, 2012, pp. 3–20.
L. Hinton, “Language revitalization: An overview,” The green book of language revitalization in practice, vol. 1, p. 18, 2001.
C. K. Galla, “Indigenous language revitalization, promotion, and education: Function of digital technology,” Computer Assisted Language Learning, vol. 29, no. 7, pp. 1137–1151, 2016.
R. Pálsson and J. Guðnason, “Computer-assisted pronunciation training in Icelandic (CAPTinI): developing a method for quantifying mispronunciation in L2 speech,” Intelligent CALL, granular systems and learner data: short papers from EUROCALL 2022, p. 334, 2022.
Y.-L. Chuang, H.-R. Hsu, Y.-W. Liu, C.-T. Hsin et al., “Computer-assisted pronunciation training system for Atayal, an Indigenous language in Taiwan,” in 2024 27th Conference of the Oriental COCOSDA International Committee for the Co-ordination and Standardisation of Speech Databases and Assessment Techniques (O-COCOSDA). IEEE, 2024, pp. 1–6.
Y.-C. Huang, Y.-H. Chen, C.-C. Kuo, C.-S. Huang, and Y.-F. Liao, “A GOP-based automatic pronunciation scoring system for Taiwanese Hakka using transformer regression models,” in 2025 28th Conference of the Oriental COCOSDA International Committee for the Coordination and Standardisation of Speech Databases and Assessment Techniques (O-COCOSDA). IEEE, 2025, pp. 1–6.
E. Ramos-Aguilar, A. Olvera-López, and I. OlmosPineda, “WavLM-based automatic pronunciation assessment for Yuhmu speech: A low-resource language,” Computación y Sistemas, vol. 29, no. 3, 2025.
H. Hamada, S. Miki, and R. Nakatsu, “Automatic evaluation of english pronunciation based on speech recognition techniques,” in Proc. Eurospeech 1989, 1989, pp. 1421–1424.
T. Konno, A. Ito, M. Ito, S. Makino, and M. Suzuki, “Intonation evaluation of english utterances using synthesized speech for computer-assisted language learning,” in 2008 International Conference on Natural Language Processing and Knowledge Engineering. IEEE, 2008, pp. 1–7.
K. Qian, Y. Zhang, H. Gao, J. Ni, C.-I. Lai, D. Cox, M. Hasegawa-Johnson, and S. Chang, “ContentVec: An improved self-supervised speech representation by disentangling speakers,” in International conference on machine learning. PMLR, 2022, pp. 18 003–18 017.
T. Koshikawa, A. Ito, and T. Nose, “Fast and speakerindependent utterance selection for ASR-Free CALL systems of minority languages,” in 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), 2025, pp. 825– 830.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Selfsupervised speech representation learning by masked prediction of hidden units,” IEEE/ACM transactions on audio, speech, and language processing, vol. 29, pp. 3451–3460, 2021.
S. M. Witt and S. J. Young, “Phone-level pronunciation scoring and assessment for interactive language learning,” Speech communication, vol. 30, no. 2-3, pp. 95–108, 2000.
J. Fu, Y. Chiba, T. Nose, and A. Ito, “Automatic assessment of English proficiency for Japanese learners without reference sentences based on deep neural network acoustic models,” Speech Communication, vol. 116, pp. 86–97, 2020.
A. I. Zahran, A. A. Fahmy, K. T. Wassif, and H. Bayomi, “Fine-tuning self-supervised learning models for end-toend pronunciation scoring,” IEEE Access, vol. 11, pp. 112 650–112 663, 2023.
K. Fu, L. Peng, N. Yang, and S. Zhou, “Pronunciation assessment with multi-modal large language models,” arXiv preprint arXiv:2407.09209, 2024.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in neural information processing systems, vol. 33, pp. 12 449–12 460, 2020.
N. Anand, M. Sirigiraju, and C. Yarra, “Unsupervised speech intelligibility assessment with utterance level alignment distance between teacher and learner Wav2Vec-2.0 representations,” arXiv preprint arXiv:2306.08845, 2023.
M. Bartelds, C. Richter, M. Liberman, and M. Wieling, “A new acoustic-based pronunciation distance measure,” Frontiers in Artificial Intelligence, vol. Volume 3 - 2020, 2020. 9
F.-A. Chao, T.-H. Lo, T.-I. Wu, Y.-T. Sung, and B. Chen, “A hierarchical context-aware modeling approach for multi-aspect and multi-granular pronunciation assessment,” in Proc. Interspeech 2023, 2023, pp. 974–978.
S. Salvador and P. Chan, “Toward accurate dynamic time warping in linear time and space,” Intelligent data analysis, vol. 11, no. 5, pp. 561–580, 2007.
N. Ketkar and J. Moolayil, “Introduction to PyTorch,” in Deep learning with python: learn best practices of deep learning models with PyTorch. Springer, 2021, pp. 27–91.
N. Minematsu, Y. Tomiyama, K. Yoshimoto, K. Shimizu, S. Nakagawa, M. Dantsuji, and S. Makino, “Development of English speech database read by Japanese to support CALL research,” in Proc. International Congress on Acoustics (ICA), vol. 1, no. 2004, 2004, pp. 557–560.
Speech Database Committee of the Priority Areas Project on “Advanced Utilization of Multimedia to Promote Higher Education Reform”, “English speech database read by Japanese students.”
[Online]. Available: https://doi.org/10.32130/src.UME-ERJ
J. Zhang, Z. Zhang, Y. Wang, Z. Yan, Q. Song, Y. Huang, K. Li, D. Povey, and Y. Wang, “Speechocean762: An open-source non-native English speech corpus for pronunciation assessment,” in Proc. Interspeech 2021, 2021, pp. 3710–3714.
J. Zhang, “Speechocean762: A non-native English corpus for pronunciation scoring task.”
[Online]. Available: https://github.com/jimbozhang/speechocean762
W.-W. Hsieh, H.-W. Chi, K.-C. Wang, P.-C. Yeh, T.-h. Liu, and C.-Y. Chiang, “OMPAL: Bridging speech and learning with an open-source Mandarin pronunciation assessment corpus for global learners.” in Proc. Interspeech 2025, 2025, pp. 2415–2419.
P. Hsieh, “OMPAL corpus.”
[Online]. Available: https://github.com/phantomhsieh/OMPAL-corpus
K. Nishina, “Speech database construction for Japanese as second language learning,” in Proceedings of SNLPOriental COCOSDA 2002, 2002, pp. 187–192.
Y. Xiong, J. Park, and O. J. Lee, “Prosodic fluency measures in Korean learners’ Mandarin,” Linguistic Research, vol. 43, no. 1, p. 295, 2026.
K. Idemaru, P. Wei, and L. Gubbins, “Acoustic sources of accent in second language japanese speech,” Language and Speech, vol. 62, no. 2, pp. 333–357, 2019.
Y. Yamashita, K. Kato, and K. Nozawa, “Automatic scoring for prosodic proficiency of english sentences spoken by japanese based on utterance comparison,” IEICE transactions on information and systems, vol. 88, no. 3, pp. 496–501, 2005.
K. Hirabayashi and S. Nakagawa, “Automatic evaluation of english pronunciation by japanese speakers using various acoustic features and pattern recognition techniques.” in Interspeech, 2010, pp. 598–601.
Y. Gong, Z. Chen, I.-H. Chu, P. Chang, and J. Glass, “Transformer-based multi-aspect multi-granularity nonnative english speaker pronunciation assessment,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, pp. 7262–7266.
K. Fu, S. Gao, S. Shi, X. Tian, W. Li, and Z. Ma, “Phonetic and prosody-aware self-supervised learning approach for non-native fluency scoring,” in Proc. Interspeech 2023, 2023, pp. 949–953.
E. Kim, J.-J. Jeon, H. Seo, and H. Kim, “Automatic pronunciation assessment using self-supervised speech representation learning,” in Proc. Interspeech 2022, 2022, pp. 1411–1415.
ダウンロード
公開済
投稿日時: 2026-07-02 05:32:47 UTC
公開日時: 2026-07-14 05:13:26 UTC
ライセンス
Copyright(c)2026
Ito, Akinori
Takashi Nose
この作品は、Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International Licenseの下でライセンスされています。
