Characterizing temporal organization in presentation-like Japanese speech using the normalized pairwise variability index
DOI:
https://doi.org/10.51094/jxiv.5974キーワード:
presentation、 speech、 rhythm、 nPVI抄録
Objective acoustic measures for presentation-like speech remain limited. We tested whether the normalized pairwise variability index (nPVI), calculated from vowel-onset intervals, distinguishes Japanese speaking contexts. Recordings of presentation-like, dialogue, and rereading speech in the Corpus of Spontaneous Japanese were analyzed. Speech rate was higher in presentation-like speech than in rereading speech, whereas nPVI was higher in dialogue than in the other contexts. Speech rate and nPVI were not clearly correlated. Thus, vowel-onset interval-based nPVI captures local temporal variation beyond overall speech rate and may provide an objective index for describing local temporal organization across speaking contexts, including presentation-like speech.
利益相反に関する開示
No COIダウンロード *前日までの集計結果を表示します
引用文献
T. Lintner and B. Belovecová, “Demographic pre-dictors of public speaking anxiety among univer-sity students,” Curr Psychol, vol. 43, no. 30, pp. 25215–25223, Aug. 2024, doi: 10.1007/s12144-024-06216-w.
M. M. Laske and F. D. DiGennaro Reed, “The ef-ficacy of remote video-based training on public speaking,” Journal of Applied Behavior Analysis, vol. 55, no. 4, pp. 1124–1143, 2022, doi: 10.1002/jaba.947.
L. Chen, G. Feng, J. Joe, C. W. Leong, C. Kitchen, and C. M. Lee, “Towards automated assessment of public speaking skills using multimodal cues,” in Proceedings of the 16th International Conference on Multimodal Interaction, in ICMI ’14. New York, NY, USA: Association for Computing Ma-chinery, Nov. 2014, pp. 200–203. doi: 10.1145/2663204.2663265.
A.-L. Giraud and D. Poeppel, “Cortical oscilla-tions and speech processing: emerging computa-tional principles and operations,” Nat Neurosci, vol. 15, no. 4, pp. 511–517, Apr. 2012, doi: 10.1038/nn.3063.
U. Goswami, “A temporal sampling framework for developmental dyslexia,” Trends in Cognitive Sciences, vol. 15, no. 1, pp. 3–10, Jan. 2011, doi: 10.1016/j.tics.2010.10.001.
A. J. Power, N. Mead, L. Barnes, and U. Goswami, “Neural entrainment to rhythmic speech in chil-dren with developmental dyslexia,” Front. Hum. Neurosci., vol. 7, Nov. 2013, doi: 10.3389/fnhum.2013.00777.
E. Grabe and E. L. Low, “Durational variability in speech and the rhythm class hypothesis,” in La-boratory Phonology 7, C. Gussenhoven and N. Warner, Eds., De Gruyter Mouton, 2008, pp. 515–546.
L. E. Ling, E. Grabe, and F. Nolan, “Quantitative characterizations of speech rhythm: sylla-ble-timing in singapore english,” Lang Speech, vol. 43, no. 4, pp. 377–401, Dec. 2000, doi: 10.1177/00238309000430040301.
J. M. Liss et al., “Quantifying speech rhythm ab-normalities in the dysarthrias,” Journal of Speech, Language, and Hearing Research, vol. 52, no. 5, pp. 1334–1352, Oct. 2009, doi: 10.1044/1092-4388(2009/08-0208).
S. M. Marcus, “Acoustic determinants of percep-tual center (P-center) location,” Perception & Psychophysics, vol. 30, no. 3, pp. 247–256, 1981.
T. V. Rathcke, E. A. Smit, C.-Y. Lin, and H. Kubo-zono, “Testing an acoustic model of the P-center in English and Japanese,” J. Acoust. Soc. Am., vol. 155, no. 4, pp. 2698–2706, Apr. 2024, doi: 10.1121/10.0025777.
K. Maekawa, “Corpus of spontaneous Japanese: its design and evaluation,” in Proceedings of the ISCA & IEEE Workshop on Spontaneous Speech Processing and Recognition (SSPR2003), 2003.
Julius-speech, segmentation-kit. (Dec. 11, 2025). Perl. julius-speech. Accessed: Aug. 01, 2026. [Online]. Available: https://github.com/julius-speech/segmentation-kit
A. Lee, T. Kawahara, and K. Shikano, “Julius --- an open source real-time large vocabulary recog-nition engine,” presented at the Proc. Eurospeech 2001, 2001, pp. 1691–1694. doi: 10.21437/Eurospeech.2001-396.
A. Lee and T. Kawahara, “Recent development of open-source speech recognition engine julius,” presented at the APSIPA ASC 2009 : Asia-Pacific Signal and Information Processing Association, 2009 Annual Summit and Conference, pp. 131–137.
Y. Mukai, D. Brenner, and B. V. Tucker, “Dura-tional variability of spontaneous and read speech: Comparison between English and Japanese,” Proc. Mtgs. Acoust., vol. 56, no. 1, p. 060004, Oct. 2025, doi: 10.1121/2.0002121.
J. Morton, S. Marcus, and C. Frankish, “Perceptual centers (P-centers),” Psychological Review, vol. 83, no. 5, pp. 405–408, 1976, doi: 10.1037/0033-295X.83.5.405.
K. Maekawa and H. Kikuchi, “Corpus-based analysis of vowel devoicing in spontaneous Jap-anese: an interim report,” in Voicing in Japanese, J. van de Weijer, K. Nanjo, and T. Nishihara, Eds., De Gruyter Mouton, 2005, pp. 205–228.
A. Martin, A. Utsugi, and R. Mazuka, “The multi-dimensional nature of hyperspeech: Evidence from Japanese vowel devoicing,” Cognition, vol. 132, no. 2, pp. 216–228, Aug. 2014, doi: 10.1016/j.cognition.2014.04.003.
ダウンロード
公開済
投稿日時: 2026-08-07 02:36:56 UTC
公開日時: 2026-08-18 06:45:52 UTC
ライセンス
Copyright(c)2026
Yudai Sakurai
Oli Jan
Socihiro Matuda
橘, 亮輔
この作品は、Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International Licenseの下でライセンスされています。
