プレプリント / バージョン1

Characterizing temporal organization in presentation-like Japanese speech using the normalized pairwise variability index

##article.authors##

  • Yudai Sakurai Human Informatics and Interaction Research Institute, National Institute of Advanced Industrial Science and Technology
  • Oli Jan Human Informatics and Interaction Research Institute, National Institute of Advanced Industrial Science and Technology
  • Socihiro Matuda Institute of Human Sciences, University of Tsukuba
  • 橘, 亮輔 Human Informatics and Interaction Research Institute, National Institute of Advanced Industrial Science and Technology https://orcid.org/0000-0002-4766-4504 https://researchmap.jp/rtachi

DOI:

https://doi.org/10.51094/jxiv.5974

キーワード:

presentation、 speech、 rhythm、 nPVI

抄録

Objective acoustic measures for presentation-like speech remain limited. We tested whether the normalized pairwise variability index (nPVI), calculated from vowel-onset intervals, distinguishes Japanese speaking contexts. Recordings of presentation-like, dialogue, and rereading speech in the Corpus of Spontaneous Japanese were analyzed. Speech rate was higher in presentation-like speech than in rereading speech, whereas nPVI was higher in dialogue than in the other contexts. Speech rate and nPVI were not clearly correlated. Thus, vowel-onset interval-based nPVI captures local temporal variation beyond overall speech rate and may provide an objective index for describing local temporal organization across speaking contexts, including presentation-like speech.

利益相反に関する開示

No COI

ダウンロード *前日までの集計結果を表示します

ダウンロード実績データは、公開の翌日以降に作成されます。

著者の経歴

橘, 亮輔、Human Informatics and Interaction Research Institute, National Institute of Advanced Industrial Science and Technology

産業技術総合研究所 人間情報インタラクション研究部門

引用文献

T. Lintner and B. Belovecová, “Demographic pre-dictors of public speaking anxiety among univer-sity students,” Curr Psychol, vol. 43, no. 30, pp. 25215–25223, Aug. 2024, doi: 10.1007/s12144-024-06216-w.

M. M. Laske and F. D. DiGennaro Reed, “The ef-ficacy of remote video-based training on public speaking,” Journal of Applied Behavior Analysis, vol. 55, no. 4, pp. 1124–1143, 2022, doi: 10.1002/jaba.947.

L. Chen, G. Feng, J. Joe, C. W. Leong, C. Kitchen, and C. M. Lee, “Towards automated assessment of public speaking skills using multimodal cues,” in Proceedings of the 16th International Conference on Multimodal Interaction, in ICMI ’14. New York, NY, USA: Association for Computing Ma-chinery, Nov. 2014, pp. 200–203. doi: 10.1145/2663204.2663265.

A.-L. Giraud and D. Poeppel, “Cortical oscilla-tions and speech processing: emerging computa-tional principles and operations,” Nat Neurosci, vol. 15, no. 4, pp. 511–517, Apr. 2012, doi: 10.1038/nn.3063.

U. Goswami, “A temporal sampling framework for developmental dyslexia,” Trends in Cognitive Sciences, vol. 15, no. 1, pp. 3–10, Jan. 2011, doi: 10.1016/j.tics.2010.10.001.

A. J. Power, N. Mead, L. Barnes, and U. Goswami, “Neural entrainment to rhythmic speech in chil-dren with developmental dyslexia,” Front. Hum. Neurosci., vol. 7, Nov. 2013, doi: 10.3389/fnhum.2013.00777.

E. Grabe and E. L. Low, “Durational variability in speech and the rhythm class hypothesis,” in La-boratory Phonology 7, C. Gussenhoven and N. Warner, Eds., De Gruyter Mouton, 2008, pp. 515–546.

L. E. Ling, E. Grabe, and F. Nolan, “Quantitative characterizations of speech rhythm: sylla-ble-timing in singapore english,” Lang Speech, vol. 43, no. 4, pp. 377–401, Dec. 2000, doi: 10.1177/00238309000430040301.

J. M. Liss et al., “Quantifying speech rhythm ab-normalities in the dysarthrias,” Journal of Speech, Language, and Hearing Research, vol. 52, no. 5, pp. 1334–1352, Oct. 2009, doi: 10.1044/1092-4388(2009/08-0208).

S. M. Marcus, “Acoustic determinants of percep-tual center (P-center) location,” Perception & Psychophysics, vol. 30, no. 3, pp. 247–256, 1981.

T. V. Rathcke, E. A. Smit, C.-Y. Lin, and H. Kubo-zono, “Testing an acoustic model of the P-center in English and Japanese,” J. Acoust. Soc. Am., vol. 155, no. 4, pp. 2698–2706, Apr. 2024, doi: 10.1121/10.0025777.

K. Maekawa, “Corpus of spontaneous Japanese: its design and evaluation,” in Proceedings of the ISCA & IEEE Workshop on Spontaneous Speech Processing and Recognition (SSPR2003), 2003.

Julius-speech, segmentation-kit. (Dec. 11, 2025). Perl. julius-speech. Accessed: Aug. 01, 2026. [Online]. Available: https://github.com/julius-speech/segmentation-kit

A. Lee, T. Kawahara, and K. Shikano, “Julius --- an open source real-time large vocabulary recog-nition engine,” presented at the Proc. Eurospeech 2001, 2001, pp. 1691–1694. doi: 10.21437/Eurospeech.2001-396.

A. Lee and T. Kawahara, “Recent development of open-source speech recognition engine julius,” presented at the APSIPA ASC 2009 : Asia-Pacific Signal and Information Processing Association, 2009 Annual Summit and Conference, pp. 131–137.

Y. Mukai, D. Brenner, and B. V. Tucker, “Dura-tional variability of spontaneous and read speech: Comparison between English and Japanese,” Proc. Mtgs. Acoust., vol. 56, no. 1, p. 060004, Oct. 2025, doi: 10.1121/2.0002121.

J. Morton, S. Marcus, and C. Frankish, “Perceptual centers (P-centers),” Psychological Review, vol. 83, no. 5, pp. 405–408, 1976, doi: 10.1037/0033-295X.83.5.405.

K. Maekawa and H. Kikuchi, “Corpus-based analysis of vowel devoicing in spontaneous Jap-anese: an interim report,” in Voicing in Japanese, J. van de Weijer, K. Nanjo, and T. Nishihara, Eds., De Gruyter Mouton, 2005, pp. 205–228.

A. Martin, A. Utsugi, and R. Mazuka, “The multi-dimensional nature of hyperspeech: Evidence from Japanese vowel devoicing,” Cognition, vol. 132, no. 2, pp. 216–228, Aug. 2014, doi: 10.1016/j.cognition.2014.04.003.

ダウンロード

公開済


投稿日時: 2026-08-07 02:36:56 UTC

公開日時: 2026-08-18 06:45:52 UTC
研究分野
心理学・教育学