プレプリント / バージョン1

日本手話アバターの物理補正結果を確認可能にする診断インタフェースの試作

物理妥当性と言語情報保存性の可視化

##article.authors##

DOI:

https://doi.org/10.51094/jxiv.5962

キーワード:

日本手話、 手話アバター、 物理ベース動作補正、 診断インタフェース、 ヒューマン・イン・ザ・ループ

抄録

手話CGアバターの動作生成において、物理シミュレーションを介した追従再生成(物理追従)は自己貫通等の物理的不整合を低減しうる一方、手の位置や運動の平滑性を変えて言語情報を歪める危険を伴う。本稿では、日本手話単語モーションの物理補正結果を対象に、補正前後の並置再生と物理妥当性・言語情報保存性指標の時系列可視化で問題箇所の特定を支援する診断インタフェースを試作する。語彙カテゴリを考慮して選定した10語では、自己貫通は6語中5語で減少した一方、jerkは全語で悪化し、頭部近接は系統的に希釈され、既定閾値の警告は接触喪失を見逃した。この結果は層別の連続値表示と対話的閾値調整の必要性を裏付ける。ろう当事者による確認・修正インタフェースを見据えた設計仮説もあわせて整理する。

利益相反に関する開示

なし

ダウンロード *前日までの集計結果を表示します

ダウンロード実績データは、公開の翌日以降に作成されます。

引用文献

手話に関する施策の推進に関する法律(令和7年法律第78号,2025-06-25施行).

NHKエンタープライズ:手話CGとデジタルヒューマンKIKI.https://sdgs.nhk-ep.co.jp/program/kiki/(2026-07-27参照).

Saunders, B. et al.: Progressive Transformers for End-to-End Sign Language Production, Proc. ECCV (2020).

Zuo, R. et al.: MaDiS: Taming Masked Diffusion Language Models for Sign Language Generation, arXiv:2601.19577 (2026).

Khan, N. et al.: SignFlow: End-to-End Sign Language Generation for One-to-Many Modeling using Conditional Flow Matching, Proc. ICMI (2025).

Stoll, S. et al.: Text2Sign: Towards Sign Language Production Using Neural Machine Translation and Generative Adversarial Networks, IJCV, Vol. 128 (2020).

Kim, J.-H. et al.: SignBLEU: Automatic Evaluation of Multi-channel Sign Language Translation, Proc. LREC-COLING (2024).

Jiang, Z. et al.: Meaningful Pose-Based Sign Language Evaluation, Proc. WMT (2025).

Cory, O. et al.: BackTranslation2.0 – A Linguistically Motivated Metric to Assess Sign Language Production, Proc. ECCV (to appear), arXiv:2606.28673 (2026).

O'Brien, C. et al.: Evaluation of Pose Estimation Systems for Sign Language Translation, Proc. LREC Workshop on the Representation and Processing of Sign Languages (2026).

Khan, N. et al.: Motion Inbetweening Based on Body Parts Integration for Sign Language Generation, ヒューマンインタフェース学会論文誌,Vol. 26, No. 4 (2024).

Lee, T. et al.: SIGNER: Temporally Grounded Sign Language Generation via Time-Resolved Conditioning, Proc. ECCV (to appear), arXiv:2506.07460 (2026).

戴梓軒・酒向慎司:拡散モデルに基づく3D手話動作の匿名化―意味保持と身元混乱の実現可能性検証―,HCGシンポジウム2025,B-5-1 (2025).

Peng, X. B. et al.: DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills, ACM Trans. Graph., Vol. 37, No. 4 (2018).

Luo, Z. et al.: Perpetual Humanoid Control for Real-time Simulated Avatars, Proc. ICCV (2023).

Luo, Z. et al.: Omnigrasp: Grasping Diverse Objects with Simulated Humanoids, Proc. NeurIPS (2024).

Yuan, Y. et al.: PhysDiff: Physics-Guided Human Motion Diffusion Model, Proc. ICCV (2023).

Li, Z. et al.: Morph: A Motion-free Physics Optimization Framework for Human Motion Generation, Proc. ICCV (2025).

Zhang, Y. et al.: PhysMoDPO: Physically-Plausible Humanoid Motion with Preference Optimization, arXiv:2603.13228 (2026).

Tao, T. et al.: The Last Mile to Production Readiness: Physics-Based Motion Refinement for Video-Based Capture, Proc. SIGGRAPH Asia Technical Communications (2025).

Tavella, F. et al.: Bridging the Communication Gap: Artificial Agents Learning Sign Language through Imitation, Proc. ICSR 2024, LNAI, Vol. 15561 (2025).

Qiao, G. et al.: SignBot: Learning Human-to-Humanoid Sign Language Interaction, Proc. ICRA (2026).

Khan, N. et al.: From Sign Language Generation to Humanoid Execution: Vision-Language Guided Retargeting with Collision Mitigation, arXiv:2607.17769 (2026).

村上智哉ほか:関節回転制約を考慮したメッシュ回帰による手指姿勢推定の検討,FIT2024,H-038 (2024).

箱崎浩平ほか:口型動作を修正した手話CGニュース文の評価と分析,信学技報,WIT2024-33 (2025).

Crasborn, O. et al.: Combining Video and Numeric Data in the Analysis of Sign Languages within the ELAN Annotation Software, Proc. LREC Workshop on Representation and Processing of Sign Languages (2006).

Kipp, M.: ANVIL: The Video Annotation Research Tool, The Oxford Handbook of Corpus Phonology, Oxford University Press (2014).

Nunnari, F. et al.: MMS Player: an open source software for parametric data-driven animation of Sign Language avatars, Adj. Proc. IVA (2025).

Uchida, T. et al.: Motion Editing Tool for Reproducing Grammatical Elements of Japanese Sign Language Avatar Animation, Proc. IEEE ICASSPW (2023).

Ranjbar, H. et al.: Towards an AI-based Sign Language Video Editing Interface, Adjunct Proc. IVA '25 (SLTAT 2025).

Sidenmark, L. et al.: AnimationDiff: A Visual Comparison Tool for Generated 3D Character Animations, Proc. ACM DIS (2026).

長嶋祐二 (2022): 工学院大学 多用途型日本手話言語データベース(KoSign). 国立情報学研究所情報学研究データリポジトリ. (データセット). https://doi.org/10.32130/rdata.5.1

Pavlakos, G. et al.: Expressive Body Capture: 3D Hands, Face, and Body from a Single Image, Proc. CVPR (2019).

Tessler, C. et al.: ProtoMotions3: An Open-source Framework for Humanoid Simulation and Control, NVlabs/ProtoMotions, v3.1 (2025).

Ivashechkin, M. et al.: Two Hands Are Better Than One: Resolving Hand to Hand Intersections via Occupancy Networks, Proc. IEEE FG (2024).

Shitara, A. and Shiraishi, Y.: Improving Continuous Japanese Fingerspelling Recognition with Transformers: A Comparative Study against CNN-LSTM Hybrids, Proc. ACHI 2025, pp. 13–19 (2025).

Liddell, S. K. and Johnson, R. E.: American Sign Language: The Phonological Base, Sign Language Studies, Vol. 64, pp. 195–277 (1989).

ダウンロード

公開済


投稿日時: 2026-08-06 14:36:28 UTC

公開日時: 2026-08-13 07:20:56 UTC
研究分野
情報科学