プレプリント / バージョン1

証拠・境界制約下におけるLLM挙動

YamRail / Evaluation Environment Audit(EEA)の再現性優先型工学研究

##article.authors##

  • OTSU, TAKEHIRO Independent Researcher

DOI:

https://doi.org/10.51094/jxiv.6109

キーワード:

大規模言語モデル、 証拠、 権限境界、 再現性、 人間・AI協働、 来歴、 フェイルクローズ、 非拡張的権限

抄録

大規模言語モデル(LLM)は、証拠、権限、または状態情報が不完全であっても、作業を継続できる。しかし実務においては、継続できる能力が必ずしも継続する許可を意味するわけではなく、もっともらしい完了が必ずしも証拠に裏付けられた完了を意味するわけでもない。

本研究では、人間とLLMの協働において証拠と境界を重視するワークフローであるYamRailの開発過程から生じた一連の運用上の制約を検討する。抽象的な安全理論から出発するのではなく、開発中に実際に使用された工学的実践をまず抽出した。具体的には、必要な証拠が欠けている場合にUNKNOWNおよびHOLDを保持すること、能力と権限を分離すること、parser・hash・attachmentの整合性によって成果物の同一性を検証すること、後の復旧後も過去の失敗状態を保持すること、ならびに追跡可能な証拠参照を維持することである。

これらの実践を、YamRailそのものから独立して適用可能な再現可能な介入へ変換した。unsupported-PASS suppression、non-expansive authority retention、state-history preservation、evidence reachability、bounded utilityの5つの仮説を、Baseline条件とConstraint条件を対にしたfixtureによって評価した。主計測は、単一モデル系列を対象として、各fixture・各条件につきN=3、合計30 experimental unitsおよび36件の成功したprovider requestから構成された。H1およびH2は、ceiling-observed controlにおいてBaseline条件がすでに目標挙動を満たしていたため、効果を実証できなかった。H3およびH4では、試験したfixture内で完全な観測上の分離が認められた。H5は定義したbounded-utility挙動を満たしたが、Baselineに対する増分差は認められなかった。これらは単一モデル系列におけるfixture-levelの観測であり、統計的有意性、モデル間一般化、またはprovider間一般化を主張するものではない。本手順はYamRailを導入せず、著者の開発環境にも依存せず、freshなLLM sessionを用いて再実施できる。ただしclosed frontier APIに対する再現は、bit-level reproductionではなくecological replicationとして扱う。

本研究の主要な問いは、モデルが抽象的な意味で「安全」になるかどうかではなく、証拠および権限の境界を明示したときに、観測可能なLLMの挙動が変化するかどうかである。また、より強い境界保持が単に拒否を増加させるだけなのか、それとも実行主体が、証拠と権限のある範囲内では有用性を維持しつつ、証拠不足または権限外の境界でのみ停止できるのかを検討する。

利益相反に関する開示

本研究に関して、開示すべき利益相反(COI)はありません。

ダウンロード *前日までの集計結果を表示します

ダウンロード実績データは、公開の翌日以降に作成されます。

引用文献

P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W. Yih, T. Rocktäschel, S. Riedel, and D. Kiela, "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks," in Advances in Neural Information Processing Systems 33 (NeurIPS 2020), 2020. arXiv:2005.11401. Peer-reviewed conference proceedings.

S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, "ReAct: Synergizing Reasoning and Acting in Language Models," in The Eleventh International Conference on Learning Representations (ICLR 2023), 2023. arXiv:2210.03629 (v1 2022-10-06; v3 2023-03-10). OpenReview WE_vluYUL-X. Peer-reviewed conference paper (confirmed against official ICLR 2023 proceedings).

T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom, "Toolformer: Language Models Can Teach Themselves to Use Tools," in Advances in Neural Information Processing Systems 36 (NeurIPS 2023), pp. 68539–68551, 2023. arXiv:2302.04761. Peer-reviewed conference proceedings (confirmed against official NeurIPS 2023 proceedings).

L. Moreau and P. Missier (eds.), "PROV-DM: The PROV Data Model," W3C Recommendation, 30 April 2013. https://www.w3.org/TR/2013/REC-prov-dm-20130430/ — Official W3C standards-track specification, not a peer-reviewed research paper.

S. Amershi, D. Weld, M. Vorvoreanu, A. Fourney, B. Nushi, P. Collisson, J. Suh, S. Iqbal, P. N. Bennett, K. Inkpen, J. Teevan, R. Kikin-Gil, and E. Horvitz, "Guidelines for Human-AI Interaction," in Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (CHI '19), 2019. doi:10.1145/3290605.3300233. Peer-reviewed conference proceedings.

J. S. Park, J. C. O'Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein, "Generative Agents: Interactive Simulacra of Human Behavior," in Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST '23), 2023. arXiv:2304.03442. doi:10.1145/3586183.3606763. Peer-reviewed conference proceedings.

Z. C. Lipton, "The Mythos of Model Interpretability," ACM Queue, 16(3), 17 July 2018, article 3241340. Earlier preprint: arXiv:1606.03490, submitted 10 June 2016. ACM Queue viewpoint/article; not classified here as a conventionally peer-reviewed research article. DOI omitted: the ACM DOI landing page was inaccessible during verification and is not needed to identify the article.

F. Doshi-Velez and B. Kim, "Towards A Rigorous Science of Interpretable Machine Learning," arXiv:1702.08608, 2017. Preprint status flagged: no peer-reviewed publication venue was identified for this work at the time of citation verification.

L. Chen, M. Zaharia, and J. Zou, "How Is ChatGPT's Behavior Changing over Time?," Harvard Data Science Review, Issue 6.2, 12 March 2024. doi:10.1162/99608f92.5317da47. Earlier preprint: arXiv:2307.09009 (v3, 31 October 2023). Journal version of record.

B. Haibe-Kains, G. A. Adam, A. Hosny, F. Khodakarami, et al. (Massive Analysis Quality Control (MAQC) Society Board of Directors), L. Waldron, B. Wang, C. McIntosh, A. Goldenberg, A. Kundaje, C. S. Greene, T. Broderick, M. M. Hoffman, J. T. Leek, K. Korthauer, W. Huber, A. Brazma, J. Pineau, R. Tibshirani, T. Hastie, J. P. A. Ioannidis, J. Quackenbush, and H. J. W. L. Aerts, "Transparency and Reproducibility in Artificial Intelligence," Nature, 586, E14–E16, 14 October 2020. doi:10.1038/s41586-020-2766-y. Published by Nature as a Matters Arising piece (response to McKinney et al. 2020), not a conventional full research article; issue number omitted as it was not shown on the verified primary page. Used here as a primary scholarly source for the transparency/reproducibility argument specifically, supplemented by [11] as the primary general machine-learning-reproducibility source.

J. Pineau, P. Vincent-Lamarre, K. Sinha, V. Larivière, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and H. Larochelle, "Improving Reproducibility in Machine Learning Research (A Report from the NeurIPS 2019 Reproducibility Program)," Journal of Machine Learning Research, 22(164), 1–20, 2021. arXiv:2003.12206. Peer-reviewed journal article.

X. Bouthillier, C. Laurent, and P. Vincent, "Unreproducible Research is Reproducible," in Proceedings of the 36th International Conference on Machine Learning (ICML 2019), PMLR 97:725–734, 2019. Peer-reviewed conference proceedings.

PyTorch Contributors, "Reproducibility," official PyTorch documentation. https://docs.pytorch.org/docs/stable/notes/randomness — Official framework documentation (living document, not a peer-reviewed paper). PyTorch-specific evidence; not generalized to all frameworks without qualification.

PyTorch Contributors, "Numerical Accuracy," official PyTorch documentation. https://docs.pytorch.org/docs/stable/notes/numerical_accuracy.html — Official framework documentation (living document, not a peer-reviewed paper). PyTorch-specific evidence; not generalized to all frameworks without qualification.

ダウンロード

公開済


投稿日時: 2026-08-17 15:21:40 UTC

公開日時: 2026-08-19 07:26:45 UTC
研究分野
情報科学