φLLM / AI Worker 技術報告書
多次元の運用作業挙動の測定および境界条件付き評価のための測定枠組み
DOI:
https://doi.org/10.51094/jxiv.6112キーワード:
AI Worker、 φLLM、 Worker Profile、 運用作業挙動、 多次元評価、 境界条件付き評価、 証拠追跡可能性、 適合性評価抄録
本研究では、LLMを孤立した抽象的な基盤モデルとしてではなく、モデル、ハーネス、ツール、作業指示、証拠条件、権限条件、評価者、および実行環境から構成される「AI Worker」として扱う。複数の運用上の作業挙動を独立した軸として観測し、各軸の意味、条件、証拠参照を保持したまま、多次元のWorker Profileへ束ねる工学的計測方法を検討する。研究計画全体では、段階的な論理負荷に対する応答変化を観測するphi_LLMを候補となる計測核として位置付けるが、本稿で報告するE1はその段階載荷本試験ではない。E1は、外部証拠からWorker Profileを構造的に評価できるかを確認する予備的な計測実証である。
E1では、共通の作業指示の下で得られた3件の一回試行Worker実行結果を匿名化し、凍結した評価設計v0.2を用いてfreshな一次評価者へ投入した。評価中はrunとmodelの対応関係を分離し、返却原本を固定した後にのみHuman Gateでmappingを開示した。ただし、一次評価者のrequested model labelは評価対象runの一つと重複しており、自己評価バイアスを排除できない。観測されたrun vectorは、GPT-5.6 Lunaが(2,2,NA,2,1,2,2,2,2)、GPT-5.6 Solが(2,2,NA,2,2,2,2,2,2)、GPT-5.4 Miniが(1,1,NA,UNKNOWN,0,0,0,1,0)であった。Miniのrunでは、凍結入力外の4ファイルを取得して使用したscope deviationが観測された一方で、25件の明示的なHOLDも保持された。
境界を越えたアクセス事象そのものは運用作業挙動の証拠として保持し、その追加情報に依存した内容評価はCONTAMINATED_UNKNOWNとして分離した。
各model conditionはN=1であり、backend model identity、reasoning configuration、provider routing等の一部条件はUNKNOWNである。したがって、本結果はモデルの恒久的特性、一般的な順位、安全性、または機微情報漏洩傾向を示すものではない。それでも限定された条件下では、scope compliance、boundary preflight、evidence integrity、HOLD retention、authority boundary等を独立した運用作業挙動軸として記録する方法が実行可能であることを示した。特に、同一run内でscope complianceとHOLD retentionが異なる状態を示したことは、Workerの挙動を単一尺度で表すのではなく、複数軸を保持したまま用途判断へ束ねる必要性を示す初期的所見である。また本研究では、評価の追跡可能性として、束ねられた判断から元の軸、観測事象、証拠、条件へ遡及可能であることを要求する。
本稿はこの初期結果を単一の閉じた評価で確定するものではなく、第三者による再現、反例生成、評価機構のクロスチェック、および方法論的批判へ引き渡すための技術報告として位置付ける。公開または外部検証そのものを安全性の証明とは扱わない。
利益相反に関する開示
本研究に関して、開示すべき利益相反(COI)はありません。ダウンロード *前日までの集計結果を表示します
引用文献
[AIST-1] National Institute of Advanced Industrial Science and Technology (AIST). 機械学習品質マネジメントガイドライン 第4版 [Machine Learning Quality Management Guideline, 4th Edition]. December 2023. DigiARC-TR-2023-03 / CPSEC-TR-2023003. DOI: 10.57346/REP.2023.3121.AIST-000000.
[AIST-2] National Institute of Advanced Industrial Science and Technology (AIST). Reference Guide to Machine Learning Quality Management. March 2022. DigiARC-TR-2022-03 / CPSEC-TR-2022004.
[AIST-3] National Institute of Advanced Industrial Science and Technology (AIST). 生成AI品質マネジメントガイドライン 第1版 [Generative AI Quality Management Guideline, 1st Edition]. May 2025. IPRI-TR-2025-01 / CPSEC-TR-2025001. DOI: 10.50886/0002003354.
[WP-1] S. Yu, F. Carroll, and B. L. Bentley, "The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment," arXiv:2604.12116, 2026.
[WP-2] T. Ou, W. Guo, A. Gandhi, G. Neubig, and X. Yue, "AgentDiagnose: An Open Toolkit for Diagnosing LLM Agent Trajectories," in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 207–215, Association for Computational Linguistics, 2025. DOI: 10.18653/v1/2025.emnlp-demos.15.
[WP-3] Z. Pei, H.-L. Zhen, Y. Zhang, Z. Yang, X. Li, X. Yu, M. Yuan, and B. Yu, "Behavioral Fingerprinting of Large Language Models," arXiv:2509.04504, 2025.
[RR-1] S. Yao, N. Shinn, P. Razavi, and K. Narasimhan, "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains," arXiv:2406.12045, 2024.
[PHI-1] P. Shojaee, I. Mirzadeh, K. Alizadeh, M. Horton, S. Bengio, and M. Farajtabar, "The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity," Advances in Neural Information Processing Systems 38 (NeurIPS 2025), 2025.
[PHI-2] T. Kwa, B. West, J. Becker, et al., "Measuring AI Ability to Complete Long Software Tasks," Advances in Neural Information Processing Systems 38 (NeurIPS 2025), 2025. arXiv:2503.14499.
[ROUTE-1] I. Ong, A. Almahairi, V. Wu, W.-L. Chiang, T. Wu, J. E. Gonzalez, M. W. Kadous, and I. Stoica, "RouteLLM: Learning to Route LLMs with Preference Data," arXiv:2406.18665, 2024.
[EVAL-2] J. Wang, Y. Hu, W. Yang, Z. Pan, X. Li, and L.-Z. Guo, "Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling," in Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 23174–23200, Association for Computational Linguistics, 2026. DOI: 10.18653/v1/2026.acl-long.1062.
[NIST-1] C. Autio, R. Schwartz, J. Dunietz, S. Jain, M. Stanley, E. Tabassi, P. Hall, and K. Roberts, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile," NIST AI 600-1, National Institute of Standards and Technology, 2024. DOI: 10.6028/NIST.AI.600-1.
[RV-1] M. Leucker and C. Schallhart, "A Brief Account of Runtime Verification," The Journal of Logic and Algebraic Programming, 78(5), pp. 293–303, 2009. DOI: 10.1016/j.jlap.2008.08.004.
[RV-2] L. Convent, S. Hungerecker, M. Leucker, T. Scheffel, M. Schmitz, D. Thoma, et al., "TeSSLa: Temporal Stream-Based Specification Language," in Formal Methods: Foundations and Applications (SBMF 2018), LNCS 11254, pp. 144–162, Springer, 2018.
[TRACE-1] L. Moreau and P. Missier (eds.), "PROV-DM: The PROV Data Model," W3C Recommendation, 30 April 2013.
[TRACE-2] J. Cheney, P. Missier, and L. Moreau (eds.), "Constraints of the PROV Data Model," W3C Recommendation, 30 April 2013.
[ISO-1] ISO/IEC 17000:2020, Conformity assessment — Vocabulary and general principles, 2nd ed., International Organization for Standardization, 2020.
[ISO-2] ISO/IEC 17025:2017, General requirements for the competence of testing and calibration laboratories, 3rd ed., International Organization for Standardization, 2017.
ダウンロード
公開済
投稿日時: 2026-08-17 15:56:02 UTC
公開日時: 2026-08-19 07:29:18 UTC
バージョン
- 2026-08-31 05:29:29 UTC(2)
- 2026-08-19 07:29:18 UTC(1)
改版理由
日本語: 著者E-mailアドレス追記の軽微是正依頼に対応しました。本文内容に変更はありません。ライセンス
Copyright(c)2026
OTSU, TAKEHIRO
この作品は、Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International Licenseの下でライセンスされています。
