日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
AI安全arXiv:2606.28347v1

エージェントの安全性は行動特性ではなく認識論的特性である

Agentic Safety is an Epistemic Property, Not a Behavioral One

シェア:XThreadsFacebookLINEはてブBluesky

AIシステムの安全性は現在の行動だけでなく、学習・適応・自己改善を通じて将来も修正可能性を保てるかという認識論的特性として捉えるべきだと論じ、その概念として「教示可能性」を導入する。

著者: Charles L. Wang, Keir Dorchen, Peter Jin

分類: cs.CY, cs.AI, cs.LG

原文アブストラクト

Contemporary AI safety spans pre-training interventions, post-training alignment, deployment-time controls, monitoring, and red-teaming. These methods are necessary, but they primarily certify snapshots of system behavior. As AI systems become more capable, dynamic, embodied, and self-improving, this snapshot view becomes incomplete: safety depends not only on whether a system behaves acceptably now, but whether it remains correctable as it learns, adapts, acts, and modifies itself over time. This paper argues that safety should therefore be treated as an epistemic property of the evolving learner, not merely a behavioral property of the current policy. We introduce teachability as the capacity to preserve future corrective leverage under bounded human, institutional, or environmental intervention. We argue that advanced systems can retain visible competence while eroding the representational, algorithmic, or meta-decision conditions needed for future correction. Safe advanced AI systems must not only behave acceptably now; they must remain teachable later.