日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
人間AI協調arXiv:2610.09043

慎重な審判:安全で効率的な人間とAIの協調的意思決定

Careful Judge: Safe and Efficient Human-AI Collaborative Decision Making

シェア:XThreadsFacebookLINEはてブBluesky

AIの判断に人間のレビューを組み合わせ、人間への問い合わせを減らしつつ安全性を保証する適応的な較正・修正パイプラインCAREを提案し、運転・言語・ロボティクス分野で有効性を示した。

詳しい要約

1. どんなもの?

- 人間とAIの協調意思決定において、安全性を保ちつつ人間のクエリを削減する手法「CARE」を提案。 - CAREはcalibrated adaptive rectification and escalationの略で、AIモデルと人間レビュアーを組み合わせ、安全で人間に沿った意思決定を保証しつつ、人間フィードバックから継続的に学習して自動化を進める。 - ブラックボックスAIモデルに適用可能で、モジュール式のパイプライン。 - 4つの安全クリティカルな実世界データセット(運転、言語、ロボティクス)で評価。

2. 先行研究と比べてどこがすごい?

- 従来はAIの棄権後の人間介入を一度きりのフォールバックとして扱い、将来のAI決定改善の機会を逃していた。 - また、選択的にクエリされた人間フィードバックからAIが適応学習すると、古いモデル用に調整された安全ガードレールが破られる問題があった。 - CAREはこれらの課題に対処し、安全を保証しつつ人間クエリを25-81%削減することを実証。

3. 技術・手法の肝は?

- CAREはAIモデルと人間レビュアーを組み合わせたエンドツーエンドパイプライン。 - 新規の適応的キャリブレーションモジュールが任意の修正モジュールに対して各時点でのリスク制御を保証。 - AIモデルが十分に訓練され、人間-AIのミスアラインメントに明確な構造がある場合、クエリ効率を改善。 - ブラックボックスAIモデルに適用可能で、モジュール式。

4. どうやって有効だと検証した?

- 運転、言語、ロボティクスにわたる4つの安全クリティカルな実世界データセットで実験。 - CAREが人間に沿った決定を達成しつつ、ベースラインと比較して人間クエリを25-81%削減することを示した。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。同分野の定番として、human-in-the-loop decision making、learning to defer、selective prediction、conformal predictionなどが関連する。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chenyu Zhang, Rachel Luo, Boyi Li, Anjali Parashar, Marco Pavone, Apoorva Sharma

分類: stat.ML, cs.AI, cs.LG

原文アブストラクト

In human-AI collaborative decision making, human review can prevent unsafe AI decisions, but each human judgment is costly. Treating human intervention after AI abstention as a one-off fallback misses the opportunity to improve future AI decisions for greater automation, yet AI adaptively learning from selectively queried human feedback breaks safety guardrails calibrated for old models. We approach this challenge with CARE---calibrated adaptive rectification and escalation---an end-to-end pipeline that combines AI models and human reviewers to guarantee safe, human-aligned decisions, while continuously learning from human feedback to achieve greater automation with fewer human queries. CARE is principled, general, modular, and works with any black-box AI model. Our novel adaptive calibration module guarantees risk control at every time step for any rectification module. We further show how CARE improves query efficiency when the AI model is well trained and the human-AI misalignment has a clear structure. Experiments on four safety-critical real-world datasets spanning driving, language, and robotics demonstrate that CARE achieves human-aligned decisions while reducing human queries by 25-81% relative to baselines.

関連論文

PR本紙発行元 EmplifAI