日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.08761

VeriFine: 身体性推論における自己改善のための検証スケーリング

VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning

シェア:XThreadsFacebookLINEはてブBluesky

方策・カリキュラム・判定器を共進化させ、人間の助言を活用して判定器を改良することで、身体性推論タスクにおける継続的な自己改善を実現するフレームワークを提案。

詳しい要約

1. どんなもの?

VeriFineは、embodied reasoningにおける自己改善のためのverificationをスケールさせるagent harness framework。policy、training curriculum、judgeの共進化を通じて検証能力を拡張する。Policy Improvement Loopではrubric judgeが繰り返し失敗を診断し、適応的curriculumを構築してpolicyを最適化。進捗が停滞しverificationがボトルネックになると、Judge Improvement Loopが情報量の多い失敗事例に対して選択的にhuman guidanceを求め、coactive calibrationでjudgeを洗練させる。改訂されたjudgeが次のデータ選択とpolicy最適化を導く。

2. 先行研究と比べてどこがすごい?

従来の固定judgeは最適化フィードバックと有用な訓練例の発見を制約し、自己改善を制限していた。特にembodied reasoningでは空間的grounding、因果推論、安全を考慮した意思決定を評価する必要があり課題が深刻。VeriFineはjudgeを固定せずpolicyやcurriculumと共進化させる点が新しい。

3. 技術・手法の肝は?

Policy Improvement LoopとJudge Improvement Loopの2段階。前者はrubric judgeで失敗を診断し適応的curriculumを構築、policyを最適化。後者は進捗停滞時にhuman guidanceを選択的に活用し、coactive calibration(人間とagentが不一致を解消し物理推論の客観的rubricに収束)でjudgeを改良。改訂judgeが次段階のデータ選択とpolicy最適化を導く。

4. どうやって有効だと検証した?

drivingおよびrobot navigationタスクでの実験により、reinforcement learningとsupervised fine-tuningの両方でpolicyとjudge能力の継続的自己改善を実証。

5. 議論はある?

policyの失敗パターンが進化するにつれ、verificationをスケールさせることが継続的自己改善を支えることを示す。ただし、human guidanceのコストやcoactive calibrationの一般性など詳細は要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究は明示されていない。関連手法としてrubric-based judging、coactive learning、self-improving policies、embodied reasoning benchmarks(例:driving, robot navigation)が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zewei Zhou, Rachel Luo, Yulong Cao, Chaowei Xiao, Chensheng Peng, Boyi Li, Thomas Tian, Zheng Lian, Yan Wang, Jiaqi Ma, Boris Ivanovic, Marco Pavone, Wenhao Ding

分類: cs.AI, cs.RO

原文アブストラクト

Self-improving policies continually expose new failure patterns, changing what their judges must be able to verify. However, current fixed judges constrain both optimization feedback and the discovery of useful training examples, limiting further self-improvement. This challenge is even more acute in embodied reasoning, where reliable evaluation must account for spatial grounding, causal reasoning, and safety-aware decision-making. We introduce VeriFine, an agent harness framework that scales verification through the co-evolution of the policy, training curriculum, and judge. The Policy Improvement Loop uses a rubric judge to diagnose recurring failures, construct an adaptive curriculum, and optimize the policy. When progress plateaus and verification becomes a bottleneck, the Judge Improvement Loop selectively queries human guidance on informative failure cases and refines the judge through coactive calibration, in which humans and agents resolve disagreements and converge toward the objective rubric of physical reasoning. The revised judge then guides the next stage of data selection and policy optimization. Experiments on driving and robot navigation tasks demonstrate continuous self-improvement in both policy and judge capability across reinforcement and supervised fine-tuning. These results show how scaling verification supports continuous self-improvement as policy failure patterns evolve.

関連論文

PR本紙発行元 EmplifAI