日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.04913

R²-WAM:World Action Modelのための修復・棄却ポストトレーニング

$R^2$-WAM: Repair-and-Reject Post-Training for World Action Models

シェア:XThreadsFacebookLINEはてブBluesky

映像予測モデルを「入力行動と整合する未来を予測するよう修復」し、その予測を使って劣った行動サンプルを棄却・負例微調整することで、環境との追加対話なしに方策を改善する手法。

詳しい要約

1. どんなもの?

- World Action Models (WAMs) を基盤としたポリシー改善フレームワーク $R^2$-WAM を提案。 - 2段階の repair-and-reject ポストトレーニング:予測未来と入力行動の整合性を高める repair 段階と、劣った行動サンプルを選択して negative fine-tuning する reject 段階からなる。 - 追加の環境相互作用や推論手続きの変更なしに、ビデオ予測を表現学習から結果に基づくポリシー改善へ拡張。

2. 先行研究と比べてどこがすごい?

- 従来の WAMs は視覚的に尤もらしい予測が入力行動を反映しない場合があり、ポリシー改善を誤導する問題があった。 - $R^2$-WAM は予測未来と入力行動の整合性を kinematic alignment score で評価し修復することで、このミスマッチを解決。 - 結果として、RoboTwin 2.0 で平均成功率 93.8%、実世界の Fold Shirt タスクで 87.5% を達成し、Fast-WAM の 0% を大幅に上回る。

3. 技術・手法の肝は?

- Repair 段階:kinematic alignment score を用いて予測と実演の動きの一致度を測定し、予測ビデオが入力行動を忠実に反映するよう修正。 - Reject 段階:修復されたビデオモデルでサンプル行動と実演行動の想像結果を比較し、予測タスク進捗が実演参照より所定のマージン以下であるサンプルに選択的に negative fine-tuning を適用。 - これにより、追加の環境相互作用なしに結果ベースのポリシー改善を実現。

4. どうやって有効だと検証した?

- RoboTwin 2.0 ベンチマークでクリーンおよびランダム化設定において平均成功率 93.8% を達成。 - 実世界の長期間タスク Fold Shirt で平均成功率 87.5% を達成し、Fast-WAM の 0% と比較。 - これらの結果から、提案手法の有効性を実験的に検証。

5. 議論はある?

- 要旨からは、手法の限界や議論の詳細は不明。 - ただし、視覚的に尤もらしい予測が入力行動を反映しない問題に対処することが主眼であり、その有効性が示されている。

6. 次に読むべき論文は?

- 要旨で参照されている Fast-WAM や World Action Models (WAMs) に関する論文。 - 関連手法として、ビデオ予測を用いたポリシー学習や negative fine-tuning に関する研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ruiyan Xu, Haisheng Su, Sixu Lin, Zhaokun Yue, Chengming Hu, Xin Jin, Guiliang Liu

分類: cs.RO

原文アブストラクト

World Action Models (WAMs) emerge as a promising foundation for policy refinement by predicting the consequences of sampled actions. However, visually plausible predictions can mislead policy refinement if they fail to reflect the input actions. To address this mismatch, we introduce $R^2$-WAM, a two-stage repair-and-reject post-training framework that first improves the consistency of predicted futures with input actions, then uses these futures to select inferior action samples for negative fine-tuning. The repair stage grounds imagination in observed robot behavior through a kinematic alignment score that measures agreement between predicted and demonstrated motion, enabling the predicted video to faithfully reflect its input actions. Using the repaired video model, the rejection stage compares imagined outcomes of sampled and demonstrated actions, selectively applying negative fine-tuning to samples whose predicted task progress falls below the demonstrated reference by a prescribed margin. Together, the two stages extend video prediction from representation learning to consequence-based policy refinement without additional environment interaction or changes to the inference procedure. $R^2$-WAM achieves 93.8% average success on RoboTwin 2.0 across clean and randomized settings. On the long-horizon real-world Fold Shirt task, it achieves 87.5% average success, compared with 0% for Fast-WAM.

関連論文

PR本紙発行元 EmplifAI