日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.04681

RoboIRS: 推論時の内部表現操作による汎用ロボットポリシーの性能改善

RoboIRS: Inference-Time Internal Representation Steering for Generalist Robot Policies

シェア:XThreadsFacebookLINEはてブBluesky

成功・失敗ロールアウトから線形分類器を学習し、ポリシーの内部表現を推論時に操作することで、パラメータ更新なしに分布外タスクでの成功率を向上させる手法を提案。

詳しい要約

1. どんなもの?

- VLA と WAM の性能劣化を推論時に回復する手法 - 名前は RoboIRS (Inference-Time Internal Representation Steering) - ポリシーのパラメータを更新せず内部表現を操作 - 成功/失敗ロールアウトから線形分類器を学習 - 結果に関連する介入位置を選びタスク固有のステアリング方向を導出 - 対象は π0.5 ポリシーと Cosmos Policy

2. 先行研究と比べてどこがすごい?

- 従来の推論時介入ベースラインを上回る - ポリシーの再学習やパラメータ更新が不要 - 推論時間の増加がわずか - シミュレーション15タスクで平均成功率 44.4%→66.2% - 実機マニピュレーションでも有効性を確認 - WAM の Cosmos Policy でも 35.4%→55.4% に改善

3. 技術・手法の肝は?

- 成功と失敗のロールアウトを収集 - それらを用いて線形分類器を訓練 - 分類器に基づき結果に関連する介入位置を選択 - タスク固有のステアリング方向を導出 - 推論時に内部表現をその方向へ操作 - ポリシーパラメータは凍結したまま適用

4. どうやって有効だと検証した?

- シミュレーション15タスクで凍結 π0.5 を使用 - 平均成功率 44.4%→66.2% を確認 - 代替の推論時介入ベースラインと比較 - 実機マニピュレーションでも同一 π0.5 で検証 - WAM の Cosmos Policy でも 35.4%→55.4% を確認 - 推論時間の増加が小さいことも示す

5. 議論はある?

- 内部表現を直接操作することで推論時性能を改善可能と主張 - ポリシーパラメータ更新なしで能力回復できる点を強調 - 推論時間増加が小さいことを利点として提示 - 限界や失敗事例、計算コストの詳細は要旨からは不明 - 一般化範囲や他モデルへの適用性は要旨からは不明

6. 次に読むべき論文は?

- π0.5 (frozen policy) - Cosmos Policy (world-action model) - 推論時介入ベースライン (inference-time intervention baselines) - VLA (vision-language-action) モデル - WAM (world-action model)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jiuzhou Lei, Chang Liu, Dayou Li, Zhiyuan Zhang, Xiao Liang, Yu She, Zhiwen Fan, Minghui Zheng

分類: cs.RO

原文アブストラクト

Vision-language-action (VLA) and world-action models (WAMs) often degrade under out-of-distribution task variations despite retaining partial task capability. To recover such capability, we propose RoboIRS, an inference-time internal representation steering method that uses successful and failed rollouts to train linear classifiers, select outcome-relevant intervention locations, and derive task-specific steering directions without updating policy parameters. On 15 simulation tasks with a frozen $π0.5$ policy, RoboIRS improves the average success rate from 44.4% to 66.2%, outperforming alternative inference-time intervention baselines while adding little inference time. We further validate RoboIRS on real-robot manipulation using the same $π0.5$ policy and demonstrate its applicability to a world-action model Cosmos Policy, where the average success rate improves from 35.4% to 55.4%. These results show that directly steering internal robot-policy representations can improve the performance of robot policies at inference time. Project website is available at https://rollingoat.github.io/roboirs/.

関連論文

PR本紙発行元 EmplifAI