日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.30965

FRAM: 軌道ガイドによる視覚特徴選択で実現するコンパクトな言語条件付きロボットマニピュレーション

FRAM: Trajectory-Guided Visual Feature Selection for Compact Language-Conditioned Robot Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

将来のエンドエフェクタ軌道を画像上のポインタとして使い、必要な局所視覚特徴だけを読む小型のVLAポリシーを提案。LIBEROで92.2%の成功率を達成し、実機UR5eでもカップ積みを実現した。

詳しい要約

1. どんなもの?

- 小型のVision-Language-ActionポリシーFRAMを提案 - 将来のend-effector軌道を現在の視覚入力に明示的に結びつける - 軌道の画像座標をspatial pointerとして使い、局所的な視覚特徴を読む - 参照位置(Where)、視覚状態(What)、将来運動(Future)に情報を整理 - パラメータ数138.7M(凍結したlanguage encoder含む) - LIBERO 4スイートで平均成功率92.2%

2. 先行研究と比べてどこがすごい?

- 従来のVLAモデルは大規模パラメータを要する - FRAMは138.7MでLIBERO平均92.2%を達成 - 3.3Bのπ0の94.2%に近い性能を大幅に少ないパラメータで実現 - 追加学習なしでLIBERO-Plusでも平均67.3% - 将来軌道と局所視覚特徴の両方が性能と頑健性を改善することをablationで確認

3. 技術・手法の肝は?

- 将来のend-effector軌道を予測し、その画像座標をspatial pointerとして利用 - 現在画像から運動に関連する局所視覚特徴を読み取る - 情報を参照位置(Where)、視覚状態(What)、将来運動(Future)に整理 - 軌道ラベルはdemonstrationとcamera geometryから自動生成 - 手動アノテーション不要 - 凍結したlanguage encoderを含む小型ポリシー

4. どうやって有効だと検証した?

- LIBEROの4標準スイートで平均成功率92.2%を達成 - π0(3.3B)の94.2%と比較 - 追加学習なしでLIBERO-Plusで平均67.3% - ablationで将来軌道と局所視覚特徴の有効性を確認 - 実機dual-arm UR5eでカップ積みを実施 - wrist cameraのみを使用し、左右アームの選択と切替を含む

5. 議論はある?

- 将来運動に基づく視覚情報選択が小型ポリシーで高性能と頑健性を両立する有効な方法であると主張 - 具体的な限界や失敗事例、計算コスト、汎化範囲の議論は要旨からは不明 - 実機実験の詳細な条件や評価指標は要旨からは不明

6. 次に読むべき論文は?

- π0(3.3BパラメータのVLAモデル) - LIBEROベンチマーク - LIBERO-Plusベンチマーク - Vision-Language-Actionモデル全般 - 関連手法としてtrajectory-guided visual feature selection

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hiroshi Ito, Hyogo Hiruma, Yoshiki Kanai, Takahiro Yoshida, Akira Kanazawa, Hiroki Yamada

分類: cs.RO

原文アブストラクト

Vision-Language-Action models achieve strong performance in robot manipulation, but often require large numbers of parameters. In this work, we propose the Future Representation Action Model (FRAM), a small policy that explicitly links the future end-effector trajectory to the current visual input. FRAM uses the image coordinates of the predicted trajectory as spatial pointers and reads local visual features related to the motion from the current image. This organizes the information for action generation into the reference position (Where), the visual state (What), and the future motion (Future). Trajectory labels are generated automatically from demonstrations and camera geometry, so no manual annotation is needed. With 138.7M parameters, including a frozen language encoder, FRAM reaches an average success rate of 92.2% over the four standard LIBERO suites, close to the 94.2% of $π_0$ with 3.3B parameters. Without extra training, it also reaches an average of 67.3% on LIBERO-Plus. Ablations confirm that both the future trajectory and the local visual features improve performance and robustness. On a real dual-arm UR5e, FRAM stacks cups using only wrist cameras, including choosing and switching between the left and right arms. These results show that selecting visual information based on future motion is an effective way to obtain both high performance and robustness in a small robot policy.

関連論文

PR本紙発行元 EmplifAI