日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
報酬モデル/マルチビューarXiv:2609.20106

AnyviewMeter: カメラ幾何とマルチビュー注意によるロボット報酬モデルの適応

AnyviewMeter: Adapting Robotic Reward Models with Camera Geometry and Multi-View Attention

シェア:XThreadsFacebookLINEはてブBluesky

カメラ視点の変化に頑健なロボット報酬モデルを実現するため、低ランク微調整とPlucker光線条件付け、同期ブロック注意を組み合わせた適応フレームワークを提案し、視点変更時の誤差を約21%削減した。

詳しい要約

1. どんなもの?

- ロボットの報酬モデルを視点変化に適応させる枠組み - タスク進捗をスカラー報酬として表現 - 事前学習済み Robometer をパラメータ効率よく適応 - 単一視点予測と多視点同時評価の両方をサポート

2. 先行研究と比べてどこがすごい?

- RGB ファインチューニングと比較して視点変化時の誤差を約21%削減 - 単一視点適応は全カメラ群で進捗予測を改善 - 多視点予測は単一視点RGB平均より誤差を41-69%削減 - 約88%のタスク・カメラ群で時間的順序が改善

3. 技術・手法の肝は?

- 低ランクファインチューニングを採用 - token-aligned Plucker rays でカメラ幾何を視覚特徴と attention の query/key に注入 - synchronous block attention で同期視点を事前学習済み decoder 内で融合 - 単一視点と多視点の両モードに対応

4. どうやって有効だと検証した?

- PickCube で単一視点適応を評価 - 視野変更時の平均絶対誤差を約21%削減 - シミュレーション操作タスクで多視点予測を評価 - 実タスクで固定カメラと手首カメラを用いて検証 - 平均絶対誤差を約21%削減

5. 議論はある?

- カメラ幾何と統合視覚証拠がタスク特化型報酬適応に有効と結論 - 要旨からは不明:計算コスト、実世界の多様な視点への汎化、失敗事例

6. 次に読むべき論文は?

- Robometer(事前学習済み報酬モデル) - Plucker rays を用いた視覚モデル - 低ランク適応(LoRA 等) - 多視点 attention 融合手法

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yuang Tu, Runjia Tan, Yujie Yan, Jinghan Hu, Chen Lv

分類: cs.CV

原文アブストラクト

Robotic reward models evaluate task execution from visual observations, but their predictions can change with camera viewpoint and occlusion even when the underlying task state is unchanged. Adapting a pretrained reward model to a local task therefore requires accounting for how that task is observed. We introduce AnyviewMeter, a geometry-conditioned adaptation framework for robotic reward models that represent task progress as a scalar reward signal. It combines low-rank fine-tuning with token-aligned Plucker rays and synchronous block attention: ray conditioning incorporates camera geometry into visual features and attention queries and keys, while block attention fuses synchronized views inside the pretrained decoder. The framework supports both single-view reward prediction and joint multi-view evaluation through parameter-efficient adaptation of a pretrained Robometer model. On PickCube, single-view adaptation improves progress prediction in every camera group and reduces mean absolute error under a changed field of view by approximately 21% relative to RGB fine-tuning. Across simulated manipulation tasks, joint multi-view prediction reduces progress error by 41-69% compared with averaging single-view RGB predictions and improves temporal ordering in approximately 88% of task-camera groups. On real tasks with fixed and wrist-mounted cameras, mean absolute error decreases by approximately 21% relative to averaged RGB fine-tuning. These results support camera geometry and joint visual evidence as useful components of task-specific robotic reward adaptation.

PR本紙発行元 EmplifAI