日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.03710

EyeRobot 2.0:手首カメラなしで精密操作を可能にする能動的視線制御

EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras

シェア:XThreadsFacebookLINEはてブBluesky

単一のステレオカメラのみで、人間の視覚のように視線を能動的に動かして注視点を定め、両手の精密操作を実現するフレームワークを提案。実世界とシミュレーションで評価し、手首カメラなしでも高性能を示した。

詳しい要約

1. どんなもの?

単一の stereo camera のみで細かい bimanual manipulation を可能にするフレームワーク。 - 人間の視覚に着想を得た active gaze を導入。 - 2つの eye viewpoint を swivel させて 3D fixation point に視線を合わせる。 - 得られた画像を foveal に処理し、画像中心により多くの visual tokens を割り当てる。 - これを Active Visual Fixation (AVF) と呼ぶ。

2. 先行研究と比べてどこがすごい?

wrist camera を不要にしつつ、passive stereo や ego + wrist camera の policy と比較して性能を向上。 - wrist camera を外すと標準 policy は real-world 成功率が 52% から 27% に低下。 - EyeRobot 2.0 は stereo のみで passive stereo を real で 40%、sim で 20% 上回る。 - wrist view が明確な場合の ego + wrist policy に匹敵 (69% vs. 64%)。 - 把持物体が wrist camera を遮る場合、成功率を2倍以上に (48% vs. 22%)。

3. 技術・手法の肝は?

階層的な gaze 制御と fixation を利用した表現。 - 低レベルの gaze servoing policy を goal object に条件付けて訓練。 - タスク進捗に基づき fixation goal を出す target selector を訓練。 - 両モジュールを real-world data で RL により訓練。 - 最初は dense geometric reward で訓練。 - 2つ目は BC gripper policy と共訓練し、人間に似た fixation sequence を発見。 - gripper 情報を fixation-relative SE(3) frame に canonicalize し、action distribution をコンパクト化。

4. どうやって有効だと検証した?

7つの real-world タスクと6つの simulated タスクの teleoperation data を収集。 - 1000以上の physical と 1800の simulated robot trials を実施。 - EyeRobot 2.0 を passive stereo および ego + wrist camera policies と比較。 - 同一データで訓練した policy 間で評価。

5. 議論はある?

wrist camera の除去が標準 policy に与える影響と、active gaze による補償を議論。 - occlusion 時の性能差が顕著。 - fixation に基づく SE(3) frame の有効性を示唆。 - 詳細な議論や限界は要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究: passive stereo policies, ego + wrist camera policies, BC gripper policy。 - 関連手法: Active Visual Fixation (AVF), gaze servoing policy, target selector。 - 同分野の定番: bimanual manipulation, stereo vision, reinforcement learning (RL), behavior cloning (BC)。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Kush Hari, Justin Kerr, Nidhya Shivakumar, Samarth Mahapatra, Carmelo Sferrazza, Jiahui Lei, Jitendra Malik, C. Karen Liu, Ken Goldberg, Angjoo Kanazawa

分類: cs.RO, cs.AI

原文アブストラクト

Inspired by human vision, we introduce a framework using active gaze to enable fine-grained bimanual manipulation with only a single stereo camera. EyeRobot 2.0 physically attends to a 3D fixation point in the scene by swiveling two eye viewpoints to center their gaze on it. The resulting images are processed foveally by allocating more visual tokens to the image centers, focusing computation on task-relevant features. Such Active Visual Fixation (AVF) requires carefully coordinated gaze during task execution, which we accomplish hierarchically by first training a low-level gaze servoing policy conditioned on a goal object, then training a target selector which emits fixation goals based on task progress. Both modules are trained with RL on real-world data: the first is trained with a dense geometric reward and the second co-trains with the BC gripper policy which allows it to discover fixation sequences that can resemble a human's fixation sequence while performing the task. EyeRobot 2.0 further takes advantage of fixation by canonicalizing gripper information into a fixation-relative SE(3) frame, which compacts the size of the action distribution to learn. We collect teleoperation data for 7 real-world and 6 simulated tasks, and conduct over 1000 physical and 1800 simulated robot trials comparing EyeRobot 2.0 against passive stereo and ego + wrist camera policies trained on the same data. Removing wrist cameras is costly for standard policies: with only passive stereo, real-world success drops from 52% to 27%. EyeRobot 2.0 closes this gap with only stereo, outperforming passive stereo by 40% in real and 20% in sim. It matches ego + wrist policies when their wrist views are clear (69% vs. 64%), and more than doubles their success when grasped objects occlude the wrist cameras (48% vs. 22%)

関連論文

PR本紙発行元 EmplifAI