日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.20107

AnyViewDex: 単眼RGB観測による視点不変な巧みなマニピュレーション

AnyViewDex: View-Invariant Dexterous Manipulation from RGB Observations

シェア:XThreadsFacebookLINEはてブBluesky

シミュレーション訓練時に多視点コントラスト学習と3D座標回帰を組み合わせることで、テスト時には未キャリブレーションの単眼RGBと固有感覚のみで視点に依存しない多指ハンド操作を実現した手法。

詳しい要約

1. どんなもの?

- 多指dexterous manipulationのvisuomotor policyを対象とした研究 - カメラ視点の変化に対して頑健な(view-invariant)制御を目指す - テスト時にRGB-Dやpoint cloudsなどの明示的3Dセンシングを使わず、単眼RGBとproprioceptionのみで動作 - 提案手法AnyViewDexは、simulation中に幾何知識を視覚表現へ埋め込むasymmetric training pipeline - 実機xArm7 + 16-DoF LEAP Handで評価

2. 先行研究と比べてどこがすごい?

- 従来のview invariance手法はRGB-Dやpoint cloudsなど明示的3Dモダリティに依存しがち - それらはhardware依存、calibration要件、sensor noise脆弱性を招く - 本研究はtest-time 3D sensingなしでview-invariant controlを実現できることを示す - simulation中に幾何知識を視覚表現へ符号化する点が新しい - 展開時はuncalibrated monocular RGBとproprioceptionのみでzero-shot動作

3. 技術・手法の肝は?

- asymmetric training pipelineを採用 - multi-view contrastive alignmentとprivileged 3D geometric supervisionを組み合わせる - simulated training中に絶対3D object coordinatesをregressする補助目的を追加 - この補助目的がgeometric grounding signalを提供 - globally pooled contrastive embeddingのspatial collapseを緩和する狙い - reinforcement learningとstudent-teacher distillationの両方で検証

4. どうやって有効だと検証した?

- reinforcement learningとstudent-teacher distillationの両設定で検証 - 実機評価はxArm7 + 16-DoF LEAP Handで実施 - 8つのunseen objectsと6つのuncalibrated viewpointsで評価 - 480 trialsで76.7%のgrasping successを達成 - 全ablation条件を含めると2,400 trials - test-time depthなしでzero-shot転移することを示す

5. 議論はある?

- 明示的3Dセンシングなしでview-invariant制御が可能であることを主張 - 幾何知識をsimulation中に視覚表現へ埋め込む設計の有効性を示唆 - 実世界展開時のhardware依存・calibration・sensor noise問題を回避できる可能性 - ただし要旨からは失敗事例や限界、計算コストの詳細は不明 - 一般化範囲や他タスクへの適用性は要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている具体的な先行研究名は記載なし - 関連手法としてRGB-Dやpoint cloudsを用いるview-invariant visuomotor policy研究 - multi-view contrastive learningを用いた視覚表現学習 - student-teacher distillationによるdexterous manipulation - privileged learning / asymmetric actor-critic - 同分野の定番としてdexterous manipulation向けvisuomotor policy学習

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Soham Patil, Om Sanjay Gunjal, Sourabh Bhosale, Arhan Chavare, Ramandeep Singh Hora, Spandan Roy

分類: cs.RO

原文アブストラクト

Visuomotor policies for multi-fingered dexterous manipulation are highly sensitive to camera viewpoint shifts. To achieve view invariance, recent methods increasingly rely on explicit 3D modalities like RGB-D or point clouds, which can introduce hardware dependencies, calibration requirements, and vulnerability to sensor noise during real-world deployment. In this work, we show that view-invariant control can be achieved without explicit test-time 3D sensing by encoding geometric knowledge into the visual representation during simulation. We present AnyViewDex, an asymmetric training pipeline that combines multi-view contrastive alignment with privileged 3D geometric supervision. By regressing absolute 3D object coordinates during simulated training, this auxiliary objective provides a geometric grounding signal that mitigates the spatial collapse of the globally pooled contrastive embedding. At deployment, the policy operates zero-shot using only uncalibrated monocular RGB and proprioception. We validate this approach across both reinforcement learning and student-teacher distillation. In hardware evaluation on an xArm7 with a 16-DoF LEAP Hand, AnyViewDex reaches 76.7% grasping success across eight unseen objects and six uncalibrated viewpoints (480 trials; 2,400 across all ablation conditions), indicating that geometrically grounded monocular policies transfer zero-shot without test-time depth. Project Page: https://anyviewdex.github.io/

関連論文

PR本紙発行元 EmplifAI