日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.27734

InfiNoVA: 視点不変なロボットポリシーのための無限新規視点拡張

InfiNoVA: Infinite Novel View Augmentation for Viewpoint Invariant Robot Policies

シェア:XThreadsFacebookLINEはてブBluesky

マルチカメラの実演を時間変化する3Dガウス表現として再構成し、任意の視点から幾何学的に一貫した新規視点画像を生成するデータ拡張手法を提案。未見視点での操作タスク成功率を大幅に向上させた。

詳しい要約

1. どんなもの?

- Vision-Language-Action (VLA) policies の視点依存性を緩和するデータ拡張フレームワーク InfiNoVA を提案。 - 同期 multi-camera デモを time-varying 3D Gaussian representation として再構成し、任意の camera pose から novel view をレンダリング。 - 元の state-action 対応を保持したまま、密な視点分布の学習データを生成する。

2. 先行研究と比べてどこがすごい?

- 従来は多様な物理視点からのデモ収集が高コストで、視点空間の被覆は疎であった。 - 生成的な novel-view synthesis で問題となる task-critical hallucinations を、明示的 scene representation により低減。 - 未見のランダム視点で、VISTA-based augmentation と未拡張 policy に対し平均成功率 5.4 倍。 - 5 つの物理カメラ視点すべてで直接学習する場合より 1.7 倍高い成功率。

3. 技術・手法の肝は?

- 各 manipulation trajectory を time-varying 3D Gaussian representation として再構成。 - サンプリングした camera pose から novel observation をレンダリングし、元の state-action 対応を維持。 - 明示的 scene representation により frame-level fidelity と temporal consistency を改善。 - policy architecture を変更せずに camera-robust 化を実現。

4. どうやって有効だと検証した?

- 4 つの real-world manipulation tasks で評価。 - 未見のランダム化視点における平均成功率を比較。 - 比較対象は VISTA-based augmentation、未拡張 policy、5 つの物理カメラ視点すべてでの直接学習。 - InfiNoVA は VISTA と未拡張に対し 5.4 倍、全物理視点学習に対し 1.7 倍の成功率を達成。

5. 議論はある?

- 密で幾何学的に根拠のある視点拡張が、policy architecture を変えずに camera-robust な robot policies への実用的経路となることを示す。 - 生成的手法で生じる task-critical hallucinations の低減が利点として議論されている。 - 限界や失敗事例、計算コスト、他タスクへの一般化については要旨からは不明。

6. 次に読むべき論文は?

- VISTA-based augmentation(要旨で比較対象として明示) - Vision-Language-Action (VLA) policies(基盤となる政策) - 3D Gaussian Splatting 系の novel-view synthesis(関連手法として一般名で挙げる)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Sai Puneeth Reddy Gottam, Elmar Rueckert, Vedant Dave

分類: cs.RO, cs.AI

原文アブストラクト

Vision-Language-Action (VLA) policies often rely strongly on the camera viewpoints seen during training, causing substantial performance degradation when deployed from unseen perspectives. Collecting demonstrations from sufficiently diverse physical viewpoints is expensive and still provides only sparse coverage of the viewpoint space. We introduce InfiNoVA, a data-augmentation framework that converts synchronized multi-camera demonstrations into a dense distribution of geometrically consistent training views. InfiNoVA reconstructs each manipulation trajectory as a time-varying 3D Gaussian representation and renders novel observations from sampled camera poses while preserving the original state-action correspondence. This explicit scene representation improves frame-level fidelity and temporal consistency while reducing task-critical hallucinations observed in generative novel-view synthesis. Across four real-world manipulation tasks, policies trained with InfiNoVA achieve 5.4x higher average success under unseen randomized viewpoints than both VISTA-based augmentation and the unaugmented policy. InfiNoVA further achieves 1.7x higher success than training directly on all five physical camera views. These results show that dense, geometrically grounded viewpoint augmentation provides a practical route toward camera-robust robot policies without modifying the underlying policy architecture.

関連論文

PR本紙発行元 EmplifAI