日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.03607

相互作用中心のスペクトル潜在誘導による世界行動学習

World Action Learning via Interaction-Centric Spectral Latent Guidance

シェア:XThreadsFacebookLINEはてブBluesky

一人称視点動画から手と物体の相互作用に着目した潜在行動を抽出し、周波数領域でロボット行動と共有される低周波成分を利用して操作ポリシーへ転移する手法WINGを提案。

詳しい要約

1. どんなもの?

- Egocentric videoからrobot policyへ相互作用知識を転移するWINGを提案。 - 観察者起因の動きとhand-object interactionを分離し、interaction-centricなlatent actionを蒸留。 - cross-embodimentのtask semanticsは緩やかに変化する時間構造に集中すると捉え、spectral domainで共有低周波成分を特定しaction生成を誘導。 - LIBERO 99.20%、RoboTwin 2.0 93.80%、RoboCasa-GR1 57.7%の平均成功率。 - 多様なgeneralization設定の4つのreal-world manipulationタスクでも高成績。

2. 先行研究と比べてどこがすごい?

- 従来のegocentric videoからの転移は、frame reconstruction由来のlatent actionがego-camera motion等のnuisance variationに支配されやすい。 - またhumanとrobotのbehaviorは時間的dynamicsが異なるため直接転移が困難。 - WINGはobserver-induced motionを分離しinteraction-centric成分のみを蒸留することでこの問題に対処。 - さらにspectral domainで共有低周波成分を利用し、cross-embodimentのtask semanticsを捉える点が新しい。 - 具体的な先行研究名は要旨からは不明。

3. 技術・手法の肝は?

- 第一段階でobserver-induced motionとhand-object interactionを分離し、interaction-centric componentをlatent actionへ蒸留。 - 第二段階でegocentric latent actionとrobot behaviorの間で、spectral domain上の共有低周波成分を特定。 - その共有低周波成分をguidanceとしてaction generationを誘導。 - これによりcross-embodimentのtask semanticsを緩やかな時間構造として活用。 - 詳細なnetwork構成やloss設計は要旨からは不明。

4. どうやって有効だと検証した?

- LIBEROで平均成功率99.20%を達成。 - RoboTwin 2.0で93.80%を達成。 - RoboCasa-GR1で57.7%を達成。 - さらに4つのreal-world manipulationタスクを多様なgeneralization設定で評価し、強い性能を示した。 - 具体的なbaseline比較やablationの有無は要旨からは不明。

5. 議論はある?

- interaction-centric spectral guidanceが、human egocentric videoからrobot controlへ物理的相互作用知識を転移する有効かつscalableな方法であると主張。 - 大規模real-world interaction data収集の高コスト・スケール困難性に対する代替としてegocentric videoの活用を位置づけ。 - 限界や失敗事例、計算コスト、spectral成分選択の感度などは要旨からは不明。 - 倫理面やデータバイアスへの言及は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照・比較されている個別研究は明示されていない。 - 関連手法として、egocentric videoからのlatent action学習、frame reconstructionベースのaction表現、cross-embodiment transfer、spectral domainでの時間構造解析が挙げられる。 - 同分野の定番として、robot manipulation benchmarkのLIBERO、RoboTwin 2.0、RoboCasa-GR1に関連する論文を次に読むとよい。 - 具体的な論文タイトルは要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhiming Liu, Yikun Miao, Ying Chen, Hongrui Yin, Fangqi Zhu, Xiaoyi Pang, Quanxin Shou, Zhengyang Yan, Haodong Wang, Song Guo

分類: cs.RO

原文アブストラクト

Learning general-purpose robot policies requires large-scale real-world interaction data, yet collecting robot demonstrations remains expensive and difficult to scale. Egocentric videos offer abundant human interaction experience with task-relevant semantics for robotic manipulation, but direct transfer is challenging for two reasons: latent actions inferred from frame reconstruction can be dominated by nuisance variation such as ego-camera motion, and human and robot behaviors often exhibit different temporal dynamics. We propose WING (World Action Learning via INteraction-Centric Spectral Latent Guidance), a framework for transferring interaction knowledge from egocentric videos to robot policies. WING first separates observer-induced motion from hand-object interaction and distills the interaction-centric component into latent actions. It then exploits the observation that cross-embodiment task semantics are concentrated in slowly varying temporal structures, identifying shared low-frequency components between egocentric latent actions and robot behaviors in the spectral domain and using them to guide action generation. WING achieves average success rates of 99.20% on LIBERO, 93.80% on RoboTwin 2.0, and 57.7% on RoboCasa-GR1, and also performs strongly across four real-world manipulation tasks under diverse generalization settings. These results show that interaction-centric spectral guidance provides an effective and scalable way to transfer physical interaction knowledge from human egocentric video to robot control. Project page: https://mikuz12.github.io/wing/

関連論文

PR本紙発行元 EmplifAI