日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.14310

VLBiMan++: 視覚言語アンカーによるワンショット両腕マニピュレーションの汎化境界の拡張

VLBiMan++: Expanding the Generalization Boundary of Vision-Language Anchored One-Shot Bimanual Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

単一の人間デモンストレーションから両腕ロボットの操作スキルを抽出し、視覚言語に基づく幾何適応で再学習なしに多様なタスク・物体・環境・機体へ汎化させるフレームワークを拡張した。

詳しい要約

1. どんなもの?

- 単一の人間によるデモンストレーションから、両腕ロボットの汎化可能な操作を実現するフレームワーク VLBiMan++ を提案。 - タスク認識分解と vision-language に基づく幾何適応により、再学習なしで新規構成へスキルを転移。 - タスク・物体・シーン・embodiment・deployment の5次元で汎化境界を拡張。 - 物体状態認識適応と軽量軌道最適化を導入し、剛体6-DoF姿勢変化を超える変化に対応。

2. 先行研究と比べてどこがすごい?

- 従来の one-shot 両腕操作は限定的な転移性の実証に留まっていた。 - VLBiMan++ はタスク・物体・シーン・embodiment・deployment の5次元で体系的に汎化を拡張。 - 大規模 teleoperated demonstrations やポリシー再学習のコストを回避。 - 単一デモから再学習なしで多様な設定に適応可能。

3. 技術・手法の肝は?

- 単一の人間デモからタスク認識分解を行い、再利用可能で適応可能なスキル要素を特定。 - vision-language grounded geometric adaptation により、新規構成へスキルを転移。 - 物体状態認識適応と軽量軌道最適化を導入。 - 剛体6-DoF姿勢変化を超える変化に対応しつつ、両腕協調を維持。

4. どうやって有効だと検証した?

- 広範な実世界実験を実施。 - タスク・物体・シーン・embodiment・deployment の各次元で、ますます困難な設定において高いタスク成功率と適応能力を維持することを実証。 - 具体的な評価指標やベースラインとの比較数値は要旨からは不明。

5. 議論はある?

- 5次元の汎化を体系的に拡張する枠組みを提示し、one-shot 両腕操作を孤立した転移性の実証からスケーラブルな枠組みへ前進させたと主張。 - 限界や失敗事例、計算コスト、安全性に関する議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照・比較されている個別研究は明示されていない。 - 関連手法として one-shot imitation learning、vision-language models、bimanual manipulation、trajectory optimization の定番文献を挙げる。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Huayi Zhou, Wei Gao, Yiyang Han, Kui Jia, Hui Huang

分類: cs.RO

原文アブストラクト

Generalizable bimanual robotic manipulation requires a reusable task prior that can persist across increasingly diverse tasks, objects, scenes, embodiments, and execution conditions, thus avoiding the prohibitive cost of large-scale teleoperated demonstrations and policy retraining. In this work, we present VLBiMan++, an extended framework that expands the generalization boundary of vision-language anchored one-shot bimanual manipulation. Starting from a single human demonstration, VLBiMan++ performs task-aware decomposition to identify reusable and adaptable skill components, and employs vision-language grounded geometric adaptation to transfer these skills to novel configurations without retraining. Building on this foundation, we systematically extend generalization along five dimensions: task generalization through diverse and long-horizon skill compositions; object generalization across unseen categories, varying geometries, and more complex articulated or deformable objects; scene generalization under clutter, occlusion, and dynamic interference; embodiment generalization across heterogeneous dual-arm robotic platforms; and deployment generalization through prolonged closed-loop execution under repeated external perturbations. To support this broader scope, we further introduce object-state-aware adaptation and lightweight trajectory optimization mechanisms that accommodate changes beyond simple rigid 6-DoF pose variations while preserving reliable bimanual coordination. Extensive real-world experiments demonstrate that VLBiMan++ maintains strong task success and adaptation capability across these increasingly challenging settings. Overall, VLBiMan++ advances one-shot bimanual manipulation from demonstrating isolated transferability toward a more systematic and scalable framework for generalization across tasks, objects, scenes, embodiments, and long-term deployment conditions.

関連論文