ビールを取ってきて:スムーズなピックアンドプレースのための合成から実への階層型ポリシー
Fetch My Beer: Synthetic-to-real Hierarchical Policy for Smooth Pick-and-place
液体入り容器をこぼさずに運ぶため、流体シミュレーションで検証した合成データと、言語・視覚からSE(3)目標を生成する高レベルモジュール+潜在拡散コントローラの階層型ポリシーを組み合わせた合成→実機フレームワークを提案。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Yingyue Li, Chenyangguang Zhang, Ruida Zhang, Bowen Fu, Guangyao Zhai, Xiangyang Ji
分類: cs.RO
原文アブストラクト
Many real-world robotic applications require dynamically sensitive manipulation, where success depends not only on reaching a target state but on maintaining stable object dynamics throughout execution. We study the stable transport of liquid-filled containers, where a robot must move objects to target locations while suppressing sloshing and preventing spillage. Unlike conventional pick-and-place, this task imposes stringent requirements on motion smoothness and trajectory-level stability, exposing clear limitations in existing systems. Specifically, fluid simulation remains too costly for online reinforcement learning; human teleoperation introduces unintended accelerations that induce sloshing during imitation learning; and current policy pipelines optimize for task completion rather than dynamic stability. We propose a synthetic-to-real framework coupling physically validated data generation with a hierarchical, diffusion-based controller. The scalable data pipeline synthesizes grasps, filters unstable poses via a vision-language model, and validates transport trajectories through fluid simulation. The policy is organized with a high-level module that translates language and visual observations into SE(3) control targets, and a latent diffusion controller that first plans efficiently in a compact latent space and then decodes dense action chunks, enabling the high control frequency needed for smooth and stable motion. Extensive experiments show our system outperforms state-of-the-art manipulation policies in transport smoothness and dynamic stability. Our project page: https://fetch-my-beer.github.io/
関連論文
- チャンク型VLAマニピュレーションポリシーの学習と実機展開のためのSim-to-Real統合パイプラインsim2real
- 運動学を超えて:筋駆動模倣学習のためのシミュレーション忠実度ベンチマークsim2real
- CRISP: 多様な形状と接触ソルバを備えた接触リッチロボットシミュレーション基盤sim2real
- 同じ世界、異なる知識:孤立評価が世界モデルの修復を誤判定するときsim2real
- DEXTERA: 単一画像から実機展開可能な巧みなマニピュレーションへ向けたReal-to-Sim-to-Realsim2real
- 単一スキャンからのガウシアンスプラッティングによる実演合成と視覚運動ポリシー学習sim2real