日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2609.18119

ビールを取ってきて:スムーズなピックアンドプレースのための合成から実への階層型ポリシー

Fetch My Beer: Synthetic-to-real Hierarchical Policy for Smooth Pick-and-place

シェア:XThreadsFacebookLINEはてブBluesky

液体入り容器をこぼさずに運ぶため、流体シミュレーションで検証した合成データと、言語・視覚からSE(3)目標を生成する高レベルモジュール+潜在拡散コントローラの階層型ポリシーを組み合わせた合成→実機フレームワークを提案。

詳しい要約

1. どんなもの?

- 液体入り容器をこぼさず安定輸送するpick-and-placeを扱う研究。 - 従来のpick-and-placeと異なり、目標到達だけでなく実行中の物体動態の安定性(揺れ抑制・溢出防止)を要求。 - 運動の滑らかさと軌道レベルの安定性が成功条件となるdynamically sensitive manipulation。 - synthetic-to-realフレームワークを提案。 - 物理的に検証されたデータ生成。 - hierarchical, diffusion-based controller。 - プロジェクトページ: https://fetch-my-beer.github.io/

2. 先行研究と比べてどこがすごい?

- 既存システムの限界を指摘。 - fluid simulationはオンライン強化学習には計算コストが高すぎる。 - 人間のteleoperationは意図しない加速度を生み、imitation learning中にsloshingを誘発。 - 現在のpolicy pipelineは動的安定性よりタスク完了を最適化。 - 提案手法は輸送の滑らかさと動的安定性でstate-of-the-art manipulation policiesを上回ると主張。 - 具体的な比較対象名は要旨からは不明。

3. 技術・手法の肝は?

- synthetic-to-realフレームワーク。 - 物理的に検証されたデータ生成と階層的diffusion-based controllerを結合。 - scalableなデータパイプライン。 - graspsを合成。 - vision-language modelで不安定なposeをフィルタリング。 - fluid simulationで輸送trajectoryを検証。 - policy構成。 - high-level module: 言語と視覚観測をSE(3)制御ターゲットに変換。 - latent diffusion controller: コンパクトなlatent spaceで効率的に計画し、dense action chunksをデコード。 - 高制御周波数を実現し、滑らかで安定した運動を可能にする。

4. どうやって有効だと検証した?

- Extensive experimentsを実施。 - 提案システムがstate-of-the-art manipulation policiesと比較して、輸送の滑らかさと動的安定性で優れることを示した。 - 具体的な評価指標・データセット・実験設定は要旨からは不明。

5. 議論はある?

- 既存手法の課題として、fluid simulationの計算コスト、teleoperationによる意図しない加速度、policy pipelineの最適化対象のずれを議論。 - 提案手法の限界や失敗ケース、一般化性、実機での詳細な議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている具体的な研究名は明示されていない。 - 関連手法として、imitation learning、reinforcement learning、diffusion policy、vision-language model、fluid simulation、hierarchical controlが挙げられる。 - 同分野の定番として、pick-and-place manipulation、dynamic manipulation、sim-to-real transferの文献を読むべき。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yingyue Li, Chenyangguang Zhang, Ruida Zhang, Bowen Fu, Guangyao Zhai, Xiangyang Ji

分類: cs.RO

原文アブストラクト

Many real-world robotic applications require dynamically sensitive manipulation, where success depends not only on reaching a target state but on maintaining stable object dynamics throughout execution. We study the stable transport of liquid-filled containers, where a robot must move objects to target locations while suppressing sloshing and preventing spillage. Unlike conventional pick-and-place, this task imposes stringent requirements on motion smoothness and trajectory-level stability, exposing clear limitations in existing systems. Specifically, fluid simulation remains too costly for online reinforcement learning; human teleoperation introduces unintended accelerations that induce sloshing during imitation learning; and current policy pipelines optimize for task completion rather than dynamic stability. We propose a synthetic-to-real framework coupling physically validated data generation with a hierarchical, diffusion-based controller. The scalable data pipeline synthesizes grasps, filters unstable poses via a vision-language model, and validates transport trajectories through fluid simulation. The policy is organized with a high-level module that translates language and visual observations into SE(3) control targets, and a latent diffusion controller that first plans efficiently in a compact latent space and then decodes dense action chunks, enabling the high control frequency needed for smooth and stable motion. Extensive experiments show our system outperforms state-of-the-art manipulation policies in transport smoothness and dynamic stability. Our project page: https://fetch-my-beer.github.io/

関連論文

PR本紙発行元 EmplifAI