日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ソフトロボティクスarXiv:2609.17035

SWIM: 視覚言語に基づくソフト全身インタラクティブマニピュレーション

SWIM: Vision-Language-Grounded Soft Whole-Body Interactive Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

RGB画像と言語指示からソフトロボットの全身駆動指令列を生成するVLAフレームワークSWIMを提案し、パッキング・リーチング・把持タスクで高成功率を達成した。

詳しい要約

1. どんなもの?

- ソフト・連続体ロボットの全身操作を言語と視覚から実行するフレームワークSWIMを提案。 - 初期RGB観察と言語指示から完全なアクチュエーションコマンド列を生成。 - VLAポリシーSWIM-VLAは拡散アクションヘッドとVisual Soft Proprioception (VSP)を統合。 - 平面腱駆動ソフトロボットでpacking, reaching, graspingを評価。

2. 先行研究と比べてどこがすごい?

- 従来のソフトロボット操作は言語・視覚文脈を全身アクチュエーションに変換するのが困難。 - 適応したOpenVLA-OFTベースラインや制御アブレーションを上回る性能。 - シミュレーションで成功率100%, 96%, 88%を達成。 - ハードウェアで100%, 80%, 75%を達成し、直接オンライン展開の75%, 40%, 25%を大幅に改善。

3. 技術・手法の肝は?

- SWIM-VLAは拡散アクションヘッドとVSPを共有表現で結合。 - 拡散ヘッドは専門家コマンドチャンクの条件付き分布をモデル化。 - VSPはシミュレーションのグラウンドトゥルースを用いて順序付きボディアンカー予測を監督。 - 限られたデモンストレーションから身体幾何学を保持する表現を学習。 - 進化するシミュレーション観察からの反復仮想ロールアウトでコマンド列を生成。 - 内在的コンプライアンスがオンラインポリシークエリなしで局所接触適応を提供。

4. どうやって有効だと検証した?

- 平面腱駆動ソフトロボットでpacking, reaching, graspingを評価。 - シミュレーションでSWIM-VLAが成功率100%, 96%, 88%を達成。 - 適応したOpenVLA-OFTベースラインと制御アブレーションを上回る。 - ハードウェアで成功率100%, 80%, 75%を達成。 - 同じポリシーチェックポイントの直接オンライン展開(75%, 40%, 25%)と比較。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- OpenVLA-OFT(比較ベースラインとして参照)。 - その他の関連手法は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Tingcong Liu, Aye Phyu Phyu Aung, Junjie Xiong, Siyi Ma, Bo An, Ke Wu, Senthilnath Jayavelu

分類: cs.RO

原文アブストラクト

Soft and continuum robots enable manipulation through distributed body deformation and contact, yet translating language and visual context into executable whole-body actuation remains a fundamental challenge. We present SWIM, a framework that maps an initial RGB observation and a language instruction to a complete actuation-command sequence. Its vision-language-action (VLA) policy, SWIM-VLA, combines a diffusion action head with Visual Soft Proprioception (VSP) through a shared representation of RGB observations, language instructions, and tendon states. The diffusion head models conditional distributions of expert command chunks, while VSP supervises ordered body-anchor predictions using simulation ground truth, encouraging the representation to retain body geometry when learning from limited demonstrations. Embodied mechanical intelligence supports physical execution of command sequences generated through iterative virtual rollout from evolving simulated observations, with intrinsic compliance providing local contact adaptation without online policy queries. We evaluate SWIM on packing, reaching, and grasping on a planar tendon-driven soft robot, with grasping targets anchored. In simulation, SWIM-VLA achieves success rates of 100\%, 96\%, and 88\%, respectively, outperforming an adapted OpenVLA-OFT baseline and controlled ablations. On hardware, SWIM achieves success rates of 100\%, 80\%, and 75\%, compared with 75\%, 40\%, and 25\% for direct online deployment of the same policy checkpoint.

関連論文