日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ソフトロボティクス/強化学習arXiv:2608.30773v1

ソフトロボットにおける全身インタラクションによる推論と操作の学習

Learning to infer and manipulate through distributed whole-arm interaction in a soft robot

シェア:XThreadsFacebookLINEはてブBluesky

ソフトロボットの腕全体を使った接触から物体情報を推論し、把持操作を行う強化学習フレームワークを提案。探索と把持を統合したリカレントポリシーとsim-to-real適応により、視覚なしでの把持を実現。

詳しい要約

1. どんなもの?

本論文は、柔らかいロボットアームを用いて、視覚情報なしに物体を探索・把握するための物理的知能フレームワークを提案する。分散したコンプライアントな相互作用を通じて、タスク関連情報の推論と操作行動の組織化を同時に行う。IMUを内蔵したハイブリッド剛軟ロボットアームを用い、強化学習によりメモリベースの制御ポリシーをエンドツーエンドで学習する。

2. 先行研究と比べてどこがすごい?

従来のロボット知能は物理的相互作用を外乱として扱うか、物体の位置ずれ補正に限定していた。本手法は、相互作用そのものを情報獲得と操作の手段として積極的に利用し、視覚なしで物体の把握を実現する点が新しい。また、探索と把握を単一のリカレントポリシーで統合し、sim-to-real適応を2段階で行う点も独自性が高い。

3. 技術・手法の肝は?

手法の核は、(i)広い作業空間探索のための事前学習済み探索ポリシー、(ii)探索と把握の目的を統合したリカレントポリシーの共同最適化、(iii)観測マッピングとポリシー微調整を含む2段階のsim-to-real適応。部分観測問題に対処するため、メモリベースの制御ポリシーを強化学習で学習する。

4. どうやって有効だと検証した?

ハイブリッド剛軟ロボットアームにIMUを埋め込み、視覚情報なしで様々な物体を識別・把握できることを実証した。学習されたポリシーは、作業空間の探索、物体との遭遇・位置特定、把握関連特性の推論、安定した全腕ラッピングを自律的に調整できることを示した。

5. 議論はある?

要旨からは、提案手法の限界や課題についての議論は不明。ただし、部分観測問題への対処やsim-to-real適応の複雑さが今後の課題となる可能性が示唆される。

6. 次に読むべき論文は?

要旨で参照されている関連研究として、動物の知覚と操作の統合に関する研究(例:elephants, octopuses)、ソフトロボティクス、強化学習による操作、sim-to-real適応に関する論文が挙げられる。具体的には、soft robot manipulation, reinforcement learning for control, sim-to-real transferの分野の定番論文が該当する。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Chuhan Zhang, Ebrahim Shahabi, Kseniia Khomenko, Wei Pan, Cosimo Della Santina

分類: cs.RO

原文アブストラクト

In animals such as elephants and octopuses, acquiring non-visual information about an object and physically engaging with it are inseparable processes mediated by rich, large-area interactions between compliant appendages and the environment. Soft robots provide a natural platform for translating this principle into engineered systems. Yet current robotic intelligence makes limited use of physical interaction, treating it primarily as a disturbance to be rejected or, at best, as a means of compensating for object misalignment. Here, we introduce a physical intelligence framework in which distributed compliant interactions jointly reveal task-relevant information and organize manipulation behavior. This results in an intrinsically partially observable problem: key task-relevant information is never measured directly, but must instead be inferred from the history of physical interactions. We propose a reinforcement-learning architecture that addresses this challenge by learning a memory-based control policy end-to-end. The key innovations making this possible are (i) a pretrained exploration policy that provides a reference for broad workspace exploration, (ii) joint optimization that integrates exploration and grasping objectives within a single recurrent policy, and (iii) a two-stage sim-to-real adaptation including observation mapping and policy fine-tuning. We demonstrate this principle through blind whole-arm grasping with a hybrid rigid-soft robotic arm that we equip with IMUs embedded directly within its compliant structure, providing its only source of proprioceptive sensing. The learned policy successfully identifies and grasps various objects by autonomously coordinating workspace exploration, object encounter and localization, inference of grasp-relevant properties, and stable whole-arm wrapping.