日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
触覚arXiv:2610.08784

PEARS: 物理事前知識と触覚フィードバックによる失敗推論・拡散誘導を用いた効率的適応

PEARS: Physical-Prior-Guided Efficient Adaptation via Failure Reasoning and Diffusion Steering for Tactile Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

視覚言語モデルの物理事前知識で失敗を診断し接触力制約を更新、触覚条件付き拡散誘導で事前学習方策をオンライン適応させる、サンプル効率の高い触覚マニピュレーション手法。

詳しい要約

1. どんなもの?

PEARSは、触覚フィードバックを備えた事前学習済みロボットポリシーを、実世界でのOOD状況に適応させるための、物理事前知識に基づくハイブリッドRLフレームワーク。 - 目的: 展開時の性能劣化を、実世界インタラクションによるpost-trainingで改善する。 - 課題: RLベースのpost-trainingは環境インタラクションを大量に必要とし、操作では各試行が遅く高コスト・破壊的になりやすい。 - 構成: physics-guided force reasoning (PFR) モジュールと、tactile-conditioned diffusion steering reinforcement learningを組み合わせる。 - 対象: 触覚フィードバックを用いたmanipulationタスク。

2. 先行研究と比べてどこがすごい?

RLベースのpost-trainingは多くの環境インタラクションを要する点が課題だった。 - 提案手法は、物理事前知識と触覚を活用し、サンプル効率の高いオンライン適応を実現する。 - シミュレーションで、最強のper-taskベースラインに対し成功率を12.4-37.4 percentage points改善。 - 一定成功率閾値に必要なインタラクションepisode数を、最速ベースライン比で最大53.2%削減。 - 実世界でWhiteboard Erasing 95%、Pipette Liquid Aspiration 90%の成功率を達成。

3. 技術・手法の肝は?

物理事前知識に基づくハイブリッドRLフレームワーク。 - physics-guided force reasoning (PFR): 各episode後、vision-language model (VLM)にエンコードされた物理事前知識を使い、視覚結果と触覚インタラクション履歴から失敗を診断し、タスクに適した接触力のboundsを更新。 - high-frequency hybrid force-position controller: 接触中にこれらのboundsを強制。 - tactile-conditioned diffusion steering reinforcement learning: 凍結したflow-matching policyの潜在ノイズを調整し、自由空間運動と接触タイミングの誤りを修正。ベースモデルは更新しない。

4. どうやって有効だと検証した?

シミュレーションと実世界実験で検証。 - シミュレーション: 最強のper-taskベースラインに対し成功率を12.4-37.4 percentage points改善。一定成功率閾値までのインタラクションepisode数を最速ベースライン比で最大53.2%削減。 - 実世界: Whiteboard Erasingで95%、Pipette Liquid Aspirationで90%の成功率。 - プロジェクトサイト: https://song-kun.github.io/pears

5. 議論はある?

PFRモジュールとpolicy steeringの組み合わせが、適応を加速しつつ高コストなインタラクションを削減できることを示す。 - 詳細な限界や失敗事例、計算コスト、VLMの誤診断リスクなどは要旨からは不明。 - 実世界タスクは2種類のみで、一般化可能性の議論は要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究は明示されていない。 - 関連手法として、RL-based post-training、vision-language model (VLM)、flow-matching policy、diffusion steering、hybrid force-position control、tactile manipulationが挙げられる。 - 同分野の定番として、robot manipulationにおけるsim-to-real transferやtactile reinforcement learningの論文を読むとよい。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Kun Song, Yiming Wang, Yilin Chen, Tianyi Ding, Jiaxin Tian, Tianqi Gong, Daolin Ma, Jia Pan

分類: cs.RO

原文アブストラクト

Pretrained robotic policies can suffer substantial performance degradation under out-of-distribution (OOD) conditions encountered during deployment, motivating post-training through real-world interaction. However, reinforcement-learning (RL)-based post-training typically requires substantial environment interactions, a burden that is especially significant in manipulation, where each trial can be slow, costly, or destructive. Therefore, we present PEARS, a physics-prior-guided hybrid RL framework for sample-efficient online adaptation of pretrained policies with tactile feedback. After each episode, its physics-guided force reasoning (PFR) module uses physical priors encoded in a vision-language model (VLM) to diagnose failures from the visual outcome and tactile interaction history and update task-appropriate contact-force bounds. A high-frequency hybrid force-position controller then enforces these bounds during contact. Complementarily, tactile-conditioned diffusion steering reinforcement learning adjusts the latent noise of the frozen flow-matching policy to correct errors in free-space motion and contact timing without updating the base model. In simulation, PEARS improves success rates by 12.4-37.4 percentage points over the strongest per-task baselines. PEARS also reduces the number of interaction episodes required for a certain success threshold by up to 53.2% relative to the fastest baseline. In real-world experiments, PEARS achieves success rates of 95% on Whiteboard Erasing and 90% on Pipette Liquid Aspiration. These results show that combining the PFR module with policy steering can accelerate adaptation while reducing costly interactions. The project website is available at https://song-kun.github.io/pears.

関連論文

PR本紙発行元 EmplifAI