日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.18242

ForceDelta-VLA: 接触の多いマニピュレーションのための力条件付き行動補正の蒸留

ForceDelta-VLA: Distilling Force-Conditioned ActionCorrections for Contact-Rich Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

力覚対応VLAポリシーを、参照行動と力補正に分解して蒸留するフレームワークを提案し、接触の多いタスクで成功率を54.4%から82.2%に向上させた。

詳しい要約

1. どんなもの?

- ForceDelta-VLAは、contact-rich manipulationのためのforce-aware VLAポリシー。 - タスクレベルの動作と接触依存の調整を単一のaction predictionで行う従来手法に対し、参照actionと補正を分離するcorrection-distillationフレームワーク。 - 軽量ポリシーが参照actionを最近のforce historyとrobot stateで調整し、参照action更新間の接触変化に応答する。

2. 先行研究と比べてどこがすごい?

- 従来のforce-aware VLAはaction predictionを参照actionと補正に分解する明示的ラベルを持たない。 - ForceDelta-VLAはfrozen teacherのforce-conditionedモードとforce-agnosticモードのペア予測から明示的なforce-correction targetを構築。 - 9つのsingle-arm/bimanual contact-richタスクで平均成功率82.2%を達成し、ForceVLAベースラインの54.4%を大幅に上回る。 - 成功試行における平均ピーク接触力も約26%低減。

3. 技術・手法の肝は?

- frozen teacherのforce-conditionedおよびforce-agnosticモードのペア予測からforce-correction targetを構築。 - 参照actionの不一致と参照状態の変化に対応するdelay-correction targetを別途用意。 - トレーニングにはasynchronous schedule replayを使用し、実行時にキャッシュされたtask contextを利用。 - 軽量ポリシーが最近のforce historyとrobot stateで参照actionを調整し、完全なaction chunkを再生成せずに接触変化に応答。

4. どうやって有効だと検証した?

- 9つのsingle-armおよびbimanual contact-richタスクで評価。 - 平均成功率82.2%を達成し、ForceVLAベースラインの54.4%と比較。 - Stage-1 Temporal Teacherの直接実行は70.6%の成功率。 - 成功試行における平均ピーク接触力がForceVLA比で約26%減少(両プラットフォーム)。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- ForceVLA(ベースラインとして比較) - Force-aware VLAポリシー(関連手法) - Vision-Language-Action (VLA) policies(同分野の定番)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ju Dong, Yu Fu, Jian Chen, Yimeng Liu, Haocheng Zhao, Lei Zhang, Kaixin Bai, Liding Zhang, Diwen Zheng, Alois Christian Knoll, Angela P. Schoellig, Jianwei Zhang

分類: cs.RO

原文アブストラクト

Force-aware Vision-Language-Action (VLA) policies improve contact-rich manipulation, but typically combine task-level motion and contact-dependent adjustment in a single action prediction. Demonstrations provide no explicit labels for decomposing that prediction into a reusable reference action and a correction. We present ForceDelta-VLA, a correction-distillation framework that constructs an explicit force-correction target using paired predictions from a frozen teacher's force-conditioned and learned force-agnostic modes. A separate delay-correction target accounts for reference-action mismatch and the change in reference state. Training uses asynchronous schedule replay with the cached task context available during execution. The resulting lightweight policy adjusts the reference actions using recent force history and robot state, responding to contact changes between reference-action updates without regenerating complete action chunks. Across nine single-arm and bimanual contact-rich tasks, ForceDelta-VLA achieves an 82.2% mean success rate, compared with 54.4% for the original ForceVLA baseline. Direct execution of our Stage-1 Temporal Teacher achieves 70.6%. Relative to ForceVLA, the complete system reduces mean peak contact force over successful trials by approximately 26% on both platforms.

関連論文

PR本紙発行元 EmplifAI