日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.23968

Opt2VLA: 接触の多いヒューマノイド全身マニピュレーションのための力認識型視覚-言語-行動モデル

Opt2VLA: Force-Aware Vision-Language-Action for Contact-Rich Humanoid Whole-Body Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

ヒューマノイドの全身操作において、VLAモデルが運動目標と接触力の指令を同時に予測し、強化学習ベースの全身制御器で追従する力認識型フレームワークを提案。軌道最適化で接触を考慮した訓練データを生成し、接触の多い3タスクで力の調整精度が向上することを示した。

詳しい要約

1. どんなもの?

- 人型ロボットの接触を伴う全身マニピュレーションのための force-aware VLA フレームワーク Opt2VLA を提案。 - VLA-to-control インターフェースに明示的な force command を導入し、単一の multi-task VLA policy が幾何学的な motion goal と連続的な contact-force reference を同時に予測。 - 予測された目標はタスク固有の RL-based whole-body controller によって追従される。 - 訓練データは whole-body trajectory optimization (TO) により、動的に実行可能かつ接触整合的な force reference 付きで生成。 - 3つの contact-rich な人型タスクで評価し、シミュレーションと実機で language-conditioned force modulation を実証。

2. 先行研究と比べてどこがすごい?

- 既存の VLA モデルは semantic planning や visuomotor control に有望だが、人型システムでは主に幾何学的な motion goal で行動を表現し、motion tracking に焦点を当てた whole-body controller に依存。 - そのため interaction force の明示的な推論や制御が限定的であった。 - Opt2VLA は VLA-to-control インターフェースに明示的な force command を導入することで、接触後は視覚観測が信頼できなくなる状況や、幾何学的に類似した動作でもタスク文脈に応じて異なる力レジームが必要な contact-rich タスクに対処。 - 明示的な force conditioning により、motion-only control よりも正確で一貫した力調整を可能にする点が優位。

3. 技術・手法の肝は?

- 単一の multi-task VLA policy が幾何学的 motion goal と連続的な contact-force reference を同時に予測。 - 予測された目標はタスク固有の RL-based whole-body controller が追従。 - スケーラブルで物理的に根拠のある監督を提供するため、whole-body trajectory optimization (TO) を明示的な force reference 付きで用いて、動的に実行可能かつ接触整合的な訓練データを生成。 - TO からの物理的に根拠のある torque supervision が force tracking の精度と安定性をさらに向上させる。

4. どうやって有効だと検証した?

- 3つの contact-rich な人型タスクで Opt2VLA を評価。 - 明示的な force conditioning が motion-only control よりも正確で一貫した力調整を可能にすることを示した。 - TO からの物理的に根拠のある torque supervision が force tracking の精度と安定性をさらに改善することを示した。 - 閉ループ評価により、シミュレーションと人型ハードウェア上で language-conditioned force modulation を実証。

5. 議論はある?

- 要旨からは不明。 - ただし、contact-rich タスクにおける力の明示的制御の重要性、接触後の視覚観測の信頼性低下、幾何学的に類似した動作でも異なる力レジームが必要となる点が動機として述べられている。 - 限界や今後の課題についての具体的な議論は要旨には記載されていない。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として vision-language-action (VLA) モデル、whole-body controller、reinforcement learning (RL)、trajectory optimization (TO) が挙げられる。 - 同分野の定番として、人型ロボットの whole-body manipulation や contact-rich manipulation に関する VLA 研究、force-aware control の論文を読むことが推奨される。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Fukang Liu, Yipu Chen, Jaehwi Jang, Danfei Xu, Zsolt Kira, Ye Zhao

分類: cs.RO

原文アブストラクト

Humanoid robots are expected to perform diverse human-level tasks in daily environments, many of which require precise regulation of interaction forces. While recent vision-language-action (VLA) models have shown promise for semantic planning and visuomotor control, existing humanoid systems primarily represent actions through geometric motion goals and rely on whole-body controllers focused on motion tracking, with limited explicit reasoning or control of interaction forces. This limitation is particularly relevant in contact-rich tasks, where geometrically similar motions may require different force regimes depending on the task context and where visual observations may become unreliable after contact. In this work, we present Opt2VLA, a force-aware VLA framework that introduces explicit force commands at the VLA-to-control interface for humanoid whole-body manipulation. A single multi-task VLA policy jointly predicts both geometric motion goals and continuous contact-force references, which are tracked by task-specific reinforcement learning (RL)-based whole-body controllers. To provide scalable and physically grounded supervision, we generate dynamically feasible and contact-consistent training data via whole-body trajectory optimization (TO) with explicit force references. We evaluate Opt2VLA on three contact-rich humanoid tasks and show that explicit force conditioning enables more accurate and consistent force regulation than motion-only control, while physically grounded torque supervision from TO further improves force tracking accuracy and stability. Closed-loop evaluations further demonstrate language-conditioned force modulation with Opt2VLA in simulation and on humanoid hardware.

関連論文

PR本紙発行元 EmplifAI