日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.10534

RoboPrompt:スパースな人間入力による直感的ロボット方策ステアリング

RoboPrompt: Intuitive Robot Policy Steering with Sparse Human Input

シェア:XThreadsFacebookLINEはてブBluesky

描いた軌跡や目標点などの簡潔な人間入力から行動の下書きを作り、ベース方策の拡散・フローマッチング過程で洗練させることで、方策を再学習せずに操舵可能にする汎用システム。

詳しい要約

1. どんなもの?

RoboPromptは、描いた軌跡・目標点・粗い方向指示といった直感的でsparseな人間入力により、ロボットのend-to-end policyの挙動を誘導する汎用かつ軽量なpolicy steeringシステム。 - 人間の意図変換を基盤policyから分離する設計。 - 再利用可能なモジュールが人間のguidanceをaction draftsへ変換。 - それをbase policyのdiffusionまたはflow-matching dynamicsで洗練。 - base policyのアーキテクチャ変更やsteerability向けfine-tuningを不要とする。

2. 先行研究と比べてどこがすごい?

先行研究と比べた利点は以下の通り。 - 従来のshared-autonomyはteleoperationによる修正を可能にするが、専用ハードウェアと操作者訓練がscale展開を妨げる。 - 人間guidanceを追加のpolicy入力とする手法は、アーキテクチャ変更とsteerability専用訓練を要し、適用可能なpolicyが限られる。 - RoboPromptはsparseで直感的な入力のみでsteeringを実現し、base policyの変更やfine-tuningを不要とする。 - 複数のpolicy(Diffusion Policy, π_{0.5}, FastWAM)にまたがって有効性を示す。

3. 技術・手法の肝は?

技術の肝は以下の通り。 - 人間の意図変換を基盤policyからdecoupleする点。 - 再利用可能なモジュールが人間guidanceをaction draftsへ変換。 - そのdraftsをbase policyのdiffusionまたはflow-matching dynamicsで洗練。 - noise spaceでaction生成を制御することで、人間の意図とpolicy priorのバランスを取る。 - base policyのアーキテクチャ変更やsteerability向けfine-tuningを必要としない。

4. どうやって有効だと検証した?

有効性の検証は以下の通り。 - Diffusion Policy, π_{0.5}, FastWAMを対象にsteeringの有効性を実験で示す。 - steered rolloutsをDAggerによるonline policy improvementに利用。 - 2-3ラウンドの反復後、π_{0.5}では3タスクで平均成功率が15.5%向上。 - Insert Breadタスクでは3つのpolicy(Diffusion Policy, π_{0.5}, FastWAM)で平均成功率が21.3%向上。 - 平均人間介入回数はそれぞれ44.0%(2.86から1.60)と81.9%(2.60から0.47)減少。

5. 議論はある?

議論の詳細は要旨からは不明。 - 限界や失敗事例、計算コスト、他タスクへの一般化については要旨に記述がない。 - 人間入力の種類や量に対する感度、policy間の性能差の要因も要旨からは不明。

6. 次に読むべき論文は?

要旨で参照・比較されている研究や関連手法は以下の通り。 - Diffusion Policy - π_{0.5} - FastWAM - DAgger - shared-autonomyによるteleoperationベースの修正手法 - 人間guidanceを追加のpolicy入力とする手法 - imitation learningによるend-to-end robot policy

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yanwen Zou, Chenyang Shi, Guoxuan Xu, Wenye Yu, Wendi Chen, Ye Pan, Cewu Lu, Chuan Wen

分類: cs.RO

原文アブストラクト

End-to-end robot policies trained through imitation learning remain constrained by limited data diversity, making reliable zero-shot deployment in real-world settings challenging. Shared-autonomy methods enable human correction through teleoperation, but specialized hardware and operator training hinder deployment at scale. Other approaches incorporate human guidance as additional policy inputs, often requiring architectural changes and dedicated training for steerability, which limits their applicability across policies. We present RoboPrompt, a general-purpose, lightweight robot policy steering system that enables users to guide policy behavior through intuitive, sparse inputs, including drawn traces, target points, and coarse directional instructions. RoboPrompt decouples human-intention translation from the underlying policy: a reusable module converts human guidance into action drafts, which are refined through the diffusion or flow-matching dynamics of the base policy. By controlling action generation in noise space, RoboPrompt balances human intent with the policy prior without modifying the base policy architecture or fine-tuning it for steerability. Experiments demonstrate effective steering across Diffusion Policy, $π_{0.5}$, and FastWAM. We further use steered rollouts for online policy improvement through DAgger. After 2-3 rounds of iteration, average success rates increase by 15.5\% for $π_{0.5}$ across three tasks and by 21.3\% across three policies(Diffusion Policy, $π_{0.5}$, FastWAM) on the Insert Bread task, while average human intervention counts decrease by 44.0\% (2.86 to 1.60) and 81.9\% (2.60 to 0.47), respectively.

関連論文

PR本紙発行元 EmplifAI