日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.03615

フロンティア推論とロボット実行の橋渡し:自律的デモ生成から密な言語監督まで

Bridging Frontier Reasoning and Robot Execution: From Autonomous Demonstration Generation to Dense Language Supervision

シェア:XThreadsFacebookLINEはてブBluesky

フロンティアモデルでデモを自動生成し高速なローカル方策を訓練するとともに、プリミティブ・原子・複合の3階層にわたる密な言語監督を導入することで、長期的タスクにおける指示追従の信頼性を高めた研究。

詳しい要約

1. どんなもの?

frontier modelの推論を低遅延なロボット実行に接続する2つの相補的アプローチを研究。 - 第1: frontier modelで自律的にdemonstrationを生成し、高速なlocal policyの訓練を補完。 - 第2: primitive/atomic/compositeの3階層にわたるdense language supervisionを導入。 - 両者をcrosswordタスクで統合評価。

2. 先行研究と比べてどこがすごい?

frontier modelは少数demonstrationで操作可能だが推論遅延が実時間制御を制限。 - 本研究はfrontier reasoningと低遅延local executionを橋渡し。 - 生成時間とコストが成功例の蓄積で減少する点を示唆。 - dense language supervisionで指示追従の信頼性を向上。

3. 技術・手法の肝は?

第1の橋: frontier modelでdemonstrationを自律生成。 - in-context例に物理エラーからの回復を示すcorrective demonstration segmentsを追加し生成信頼性を改善。 - 展開時はharnessがfrontier生成指示と高速local policyを統合。 - 第2の橋: primitive/atomic/compositeの3階層で多面的記述を持つdense language supervision。

4. どうやって有効だと検証した?

RoboCasa 365とBEHAVIOR-1Kのlong-horizonタスクで検証。 - 実行進行に伴い指示が変化する設定。 - oracleとfrontier-model instructorの両方でcombined supervisionが最高性能。 - crosswordタスクで意味計画と操作を固定時間予算内で統合評価。

5. 議論はある?

低遅延実行をlocal policyに委ねると、ボトルネックはfrontier modelの多様な指示に確実に従う能力へ移行。 - 生成時間とコストが成功例蓄積で減少する観察は効率的データ収集への道を示唆。 - 両橋の相補性を支持する結果。

6. 次に読むべき論文は?

RoboCasa 365、BEHAVIOR-1K、crosswordタスク。 - frontier model、local policy、dense language supervision関連の研究。 - 具体的な参照論文は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Bosung Kim, Alexander Trevithick, Ruiyi Wang, Prithviraj Ammanabrolu

分類: cs.RO

原文アブストラクト

Recent advances in frontier models enable robot manipulation from only a few demonstrations, but high inference latency limits their use for real-time robot control. To bridge this gap, we study two complementary approaches that connect frontier reasoning with low-latency local execution. First, we use a frontier model to autonomously generate demonstrations that supplement human demonstrations for training a fast local policy. We augment its in-context examples with corrective demonstration segments that show how to recover from physical errors, improving generation reliability. Generation time and cost decrease as successful examples accumulate in context, suggesting a path toward more efficient data collection. At deployment, a harness combines frontier-generated instructions with a fast local policy, enabling efficient execution while preserving the frontier model's ability to guide and correct actions. With low-latency execution delegated to the local policy, the bottleneck shifts to its capacity to reliably follow the frontier model's diverse instructions. Our second bridge introduces dense language supervision across three nested granularities---primitive, atomic, and composite---with multi-aspect descriptions at each level. Across long-horizon tasks in RoboCasa 365 and BEHAVIOR-1K, where instructions change as execution progresses, the combined supervision achieves the highest performance under both oracle and frontier-model instructors, demonstrating more reliable instruction following through the policy's language interface. Finally, we evaluate both bridges together on a crossword task that combines semantic planning and manipulation within a fixed time budget. These results support autonomous demonstration generation and dense language supervision as complementary components for connecting frontier reasoning to low-latency local execution.

関連論文

PR本紙発行元 EmplifAI