日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.12172

空間言語モデリングによる方策学習と状態予測の統合

Unifying Policy Learning and State Prediction through Spatial Language Modeling

シェア:XThreadsFacebookLINEはてブBluesky

シーン輪郭・目標・行動・将来状態を離散座標と意味トークンの共通語彙で表現し、単一の自己回帰Transformerで行動生成と行動条件付き状態予測を同時に学習する手法を提案。Push-Tと実機で評価。

詳しい要約

1. どんなもの?

- 本論文は、Spatial Language Modelingを提案する。 - シーン輪郭、目標、行動ターゲット、将来状態を離散座標と意味トークンの共有語彙で表現する。 - タスク固有の文法で空間シーケンスを構成し、単一の自己回帰Transformerが次トークン予測目的で行動生成と行動条件付き状態予測を学習する。 - ランダムプレイ遷移事前学習後、専門家デモで行動と状態を共同訓練する。 - 制御時は実行可能な行動ターゲットのみをデコードし、観測状態で履歴を更新する。

2. 先行研究と比べてどこがすごい?

- 行動がシーン形状をどう変えるかの学習が、目標指向マニピュレーションの補完的監督になる点を示す。 - 従来の行動生成と状態予測を別々に扱う手法と異なり、共有語彙と文法で統合する。 - 単一の自己回帰Transformerで両タスクを共通の次トークン目的で学習する点が新しい。 - 実ロボットで評価したポリシーベースラインよりタスク成功率とターゲットカバレッジが高い。 - 共同行動・状態シーケンス訓練とランダムプレイ事前学習の利点をアブレーションで示す。

3. 技術・手法の肝は?

- シーン輪郭、目標、行動ターゲット、将来状態を離散座標と意味トークンの共有語彙で表現する。 - タスク固有の文法がこれらを空間シーケンスに組織化する。 - 自己回帰Transformerが次トークン予測目的で行動生成と行動条件付き状態予測を学習する。 - 訓練はランダムプレイ遷移事前学習後、専門家デモで行動と状態を共同訓練する。 - 事前学習では記録された行動座標が後続状態予測を条件付け、予測損失から除外される。 - 制御時は実行可能な行動ターゲットのみをデコードし、新観測状態で履歴を更新する。

4. どうやって有効だと検証した?

- Push-Tのシミュレーションと実ロボットで評価した。 - シミュレーションで競争力のある性能を達成した。 - 実ロボットで評価したポリシーベースラインよりタスク成功率とターゲットカバレッジが高かった。 - 訓練アブレーションで共同行動・状態シーケンスによる制御改善と、ランダムプレイ事前学習によるさらなる利得を示した。 - 与えられた行動軌跡に対し、同じモデルが連続するシーン状態を予測し、押すことの幾何学的効果を捉えた。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- Push-T - 自己回帰Transformer - ランダムプレイ事前学習 - 行動条件付き状態予測 - 目標指向マニピュレーション

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Minye Wu, Zehao Wang, Tinne Tuytelaars

分類: cs.RO, cs.AI

原文アブストラクト

Learning how actions change scene geometry can provide complementary supervision for goal-directed manipulation. We introduce Spatial Language Modeling, which represents scene contours, goals, action targets, and future states with a shared vocabulary of discrete coordinates and semantic tokens. A task-specific grammar organizes these elements into spatial sequences, allowing one autoregressive Transformer to learn action generation and action-conditioned state prediction through a common next-token objective. We train the model from scratch using random-play transition pretraining followed by joint action and state training on expert demonstrations. During pretraining, recorded action coordinates condition subsequent state predictions and are excluded from the prediction loss. During control, the model decodes only executable action targets and updates its history with newly observed states. We evaluate the approach on Push-T in simulation and on a real robot. The model achieves competitive simulation performance and higher task success and target coverage than the evaluated real-robot policy baselines. Training ablations show improved control with joint action and state sequences, with further gains from random-play pretraining. Given supplied action trajectories, the same model also predicts successive scene states, capturing the geometric effects of pushing.

関連論文

PR本紙発行元 EmplifAI