日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.07785

Attacca: 状態連続性下での長期身体エージェントのための目標指向制御

Attacca: Goal-Directed Control under State Continuity for Long-Horizon Embodied Agents

シェア:XThreadsFacebookLINEはてブBluesky

長期タスク実行中の状態連続性に対応するため、検索から相互作用までの軌跡で視覚目標条件付きポリシーを訓練するAttaccaを提案。目標画像を実行環境から切り離し、行動フェーズ条件付けを導入。

詳しい要約

1. どんなもの?

- 長期的な embodied agent の連続タスク実行を対象とした研究。 - 既存の visual goal-conditioned policy は目標が可視な単発相互作用で評価され、連続実行の状態継続を扱えない。 - 各タスクは前タスクの状態から始まり、位置・姿勢の変化、世界の改変、視野外の目標が生じる。 - 提案手法 Attacca は search-to-interact の完全軌跡で policy を訓練する。 - Minecraft の短・長 horizon タスクで評価。

2. 先行研究と比べてどこがすごい?

- 従来の visual goal-conditioned policy は目標が既に見える単発相互作用を前提とする。 - Attacca は目標画像を実行環境から切り離し、状態継続下の長期的タスク実行を扱う。 - 最強 baseline に対し clean success で 1.7-2.4x 改善。 - 長 horizon タスクで最大 7x の改善を達成。 - 目標を現在の観測に ground できない場合の失敗を緩和。

3. 技術・手法の肝は?

- context-decoupled goal sampling: 各デモに別世界のクラス互換な masked goal image を対応させ、直接の scene/pose 対応を除去。 - target-mask prediction head により dense な current-view grounding を学習。 - action imitation を超える補助 supervision を提供。 - behavioral-phase conditioning: Search, Approach, Interact の段階を区別し、実行進行に応じて制御を適応。 - 完全な search-to-interact 軌跡で訓練。

4. どうやって有効だと検証した?

- Minecraft の複数の短・長 horizon embodied タスクで評価。 - clean success 39.0-47.5% を達成し、最強 baseline を 1.7-2.4x 上回る。 - 長 horizon タスクで 54%, 30%, 28% の completion を達成。 - 最大 7x の改善を確認。 - 具体的な ablation や評価プロトコルの詳細は要旨からは不明。

5. 議論はある?

- 状態継続下での長期的タスク実行の難しさを指摘。 - 目標が視野外にある場合の grounding 問題に対処。 - 提案手法の有効性を示す一方、限界や失敗事例の議論は要旨からは不明。 - 他の環境やタスクへの一般化可能性は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として visual goal-conditioned policy, goal-conditioned imitation learning, behavioral cloning, masked goal image を用いる手法が挙げられる。 - 同分野の定番として Minecraft における embodied agent 研究 (MineRL, VPT など) や long-horizon task の hierarchical RL が考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Gyusik Seo, Jaehong Yoon

分類: cs.AI

原文アブストラクト

A central capability of embodied agents is to accomplish complex objectives through sequences of interdependent tasks. Yet existing visual goal-conditioned policies underlying these agents are typically evaluated on isolated interactions where the target is already visible, and thus do not capture the conditions that arise during continuous long-horizon task execution. In such settings, each task begins from the state left by the previous one: the agent may end at a different position and orientation, the world may have been modified, and the next interaction target may lie outside the current field of view. As a result, agents relying on such policies may struggle to proceed to the next task when they cannot ground their target in the current observation. To address this challenge, we propose Attacca, a new approach that trains visual goal-conditioned policies on complete search-to-interact trajectories using goal images decoupled from the execution environment. Attacca uses context-decoupled goal sampling to pair each demonstration with a class-compatible masked goal image from another world, removing direct scene and pose correspondence. It learns dense current-view grounding through a target-mask prediction head, providing auxiliary supervision beyond action imitation. We further introduce behavioral-phase conditioning that teaches the policy to distinguish Search, Approach, and Interact stages and adapt its control as execution progresses. We evaluate Attacca on multiple short- and long-horizon embodied tasks in Minecraft. Our method achieves 39.0-47.5% clean success, improving over the strongest baseline by 1.7-2.4x. On long-horizon tasks, it attains 54%, 30%, and 28% completion, yielding up to a 7x improvement.

関連論文

PR本紙発行元 EmplifAI