日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.09509

TERRA: 時間的効果表現と関係的整合による可搬な潜在行動の学習

TERRA: Learning Transportable Latent Actions through Temporal Effect Representation and Relational Alignment

シェア:XThreadsFacebookLINEはてブBluesky

視覚遷移から潜在行動を学習する際、遷移の「時間的効果」を表現し、異なる初期状態でも同じ意味を持つよう整合させる手法TERRAを提案。

詳しい要約

1. どんなもの?

- 視覚遷移から推論されるaction-like codes(latent actions)でロボットポリシーを監督する手法。 - コードが遷移から何を保持するか、異なる初期状態で再利用しても同じ意味を持つかが課題。 - TERRAは遷移をcompact temporal effect(net feature change + low-order within-window dynamics component)で記述し、連続latentを学習。 - 同じeffect spaceを再利用の参照とし、Effect-Anchored Transport (EAT)で他初期状態にデコードし、effectを元の観測にアンカー。 - これによりlatentは遷移元だけでなく、文脈間での振る舞いで形成される。

2. 先行研究と比べてどこがすごい?

- 従来のlatent actionsは再構成に依存し、latentとその元状態のペアしか観測しないため、異なる初期状態での意味の一貫性が未検証。 - TERRAは時間的効果表現と関係的アラインメントを統合し、両問題を同一のeffect spaceで解決。 - 凍結線形リーダーで、UniVLAやLAPA-styleベースラインより行動予測が正確。 - 視覚的distractor下での性能劣化が緩やか。 - 受容文脈が遠ざかっても、輸送された遷移がドナー行動に忠実。 - 同予算の対照実験で、利得の大部分がEATに由来することを示す。 - 同一事前学習規模でLIBERO平均成功率93.4%(UniVLAは91.8%)。

3. 技術・手法の肝は?

- 遷移をcompact temporal effectで表現:net feature changeと低次within-window dynamics componentの組み合わせ。 - この効果から連続latentを学習。 - Effect-Anchored Transport (EAT):latentを他の初期状態でデコードし、得られた効果を元の観測効果にアンカー。 - これによりlatentは文脈を跨いだ振る舞いで形成され、再構成に依存しない。 - 凍結線形リーダーで行動予測を評価。

4. どうやって有効だと検証した?

- 凍結線形リーダーを用いて行動予測精度を評価。 - UniVLAおよびLAPA-styleベースラインと比較。 - 視覚的distractor下での性能劣化を評価。 - 受容文脈が遠ざかる際の輸送遷移の忠実性を検証。 - 同予算の対照実験でEATの寄与を分析。 - 同一事前学習規模でLIBEROベンチマークを評価(平均成功率93.4% vs UniVLA 91.8%)。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- UniVLA - LAPA-style baseline - LIBEROベンチマーク - latent actionsに関する研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Tianxingjian Ding, Mubarak Shah, Yu Tian

分類: cs.CV, cs.LG, cs.RO

原文アブストラクト

Latent actions supervise robot policies with action-like codes inferred from visual transitions, and their usefulness hinges on two questions: what a code keeps from a transition, and whether it still means the same thing when reused in a different initial state. The first is a tension in time: an endpoint difference discards how motion unfolds, while the full sequence admits nuisance variation. The second is left open by reconstruction, which only ever observes a latent together with the state it came from. We argue that both questions can be answered in the same place. TERRA (Temporal Effect Representation and Relational Alignment) describes a transition by a compact temporal effect, its net feature change together with a low-order within-window dynamics component, and learns a continuous latent from this effect. The same effect space then serves as the reference for reuse: Effect-Anchored Transport (EAT) decodes a latent in other initial states and anchors the resulting effect to the one observed at its source, so that the latent is shaped by what it does across contexts rather than only by the transition it came from. With frozen linear readers, TERRA predicts actions more accurately than UniVLA and a LAPA-style baseline, degrades more slowly under visual distractors, and keeps transported transitions faithful to the donor action as the recipient context moves farther away; a same-budget control shows that these gains come largely from EAT. At matched pretraining scale, the complete system reaches 93.4% average success on LIBERO, compared with 91.8% for UniVLA.

関連論文

PR本紙発行元 EmplifAI