日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
sim2realarXiv:2610.04372

人間からロボットへの視覚適応のためのフレームレベル時間アラインメント

Frame-Level Temporal Alignment for Human-to-Robot Visual Adaptation

シェア:XThreadsFacebookLINEはてブBluesky

人間動画で事前学習した視覚エンコーダをロボット操作に適応させる際、実行速度や非キーフレームの割合の違いによる対応ずれを、フレームレベルのアノテーションなしで時間的先験を用いて解決するFLTAを提案した。

詳しい要約

1. どんなもの?

- 人間動画で事前学習した視覚表現をロボットマニピュレーションへ転移する枠組み。 - 人間とロボットのデモ間の対応関係を学習する必要がある。 - 実行速度や非キーフレームの割合の違いにより、同じ相対時刻のフレームが異なるタスク段階を表す問題に対処。 - Frame-Level Temporal Alignment (FLTA) を提案。 - 2つの時間的事前知識を用いて視覚エンコーダを適応。

2. 先行研究と比べてどこがすごい?

- フレームレベルの対応アノテーションなしで共有タスク進捗表現を学習。 - 異なる相対時間位置のフレーム同士をマッチング可能。 - 実行速度の違いを許容する点が先行研究と異なる。 - ResNet-50とViTで平均シミュレーション成功率が最良ベースライン比で46.93%、65.96%の相対改善。 - 実世界タスクでも高い成功率。 - 有効な人間-ロボット適応は更新パラメータ数より選択するパラメータに依存することを示唆。

3. 技術・手法の肝は?

- 2つの時間的事前知識を利用。 - グローバル進捗事前知識: 正規化時間位置と視覚類似度を組み合わせ、ソフト対応ターゲットを構築。 - ローカル時間順序事前知識: 後退遷移にペナルティを与え、停滞と可変の前進速度を許容。 - これにより実行速度差に対応。 - フレームレベル対応アノテーション不要。

4. どうやって有効だと検証した?

- ResNet-50とViTエンコーダで評価。 - 平均シミュレーション成功率で最良ベースライン比46.93%、65.96%の相対改善。 - 実世界マニピュレーションタスクで高い成功率を達成。 - プロジェクトページ: https://rtx5090ultra.github.io/FLTA-Project-Page/

5. 議論はある?

- 有効な人間-ロボット適応は更新パラメータ数より選択するパラメータに依存することを示唆。 - その他の議論や限界は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 同分野の定番として human-to-robot visual adaptation、visual representation transfer、robot manipulation に関する研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xizhe Zhang, Jingfeng Zhang, Zirun Zhou, Hong Jia

分類: cs.RO, cs.LG

原文アブストラクト

Transferring visual representations pretrained on human videos to robot manipulation requires learning reliable correspondences between human and robot demonstrations. However, paired demonstrations can differ in execution rate and in the proportion of non-key frames that do not directly reflect task progress. Frames at the same relative timestamp may therefore represent different task stages, which can cause correspondence learning to fail. To address these issues, we propose Frame-Level Temporal Alignment (FLTA), a framework that uses two temporal priors to adapt visual encoders pretrained on human videos for robot manipulation. It learns shared task-progress representations without frame-level correspondence annotations, allowing frames at different relative temporal positions to match. A global progress prior combines normalized temporal positions with visual similarity to construct a soft correspondence target. A local temporal order prior penalizes backward transitions while allowing stays and varying forward rates to accommodate execution-rate differences. With ResNet-50 and ViT encoders, our method achieves relative improvements of 46.93% and 65.96%, respectively, over the best baselines in average simulation success rates and achieves higher task success rates on real-world manipulation tasks. These results also suggest that effective human-robot adaptation depends less on the number of parameters updated than on which parameters are selected. Our project page is available at https://rtx5090ultra.github.io/FLTA-Project-Page/.

関連論文

PR本紙発行元 EmplifAI