日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.08183

コンパクトなロボットポリシーに必要なのは細粒度の視覚表現

Compact Robot Policies Need Fine-Grained Visual Representations

シェア:XThreadsFacebookLINEはてブBluesky

視覚表現に注目し、VLMや動画生成事前学習を使わない48.9Mパラメータの小型ポリシーCoRPを構築。大規模モデルに匹敵する性能を達成し、事前学習初期化やトークン圧縮が性能を左右することを示した。

詳しい要約

1. どんなもの?

多タスクmanipulation policyの性能差がarchitecture・scale・pretrained priorsのどれに由来するか切り分けられない問題に対し、visual representationが主因だと主張する研究。 - CoRP (Compressed Representation Policy) を提案。 - 48.9M parametersの意図的にcompactなpolicy。 - vision-language modelもvideo-generative priorも使わない。 - representation extractorとflow-matching action generatorに分解。 - 性能: LIBEROで97.0%、RoboTwin 2.0 Clean/Randomizedで75.78%/73.36%。 - これは40.9-163.6倍大きいシステムに匹敵する。

2. 先行研究と比べてどこがすごい?

従来のmulti-task manipulation policy比較はarchitecture・scale・pretrained priorsが同時に異なり、性能要因を特定できなかった。 - 本研究はaction generatorを固定し、extractorの性質を1つずつ変える統制実験で要因を分離。 - その結果、性能の大半はvisual representationに由来し、parameter scaleやgenerative priorsはほぼ偶発的だと主張。 - 40.9-163.6倍大きいシステムに匹敵する性能を48.9M parametersで達成。

3. 技術・手法の肝は?

CoRPはrepresentation extractorとflow-matching action generatorにfactorizeされるcompact policy。 - 統制実験でextractorの性質を1つずつ変化させる。 - pretrained initialization: random ViT-S/14はLIBEROで78.1%、ImageNet ResNet-34は74.5%に低下。 - pretrainingだけでは不十分で、encoderをfreezeすると19.8ポイント低下。 - compression: 各viewを48 tokensにresamplingすると全patch tokensより良い (97.0% vs 83.2%)。 - variational information bottleneckはhard token budgetより悪く、LIBERO-Goalを95.8%から33.0%に低下させる。 - language conditioningはobservationがgoalを曖昧にする場合のみ寄与 (LIBERO-…

4. どうやって有効だと検証した?

LIBEROとRoboTwin 2.0 (Clean/Randomized) で評価。 - CoRPはLIBERO 97.0%、RoboTwin 2.0 Clean 75.78%、Randomized 73.36%。 - action generatorを固定しextractorの性質を1つずつ変えるablationを実施。 - pretrained initialization、freezing、compression、variational information bottleneck、language conditioningの効果を測定。 - これらの統制実験でvisual representationの寄与を検証。

5. 議論はある?

compact policyが機能する条件として、representationがpretrained・task-adapted・compressedであることを主張。 - parameter scaleやgenerative priorsはほぼ偶発的と議論。 - language conditioningの寄与はobservationがgoalを曖昧にする場合に限られ、RoboTwin 2.0ではobservationが明確なため除去するとわずかに改善。 - variational information bottleneckがinstruction-dependent token selectionを抑制し性能を大きく下げる点を議論。 - その他の限界や議論は要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究・手法を挙げる。 - LIBERO benchmark。 - RoboTwin 2.0 benchmark。 - ViT-S/14。 - ImageNet ResNet-34。 - flow-matching action generator。 - variational information bottleneck。 - vision-language model。 - video-generative prior。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Nanhe Chen, Runqiu Yang, Jiawei Tang, Sichao Liu, Yuquan Wang

分類: cs.RO, cs.AI, cs.LG

原文アブストラクト

Multi-task manipulation policies differ in architecture, scale, and pretrained priors all at once, so published comparisons cannot attribute performance to any single component. We argue that most of it comes from the visual representation, and that parameter scale and generative priors are largely incidental. To test this, we build CoRP (Compressed Representation Policy), a deliberately compact policy (48.9M parameters, no vision-language model and no video-generative prior) that factorizes into a representation extractor and a flow-matching action generator. It reaches 97.0% on LIBERO and 75.78%/73.36% on RoboTwin 2.0 Clean/Randomized, matching systems 40.9-163.6x larger. Holding the action generator fixed, we then vary one extractor property at a time. Pretrained initialization is decisive: a random ViT-S/14 drops to 78.1% and an ImageNet ResNet-34 to 74.5% on LIBERO. Pretraining alone is not enough, as freezing the encoder costs 19.8 points. Compression matters as much: resampling each view to 48 tokens beats passing all patch tokens (97.0% vs 83.2%), and a variational information bottleneck over those tokens is worse than a hard token budget, cutting LIBERO-Goal from 95.8% to 33.0% by suppressing the instruction-dependent token selection the policy relies on. Language conditioning contributes only where the observation leaves the goal ambiguous (LIBERO-Goal: 9.2% to 95.8%), while on RoboTwin 2.0, where observations are unambiguous, removing it slightly improves success. Therefore, we argue that a compact policy works when its representation is pretrained, task-adapted, and compressed. Project page: https://corp-policy.github.io/

関連論文

PR本紙発行元 EmplifAI