日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
操作学習arXiv:2609.03715v1

MINERVA: 操作ポリシーはどれだけ小さくできるか、LIBEROを解くために

MINERVA: How Small Can a Manipulation Policy Be and Still Solve LIBERO?

シェア:XThreadsFacebookLINEはてブBluesky

LIBEROベンチマークを解くのに必要な最小のモデル容量を探るため、わずか0.54MパラメータのコンパクトなポリシーMINERVAを提案し、95.1%の成功率を達成。大規模VLAモデルと同等性能を桁違いに少ないパラメータで実現した。

詳しい要約

1. どんなもの?

MINERVAは、LIBERO操作ベンチマークにおけるタスク固有の容量下限を測定するために設計された、意図的にコンパクトな視覚運動ポリシーのファミリーである。0.54Mパラメータのポリシーが、4つの標準LIBEROスイートでの2,000ロールアウトにおいて95.1%の平均成功率を達成し、LeRobot π0.5の報告結果よりわずか2.4ポイント低いものの、パラメータ数は7,700倍少ない。性能は約1Mパラメータで飽和し、0.25M未満で崩壊する。

2. 先行研究と比べてどこがすごい?

従来のVLAモデルは数十億パラメータでLIBEROを支配しているが、MINERVAは極めて小さなポリシーで同等の性能を達成し、ベンチマークの実際の容量要求を明らかにした。また、Flow Matchingが直接L1回帰に対して優位性がないこと、標準的な命令条件付けが主にタスクIDの記憶に依存していることなどを示し、既存の大規模モデルの有効性に疑問を投げかけている。

3. 技術・手法の肝は?

手法の肝は、アーキテクチャ、トレーニング、推論の広範なスイープを通じて、アクションチャンク長と視覚容量のみが一貫してトレーニングシードの±1ポイント帯を超える影響を持つことを特定した点。また、タスクID置換プローブを用いて、標準的なLIBERO命令条件付けが記憶されたタスクの選択に過ぎないことを示した。

4. どうやって有効だと検証した?

4つの標準LIBEROスイートで2,000ロールアウトを実施し、平均成功率95.1%を達成。さらに、89タスクのLIBERO-90で94.6%の成功率を確認。LIBERO-Plus摂動では46〜56%に低下し、フォトメトリックシフトに対するロバスト性がほぼゼロであることを示した。また、ラップトップCPU上で制御ステップごとに5〜9msの再計画時間を達成し、SmolVLAより113倍、π0.5より1,400倍高速であることを実証した。

5. 議論はある?

議論として、性能が1Mパラメータで飽和し、0.25M未満で崩壊することから、LIBEROのタスク固有の容量下限が存在することが示唆される。また、標準的な命令条件付けがタスクIDの記憶に依存しているという発見は、ベンチマークの評価方法に疑問を投げかける。さらに、LIBERO-Plusでのロバスト性の低さは、実世界展開への課題を示している。

6. 次に読むべき論文は?

要旨で参照されている研究は、LeRobot π0.5、SmolVLA、LIBERO、LIBERO-Plusである。次に読むべき論文としては、これらの関連研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Kohei Sendai, Tatsuya Matsushima, Yusuke Iwasawa

分類: cs.RO

原文アブストラクト

Vision-language-action (VLA) models with billions of parameters now dominate the LIBERO manipulation benchmark, but the model capacity actually required by the benchmark remains unclear. We introduce MINERVA (MINimal Efficient Robotic Vision-Action policy), a family of deliberately compact visuomotor policies designed to measure this task-specific capacity floor. A 0.54M-parameter policy achieves 95.1% average success over 2,000 rollouts on the four standard LIBERO suites, only 2.4 points below the reported LeRobot $π_{0.5}$ result despite using 7,700$\times$ fewer parameters. Performance saturates near 1M parameters and collapses below 0.25M. Across broad architectural, training, and inference sweeps, only action-chunk length and vision capacity consistently exceed a $\pm$1-point training-seed band. Flow matching provides no detectable advantage over direct L1 regression across three seeds, while regression is up to 3.8$\times$ faster on GPU. A task-ID permutation probe shows that standard LIBERO instruction conditioning primarily selects among memorized tasks: changing only the task-ID mapping reduces success to near chance. The same recipe achieves 94.6% success across 89 LIBERO-90 tasks, while LIBERO-Plus perturbations reduce performance to 46--56%, with near-zero robustness to photometric shifts. The 0.54M policy replans every control step in 5--9 ms per chunk on a laptop CPU, 113$\times$ faster than SmolVLA and 1,400$\times$ faster than $π_{0.5}$, without a GPU. These results establish a first empirical estimate of LIBERO's task-specific capacity floor and motivate capacity-aware design and distillation for deployment-efficient robot policies.

関連論文