日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.38164

Rho: 効率的に適応可能なVLAモデルの基盤

Rho: A Foundation for Efficiently Adaptable VLA Models

シェア:XThreadsFacebookLINEはてブBluesky

双腕マニピュレーション向けのオープンウェイトVLAモデル群「Rho」を提案し、3つの実機での少量データ適応と、15エピソード程度の修正フィードバックによるオンライン適応を実現した。

詳しい要約

1. どんなもの?

- 汎用フィジカルAI向けのopen-weights VLAモデル群「Rho」 - 双腕マニピュレーション用に設計 - 研究・産業向けdual-armロボット3機種(YAM Box, UR AI Trainer, FR3 Duo)に対応 - データ少量での下流タスク適応を狙う - 基盤モデルとembodiment別checkpointを公開 - オンライン適応機能も内蔵

2. 先行研究と比べてどこがすごい?

- 既存のopen-weights VLAと比較して同等以上 - 評価したタスク・embodiment・baseline全体で最強の総合性能 - embodiment midtrainingが下流適応を改善することを実験で示す - 15件程度の修正episodeで軽量モジュールを適応可能 - offline finetuning分布の周縁状況に対応 - 具体的な先行研究名は要旨からは不明

3. 技術・手法の肝は?

- action-expertアーキテクチャと学習レシピを体系的にablate - embodiment midtrainingを採用 - オンライン適応: - 内部latent policyがcorrective feedbackから学習 - 観測条件付きnoise入力を選択 - frozen flow-matching action expertへ入力 - 軽量モジュールのみ適応

4. どうやって有効だと検証した?

- 制御されたsimulation実験 - 物理ロボット実験 - 3 embodiment(YAM Box, UR AI Trainer, FR3 Duo)で評価 - 既存open-weights VLAおよび複数baselineと比較 - オンライン適応は15 corrected episodesで検証

5. 議論はある?

- embodiment midtrainingの有効性を主張 - オンライン適応によりoffline finetuning分布の周縁状況へ対応可能と主張 - 限界・失敗事例・計算コスト・安全性などの議論は要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている具体的な先行研究は不明 - 同分野の関連手法としてOpenVLA, RT-2, Octo, π0, diffusion policy, flow matchingなどが候補

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Rho Team, Simran Bagaria, Daphne Chen, Dean Fortier, Jianlong Fu, Michael Harrison, Tess Hellebrekers, Neel Joshi, Andrey Kolobov, Dalton Moore, Galen Mullins, Michael Murray, Eduardo Salinas, Reuben Tan

分類: cs.RO

原文アブストラクト

General-purpose physical AI models must combine broad visual and linguistic capabilities with precise control across robot embodiments and efficient adaptation to downstream tasks. We introduce Rho, a family of open-weights VLA models for bimanual manipulation designed for data-light task adaptation on 3 embodiments representative of dual-arm robots across research labs and the industry -- YAM Box, UR AI Trainer, and FR3 Duo. We systematically ablate Rho's action-expert architecture and training recipe, and show in controlled simulation and physical-robot experiments that embodiment midtraining improves downstream adaptation. The resulting Rho variants for YAM Box, UR AI Trainer, and FR3 Duo match or outperform existing open-weights VLAs and achieve the strongest overall performance across the tasks, embodiments, and baselines evaluated in this report. We further demonstrate the Rho model family's built-in capacity for online adaptation: an internal latent policy learns from corrective feedback to select observation-conditioned noise inputs for the frozen flow-matching action expert. With as few as 15 corrected episodes, adapting this lightweight module enables Rho to handle task situations at the fringe of its offline finetuning distribution. Together, these results position Rho as both a strong general-purpose robotic manipulation model and a practical foundation for adaptation. We release the base Rho model and the embodiment-specific checkpoints to facilitate Rho's deployment in research experiments and practical industrial use cases.

関連論文

PR本紙発行元 EmplifAI