Sergey Levine
UC Berkeley
収録論文 314本 ・ フィジカルAI/ロボット学習
※arXiv著者名で収集。同姓同名の別人の論文が含まれる場合があります。
論文
- 待ち時間に学習する:推論遅延下での汎用ロボットポリシーのRLファインチューニングRL/遅延対応2026/8/24
推論遅延がRLのマルコフ性を壊す問題に対し、非同期推論と状態拡張で遅延を隠蔽しつつRL改善を可能にするフレームワークARLIを提案。シミュレーションと実機で有効性を実証した。
- 行動チャンキングがロボット制御における行動クローニングの性能を向上させる理由ロボット制御2026/8/3
行動チャンキングの性能向上の理由を実験的に検証し、既存の仮説が不十分であることを示し、遅延ポリシーや暗黙のアンサンブル効果が重要であることを明らかにした。
- 意味的強化学習による汎用ロボットポリシーの適応強化学習2026/6/1
汎用ロボットポリシーを強化学習で適応させる際、アクション空間ではなく言語プロンプト空間を最適化する手法を提案し、複雑な長期的タスクを解決できることを示した。
- 成功訪問マッチングによるプロセス報酬学習でRLを効率化強化学習2026/6/1
スパースな報酬を、成功・失敗エピソードを判別する識別器を使って密なプロセス報酬に変換し、ロボット操作タスクのRL微調整を高速化する手法を提案した。
- SC3-Eval: 自己整合的なビデオ生成によるロボット基盤モデルの評価評価2026/6/1
ロボット操作ポリシーの実世界評価を、アクション条件付きビデオ世界モデルを用いてスケーラブルに代替する手法を提案。前方・逆方向ダイナミクス整合性、クロスビュー整合性、テスト時整合性の3つの整合性を課すことで、正確なポリシー評価を実現する。
- フロー逆転ステアリングによるロボット汎用ポリシーの性能向上操作2026/6/1
フローマッチングに基づく汎用ロボットポリシーに対し、不完全な行動を逆フローで潜在ノイズに変換し、近傍の高品質な行動モードへ写像する手法を提案。シミュレーションと実機でゼロショット制御や強化学習の改善に効果を示した。
- OGPO: 生成制御ポリシーのサンプル効率的な完全ファインチューニングロボット学習2026/5/1
拡散やフローベースの生成制御ポリシーを、オフポリシーな批判ネットワークとPPO目的関数を用いて効率的にファインチューニングする手法OGPOを提案し、操作タスクで最先端の性能を達成した。
- RLトークン:視覚言語行動モデルを用いたオンライン強化学習のブートストラップVLA/強化学習2026/4/1
事前学習済みの視覚言語行動モデル(VLA)にRLトークンと呼ばれる小型の表現を追加し、その上で強化学習を行うことで、実ロボット上で数時間の練習だけで効率的に微調整できる手法を提案した。
- π0.7: 操縦可能な汎用ロボット基盤モデルと創発的能力VLA2026/4/1
多様なコンテキスト条件付けを用いて、未見環境での言語指示追従やゼロショットの身体汎化を実現するロボット基盤モデルπ0.7を提案した。
- MEM: 視覚言語行動モデルのためのマルチスケール具現化メモリVLA2026/3/1
ロボットの長期タスク遂行のために、ビデオベースの短期記憶とテキストベースの長期記憶を組み合わせたマルチスケールメモリアーキテクチャを提案し、15分に及ぶ複雑なタスクを可能にした。
- AsyncVLA:エッジ上での高速・堅牢なナビゲーションのための非同期VLAVLA2026/2/1
大規模基盤モデルを遠隔ワークステーションで動かし、軽量なオンボードアダプタが高頻度で動作を補正する非同期制御フレームワークを提案し、最大6秒の通信遅延下でも高いナビゲーション成功率を実現した。
- 操縦可能な視覚-言語-行動ポリシーによる身体性推論と階層制御VLA2026/2/1
サブタスクや動作、ピクセル座標など多様な抽象度の合成コマンドでVLAを訓練し、VLMの知識を低レベル行動に反映させて汎化性能を高める手法を提案。
- SteerVLA:長尾運転シーンで視覚言語行動モデルを操舵する自動運転/VLA2026/2/1
VLMの常識推論で細かい言語指示を生成し、VLA運転ポリシーを操舵する手法を提案。閉ループベンチマークで長尾シーンの性能を大幅に改善した。
- RoboReward: ロボティクス向け汎用視覚言語報酬モデルVLA2026/1/1
実ロボットの大規模データセットから報酬モデル用データセットとベンチマークを構築し、4B/8Bの視覚言語報酬モデルを訓練。短期的なロボットタスクで大規模VLMを上回る性能を達成し、実機強化学習に適用した。
- 随伴マッチングによるQ学習強化学習2026/1/1
拡散・フローマッチング方策をQ関数で効率的に最適化するため、随伴マッチングを活用したTDベース強化学習手法QAMを提案し、疎な報酬タスクで従来法を上回る性能を示した。
- 事後行動クローニング:効率的なRLファインチューニングのためのBC方策の事前学習模倣学習/RLファインチューニング2025/12/1
模倣学習で事前学習した方策がRLファインチューニングの初期値として不十分な場合があることを理論的に示し、デモンストレータ行動の事後分布をモデル化するPostBCを提案して効率的なファインチューニングを可能にした。
- 学習時アクション条件付けによる効率的なリアルタイムチャンキングVLA2025/12/1
推論時のインペインティングの代わりに、学習時に推論遅延を模擬してアクションプレフィックスで条件付けることで、VLAのリアルタイムチャンキングを計算コストを抑えつつ実現する手法を提案。
- パラメータマージによる視覚-言語-行動ロボット方策の頑健なファインチューニングVLA2025/12/1
ファインチューニング時に事前学習済みモデルと重みを補間することで、汎用能力を保ちつつ新しいタスクを頑健に学習する手法を提案した。
- 視覚-言語-行動モデルにおける人間からロボットへの転移の創発VLA2025/12/1
多様なロボットデータで事前学習したVLAモデルに人間の動画を共訓練すると、人間からロボットへのスキル転移が自然に生じることを示し、その要因が身体非依存の表現獲得にあると分析した。
- PolaRiS: 汎用ロボットポリシーのためのスケーラブルな実世界→シミュレーション評価sim2real2025/12/1
実世界の短い動画スキャンからニューラル再構成で高忠実度なシミュレーション環境を構築し、汎用ロボットポリシーを実世界と強く相関する形で評価できるようにしたフレームワーク。
- 分離型Qチャンキング:批評家と方策のチャンク長を分離した強化学習強化学習2025/12/1
批評家のチャンク長と方策のチャンク長を分離し、短い行動チャンクで方策を最適化することで、長期的な目標条件付きタスクにおけるオフライン強化学習の性能を向上させる手法を提案した。
- スケーラブルな目標条件付き強化学習のためのマルチステップ準距離学習目標条件付き強化学習2025/11/1
マルチステップモンテカルロリターンを用いて準距離を学習するオフライン目標条件付き強化学習手法を提案し、長期的な視覚タスクや実世界のロボット操作で有効性を示した。
- π*0.6:経験から学ぶVLAVLA2025/11/1
実世界での経験と人間の修正を活用した強化学習手法RECAPにより、洗濯物畳みや箱組み立て、エスプレッソ抽出などのタスクを高成功率で実行できる汎用VLAモデルπ*0.6を開発した。
- 推論時経験からアフォーダンスを学ぶ視覚-言語-行動モデルVLA2025/10/1
VLAモデルの低レベル方策を高レベルVLMと接続し、実行時の失敗経験を文脈に取り込んで反省・再計画することで、長期的タスクを遂行できるようにする手法LITENを提案。
- OmniVLA:ロボットナビゲーションのためのオムニモーダル視覚言語行動モデルVLA2025/9/1
2D姿勢・自己中心画像・自然言語などの複数のゴール指定をランダムに融合して学習し、未知環境への汎化や新しい言語指示への追従を可能にしたナビゲーション用VLAモデルを提案。
- CAST: 反事実ラベルによる視覚言語行動モデルの指示追従性向上VLA2025/8/1
視覚言語モデルを用いて既存のロボットデータセットに反事実ラベルを付与し、言語指示への追従性能を向上させる手法を提案。
- アクションチャンキングを用いた強化学習強化学習2025/7/1
将来の行動列をまとめて予測するアクションチャンキングをTDベースの強化学習に導入し、長期的で報酬が疎なタスクにおける探索とサンプル効率を改善するQ-chunkingを提案した。
- 行動的探索:インコンテキスト適応による探索の学習VLA2025/7/1
専門家のデモデータを用いて、過去の観測と探索度合いを条件に行動を予測する長文脈生成モデルを訓練し、オンラインでの高速適応と効率的な探索を実現する手法を提案。
- RoboArena:汎用ロボットポリシーの分散型実世界評価ロボット評価/ベンチマーク2025/6/1
評価者ネットワークが自由にタスクと環境を選び、ペア比較の二重盲検評価を集約することで、汎用ロボットポリシーをスケーラブルかつ公平にランキングする手法を提案し、7機関・600以上の実機評価で有効性を示した。
- アクションチャンキング方策のリアルタイム実行VLA2025/6/1
拡散・フロー系VLAの推論遅延に対応するため、次チャンクを生成しつつ現在のチャンクを非同期実行するRTCを提案し、動的タスクでの有効性を示した。
- 拡散ポリシーを潜在空間強化学習で操縦するVLA2025/6/1
拡散ポリシーの潜在ノイズ空間に強化学習を適用し、実世界で効率的に自律適応させる手法DSRLを提案。
- モデルベース再アノテーションによるどこでも走行学習ナビゲーション2025/5/1
クラウドソースの遠隔操作データや未ラベルのYouTube動画を、学習したモデルベース専門家で再ラベル付けし、長距離ナビゲーション方策LogoNavを訓練して未知環境でのロバストな走行を実現した。
- 効率的な身体性推論のための学習戦略VLA2025/5/1
ロボットのChain-of-Thought推論がなぜ有効かを分析し、軽量で高速な代替推論手法を提案した論文。
- 知識を保護する視覚-言語-行動モデル:高速学習・高速推論・より良い汎化VLA2025/5/1
連続制御用のアクションエキスパートを組み込んだVLAモデルにおいて、学習速度と知識転移が損なわれる問題を分析し、VLMバックボーンを保護する学習手法を提案した。
- π0.5: オープンワールド汎化を実現する視覚-言語-行動モデルVLA2025/4/1
異種タスクの共同学習により、未知の家庭環境でもキッチンや寝室の掃除などの長期的で器用な操作を実行できるVLAモデルπ0.5を提案した。
- AutoEval: 実世界における汎用ロボットマニピュレーションポリシーの自律評価マニピュレーション2025/3/1
ロボットポリシーの実世界評価を人手なしで24時間自律的に行うシステムAutoEvalを提案し、自動成功判定とシーンリセットにより人手による評価と近い結果が得られることを示した。
- 時間表現アラインメント:後継特徴がロボット指示追従における創発的な構成性を可能にするVLA2025/2/1
現在と将来の状態表現を時間的に整合させる損失で学習することで、明示的なサブタスク計画や強化学習なしに、基本タスクの組み合わせからなる複合タスクをこなせる構成性がロボット操作タスクで向上することを示した論文。
- 反射的計画:多段階長期的ロボットマニピュレーションのための視覚言語モデルマニピュレーション2025/2/1
視覚言語モデルに「反射」機構を組み込み、未来の世界状態を想像しながら推論を反復的に改善することで、多段階の長期ロボットマニピュレーション計画を可能にした。
- Hi Robot: 階層型視覚言語行動モデルによるオープンエンドな指示追従VLA2025/2/1
視覚言語モデルを階層的に用い、複雑な指示や実行中のフィードバックを解釈してロボットの低レベル行動に変換するシステムを提案し、片腕・双腕・移動ロボットで評価した。
- 視覚を超えて:言語グラウンディングによる異種センサーでの汎用ロボット方策のファインチューニングVLA2025/1/1
自然言語を共通のクロスモーダル基盤として用い、視覚・触覚・音声などの異種センサー入力を扱えるよう汎用ロボット方策をファインチューニングする手法FuSeを提案。実世界実験で成功率を20%以上向上させた。
- FAST: 視覚言語行動モデルのための効率的な行動トークン化VLA2025/1/1
離散コサイン変換に基づく新しい行動トークン化手法FASTを提案し、高頻度で器用なロボットタスクにおける自己回帰型VLAの学習を可能にした。
- PAE: 基盤モデルエージェントのための自律的スキル発見フレームワークVLA2024/12/17
基盤モデルエージェントが環境の文脈情報からタスクを自律的に提案し、VLMによる成功評価を報酬として強化学習でスキルを獲得する学習システムPAEを提案した。
- RLDG: 強化学習によるロボット汎用方策の蒸留マニピュレーション2024/12/1
強化学習で生成した高品質なデータを用いて汎用ロボット方策を微調整する手法を提案し、人間のデモより最大40%高い成功率を達成した。
- 汎用ロボット基盤モデルを価値誘導で操る:V-GPSVLA2024/10/1
オフライン強化学習で学習した価値関数を用いて汎用ロボット方策の行動を再ランク付けし、微調整なしで性能を向上させる手法を提案。
- GHIL-Glue: フィルタリングされたサブゴール画像による階層制御VLA2024/10/1
生成モデルが作るサブゴール画像をフィルタリングし、下位方策と効果的に繋ぐことで、言語条件付きロボット操作の汎化性能を向上させた研究。
- ロボット拡散トランスフォーマーの構成要素VLA2024/10/1
拡散モデルとトランスフォーマーを組み合わせたロボット政策の設計指針を検討し、長期的な器用タスクで高性能を発揮する新アーキテクチャを提案した。
- π0: 汎用ロボット制御のための視覚-言語-行動フローモデルVLA2024/10/1
事前学習済み視覚言語モデルにフローマッチングを組み合わせ、多様なロボットの大規模データで訓練することで、洗濯物畳みや箱組み立てなどの器用なタスクをゼロショットや言語指示で実行できる汎用ロボット基盤モデルを提案した。
- 人間参加型強化学習による精密で器用なロボットマニピュレーションマニピュレーション2024/10/1
人間のデモと修正を組み込んだ視覚ベース強化学習システムを提案し、動的操り・精密組立・両腕協調などの多様なタスクを1〜2.5時間の訓練でほぼ完璧な成功率かつ高速に学習できることを示した。
- 実世界の視覚データから学習する走破性認識型脚式ナビゲーション歩行2024/10/1
脚式ロボットの制御器の価値関数に基づいて走破性コストをロボット中心に推定し、RGBDナビゲーション計画に統合して実世界で効率的に強化学習する手法を提案した。
- LeLaN: 実世界動画から言語条件付きナビゲーション方策を学習ナビゲーション2024/10/1
大規模視覚言語モデルとロボット基盤モデルを活用し、ラベルなし・行動なしの一人称視点動画から言語で指示された物体ナビゲーション方策を学習する手法を提案。1000回以上の実世界実験で既存手法を上回り、エッジ計算で4倍高速に推論可能。
- KALIE: ロボットデータなしで視覚言語モデルを微調整するオープンワールドマニピュレーションマニピュレーション2024/9/1
人間がラベル付けした2D画像のみで視覚言語モデルを微調整し、自然言語指示と視覚からキーポイントアフォーダンスを予測してロボットを制御する手法を提案。50例のデータで未知物体の新規タスクを頑健に遂行できる。
- 異なる身体を超えて学ぶ:操作・ナビゲーション・歩行・飛行を1つのポリシーでVLA2024/8/1
20種類のロボットの90万軌道をTransformerで学習し、単一のポリシーでマニピュレーションから歩行・飛行まで制御できるCrossFormerを提案した。
- D5RL: データ駆動型深層強化学習のための多様なデータセットオフライン強化学習2024/8/1
実世界のロボット操作・移動環境を模したシミュレーションで、スクリプト・人間操作・その他多様なデータ源を含むオフライン強化学習の新ベンチマークを提案。
- 言語最適化による方策適応:少数ショット模倣のためのタスク分解模倣学習2024/8/1
視覚言語モデルによるタスク分解を活用し、少数の実演から未知の長期的ロボット操作タスクに適応する手法PALOを提案した。
- Mobility VLA:長文脈VLMとトポロジカルグラフによるマルチモーダル指示ナビゲーションナビゲーション/VLA2024/7/1
デモ動画と自然言語・画像の指示を入力とし、長文脈VLMで目標フレームを特定、トポロジカルグラフに基づく低レベル方策で実環境をナビゲートする階層型VLAを提案した。
- HiLMa-Res: 残差強化学習による四脚ロボットの移動と操作を統合する汎用階層フレームワーク移動操作2024/7/1
四脚ロボットが歩行しながら脚で操作を行うための階層型強化学習フレームワークを提案し、実機で複数の移動操作タスクにおいて既存手法より優れた性能を示した。
- 視覚言語モデルの常識推論による脚式ロボットの適応脚式ロボット2024/7/1
視覚言語モデル(VLM)の常識推論を活用し、過去の相互作用と将来の計画を考慮して脚式ロボットが未知の障害物に対処するシステムVLM-PCを提案。
- 身体性を伴う連鎖的思考推論によるロボット制御VLA2024/7/1
視覚言語行動モデルに、計画・サブタスク・動作・物体位置などの推論を段階的に行わせてから行動を予測するECoTを導入し、大規模データセット向けの合成訓練データ生成パイプラインを構築して成功率を向上させた。
- 基盤モデルによる指示追従スキルの自律的改善VLA2024/7/1
視覚言語モデルを活用してロボットが自律的に多様なデータを収集・評価し、人間の注釈なしで指示追従ポリシーを改善する手法を提案。未知環境での性能を2倍に向上させた。
- OpenVLA: オープンソースの視覚-言語-行動モデルVLA2024/6/1
97万件の実機ロボット実演で学習した7BパラメータのオープンソースVLAモデルを提案し、閉源モデルRT-2-Xを上回る汎用マニピュレーション性能と効率的なファインチューニングを実現した。
- 言語ガイド付きスキル発見スキル発見2024/6/1
大規模言語モデルの意味知識を活用し、ユーザープロンプトに基づいて意味的に多様なスキルを発見するフレームワークを提案。脚ロボットやマニピュレーションで有効性を示した。
- RACER: 認識的不確実性に基づくリスク感受型強化学習による高速走行と衝突低減強化学習/自動運転2024/5/1
リスク感受型制御と適応的行動空間カリキュラムを組み合わせ、認識的不確実性推定により分布外状態を自動回避する強化学習フレームワークを提案し、実車のラリーカーで高速オフロード走行を実現した。
- Octo: オープンソースの汎用ロボットポリシーVLA2024/5/1
80万件の軌道データで学習したTransformerベースの汎用ロボットマニピュレーションポリシーを提案し、言語や目標画像で指示でき、新しいセンサーや行動空間にも短時間でファインチューニング可能であることを示した。
- 実世界ロボットマニピュレーションポリシーのシミュレーション評価sim2real2024/5/1
実環境とシミュレーションの制御・視覚ギャップを緩和し、実機セットアップに対応した評価環境SIMPLERを構築。実機とシミュレーションの性能が強く相関することを示した。
- 視触覚データを用いた双腕多指ハンドによる技能学習マニピュレーション2024/4/1
低コストな双腕多指ハンドの遠隔操作システムHATOを開発し、触覚センサ付き義手を転用して視触覚データを収集、長期的で高精度なタスクを学習させた。
- SELFI: 強化学習による社会ナビゲーションのための自律的自己改善社会ナビゲーション2024/3/1
オフラインのモデルベース学習で事前学習した方策を、オンラインのモデルフリー強化学習で微調整する手法を提案し、実環境での衝突回避と社会的配慮行動を改善した。
- ロボットに怒鳴って教える:言語訂正によるオンザフライ改善VLA2024/3/1
人間がロボットに言語で訂正を与えることで、高レベル方策がリアルタイムに適応し、反復学習を通じて長期的タスクの成功率を向上させるフレームワークを提案。
- DROID: 大規模実環境ロボットマニピュレーションデータセットマニピュレーション2024/3/1
北米・アジア・欧州の50名が12ヶ月かけて564シーン・84タスクで収集した76k軌道・350時間のロボット操作データセットを構築し、学習ポリシーの性能と汎化性能が向上することを示した。
- MOKA: マークベース視覚プロンプティングによるオープンワールドロボットマニピュレーションマニピュレーション2024/3/1
視覚言語モデルに画像上のマークを手がかりに質問応答させ、自由文指示からアフォーダンスを予測してロボットの卓上マニピュレーションを実現する手法を提案。
- PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs2024/2/1
- Foundation Policies with Hilbert Representations2024/2/1
- Pushing the Limits of Cross-Embodiment Learning for Manipulation and Navigation2024/2/1
- Reinforcement Learning for Versatile, Dynamic, and Robust Bipedal Locomotion Control2024/1/1
- AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents2024/1/1
- FMB: a Functional Manipulation Benchmark for Generalizable Robotic Learning2024/1/1
- SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning2024/1/1
- Chain of Code: Reasoning with a Language Model-Augmented Code Emulator2023/12/1
- RLIF: Interactive Imitation Learning as Reinforcement Learning2023/11/1
- Adapt On-the-Go: Behavior Modulation for Single-Life Robot Deployment2023/11/1
- Grow Your Limits: Continuous Improvement with Real-World RL for Robotic Locomotion2023/10/1
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models2023/10/1
- Navigation with Large Language Models: Semantic Guesswork as a Heuristic for Planning2023/10/1
- NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration2023/10/1
- Offline Retraining for Online RL: Decoupled Policy Learning to Mitigate Exploration Bias2023/10/1
- Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models2023/10/1
- METRA: Scalable Unsupervised RL with Metric-Aware Abstraction2023/10/1
- REBOOT: Reuse Data for Bootstrapping Efficient Real-World Dexterous Manipulation2023/9/1
- Bootstrapping Adaptive Human-Machine Interfaces with Offline Reinforcement Learning2023/9/1
- Robotic Offline RL from Internet Videos via Value-Function Pre-Training2023/9/1
- Q-Transformer: Scalable Offline Reinforcement Learning via Autoregressive Q-Functions2023/9/1
- BridgeData V2: A Dataset for Robot Learning at Scale2023/8/1
- Multi-Stage Cable Routing through Hierarchical Imitation Learning2023/7/1
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control2023/7/1
- Goal Representations for Instruction Following: A Semi-Supervised Language Interface to Control2023/7/1
- HIQL: Offline Goal-Conditioned RL with Latent States as Actions2023/7/1
- Contrastive Example-Based Control2023/7/1
- ViNT: A Foundation Model for Visual Navigation2023/6/1
- SACSoN: Scalable Autonomous Control for Social Navigation2023/6/1
- Deep RL at Scale: Sorting Waste in Office Buildings with a Fleet of Mobile Manipulators2023/5/1
- FastRLAP: A System for Learning High-Speed Driving via Deep RL and Autonomous Practicing2023/4/1
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware2023/4/1
- Learning and Adapting Agile Locomotion Skills by Transferring Experience2023/4/1
- Neural Constraint Satisfaction: Hierarchical Abstraction for Combinatorial Generalization in Object Rearrangement2023/3/1
- Grounded Decoding: Guiding Text Generation with Grounded Models for Embodied Agents2023/3/1
- PaLM-E: An Embodied Multimodal Language Model2023/3/1
- Robust and Versatile Bipedal Jumping Control through Reinforcement Learning2023/2/1
- Learning Robotic Navigation from Experience: Principles, Methods, and Recent Results2022/12/1
- RT-1: Robotics Transformer for Real-World Control at Scale2022/12/1
- Imitation Is Not Enough: Robustifying Imitation with Reinforcement Learning for Challenging Driving Scenarios2022/12/1
- Dexterous Manipulation from Images: Autonomous Real-World RL via Substep Guidance2022/12/1
- Offline Reinforcement Learning for Visual Navigation2022/12/1
- Robotic Skill Acquisition via Instruction Augmentation with Vision-Language Models2022/11/1
- GNM: A General Navigation Model to Drive Any Robot2022/10/1
- Learning on the Job: Self-Rewarding Offline-to-Online Finetuning for Industrial Insertion of Novel Connectors from Vision2022/10/1
- ExAug: Robot-Conditioned Navigation Policies via Geometric Experience Augmentation2022/10/1
- Pre-Training for Robots: Offline RL Enables Learning New Tasks from a Handful of Trials2022/10/1
- Generalization with Lossy Affordances: Leveraging Broad Offline Data for Learning Visuomotor Tasks2022/10/1
- Simplifying Model-based RL: Learning Representations, Latent-space Models, and Policies with One Objective2022/9/1
- GenLoco: Generalized Locomotion Controllers for Quadrupedal Robots2022/9/1
- Hierarchical Reinforcement Learning for Precise Soccer Shooting Skills using a Quadrupedal Robot2022/8/1
- A Walk in the Park: Learning to Walk in 20 Minutes With Model-Free Reinforcement Learning2022/8/1
- LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action2022/7/1
- Don't Start From Scratch: Leveraging Prior Data to Automate Robotic Reinforcement Learning2022/7/1
- Inner Monologue: Embodied Reasoning through Planning with Language Models2022/7/1
- Planning to Practice: Efficient Online Fine-Tuning by Composing Goals in Latent Space2022/5/1
- First Contact: Unsupervised Human-Machine Co-Adaptation via Mutual Information Maximization2022/5/1
- INFOrmation Prioritization through EmPOWERment in Visual Model-Based RL2022/4/1
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances2022/4/1
- Control-Aware Prediction Objectives for Autonomous Driving2022/4/1
- Demonstration-Bootstrapped Autonomous Practicing via Multi-Task Reinforcement Learning2022/3/1
- How to Leverage Unlabeled Data in Offline Reinforcement Learning2022/2/1
- BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning2022/2/1
- ASHA: Assistive Teleoperation via Human-in-the-Loop Reinforcement Learning2022/2/1
- ViKiNG: Vision-Based Kilometer-Scale Navigation with Geographic Hints2022/2/1
- CoMPS: Continual Meta Policy Search2021/12/1
- Autonomous Reinforcement Learning: Formalism and Benchmarking2021/12/1
- Value Function Spaces: Skill-Centric State Abstractions for Long-Horizon Reasoning2021/11/1
- AW-Opt: Learning Robotic Skills with Imitation and Reinforcement at Scale2021/11/1
- Hybrid Imitative Planning with Geometric and Predictive Costs in Off-road Environments2021/11/1
- Mismatched No More: Joint Model-Policy Optimization for Model-Based RL2021/10/1
- Legged Robots that Keep on Learning: Fine-Tuning Locomotion Policies in the Real World2021/10/1
- Offline Meta-Reinforcement Learning for Industrial Insertion2021/10/1
- Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets2021/9/1
- Conservative Data Sharing for Multi-Task Offline Reinforcement Learning2021/9/1
- Autonomous Reinforcement Learning via Subgoal Curricula2021/7/1
- Fully Autonomous Real-World Reinforcement Learning with Applications to Mobile Manipulation2021/7/1
- Offline Meta-Reinforcement Learning with Online Self-Supervision2021/7/1
- MURAL: Meta-Learning Uncertainty-Aware Rewards for Outcome-Driven Reinforcement Learning2021/7/1
- What Can I Do Here? Learning New Skills by Imagining Visual Affordances2021/6/1
- Model-Based Reinforcement Learning via Latent-Space Collocation2021/6/1
- Hierarchically Integrated Models: Learning to Navigate from Heterogeneous Robots2021/6/1
- MT-Opt: Continuous Multi-Task Robotic Reinforcement Learning at Scale2021/4/1
- Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic Skills2021/4/1
- DisCo RL: Distribution-Conditioned Reinforcement Learning for General-Purpose Policies2021/4/1
- Reset-Free Reinforcement Learning via Multi-Task Learning: Learning Dexterous Manipulation Behaviors without Human Intervention2021/4/1
- Contingencies from Observations: Tractable Contingency Planning with Learned Behavior Models2021/4/1
- Rapid Exploration for Open-World Navigation with Latent Goal Models2021/4/1
- Reinforcement Learning for Robust Parameterized Locomotion Control of Bipedal Robots2021/3/1
- Maximum Entropy RL (Provably) Solves Some Robust RL Problems2021/3/1
- Replacing Rewards with Examples: Example-Based Policy Search via Recursive Classification2021/3/1
- How to Train Your Robot with Deep Reinforcement Learning; Lessons We've Learned2021/2/1
- COMBO: Conservative Offline Model-Based Policy Optimization2021/2/1
- SimGAN: Hybrid Simulator Identification for Domain Adaptation via Adversarial Reinforcement Learning2021/1/1
- ViNG: Learning Open-World Navigation with Visual Goals2020/12/1
- Model-Based Visual Planning with Self-Supervised Functional Distances2020/12/1
- Parrot: Data-Driven Behavioral Priors for Reinforcement Learning2020/11/1
- Reinforcement Learning with Videos: Combining Offline Observations with Interaction2020/11/1
- Rearrangement: A Challenge for Embodied AI2020/11/1
- MELD: Meta-Reinforcement Learning from Images via Latent State Models2020/10/1
- LaND: Learning to Navigate from Disengagements2020/10/1
- COG: Connecting New Skills to Past Experience with Offline Reinforcement Learning2020/10/1
- Conservative Safety Critics for Exploration2020/10/1
- Assisted Perception: Optimizing Observations to Communicate State2020/8/1
- Off-Dynamics Reinforcement Learning: Training for Transfer with Domain Classifiers2020/6/1
- AWAC: Accelerating Online Reinforcement Learning with Offline Datasets2020/6/1
- RL-CycleGAN: Reinforcement Learning Aware Simulation-To-Real2020/6/1
- Long-Horizon Visual Planning with Goal-Conditioned Hierarchical Predictors2020/6/1
- Can Autonomous Vehicles Identify, Recover From, and Adapt to Distribution Shifts?2020/6/1
- Never Stop Learning: The Effectiveness of Fine-Tuning in Robotic Reinforcement Learning2020/4/1
- Model-Based Meta-Reinforcement Learning for Flight with Suspended Payloads2020/4/1
- Learning Agile Robotic Locomotion Skills by Imitating Animals2020/4/1
- The Ingredients of Real-World Robotic Reinforcement Learning2020/4/1
- Emergent Real-World Robotic Skills via Unsupervised Off-Policy Reinforcement Learning2020/4/1
- Meta-Reinforcement Learning for Robotic Industrial Insertion Tasks2020/4/1
- Thinking While Moving: Deep Reinforcement Learning with Concurrent Control2020/4/1
- Scalable Multi-Task Imitation Learning with Autonomous Improvement2020/3/1
- Inverting the Pose Forecasting Pipeline with SPF2: Sequential Pointcloud Forecasting for Sequential Pose Forecasting2020/3/1
- OmniTact: A Multi-Directional High Resolution Touch Sensor2020/3/1
- Learning to Walk in the Real World with Minimal Human Effort2020/2/1
- Rewriting History with Inverse RL: Hindsight Inference for Policy Improvement2020/2/1
- BADGR: An Autonomous Self-Supervised Learning-Based Navigation System2020/2/1
- Gradient Surgery for Multi-Task Learning2020/1/1
- Morphology-Agnostic Visual Robotic Control2019/12/1
- AVID: Learning Multi-Stage Tasks via Pixel-Level Translation of Human Videos2019/12/1
- Learning Predictive Models From Observation and Interaction2019/12/1
- Planning with Goal-Conditioned Policies2019/11/1
- Scaled Autonomy: Enabling Human Operators to Control Robot Fleets2019/10/1
- Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning2019/10/1
- Contextual Imagined Goals for Self-Supervised Robotic Learning2019/10/1
- Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning2019/10/1
- RoboNet: Large-Scale Multi-Robot Learning2019/10/1
- Plan Arithmetic: Compositional Plan Vectors for Multi-Task Control2019/10/1
- Deep Dynamics Models for Learning Dexterous Manipulation2019/9/1
- ROBEL: Robotics Benchmarks for Learning with Low-Cost Robots2019/9/1
- Dynamics-Aware Unsupervised Discovery of Skills2019/7/1
- Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill Discovery2019/7/1
- Search on the Replay Buffer: Bridging Planning and Reinforcement Learning2019/6/1
- Deep Reinforcement Learning for Industrial Insertion Tasks with Visual Inputs and Natural Rewards2019/6/1
- Efficient Exploration via State Marginal Matching2019/6/1
- Off-Policy Evaluation via Off-Policy Classification2019/6/1
- PRECOG: PREdiction Conditioned On Goals in Visual Multi-Agent Settings2019/5/1
- Data-efficient Learning of Morphology and Controller for a Microrobot2019/5/1
- Safety Augmented Value Estimation from Demonstrations (SAVED): Safe Deep Model-Based RL for Sparse Cost Robotic Tasks2019/5/1
- REPLAB: A Reproducible Low-Cost Arm Benchmark Platform for Robotic Learning2019/5/1
- Improvisation through Physical Understanding: Using Novel Objects as Tools with Visual Foresight2019/4/1
- Guided Meta-Policy Search2019/4/1
- End-to-End Robotic Reinforcement Learning without Reward Engineering2019/4/1
- Learning to Identify Object Instances by Touch: Tactile Recognition via Multimodal Matching2019/3/1
- Manipulation by Feel: Touch-Based Control with Deep Predictive Models2019/3/1
- Skew-Fit: State-Covering Self-Supervised Reinforcement Learning2019/3/1
- Learning Latent Plans from Play2019/3/1
- Artificial Intelligence for Prosthetics - challenge solutions2019/2/1
- Generalization through Simulation: Integrating Simulated and Real Data into Deep Reinforcement Learning for Vision-Based Autonomous Flight2019/2/1
- Low Level Control of a Quadrotor with Deep Model-Based Reinforcement Learning2019/1/1
- Sim-to-Real via Sim-to-Sim: Data-efficient Robotic Grasping via Randomized-to-Canonical Adaptation Networks2018/12/1
- Residual Reinforcement Learning for Robot Control2018/12/1
- Visual Foresight: Model-Based Deep Reinforcement Learning for Vision-Based Robotic Control2018/12/1
- Reasoning About Physical Interactions with Object-Oriented Prediction and Planning2018/12/1
- Soft Actor-Critic Algorithms and Applications2018/12/1
- Visual Memory for Robust Path Following2018/12/1
- Learning to Walk via Deep Reinforcement Learning2018/12/1
- Deep Online Learning via Meta-Learning: Continual Adaptation for Model-Based RL2018/12/1
- Grasp2Vec: Learning Object Representations from Self-Supervised Grasping2018/11/1
- Hierarchical Policy Design for Sample-Efficient Learning of Robot Table Tennis Through Self-Play2018/11/1
- Few-Shot Goal Inference for Visuomotor Learning and Planning2018/10/1
- Deep Imitative Models for Flexible Inference, Planning, and Control2018/10/1
- One-Shot Hierarchical Imitation Learning of Compound Visuomotor Tasks2018/10/1
- Dexterous Manipulation with Deep Reinforcement Learning: Efficient, General, and Low-Cost2018/10/1
- Robustness via Retrying: Closed-Loop Robotic Manipulation with Self-Supervised Learning2018/10/1
- Composable Action-Conditioned Predictors: Flexible Off-Policy Learning for Robot Navigation2018/10/1
- Time Reversal as Self-Supervision2018/10/1
- SOLAR: Deep Structured Representations for Model-Based Reinforcement Learning2018/8/1
- Visual Reinforcement Learning with Imagined Goals2018/7/1
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation2018/6/1
- Learning Instance Segmentation by Interaction2018/6/1
- Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review2018/5/1
- More Than a Feeling: Learning to Grasp and Regrasp using Vision and Touch2018/5/1
- Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models2018/5/1
- Stochastic Adversarial Video Prediction2018/4/1
- Universal Planning Networks2018/4/1
- Learning to Adapt in Dynamic, Real-World Environments Through Meta-Reinforcement Learning2018/3/1
- Composable Deep Reinforcement Learning for Robotic Manipulation2018/3/1
- Learning Flexible and Reusable Locomotion Primitives for a Microrobot2018/3/1
- Shared Autonomy via Deep Reinforcement Learning2018/2/1
- Deep Reinforcement Learning for Vision-Based Robotic Grasping: A Simulated Comparative Evaluation of Off-Policy Methods2018/2/1
- One-Shot Imitation from Observing Humans via Domain-Adaptive Meta-Learning2018/2/1
- Diversity is All You Need: Learning Skills without a Reward Function2018/2/1
- Unifying Map and Landmark Based Representations for Visual Navigation2017/12/1
- Sim2Real View Invariant Visual Servoing by Recurrent Control2017/12/1
- Learning Image-Conditioned Dynamics Models for Control of Under-actuated Legged Millirobots2017/11/1
- Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning2017/11/1
- Divide-and-Conquer Reinforcement Learning2017/11/1
- The Feeling of Success: Does Touch Sensing Help Predict Grasp Outcomes?2017/10/1
- Self-Supervised Visual Planning with Temporal Skip Connections2017/10/1
- Stochastic Variational Video Prediction2017/10/1
- One-Shot Visual Imitation Learning via Meta-Learning2017/9/1
- Using Simulation and Domain Adaptation to Improve Efficiency of Deep Robotic Grasping2017/9/1
- Self-supervised Deep Reinforcement Learning with Generalized Computation Graphs for Robot Navigation2017/9/1
- Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations2017/9/1
- MBMF: Model-Based Priors for Model-Free Reinforcement Learning2017/9/1
- Learning Robotic Manipulation of Granular Media2017/9/1
- Deep Object-Centric Representations for Generalizable Robot Learning2017/8/1
- Neural Network Dynamics for Model-Based Deep Reinforcement Learning with Model-Free Fine-Tuning2017/8/1
- GPLAC: Generalizing Vision-Based Robotic Skills using Weakly Labeled Images2017/8/1
- Imitation from Observation: Learning to Imitate Behaviors from Raw Video via Context Translation2017/7/1
- End-to-End Learning of Semantic Grasping2017/7/1
- Vision-Based Multi-Task Manipulation for Inexpensive Robots Using End-To-End Learning from Demonstration2017/7/1
- Interpolated Policy Gradient: Merging On-Policy and Off-Policy Gradient Estimation for Deep Reinforcement Learning2017/6/1
- Time-Contrastive Networks: Self-Supervised Learning from Video2017/4/1
- Learning Visual Servoing with Deep Features and Fitted Q-Iteration2017/3/1
- Combining Model-Based and Model-Free Updates for Trajectory-Centric Reinforcement Learning2017/3/1
- Learning Invariant Feature Spaces to Transfer Skills with Reinforcement Learning2017/3/1
- Combining Self-Supervised Learning and Imitation for Vision-Based Rope Manipulation2017/3/1
- Cognitive Mapping and Planning for Visual Navigation2017/2/1
- Uncertainty-Aware Reinforcement Learning for Collision Avoidance2017/2/1
- Unsupervised Perceptual Rewards for Imitation Learning2016/12/1
- Generalizing Skills with Semi-Supervised Reinforcement Learning2016/12/1
- Learning Dexterous Manipulation Policies from Experience and Imitation2016/11/1
- CAD2RL: Real Single-Image Flight without a Single Real Image2016/11/1
- Reset-Free Guided Policy Search: Efficient Deep Reinforcement Learning with Stochastic Initial States2016/10/1
- Deep Reinforcement Learning for Robotic Manipulation with Asynchronous Off-Policy Updates2016/10/1
- EPOpt: Learning Robust Neural Network Policies Using Model Ensembles2016/10/1
- Collective Robot Reinforcement Learning with Distributed Asynchronous Guided Policy Search2016/10/1
- Deep Visual Foresight for Planning Robot Motion2016/10/1
- Path Integral Guided Policy Search2016/10/1
- Deep Reinforcement Learning for Tensegrity Robot Locomotion2016/9/1
- Learning Modular Neural Network Policies for Multi-Task and Multi-Robot Transfer2016/9/1
- Learning from the Hindsight Plan -- Episodic MPC Improvement2016/9/1
- Guided Policy Search as Approximate Mirror Descent2016/7/1
- Learning to Poke by Poking: Experiential Learning of Intuitive Physics2016/6/1
- Unsupervised Learning for Physical Interaction through Video Prediction2016/5/1
- Guided Cost Learning: Deep Inverse Optimal Control via Policy Optimization2016/3/1
- Continuous Deep Q-Learning with Model-based Acceleration2016/3/1
- Learning Hand-Eye Coordination for Robotic Grasping with Deep Learning and Large-Scale Data Collection2016/3/1
- Learning Dexterous Manipulation for a Soft Robotic Hand from Human Demonstration2016/3/1
- Model-based Reinforcement Learning with Parametrized Physical Models and Optimism-Driven Exploration2015/9/1
- Learning Deep Control Policies for Autonomous Aerial Vehicles with MPC-Guided Policy Search2015/9/1
- One-Shot Learning of Manipulation Skills with Online Dynamics Adaptation and Neural Network Priors2015/9/1
- Deep Spatial Autoencoders for Visuomotor Learning2015/9/1
- Learning Deep Neural Network Policies with Continuous Memory States2015/7/1
- High-Dimensional Continuous Control Using Generalized Advantage Estimation2015/6/1
- End-to-End Training of Deep Visuomotor Policies2015/4/1
- Learning Contact-Rich Manipulation Skills with Guided Policy Search2015/1/1
- Exploring Deep and Recurrent Architectures for Optimal Control2013/11/1