Markus Wulfmeier
収録論文 41本 ・ フィジカルAI/ロボット学習
※arXiv著者名で収集。同姓同名の別人の論文が含まれる場合があります。
論文
- HuGo: ヒューマノイドの移動操作のための全身ポリシーコードを設計するLLMVLA2026/9/24
LLMがタスク記述から高レベルポリシーコードを生成し、ロールアウトで改良することで、報酬設計や実演なしにヒューマノイドの移動操作を実現する階層的手法を提案。
- EXIMO: VLMによるVLAポリシー探索のガイドVLA/強化学習2026/8/20
VLAポリシーの効率的なファインチューニング手法EXIMOを提案。VLMをプランナーとして使い、長期的なタスクを分解してデータ収集し、模倣学習とオフポリシーRLで最適化する。
- マルチエージェント強化学習による超人的で安全なアジャイルレーシングマルチエージェント強化学習2026/5/1
マルチエージェント強化学習を用いて、高速クアッドローターレースで人間のチャンピオンパイロットを上回る性能と安全性を両立させ、ゼロショットで人間との安全な相互作用を実現した。
- 視覚的ストーリーテリングのための具現化されたコンパニオンヒューマンロボットインタラクション2026/3/1
描画ロボットと大規模言語モデルを統合し、人間と機械が対話しながら共同で視覚的な物語を創り出す芸術的システムを提案した。専門家による評価で、独自の美的アイデンティティと展示価値が確認された。
- 実機ロボットのオンライン強化学習で重要な設計選択とは何か強化学習2026/2/1
3種類のロボットで100回の実機学習を行い、アルゴリズムやシステム、実験設計の各選択を系統的に検証し、安定した学習を実現する設計指針を明らかにした。
- Gemini Robotics 1.5:高度な身体性推論・思考・動作転移で汎用ロボットの最前線を押し広げるVLA2025/10/1
異種ロボットデータから学ぶ動作転移機構と自然言語での多段階推論を組み合わせたVLAモデルと、身体性推論に特化したモデルを発表し、複雑な多段階タスクの実行と解釈性を向上させた。
- 巧みな操作のための政策アイドリングの活用マニピュレーション2025/8/1
ロボットの巧みな操作において、政策が特定の状態で動かなくなる「アイドリング」を検出し、その状態に摂動を加えることで探索を促し、性能を向上させる手法を提案。
- 自己中心視覚と深層強化学習によるロボットサッカーの学習マルチエージェント強化学習2024/5/1
自己中心的なRGB視覚のみを入力とし、関節レベルの行動を出力するマルチエージェントロボットサッカーの方策を深層強化学習で訓練し、シミュレーションから実機へ転移することに成功した。
- 成長するQネットワーク:適応的制御解像度による連続制御タスクの解決強化学習2024/4/1
離散行動空間を粗い解像度から細かい解像度へと成長させることで、連続制御タスクにおける強化学習の性能を向上させる手法を提案。
- Real-World Fluid Directed Rigid Body Control via Deep Reinforcement Learning2024/2/1
- Mastering Stacking of Diverse Shapes with Large-Scale Iterative Reinforcement Learning on Real Robots2023/12/1
- Foundations for Transfer in Reinforcement Learning: A Taxonomy of Knowledge Modalities2023/12/1
- Replay across Experiments: A Natural Extension of Off-Policy RL2023/11/1
- Equivariant Data Augmentation for Generalization in Offline Reinforcement Learning2023/9/1
- Real Robot Challenge 2022: Learning Dexterous Manipulation from Offline Data in the Real World2023/8/1
- Towards A Unified Agent with Foundation Models2023/7/1
- Learning Agile Soccer Skills for a Bipedal Robot with Deep Reinforcement Learning2023/4/1
- SkillS: Adaptive Skill Sequencing for Efficient Temporally-Extended Exploration2022/11/1
- Solving Continuous Control via Q-learning2022/10/1
- MO2: Model-Based Offline Options2022/9/1
- Forgetting and Imbalance in Robot Lifelong Learning with Off-policy Data2022/4/1
- Imitate and Repurpose: Learning Reusable Robot Movement Skills From Human and Animal Behaviors2022/3/1
- Wish you were here: Hindsight Goal Selection for long-horizon dexterous manipulation2021/12/1
- Learning Transferable Motor Skills with Hierarchical Latent Mixture Policies2021/12/1
- Is Bang-Bang Control All You Need? Solving Continuous Control with Bernoulli Policies2021/11/1
- Is Curiosity All You Need? On the Utility of Emergent Behaviours from Curious Exploration2021/9/1
- From Motor Control to Team Play in Simulated Humanoid Football2021/5/1
- Representation Matters: Improving Perception and Exploration for Robotics2020/11/1
- "What, not how": Solving an under-actuated insertion task from scratch2020/10/1
- Towards General and Autonomous Learning of Core Skills: A Case Study in Locomotion2020/8/1
- Data-efficient Hindsight Off-policy Option Learning2020/7/1
- Simple Sensor Intentions for Exploration2020/5/1
- Continuous-Discrete Reinforcement Learning for Hybrid Control in Robotics2020/1/1
- Compositional Transfer in Hierarchical Reinforcement Learning2019/6/1
- Efficient Supervision for Robot Learning via Imitation, Simulation, and Adaptation2019/4/1
- On Machine Learning and Structure for Mobile Robots2018/6/1
- Incremental Adversarial Domain Adaptation for Continually Changing Environments2017/12/1
- Reverse Curriculum Generation for Reinforcement Learning2017/7/1
- Addressing Appearance Change in Outdoor Robotics with Adversarial Domain Adaptation2017/3/1
- Incorporating Human Domain Knowledge into Large Scale Cost Function Learning2016/12/1
- Watch This: Scalable Cost-Function Learning for Path Planning in Urban Environments2016/7/1