Yao Lu
収録論文 28本 ・ フィジカルAI/ロボット学習
セグメンテーションVLA
※arXiv著者名で収集。同姓同名の別人の論文が含まれる場合があります。
論文
- Hyp2Former: オープンセットパノプティックセグメンテーションのための階層認識双曲埋め込みセグメンテーション2026/5/1
既知カテゴリの階層構造を双曲空間で学習し、未知物体を高次概念に近づけて検出するオープンセットパノプティックセグメンテーション手法を提案。
- π0.7: 操縦可能な汎用ロボット基盤モデルと創発的能力VLA2026/4/1
多様なコンテキスト条件付けを用いて、未見環境での言語指示追従やゼロショットの身体汎化を実現するロボット基盤モデルπ0.7を提案した。
- ForeAct: 効率的な視覚的先読み計画でVLAを誘導するVLA2026/2/1
将来の観測画像とサブタスク記述を生成してVLAの視覚入力に追加するだけで、アーキテクチャ変更なしに実世界タスクの成功率を大幅に向上させる計画モジュールを提案。
- VLASH: 未来状態を考慮した非同期推論によるリアルタイムVLAVLA2025/12/1
ロボットの状態を前回の行動チャンクで先読みし、予測と実行の時間ずれを解消することで、追加コストなしにVLAの非同期推論を高速化・安定化する手法を提案。
- π*0.6:経験から学ぶVLAVLA2025/11/1
実世界での経験と人間の修正を活用した強化学習手法RECAPにより、洗濯物畳みや箱組み立て、エスプレッソ抽出などのタスクを高成功率で実行できる汎用VLAモデルπ*0.6を開発した。
- EgoVLA:自己中心視点の人間動画から視覚-言語-行動モデルを学習VLA2025/7/1
一人称視点の人間動画でVLAモデルを訓練し、逆運動学とリターゲティングで人間の手の動きをロボット動作に変換、少数のロボット実演で微調整してロボットポリシーを獲得する手法を提案。
- CoT-VLA:視覚的思考連鎖推論を用いた視覚-言語-行動モデルVLA2025/3/1
将来の画像フレームを視覚的目標として自己回帰的に予測してから行動系列を生成することで、視覚的思考連鎖推論をVLAに組み込み、実世界タスクで17%の性能向上を達成した。
- AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents2024/1/1
- SARA-RT: Scaling up Robotics Transformers with Self-Adaptive Robust Attention2023/12/1
- RoboVQA: Multimodal Long-Horizon Reasoning for Robotics2023/11/1
- RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches2023/11/1
- Waymax: An Accelerated, Data-Driven Simulator for Large-Scale Autonomous Driving Research2023/10/1
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models2023/10/1
- Q-Transformer: Scalable Offline Reinforcement Learning via Autoregressive Q-Functions2023/9/1
- Alexa, play with robot: Introducing the First Alexa Prize SimBot Challenge on Embodied AI2023/8/1
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control2023/7/1
- Deep RL at Scale: Sorting Waste in Office Buildings with a Fleet of Mobile Manipulators2023/5/1
- Grounded Decoding: Guiding Text Generation with Grounded Models for Embodied Agents2023/3/1
- Open-World Object Manipulation using Pre-trained Vision-Language Models2023/3/1
- RT-1: Robotics Transformer for Real-World Control at Scale2022/12/1
- Token Turing Machines2022/11/1
- PI-QT-Opt: Predictive Information Improves Multi-Task Robotic Reinforcement Learning at Scale2022/10/1
- Robotic Table Wiping via Reinforcement Learning and Whole-body Trajectory Optimization2022/10/1
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances2022/4/1
- AW-Opt: Learning Robotic Skills with Imitation and Reinforcement at Scale2021/11/1
- Value Function Spaces: Skill-Centric State Abstractions for Long-Horizon Reasoning2021/11/1
- Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic Skills2021/4/1
- Visionary: Vision architecture discovery for robot learning2021/3/1