Arsalan Mousavian
収録論文 41本 ・ フィジカルAI/ロボット学習
マニピュレーション模倣学習世界モデル動作計画VLA
※arXiv著者名で収集。同姓同名の別人の論文が含まれる場合があります。
論文
- PointWorld: 実世界ロボット操作のための3D世界モデルのスケーリングマニピュレーション2026/1/1
RGB-D画像とロボットの行動コマンドから3D点群の動きを予測する大規模事前学習済み3D世界モデルを提案し、実機ロボットのモデル予測制御に応用した。
- VT-Refine: 視触覚フィードバックによる両手組立のシミュレーション微調整学習マニピュレーション2025/10/1
少数の実演と高忠実度の触覚シミュレーションを組み合わせ、強化学習で両手組立ポリシーを微調整することで、接触の多い精密作業の実世界性能を向上させた研究。
- Dexplore: 参照スコープ探索による器用な操作のためのスケーラブルなニューラル制御マニピュレーション2025/9/1
人間の手のモーションキャプチャをソフトなガイドとして活用し、リターゲティングと追従を単一ループで統合最適化することで、器用なロボット操作の制御方策を大規模に学習する手法を提案。
- 3D FlowMatch Actor:単腕・双腕マニピュレーションのための統合3Dポリシーマニピュレーション2025/8/1
流れマッチングと3D事前学習済み視覚表現を組み合わせ、単腕・双腕ロボット操作を高速かつ高精度に行う3Dポリシーを提案。
- 人間の単一動画からの視覚的模倣によるスロット単位のロボット配置模倣学習2025/4/1
人間の作業動画を1本見せるだけで、どの物体をどこに置くかを理解し、ロボットが同じ配置作業を再現できるモジュール式システムSLeRPを提案した。
- 物理AIのためのCosmos世界基盤モデルプラットフォーム世界モデル2025/1/1
物理AI開発者向けに、カスタマイズ可能な世界モデルを構築するための基盤モデル・動画キュレーション・トークナイザを提供するオープンソースプラットフォーム。
- DiffusionSeeder: 拡散モデルによる運動最適化のシード生成で高速動作計画動作計画2024/10/1
拡散モデルで深度画像から多様な軌道を生成し、GPU加速運動最適化の初期シードとして使うことで、混雑環境でのロボット動作計画を大幅に高速化・高成功率化する手法。
- RoboPoint: ロボティクス向け空間アフォーダンス予測の視覚言語モデルVLA2024/6/1
合成データで視覚言語モデルをロボット向けに微調整し、言語指示から画像内のキーポイントアフォーダンスを予測するRoboPointを開発。実世界データ不要でスケーラブルに学習でき、ナビゲーションや操作などのタスクで既存手法を上回る。
- M2T2: Multi-Task Masked Transformer for Object-centric Pick and Place2023/11/1
- CabiNet: Scaling Neural Collision Detection for Object Rearrangement with Procedural Scene Generation2023/4/1
- Constrained Generative Sampling of 6-DoF Grasps2023/2/1
- MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare2022/12/1
- Learning Robust Real-World Dexterous Grasping Policies via Implicit Shape Augmentation2022/10/1
- ProgPrompt: Generating Situated Robot Task Plans using Large Language Models2022/9/1
- Deep Learning Approaches to Grasp Synthesis: A Review2022/7/1
- IFOR: Iterative Flow Minimization for Robotic Object Rearrangement2022/2/1
- RICE: Refining Instance Masks in Cluttered Environments with Graph Neural Networks2021/6/1
- NeRP: Neural Rearrangement Planning for Unknown Objects2021/6/1
- STORM: An Integrated Framework for Fast Joint-Space Model-Predictive Control for Reactive Manipulation2021/4/1
- RGB-D Local Implicit Function for Depth Completion of Transparent Objects2021/4/1
- Sim-to-Real for Robotic Tactile Sensing via Physics-Based Simulation and Learned Latent Projections2021/3/1
- Contact-GraspNet: Efficient 6-DoF Grasp Generation in Cluttered Scenes2021/3/1
- Interpreting and Predicting Tactile Signals for the SynTouch BioTac2021/1/1
- Object Rearrangement Using Learned Implicit Collision Functions2020/11/1
- ACRONYM: A Large-Scale Grasp Dataset Based on Simulation2020/11/1
- Reactive Long Horizon Task Execution via Visual Skill and Precondition Models2020/11/1
- Reactive Human-to-Robot Handovers of Arbitrary Objects2020/11/1
- Goal-Auxiliary Actor-Critic for 6D Robotic Grasping with Point Clouds2020/10/1
- Unseen Object Instance Segmentation for Robotic Environments2020/7/1
- Learning RGB-D Feature Embeddings for Unseen Object Instance Segmentation2020/7/1
- Interpreting and Predicting Tactile Signals via a Physics-Based and Data-Driven Framework2020/6/1
- LatentFusion: End-to-End Differentiable Reconstruction and Rendering for Unseen Object Pose Estimation2019/12/1
- A Billion Ways to Grasp: An Evaluation of Grasp Sampling Schemes on a Dense, Physics-based Grasp Data Set2019/12/1
- 6-DOF Grasping for Target-driven Object Manipulation in Clutter2019/12/1
- Self-supervised 6D Object Pose Estimation for Robot Manipulation2019/9/1
- The Best of Both Modes: Separately Leveraging RGB and Depth for Unseen Object Instance Segmentation2019/7/1
- PoseRBPF: A Rao-Blackwellized Particle Filter for 6D Object Pose Tracking2019/5/1
- 6-DOF GraspNet: Variational Grasp Generation for Object Manipulation2019/5/1
- Synthesizing Training Data for Object Detection in Indoor Scenes2017/2/1
- Multiview RGB-D Dataset for Object Instance Detection2016/9/1
- Semantic Image Based Geolocation Given a Map2016/9/1