Mudit Verma
収録論文 6本 ・ フィジカルAI/ロボット学習
強化学習
※arXiv著者名で収集。同姓同名の別人の論文が含まれる場合があります。
論文
- 人間の選好に基づく報酬学習のための後知恵PRIOR強化学習2024/4/12
選好ベース強化学習において、世界モデルで状態の重要度を推定し、報酬を状態重要度に比例させる補助目的を導入することで、クレジット割り当て問題を改善し、歩行・操作タスクでの性能と報酬復元を向上させた。
- Theory of Mind abilities of Large Language Models in Human-Robot Interaction : An Illusion?2024/1/1
- Incorporating Human Flexibility through Reward Preferences in Human-AI Teaming2023/12/21
- A State Augmentation based approach to Reinforcement Learning from Human Preferences2023/2/17
- Exploiting Unlabeled Data for Feedback Efficient Human Preference based Reinforcement Learning2023/2/1
- Trust-Aware Planning: Modeling Trust Evolution in Iterated Human-Robot Interaction2021/5/1