Stefan Wermter
University of Hamburg
収録論文 84本 ・ フィジカルAI/ロボット学習
※arXiv著者名で収集。同姓同名の別人の論文が含まれる場合があります。
論文
- VLA-Dreamerに向けて:世界モデルによるVLA行動の洗練VLA2026/9/25
VLAの視覚エンコーダの埋め込み空間で世界モデルを学習し、データ効率の改善と推論時の短期計画を目指す構想論文。
- MorphIK: 未知のロボットに対する形態条件付きニューラル逆運動学逆運動学2026/9/24
フローマッチングとトランスフォーマーを用いて、訓練で見たことのない多様なロボットの逆運動学を高精度に解き、さらにヌル空間サンプリングも可能にする手法を提案。
- SmoLSTM: エピソード記憶を持続するコンパクトな視覚-言語-行動モデルVLA2026/9/19
凍結したSmolVLMとLSTM制御層を組み合わせ、エピソード全体でリセットしない再帰状態により、遮蔽や見分けがつかない物体を含むタスクを少ないパラメータで高精度に実行するVLAモデルを提案。
- 視覚・言語・行動モデルの潜在クラスタ解析VLA/解釈可能性2026/9/2
VLAモデルの内部表現を層ごとに解析し、行動デコーダの潜在空間をクラスタリングして解釈可能な概念を抽出するフレームワークを提案した。
- 未知の物体を識別する:ロボット物体操作のためのプロンプト不要なオープンボキャブラリ異常認識物体認識2026/6/25
ロボットが未知の物体を認識できるように、異常検出用MAEとプロンプト不要のオープンボキャブラリ分類器NOVICを組み合わせた二段階フレームワークAnomNOVICを提案し、テーブルトップ環境で高い精度を達成した。
- いつ手助けしましょうか?グループ人間-ロボット協調におけるプロアクティブ性の効果についてHRI/グループ協調2026/6/1
ペア参加者がヒューマノイドロボットと脱出ゲームを行う実験で、ロボットが呼ばれた時だけ反応するリアクティブモデルと、常時聞いて自発的に貢献するプロアクティブモデルを比較し、相互作用頻度や成功率、参加者の評価に与える影響を調べた。
- 人間とロボットの相互作用における不確実性、曖昧さ、多義性:概念化の重要性HRI2026/4/1
この論文は、人間とロボットの相互作用(HRI)における不確実性、曖昧さ、多義性の概念を整理し、一貫した定義と相互関係を提案するものである。
- 2D観測から汎化可能な3Dシーン表現を学習する手法3D認識2026/2/1
ロボットの視点画像から作業空間の3D占有を予測する汎化可能なNeRFを提案し、ヒューマノイドで26mmの精度を達成した。
- 説明可能な人間-ロボットインタラクションのための心の理論人間-ロボットインタラクション2025/12/1
ロボットの心の理論(ToM)を説明可能AI(XAI)の一種として捉え、XAIの評価枠組みで評価することを提案し、ユーザー中心の説明の重要性を論じた論文。
- 指差し誘導による対象推定:Transformerベースの注意機構を用いてHRI2025/9/1
人間の指差しジェスチャーからテーブル上の対象物を推定するため、モダリティ間注意機構を活用したモジュール型アーキテクチャMM-ITFを提案し、単眼RGB画像で高精度に対象を予測できることを示した。
- NICOLロボットにおけるキーポイントベース拡散モデルによる動作計画動作計画2025/9/1
数値計画手法のデータセットを用いて拡散モデルを学習し、NICOLロボットの動作計画を高速化。点群入力を用いなくても、衝突のない解を最大90%の成功率で生成し、実行時間を桁違いに短縮した。
- ロボットと話す:HRI応用のための音声基盤モデルの実践的検討音声認識2025/8/1
人間とロボットの対話で使われる音声認識システム4種を、雑音・訛り・子供/高齢者・障害・自発発話など6つの難しさを持つ8データセットで評価し、性能差や幻覚・バイアスの問題を明らかにした。
- 長期的な人間とロボットの相互作用における個別化された説明XHRI2025/7/1
ロボットが人間に説明を行う際に、ユーザーの知識モデルを更新・参照して説明の詳細度を個別化するフレームワークを提案し、LLMを用いた3つのアーキテクチャを病院巡回ロボットとキッチンアシスタントロボットのシナリオで評価した。
- LLMを対話的教師として用いるロボットマニピュレーションの模倣学習マニピュレーション2025/4/1
人間教師の代わりに大規模言語モデル(LLM)を対話的教師として活用し、ロボットマニピュレーションの模倣学習を効率化するフレームワークLLM-iTeachを提案した。
- LLM+MAP: 大規模言語モデルと計画領域定義言語を用いた両腕ロボットのタスク計画タスク計画2025/3/1
LLMの推論とマルチエージェント計画を組み合わせ、両腕ロボットの長期的なタスクを効率的に分解・割り当てする計画フレームワークを提案。
- 振っても混ぜるな:ヒューマンロボットバーテンディングにおけるグラスの視覚理解のための新規データセット物体検出2025/3/1
透明・反射するグラスの多様性に対応するため、RGB-Dセンサから自動ラベリングで収集した実世界グラスデータセットGlassNICOLDatasetを構築し、オープンボキャブラリ検出器を上回る性能とバーテンディングタスクでの81%の成功率を達成した。
- DIRIGENt:拡散モデルによる人間の実演からのエンドツーエンドロボット模倣模倣学習2025/1/1
人間がロボットを模倣したデータで拡散モデルを訓練し、RGB画像から直接関節値を生成してロボットが人間の動作を模倣できるエンドツーエンド手法を提案した。
- 多様なユーザー層に適応する人間-ロボットインタラクションのフレームワークHRI2024/10/1
多様なユーザー層に合わせて対話を調整し、ユーザーが小さな割り込みや大きな変更でインタラクションを制御できるROSベースの適応型HRIフレームワークを開発した。
- ロボットもマルチタスクをこなせる:記憶アーキテクチャとLLMを統合したクロスタスク行動生成VLA2024/7/1
人間の認知に着想を得た二層の記憶モデルと2つのLLMを組み合わせ、ロボットが複数タスク間を切り替えながら適応的に行動を生成する手法を提案し、5つのタスクで性能向上を示した。
- おしゃべりするロボット:マルチモーダルな人間-ロボット会話と協働の基盤構築VLA2024/7/1
LLMを中心に音声認識・音声生成・物体検出・姿勢推定・ジェスチャ検出を統合し、自然な人間-ロボット会話と協働を実現するモジュール型システムを提案した。
- 細部が違いを生む:物体状態に敏感な神経ロボティクスタスクプランニングタスクプランニング2024/6/1
物体の状態を考慮したタスク計画を行うエージェントOSSAを提案し、モジュール型とモノリシック型の2手法を比較した。テーブル上の物体を片付けるタスクで評価し、モノリシック型が優位であることを示した。
- 言語モデルによる強化学習エージェントのメンタルモデリングVLA2024/6/1
大規模言語モデルがエージェントの行動履歴からその意思決定を推論し、メンタルモデルを構築できるかを評価指標を提案して検証した研究。
- LLM駆動によるロボットスキルの自動発見フレームワークスキル獲得2024/5/1
LLMがシーン記述とロボット構成からタスクを提案し、強化学習と視覚言語モデルによる検証を通じて、ゼロから多様な基本スキルを段階的に獲得・拡張する手法を提案。
- 他人の靴で拡散する:拡散モデルによるロボットの視点取得模倣学習2024/4/1
三人称視点の実演から一人称視点の画像を生成する拡散モデルを提案し、ヒューマノイドロボットが直接三人称のデモンストレーションから模倣学習できるようにした。
- 大規模言語モデルによる双腕ロボットの制御マニピュレーション2024/4/1
LLMを使って双腕ロボットの長期的な協調作業を計画・制御するエージェントLABORを提案し、シミュレーションで成功率が向上することを示した論文。
- リンゴとオレンジの比較:LLMを活用した物体分類タスクにおけるマルチモーダル意図予測意図予測2024/4/1
物体分類タスクで協働するロボットが、ユーザーのジェスチャー・姿勢・表情・発話・環境状態をLLMで統合し、意図を階層的に予測する手法を提案・評価した。
- ヒューマノイド実体エージェントによる神経ロボティクス把持のための逆運動学マニピュレーション2024/4/1
ベジェ曲線ベースのデカルト軌道計画を、プラットフォーム非依存な神経模倣逆運動学CycleIKで関節空間軌道に変換し、LLMを核とする実体エージェントを通じてNICO/NICOLでの把持を実現した。
- Human Impression of Humanoid Robots Mirroring Social Cues2024/1/1
- Robotic Imitation of Human Actions2024/1/1
- Accelerating Reinforcement Learning of Robotic Manipulations via Feedback from Large Language Models2023/11/1
- Continual Robot Learning using Self-Supervised Task Inference2023/9/1
- The Emotional Dilemma: Influence of a Human-like Robot on Trust and Cooperation2023/7/1
- Clarifying the Half Full or Half Empty Question: Multimodal Container Classification2023/7/1
- CycleIK: Neuro-inspired Inverse Kinematics2023/7/1
- NICOL: A Neuro-inspired Collaborative Semi-humanoid Robot that Bridges Social Interaction and Reliable Manipulation2023/5/1
- Map-based Experience Replay: A Memory-Efficient Solution to Catastrophic Forgetting in Reinforcement Learning2023/5/1
- A Closer Look at Reward Decomposition for High-Level Robotic Explanations2023/4/1
- Sample-efficient Real-time Planning with Curiosity Cross-Entropy Method and Contrastive Learning2023/3/7
- Chat with the Environment: Interactive Multimodal Perception Using Large Language Models2023/3/1
- The Robot in the Room: Influence of Robot Facial Expressions and Gaze on Human-Human-Robot Collaboration2023/3/1
- Partially Adaptive Multichannel Joint Reduction of Ego-noise and Environmental Noise2023/3/1
- Wrapyfi: A Python Wrapper for Integrating Robots, Sensors, and Applications across Multiple Middleware2023/2/1
- Learning Bidirectional Action-Language Translation with Limited Supervision and Incongruent Input2023/1/1
- Introspection-based Explainable Reinforcement Learning in Episodic and Non-episodic Scenarios2022/11/1
- Learning to Autonomously Reach Objects with NICO and Grow-When-Required Networks2022/10/1
- Intelligent problem-solving as integrated hierarchical reinforcement learning2022/8/1
- Judging by the Look: The Impact of Robot Gaze Strategies on Human Cooperation2022/8/1
- Impact Makes a Sound and Sound Makes an Impact: Sound Guides Representations and Explorations2022/8/1
- Learning Flexible Translation between Robot Actions and Language Descriptions2022/7/1
- Explain yourself! Effects of Explanations in Human-Robot Interaction2022/4/1
- Language Model-Based Paired Variational Autoencoders for Robotic Language Learning2022/1/1
- A trained humanoid robot can perform human-like crossmodal social attention and conflict resolution2021/11/1
- Robotic Occlusion Reasoning for Efficient Object Existence Prediction2021/7/1
- Model Mediated Teleoperation with a Hand-Arm Exoskeleton in Long Time Delays Using Reinforcement Learning2021/7/1
- Behavior Self-Organization Supports Task Inference for Continual Robot Learning2021/7/1
- Exercise with Social Robots: Companion or Coach?2021/3/1
- Disambiguating Affective Stimulus Associations for Robot Perception and Dialogue2021/3/1
- Continual Learning from Synthetic Data for a Humanoid Exercise Robot2021/2/1
- Reinforcement Learning with Time-dependent Goals for Robotic Musicians2020/11/1
- Sensorimotor representation learning for an "active self" in robots: A model survey2020/11/1
- Robotic self-representation improves manipulation skills and transfer learning2020/11/1
- Integrating Intrinsic and Extrinsic Explainability: The Relevance of Understanding Neural Networks for Human-Robot Interaction2020/10/1
- Affect-Driven Modelling of Robot Personality for Collaborative Human-Robot Interactions2020/10/1
- Enhancing a Neurocognitive Shared Visuomotor Model for Object Identification, Localization, and Grasping With Learning From Auxiliary Tasks2020/9/1
- Curious Hierarchical Actor-Critic Reinforcement Learning2020/5/1
- Explainable Goal-Driven Agents and Robots -- A Comprehensive Review2020/4/1
- Improving Robot Dual-System Motor Learning with Intrinsically Motivated Meta-Control and Latent-Space Experience Imagination2020/4/1
- Solving Visual Object Ambiguities when Pointing: An Unsupervised Learning Approach2019/12/1
- EDA: Enriching Emotional Dialogue Acts using an Ensemble of Neural Annotators2019/12/1
- Efficient Intrinsically Motivated Robotic Grasping with Learning-Adaptive Imagination in Latent Space2019/10/1
- Hierarchical Control for Bipedal Locomotion using Central Pattern Generators and Neural Networks2019/9/1
- Towards Learning How to Properly Play UNO with the iCub Robot2019/8/1
- From semantics to execution: Integrating action planning with reinforcement learning for robotic causal problem-solving2019/5/1
- Curious Meta-Controller: Adaptive Alternation between Model-Based and Model-Free Control in Deep Reinforcement Learning2019/5/1
- Improving interactive reinforcement learning: What makes a good teacher?2019/4/1
- Deep Intrinsically Motivated Continuous Actor-Critic for Efficient Robotic Visuomotor Skill Learning2018/10/1
- Towards Dialogue-based Navigation with Multivariate Adaptation driven by Intention and Politeness for Social Robots2018/9/1
- Deep Neural Object Analysis by Interactive Auditory Exploration with a Humanoid Robot2018/7/1
- Multi-modal Feedback for Affordance-driven Interactive Reinforcement Learning2018/7/1
- Potentials and Limitations of Deep Neural Networks for Cognitive Robots2018/5/1
- On the Robustness of Speech Emotion Recognition for Human-Robot Interaction with Deep Neural Networks2018/4/1
- EmoRL: Continuous Acoustic Emotion Classification using Deep Reinforcement Learning2018/4/1
- A Neurorobotic Experiment for Crossmodal Conflict Resolution in Complex Environments2018/2/1
- An Incremental Self-Organizing Architecture for Sensorimotor Learning and Prediction2017/12/1