日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
HRIarXiv:2609.34220

mmHRI: ミリ波レーダによるプライバシー保護型ヒューマンロボットインタラクション

mmHRI: Towards Privacy-Preserving Human-Robot Interaction with Millimeter-Wave Radar

シェア:XThreadsFacebookLINEはてブBluesky

ミリ波レーダで人の動作と3D姿勢を推定し、それをテキスト指示に変換してVLAポリシーを制御する、プライバシー保護型のロボットマニピュレーション枠組みを提案した。

詳しい要約

1. どんなもの?

- 支援ロボットが人間中心環境で行うHRIタスク(物体配送など)を対象とする。 - 既存のHRIはRGBカメラで人間を継続観察し、手ジェスチャー等の非言語コマンドに応答するが、病棟やレストランなどプライバシー重視環境ではカメラ観察が制限される。 - 本研究はmmWave radarを用い、識別可能な画像を取得せずに人間の動きを感知するプライバシー保護HRIを実現する。 - 提案するmmHRIは、mmWave radar誘導のプライバシー保護HRIを達成する初のマルチモーダルロボットマニピュレーションフレームワークである。

2. 先行研究と比べてどこがすごい?

- 既存HRIはRGBカメラ依存で、プライバシー制約環境に適用困難。 - 既存のradarベース代替手法と比較して、プライバシー保護カーテン設定下で85.09%の行動認識精度を達成し、優位性を示す。 - ロボット試行では視覚遮蔽下でも配送・回収に成功し、未見被験者、 clutter 構成、環境をまたいで安定したタスク性能を実証。 - 詳細な先行研究との比較は要旨からは不明。

3. 技術・手法の肝は?

- クラッタの多いロボットマニピュレーション環境におけるradarデータの疎性と時間的不整合を緩和する2つの設計を導入。 - 第一に、未フィルタの生radar tensorとradar point cloudの両方から共同学習するdual-stream architectureを提案し、人間の行動と3D poseを推定。 - 第二に、信号不整合を緩和するため、履歴radar特徴を保持するmemory-based state-space model (MSSM)を組み込み、pose/actionの急変を低減。 - 推定された人間状態を構造化テキストのロボット指示に変換し、vision-language-action (VLA) policyを制御して閉ループのロボットマニピュレーションと人間認識反応を実現。

4. どうやって有効だと検証した?

- 評価は人間行動認識と閉ループ配送・回収を対象。 - プライバシー保護カーテン設定で85.09%の行動認識精度を達成し、既存radarベース代替手法を上回る。 - ロボット試行により視覚遮蔽下での配送・回収成功を実証。 - 未見被験者、 clutter 構成、環境をまたいで安定したタスク性能を示す。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、radar-based HRI、vision-language-action (VLA) policy、memory-based state-space model (MSSM) が挙げられる。 - 同分野の定番として、RGBカメラベースHRI、mmWave radarによる人間行動認識、ロボットマニピュレーションのVLA手法が考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Junqiao Fan, Yuxuan Hu, Bofan Lyu, Yanshuo Lu, Pengfei Liu, Jiarui Zhang, Fangqiang Ding, Lihua Xie, Gen Li, Jianfei Yang

分類: cs.RO, cs.CV

原文アブストラクト

Assistive robots increasingly operate in many human-centered environments and perform various human-robot interaction (HRI) tasks, such as object delivery. However, most existing HRI systems rely on RGB cameras that continuously observe humans to respond to non-verbal commands, such as hand gestures. This raises privacy concerns in privacy- critical environments, such as hospital wards or restaurants, where direct camera observation of humans is restricted. To develop privacy-preserving HRI, we leverage millimeter-wave (mmWave) radar, which can sense human motion through privacy barriers without identifiable imagery. We propose mmHRI, the first multi-modal robot manipulation framework that achieves mmWave radar-guided privacy-preserving HRI. mmHRI introduces two key designs to mitigate the sparsity and temporal inconsistency of radar data in cluttered robot manipulation environments. First, we propose a dual-stream architecture that jointly learns from unfiltered raw radar tensors and radar point clouds to estimate both human actions and 3D poses. To mitigate signal inconsistency, mmHRI further incorporates a memory-based state-space model (MSSM) that retains historical radar features to reduce abrupt changes in pose/action. These estimated human states are then converted into structured textual robot instructions, which control a vision-language-action (VLA) policy for closed-loop robot manipulation and human-aware reactions. Our evaluation covers human action recognition and closed-loop delivery and retrieval. In the privacy-preserving curtain setting, mmHRI achieves 85.09% action-recognition accuracy, outperforming existing radar-based alternatives. Robot trials further demonstrate successful delivery and retrieval under visual occlusion, with stable task performance across unseen subjects, clutter configurations, and environments.

関連論文

PR本紙発行元 EmplifAI