日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ソーシャルナビゲーションarXiv:2609.28467

どこに加わるべきか?言語誘導型目標予測によるロボットのグループ参加

Where Should I Join? Robot Group Joining via Language-Guided Goal Prediction

シェア:XThreadsFacebookLINEはてブBluesky

自然言語で指定されたグループにロボットが参加するため、言語条件付きモデルでグループを特定し、人間の隊形事前分布を用いて社会的に適切な参加位置を予測する手法を提案した。

詳しい要約

1. どんなもの?

- 言語指示に基づくロボットのグループ参加を定式化 - 観測と自然言語記述から対象グループを特定し、社会的に適切な参加位置を予測 - 会話、行列、観客などの状況で評価

2. 先行研究と比べてどこがすごい?

- 従来のsocial navigationは目標が指定される前提 - 本研究はグループのリアルタイム活動と形成から参加位置を予測する必要 - 言語接地と参加位置予測を統合し、ベースラインを上回る性能

3. 技術・手法の肝は?

- 再帰的スペクトル分割で候補サブセットを生成 - 言語条件付き画像-幾何モデルでランキング - 人間形成の事前分布を利用したマルチモーダルエネルギー-オリエンテーションマップで参加位置を予測

4. どうやって有効だと検証した?

- 会話、行列、観客のデータセットで評価 - グループサイズ、群集密度、視覚的曖昧さを変化させて検証 - 接地精度で競争力、参加位置予測で全ベースラインを上回る - 実ロボット実験で静的・動的相互作用を実証

5. 議論はある?

- 要旨からは不明

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 関連手法としてsocial navigation、language grounding、spectral partitioning、energy-based pose predictionが挙げられる

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zilin Fang, Zishuo Wang, Gim Hee Lee, David Hsu

分類: cs.RO, cs.AI

原文アブストラクト

Social navigation typically assumes a specified goal and focuses on reaching it while respecting social conventions, whereas robot group joining requires predicting where to join based on the group's real-time activity and formation. This is a highly semantic task, yet an important capability for applications such as robotic guide dogs and autonomous mobility scooters. We formulate language-grounded robot group joining: given an observation and a natural-language description of a target group, the robot identifies the relevant group members and predicts socially compliant joining poses. For grounding, we generate structured candidate subsets through recursive spectral partitioning and rank them with a language-conditioned image--geometry model. Given the grounded group, a goal predictor leverages human-formation priors to produce a multimodal energy--orientation map over feasible robot poses. Experiments on conversations, queues, and audiences across varying group sizes, crowd densities, and visual ambiguities show that our method achieves competitive grounding accuracy with sub-second inference and outperforms all baselines in joining-pose prediction. Real-robot experiments further demonstrate group joining in both static and dynamically changing interactions.

関連論文

PR本紙発行元 EmplifAI