日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
人間-ロボット相互作用arXiv:2610.12245

固定参照ポーズ残差による人間-ロボット相互作用予測のデータセット間手がかり転移の測定

Fixed-Reference Pose Residuals for Measuring Cross-Dataset Cue Transfer in Human-Robot Interaction Anticipation

シェア:XThreadsFacebookLINEはてブBluesky

人とロボットの接触前予測において、あるロボットのデータで学習したモデルを別のロボットに適用した際の手がかりの転移を、幾何情報とポーズ補正を分離する固定参照ポーズ残差モデルで測定した研究。

詳しい要約

1. どんなもの?

本論文は、公共空間のsocial/service robotが、近くの人物が接触前に接近・接触するかを予測するanticipationを対象とする。異なるrobot・異なるsite間で、あるrobotで学習したモデルを別のrobotに適用した際のcue transferを測定するため、fixed-reference pose residual (FRPR) モデルを提案する。FRPRは、人物のbounding boxとmaskから作るgeometry predictorを学習後にfreezeし、temporal networkがbody poseからlogitへのadditive correctionを学習する。これにより各予測がgeometry termとpose termに厳密に分解される。2つのpublic egocentric datasets、HUI360とSSUP-A間で評価する。

2. 先行研究と比べてどこがすごい?

先行研究と比べてどこがすごいかは要旨からは不明。ただし、あるrobotで学習したモデルを別robot・別siteに適用した際のcue transferは largely unknown と述べ、FRPRで予測をgeometry termとpose termに厳密に分解して測定可能にした点が特徴。SSUP-AからHUI360ではpose correctionがAPを0.277から0.321に上げたが、逆方向ではmeasurable gainなしという非対称性を示した。また、freezingはjoint trainingに対しAP advantageなし、simple geometric baselinesやtree ensemblesがcompetitive or betterであり、本構成はprediction accuracyではなくmeasurementのためのものと位置づける。

3. 技術・手法の肝は?

FRPRは、人物のbounding boxとmaskからgeometry predictorを構築し、学習後にfreezeする。次にtemporal networkがbody poseからlogitへのadditive correctionを学習し、予測をgeometry termとpose termに厳密に分解する。source data上で全ての選択を行い、source-selected geometry referenceよりも強いreferenceに対しても同様の非対称性を確認する。さらにhead-orientation residualを加えると両方向でsmall gainsが得られた。post hocでは、personがcameraを向いているかはdatasets間でdiscriminative directionを保ったが、head pitchは反転した。

4. どうやって有効だと検証した?

2つのpublic egocentric datasets、HUI360とSSUP-Aを異なるrobotで記録したものを用い、source data上で全ての選択を行い、SSUP-AからHUI360およびその逆方向で評価した。SSUP-AからHUI360ではpose correctionがAPを0.277から0.321に上げたが、逆方向ではmeasurable gainなし。source-selected geometry referenceよりも強いreferenceでも同じ非対称性。freezingはjoint trainingに対しAP advantageなし。simple geometric baselinesとtree ensemblesはcompetitive or better。head-orientation residualは両方向でsmall gains。source dataでthresholdsを選んだgeometryを使うneural modelsはtarget interactionsのat most 17%しか検出しなかった。

5. 議論はある?

本構成はprediction accuracyではなくmeasurementのために機能する。cue transferには方向依存の非対称性があり、SSUP-AからHUI360ではpose correctionが有効だが逆方向ではmeasurable gainなし。freezingのAP advantageはなく、simple geometric baselinesやtree ensemblesがcompetitive or better。post hocでpersonがcameraを向いているかはdatasets間でdiscriminative directionを保つが、head pitchは反転する。source dataでthresholdsを選ぶとtarget interactionsの検出率はat most 17%と低い。

6. 次に読むべき論文は?

要旨で参照/比較されている研究として、HUI360とSSUP-Aの2つのpublic egocentric datasets、simple geometric baselines、tree ensembles、joint training、source-selected geometry reference、head-orientation residualが挙げられる。関連手法として、human-robot interaction anticipation、cross-dataset cue transfer、pose-based temporal network、geometry predictor、bounding boxとmaskを用いる手法が次の読むべき候補。コードと処理済みデータは https://github.com/WeiZhou96/FRPR-interaction-anticipation で公開。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Bowen Yang, Xinliang Xiao, Wenjing Zhang, Li Yang, Wei Zhou

分類: cs.RO

原文アブストラクト

Social and service robots in public spaces need to anticipate which nearby person is about to approach and touch them, so that a response can be prepared before contact. It is largely unknown which cues support this anticipation when a model trained with one robot is used on another robot at a different site. We study this question with a fixed-reference pose residual (FRPR) model: a geometry predictor built from the person's bounding box and mask is trained and frozen, and a temporal network then learns from body pose an additive correction to its logit, so that every prediction splits exactly into a geometry term and a pose term. Between two public egocentric datasets recorded by different robots, HUI360 and SSUP-A, with every choice made on source data, the pose correction raised average precision (AP) from 0.277 to 0.321 from SSUP-A to HUI360 and gave no measurable gain in the opposite direction; the same asymmetry held over a stronger, source-selected geometry reference. Freezing gave no AP advantage over joint training, and simple geometric baselines and tree ensembles remained competitive or better, so the construction serves measurement rather than prediction accuracy. A head-orientation residual added small gains in both directions. Post hoc, whether a person faces the camera kept its discriminative direction across datasets, whereas head pitch reversed. With thresholds chosen on source data, the neural models that use geometry detected at most 17% of target interactions. Code and processed data are available at https://github.com/WeiZhou96/FRPR-interaction-anticipation.

関連論文

PR本紙発行元 EmplifAI