日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
共有制御arXiv:2609.10215

オンライン有界合理性人間行動推定を用いた適応的共有制御

Adaptive Shared Control with Online Bounded-Rational Human Behavior Estimation

シェア:XThreadsFacebookLINEはてブBluesky

人間の完全合理性を仮定せず、レベルk有界合理性モデルと適応動的計画法で候補方策を構築し、状態遷移残差から人間行動分布をオンライン推定してロボットが期待コスト最小の応答を計算する共有制御手法を提案した。

詳しい要約

1. どんなもの?

- 非線形制御アフィン系における適応型共有人間-ロボット制御を扱う研究。 - 人間が完全に合理的とは限らないという仮定を緩和し、観測された限定合理的な人間行動にロボットが適応する。 - level-k bounded-rationality model を用いて2プレイヤーゲームを構成し、候補となる人間とロボットのポリシーを生成。 - 共有制御中に状態遷移残差を蓄積し、人間行動の確率モデルを推定。 - 推定分布に基づく分布認識型の1ステップ最適応答を計算する。

2. 先行研究と比べてどこがすごい?

- 従来の共有制御では人間の完全合理性を仮定することが多いが、本研究は限定合理性を考慮。 - 単一の候補選択や保存されたロボットポリシーの平均化ではなく、推定された人間行動分布全体に対する期待協調コストを最小化する1ステップ最適応答を計算。 - 最大確率ポリシーや確率重み付きポリシーといったベースラインと比較して、累積実行コストが低いことを示した。

3. 技術・手法の肝は?

- level-k bounded-rationality model に基づき、交互最適応答計算で有限の候補人間・ロボットポリシー集合を構築。 - 価値関数とポリシーは adaptive dynamic programming で近似。 - 共有制御中、状態遷移残差を測定システム進化と候補人間ポリシー予測軌道の比較で算出。 - 残差を忘却係数で累積し、有限候補集合上の確率的人間行動モデルにマッピング。 - ロボットは推定人間行動分布全体で期待協調コストを最小化する分布認識型1ステップ最適応答を計算。 - 二次終端価値近似と Euler 状態伝播の場合、期待人間入力で表される閉形式解が得られる。

4. どうやって有効だと検証した?

- ベンチマーク非線形システム安定化タスクと平面マニピュレータ共有制御セットアップのシミュレーションで評価。 - 推定人間行動分布とシミュレートされた人間行動分布の間の Kullback-Leibler divergence が減少することを報告。 - 共有制御相互作用期間中のロボットエージェントの累積実行コストが、最大確率および確率重み付き代替ポリシーのベースラインよりも低いことを示した。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- level-k bounded-rationality model に関する研究。 - adaptive dynamic programming を用いた共有制御に関する研究。 - 共有人間-ロボット制御におけるベースラインとして言及された maximum-probability および probability-weighted alternative policies に関連する研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Henry Ascencio Trejo, Roel Pieters, Gokhan Alcan

分類: cs.RO, eess.SY

原文アブストラクト

This work considers adaptive shared human-robot control for nonlinear control-affine systems, where the assumption of a fully rational human is relaxed and the robot adapts its assistance to observed boundedly rational human behavior. We use a level-k bounded-rationality model of the two-player game to construct a finite bank of candidate human and robot policies through alternating best-response computations, with the associated value functions and policies approximated using adaptive dynamic programming. During the shared-control interaction, state-transition residuals compare the measured system evolution with the trajectories predicted by the candidate human policies. The residuals are accumulated using a forgetting factor and mapped to a probabilistic human-behavior model over the finite candidate bank. Rather than selecting a single candidate or averaging stored robot policies, the robot computes a distribution-aware one-step best response by minimizing an expected cooperative cost over the complete estimated human behavior distribution. For a quadratic terminal-value approximation and Euler state propagation, this response admits a closed-form solution expressed in terms of the expected human input. The proposed methods are evaluated in simulations of a benchmark nonlinear system stabilization task, and of a planar manipulator shared control setup. The reported results show decreasing Kullback-Leibler divergence between the estimated and simulated human behavior distributions, and a lower accumulated running cost for the robot agent over the shared control interaction period, than the maximum-probability and probability-weighted alternative policies baseline.

関連論文