日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
強化学習/ロバスト制御arXiv:2610.03249

自己修復型リカレントアンサンブルによる分布シフトからのリアルタイム回復

Self-Repairing Recurrent Ensembles for Real-Time Recovery from Distribution Shift

シェア:XThreadsFacebookLINEはてブBluesky

センサドリフトや故障による分布シフトに対し、マスク付きリカレントネットのアンサンブルとカルマン融合で自己教師ラベルを作り、RFLOでオンライン微調整して回復する手法。

詳しい要約

1. どんなもの?

- 事前学習済みコントローラが、センサドリフト、センサ故障、測定ノイズなどによる分布シフトに直面した際、オンラインかつ教師なしでポリシーを回復させる手法。 - コントローラはリカレントネットワークのアンサンブルで、各メンバーは観測ベクトルのランダムにマスクされたサブセットを観測し、ガウス出力を逐次Kalman fusionで統合する。 - 展開時、コンセンサスを自己教師ラベルとして各メンバーを微調整し、各メンバーの寄与をKalman gainの二乗の補数に比例してスケーリングする。 - 勾配はRFLO(Real-Time Recurrent Learningの効率的で生物学的に妥当な近似)を用いて計算され、環境ステップごとにパラメータ更新が行われる。 - シミュレーション連続制御タスクで、センサシフト後に元の性能に近い回復を達成。完全な観測を見るアンサンブルは回復できない。 - 同じフレームワークは完全オンライン対話型模倣学習も包含し、専門家が存在する場合、コンセンサスラベルを専門家の行動に置き換え、同一の更新規則でテレオペレーション中にポリシーを洗練する。

2. 先行研究と比べてどこがすごい?

- 従来のアンサンブル手法は完全な観測を見るため、分布シフト後に回復できないのに対し、本手法はランダムマスクされたサブセットを観測するメンバーとKalman fusionにより、シフトにオンラインで適応できる。 - 教師なしで自己教師ラベルを用いるため、専門家の修正ラベルが利用できない状況でも回復可能。 - RFLOにより、Real-Time Recurrent Learningの効率的な近似で、各環境ステップ後にパラメータ更新が可能となり、シフトが展開するにつれてポリシーが反応する。 - 完全オンライン対話型模倣学習も同一フレームワークで扱える点が新しい。

3. 技術・手法の肝は?

- リカレントネットワークのアンサンブル:各メンバーは観測ベクトルのランダムにマスクされたサブセットを観測。 - ガウス出力を逐次Kalman fusionで統合し、自信のあるメンバーがコンセンサス行動を支配。 - 展開時、コンセンサスを自己教師ラベルとして各メンバーを微調整。各メンバーの寄与はKalman gainの二乗の補数に比例してスケーリング。 - 勾配計算にRFLO(Real-Time Recurrent Learningの効率的で生物学的に妥当な近似)を使用。環境ステップごとにパラメータ更新。 - 専門家が存在する場合、コンセンサスラベルを専門家行動に置き換え、同一更新規則でテレオペレーション中にポリシーを洗練。

4. どうやって有効だと検証した?

- シミュレーション連続制御タスクで評価。 - センサシフト後、本手法は元の性能に近い回復を達成。 - 完全な観測を見るアンサンブルは回復できないことを確認。 - 完全オンライン対話型模倣学習の枠組みも同一フレームワークで機能することを示唆。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:Real-Time Recurrent Learning (RTRL)、RFLO、Kalman fusion、オンライン対話型模倣学習。 - 関連手法:リカレントネットワークのアンサンブル、自己教師あり学習、分布シフトへの適応。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Julian Lemmel, Pedro D. Wendel Garcia, Taisuke Kobayashi, Radu Grosu

分類: cs.NE, cs.RO, eess.SY

原文アブストラクト

Deploying a pretrained controller exposes it to conditions that are absent from its training data. Sensor drift, outright sensor failure and accumulating measurement noise all induce a distribution shift that can collapse an otherwise competent policy; typically at a point in time where no expert is available to supply corrective labels. We present a method that lets a policy recover from such shifts online and without supervision. Our controller is an ensemble of recurrent networks, each of which observes a randomly masked subset of the observation vector, and whose Gaussian outputs are combined through sequential Kalman fusion so that confident members dominate the consensus action. At deployment, we treat this consensus as a self-supervised label and fine-tune each member towards it, scaling each member's contribution proportional to the complement of its squared Kalman gain. Gradients are computed using RFLO, an efficient and biologically plausible approximation of Real-Time Recurrent Learning, so that a parameter update follows every environment step and the policy reacts to a shift as it unfolds. On a range of simulated continuous control tasks, our approach recovers close to the original performance after a sensor shift, while ensembles that see the full observation are unable to recover. The same framework subsumes fully online interactive imitation learning: when an expert is present, the consensus label is replaced by the expert action and the identical update rule refines the policy during teleoperation.

関連論文

PR本紙発行元 EmplifAI