日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.07745

ずれた座標系を見抜く:視覚・力覚精密組立のための特権ノイズ蒸留

Seeing Through the Displaced Frame: Privileged Noise Distillation for Vision-Force Precision Assembly

シェア:XThreadsFacebookLINEはてブBluesky

精密組立では姿勢誤差が観測と行動の座標系を同時にずらす問題に対し、訓練時のみ特権情報を使う蒸留でノイズに頑健な視覚・力覚方策を学習し、実機Frankaでも高い成功率を達成した。

詳しい要約

1. どんなもの?

視覚と力覚を用いた精密組立において、pose error が観測だけでなく行動の座標系 (coordinate frame) を歪める問題を扱う研究。FORGE benchmark では真の pose なら 97-99% 成功する state-based policy が σ=5 mm の pose-noise で 32-60% に低下する。訓練時のみ得られる privileged 情報 (offset) を student に蒸留し、テスト時は noisy state、raw wrench window、2 台の RGB camera のみで動作する手法を提案。

2. 先行研究と比べてどこがすごい?

- 従来の state-based policy は σ=5 mm で 32-60% 成功にとどまる。 - 提案手法の teacher-route student は σ=0-5 mm で 92-99% 成功を維持。 - 6 つの non-privileged baseline は 2-80% に低下するのに対し、提案は頑健。 - 同一 student 構成・データ量・訓練手順の behavior cloning 比較で、offset 情報なしのデモは σ=5 mm で 31.5% 成功、privileged デモは 88.7%、relabelled デモは 95.8% に達し、頑健性の源泉が supervision にあることを分離。 - Franka への zero-shot 展開で σ=5 mm の pooled 成功 83.3% に対し、最強の state-based policy は 34.4%。

3. 技術・手法の肝は?

- 訓練時のみ利用可能な privileged teacher が simulation で offset を観測。 - あるいは clean demonstration を displaced frame に closed form で relabel。 - 展開する student は behavior cloning の後、1 ラウンドの DAgger で訓練。 - テスト時入力は noisy state、raw wrench window、2 台の RGB camera のみ。 - 展開可能な sensing だけでは不十分で、latent frame offset の補償を符号化する supervision が頑健性に必要。

4. どうやって有効だと検証した?

- FORGE benchmark の未改変タスクで評価。 - teacher-route student は σ=0-5 mm で 92-99% 成功。 - 6 つの non-privileged baseline は 2-80% に低下。 - 同一 student 構成・データ量・訓練手順の behavior cloning 実験で、offset なしデモ 31.5%、privileged デモ 88.7%、relabelled デモ 95.8% (σ=5 mm)。 - Franka に zero-shot 展開し、σ=5 mm で pooled 成功 83.3% (最強 state-based policy は 34.4%)。

5. 議論はある?

- 展開可能な sensing だけでは不十分で、latent frame offset の補償を符号化する supervision が頑健性に必要と主張。 - 訓練時のみ offset にアクセスする privileged 情報の利用が鍵。 - 要旨からは、sim-to-real gap や他のノイズ源、計算コスト、一般化限界についての議論は不明。

6. 次に読むべき論文は?

- FORGE benchmark (本論文で使用) - DAgger (訓練手法) - behavior cloning (比較手法) - state-based policy (比較対象) - privileged learning / privileged teacher (関連手法) - Franka (展開プラットフォーム)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ching-Hsiang Chang, Tzu-Yu Chuang, Yi-Hsiu Lee, Yi-Ting Chen, Yuan-Fu Yang, Min Sun

分類: cs.RO

原文アブストラクト

Pose error in precision assembly can corrupt not only what a robot observes but also the coordinate frame in which it acts. On the FORGE benchmark, the official state-based policy succeeds in 97% to 99% of episodes with the true pose but only 32% to 60% at the benchmark's $σ=5$ mm pose-noise setting. The same estimated pose enters the observation and anchors the action frame, making the offset unidentifiable from proprioceptive state alone before contact. We supply this missing information during training in two ways. A privileged teacher observes the offset in simulation, while clean demonstrations can instead be relabelled into the displaced frame in closed form. The deployed student is trained with behaviour cloning followed by one DAgger round and receives only noisy state, a raw wrench window, and two RGB cameras at test time. On the unmodified FORGE tasks, the teacher-route student maintains 92% to 99% success across $σ=0$ to 5 mm, while six non-privileged baselines fall to 2% to 80%. A matched behaviour-cloning experiment isolates the source of this robustness. With the same student architecture, data budget, and training procedure, demonstrations generated without offset access yield only 31.5% success at $σ=5$ mm, whereas privileged and relabelled demonstrations reach 88.7% and 95.8%. Deployed zero-shot on a Franka, the student reaches 83.3% pooled success at $σ=5$ mm against 34.4% for the strongest state-based policy. Deployable sensing alone is insufficient. Robustness requires supervision that encodes compensation for the latent frame offset.Project page: https://drychang.github.io/displaced-frame/

関連論文

PR本紙発行元 EmplifAI