日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.36182

アクチュエータ劣化下におけるマニピュレーションポリシーのテスト時適応

Test-Time Adaptation of Manipulation Policies Under Actuator Degradation

シェア:XThreadsFacebookLINEはてブBluesky

関節温度やモータ電流などのテレメトリを用いて、凍結したマニピュレーションポリシーの出力アクションを補正する軽量Transformerを学習し、アクチュエータ劣化時でも成功率を向上させる手法を提案。

詳しい要約

1. どんなもの?

- アクチュエータ劣化下でのロボットマニピュレーション政策の適応手法 - 提案手法は Telemetry-Aware Action Rectification (TeAR) - 凍結した政策をテレメトリ条件付き政策に変換 - 低レベルコントローラに送る前の行動を修正 - 軽量 Transformer で行動とテレメトリを統合 - 個々の行動成分を増幅・減衰・バイアス - 18 の政策-タスクペア、8 政策ファミリ、5 タスクで評価

2. 先行研究と比べてどこがすごい?

- 従来はテレメトリをログや安全チェックのみに使用 - 政策適応には活用されていなかった - ベース政策や仮定モデル逆変換と比較 - 劣化モデル不一致のペア評価で成功率 31.8% - ベース政策 25.6%、仮定モデル逆 30.6% を上回る - 物理アームで加熱下の成功率を 10-15% 改善 - ロボット上でのファインチューニング不要

3. 技術・手法の肝は?

- 政策非依存の手法 - 凍結した政策の出力行動を修正 - 軽量 Transformer を学習 - 提案行動とライブアクチュエータテレメトリを結合 - テレメトリには関節温度、モーター電流、供給電圧を含む - 行動成分ごとに増幅、減衰、バイアスを適用 - 低レベルコントローラに到達する前に修正

4. どうやって有効だと検証した?

- 18 の政策-タスクペアで評価 - 8 政策ファミリ、5 マニピュレーションタスク - 劣化モデル不一致の追加ペア評価 - 成功率をベース政策、仮定モデル逆と比較 - 物理アームで加熱下の評価 - ロボット上ファインチューニングなしで 10-15% 改善

5. 議論はある?

- 要旨からは不明 - 限界や失敗事例についての記述なし - 計算コストやリアルタイム性の議論なし - 一般化可能性の議論なし

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない - 関連手法として強化学習、模倣学習、ドメイン適応、Test-Time Adaptation が考えられる - 同分野の定番として Behavior Cloning、Diffusion Policy、ACT などが挙げられる

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Som Sagar, Ransalu Senanayake

分類: cs.RO

原文アブストラクト

Robot manipulation policies are usually trained under the assumption that a commanded action produces the same motion as it did during training even after hours of operation. Real hardware violates this assumption as the motors gradually heat up, current saturates near contact, voltage sags under load, thus the same policy action can produce a weaker, delayed, or noisier motion. These conditions are already measured by onboard telemetry, such as joint temperature, motor current, and supply voltage, yet this signal is typically used only for logging or safety checks rather than policy adaptation. We introduce Telemetry-Aware Action Rectification (TeAR), a policy-agnostic method that turns a frozen manipulation policy into a telemetry-conditioned policy by rectifying its outgoing action before it reaches the low-level controller. TeAR learns a lightweight Transformer that combines the proposed action with live actuator telemetry and amplifies, damps, or biases individual action components. We evaluate TeAR across 18 policy-task pairs spanning 8 policy families and 5 manipulation tasks. In an additional paired evaluation with degradation-model mismatch, TeAR achieves 31.8% success, compared with 25.6% for the base policy and 30.6% for an assumed-model inverse. On a physical arm, TeAR improves success under heating by 10-15% without on-robot fine-tuning.

関連論文

PR本紙発行元 EmplifAI