日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.00904

スクリューアテンション:Transformer内部に剛体代数を組み込む

Screw Attention: Rigid-Body Algebra Inside a Transformer

シェア:XThreadsFacebookLINEはてブBluesky

各トークンを姿勢を持つ剛体として扱い、ペア間の相対姿勢と関節スクリューをメッセージ伝達に用いるTransformer層を提案。フレーム不変な注意機構により、少ないパラメータで操作タスクの性能と幾何変化への頑健性を大幅に向上させた。

詳しい要約

1. どんなもの?

- Transformer layer の新手法「Screw Attention」を提案。 - 各 token を pose を持つ body とし、token 間の関係を graph edge ではなく spatial transform として扱う。 - 各 token ペアは relative pose と、robot joint の場合は joint screw を持つ。 - message はこの関係に沿って receiver の frame に輸送され、attention score は frame-invariant な量のみを見る。 - 構成上、各 token での独立な frame 変更に対して message が equivariant で、単一層で rigid-body mechanics の velocity recursion を表現できる。

2. 先行研究と比べてどこがすごい?

- 学習された manipulation policy は rigid-body mechanics が閉形式で与える spatial relation をデータから再発見しており、データコストと幾何変化への脆弱性がある。 - Screw Attention は relation を spatial transform として組み込む。 - 同サイズの graph、transformer、flat network と比較して同等以上。 - 16,162 パラメータで LIBERO-Spatial の 97.3% を達成し、27 倍のパラメータを持つ flat network を上回る。 - per-link frame convention 変更下でも成功率が不変で、他の学習ネットワークは 3% 未満に低下。 - analytic controller への gated residual として挿入すると insertion 成功率が 17.3 ポイント向上。

3. 技術・手法の肝は?

- Transformer layer 内で、2 body 間の関係を spatial transform として表現。 - 各 token は pose を持つ body。 - token ペアごとに relative pose と、robot joint の場合は joint screw を保持。 - message を relation に沿って receiver の frame に輸送。 - attention score は frame-invariant な量のみを参照。 - 構成上、各 token での独立な frame 変更に対して message が equivariant。 - 単一層で rigid-body mechanics の velocity recursion を表現可能。

4. どうやって有効だと検証した?

- シミュレーション manipulation タスクで検証。 - LIBERO-Spatial において graph、transformer、flat network と比較。 - 16,162 パラメータで 97.3% を達成(画像・言語なし、object pose から)。 - per-link frame convention 変更下で成功率不変、他の学習ネットワークは 3% 未満。 - analytic controller への gated residual として insertion 成功率が 17.3 ポイント向上。 - pose noise 10mm までと、Franka arm の factory calibration 内の joint offset の影響を受けない。

5. 議論はある?

- 結果から「geometry は、システムの他の部分が供給しない frame 間の関係をタスクが要求するときに決定的である」という基準を提案。 - その他の議論や限界については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:graph、transformer、flat network、LIBERO-Spatial、analytic controller、Franka arm。 - 関連手法として rigid-body mechanics、screw theory、equivariant transformer が挙げられる。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Aly Magassouba

分類: cs.RO, cs.AI

原文アブストラクト

Learned manipulation policies rediscover from data the spatial relations that rigid-body mechanics supplies in closed form. This costs data, and it leaves the policies fragile to geometric changes in the scene. We present Screw Attention, a transformer layer in which the relation between two bodies is a spatial transform rather than a graph edge. Every token is a body with a pose. Each pair of tokens carries the relative pose and, for robot joints, the joint screw. Messages are transported along this relation into the receiver's frame, while the attention scores see only frame-invariant quantities. By construction, the messages are equivariant to an independent change of frame at every token, and a single layer can express the velocity recursion of rigid-body mechanics. On simulated manipulation tasks, Screw Attention matches or outperforms controls of the same size, including graph, transformer and flat networks on LIBERO-Spatial. With 16,162 parameters it reaches 97.3% on LIBERO-Spatial from object poses (without images or language), above a flat network with 27x more parameters. Under a change of per-link frame convention its success is unchanged, while every other learned network falls below 3%. Placed on an analytic controller as a gated residual, it raises insertion success by 17.3 points. It is unaffected by pose noise up to 10,mm and by joint offsets within the factory calibration of a Franka arm. These results suggest a criterion: geometry is decisive when the task requires relations between frames that no other part of the system supplies. Code and trained policies will be released.

関連論文

PR本紙発行元 EmplifAI