スクリューアテンション:Transformer内部に剛体代数を組み込む
Screw Attention: Rigid-Body Algebra Inside a Transformer
各トークンを姿勢を持つ剛体として扱い、ペア間の相対姿勢と関節スクリューをメッセージ伝達に用いるTransformer層を提案。フレーム不変な注意機構により、少ないパラメータで操作タスクの性能と幾何変化への頑健性を大幅に向上させた。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Aly Magassouba
分類: cs.RO, cs.AI
原文アブストラクト
Learned manipulation policies rediscover from data the spatial relations that rigid-body mechanics supplies in closed form. This costs data, and it leaves the policies fragile to geometric changes in the scene. We present Screw Attention, a transformer layer in which the relation between two bodies is a spatial transform rather than a graph edge. Every token is a body with a pose. Each pair of tokens carries the relative pose and, for robot joints, the joint screw. Messages are transported along this relation into the receiver's frame, while the attention scores see only frame-invariant quantities. By construction, the messages are equivariant to an independent change of frame at every token, and a single layer can express the velocity recursion of rigid-body mechanics. On simulated manipulation tasks, Screw Attention matches or outperforms controls of the same size, including graph, transformer and flat networks on LIBERO-Spatial. With 16,162 parameters it reaches 97.3% on LIBERO-Spatial from object poses (without images or language), above a flat network with 27x more parameters. Under a change of per-link frame convention its success is unchanged, while every other learned network falls below 3%. Placed on an analytic controller as a gated residual, it raises insertion success by 17.3 points. It is unaffected by pose noise up to 10,mm and by joint offsets within the factory calibration of a Franka arm. These results suggest a criterion: geometry is decisive when the task requires relations between frames that no other part of the system supplies. Code and trained policies will be released.
関連論文
- SkeleWAM: 骨格ワールドアクションモデリングによる効率的なロボットマニピュレーションマニピュレーション
- FlashDexRetarget: 多動作リターゲティングによる器用操作データ生成の高速化マニピュレーション
- 解像度に一貫したヤコビアン場を学習する生体模倣剛柔指マニピュレーション
- Recova: 自律ロボットマニピュレーションのためのエージェント誘導型失敗回復マニピュレーション
- 経験と実演による6自由度把持合成の継続学習マニピュレーション
- 再構成・練習・実世界展開:身体性エージェントのためのガイド付き自己改善マニピュレーション