GLASS: エッジロボティクス向けのアーキテクチャ適応型・構成可能なデバイスサイド線形代数ライブラリ
GLASS: Architecture-Tuned, Composable, Device-Side Linear Algebra for Edge Robotics and Beyond
GPU上でロボティクス向け線形代数・幾何計算をスレッド/ワープ/ブロック単位で実行できるヘッダオンリーCUDA C++ライブラリGLASSを提案し、実装選択をアーキテクチャごとにオフライン計測で最適化することでエッジGPUで最大73倍の高速化を実現した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Brian Plancher
分類: cs.RO, cs.DC, cs.MS
原文アブストラクト
GPU robotics lacks the reusable numerical infrastructure of mature CPU stacks, instead relying on compiler frameworks that introduce overhead or repeatedly reimplementing numerical libraries. To address this, we introduce GLASS (GPU Linear Algebra Simple Subroutines), a header-only CUDA C++ library that provides thread-, warp-, block-, and NVIDIA-backed implementations of robotics-scale linear algebra and geometric computations under one composable device API. GLASS treats implementation choice, execution scope, and launch packing as architecture-specific placement decisions determined by offline measurement and resolved statically at compile time. This is critical as the best and worst placements differ by a median of 4.9x (max 81x), with 145 of 396 recommended placements changing between a Jetson AGX Orin and an RTX 5090, and 162 of 396 versus an AGX Xavier. These stakes are highest at the edge as GLASS's advantage over the best of PyTorch and JAX is as much as 73x on the Orin versus 12x on the RTX 5090. GLASS is released open source with independent numerical oracles and source-bound local-GPU test attestation. Finally, integrating GLASS with published robotics systems both exposed a pre-existing numerical bug and improved embedded runtimes by up to 1.5x.