日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
線形代数/エッジGPUarXiv:2609.28179

GLASS: エッジロボティクス向けのアーキテクチャ適応型・構成可能なデバイスサイド線形代数ライブラリ

GLASS: Architecture-Tuned, Composable, Device-Side Linear Algebra for Edge Robotics and Beyond

シェア:XThreadsFacebookLINEはてブBluesky

GPU上でロボティクス向け線形代数・幾何計算をスレッド/ワープ/ブロック単位で実行できるヘッダオンリーCUDA C++ライブラリGLASSを提案し、実装選択をアーキテクチャごとにオフライン計測で最適化することでエッジGPUで最大73倍の高速化を実現した。

詳しい要約

1. どんなもの?

- GLASS (GPU Linear Algebra Simple Subroutines) は、edge robotics 向けの header-only CUDA C++ ライブラリ。 - thread-, warp-, block-, NVIDIA-backed の実装を一つの composable device API で提供。 - robotics-scale の線形代数と幾何計算を対象とする。 - 実装選択・実行スコープ・launch packing をアーキテクチャ固有の配置決定として扱い、オフライン計測に基づきコンパイル時に静的に解決する。

2. 先行研究と比べてどこがすごい?

- GPU robotics は成熟した CPU スタックのような再利用可能な数値基盤がなく、compiler frameworks のオーバーヘッドや数値ライブラリの再実装に依存していた。 - GLASS は単一の composable device API の下で複数スコープの実装を提供し、配置決定を静的に解決する点で異なる。 - 最良と最悪の配置の差は中央値 4.9x (最大 81x) に達し、Jetson AGX Orin と RTX 5090 の間で 396 中 145 の推奨配置が変化、AGX Xavier との間では 396 中 162 が変化する。 - edge での優位性は PyTorch と JAX の最良と比べ Orin で最大 73x、RTX 5090 で 12x。

3. 技術・手法の肝は?

- header-only CUDA C++ ライブラリとして実装。 - thread-, warp-, block-, NVIDIA-backed の各スコープの線形代数・幾何計算カーネルを提供。 - 実装選択、実行スコープ、launch packing をアーキテクチャ固有の placement 決定とみなし、オフライン計測で決定してコンパイル時に静的に解決する。 - 独立した numerical oracles と source-bound local-GPU test attestation を備える。

4. どうやって有効だと検証した?

- 独立した numerical oracles と source-bound local-GPU test attestation を提供。 - Jetson AGX Orin、RTX 5090、AGX Xavier 上で配置の推奨が変化することを計測。 - PyTorch と JAX の最良実装との比較で Orin 最大 73x、RTX 5090 で 12x の優位を確認。 - 公開された robotics systems に統合し、既存の数値バグを発見し、embedded runtime を最大 1.5x 改善。

5. 議論はある?

- 最良と最悪の配置の差が大きく、アーキテクチャ間で推奨配置が変化するため、静的な配置決定の重要性が議論される。 - edge での優位性が特に高いことが示される。 - 統合により既存の数値バグが露呈した点も議論の対象。 - 要旨からは、限界や今後の課題についての明示的な議論は不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究: PyTorch, JAX。 - 関連手法: compiler frameworks, 数値ライブラリの再実装。 - 同分野の定番: CUDA, cuBLAS, Eigen, Ceres Solver などが考えられるが、要旨での直接参照はない。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Brian Plancher

分類: cs.RO, cs.DC, cs.MS

原文アブストラクト

GPU robotics lacks the reusable numerical infrastructure of mature CPU stacks, instead relying on compiler frameworks that introduce overhead or repeatedly reimplementing numerical libraries. To address this, we introduce GLASS (GPU Linear Algebra Simple Subroutines), a header-only CUDA C++ library that provides thread-, warp-, block-, and NVIDIA-backed implementations of robotics-scale linear algebra and geometric computations under one composable device API. GLASS treats implementation choice, execution scope, and launch packing as architecture-specific placement decisions determined by offline measurement and resolved statically at compile time. This is critical as the best and worst placements differ by a median of 4.9x (max 81x), with 145 of 396 recommended placements changing between a Jetson AGX Orin and an RTX 5090, and 162 of 396 versus an AGX Xavier. These stakes are highest at the edge as GLASS's advantage over the best of PyTorch and JAX is as much as 73x on the Orin versus 12x on the RTX 5090. GLASS is released open source with independent numerical oracles and source-bound local-GPU test attestation. Finally, integrating GLASS with published robotics systems both exposed a pre-existing numerical bug and improved embedded runtimes by up to 1.5x.

PR本紙発行元 EmplifAI