日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
アニメーション圧縮arXiv:2610.04211

CurveCodec 2: 学習型エントロピーモデルによるスケルトン非依存のアニメーション圧縮

CurveCodec 2: Skeleton-agnostic animation compression with a learned entropy model

シェア:XThreadsFacebookLINEはてブBluesky

骨格アニメーションをlog map上の曲線として符号化し、学習したエントロピーモデルで残差を圧縮することで、任意のスケルトンに対して誤差上限を保証しつつACLの最悪ケースを達成するコーデックを提案。

詳しい要約

1. どんなもの?

- 骨格アニメーションの新しい圧縮コーデック CurveCodec 2 を提案。 - 各 joint の transform を全フレーム保存する冗長性を削減する。 - 任意の skeleton に対応する skeleton-agnostic な設計。 - 誤差上限を明示した契約 (worst-case / mean) を満たす。 - 学習した entropy model で residual を符号化する。

2. 先行研究と比べてどこがすごい?

- 先行の CurveCodec は ACL の平均誤差に並ぶが最悪誤差で劣り、payload を float で数えていた。 - 本手法は ACL の worst case を許容内で満たす契約を検証。 - 既定精度 0.01 cm で ACL の 0.37x バイト、0.1 cm で 0.22x。 - 学習モデルの整数推論がプラットフォーム間で bit-exact。 - 再学習なしで未学習の species に転移する。

3. 技術・手法の肝は?

- 各 sub-track を log map 上の curve として符号化。 - closed loop で量子化し、rate-distortion 選択した key に間引く。 - 階層を通した closed loop で joint ごとに符号化しない sample を選ぶ。 - residual を小さな学習 entropy model で entropy coding。 - 学習モデルの整数推論を bit-exact に実装。

4. どうやって有効だと検証した?

- 33 データセット、4,472 クリップの held-out test 側で評価。 - 各復号クリップで ACL の worst case または mean 契約を検証。 - 既定精度 0.01 cm で 0.37x、0.1 cm で 0.22x のバイト数。 - 1 CPU コアで復号できることを確認。 - 未学習 species への再学習なし転移を確認。

5. 議論はある?

- 冗長性の所在と学習モデルの役割を測定で分析。 - 最大の削減は各量子化 curve を自身の過去から予測することから得る。 - 次に joint ごとの sample 選択が寄与する。 - 数百万サンプルの nearest-neighbour oracle は線形補間を上回らない。 - 試した learned in-betweener はコストに見合わなかった。

6. 次に読むべき論文は?

- ACL (production library of modern game engines)。 - CurveCodec (先行 codec)。 - log map を用いた curve 表現。 - rate-distortion 選択と entropy coding の関連手法。 - nearest-neighbour oracle と linear interpolation の比較。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Mingyi Shi, Huancheng Lin, Xuelin Chen, Taku Komura

分類: cs.GR, cs.AI, cs.RO

原文アブストラクト

Skeletal motion is stored as every joint's transform at every frame, yet most of it is implied by the body rather than by what the motion is about. Compression is one way to ask what a motion must still say once the body is known, and a production codec must answer it for any skeleton with a stated error bound. Our earlier codec, CurveCodec, matched the mean error of ACL, the production library of modern game engines, with a learned prior over sparse anchors, but not ACL's worst case, and it counted its payload as floats rather than bits. Here we ask where the redundancy of skeletal motion lies and which part of a codec a learned model should take over. Measurements give three answers. At production precision the largest saving comes from predicting each quantized curve from its own past, the second from choosing per joint, in closed loop through the hierarchy, which samples not to code. On the gaps such an encoder leaves, a nearest-neighbour oracle over millions of training samples is no better than linear interpolation, and no learned in-betweener we tried paid for itself. What a network does learn is the distribution of the residuals the codec must send. CurveCodec 2 codes every sub-track as a curve in the log map, quantized in closed loop and thinned to rate-distortion-selected keys, with residuals entropy-coded under a small learned model whose integer inference is bit-exact across platforms. Two contracts are verified on every decoded clip: ACL's own worst case per joint within a stated tolerance, or ACL's mean error per clip. On a held-out test side of 4,472 clips from 33 datasets, CurveCodec 2 needs 0.37x ACL's bytes at ACL's default precision of 0.01 cm under the worst-case contract and 0.22x at 0.1 cm under the mean contract, decodes on one CPU core, and transfers without retraining to a species absent from training. Project page: https://rubbly.cn/publications/curvecodec/

PR本紙発行元 EmplifAI