日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
アフォーダンス知覚arXiv:2609.37264

UniAfford: トークンルーティング型マルチタスク学習による汎用可能な2D-3Dアフォーダンス知覚

UniAfford: Token-Routed Multitask Learning for Generalizable 2D-3D Affordance Perception

シェア:XThreadsFacebookLINEはてブBluesky

MLLMを共通の意味ハブとし、タスク別トークンルーターで画像・点群のアフォーダンスクエリを生成することで、2D・3D・統合アフォーダンス知覚を単一枠組みで扱う手法を提案。

詳しい要約

1. どんなもの?

- MLLMベースの2D-3D affordance perceptionを統合する枠組み - 2Dと3Dのaffordance groundingを別問題として扱う分断を解消 - UniAfford: 汎化可能な2D-3D affordance perceptionの統一フレームワーク - UniAfford-Data: pixel-level 2D注釈、point-level 3D注釈、言語指示を統合したデータセット - 共有object-affordance taxonomyの下でheterogeneous supervisionを支援 - image-only、point-cloud-only、paired multimodal入力から2D/3D/合同推論が可能

2. 先行研究と比べてどこがすごい?

- 従来は2Dと3Dのaffordance groundingが別々に発展 - タスク定義、supervision形式、データセット、評価プロトコルが異なる - この分断が視覚空間と幾何空間をまたぐ転移可能なobject-affordance semanticsの学習を制限 - UniAffordは両者を統合し、target-specific fine-tuningなしでzero-shot汎化を実現 - modality-isolated protocols下でbranch-wiseのstate-of-the-art性能も達成

3. 技術・手法の肝は?

- Token Router for Tasks: MLLMベースシステム向けmultitask学習パラダイム - 言語headに事前定義マーカーを生成させず、contextual hidden statesをタスク固有branchへルーティング - routed statesをbranch固有目的関数で直接supervisedし、dense prediction lossが共有MLLM表現を形成 - MLLMを共有semantic hubとして採用 - modality-aware token routerがimage/point-cloud affordance queriesを生成 - 各queryがSAM-style pixel decoderとSONATA-based point decoderを条件付け - semantic-level 2D-3D pairingでheterogeneous supervisionを支援

4. どうやって有効だと検証した?

- 2Dおよび3D affordanceベンチマークでzero-shot汎化を評価 - target-specific fine-tuningなしで強いzero-shot汎化を実証 - modality-isolated protocols下でbranch-wiseのstate-of-the-art性能を達成 - ablationでtoken routing、joint 2D-3D supervision、decoder couplingを検証 - language-head diagnosticsによりrouted latent statesが意味あるobject-affordance semanticsを保持することを確認

5. 議論はある?

- 2D-3D affordance perceptionの分断が転移可能なsemantics学習を制限する点を議論 - token routing、joint 2D-3D supervision、decoder couplingの有効性をablationで検証 - language-head diagnosticsでrouted latent statesの意味的妥当性を確認 - 具体的な限界や失敗事例、計算コスト、データセットの偏りなどは要旨からは不明

6. 次に読むべき論文は?

- SAM-style pixel decoder (SAM) - SONATA-based point decoder (SONATA) - MLLMベースのaffordance grounding関連研究 - 2D affordance groundingの先行研究 - 3D affordance groundingの先行研究 - 具体的な参照論文名は要旨からは不明

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yuhao Liu, Yiming Zhong, Hanqing Wang, Shaocheng Yan, Yuhang Zhang, Wenzhou Lyu, Ziyang Ding, Wei Zhang, Xue Zhao, Jin Pan, Yuexin Ma, Xinge Zhu

分類: cs.CV, cs.AI

原文アブストラクト

Affordance perception aims to localize actionable regions supporting embodied interaction, yet 2D and 3D affordance grounding have evolved as separate problems, with different task definitions, supervision formats, datasets, and evaluation protocols. This fragmentation limits the learning of transferable object-affordance semantics across visual and geometric spaces. We propose Token Router for Tasks, a multitask training paradigm for MLLM-based systems that routes contextual hidden states to task-specific branches without requiring the language head to generate predefined markers. Routed states are supervised directly by branch-specific objectives, enabling dense prediction losses to shape shared MLLM representations. We instantiate this paradigm as UniAfford, a unified framework for generalizable 2D-3D affordance perception, together with UniAfford-Data, a dataset integrating pixel-level 2D annotations, point-level 3D annotations, and language instructions under a shared object-affordance taxonomy, supporting heterogeneous supervision through semantic-level 2D-3D pairing. UniAfford adopts an MLLM as a shared semantic hub and a modality-aware token router to produce image- and point-cloud-affordance queries. These queries respectively condition a SAM-style pixel decoder and a SONATA-based point decoder, enabling flexible 2D, 3D, and joint affordance inference from image-only, point-cloud-only, or paired multimodal inputs. Experiments demonstrate strong zero-shot generalization across 2D and 3D affordance benchmarks without target-specific fine-tuning, alongside state-of-the-art branch-wise performance under modality-isolated protocols. Ablations validate token routing, joint 2D-3D supervision, and decoder coupling, while language-head diagnostics show that routed latent states carry meaningful object-affordance semantics. Project page: https://4dvlab.github.io/UniAfford

PR本紙発行元 EmplifAI