日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.01726

単一画像からのクエリ条件付き関節構造推定

Query-Conditioned Articulation Estimation from a Single Image

シェア:XThreadsFacebookLINEはてブBluesky

単一のRGB画像とクエリ点から関節物体の運動学パラメータを推定するモデルQueryArtを提案し、実機マニピュレータで70.2%の成功率を達成した。

詳しい要約

1. どんなもの?

- 単一の RGB 画像から物体の articulation parameters を推定するモデル QueryArt を提案。 - 入力は RGB 画像、2D query point、camera intrinsics。 - クエリ点に対する 3D articulation geometry を、その点の depth を単位として推定。 - 単一の depth 測定で scale を復元し metric parameters を得る。 - 合成と実世界の articulation datasets を混合して学習。

2. 先行研究と比べてどこがすごい?

- 従来の単一画像手法は articulation part segmentation と articulation estimation を結合し、検出漏れや誤った part association に弱い。 - また単一視点では scale が不定な metric 3D geometry を回帰していた。 - QueryArt は query point 相対かつ depth 単位で推定するため、画像のみから同定可能。 - 複数の benchmark で最近の baselines を上回り、out-of-distribution データでも優位。

3. 技術・手法の肝は?

- 単一 RGB 画像、2D query point、camera intrinsics を入力とする。 - クエリ点に対する 3D articulation geometry を、その点の depth を単位として推定するよう学習。 - 単一の depth 測定をクエリ点で行い scale を供給して metric parameters を復元。 - 合成と実世界の articulation datasets をキュレーションした混合データで学習。

4. どうやって有効だと検証した?

- 複数の benchmarks で最近の baselines と比較。 - ほとんどの articulation metrics で baselines を上回り、out-of-distribution データでも優位。 - 実世界設定の検証として mobile manipulator 上で 57 の manipulation trials を実施。 - 16 object parts、5 viewpoint classes をカバーし、70.2% の成功率を達成。

5. 議論はある?

- 要旨からは不明。 - 限界や失敗事例、計算コスト、depth 測定の誤差影響などについての議論は要旨に記載なし。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として articulation estimation、part segmentation、single-image 3D reconstruction、mobile manipulation の定番研究を挙げる。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Abdelrhman Werby, Fabio Scaparro, Kai O. Arra

分類: cs.RO

原文アブストラクト

Enabling robots to estimate the kinematic parameters of articulated objects unlocks a wide range of capabilities for interaction and manipulation. The estimation has to happen from the information the robot currently observes, often just a single RGB image of an object it has never seen before. Current single-image approaches couple articulation part segmentation with articulation estimation, making their predictions vulnerable to missed detections and incorrect part associations, and they regress metric 3D geometry that a single view fixes only up to scale. We present QueryArt, a model that estimates articulation parameters from a single RGB image, a 2D query point, and camera intrinsics. QueryArt is trained to estimate the 3D articulation geometry relative to the queried point and in units of its depth, which keeps its target identifiable from the image alone. A single depth measurement at the query point then supplies the scale and recovers the metric parameters. We train QueryArt on a curated mixture of synthetic and real-world articulation datasets. We evaluate QueryArt on several benchmarks and compare it against recent baselines. QueryArt outperforms recent baselines on most articulation metrics, including on out-of-distribution data. To demonstrate the model's capabilities in real-world settings, we evaluate QueryArt on a mobile manipulator across 57 manipulation trials spanning 16 object parts and five viewpoint classes, achieving a 70.2% success rate. We provide code and videos at: https://abwerby.github.io/queryart/

関連論文

PR本紙発行元 EmplifAI