日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.24059

双腕モバイルマニピュレーションのための自動ラベリング

Automatic Labelling for Bimanual Mobile Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

同期した運動信号を位相に分割し、視覚言語推論で意味づけを行うことで、双腕モバイルマニピュレーションのサブタスクに自動でラベルを付与するパイプラインを提案し、実タスクで有効性を検証した。

詳しい要約

1. どんなもの?

- 双腕モバイルマニピュレーションのための自動ラベリングパイプラインを提案。 - 決定論的軌道解析で時間的境界を、vision-language (VL) 推論で意味解釈を割り当てる。 - 同期した運動信号をフェーズに分割し、フェーズ局所的なVL推論で内容を記述。 - base、left arm、right armの行動について出力を集約する。 - 主に29の実Galaxea双腕モバイルマニピュレーションタスクで評価。

2. 先行研究と比べてどこがすごい?

- 従来は長期horizonポリシー向けの意味的サブタスクラベルにおいて、信頼できる時間的境界と広い意味記述の自動特定が困難だった。 - 本手法は時間的局所化を決定論的軌道解析に、意味解釈をVL推論に分離することで、非同期な双腕行動を保持しつつ構造化アノテーションを生成できる点が新しい。 - 具体的な先行研究との比較は要旨からは不明。

3. 技術・手法の肝は?

- 同期したkinematic signalsをフェーズに分割する決定論的軌道解析。 - フェーズごとに局所化したVL推論を実行し、内容を記述。 - base、left arm、right armの行動ごとに出力を集約。 - これにより時間的境界と意味記述を組み合わせた自動ラベリングを実現。

4. どうやって有効だと検証した?

- 29の実Galaxea双腕モバイルマニピュレーションタスクで評価。 - VL推論を3回繰り返し、選択タスクで87.4%が同じ出力値。 - 9名の参加者による全29タスクのレビューで、時間分割(90.5%)、身体ラベル(90.7%)、腕ラベル(78.7%)に肯定的受容。

5. 議論はある?

- セグメンテーション-VL設計が非同期な双腕行動を保持しつつ構造化アノテーションを生成できることを示唆。 - より豊かな意味的サブタスク識別と状態ベース検証の基盤を提供。 - 限界や課題についての具体的な議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 同分野の定番として、vision-language models (VLMs) を用いたロボティクス、長期horizonポリシー、双腕マニピュレーションに関する研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yupu Lu, Jia Pan

分類: cs.RO

原文アブストラクト

Semantically meaningful subtask labels can provide useful contexts for long-horizon policies, but automatically identifying both reliable temporal boundaries and broad semantic descriptions for annotations remains difficult. We present an automatic labelling pipeline that assigns temporal localisation to deterministic trajectory analysis and semantic interpretation to vision-language (VL) reasoning. The pipeline segments synchronised kinematic signals into phases, performs phase-localised VL reasoning to describe the contents, and aggregates the outputs for the base, left arm, and right arm actions. We evaluate this pipeline primarily on 29 real Galaxea bimanual mobile-manipulation tasks. Repeating the VL reasoning three times first produces the same output value for 87.4% on selected tasks. A review by nine participants across all 29 tasks then judgements on the labelled phases and shows positive acceptance of temporal divisions (90.5%), body labels (90.7%), and arm labels (78.7%). The results indicate that the segmentation-VL design can produce structured annotations while preserving asynchronous bimanual behaviour, providing a basis for richer semantic subtask identification and state-based verification.

関連論文

PR本紙発行元 EmplifAI