日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.12089

ManiUnit: 長期的タスクのためのマニピュレーションスキルデータセットとベンチマーク

ManiUnit: A Manipulation Skill Dataset and Benchmark for Long-Horizon Tasks

シェア:XThreadsFacebookLINEはてブBluesky

50のBEHAVIOR-1K活動から構築した21スキル種別・137,899セグメントのデータセットと1,260テストインスタンスのベンチマークを提案し、VLAポリシーのスキル別評価と初期状態摂動への感度を測定した。

詳しい要約

1. どんなもの?

- 長期的な mobile manipulation を対象とした Manipulation Skill Dataset と Benchmark である ManiUnit を提案 - BEHAVIOR-1K の 50 活動から構築 - dataset は 21 skill types、417 subtasks、137,899 segments を含む - benchmark は 1,260 test instances を含む - 各 segment に subtask instruction を付与し、skill 単位の評価を可能にする

2. 先行研究と比べてどこがすごい?

- 従来の task-level 評価では skill 固有の診断が難しく、早期失敗で後続 skill が未検証になる問題を指摘 - ManiUnit は各 segment に明示的な subtask instruction を付与 - 開始 base position や joint configuration の摂動に対する感度を測定 - 中間 simulator states を復元し local success conditions を定義することで、先行段階なしに各 skill を評価可能 - 集約スコアが skill ごとの大きな差を隠しうることを示す

3. 技術・手法の肝は?

- BEHAVIOR-1K の 50 活動から 137,899 segments を抽出し、21 skill types、417 subtasks に整理 - 各 segment に subtask instruction を対応付ける - robot の開始 base position または joint configuration の摂動に対する感度を測定 - 中間 simulator states を復元し、local success conditions を定義 - 代表的な vision-language-action (VLA) policies を評価

4. どうやって有効だと検証した?

- 代表的な VLA policies を ManiUnit benchmark で評価 - 類似した aggregate scores が per-skill の大きな差を隠すことを確認 - 開始状態摂動により実行が劣化し、full benchmark で joint perturbations は demonstrated starting states 比で成功率を約 56% 低下 - 2 つの long-horizon activities で、ManiUnit segments で訓練した skill policy は local manipulation success 78.7%、complete demonstrations で訓練した task policy は 49.3% - task と skill policies を planner で調整すると full-task success が 4.0% から 18.0% に上昇

5. 議論はある?

- 類似の aggregate scores が skill ごとの性能差を隠す可能性を指摘 - 開始状態の摂動が実行性能を大きく低下させることを示す - task-level metrics では skill-specific diagnosis が困難で、早期失敗が後続 skill の評価を妨げる問題を提起 - その他の議論や限界は要旨からは不明

6. 次に読むべき論文は?

- BEHAVIOR-1K - vision-language-action (VLA) policies - long-horizon mobile manipulation - task policy と skill policy を調整する planner 関連研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Guoting Wei, Dawei Yan, Xia Yuan, Gengming Zhang, Yelin He, Guodong Du, Jiaquan Ye, Heng Zhang, Xinming Wei, Xianbiao Qi, Chunxia Zhao, Haokui Zhang, Rong Xiao

分類: cs.RO

原文アブストラクト

Long-horizon mobile manipulation requires a robot to navigate multi-room environments and execute a sequence of manipulation skills under a single natural language instruction. Learning and evaluating these skills present three challenges: similar observations under a fixed task instruction may make skill selection ambiguous; even when a preceding skill succeeds, the robot state inherited by the next skill may deviate from its demonstrated starting states and affect execution; and task-level metrics hinder skill-specific diagnosis, while early failures leave later skills untested. We therefore introduce ManiUnit, a manipulation skill dataset and benchmark built from 50 BEHAVIOR-1K activities. Its dataset contains 137,899 segments across 21 skill types and 417 subtasks, and its benchmark contains 1,260 test instances. Correspondingly, ManiUnit pairs each segment with an explicit subtask instruction; measures sensitivity to perturbations of the robot's starting base position or joint configuration; and restores intermediate simulator states and defines local success conditions so that each skill can be evaluated without executing preceding stages. Evaluations of representative vision-language-action (VLA) policies show that similar aggregate scores can hide substantial per-skill differences. The tested starting-state perturbations also degrade execution: on the full benchmark, joint perturbations reduce success rates by approximately 56% relative to those from demonstrated starting states. On two long-horizon activities, a skill policy trained on ManiUnit segments achieves 78.7% local manipulation success, compared with 49.3% for a task policy trained on complete demonstrations. The trained skills further support complete-task execution on these activities, as coordinating the task and skill policies through a planner raises full-task success from 4.0% to 18.0%.

関連論文

PR本紙発行元 EmplifAI