日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.03278

DexJoCo-X: 多様なロボットハンド操作のための行動表現ベンチマーク

DexJoCo-X: Benchmarking Action Representations for Multi-Hand Dexterous Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

7種類の器用なハンド、6つのタスク、2100のデモを用いて、異なるハンド形態に共通する行動表現を比較評価するベンチマークとツールキットを提案し、統一行動空間と身体性認識エキスパートの有効性を示した。

詳しい要約

1. どんなもの?

- 複数の dexterous hand を対象とした cross-embodiment learning のための benchmark と toolkit である DexJoCo-X を提案 - 7種類の代表的な dexterous hand、6つの single-arm および bimanual タスク、2,100の balanced demonstrations を提供 - 共通の scene、success criteria、execution interface を備えた multi-hand, multi-task protocol を整備 - glove-to-hand mapping を再設計し、reviewed demonstrations を randomized scene に拡張する自動 pipeline を構築 - π_{0.5}, Ego-Pi, Being-H0.5 を用いて、shared action interface が multi-hand learning に十分かを検証

2. 先行研究と比べてどこがすごい?

- 既存研究では hand、task、dataset、control interface の違いにより、representation、pretraining、architecture の効果を分離できなかった - DexJoCo-X は共通 scene、success criteria、execution interface を備えた matched multi-hand, multi-task protocol を提供し、制御された比較を可能にする - 7種類の hand と 6タスク、2,100 demonstrations を揃え、cross-embodiment 表現の効果を分離して評価できる点が新しい - 単一の policy で全7 hand を制御できるかを体系的に検証する benchmark はこれまでにない

3. 技術・手法の肝は?

- 共通の scene、success criteria、execution interface を備えた multi-hand, multi-task protocol を設計 - glove-to-hand mapping を再設計し、reviewed demonstrations を randomized scene に拡張する自動 pipeline を構築 - π_{0.5} を 80次元 bimanual output に拡張して評価 - Ego-Pi は interleaved prediction で pretrained action head を保持し、per-hand multi-task learning を実現 - Being-H0.5 は cross-embodiment pretraining、unified action space、embodiment-aware experts を組み合わせ、function-aligned action slots を導入

4. どうやって有効だと検証した?

- π_{0.5} を 80次元 bimanual output に拡張すると success がほぼゼロになることを確認 - Ego-Pi は per-hand multi-task learning を支えるが、7 hand の joint training には不十分であることを確認 - Being-H0.5 は cross-embodiment pretraining、unified action space、embodiment-aware experts により、1つの policy で全7 hand を制御できることを確認 - function-aligned action slots は mean success 47.7% を達成し、native coordinates の 47.0%、DexLatent の 33.1% と比較して有効性を検証

5. 議論はある?

- cross-embodiment representation は action coordinates、pretraining、architecture が共同で shared manipulation structure と embodiment-specific control を分離することに依存する - shared action interface だけでは multi-hand learning に不十分である可能性を示唆 - 7 hand の joint training には architecture や pretraining の工夫が不可欠であることを議論 - ただし、各手法の詳細な限界や failure case の分析は要旨からは不明

6. 次に読むべき論文は?

- π_{0.5} (pi-zero-point-five) - Ego-Pi - Being-H0.5 - DexLatent - 関連する cross-embodiment learning や dexterous manipulation の研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xiangwei Jiang, Yao Mu, Lixin Duan, Wen Li

分類: cs.RO

原文アブストラクト

As dexterous hands proliferate, collecting data and training policies separately for every morphology becomes increasingly impractical. Scalable cross-embodiment learning therefore requires a unified representation that captures shared manipulation structure while preserving morphology-specific control. Differences in hands, tasks, datasets, and control interfaces prevent existing studies from isolating the effects of representation, pretraining, and architecture. We introduce DexJoCo-X, a benchmark and toolkit for controlled comparison across seven representative dexterous hands, six single-arm and bimanual tasks, and 2,100 balanced demonstrations. DexJoCo-X provides a matched multi-hand, multi-task protocol with common scenes, success criteria, and execution interfaces, redesigned glove-to-hand mappings, and an automated pipeline that expands reviewed demonstrations across randomized scenes. Using $π_{0.5}$, Ego-Pi, and Being-H0.5, we examine whether a shared action interface is sufficient for multi-hand learning. Expanding $π_{0.5}$ to an 80-dimensional bimanual output yields near-zero success. Ego-Pi preserves the pretrained action head through interleaved prediction and supports per-hand multi-task learning, but remains ineffective for seven-hand joint training. By contrast, Being-H0.5 combines cross-embodiment pretraining, a unified action space, and embodiment-aware experts, enabling one policy to control all seven hands. Within this architecture, function-aligned action slots achieve 47.7% mean success, compared with 47.0% for native coordinates and 33.1% for DexLatent. These results show that cross-embodiment representation depends on the entire learning system: action coordinates, pretraining, and architecture must jointly separate shared manipulation structure from embodiment-specific control.

関連論文

PR本紙発行元 EmplifAI