日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
変形物操作arXiv:2608.10718v1

変形物操作のためのTCAM:WBCD 2026トラック4におけるRMC2チャンピオンシステム

TCAM for Autonomous Deformable Manipulation: The RMC2 Champion System for WBCD 2026 Track 4

シェア:XThreadsFacebookLINEはてブBluesky

この論文は、Tシャツのピッキングから整列、平滑化までの変形物操作タスクを完全自律で行うシステムを、TCAMフレームワークを用いて構築した技術報告です。

詳しい要約

1. どんなもの?

本報告は、WBCD 2026 Track 4: Deformable Manipulation Challenge における RMC2 チームの優勝ソリューションを記述する。タスクは、スタックから単一の T シャツをピックし、印刷用パレットに載せ、カラーをターゲット領域に合わせ、印刷領域を滑らかにする一連の操作であり、単層分離、変形物の搬送、精密配置、接触を伴う表面調整を含む。完全自律実行が強く動機付けられており、TCAM (TermiBrain Causal Action Model) フレームワークを中心に、ハードウェア、知覚、データ、学習がポリシーが扱う物理的相互作用の複雑さを低減するように設計された完全自律システムを構築した。

2. 先行研究と比べてどこがすごい?

先行研究と比べて、ハードウェア、知覚、データ、学習を統合的に設計し、物理的相互作用の複雑さを低減する点が優れている。特に、単層分離用のカスタム 3D プリントグリッパー、手首中心の 4 カメラ構成、UMI スタイルのデモと実ロボットデモの組み合わせ、TCAM による因果分析に基づくデータ再収集とポリシー微調整のクローズドループが特徴的である。

3. 技術・手法の肝は?

手法の肝は、TCAM フレームワークによる閉ループシステムである。各軌道を分析して結果に寄与する物理的要因を特定し、データ再収集とポリシー微調整を駆動する。ポリシーはマルチビュー VLA バックボーンから 30 ステップのエンドエフェクタのデルタポーズアクションチャンクを出力する。また、単層分離用のカスタムグリッパー、タスクコンテキスト用の上部魚眼カメラと近接接触観察用の下部 RGB カメラを組み合わせた手首中心の 4 カメラ構成、UMI スタイルのデモと実ロボットデモの併用が重要である。

4. どうやって有効だと検証した?

最終競技会で、システムは 25 枚の T シャツをロードし、1 回あたり平均約 23 秒で、22 枚が要求される表面平滑性を達成し、Track 4 で第 1 位を獲得した。

5. 議論はある?

要旨からは、システムの限界や競技外での一般化に関する議論は不明。競技の成功は示されたが、異なる布地や環境での性能、TCAM の因果分析の詳細、データ収集のコストなどについては言及がない。

6. 次に読むべき論文は?

要旨で参照されている関連研究として、UMI (Universal Manipulation Interface) スタイルのデモ、VLA (Vision-Language-Action) モデル、TCAM (TermiBrain Causal Action Model) が挙げられる。次に読むべき論文としては、これらの基盤となった研究や、変形物操作における同様の課題を扱った論文が考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Guangrui Shen, Zhili He, Shigang Wang, Yuanjun Sun, Qing Yu

分類: cs.RO

原文アブストラクト

This technical report describes the RMC2 Team's champion solution for the WBCD 2026 Track 4: Deformable Manipulation Challenge. The task requires a robot to pick a single T-shirt from a stack, load it onto a printing pallet, align the collar with a target area, and smooth the printing region, a sequence that involves single-layer separation, deformable transport, precise placement, and contact-rich surface adjustment. The competition strongly incentivizes fully autonomous execution, motivating the development of an autonomous solution. We built a fully autonomous system around the TCAM (TermiBrain Causal Action Model) framework, with the design principle that hardware, perception, data, and learning should jointly reduce the physical interaction complexity the policy must handle. A custom 3D-printed gripper designed for single-layer fabric separation improves picking reliability on a dual-arm ARX X5 platform. A wrist-centric four-camera setup pairs upper fisheye cameras for task-level context with lower RGB cameras for close-range gripper-cloth contact observation. We combine portable UMI-style demonstrations with real-robot demonstrations collected on the deployable platform to provide both broad manipulation priors and deployment-specific dynamics. TCAM ties these components into a closed loop: each trajectory is analyzed to identify the physical factors contributing to its outcome, driving targeted data recollection and policy fine-tuning. The policy outputs 30-step end-effector delta-pose action chunks from a multi-view VLA backbone. In the final competition, our system loaded 25 T-shirts at an average of approximately 23 seconds per attempt, with 22 achieving the required surface smoothness, securing first place in Track 4.