日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
世界モデルarXiv:2609.30214

水中C3-JEPA:ROVサルベージのための物体中心クロスビュー世界モデル

Underwater C3-JEPA: An Object-Centric Cross-View World Model for ROV Salvage

シェア:XThreadsFacebookLINEはてブBluesky

複数視点のRGB映像と制御信号から、接触や水流の遅れを考慮して対象物の状態変化を潜在空間で予測する物体中心の世界モデルを提案し、実水中映像で有効性を示した。

詳しい要約

1. どんなもの?

- 近距離の重量物を扱う水中 ROV salvage 向けの、object-centric な multi-view 予測 world model。 - 名称は Underwater C3-JEPA(cross-view, control-conditioned, context-extended)。 - 接触センサ無しで、同期した multi-view RGB 観測と vehicle 制御信号から、contact interaction と hydrodynamic lag 下での task-object 状態の潜在空間での進展を予測する。 - 予測インタフェースは MPC 候補評価や imagined-rollout による behavior-agent 訓練を支える。

2. 先行研究と比べてどこがすごい?

- reconstruction-free な latent baseline と比較し、学習表現が下流 probe へより多くの task-relevant 情報を転移する。 - predictor は軽量に保たれる。 - 実水中映像で、同一アーキテクチャが withheld camera の object state を復元し persistence を上回る。 - この recipe は simulation を超えて転移する。

3. 技術・手法の肝は?

- 複数カメラ観測を task-object token と context token に encode する。 - held-out-view attention で cross-camera の evidence を fuse する。 - control を条件として未来状態を直接予測する。 - weak binding により低 annotation コストで target と gripper を anchor する。 - SIGReg で幾何表現を sharp にする。

4. どうやって有効だと検証した?

- 学習表現を下流 probe に転移させ、reconstruction-free latent baseline と比較した。 - 実水中ビデオで、withheld camera の object state 復元と persistence に対する優位を検証した。 - 予測インタフェースが MPC 候補評価と imagined-rollout behavior-agent 訓練を支えることを示した。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- reconstruction-free latent baseline(要旨で比較対象として言及)。 - JEPA 系の world model(一般名)。 - model-predictive-control(MPC)を用いるロボティクス研究(一般名)。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yuncong Yang, Jinlong Li, Yulong Xue, Feng Wu, Chunwen Zhang, Lei Qiao, Xuyang Wang

分類: cs.RO, cs.AI

原文アブストラクト

We present Underwater C$^{3}$-JEPA (cross-view, control-conditioned, context-extended), an object-centric multi-view predictive world model for near-field heavy-load underwater ROV salvage. Without contact sensors, it predicts in latent space how the task-object state evolves through contact interaction and under the hydrodynamic lag of the vehicle, from synchronized multi-view RGB observations and vehicle control signals. C$^{3}$-JEPA encodes multi-camera observations into task-object and context tokens, fuses cross-camera evidence through held-out-view attention, and directly predicts future states conditioned on control. Weak binding anchors the target and gripper at low annotation cost, while SIGReg sharpens the geometric representation. Experiments show that the learned representation transfers substantially more task-relevant information to downstream probes than a reconstruction-free latent baseline, while keeping the predictor lightweight. The resulting predictive interface supports model-predictive-control (MPC) candidate evaluation and imagined-rollout behavior-agent training. Validation on real underwater video shows the same architecture recovering a withheld camera's object state and staying ahead of persistence, so the recipe transfers beyond simulation.

関連論文

PR本紙発行元 EmplifAI