日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
微細操作arXiv:2610.12241

言語から動作へ:顕微鏡ロボットのためのタスク条件付きフォーカルスタック軌道統合

From Language to Motion: Task-Conditioned Focal-Stack Trajectory Integration for Microscopic Robots

シェア:XThreadsFacebookLINEはてブBluesky

言語指示を幾何演算子に変換し、焦点面軌道を信頼度重み付けと動的計画法で統合することで、顕微鏡ロボットの高精度なタスク実行を実現した。

詳しい要約

1. どんなもの?

- 微細ロボットのタスク幾何を言語・部品・焦点変化に対応させるsemantic-to-physicalフレームワーク。 - 指示を制約付き幾何演算子に変換し、frozen open-vocabulary perceptionを再利用。 - 局所的に信頼できるfocal-plane trajectoriesをconfidence weightingとdynamic programmingで統合。 - 2-D経路を物理実行に接続するcalibrated multi-view geometry。

2. 先行研究と比べてどこがすごい?

- part-specific U-Net training(20-100 labels)と比較し、zero-new-label構成で15分(従来72-165分)。 - image-first multi-focus fusionと比較し、trajectory-space integrationでRMSE 14.41→6.28 pixels(56.4%減)、P95 error 20.07→8.13 pixels(59.5%減)。 - 9つのpart-illumination条件下で有効性を確認。

3. 技術・手法の肝は?

- 指示をconstrained geometric operatorsにマッピング。 - frozen open-vocabulary perceptionを再利用。 - focal-plane trajectoriesをconfidence weightingとdynamic programmingで統合。 - calibrated multi-view geometryで2-D pathsを物理実行に接続。

4. どうやって有効だと検証した?

- prompt、unseen-part、geometry reconfigurationテストで6.30-6.59-pixel RMSE。 - 9つのpart-illumination条件でtrajectory-space integrationのRMSEとP95 errorを評価。 - ablationでconfidenceとpath-wise selectionの役割を分離。 - 代表的なロボット実験でtarget-region coverageが83.5%→92.9%に改善。

5. 議論はある?

- dispensingはmeasurable physical traceを提供し、手法のtask-specific limitationではない。 - その他の議論や限界は要旨からは不明。

6. 次に読むべき論文は?

- part-specific U-Net training - image-first multi-focus fusion - open-vocabulary perception - dynamic programming - calibrated multi-view geometry

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Junjie Xie, Chuxuan He, Junkai Huang, Heng Zhang, Angen Ye, Yujia Song, Yuqing Li, Pengsong Zhang, Dapeng Zhang

分類: cs.RO

原文アブストラクト

Microscopic robots require accurate task geometry despite changes in language, parts, and focus. We present a semantic-to-physical framework that maps instructions to constrained geometric operators, reuses frozen open-vocabulary perception, and integrates locally reliable focal-plane trajectories by confidence weighting and dynamic programming. Calibrated multi-view geometry connects 2-D paths to physical execution. Prompt, unseen-part, and geometry reconfiguration tests yield 6.30-6.59-pixel RMSE. Relative to part-specific U-Net training with 20-100 labels, the proposed zero-new-label configuration takes 15 rather than 72-165 min. Across nine part-illumination conditions, trajectory-space integration reduces RMSE from 14.41 to 6.28 pixels (56.4%) and P95 error from 20.07 to 8.13 pixels (59.5%) compared with image-first multi-focus fusion. An ablation isolates the roles of confidence and path-wise selection. In representative robot experiments, target-region coverage improves from 83.5% to 92.9%. Dispensing provides a measurable physical trace, not a task-specific limitation of the method.

関連論文

PR本紙発行元 EmplifAI