日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
触覚/視覚統合arXiv:2607.16146v1

VTLoc: 視覚点群における学習ベースの触覚接触位置推定

VTLoc: Learning-based Tactile Contact Localization in Visual Point Clouds

シェア:XThreadsFacebookLINEはてブBluesky

触覚データと視覚点群を統合し、物体表面での接触位置を高精度に推定する新しいフレームワークVTLocを提案した。幾何学的マルチモーダル整列と反復更新機構により、実世界100物体のベンチマークで精度を向上させた。

著者: Zhiyuan Wu, Zhuo Chen, Shan Luo

分類: cs.RO, cs.CV

原文アブストラクト

Vision and touch are complementary modalities essential for robotic perception and manipulation. While vision provides global object context, touch offers precise local information at contact points. Integrating these modalities for contact localization, i.e., predicting the location of touch on an object's surface, poses significant challenges due to the need for accurate spatial alignment between tactile data and visual geometry. To address this challenge, we propose VTLoc, a novel visual-tactile framework that localizes contact points from tactile readings using a 3D point cloud as visual input. VTLoc introduces two key components: a geometric multi-modal alignment module, which reconstructs a pseudo-point cloud from fused visual-tactile features and aligns it with the visual point cloud to enforce spatial consistencies across modalities; and an iterative localizing updater, which iteratively refines the predicted contact location using fused visual-tactile features. Evaluated on a new benchmark of 100 real-world objects, VTLoc improves single-touch contact localization by reducing local-to-global correspondence ambiguity.