PuzzleMate: 一人称視点のパズル支援に向けたMLLMベンチマーク
PuzzleMate: Benchmarking MLLMs for Egocentric Puzzle Assistance
ジグソーパズルを題材に、MLLMが一人称視点で現在のパズル状態を認識し次の手順を指示できるかを評価するベンチマークを提案し、既存モデルの限界を明らかにした。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Avijit Dasgupta, Shayon Dasgupta, Zakaria Laskar, C. V. Jawahar, Karteek Alahari
分類: cs.CV
原文アブストラクト
Personal AI assistants hold the potential to evolve from digital interfaces into embodied companions capable of guiding users through complex physical activities. For these assistants to become integral to daily life, they must do more than identify objects; they must provide precise, step-by-step instructions that align with a user's real-time progress. While Multimodal Large Language Models (MLLMs) show promise in general visual understanding, their ability to deliver grounded, sequential guidance for fine-grained manipulation tasks remains largely unverified. In this paper, we choose the jigsaw puzzle as a strategic testbed for this capability. Unlike general object recognition, puzzle solving demands high-precision spatial reasoning, the ability to distinguish between minute geometric variations, and a rigorous adherence to sequential logic. We investigate this capability through PuzzleMate, a novel framework focused on jigsaw puzzle solving captured through an egocentric viewpoint. We deploy PuzzleMate in a user-in-the-loop study to evaluate how well state-of-the-art MLLMs perceive the current puzzle state and generate actionable next-step instructions. Our analysis reveals seven key bottlenecks that limit their effectiveness. Building on these insights, we propose a benchmark that enables systematic evaluation of MLLMs' reasoning capabilities for puzzle solving. Our findings reveal a substantial performance gap in current models like GPT-5.2 and Gemini-2.5-Pro; while these MLLMs are highly capable, they struggle to navigate the intricate reasoning and sequential logic essential for jigsaw puzzle assistance.