日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.10962

iAm.md: 未知のオープンボキャブラリ領域におけるエージェント内省によるロボットスキル自己評価

iAm.md: Robot Skill Self-Assessment through Agentic Introspection for Unknown Open-Vocabulary Domains

シェア:XThreadsFacebookLINEはてブBluesky

LLMエージェントがロボットのスキルと環境を内省し、実行可能な計画を生成するためのMarkdown標準とフレームワークを提案。シミュレーションでナビゲーションとマニピュレーションのタスクにおける有効性を示した。

著者: Vincenzo Guarino, Emanuele Musumeci, Vincenzo Suriani, Daniele Nardi

分類: cs.RO, cs.AI

原文アブストラクト

Agentic AI based on Large Language Model generalization capabilities offers a wide range of potential applications, including planning for embodied tasks. For example, embodied agents based on Foundation models can generate plausible plans in autonomous robotics scenarios. Due to limited context windows or hallucinatory phenomena in the next-token prediction formulation, behaviors may be generated without establishing whether the deployed robot and the observed environment actually support the requested operation, in what we call a "grounding failure". Thanks to the recent improvements in reasoning capabilities of foundation models, autonomous robot behavior generation problem can be formulated as a code generation problem. We present iAm.md, a Markdown standard and generation framework, that allows anchoring this process in complementary forms of deployment evidence. Through open-vocabulary semantic mapping, we combine local vision-language detections and object segmentation and refer them to persistent object records in this intermediate standardized representation, allowing agentic introspection. We then study this new technique on a simulated TIAGo, on navigation-and-manipulation tasks, showing how this standardized representation jointly supports skill self-assessment and executable task generalization.

関連論文

PR本紙発行元 EmplifAI