日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2608.22035

Ludi 0.1: 社会的知能を持つロボットのためのエージェントシステム

Ludi${}_{\scriptscriptstyle 0.1}$: An Agentic System for Socially Intelligent Robots

シェア:XThreadsFacebookLINEはてブBluesky

ロボットが曖昧な指示を理解し、文脈を保持し、意図の変化に応じて行動を修正しながら人間と協調するための、対話・推論・記憶・ナビゲーション・操作を統合したエージェントシステムを提案した。

詳しい要約

1. どんなもの?

Ludi 0.1は、社会的知能を持つロボットのためのエージェントシステムである。音声対話、マルチモーダル推論、記憶、ナビゲーション、学習された操作を統合し、曖昧な要求、明確化、修正、割り込み、社会的・タスク対話の混合、マルチステップタスクなどに対応する。意思決定の中核は、マルチターン対話トレースで微調整されたvision-language modelであり、専用のハーネスがモデルとツールの相互作用ループを管理し、専門のナビゲーション・操作ポリシーが物理的スキルを実行する。

2. 先行研究と比べてどこがすごい?

従来のロボット基盤モデルは知覚と制御を大幅に進歩させたが、孤立したコマンドの実行に留まっていた。Ludi 0.1は、曖昧さの認識、文脈の維持、意図の伝達、ユーザーの意図変化に応じた行動の修正といった、自然な人間とロボットの協調に必要な能力を統合したエージェントシステムを提案する点で優れている。

3. 技術・手法の肝は?

手法の肝は、マルチターン対話トレースで微調整されたvision-language modelを意思決定コアとし、専用ハーネスでモデルとツールの相互作用ループを管理すること。さらに、専門のナビゲーション・操作ポリシーが物理的スキルを実行する。これにより、対話と物理的行動を統合する。

4. どうやって有効だと検証した?

要旨からは具体的な検証方法は不明。ただし、システムが実用的な経路を示し、将来の統合基盤モデル開発に必要なマルチモーダル対話トレースを生成することが述べられている。

5. 議論はある?

要旨からは議論は不明。ただし、現在のシステムは実用的な経路を示すが、より深く統合された基盤モデルへの発展が課題として示唆されている。

6. 次に読むべき論文は?

要旨で参照されている研究は明示されていないが、関連分野としてrobot foundation models、vision-language models、human-robot collaboration、interactive speech、multimodal reasoning、navigation、learned manipulationに関する論文が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Wooseong Chung, William Cong, Jakub Dworakowski, Ethan Ewer, Tri Wahyu Guntara, Yeonwoo Jeong, Tianchong Jiang, Chaewon Kim, Hyunseo Kim, Jinwoo Kim, Jinyeon Kim, Yea-Seul Kim, Jack Kunde, Kangwook Lee, Sangheon Lee, Robert Nowak, Junha Roh

分類: cs.RO

原文アブストラクト

Robot foundation models have substantially advanced perception and control, but natural human-robot collaboration requires more than executing isolated commands. A robot must recognize ambiguity, maintain context across turns, communicate its intentions, and revise ongoing behavior as the user's intent changes. We present $\scriptstyle\mathsf{Ludi}_{\scriptscriptstyle 0.1}$, an agentic system for socially intelligent robots that integrates interactive speech, multimodal reasoning, memory, navigation, and learned manipulation. Its decision-making core is a fine-tuned vision-language model trained on multi-turn interaction traces spanning ambiguous requests, clarifications, corrections, interruptions, mixed social and task dialogue, and multi-step tasks. A purpose-built harness manages the model-tool interaction loop, while specialized navigation and manipulation policies execute physical skills. Ludi${}_{\scriptscriptstyle 0.1}$ demonstrates a practical path toward fluid human-robot collaboration today while producing the multimodal interaction traces needed to develop a more deeply integrated foundation model for robots and people.

関連論文