日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.10812

Skill-SLM: エージェントスキル駆動型小型言語モデルによる信頼性の高いロボット操作

Skill-SLM: Agent Skill-driven Small Language Models for Reliable Robot Operation

シェア:XThreadsFacebookLINEはてブBluesky

自然言語指示をサブタスクに分解し、スキルライブラリから適切なスキルを選んで実行可能なロボット操作に組み立てる小型言語モデルフレームワークを提案。

詳しい要約

1. どんなもの?

- 小型言語モデル(SLM)を搭載したロボット操作のためのフレームワーク「Skill-SLM」を提案。 - 自然言語のタスク指示をサブタスクに分解し、スキルライブラリから適切なスキルを選択・編成して実行可能なロボット操作を生成。 - タスク分解とスキル合成の問題として再定式化し、SLMが信頼性高く動作することを目指す。

2. 先行研究と比べてどこがすごい?

- 既存手法は蒸留指向で代表的なタスク・解ペアを列挙するため、データセット構築が困難で多様なタスクへの汎化が限定的。 - Skill-SLMはスキル駆動型ワークフローを採用し、未見タスクへの汎化能力で蒸留指向ベースラインを大幅に上回る。 - 地上車両タスクでも適用可能であり、異なるロボットプラットフォームへの拡張性を示す。

3. 技術・手法の肝は?

- ロボット操作スキルを考慮した文脈自由文法(CFG)を新規提案し、タスク達成に必要なスキルを抽出してスキルライブラリを構築。 - LLM教師を構成し、SLM用の訓練データセットを誘導・合成。SLMがタスク分解とスキル編成を信頼性高く行えるようにする。 - 漸進的スキル編成戦略を採用し、スキル実装とロボット操作全体の信頼性を向上。

4. どうやって有効だと検証した?

- UAV操作タスクでの実験により、蒸留指向ベースラインと比較してSkill-SLMが大幅に優れることを示す。 - 特に汎化能力を要する未見タスクで有効性を確認。 - 地上車両タスクでの追加実験により、異なるロボットプラットフォームへの適用可能性を検証。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 蒸留指向ベースライン(具体的名称は要旨に記載なし) - 関連手法:LLM教師によるデータ合成、文脈自由文法(CFG)を用いたスキル抽出、漸進的スキル編成戦略

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Wenhao Wang, Yanyan Li, Jiawei Yuan

分類: cs.RO

原文アブストラクト

Small language models (SLMs) have been increasingly adopted for onboard robot operation because they enable intelligent decision-making. However, existing approaches are mainly distillation-oriented and rely on enumerating representative task-solution pairs. This makes dataset construction difficult and limits generalization to diverse robot tasks whose possible forms grow rapidly. This paper proposes Skill-SLM, a framework that reformulates SLM-driven robot operation as a task-decomposition and skill-composition problem. Given a natural language task instruction, Skill-SLM decomposes the task into subtasks, selects appropriate skills from the skill library, and orchestrates the selected skills into executable robot operations. First, to support the skill-driven workflow, we propose a novel robot operational skill aware context-free grammar (CFG) to extract the skills required to accomplish tasks and build the skill library accordingly. Then, we configure LLM teachers to induce and synthesize training datasets for the SLMs, enabling SLMs to decompose tasks and orchestrate skills reliably. Additionally, we employ a progressive skill orchestration strategy to improve the reliability of skill implementation and overall robot operation. Experiments on UAV operation tasks indicate that Skill-SLM substantially outperforms distillation-oriented baselines, especially on unseen tasks that require generalization of capabilities. Additional experiments on ground vehicle tasks further demonstrate that Skill-SLM can be applied to different robot platforms.

関連論文

PR本紙発行元 EmplifAI