日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
タスク計画arXiv:2609.21221

知覚タスク計画のための完全微分可能なニューロ・ソフト・シンボリックフレームワーク

A Fully Differentiable Neuro-Soft-Symbolic Framework for Perceptual Task Planning

シェア:XThreadsFacebookLINEはてブBluesky

視覚知覚とタスク計画を単一の計算グラフで結び、微分可能なソフトシンボリック状態と遷移演算子で計画を最適化する手法を提案。不確実な知覚下でも高い成功率を達成し、ロボット動作実行との整合性も検証した。

詳しい要約

1. どんなもの?

- 知覚的不確実性を考慮したタスクプランニングのためのフレームワーク。 - 視覚知覚とタスクプランニングを単一の計算グラフで接続。 - 連続的なsoft symbolic stateを維持し、微分可能なsoft-T_P遷移演算子でルールを表現。 - 短い計画ホライズンで行動logitsを最適化。 - 計画目的からの勾配で知覚パラメータも更新可能。

2. 先行研究と比べてどこがすごい?

- 従来手法は知覚を離散シンボルに変換し、不確実性を捨て、タスクレベルのフィードバックを切断。 - 提案手法は完全微分可能で、知覚と計画を統合し、タスク関連の知覚表現を洗練。 - BlocksworldでLatPlan-40を40/40解決(LatPlanは33/40)、PlanBench-600で596/600解決(推論モデルベースラインは587/600)。 - 計算量と時間を大幅に削減。 - 知覚不確実性アブレーションで成功率が59%から83%に向上。

3. 技術・手法の肝は?

- 完全微分可能なneuro-soft-symbolicフレームワーク。 - 連続的なsoft symbolic stateを維持。 - ドメインルールを微分可能なsoft-T_P遷移演算子に持ち上げ。 - 短い計画ホライズンで行動logitsを最適化。 - 計画目的からの勾配が知覚パラメータを更新し、タスク関連知覚表現を洗練。

4. どうやって有効だと検証した?

- BlocksworldでLatPlan-40の40/40タスクを解決。 - PlanBench-600で596/600タスクを解決。 - 知覚不確実性アブレーションで成功率を59%から83%に改善。 - Blocksworldシーンでタスク・アンド・モーションシミュレーションを実施し、実行レベルの検証を提供。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- LatPlan - PlanBench - 推論モデルベースライン(具体的名称は要旨からは不明)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hongyan Wei, Wael AbdAlmageed

分類: cs.AI, cs.RO

原文アブストラクト

Perceptual planning tasks require two key capabilities: accurately perceiving uncertain scenes and planning valid action sequences following logical rules. Conventional methods convert perception into discrete symbolic facts and then plan, discarding perceptual uncertainty and severing task-level feedback to perception. We introduce a generic, fully differentiable neuro-soft-symbolic framework that connects visual perception and task planning within a single computational graph. The framework maintains a continuous soft symbolic state, lifts domain rules into a differentiable soft-$T_P$ transition operator, and optimizes action logits over a short planning horizon. Gradients from the planning objective can also update the perception parameters, allowing task-relevant perceptual representations to be refined during planning. On Blocksworld, our method solves 40/40 LatPlan-40 tasks and 596/600 PlanBench-600 tasks, compared with 33/40 for LatPlan and 587/600 for the reasoning-model baseline, while requiring substantially less computation and time. In the perceptual-uncertainty ablation, our method improves the success rate from 59\% with frozen perception to 83\%. We further conduct task-and-motion simulations on Blocksworld scenes, providing an execution-level validation of the compatibility between decoded task plans and downstream robotic motion execution.

関連論文

PR本紙発行元 EmplifAI