PhysEvo: 凍結モデルの物理的自己改善フレームワーク
PhysEvo: Astra Can Act, Let It
単一の凍結モデルを中心に、タスク実行とメタエージェントによる失敗診断・ツール改善を繰り返す物理的再帰的自己改善フレームワークを提案し、シミュレーションと実機で高い成功率を達成した。
著者: Wenqing Tian, Zeyu Zhang, Zhaocheng Liu, Fengwei Liu, Qiang Liu, Liang Wang
分類: cs.RO, cs.AI
原文アブストラクト
Astra can act, yet reliable manipulation depends on the system through which it observes and controls the world. We introduce PhysEvo, a framework for physical recursive self-improvement (RSI) around a single frozen model. A task agent executes robot tasks; a meta-agent uses the resulting trajectories to diagnose failures, revise tools and skills, and test corrections. The meta-agent can also improve its own diagnostic tools, so retained revisions support both later action and later self-improvement. This process develops joint-level control, evidence-seeking observation, and reusable manipulation skills without model-weight updates or a separately trained action policy. Across 42 RoboDojo tasks, held-out-layout evaluation of retained task-specific deployment versions yields a five-dimension average score of 68.14/100 and 62.00% success, compared with 47.17% for RoboDawn's one-shot Astra agent, the strongest published reference in our comparison. On eight manipulation tasks challenging direct Astra, PhysEvo achieves 55.00% success, compared with 1.25% for the direct-Astra reference. Deploying the simulation-evolved harness on AgileX PiPER and continuing skill revision yields 90.60/100 average score and 84.00% success across 25 trials on five real-world tasks. PhysEvo turns the consequences of action into persistent, testable changes to how a frozen model acts and improves.
関連論文
- 巧みな把持安定性のための時間的視触覚学習マニピュレーション
- FoldBack: 長期的な衣類折り畳みのための自己修正型マスク生成ポリシーマニピュレーション
- 関節特化型ハイブリッド遠隔駆動を用いた全駆動4自由度ロボット指の設計マニピュレーション
- UMI型ロボット教示のための高精度エンドエフェクタ位置推定マニピュレーション
- RoboPace: 接触を考慮した行動チャンク方策の時間最適リタイミングマニピュレーション
- 断続的な視覚喪失に頑健な実ロボットマニピュレーションのための標的モダリティドロップアウトマニピュレーション