日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
物理シミュレーション/LLMエージェントarXiv:2610.08720

WorldSolver: LLMエージェントはソルバ生成を通じて物理ダイナミクスをシミュレートできるか

WorldSolver: Can LLM Agents Simulate the Physical Dynamics via Solver Generation?

シェア:XThreadsFacebookLINEはてブBluesky

古典的CG論文由来の168の物理シミュレーションタスクを集めたベンチマークWorldSolverを提案し、LLMエージェントがソルバをコード生成して物理現象を再現できるかを実行・視覚・物理的妥当性の3軸で評価した。

詳しい要約

1. どんなもの?

- LLM-based agents が物理シミュレーションの solver を生成できるかを評価する benchmark『WorldSolver』を提案。 - 61 本の classic computer graphics 論文の物理現象から 168 の simulation tasks を構築し、7 つの physical domains を網羅。 - 各 task は固定の simulation environment を提供する code scaffold を含み、agent が solver 実装を補完する形式。 - 評価軸は Execution Checks、Visual Fidelity、Physical Plausibility の 3 次元。

2. 先行研究と比べてどこがすごい?

- 従来の LLM agent 研究は科学的・工学的問題解決に進展があるが、物理 simulation の solver 生成能力は未探索だった。 - 既存 benchmark と異なり、物理理解・数学的定式化・software engineering を統合的に要求する点が新しい。 - 実行可能性だけでなく、視覚的忠実性と物理的妥当性まで評価する枠組みを提示。

3. 技術・手法の肝は?

- 各 task に code scaffold を用意し、simulation environment を固定して agent は solver のみを実装。 - solver は dynamic system の状態が時間発展する計算を担う。 - 評価は Execution Checks(実行成功)、Visual Fidelity(意図した動的挙動の再現)、Physical Plausibility(物理に基づく検証)の 3 次元。

4. どうやって有効だと検証した?

- frontier agents を対象に実験を実施。 - 実行可能な solver の生成自体が難しく、visual/physical correctness の達成はさらに困難と判明。 - GPT-5.6-Sol と Claude-Opus-5 が相対的に良好だが、overall scores はそれぞれ 48.7% と 46.7% に留まる。

5. 議論はある?

- WorldSolver は agentic solver generation への初期ステップと位置づけられる。 - 現状の frontier agents でも物理 simulation の忠実な再現は未達成であり、改善の余地が大きい。 - 動的物理世界を忠実に simulate できる agent への進展を促すことを期待。

6. 次に読むべき論文は?

- 要旨で参照・比較されている個別研究は明示されていない。 - 関連手法として、LLM-based agents による scientific/engineering problem solving、physics simulation、embodied AI、computer graphics の classic papers が挙げられる。 - 同分野の定番として、LLM agent の code generation benchmark や physics simulation benchmark を参照するのが有用。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Siru Jiang, Yongzhe Lyu, Shuo Lu, Yubin Wang, Yuxiang Zhang, Yue Liao, Bin Wang, Jian Liang, Tieniu Tan

分類: cs.AI

原文アブストラクト

LLM-based agents are increasingly advancing scientific and engineering problem solving, with physics simulation emerging as a challenging yet practical testbed for reproducing complex physical phenomena with application in embodied AI, games and films. As the workhorse of such simulation, a solver computes how the state of a dynamic system evolves over time. Building such solvers requires physical understanding to identify appropriate models, mathematical reasoning to formulate the underlying dynamics, and software engineering to implement them as executable code, yet this capability of LLM agents remains underexplored. To this end, we introduce WorldSolver, a benchmark of 168 simulation tasks derived from physical phenomena in 61 classic computer graphics papers, spanning 7 physical domains. Each task contains a code scaffold that provides a fixed simulation environment for the scene, with the solver implementation left for the agent to complete. Specifically, we evaluate them along three dimensions: Execution Checks for successful execution, Visual Fidelity for reproducing the intended dynamic behavior in the rendered simulation, and Physical Plausibility for physics-grounded verification of the generated dynamics. Experiments on frontier agents reveal that producing executable solvers is difficult itself, and satisfying visual and physical correctness is even harder. GPT-5.6-Sol and Claude-Opus-5 perform comparatively better than the other evaluated agents, yet achieve overall scores of only 48.7% and 46.7%, respectively. WorldSolver is an early step toward agentic solver generation, and we hope it helps drive progress toward agents that can faithfully simulate the dynamic physical world. Code is available at https://github.com/sirujiang/WorldSolver.

PR本紙発行元 EmplifAI