日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ベンチマーク/身体性エージェントarXiv:2609.30971

SciHorizon-eLab: 科学実験用身体性エージェントのスケーラブルなベンチマークのためのエージェント型プロトコル・タスクコンパイラ

SciHorizon-eLab: An Agentic Protocol-to-Task Compiler for Scalable Benchmarking of Scientific Embodied Agents

シェア:XThreadsFacebookLINEはてブBluesky

自然言語の実験プロトコルを意味接地・実行可能タスク合成・シミュレーション検証を通じて身体性タスクへ自動コンパイルし、300タスクのベンチマークを構築した研究。

詳しい要約

1. どんなもの?

- 科学実験の自動化を目指すembodied agentsの評価環境を自動構築する、agentic protocol-to-task compiler「SciHorizon-eLab」を提案。 - 自然言語の実験protocolを、意味を保ったembodied taskへと段階的にコンパイルする。 - semantic grounding、executable task synthesis、multi-stage simulation-based certificationの3段階で構成。 - 300のcertified tasksからなるbenchmark「BenchName」も構築し、HIL実行やexpert demonstration生成、step-level評価を支援。

2. 先行研究と比べてどこがすごい?

- 既存のsimulation-based laboratory benchmarksは手動のtask engineeringに大きく依存していた。 - 本手法はprotocol-to-task compilation問題として定式化し、多様なscientific protocolsを実行可能かつ検証可能なembodied tasksへ大規模に自動コンパイル可能にした。 - 再現可能なexpert demonstrationsとexecution tracesの生成も可能にした点が先行研究と異なる。

3. 技術・手法の肝は?

- 自然言語protocolを入力とし、semantic groundingで意味的に接地した環境を生成。 - executable task synthesisにより実行可能なmanipulation programsを合成。 - multi-stage simulation-based certificationでstep-level success specificationsを付与し、検証済みタスクを出力。 - これにより再現可能なexpert demonstrationsとexecution tracesの生成を実現。

4. どうやって有効だと検証した?

- 構築したBenchNameは多様なlaboratory operationsにわたる300のcertified tasksを含む。 - HIL task execution、再現可能なexpert-demonstration生成、ordered step-level evaluationをサポート。 - 代表的なタスクで最強のpolicyでも平均成功率は49.7%にとどまり、humanとembodied agentのcoordinationに顕著な弱点があることを示した。

5. 議論はある?

- 最強policyの平均成功率が49.7%と低く、humanとembodied agentのcoordinationに顕著な弱点があると報告。 - その他の議論や限界については要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照・比較されている個別の先行研究は明示されていない。 - 同分野の関連手法として、simulation-based laboratory benchmarksやembodied agent benchmarks(例: RLBench, ALFRED, BEHAVIOR)が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Maokai Qin, Chuan Qin, Qi Zhang, Dianyu Liu, Zirui Liu, Hongting Niu, Yuanchun Zhou, Hengshu Zhu

分類: cs.AI

原文アブストラクト

Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments. Existing simulation-based laboratory benchmarks rely heavily on manual task engineering, making it challenging to systematically compile diverse scientific protocols into executable and verifiable embodied tasks at scale. To address this challenge, we introduce SciHorizon-eLab, an agentic protocol-to-task compiler that formulates scientific embodied task construction as a compilation problem. Given a natural-language protocol of scientific experiments, SciHorizon-eLab progressively compiles laboratory protocols into semantic-preserving embodied tasks through semantic grounding, executable task synthesis, and multi-stage simulation-based certification. The system generates semantically grounded environments, executable manipulation programs, and step-level success specifications, while enabling reproducible generation of expert demonstrations and execution traces. Using this pipeline, we further construct \BenchName, a ready-to-use benchmark comprising 300 certified tasks across diverse laboratory operations. It supports HIL task execution, reproducible expert-demonstration generation, and ordered step-level evaluation. Across representative tasks, the strongest policy attains an average success rate of only 49.7%, with further evaluations revealing pronounced weaknesses in human and embodied agent coordination. We publicly release the code, benchmark data, and evaluation toolkit at https://github.com/SciHorizon-elab/SciHorizon-elab.

PR本紙発行元 EmplifAI