日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.30554

ロボット制御のためのプライバシー保護型プロンプト方策探索

Privacy-Preserving Prompted Policy Search for Robotic Control

シェア:XThreadsFacebookLINEはてブBluesky

クラウドLLMに方策パラメータや報酬履歴を秘匿したまま送信し、LLM誘導の方策最適化を可能にするフレームワークPP-ProPSを提案した。

詳しい要約

1. どんなもの?

- LLMをin-context policy optimizerとして用いる強化学習(RL)の枠組み - クラウドLLM APIへ生のpolicy parametersとreward historyを送る際の機密漏洩を防ぐ - Privacy-Preserving Prompted Policy Search (PP-ProPS)を提案 - client-sideの秘密変換でpolicy parametersとreward値を符号化し、LLMには符号化済み情報のみ観測させる - Vanilla ProPSと異なり真の最適episodic returnの開示を不要とする - MuJoCo locomotion、classic control、highway driving、robotic arm manipulationで評価

2. 先行研究と比べてどこがすごい?

- 従来のLLM-guided policy optimizationは生のpolicy parametersとreward historyをクラウドAPIへ送信し、制御戦略が第三者に露出する問題があった - PP-ProPSはclient-sideの秘密変換により、LLM提供者には符号化済みpolicy parametersとscaled reward情報のみが見える - Vanilla ProPSでは必要だった真の最適episodic returnの知識・開示が不要 - 単一のtotal returnではなく個別のreward componentsをLLMへ与え、各候補policyについてより情報量の多いfeedbackを提供 - bounded historyによりpromptの無限増大を防ぎ、高次元policyやopen-weight LLMの利用を可能にする - 10タスク中7タスクでVanilla ProPSを上回り、6タスク中5タスクでPPO、SAC、TRPOを上回る

3. 技術・手法の肝は?

- client-sideのsecret transformationsによりpolicy parametersとreward値を符号化 - 各API requestに符号化済みpolicy parametersとscaled reward情報のみを含める - LLM提供者は符号化済み情報しか観測できない - 真の最適episodic returnを既知とせず、開示も不要 - 個別のreward componentsをLLMへ提示し、候補policyごとのfeedbackを豊富化 - bounded historyを採用し、promptの無限成長を防止 - 高次元policyの探索を改善し、open-weight LLMの利用を支援

4. どうやって有効だと検証した?

- continuousおよびdiscrete control problemsで評価 - 対象はMuJoCo locomotion、classic control、highway driving、robotic arm manipulation - Vanilla ProPSと比較し、評価した10タスク中7タスクでPP-ProPSが優位 - PPO、SAC、TRPOなどの従来RL手法と比較し、6タスク中5タスクでPP-ProPSが優位 - 具体的な評価指標や統計的有意性は要旨からは不明

5. 議論はある?

- 生のpolicy parametersとreward historyをクラウドLLM APIへ送る際の機密漏洩リスクを指摘 - 符号化によりLLM提供者への露出を限定する設計を主張 - 真の最適episodic returnを不要とすることで実運用上の制約を緩和 - 個別reward componentsとbounded historyによる探索改善を主張 - 限界、失敗事例、計算コスト、符号化の安全性証明などは要旨からは不明

6. 次に読むべき論文は?

- Vanilla ProPS - PPO - SAC - TRPO - MuJoCo locomotion - classic control - highway driving - robotic arm manipulation - LLM as in-context policy optimizer for RL

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Ali Irshayyid, Feng Lin, Chong Li, Jun Chen

分類: cs.RO, eess.SY

原文アブストラクト

Large language models (LLMs) have recently demonstrated promising capabilities as in-context policy optimizers for Reinforcement Learning (RL), enabling policy search driven by both numerical reward signals and natural language reasoning. However, deploying such methods in practice requires transmitting raw policy parameters and rewards history to cloud-based LLM APIs, exposing proprietary control strategies to third-party service providers. To address this issue, this paper introduces Privacy-Preserving Prompted Policy Search (PP-ProPS), a framework that enables LLM-guided policy optimization while keeping policy and environmental parameters confidential. PP-ProPS encodes policy parameters and reward values using secret client-side transformations before they are included in each API request, ensuring that the LLM provider observes only encoded policy parameters and scaled reward information. Furthermore, unlike Vanilla ProPS, the proposed framework does not require the true optimal episodic return to be known or disclosed to the LLM. Beyond protecting the optimization data, PP-ProPS improves the search process in two ways. First, it provides the LLM with individual reward components instead of only a single total return, offering more informative feedback about each candidate policy. Second, it uses a bounded history that prevents the prompt from growing indefinitely, improving search with high-dimensional policies and supporting the use of open-weight LLMs. The proposed PP-ProPS is evaluated on both continuous and discrete control problems spanning Multi-Joint dynamics with Contact (MuJoCo) locomotion, classic control, highway driving, and robotic arm manipulation. Compared to Vanilla ProPS, the proposed PP-ProPS outperforms ProPS in seven of the ten evaluated tasks, and surpasses conventional RL methods including PPO, SAC, and TRPO, in five of the six tasks.

関連論文

PR本紙発行元 EmplifAI