日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.27247

行動を変える記憶は行動を導く記憶ではない:履歴条件付きロボット方策の反事実的監査

Memory That Changes Action Is Not Memory That Guides It: Counterfactual Auditing of History-Conditioned Robot Policies

シェア:XThreadsFacebookLINEはてブBluesky

ロボットの記憶が意思決定を導いているかを、同一の現在で二つの履歴を交差させる反事実的監査(CMA)で評価し、記憶への感度と信頼性を分離して検証した。

詳しい要約

1. どんなもの?

履歴条件付きロボットポリシー(history-conditioned robot policies)の記憶(memory)が意思決定を導いているかを評価する手法。Counterfactual Memory Audit (CMA) を提案。同一の現在入力に再収束する2つの履歴を交差させ、凍結ポリシーを共通乱数下でクエリし、各保存行動を両履歴で評価する。これにより memory sensitivity、warranted choice、matched-world physical value、per-pair reliability を分離する。

2. 先行研究と比べてどこがすごい?

従来の memory-policy 評価はタスク成功や記憶摂動下の行動変化に依存し、記憶が決定を導くことを確立できない。CMA は「記憶が行動を変えること」と「記憶が行動を導くこと」を区別する decision-level audit を提供する点が新しい。

3. 技術・手法の肝は?

2つの履歴を verified-identical present で交差させ、凍結ポリシーを common randomness 下でクエリし、各 saved action を両履歴で評価。これにより memory sensitivity、warranted choice、matched-world physical value、per-pair reliability を分離。native interventions で closed-loop influence も検証。

4. どうやって有効だと検証した?

Mem-0 で全 audited Put Back pair が行動を変えるが、完全に信頼できるのは 20/64 pair のみ。後の Swap 決定では全 paired action が変化するが両記憶が同じ分岐を選択。native interventions で history bank 置換が行動を置換内容へ誘導し、4096-byte protected anchor の復元が注入 bank fault で失われた Swap 成功の 38.9 ポイントを回復。dual-arm physical platform では記憶が saved action を変えるが、完了した Put Back 操作 9 件中 5 件が誤ったターゲットに到達。

5. 議論はある?

ロボットは記憶し反応できるが、その過去が正当化する行動を選ぶために記憶を確実に使っているわけではない。CMA はこれらのケースを区別する decision-level audit を提供する。

6. 次に読むべき論文は?

要旨からは不明(参照/比較されている研究や関連手法が明示されていない)。同分野の定番として memory-augmented robot policies、history-conditioned policies、counterfactual evaluation などが考えられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Jiajie Zhang, Yankai Xiang, Changhao Chen

分類: cs.RO

原文アブストラクト

A robot returning a block to its origin tray may encounter two task-consistent pasts that reconverge to the same current input but warrant different actions. Yet memory-policy evaluations often rely on task success or action change under memory perturbation, neither of which establishes that memory guides the decision. We propose the \textbf{Counterfactual Memory Audit (CMA)}, an evaluation protocol that crosses two histories at a verified-identical present, queries a frozen policy under common randomness, and evaluates each saved action under both pasts. This separates memory sensitivity, warranted choice, matched-world physical value, and per-pair reliability. On Mem-0, every audited Put Back pair changes action, but only $20/64$ pairs are fully reliable; at a later Swap decision, all paired actions change while both memories select the same branch. Native interventions further show closed-loop influence: replacing the history bank redirects behavior toward the replaced content, while restoring a 4096-byte protected anchor recovers $38.9$ points of Swap success lost to injected bank faults. On a dual-arm physical platform, memory changes saved actions, yet five of nine completed Put Back manipulations reach the wrong target. These results show that a robot can remember and react without reliably using memory to choose the behavior its past warrants. CMA provides a decision-level audit for distinguishing these cases.

関連論文

PR本紙発行元 EmplifAI