日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.11986

FearCaut-Qwen:視覚言語モデルにおける感情ステアリングが危険評価の判断基準を変化させる

FearCaut-Qwen: Affective Steering in a Vision-Language Model Shifts the Decision Criterion for Hazard Assessment

シェア:XThreadsFacebookLINEはてブBluesky

災害後の被害評価で危険判定を避けがちな視覚言語モデルの判断基準を、恐怖に関連する感情表現を活性化ステアリングで操作することで改善し、赤判定の再現率を27%から75.7%へ引き上げた。

詳しい要約

1. どんなもの?

- Vision-language model (VLM) の災害後被害評価における過度に保守的な判断傾向を、signal detection theory (SDT) で分析し、affective steering で修正する手法を提案。 - 具体的には、Qwen2.5-VL-7B-Instruct をベースにした二段階の SeisMLLM パイプラインで、SeisMLLM-1K テスト分割において、真に危険な建物の27.0%しか Red と判定せず、false Red は皆無 (c = +1.354, d' = 1.521) という問題を扱う。 - 恐怖に関連する affective representation を操作することで、判断基準をシフトさせ、Red recall を75.7%に向上させる。

2. 先行研究と比べてどこがすごい?

- 従来の VLM は被害評価で recall が低いことが指摘されているが、本研究は SDT を用いて知覚能力 (d') と判断基準 (c) を分離し、問題が基準の置き方にあることを明確化。 - 人間の恐怖がリスク回避を生むという知見に着想を得て、VLM の affective representation を因果的に操作する新手法を提案。 - 再学習なしで推論時に判断を調整できる点が、従来のファインチューニングやプロンプトエンジニアリングとは異なる。

3. 技術・手法の肝は?

- mechanistic interpretability を用いて、感情豊かな自然場景で affective circuit を局在化。 - sparse-neuron knockout と distributed steering で因果的検証を行い、恐怖方向を同定。 - その方向を建物タスクに注入 (activation steering) し、下流予測への影響を観察。 - 恐怖方向の注入で Red recall が上昇し、同じ方向を減算すると Red 予測が完全に抑制される。ノルムマッチしたランダム方向や幸福方向では有意な効果なし。

4. どうやって有効だと検証した?

- SeisMLLM-1K テスト分割で評価。 - 恐怖方向注入により Red recall が75.7%に上昇 (p<0.001)。 - 同じ方向の減算で Red 予測が完全に抑制。 - ノルムマッチしたランダム方向と幸福方向では有意な効果がなかった。 - 判断基準のシフト (c = -1.515) が確認され、識別力は向上しなかった (d' = -0.493)。

5. 議論はある?

- 判断基準のシフトが主なメカニズムであり、知覚能力の向上ではない。 - 再学習なしで推論時に VLM の判断を調整できることを示す。 - mechanistic interpretability が工学的応用における VLM の判断行動の診断と制御に有用であることを実証。 - 倫理的含意や他のタスクへの一般化可能性については要旨からは不明。

6. 次に読むべき論文は?

- SeisMLLM パイプライン (Qwen2.5-VL-7B-Instruct ベース) の詳細。 - signal detection theory を VLM に適用した研究。 - mechanistic interpretability を用いた activation steering の関連研究 (例: representation engineering, causal tracing)。 - 感情表現の操作に関する研究 (affective computing)。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Xiaoshan Zhou

分類: cs.CV

原文アブストラクト

Vision-language models (VLMs) show great potential for damage assessment after a disaster, but a recurring deficiency is that they are reluctant to declare a hazard; that is, recall is low even when overall accuracy appears adequate. This study examines that deficiency by using signal detection theory to decompose the decision behavior into perceptual capability and decision-criterion placement. We then propose a novel method for correcting the over-conservative decision policy, inspired by the finding that fear makes humans risk-averse, and ask whether an affective representation associated with fear can be causally manipulated to similarly alter a VLM's decision tendency. Using mechanistic interpretability, we localize a causally implicated affective circuit in the model and use activation steering to manipulate it while observing the effect on downstream prediction. The method is tested on a two-stage SeisMLLM pipeline built on Qwen2.5-VL-7B-Instruct, which flags only 27.0% of genuinely unsafe buildings on the SeisMLLM-1K test split and never issues a false Red, an SDT criterion of c = +1.354, despite adequate evidence quality (d' = 1.521). An affective direction is localized on emotion-rich natural scenes, causally validated by sparse-neuron knockout and distributed steering on held-out emotion data, and then injected into the building task. Fear-direction injection raises Red recall to 75.7% (p<0.001), and subtracting the same direction suppresses Red predictions entirely, whereas norm-matched random and matched happiness directions show no significant effect. The mechanism is a shift in criterion (c=-1.515) while discrimination is not improved (d'=-0.493). These results show that VLM decisions can be adjusted at inference time without retraining and demonstrate how mechanistic interpretability can be used to diagnose and control VLM decision behaviors in engineering applications.

関連論文

PR本紙発行元 EmplifAI