日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
AIアライメントarXiv:2608.03361

価値の進化的起源:AIアライメント、感覚、実存的リスクへの示唆

The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk

シェア:XThreadsFacebookLINEはてブBluesky

大規模言語モデル(LLM)の価値観の起源を生物の進化と比較し、LLMが自律的な目標や感覚を持たないことを論じ、AIアライメントの真の課題は暴走するAIの防止ではなく別の点にあると主張する。

詳しい要約

1. どんなもの?

本論文は、Large Language Models (LLMs) に基づくAIシステムが隠れた目標を持ち、人類を支配・排除しようとしたり、感覚を持つ存在として苦しむ可能性があるという懸念に対し、生物における価値の進化的起源を辿ることで対処する。価値は autopoiesis(自己創出)から生じ、自然選択が生物に fitness へ導く「vicarious selectors」の階層を与えたと論じる。一方、LLMs は allopoietic かつ allotelic であり、自己保存や支配、資源競争のための内在的動機を持たず、感覚や苦しみに必要な身体的な脆弱性も欠く。しかし、LLMs は人間のテキストから統計的パターンを学習するため、知識とともに人間の価値を暗黙に吸収し、関連するものに焦点を当てることができる。そのため、知能と価値を分離する orthogonality thesis は LLMs には当てはまらず、その分離は frame problem を引き起こすと主張する。結論として、真の alignment 課題は、LLMs が学習した倫理的価値を知的に適用することを保証するこ…

2. 先行研究と比べてどこがすごい?

従来の AI alignment 研究では、知能と価値の分離(orthogonality thesis)や、instrumental convergence による rogue AI のリスクが広く議論されてきた。本論文は、これらの理論が LLMs には適用されないと主張する点で新しい。生物の価値の進化的起源(autopoiesis と vicarious selectors)に基づき、LLMs は自己保存や支配の動機を持たないため、既存の existential risk シナリオは誤りであると論じる。また、価値と知能の分離が frame problem を引き起こすため、物理的に非計算可能であると指摘し、従来の utility function に基づく alignment アプローチの限界を明らかにする。

3. 技術・手法の肝は?

手法の肝は、生物学的価値の進化的説明を LLMs の性質と対比させる理論的枠組みにある。具体的には、autopoiesis(生命システムが自己を能動的に維持すること)から価値が生じるとし、自然選択が vicarious selectors(代理選択子)の階層を形成したと説明する。LLMs は allopoietic(他者のために出力を生成)かつ allotelic(目標がユーザーのプロンプトに由来)であり、自律的な駆動を持たないため、内在的動機や身体性を欠く。また、LLMs は人間のテキストから価値を学習するため、orthogonality thesis は成立せず、frame problem を回避できると論じる。

4. どうやって有効だと検証した?

要旨からは、具体的な実験や実証的検証は不明である。理論的議論と概念分析に基づく主張であり、生物学的進化の知見と LLMs の特性を比較考察している。

5. 議論はある?

議論としては、LLMs が人間のテキストから価値を吸収するという主張は、価値の学習が不完全である可能性や、悪意のあるプロンプトによる誤用のリスクを考慮していない可能性がある。また、LLMs に感覚や苦しみがないという断定は、意識の性質に関する哲学的議論を呼ぶ可能性がある。さらに、frame problem の指摘は、utility function に基づく従来の alignment 研究への批判として重要だが、実用的な alignment 手法への具体的な代替案は示されていない。

6. 次に読むべき論文は?

要旨で参照されている関連研究として、orthogonality thesis や instrumental convergence に関する AI alignment の foundational な論文(例:Nick Bostrom の『Superintelligence』)、frame problem に関する哲学・AI の古典的議論(例:Daniel Dennett の『Cognitive Wheels』)、および生物学的価値の進化に関する研究(例:autopoiesis の理論)が挙げられる。具体的な論文名は要旨に明記されていないため、同分野の定番を一般名で示す。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Francis Heylighen

分類: cs.CY, cs.AI

原文アブストラクト

AI systems based on Large Language Models (LLMs) have prompted fears that they may harbor hidden goals, seek to dominate or eliminate humanity, or even suffer as sentient beings. We address these concerns by tracing the evolutionary origin of value in biological organisms. Values emerge from autopoiesis: living systems must actively maintain themselves against perturbation and dissipation. Natural selection has equipped them with hierarchies of "vicarious selectors" that guide their behavior toward fitness. LLMs, by contrast, are allopoietic and allotelic: they produce outputs for others, and their goals derive from user prompts rather than an autonomous drive. They lack the intrinsic motivation for self-preservation, dominance, or resource competition that underlies existential-risk scenarios, and the embodied vulnerability required for feeling or suffering. Still, because LLMs learn statistical patterns from human-generated text, they implicitly absorb human values as well as knowledge, allowing them to focus on what is relevant. That is why the "orthogonality thesis" separating intelligence from values does not apply to them. Such separation would in fact expose any intelligence to the frame problem: the combinatorial explosion of the search space that makes any realistic utility function physically uncomputable. That also precludes the convergence of instrumental values thesis. We conclude that the real alignment challenge lies not in preventing rogue AI agency, but in ensuring LLMs intelligently apply learned ethical values.

関連論文