日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
制御・身体性arXiv:2609.34840

侵害受容を制御プリミティブとして:単一身体に展開するエージェントのための求心性チャネルと侵害受容記憶

Nociception as a Control Primitive: Afferent Channels and Nociceptive Memory for Agents Deployed in One Body

シェア:XThreadsFacebookLINEはてブBluesky

単一の身体で生涯学習できないエージェントに、負荷依存の侵害受容チャネルと記憶を持たせ、摩耗を考慮した作業配分を改善する手法を提案。シミュレーションで寿命と生産量の向上を実証。

詳しい要約

1. どんなもの?

- 単一の身体に配備された agent が、その身体の摩耗耐性を学習できない epoch-one 設定を扱う。 - 方針のパラメータは身体が与えられる前に固定され、生涯更新されない。 - agent は load-gated な nociceptive channel と、感じたことを保持する memory を持つ。 - 感じた cost が、最も優しい作業ではなく、まだ感じていない best-paid な作業へ配分を移すことを証明する。 - 保持のない agent は felt-cost 制約が束縛するのを見ないこと、channel は脅威が個別に予測不能・回避が安価・無視が高価な場合にのみ報いることを示す。

2. 先行研究と比べてどこがすごい?

- 生涯を通じて学習する従来設定ではなく、身体が引かれる前に固定される epoch-one 設定を対象とする点が異なる。 - 同じ個体で channel と memory の有無を比較し、生涯をまたいで学習した schedule を持たない点で先行研究と異なる。 - population-trained agent は同じ channel から +0.65 年を得るが -0.54 の output を伴い、単一身体には種事前分布がないため差が出る。 - 両者は identity で関連し、ablation mean は per-body value の (1-χ) を報告する(χ は blind schedule が既に捉える割合)。

3. 技術・手法の肝は?

- load-gated な nociceptive channel と、感じた cost を保持する memory を agent に持たせる。 - 固定重み方針のパラメータを身体が引かれる前に設定し、生涯更新しない。 - 感じた cost に基づき、まだ感じていない best-paid な作業へ配分を移す制御則。 - 保持・置換・感じることを組み合わせ、身体ごとに channel と memory の有無を比較する。 - 種事前分布が供給するものを単一身体は持たないため、population-trained agent との差を identity で関連付ける。

4. どうやって有効だと検証した?

- 2,000 の simulated floor-layer knees で検証。摩耗は published loss rates に基づく。 - 感じ・保持・置換により、working life が age 55.2 から 59.6 へ、career output が 33.7 から 36.1 へ増加。 - 69.3% の身体が gain し、none lose。 - 感じるが保持しない身体は +4.4 年のうち 1 年を得て、残りは retention が担う。 - population-trained agent は同じ channel から +0.65 年を得るが -0.54 output。 - regime map が価値を予測する領域では、care robot が certified service life を 6 倍にし、field-anchored fleet が 0.55 ではなく 0.15 の machine を write off。予測しない領域では rover は blind caution に比べほとんど得をしない。

5. 議論はある?

- 感じた cost が best-paid な未感覚作業へ配分を移すこと、保持のない agent は felt-cost 制約が束縛するのを見ないことを証明。 - channel は脅威が個別に予測不能・回避が安価・無視が高価な場合にのみ報いる。 - population-trained agent との差は種事前分布が既に供給するもので、単一身体にはない。 - 両者は identity で関連し、ablation mean は per-body value の (1-χ) を報告する。 - regime map は両方向で成立する。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究や関連手法は明示されていない。 - 同分野の定番として、population-based training、domain randomization、sim-to-real transfer、wear-aware control、proprioceptive sensing に関する研究が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Wolfgang Maass

分類: cs.AI

原文アブストラクト

An agent deployed in a single body cannot learn how fast that body wears, because every trial that would reveal its wear resistance wears the body it would protect. We study this \emph{epoch-one} setting, in which the parameters of a fixed-weight policy are set before the body is drawn and never updated in life. The agent carries a load-gated nociceptive channel and a memory that retains what was felt. We prove that felt cost moves the allocation to the best-\emph{paid} work not yet felt rather than the gentlest, that an agent without retention never sees the felt-cost constraint bind, and that the channel pays only where the threat is individually unpredictable, cheap to avoid and expensive to ignore. We measure per body, setting the agent with channel and memory against the same individual without them, where neither carries a schedule learned across lives. On $2{,}000$ simulated floor-layer knees, with wear anchored to published loss rates, feeling, retaining and substituting extends the working life from age $55.2$ to $59.6$ and raises career output from $33.7$ to $36.1$. $69.3\%$ of bodies gain and \textbf{none lose}. A body that feels but retains nothing past the day gains one of the $+4.4$ years, and retention carries the rest. A population-trained agent gains $+0.65$ years from the same channel at $-0.54$ output. The difference is what a species prior already supplies, and a single body has none. The two are related by an identity, the ablation mean reporting $(1-χ)$ of the per-body value with $χ$ the share a blind schedule already captures, so we report both. Where the regime map predicts value, a care robot sextuples its certified service life and a field-anchored fleet writes off $0.15$ of its machines instead of $0.55$. Where it predicts none, a rover gains little over blind caution, so the map holds in both directions.

PR本紙発行元 EmplifAI