患者ワールドモデルにおける反復測定リーク・分布シフト・部分観測下での信頼性
Repeated-Measure Leakage, Distribution Shift, and Reliability under Partial Observation in Patient World Models
PhysioNet GaitPDBのデジタルバイオマーカー予測を題材に、評価条件を段階的に厳しくして、患者間汎化や分布シフト・欠測下での予測性能と信頼性を検証した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
分類: cs.LG
原文アブストラクト
Patient world models are increasingly proposed for longitudinal prediction, intervention-aware reasoning, and clinical-trial simulation. Causal or clinical intervention validity is distinct from predictive generalization and reliability; before making stronger claims, the underlying predictive state should generalize across patients, survive realistic shifts and missing observations, and expose failure through meaningful reliability signals. We evaluate these prerequisites in a deliberately narrow setting: short-horizon digital-biomarker forecasting from PhysioNet GaitPDB, comprising 165 participants, 306 recordings, and 51,129 context-future pairs. Using persistence, ridge, MLP, GRU, Transformer, and a compact JEPA-style predictor, we build an evaluation ladder that progressively removes raw temporal overlap, same-recording familiarity, and same-patient familiarity before testing unseen-patient generalization. For GRU, NMSE rises from 0.1227 under random-window splitting to 0.1393 after eliminating raw train-test overlap and to 0.1961 under patient holdout. Among 54 participants with repeated recordings, exposure to a different recording from the same patient improves GRU NMSE from 0.2177 to 0.1556, while a recording-excluded identity hypothesis is not supported at the participant level. Under participant-held-out evaluation, MLP and Transformer are statistically indistinguishable. Study shift, a four-times-longer prediction gap, and partial observation further degrade performance; under 50% temporal masking, Transformer NMSE rises to 0.611 while MC-dropout predictive variance falls. We do not claim a longitudinal or intervention-aware simulator. Instead, the results support a prerequisite evaluation stack of patient separation, repeated-measure controls, shift, missingness, and uncertainty validation before stronger patient-world-model claims are trusted.