日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
金融時系列表現学習arXiv:2610.09048

金融市場の世界モデル構築に向けて

Towards Financial World Modeling

シェア:XThreadsFacebookLINEはてブBluesky

米国株の約1兆観測を含む大規模データセットMarket-1Tを構築し、18種のエンコーダ学習戦略を比較評価して、金融市場の世界モデルに向けた表現学習の基盤を確立した。

詳しい要約

1. どんなもの?

- 金融市場のworld model構築を目指す研究。 - 計画・意思決定に有用な状態表現の学習と評価を目的とする。 - 米国株式のほぼ1兆観測を含むMarket-1Tデータセットを導入。 - 2008–2025年、1 Hz解像度。 - 厳密な評価プロトコルを開発。 - 18のencoder学習戦略を約20年の市場レジームで比較。 - 予測有用性と潜在構造のprobeで評価。

2. 先行研究と比べてどこがすごい?

- 従来の金融表現学習は個別の予測タスクで評価されがち。 - 単一期間・狭いデータセットに偏ることが多かった。 - 本研究は大規模データと長期レジームで系統的に比較。 - 予測性能が同程度でも市場状態の組織化が異なることを示す。 - world model向けの基盤を提供。

3. 技術・手法の肝は?

- Market-1T: 米国株式の約1兆観測、2008–2025、1 Hz。 - 厳密な評価プロトコルを設計・実装。 - 18のencoder-training戦略を大規模比較。 - 予測有用性とlatent structureのprobeで評価。 - 市場全体の条件、資産別期待リターン、流動性、ボラティリティ、資産間関係を考慮。

4. どうやって有効だと検証した?

- 約20年の市場レジームにわたる大規模研究。 - 18のencoder学習戦略を比較。 - 一般的な金融タスクでの予測有用性を評価。 - 潜在構造のprobeを実施。 - 予測性能が類似でも市場状態の組織化が異なることを発見。

5. 議論はある?

- 予測性能が同程度のencoderでも市場状態の表現が大きく異なる。 - 金融表現学習の評価は個別タスク・単一期間に偏っていた。 - 大規模・長期データと厳密評価の必要性を示唆。 - 具体的な議論や限界は要旨からは不明。

6. 次に読むべき論文は?

- DINO-WM - V-JEPA 2 - LeWM - 関連するworld modelや金融表現学習の手法

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Humzah Merchant, Alec Guthrie, Simon Mahns, Randall Balestriero, Bradford Levy

分類: cs.LG

原文アブストラクト

Building a world model requires a state representation useful for planning and decision-making---potentially over tasks unknown at training time. In the context of financial markets, planning and decision-making may require a model to reason about market-wide conditions, asset-specific expected returns, liquidity, volatility, and cross-asset relationships. Yet financial representation learning has largely been evaluated on individual predictive tasks, oftentimes on a single time period using comparatively narrow datasets. We address this through three primary contributions. First, we introduce Market-1T, a dataset containing nearly one trillion observations across U.S. equities from 2008 to 2025 at 1 Hz resolution. Second, we develop and implement a rigorous evaluation protocol. Third, we conduct a systematic large-scale study of financial representation learning, comparing 18 encoder-training strategies across nearly two decades of market regimes. We evaluate learned representations both by their predictive utility on common finance tasks and through probes of latent structure. We find that encoders with similar predictive performance can organize market state very differently. Collectively, we establish a foundation for training and evaluating financial market representations in support of world models such as DINO-WM, V-JEPA 2, and LeWM.

PR本紙発行元 EmplifAI