Submitted:
14 September 2026
Posted:
14 September 2026
You are already at the latest version
Abstract
Stock movement prediction remains challenging because financial time series are nonstationary, noisy, and interdependent across heterogeneous information sources. Conventional multimodal forecasting methods often integrate price, text, and inter-firm relationships into a single representation, making it difficult to capture modality-pair-specific interactions and their varying relevance across market conditions. To address this issue, this study proposes Pairwise Cross-Modal Attention Fusion with Mixture-of-Experts (PaCMoE), which explicitly decomposes multimodal interactions into three modality pairs— price-text, price-graph, and text-graph pair. Each pair is modeled by an independent cross-modal attention expert, and a Mixture-of-Experts router adaptively integrates the resulting representations according to the input. On the ACL18 (StockNet) benchmark, PaCMoE is evaluated using chronological data splitting, validation-based model selection, and ten random seeds. The proposed method achieves an average accuracy of 60.69% and an MCC of 0.2200, outperforming the compared methods evaluated with multiple runs or random seeds in both metrics. These results demonstrate the effectiveness of explicitly modeling pairwise cross-modal interactions and adaptively integrating their contributions for robust stock movement prediction.
Keywords:
multimodal fusion
; stock movement prediction
; cross-a
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.