Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Stock return prediction with Transformer-style architectures faces a practical trade-off between expressive relational modeling and scalability on large stock universes. To address this issue, this paper proposes Domain-Aware Relational Transformer (DART), an efficient framework that improves both predictive performance and computational cost for large-scale stock prediction. DART contains two core components. First, a Multi-Scale Temporal Fusion (MSTF) module extracts short-, medium-, and longer-horizon signals through parallel convolutions and aggregates them with a Progressive Temporal Fusion (PTF) network to capture cumulative temporal dynamics. Second, a Domain-Knowledge-Guided Sparse Attention (DSA) mechanism uses industry classification priors to constrain cross-stock attention to economically relevant neighborhoods, reducing unnecessary interactions while preserving informative co-movement structure. Experiments on three real-world U.S. market datasets, NASDAQ, NYSE, and S&amp;P 500, compare DART with classical baselines and recent strong models, including STHAN-SR, CausalStock, and MambaStock, under a unified setting. The results show that DART achieves state-of-the-art performance across the evaluated benchmarks, delivering the strongest overall results in Information Coefficient (IC), Rank Information Coefficient (RIC), Precision@10, and Sharpe Ratio, with especially clear advantages in risk-adjusted returns. In addition, floating-point-operation analysis shows that the sparse attention design reduces the core computational load by about 83%-84% relative to full attention. These findings indicate that DART provides a practical, scalable, and state-of-the-art solution for large-scale stock prediction and portfolio-oriented ranking tasks.</p>

Show More

Keywords

dart stock attention prediction temporal

Related Articles

PORE

About

Connect