Abstract
<title>Abstract</title> <p> This systematic review synthesizes evidence from 46 empirical studies (2007–2025) that extract sentiment from textual data for economic forecasting, supplemented by 20 theoretical papers developing macroeconomic models of sentiment. We conduct a comprehensive quantitative synthesis establishing that sentiment-based models improve forecast accuracy by 12–20% (median RMSE reduction), with the precise estimate depending on baseline strength: studies using strong benchmarks (professional forecasts, ARIMA) report ∼12–15%, while those using weak baselines (AR(1)) report larger but less realistic gains. The unconditional median across all studies is 20%, but this pools comparisons against baselines of markedly different strength: approximately 35% of reviewed studies compare against AR(1) or near-unconditional benchmarks, while only ∼25% compare against professional forecasts (SPF, Blue Chip). Well-documented publication bias toward positive forecasting results further tempers the headline figure. We document clear methodological evolution from dictionary-based approaches (55% of studies) through traditional ML and transformer models to emerging Large Language Models (3%), characterize validation practices (83% employ out-of-sample testing, but only 33% follow real-time data protocols), and assess methodological quality via structured risk-of-bias scoring. Critical geographic and linguistic gaps persist: 56% of studies focus on the US, 84% analyze English text, and zero sentiment indices exist for Africa, Latin America, or Bangla-speaking regions (265 million speakers). We develop a best-practices framework for sentiment forecasting validation and propose expansion strategies, using the Bangla Economic Narrative Index (BENI) as a case study for bridging geographic gaps. <bold>JEL Codes:</bold> C53, C55, E37, E44, G12, G17 <bold>arXiv:</bold> econ.EM (Primary), q-fin.ST, cs.CL </p>