Relevance
7/10
The paper provides actionable insights for quantitative trading in several ways: (1) The lead-lag relationship between fast and slow information diffusion stocks can be exploited for cross-sectional momentum strategies; (2) The predictive power of retail-driven diffusion (China) and institution-driven diffusion (U.S.) over multiple quarters can inform portfolio rebalancing timing; (3) Understanding excess comovement drivers helps in constructing more effective hedging strategies and reducing correlated risk in portfolios; (4) The mechanism through trading behavior (BSI correlation) provides a real-time signal for detecting synchronized trading; (5) The cross-market comparison highlights different alpha sources in China vs. U.S. However, the paper is primarily academic and does not provide direct trading signals or backtested strategies. The quarterly frequency of key variables limits high-frequency trading applications.
Implementation Complexity
8/10
High implementation complexity due to: (1) Massive dataset construction requiring web scraping of 600M+ forum messages and identification of co-investors across subforums; (2) Stock-pair level analysis generating ~180M (China) and ~140M (U.S.) observations requiring significant computational resources; (3) Multiple identification strategies (quasi-natural experiment, PSM, IV-2SLS) each requiring careful implementation; (4) Fama-French five-factor model estimation for residual returns; (5) Buy-sell imbalance construction with market-specific investor classification algorithms; (6) Lead-lag regression and multi-period predictive models; (7) Extensive robustness checks including alternative comovement measures, ownership thresholds, high-dimensional fixed effects, and Fama-Macbeth regressions. The variable construction pipeline alone is substantial.
Reproducibility
3/5
The paper uses well-known data sources (Eastmoney/Guba, StockTwits, CSMAR, CRSP, TAQ, Wind, Compustat, IBES, RavenPack) and standard econometric methods (FF5 model, panel regressions, PSM, IV). However, the construction of retail-driven information diffusion requires custom web scraping of over 600 million forum messages, identifying co-investors across subforums, and tracking reply chains—processes that are complex and not fully documented in code. The variable construction (Flow, Inst, BSI) is described in detail but replication requires significant data engineering effort. No code or processed data repository is provided.