Rating
1767
Battle Count: 68
Relevance
7/10
Highly relevant to quantitative trading. The paper directly addresses sentiment extraction from regulatory filings (10-K) for both return and volatility prediction, which are core inputs for trading strategies. The aggregation-dependent findings (full text better at sector/portfolio level, Item 1A better at firm level) provide actionable guidance for signal construction. The Kalman filtering approach for smoothing irregular filing arrivals is directly applicable to real-time signal generation. The supervised approach outperforming dictionary baselines is practically significant. However, the paper is more of a methodology/ablation study than a backtested trading strategy, and the sample is limited to one sector and one firm at the individual level.
Implementation Complexity
5/10
Moderate complexity. The core model involves document-term matrix construction, word scoring (Eq. 2-3), least-squares estimation of topic distributions (Eq. 4), and penalized log-likelihood optimization (Eq. 5). The Kalman filter (Eqs. 6-7) adds time-series smoothing. The main complexity lies in the data pipeline: HTML parsing from EDGAR, Item 1A extraction via regex and structural heuristics, and handling of 1,383 filings across 94 firms. The model itself is relatively simple (no deep learning), but the NLP preprocessing and extraction pipeline requires careful engineering. Hyperparameter tuning for volatility labels adds some complexity.
Reproducibility
3/5
The methodology is well-described with explicit equations (Eqs. 1-7), hyperparameters (lambda=0.1, alpha+/alpha- thresholds, k minimum count, theta=65th percentile, q=0.65), and a clear data pipeline (EDGAR retrieval, HTML parsing, Item 1A extraction). However, no code or repository is mentioned. The data source (SEC EDGAR) is publicly accessible, and the universe (Nasdaq-100 constituents) is well-defined. The qualitative word-theme interpretation is subjective and not independently validated.
About this paper
Methodology: Supervised Lexicon-Learning Sentiment Score Prediction Model. Problem types: Classification, Natural Language Processing, Risk Management, Time Series Forecasting.
The interactive Everscope explorer (charts, battles, favorites) loads below.