From Word Counts to Context: Topic Models for Asset Pricing

By Kevin Foley, Jonathan Hartadi, Shivesh Prakash, Swapnil Vatsal

Rating

1500
Battle Count: 0

Relevance

8/10
Highly relevant for strategies utilizing alternative data (news text) for alpha generation. The paper demonstrates how contextual embeddings can improve topic coherence and potentially risk factor extraction, though statistical significance is currently lacking.

Implementation Complexity

7/10
Requires handling large unstructured text datasets, training/fine-tuning or using pre-trained transformers, implementing k-means clustering on high-dimensional embeddings, and integrating with Sparse IPCA factor models. The pipeline is modular but computationally intensive.

Reproducibility

4/5
Code and configuration files are available on GitHub. The FNSPID dataset is public, but CRSP data requires WRDS access. The paper details preprocessing steps, hyperparameters (e.g., temperature tau=0.07, L=60 topics), and evaluation windows.

About this paper

Methodology: Narrative Asset Pricing with Contextual Embeddings. Problem types: Topic Modeling, Clustering, Portfolio Optimization, Risk Management, Natural Language Processing.

The interactive Everscope explorer (charts, battles, favorites) loads below.