From Financial Sentiment Classification to Return Predictability: A QLoRA Benchmark of Large Language Models

By Fusheng Luo

Rating

1586
Battle Count: 120

Relevance

7/10
The paper directly addresses the core quantitative trading question of whether NLP-derived sentiment signals predict future returns. It provides a rigorous framework for evaluating cross-sectional rank ICs, long-short portfolio construction, and statistical inference (Newey-West + FDR). However, the key finding is negative: no model achieves statistically robust return predictability after multiple-testing correction. The paper is highly relevant as a benchmark and cautionary study for sentiment-based alpha research, demonstrating the gap between classification accuracy and tradable signals. The methodology (QLoRA adaptation, signal construction, portfolio evaluation) is directly applicable to quantitative trading research.

Implementation Complexity

6/10
Experiment 1 requires GPU resources for QLoRA fine-tuning of 7-8B parameter models (feasible on single consumer GPU with 4-bit quantization), plus standard NLP pipeline setup. Experiment 2 requires access to Benzinga commercial data, S&P 100 constituent history, adjusted price data, and implementation of cross-sectional IC calculations, Newey-West inference, FDR correction, and overlapping-cohort portfolio construction. The two-stage design adds complexity. Hyperparameters are well-specified but tuning and validation require iterative experimentation.

Reproducibility

4/5
The paper provides detailed hyperparameters (Tables 4, 5), fixed random seed (42), explicit dataset harmonization rules, stratified sampling protocol, and full QLoRA configuration. However, no GitHub repository or code link is provided. The Benzinga dataset access and exact preprocessing steps for the downstream evaluation are described but not fully reproducible without data access. Runtime measurements are noted as single-run observations rather than controlled benchmarks.

About this paper

Methodology: QLoRA Benchmark with Two-Stage Evaluation. Problem types: Classification, Natural Language Processing, Ranking, Transfer Learning, Zero-shot Learning, Imbalanced Learning.

The interactive Everscope explorer (charts, battles, favorites) loads below.