Optimal Stratified Allocation for Rare-Event Onset Forecasting in Dependent Sequences

By Jaskaran Singh

Rating

1573
Battle Count: 74

Relevance

4/10
The paper explicitly states it runs no backtest and makes no claim of profitability. However, it addresses a relevant screening problem: ranking 350 securities by probability of entering an explosive regime within 5-20 trading days. The PSY detector is deployed as infrastructure by the Dallas Fed. The allocation theory (Neyman optimal for rare events) and score alignment across retraining windows are directly applicable to any ML pipeline for rare-event financial prediction. The early-warning operating points (42.5% episode recall at 1.64 false alerts/security-year with 6.5-day median lead) could inform risk monitoring systems. The work is more relevant to risk management and regulatory monitoring than to direct alpha generation.

Implementation Complexity

7/10
The theoretical framework (Neyman allocation, Serfling bounds, conformal validity) is mathematically sophisticated but the implementation is straightforward: stratified sampling with class-balanced allocation, LightGBM with class weights, Platt calibration, and empirical CDF rank transformation. The PSY/BSADF detector implementation with simulated critical values adds complexity. The purged expanding-window CV with 140-day gaps, episode filtering (merge gap + minimum duration), and the full evaluation protocol (54 configurations × 5 folds × 3 seeds) require careful engineering. K-means representative selection adds O(N_c K_c d) cost. The main complexity lies in the evaluation protocol design and ensuring no temporal leakage.

Reproducibility

4/5
The paper provides a repository with manuscript source, generated tables, orchestration scripts, and all run artifacts (885 model configurations). Seeds are fixed (11, 23, 47). Data comes from a public-domain Kaggle archive. However, the specific PSY detector implementation details (2000 simulated random walks, specific merge/duration filters) and the exact LightGBM hyperparameters are documented. The 2.5% class-balanced configuration used in operational analyses was fixed before multiplicity analysis but its selection is not independently documented.

About this paper

Methodology: Neyman-optimal stratified allocation for class-weighted empirical risk with purged temporal validation. Problem types: Classification, Time Series Forecasting, Ranking, Imbalanced Learning, Anomaly Detection, Optimization.

The interactive Everscope explorer (charts, battles, favorites) loads below.