Heads, Not Backbones: Output Heads Dominate Architectures on Fat-Tailed Returns

By Sichao He, Yansong Zhang

Rating

1837
Battle Count: 76

Relevance

7/10
Highly relevant for risk management, VaR/ES estimation, and regulatory capital computation (Basel FRTB IMA). The paper demonstrates that a 2-line head change (point→GMM) yields meaningful CRPS improvements (up to +6.4%) and better calibration. However, the paper explicitly shows that distributional forecasts do NOT translate to trading alpha (naive strategy loses money). The h-split finding (head dominates at h≤3, backbone at h≥6) is practically important for horizon-specific model selection. The cross-asset boundary condition (works on return-like, fails on rates/FX) limits universal applicability.

Implementation Complexity

4/10
The core experiment is straightforward: 4 backbones × 3 heads with shared hyperparameters. The GMM head is described as a '2-line change' from point head. All models use d_hidden=64, same Adam optimizer, same training protocol. The main complexity lies in the evaluation infrastructure (walk-forward folds, bootstrap CIs, MCS tests, VaR backtests, cross-asset replication). Total compute is minimal (<15 min on RTX 4060). Code is fully open-source with single CLI reproduction.

Reproducibility

5/5
Single CLI call per experiment reproduces every number. All data, code, and results available on GitHub. Fixed seeds {0,1,2}, shared hyperparameters across all variants, no backbone-specific tuning. Total wall-clock cost under 15 minutes on NVIDIA RTX 4060. Shiller dataset and FRED data are publicly available.

About this paper

Methodology: Comparative 4-backbone × 3-head experiment with anchored walk-forward validation. Problem types: Time Series Forecasting, Density Estimation, Risk Management, Portfolio Optimization.

The interactive Everscope explorer (charts, battles, favorites) loads below.