Incorporating data drift to perform survival analysis on credit risk

By Jianwei Peng, Stefan Lessmann

Rating

1752
Battle Count: 73

Relevance

3/10
The paper is primarily focused on credit risk and mortgage default prediction rather than quantitative trading. However, the survival analysis framework, data drift handling, and calibration techniques are transferable to trading contexts involving time-to-event modelling (e.g., counterparty default risk, credit spread modelling). The drift-adaptive modelling approach is relevant for any financial ML system operating in non-stationary environments.

Implementation Complexity

5/10
The core model is implemented using standard logistic regression with scikit-learn, making it computationally lightweight. However, the full pipeline involves: (1) computing scheduled amortisation paths and balance deviations, (2) per-loan OLS regression for longitudinal markers, (3) landmark-based sample construction, (4) one-hot encoding of landmarks, (5) isotonic calibration, and (6) drift simulation for experimentation. The methodology is modular but requires careful domain knowledge of mortgage mechanics and survival analysis.

Reproducibility

4/5
The paper provides a GitHub/Zenodo repository with processed datasets and code (https://doi.org/10.5281/zenodo.18392896). Raw data is publicly available from Freddie Mac. Hyperparameters and experimental settings are described in detail. However, the drift simulation parameters and some preprocessing steps require careful replication. The paper uses scikit-learn for implementation.

About this paper

Methodology: Landmark-based Dynamic Joint Modelling with Isotonic Calibration (LMISO). Problem types: Survival Analysis, Classification, Risk Management, Imbalanced Learning, Online Learning.

The interactive Everscope explorer (charts, battles, favorites) loads below.