PPO-HRAP: Proximal Policy Optimization with a Hybrid Regime-Aware Policy for Risk-Controlled Trading

By Duong Hien Chi Kien, Huynh Thanh Trung

Rating

1524
Battle Count: 51

Relevance

9/10
Highly relevant as it addresses a core problem in RL trading: balancing upside participation with drawdown control. The hybrid approach of blending learned policies with interpretable regime priors is a practical engineering solution for risk-adjusted returns.

Implementation Complexity

6/10
Moderate complexity. Requires implementing PPO, custom reward shaping with dynamic coefficients based on VIX, regime detection logic, and action blending. The architecture is described clearly, but tuning the blend coefficient and reward weights is non-trivial.

Reproducibility

4/5
The paper provides a GitHub repository link, detailed hyperparameters (Table 3), feature engineering formulas, and specific data splits (2010-2017 train, 2018-2019 val, 2020-2022 test). Multi-seed analysis is included for SPY.

About this paper

Methodology: PPO-HRAP. Problem types: Reinforcement Learning, Portfolio Optimization, Risk Management.

The interactive Everscope explorer (charts, battles, favorites) loads below.