Rating
1790
Battle Count: 51
Relevance
9/10
Highly relevant for practitioners implementing RL-based execution algorithms. It highlights critical pitfalls in training stability, the importance of repeated-seed evaluation, and the danger of attributing performance gains to architecture when they may stem from training specification artifacts.
Implementation Complexity
6/10
Moderate complexity. Requires implementing DDQL, K-Means clustering, and a custom LOB replay simulator. The core logic is standard, but the rigorous evaluation protocol (100 seeds, paired comparisons) adds computational overhead.
Reproducibility
4/5
The paper provides detailed pseudocode for the fill mechanics, specific hyperparameters (Table 1), and states that code and configurations will be released on publication. It explicitly discusses device-dependent reproducibility issues (CPU vs CUDA) and provides a device-matched baseline for decomposition.
About this paper
Methodology: Mixture of Experts with K-Means Routing. Problem types: Algorithmic Execution, Reinforcement Learning, Clustering.
The interactive Everscope explorer (charts, battles, favorites) loads below.