Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games

By Christos S. Koulouris, Carlo Campajola

Rating

1916
Battle Count: 50

Relevance

9/10
Highly relevant for understanding how independent RL agents might inadvertently coordinate or punish each other in execution algorithms, impacting slippage and market impact. It provides critical insights for designing anti-collusive trading algorithms and regulatory compliance.

Implementation Complexity

8/10
Requires implementing a custom Almgren-Chriss simulator, a Transformer-based history encoder, and independent PPO agents with specific history masking and reward structures. The deviation testing protocol adds further complexity.

Reproducibility

4/5
The paper provides detailed hyperparameters (learning rates, network architecture, PPO parameters) and defines the simulation environment (Almgren-Chriss parameters, noise levels). However, no code repository is explicitly linked in the text provided.

About this paper

Methodology: Independent Proximal Policy Optimization (PPO) with History Encoding. Problem types: Algorithmic Execution, Reinforcement Learning, Multi-agent Learning, Game Theory.

The interactive Everscope explorer (charts, battles, favorites) loads below.