AAMAS 2026
Dynamic Action Space Reinforcement Learning for Optimal Trading Execution
Abstract
Optimal trading execution (OTE) involves dividing large parent ordersintosmallerchildorderstominimizetransactioncostsandmarket impact. While reinforcement learning (RL) has demonstrated its effectiveness in this domain, traditional RL methods for OTE are often constrained by a small-scale action space, limiting their performance. Simply increasing the action space results in exponential growth in trial-and-error. Dynamically pruning action space based on context offers a promising balance between performance and efficiency. However, existing RL models are designed for a fixedaction space, making them impractical for dynamic action spaces. Balancing exploration and exploitation in the dynamic action space also remains an unsolved challenge. To address these issues, we propose Dynamic Action Space Reinforcement Learning (DASRL). It iteratively leverages adaptive exploration and dynamic action space search to improve learning efficiency in trading tasks with large action spaces. Theoretical analysis shows that the DASRL reduces cumulative regret while ensuring the introduced optimality bias decreases over time. Experimental results on eight stocks from different sectors demonstrate that the DASRL significantly outperforms traditional strategies and RL baselines, reducing transaction costs by up to 16. 2% and improving learning efficiency by up to 14. 0%. It highlights the effectiveness of the DASRL in the OTE task with large-scale action spaces.
Authors
Keywords
Context
- Venue
- International Conference on Autonomous Agents and Multiagent Systems
- Archive span
- 2002-2026
- Indexed papers
- 8043
- Paper id
- 1104102036227465104