Arrow Research search
Back to AAMAS

AAMAS 2026

Dynamic Action Space Reinforcement Learning for Optimal Trading Execution

Conference Paper Research Paper Track Autonomous Agents and Multiagent Systems

Abstract

Optimal trading execution (OTE) involves dividing large parent ordersintosmallerchildorderstominimizetransactioncostsandmarket impact. While reinforcement learning (RL) has demonstrated its effectiveness in this domain, traditional RL methods for OTE are often constrained by a small-scale action space, limiting their performance. Simply increasing the action space results in exponential growth in trial-and-error. Dynamically pruning action space based on context offers a promising balance between performance and efficiency. However, existing RL models are designed for a fixedaction space, making them impractical for dynamic action spaces. Balancing exploration and exploitation in the dynamic action space also remains an unsolved challenge. To address these issues, we propose Dynamic Action Space Reinforcement Learning (DASRL). It iteratively leverages adaptive exploration and dynamic action space search to improve learning efficiency in trading tasks with large action spaces. Theoretical analysis shows that the DASRL reduces cumulative regret while ensuring the introduced optimality bias decreases over time. Experimental results on eight stocks from different sectors demonstrate that the DASRL significantly outperforms traditional strategies and RL baselines, reducing transaction costs by up to 16. 2% and improving learning efficiency by up to 14. 0%. It highlights the effectiveness of the DASRL in the OTE task with large-scale action spaces.

Authors

Keywords

  • ReinforcementLearning
  • AlgorithmicTrading
  • DeepLearning
  • Quantitative Finance

Context

Venue
International Conference on Autonomous Agents and Multiagent Systems
Archive span
2002-2026
Indexed papers
8043
Paper id
1104102036227465104
v2026.09.13