EAAI Journal 2026 Journal Article
High-precision multimodal vehicle trajectory prediction model based on cross-layer interleaved spatiotemporal attention mechanism
- Fei Teng
- Liqiang Jin
- Junnian Wang
- Feng Xiao
- Mengdi Guo
- Yanbo Zhou
- Jin Zhang
In increasingly complex traffic environments, spatiotemporal attention mechanisms have made remarkable advancements in scene-level interaction modelling. However, the deep and multi-scale spatiotemporal representations required for safe and efficient decision-making in intelligent vehicles remain underexplored. Aiming to address this limitation, this study proposes a multimodal trajectory prediction model based on a cross-layer interleaved spatiotemporal attention (CLISTA) mechanism. Compared with conventional spatiotemporal attention frameworks, CLISTA more effectively captures multi-scale spatiotemporal interactions in complex traffic scenes through the alternating fusion of spatial and temporal features across network layers via a cross-layer interleaving structure. Firstly, spatial, dynamic and heading conflict risks are derived from the relative motion between the target vehicle and its neighbours and aggregated into a social grid weight matrix, through which the neighbours' collective influence on the target vehicle is quantified. Secondly, spatial and temporal multi-head attention modules are designed within each layer. By integrating an interleaved ‘spatial–temporal’ stacking strategy with cross-layer skip connections, the model facilitates progressive alignment and deep fusion, ranging from local interactions to long-range dependencies. Subsequently, an intention recognition module is developed. A second-order gated bilinear fusion mechanism is introduced to adaptively model higher-order couplings between local neighbour dynamics and global interaction semantics, thereby yielding a multimodal probability distribution over the target vehicle's driving intentions. Lastly, multimodal trajectory predictions are generated by decoding the fused spatiotemporal features together with the inferred intention information. Experimental results on three benchmark datasets—NGSIM (Next Generation Simulation), AD4CHE (Aerial Dataset for China Congested Highway and Expressway), and highD—demonstrate that CLISTA consistently outperforms the baseline methods. Relative to the next-best model, it reduces average/final displacement errors by 16. 67 %/21. 23 %, 12. 99 %/21. 14 % and 10. 53 %/21. 59 % on NGSIM, AD4CHE and HighD, respectively. Overall, CLISTA offers reliable multi-hypothesis trajectory priors for safe and efficient decision-making in complex traffic scenarios.