IS Journal 2026 Journal Article
Learning Motion-Scene Disentanglement for Trajectory Prediction From Videos
- Haowen Tang
- Ping Wei
- Ziyang Ren
- Huan Li
Agent trajectory prediction plays a significant role in various intelligent systems, as it entails the accurate anticipation of the future trajectory based on the historical data. Conventional approaches often rely on ready-made trajectory coordinates as inputs, which remain inapplicable in video-based scenarios. While two-stage trajectory prediction methods based on tracking-prediction paradigms have made progress in predicting trajectories, they still suffer from the information degradation and error accumulation due to the independence stages. In this article, we propose an end-to-end model (MSDN) to directly predict future trajectories from videos. We design novel disentanglement structures to explicitly learn the motion-aware and scene-aware representations from videos to avoid information degradation. To alleviate error accumulation, we propose the temporal consistency learning structure. Extensive experiments on the ETH-UCY dataset and the Stanford Drone Dataset demonstrate that the proposed model not only outperforms the previous approaches, but also significantly enhances the inference speed.