Arrow Research search

Author name cluster

Fu Liu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

2 papers
1 author row

Possible papers

2

AAAI Conference 2026 Conference Paper

Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent Space

  • Jian Zhu
  • Zhengyu Jia
  • Tian Gao
  • Jiaxin Deng
  • Shidi Li
  • Lang Zhang
  • Fu Liu
  • Peng Jia

Advanced end-to-end autonomous driving systems predict other vehicles' motions and plan ego vehicle's trajectory. The world model that can foresee the outcome of the trajectory has been used to evaluate the end-to-end autonomous driving system. However, existing world models predominantly emphasize the trajectory of the ego vehicle and leave other vehicles uncontrollable. This limitation hinders their ability to realistically simulate the interaction between the ego vehicle and the driving scenario. In addition, it remains a challenge to match multiple trajectories with each vehicle in the video to control the video generation. To address above issues, a driving World Model named EOT-WM is proposed in this paper, unifying Ego-Other vehicle Trajectories in videos. Specifically, we first project ego and other vehicle trajectories in the BEV space into the image coordinate to match each trajectory with its corresponding vehicle in the video. Then, trajectory videos are encoded by the Spatial-Temporal Variational Auto Encoder to align with driving video latents spatially and temporally in the unified visual space. A trajectory-injected diffusion Transformer is further designed to denoise the noisy video latents for video generation with the guidance of ego-other vehicle trajectories. In addition, we propose a metric based on control latent similarity to evaluate the controllability of trajectories. Extensive experiments are conducted on the nuScenes dataset, and the proposed model outperforms the state-of-the-art method by 30% in FID and 55% in FVD. The model can also predict unseen driving scenes with self-produced trajectories.

EAAI Journal 2025 Journal Article

Speech emotion recognition based on spiking neural network and convolutional neural network

  • Chengyan Du
  • Fu Liu
  • Bing Kang
  • Tao Hou

There is an urgent need to determine emotions automatically through speech signals to promote the progress of intelligent technology. However, the low accuracy problem isn't solved so far as, this hinders potential applications of Speech Emotion Recognition (SER). One of the most critical reasons for this low accuracy is that subjective emotions are random and generate weak pulse signals; moreover, they are often hidden in audio, video, and text feature which are extracted from speech. Hence, the features may not be discriminative enough to depict subjective emotions. Therefore, a dual-path SER framework is designed in this paper. Added to the traditional Convolutional Neural Network (CNN)-based SER scheme to handle speech emotion features, the Spiking Neural Network (SNN) framework is added to identify the dynamic pulse emotion features and improve the accuracy of SER. At the same time, a Perceptual Neuron Encoding Layer (PNEL) is proposed to enhance the ability to process speech signals. Overall, the experimental results on the interactive emotional dyadic motion capture database (IEMOCAP) databases show that the proposed approach can achieve 65. 3% accuracy and excellent performance in solving the SER issues compared to other existing approaches.

v2026.09.13