Arrow Research search

Author name cluster

Yulong Liu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

9 papers
2 author rows

Possible papers

9

AAAI Conference 2026 Conference Paper

Single-Stage fMRI-to-3D Reconstruction via Viewpoint-Aware Embedding and Hierarchical Guidance

  • Xun Zhang
  • Weihao Xia
  • Yulong Liu
  • Bo Yang
  • Alessandro Bozzon
  • Pan Wang

Understanding the neural basis of three-dimensional (3D) perception is a fundamental objective in cognitive neuroscience. Despite advances in decoding 2D visual stimuli from neural data, reconstructing high-fidelity 3D objects with detailed texture and geometry remains largely unexplored. In this work, we introduce NeuroSculptor3D, the first single-stage, end-to-end framework for reconstructing textured 3D shapes directly from brain activity. NeuroSculptor3D integrates a viewpoint-aware brain embedding module that captures fine-grained spatial variations across visual perspectives, and a hierarchical guidance mechanism that aligns brain-derived features with perceptual, semantic, and structural priors. Together, these components facilitate the generation of consistent multi-view embeddings, which are then decoded via TRELLIS to produce high-quality textured 3D reconstructions. Experiments on the fMRI-Shape dataset demonstrate that NeuroSculptor3D outperforms existing baselines across multiple settings, achieving significant improvements in both structural accuracy and semantic consistency. Code will be released to facilitate further research.

EAAI Journal 2025 Journal Article

A multi-temporal granularity feature driven convolutional ensemble model for electricity theft detection

  • Mingfa Yang
  • Qinyu Huang
  • Yulong Liu
  • Xidong Zheng
  • Tao Jin
  • Mohamed A. Mohamed

Electricity theft causes substantial economic losses and safety hazards. While the widespread adoption of advanced metering infrastructure has significantly reduced electricity theft, perpetrators continue to find ways to exploit the system, employing increasingly covert and intricate methods. To address the ongoing challenge, this paper proposes an attention mechanism optimized multi-temporal granularity feature driven convolutional ensemble model for enhanced accuracy and robustness in electricity theft detection (ETD). For comprehensive feature extraction across diverse temporal scales, the proposed framework integrates two specialized feature extraction modules. The first module, a squeeze-and-excitation network-optimized temporal convolutional network, selectively focuses on informative temporal features within the electricity consumption data. The second module, a dual-dimensional attention enhanced deep residual network composed of residual blocks embedded with the convolutional block attention module, facilitates the model's concurrent learning of informative spatial and temporal features. Then, the features from each module are fused and classified through a fully connected layer. To validate the effectiveness of the proposed ETD method, this paper conducted simulation experiments using the publicly available dataset from the State Grid Corporation of China. The experimental results show that the model optimized with the attention mechanism significantly improves the performance of ETD. Compared to other ETD models, the proposed model performs excellently in various indicators under different training set ratios and sample imbalance scenarios, demonstrating good generalization and robustness. Additionally, the model was deployed on a Raspberry Pi edge computing device to further verify its feasibility in practical engineering applications.

NeurIPS Conference 2025 Conference Paper

Foundation Cures Personalization: Improving Personalized Models’ Prompt Consistency via Hidden Foundation Knowledge

  • Yiyang Cai
  • Zhengkai Jiang
  • Yulong Liu
  • Chunyang Jiang
  • Wei Xue
  • Yike Guo
  • Wenhan Luo

Facial personalization faces challenges to maintain identity fidelity without disrupting the foundation model's prompt consistency. The mainstream personalization models employ identity embedding to integrate identity information within the attention mechanisms. However, our preliminary findings reveal that identity embeddings compromise the effectiveness of other tokens in the prompt, thereby limiting high prompt consistency and attribute-level controllability. Moreover, by deactivating identity embedding, personalization models still demonstrate the underlying foundation models' ability to control facial attributes precisely. It suggests that such foundation models' knowledge can be leveraged to cure the ill-aligned prompt consistency of personalization models. Building upon these insights, we propose FreeCure, a framework that improves the prompt consistency of personalization models with their latent foundation models' knowledge. First, by setting a dual inference paradigm with/without identity embedding, we identify attributes (e. g. , hair, accessories, etc. ) for enhancements. Second, we introduce a novel foundation-aware self-attention module, coupled with an inversion-based process to bring well-aligned attribute information to the personalization process. Our approach is training-free, and can effectively enhance a wide array of facial attributes; and it can be seamlessly integrated into existing popular personalization models based on both Stable Diffusion and FLUX. FreeCure has consistently shown significant improvements in prompt consistency across these facial personalization models while maintaining the integrity of their original identity fidelity.

AAAI Conference 2025 Conference Paper

See Through Their Minds: Learning Transferable Brain Decoding Models from Cross-Subject fMRI

  • Yulong Liu
  • Yongqiang Ma
  • Guibo Zhu
  • Haodong Jing
  • Nanning Zheng

Deciphering visual content from fMRI sheds light on the human vision system, but data scarcity and noise limit brain decoding model performance. Traditional approaches rely on subject-specific models, which are sensitive to training sample size. In this paper, we address data scarcity by proposing shallow subject-specific adapters to map cross-subject fMRI data into unified representations. A shared deep decoding model then decodes these features into the target feature space. We use both visual and textual supervision for multi-modal brain decoding and integrate high-level perception decoding with pixel-wise reconstruction guided by high-level perceptions. Our extensive experiments reveal several interesting insights: 1) Training with cross-subject fMRI benefits both high-level and low-level decoding models; 2) Merging high-level and low-level information improves reconstruction performance at both levels; 3) Transfer learning is effective for new subjects with limited training data by training new adapters; 4) Decoders trained on visually-elicited brain activity can generalize to decode imagery-induced activity, though with reduced performance.

ECAI Conference 2024 Conference Paper

Hidden States in LLMs Improve EEG Representation Learning and Visual Decoding

  • Aoyang Liu
  • Haodong Jing
  • Yulong Liu
  • Yongqiang Ma
  • Nanning Zheng 0001

Analyzing brain signals and reconstructing visual stimuli from the brain can facilitate further exploration on cognitive functions of the human brain, which have attracted strong interest in neuroscience and artificial intelligence. However, due to defects such as complex noises and the lack of alignment accuracy, efficient methods for extracting information from Electroencephalogram (EEG) signals are still very limited, making it difficult to perform EEG visual decoding tasks. Our study shows a way to handle the issues by proposing a new method for EEG representation learning and visual decoding, thus completing end-to-end image reconstruction tasks from EEG signals. We utilize the ability of semantic extraction and prediction of large language models (LLMs) to enhance the performance of EEG feature extraction. For semantic representation learning, we align EEG signals with target semantic embeddings, which are obtained from hidden states of Large Language Model Meta AI 2 (LLaMa-2) by inputting descriptions of images into the model. We also extract visual features from EEG signals to improve the quantity of the reconstructed images at low levels. Then we fuse semantic features and visual features by applying a pre-trained diffusion model and finally generate the corresponding images. We are the first to incorporate the LLM into EEG visual decoding tasks. Our method achieves the state-of-the-art result of EEG classification accuracy and the quality of reconstructed images on ImageNet-EEG datasets. In one word, our work is an important step forward in the field of exploiting the relationship between language models and human visual cognition. Our codes are available at https: //github. com/lay-atsa/llm4eeg.

EAAI Journal 2024 Journal Article

Research on state-parameter estimation of unmanned Tractor—A hybrid method of DEKF and ARBFNN

  • Guangfei Xu
  • Meizhou Chen
  • Xiangkun He
  • Yulong Liu
  • Jian Wu
  • Peisong Diao

Unmanned tractor relies on multi-sensor information collection to obtain the current state or parameters. However, when driving in complex field, it will inevitably suffer from the uneven ground, the impact of crop straw and clods et al. , which usually poses considerable challenges for multi-sensor to obtain stable and accurate value of states and parameters to realize tractor automatic lane guidance control. Therefore, a novel state and parameter estimation method by mixing the dual extended Kalman filter (DEKF) technology and adaptive radial basis function neural network (ARBFNN) technology is proposed in this paper. Firstly, DEKF technique is applied to estimate key states and initial model time-varying parameters–front/rear axle cornering stiffnesses at the same time. Then, in order to further improve the accuracy of estimation value of front/rear axle cornering stiffnesses during the automatic lane guidance control process, an ARBFNN technology is investigated by taking the heading error, lateral error and initial estimation value of front/rear axle cornering stiffnesses from DEKF as inputs to approach ideal estimation value. Finally, results of automatic lane guidance control scenarios from both simulation and hardware-in-loop (HIL) implementation show that the proposed hybrid estimated method of DEKF and ARBFNN can robustly obtain satisfactory estimation value and automatic lane guidance control performance for tractor when the control system is characterized by both time-varying model parameters and uncertain external energy-bounded disturbance. A comparative study is also conducted to investigate cases to show its effectiveness when the hybrid DEKF-ARBFNN state and parameter estimation method is used and when it is not.

NeurIPS Conference 2022 Conference Paper

TaiSu: A 166M Large-scale High-Quality Dataset for Chinese Vision-Language Pre-training

  • Yulong Liu
  • Guibo Zhu
  • Bin Zhu
  • Qi Song
  • Guojing Ge
  • Haoran Chen
  • GuanHui Qiao
  • Ru Peng

Vision-Language Pre-training (VLP) has been shown to be an efficient method to improve the performance of models on different vision-and-language downstream tasks. Substantial studies have shown that neural networks may be able to learn some general rules about language and visual concepts from a large-scale weakly labeled image-text dataset. However, most of the public cross-modal datasets that contain more than 100M image-text pairs are in English; there is a lack of available large-scale and high-quality Chinese VLP datasets. In this work, we propose a new framework for automatic dataset acquisition and cleaning with which we construct a new large-scale and high-quality cross-modal dataset named as TaiSu, containing 166 million images and 219 million Chinese captions. Compared with the recently released Wukong dataset, our dataset is achieved with much stricter restrictions on the semantic correlation of image-text pairs. We also propose to combine texts collected from the web with texts generated by a pre-trained image-captioning model. To the best of our knowledge, TaiSu is currently the largest publicly accessible Chinese cross-modal dataset. Furthermore, we test our dataset on several vision-language downstream tasks. TaiSu outperforms BriVL by a large margin on the zero-shot image-text retrieval task and zero-shot image classification task. TaiSu also shows better performance than Wukong on the image-retrieval task without using image augmentation for training. Results demonstrate that TaiSu can serve as a promising VLP dataset, both for understanding and generative tasks. More information can be referred to https: //github. com/ksOAn6g5/TaiSu.

NeurIPS Conference 2021 Conference Paper

Linear Convergence of Gradient Methods for Estimating Structured Transition Matrices in High-dimensional Vector Autoregressive Models

  • Xiao Lv
  • Wei Cui
  • Yulong Liu

In this paper, we present non-asymptotic optimization guarantees of gradient descent methods for estimating structured transition matrices in high-dimensional vector autoregressive (VAR) models. We adopt the projected gradient descent (PGD) for single-structured transition matrices and the alternating projected gradient descent (AltPGD) for superposition-structured ones. Our analysis demonstrates that both gradient algorithms converge linearly to the statistical error even though the strong convexity of the objective function is absent under the high-dimensional settings. Moreover our result is sharp (up to a constant factor) in the sense of matching the phase transition theory of the corresponding model with independent samples. To the best of our knowledge, this analysis constitutes first non-asymptotic optimization guarantees of the linear rate for regularized estimation in high-dimensional VAR models. Numerical results are provided to support our theoretical analysis.

ICRA Conference 2014 Conference Paper

Simultaneous prototype selection and outlier isolation for traffic sign recognition: A collaborative sparse optimization method

  • Huaping Liu 0001
  • Yulong Liu
  • Yuanlong Yu 0001
  • Fuchun Sun 0001

Video-based traffic sign recognition is one of the most important task for unmanned autonomous vehicle. However, there always exists unavoidable outliers in the practical scenario. Therefore, robust prototype extraction from the noisy sample set is highly expected to help traffic sign recognition in video sequence. In this paper, we propose a novel approach for simultaneous prototype extraction and outlier isolation through collaborative sparse learning. The new model accounts for not only the reconstruction capability and the sparsity, but also the robustness. To solve the optimization problem, we adopt the Alternating Directional Method of Multiplier (ADMM) technology to design an iterative algorithm. Finally, the effectiveness of the approach is demonstrated by experiments on GTSRB dataset.

v2026.09.13