Arrow Research search

Author name cluster

Yu Zeng

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

12 papers
1 author row

Possible papers

12

EAAI Journal 2025 Journal Article

Confidence-aware iterative training for cross-lingual entity alignment

  • Fan Ye
  • Yu Zeng
  • Zhangling Duan
  • Zhaolong Ling
  • Yun Yang

Entity alignment (EA) aims to identify equivalent entities across knowledge graphs, serving as a critical step in integrating multi-source knowledge graphs. In recent years, EA methods relying on alignment seeds have achieved impressive performance. However, the high cost of manually labeled alignment seeds has posed a significant limitation to their practical applications in the real world. Most methods adopt iterative strategies to generate pseudo-alignment seeds automatically. However, unreliable iterative strategies introduce a large number of noisy pseudo-alignment seeds, leading to low-quality entity embeddings. In this paper, we propose a Confidence-Aware Iterative Training (CAIT) framework for unsupervised EA tasks. The framework initially extracts semantic and structural features of entities from knowledge graphs, then generates pseudo-alignment seeds by the confidence-aware iterative strategy, which limits the quantity of noisy pseudo-alignment seeds and progressively enhances the quality of entity embeddings. Extensive experiments conducted on widely used benchmark datasets demonstrate that CAIT outperforms existing state-of-the-art methods in cross-lingual EA tasks, both with and without prior alignment seeds.

JBHI Journal 2025 Journal Article

Unsupervised Feature Selection-Driven Active Learning for Semi-Supervised Automatic ECG Analysis

  • Xiao Li
  • Yongkang Zhou
  • Songyang An
  • Yu Zeng
  • Xinqi Zhang
  • Jun Wang
  • Yizhe Huang
  • Fan Lin

Automatic analysis methods of electrocardiograms (ECGs) usually required large-scale annotated training data, but the annotation process is extremely time-consuming. While semi-supervised learning can leverage unlabeled data, its performance depends heavily on the quality of the initial labeled subset. Active learning has been used to identify the most informative samples for annotation, but conventional approaches face three critical limitations: (1) dependency on manual intervention for iterative query design, (2) prohibitive computational costs during sample selection, and (3) limited compatibility with semi-supervised learning frameworks. To address these limitations, we proposed an Unsupervised Active Feature-selective Semi-Supervised Learning (UAFSSL) framework for ECG analysis, including an unsupervised feature selection-based active learning module and a semi-supervised learning module. UAFSSL captures latent data distributions via unsupervised feature extraction, selects diverse and representative samples using pseudo-label clustering, and integrates seamlessly with semi-supervised learning to eliminate human intervention. We validated our algorithm on an ECG waveform segmentation task and an atrial fibrillation detection task. In the waveform segmentation task, our method improved the F1-score for P-wave delineation by 2. 4% compared to random sampling, using only 5% of labeled samples. For the atrial fibrillation detection task, we evaluated our method on both the AFDB and a 24-hour dataset collected from 500 atrial fibrillation patients. Using only 200 labeled samples for model training, our method achieved AUC improvements of 2. 5% and 2. 2% over random sampling in five-fold cross validation. This is the first study to integrate unsupervised active learning with semi-supervised learning for automatic ECG analysis, offering a robust, automated solution to reduce annotation costs while enhancing clinical applicability.

AAAI Conference 2025 Conference Paper

VFM-Adapter: Adapting Visual Foundation Models for Dense Prediction with Dynamic Hybrid Operation Mapping

  • Zheng Chen
  • Yu Zeng
  • Zehui Chen
  • Hongzhi Gao
  • Lin Chen
  • Jiaming Liu
  • Feng Zhao

Although pre-trained large vision foundation models (VFM) yield superior results on various downstream tasks, full fine-tuning is often impractical due to its high computational cost and storage requirements. Recent advancements in parameter-efficient fine-tuning (PEFT) of VFM for image classification show significant promise. However, the application of PEFT techniques to dense prediction tasks remains largely unexplored. Our analysis of existing methods reveals that the underlying premise of utilizing low-rank parameter matrices, despite their efficacy in specific applications, may not be adequately suitable for dense prediction tasks. To this end, we propose a novel PEFT learning approach tailored for dense prediction tasks, namely VFM-Adapter. Specifically, the VFM-Adapter introduces a hybrid operation mapping technique that seamlessly integrates local information with global modeling to the adapter module. It capitalizes on the distinct inductive biases inherent in different operations. Additionally, we dynamically generate parameters for the VFM-Adapter, enabling flexibility of feature extraction given specific inputs. To validate the efficacy of VFM-Adapter, we conduct extensive experiments across object detection, semantic segmentation, and instance segmentation tasks. Results on multiple benchmarks consistently demonstrate the superiority of our method over previous approaches. Notably, with only three percent of the trainable parameters of the SAM-Base backbone, our approach achieves competitive or even superior performance compared to full fine-tuning. The code will be available.

NeurIPS Conference 2025 Conference Paper

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

  • Qiuchen Wang
  • Ruixue Ding
  • Yu Zeng
  • Zehui Chen
  • Lin Chen
  • Shihang Wang
  • Pengjun Xie
  • Fei Huang

Effectively retrieving, reasoning and understanding visually rich information remains a challenge for traditional Retrieval-Augmented Generation (RAG) methods. On the one hand, traditional text-based methods cannot handle visual-related information. On the other hand, current vision-based RAG approaches are often limited by fixed pipelines and frequently struggle to reason effectively due to the insufficient activation of the fundamental capabilities of models. As reinforcement learning (RL) has been proven to be beneficial for model reasoning, we introduce VRAG-RL, a novel RL framework tailored for complex reasoning across visually rich information. With this framework, VLMs interact with search engines, autonomously sampling single-turn or multi-turn reasoning trajectories with the help of visual perception tokens and undergoing continual optimization based on these samples. Our approach highlights key limitations of RL in RAG domains: (i) Prior Multi-modal RAG approaches tend to merely incorporate images into the context, leading to insufficient reasoning token allocation and neglecting visual-specific perception; and (ii) When models interact with search engines, their queries often fail to retrieve relevant information due to the inability to articulate requirements, thereby leading to suboptimal performance. To address these challenges, we define an action space tailored for visually rich inputs, with actions including cropping and scaling, allowing the model to gather information from a coarse-to-fine perspective. Furthermore, to bridge the gap between users' original inquiries and the retriever, we employ a simple yet effective reward that integrates query rewriting and retrieval performance with a model-based reward. Our VRAG-RL optimizes VLMs for RAG tasks using specially designed RL strategies, aligning the model with real-world applications. Extensive experiments on diverse and challenging benchmarks show that our VRAG-RL outperforms existing methods by 20\% (Qwen2. 5-VL-7B) and 30\% (Qwen2. 5-VL-3B), demonstrating the effectiveness of our approach. The code is available at https: //github. com/Alibaba-NLP/VRAG.

YNIMG Journal 2024 Journal Article

A fast dynamic causal modeling regression method for fMRI

  • Haifeng Wu
  • Xinhang Hu
  • Yu Zeng

Dynamic Causal Modeling (DCM) is a crucial tool for studying brain effective connectivity, offering valuable insights into brain network dynamics through functional magnetic resonance imaging (fMRI) and electrophysiology (EEG and MEG). However, its high computational complexity limits its applicability in large-scale network analysis. To address this issue, we propose a regression algorithm that integrates the Generalized Linear Model (GLM) with Sparse DCM, termed GSD. This algorithm enhances computational performance through three key optimizations: (1) utilizing the symmetry of the Fourier transform to convert complex frequency domain calculations into real number operations, thereby reducing computational complexity; (2) applying GLM and filtering techniques to minimize the effects of noise and confounds, enhancing parameter estimation accuracy; and (3) defining a new cost function to optimize variational inference and filter parameters, further improving parameter estimation accuracy. We validated the GSD algorithm using three public fMRI datasets: simulated Smith small-world network data, attention and motion measured data, and face recognition repetition effect measured data. The experimental results demonstrate that the GSD algorithm reduces computation time by over 50 % while maintaining parameter estimation performance comparable to traditional methods. These findings offer a new perspective on balancing model interpretability and computational efficiency, potentially broadening the application of DCM across various fields.

NeurIPS Conference 2024 Conference Paper

HairDiffusion: Vivid Multi-Colored Hair Editing via Latent Diffusion

  • Yu Zeng
  • Yang Zhang
  • Jiachen Liu
  • Linlin Shen
  • Kaijun Deng
  • Weizhao He
  • Jinbao Wang

Hair editing is a critical image synthesis task that aims to edit hair color and hairstyle using text descriptions or reference images, while preserving irrelevant attributes (e. g. , identity, background, cloth). Many existing methods are based on StyleGAN to address this task. However, due to the limited spatial distribution of StyleGAN, it struggles with multiple hair color editing and facial preservation. Considering the advancements in diffusion models, we utilize Latent Diffusion Models (LDMs) for hairstyle editing. Our approach introduces Multi-stage Hairstyle Blend (MHB), effectively separating control of hair color and hairstyle in diffusion latent space. Additionally, we train a warping module to align the hair color with the target region. To further enhance multi-color hairstyle editing, we fine-tuned a CLIP model using a multi-color hairstyle dataset. Our method not only tackles the complexity of multi-color hairstyles but also addresses the challenge of preserving original colors during diffusion editing. Extensive experiments showcase the superiority of our method in editing multi-color hairstyles while preserving facial attributes given textual descriptions and reference images.

AAAI Conference 2024 Conference Paper

Large Occluded Human Image Completion via Image-Prior Cooperating

  • Hengrun Zhao
  • Yu Zeng
  • Huchuan Lu
  • Lijun Wang

The completion of large occluded human body images poses a unique challenge for general image completion methods. The complex shape variations of human bodies make it difficult to establish a consistent understanding of their structures. Furthermore, as human vision is highly sensitive to human bodies, even slight artifacts can significantly compromise image fidelity. To address these challenges, we propose a large occluded human image completion (LOHC) model based on a novel image-prior cooperative completion strategy. Our model leverages human segmentation maps as a prior, and completes the image and prior simultaneously. Compared to the widely adopted prior-then-image completion strategy for object completion, this cooperative completion process fosters more effective interaction between the prior and image information. Our model consists of two stages. The first stage is a transformer-based auto-regressive network that predicts the overall structure of the missing area by generating a coarse completed image at a lower resolution. The second stage is a convolutional network that refines the coarse images. As the coarse result may not always be accurate, we propose a Dynamic Fusion Module (DFM) to selectively fuses the useful features from the coarse image with the original input at spatial and channel levels. Through extensive experiments, we demonstrate our method’s superior performance compared to state-of-the-art methods.

AAAI Conference 2023 Conference Paper

JR2Net: Joint Monocular 3D Face Reconstruction and Reenactment

  • Jiaxiang Shang
  • Yu Zeng
  • Xin Qiao
  • Xin Wang
  • Runze Zhang
  • Guangyuan Sun
  • Vishal Patel
  • Hongbo Fu

Face reenactment and reconstruction benefit various applications in self-media, VR, etc. Recent face reenactment methods use 2D facial landmarks to implicitly retarget facial expressions and poses from driving videos to source images, while they suffer from pose and expression preservation issues for cross-identity scenarios, i.e., when the source and the driving subjects are different. Current self-supervised face reconstruction methods also demonstrate impressive results. However, these methods do not handle large expressions well, since their training data lacks samples of large expressions, and 2D facial attributes are inaccurate on such samples. To mitigate the above problems, we propose to explore the inner connection between the two tasks, i.e., using face reconstruction to provide sufficient 3D information for reenactment, and synthesizing videos paired with captured face model parameters through face reenactment to enhance the expression module of face reconstruction. In particular, we propose a novel cascade framework named JR2Net for Joint Face Reconstruction and Reenactment, which begins with the training of a coarse reconstruction network, followed by a 3D-aware face reenactment network based on the coarse reconstruction results. In the end, we train an expression tracking network based on our synthesized videos composed by image-face model parameter pairs. Such an expression tracking network can further enhance the coarse face reconstruction. Extensive experiments show that our JR2Net outperforms the state-of-the-art methods on several face reconstruction and reenactment benchmarks.

JBHI Journal 2020 Journal Article

Multi-Task Joint Learning Model for Segmenting and Classifying Tongue Images Using a Deep Neural Network

  • Qiang Xu
  • Yu Zeng
  • Wenjun Tang
  • Wei Peng
  • Tingwei Xia
  • Zongrun Li
  • Fei Teng
  • Weihong Li

Automatic tongue image segmentation and tongue image classification are two crucial tongue characterization tasks in traditional Chinese medicine (TCM). Due to the complexity of tongue segmentation and fine-grained traits of tongue image classification, both tasks are challenging. Fortunately, from the perspective of computer vision, these two tasks are highly interrelated, making them compatible with the idea of Multi-Task Joint learning (MTL). By sharing the underlying parameters and adding two different task loss functions, an MTL method for segmenting and classifying tongue images is proposed in this paper. Moreover, two state-of-the-art deep neural network variants (UNET and Discriminative Filter Learning (DFL)) are fused into the MTL to perform these two tasks. To the best of our knowledge, our method is the first attempt to manage both tasks simultaneously with MTL. We conducted extensive experiments with the proposed method. The experimental results show that our joint method outperforms the existing tongue characterization methods. Besides, visualizations and ablation studies are provided to aid in understanding our approach, which suggest that our method is highly consistent with human perception.

IJCAI Conference 2020 Conference Paper

RECPARSER: A Recursive Semantic Parsing Framework for Text-to-SQL Task

  • Yu Zeng
  • Yan Gao
  • Jiaqi Guo
  • Bei Chen
  • Qian Liu
  • Jian-Guang Lou
  • Fei Teng
  • Dongmei Zhang

Neural semantic parsers usually fail to parse long and complicated utterances into nested SQL queries, due to the large search space. In this paper, we propose a novel recursive semantic parsing framework called RECPARSER to generate the nested SQL query layer-by-layer. It decomposes the complicated nested SQL query generation problem into several progressive non-nested SQL query generation problems. Furthermore, we propose a novel Question Decomposer module to explicitly encourage RECPARSER to focus on different components of an utterance when predicting SQL queries of different layers. Experiments on the Spider dataset show that our approach is more effective compared to the previous works at predicting the nested SQL queries. In addition, we achieve an overall accuracy that is comparable with state-of-the-art approaches.

JBHI Journal 2020 Journal Article

State Estimation of Hemodynamic Model for fMRI Under Confounds: SSM Method

  • Haifeng Wu
  • Mingzhi Lu
  • Yu Zeng

Through hemodynamic models, the change of neuronal state can be estimated from functional magnetic resonance imaging (fMRI) signals. Usually, there are confounds in the fMRI signal, which will degrade the performance of the estimation for the neuronal state change. For the reason, this paper introduces a state-space model with confounds, from a conventional hemodynamic model. In this model, a successive state estimation method requires a state value vector, an error covariance, an innovation covariance, and a cross covariance to be re-derived. Thus, a confounds square-root cubature Kalman smoothing (CSCKS) algorithm is proposed in this paper. We use a Balloon-Windkessel model to generate simulation data and add confounds signals to evaluate the performance of the proposed algorithm. The experiment results show that when the signal-to-interference ratio is less than 21 dB, the CSCKS proposed in this paper reduced estimation error to 16%, whereas the traditional algorithm reduced it to only 73%.

AAAI Conference 2019 Conference Paper

Deep Embedding Features for Salient Object Detection

  • Yunzhi Zhuge
  • Yu Zeng
  • Huchuan Lu

Benefiting from the rapid development of Convolutional Neural Networks (CNNs), some salient object detection methods have achieved remarkable results by utilizing multi-level convolutional features. However, the saliency training datasets is of limited scale due to the high cost of pixel-level labeling, which leads to a limited generalization of the trained model on new scenarios during testing. Besides, some FCN-based methods directly integrate multi-level features, ignoring the fact that the noise in some features are harmful to saliency detection. In this paper, we propose a novel approach that transforms prior information into an embedding space to select attentive features and filter out outliers for salient object detection. Our network firstly generates a coarse prediction map through an encorder-decorder structure. Then a Feature Embedding Network (FEN) is trained to embed each pixel of the coarse map into a metric space, which incorporates much attentive features that highlight salient regions and suppress the response of non-salient regions. Further, the embedded features are refined through a deep-to-shallow Recursive Feature Integration Network (RFIN) to improve the details of prediction maps. Moreover, to alleviate the blurred boundaries, we propose a Guided Filter Refinement Network (GFRN) to jointly optimize the predicted results and the learnable guidance maps. Extensive experiments on five benchmark datasets demonstrate that our method outperforms state-of-the-art results. Our proposed method is end-to-end and achieves a realtime speed of 38 FPS.

v2026.09.13