Arrow Research search

Author name cluster

Feng Yu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

14 papers
1 author row

Possible papers

14

AAAI Conference 2026 Conference Paper

Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation

  • Xinyi Tong
  • Yiran Zhu
  • Jishang Chen
  • Chunru Zhan
  • Tianle Wang
  • Sirui Zhang
  • Nian Liu
  • Tiezheng Ge

Video-to-Music generation seeks to generate musically appropriate background music that enhances audiovisual immersion for videos. However, current approaches suffer from two critical limitations: 1) incomplete representation of video details, leading to weak alignment, and 2) inadequate temporal and rhythmic correspondence, particularly in achieving precise beat synchronization. To address the challenges, we propose Video Echoed in Music (VeM), a latent music diffusion that generates high-quality soundtracks with semantic, temporal, and rhythmic alignment for input videos. To capture video details comprehensively, VeM employs a hierarchical video parsing that acts as a music conductor, orchestrating multi-level information across modalities. Modality-specific encoders, coupled with a storyboard-guided cross-attention mechanism (SG-CAtt), integrate semantic cues while maintaining temporal coherence through position and duration encoding. For rhythmic precision, the frame-level transition-beat aligner and adapter (TB-As) dynamically synchronize visual scene transitions with music beats. We further contribute a novel video-music paired dataset sourced from e-commerce advertisements and video-sharing platforms, which imposes stricter transition-beat synchronization requirements. Meanwhile, we introduce novel metrics tailored to the task. Experimental results demonstrate superiority, particularly in semantic relevance and rhythmic precision.

AAAI Conference 2025 Conference Paper

GarFast: Realistic and Fast Garment Transfer with a Simplified Parser-Free Approach

  • Chenghu Du
  • Junyin Wang
  • Yi Rong
  • Feng Yu
  • Shengwu Xiong

A good garment try-on model should learn the transfer between different types of garments while satisfying: 1) high fidelity and 2) low inference speed. Existing methods address either of these two issues, limited processing speed or low generation quality. We directly use a lightweight encoder-decoder, ensuring faster speeds. To tackle the problem of lower image quality typically generated by lighter models, we present GarFast, a simplified, parser-free framework that optimizes the same lightweight network through a two-stage transformation of real data roles (from input to supervision), thereby greatly promoting model convergence. Specifically, first, we propose a correction strategy to prevent the difficulty of convergence caused by the lack of ground truth in the first stage. Second, we propose a fine-grained domain consistency to ensure that the results generated in the unsupervised first stage are highly realistic clothed human images. Finally, we propose a skin-variant refinement loss and a skinMix regularization to amplify texture differences and enhance the realism of skin-variant regions, thereby improving the quality of the generated skin. Extensive experiments thoroughly demonstrate that our method achieves high resolution, near real-time performance, and superior reconstruction quality compared to state-of-the-art approaches, with processing times of less than 0.03 seconds on an Nvidia A100.

AAAI Conference 2025 Conference Paper

Latent Diffusion-Enhanced Virtual Try-On via Optimized Pseudo-Label Generation

  • Chenghu Du
  • Junyin Wang
  • Feng Yu
  • Shengwu Xiong

Efficiently applying fully supervised learning to virtual try-on tasks is challenging due to the lack of paired ground truth in available training samples. Recent works have achieved virtual try-ons by employing self-supervised learning-based inpainting paradigms. However, this approach is heavily dependent on the constraints of inpainting masks. An incorrect mask can mislead the generated results, while overly large mask areas can lose essential original information, thereby hindering the synthesis of high-quality results. To address these problems, we propose a latent diffusion model-based virtual try-on network that achieves fully supervised learning using the concept of cycle consistency and knowledge distillation. Specifically, we divide our approach into pretext and downstream tasks. In the pretext task, we generate a pseudo-label (pseudo-person image) to form paired training samples, which enables the downstream task to achieve fully supervised learning. To prevent the unreliable pseudo-person image from introducing irresponsible prior knowledge, we propose a noise-covering strategy, which aims at fully optimizing the pseudo-label to eliminate the impact of the incorrect inpainting mask as much as possible. Additionally, we propose a skin refinement loss to further enhance the generation of details in the skin region. Extended experiments demonstrate that our proposed method is superior to state-of-the-art methods.

JBHI Journal 2025 Journal Article

Learning to Detect Sleep Micro-Events from Coarse Sleep Stage Annotations

  • Chenhao Wang
  • Yan Pei
  • Jing Hu
  • Chengyang Han
  • Jiahui Xu
  • Lisan Zhang
  • Feng Yu
  • Bo Jin

Sleep micro-events, such as sleep spindles and K-complexes, are closely associated with neurological cognitive functions. While artificial intelligence (AI)-assisted sleep micro-event detection provides automated annotation to reduce reliance on labor-intensive expert labeling, current supervised approaches require precisely annotated datasets that remain scarce in clinical practice. To overcome this data bottleneck, this paper introduces a Weakly Supervised Sleep Micro-Event Detector (WSSMED) that leverages readily available coarse sleep stage annotations. The proposed WSSMED features a dual-branch architecture, consisting of a wave prototype module and a cluster module, designed to capture the fine-grained sleep micro-event patterns experts rely on for sleep staging. This framework infers expert logic from coarse annotations while mitigating performance degradation caused by annotation inconsistencies arising from inter-rater variability. Experiments conducted on two public datasets and one clinical dataset demonstrate that WSSMED achieves state-of-the-art performance in detecting sleep spindles and K-complexes, as evaluated at both sample-level and event-level in terms of precision, recall and F1-score metrics. Furthermore, subject-level evaluation demonstrates that the density and duration of micro-events detected by WSSMED-key metrics linked to cognitive function and neurological status-align more closely with expert annotations than those of other reported methods. These results highlight the clinical potential of WSSMED for reliable sleep micro-event analysis.

IJCAI Conference 2025 Conference Paper

NotaGen: Advancing Musicality in Symbolic Music Generation with Large Language Model Training Paradigms

  • Yashan Wang
  • Shangda Wu
  • Jianhuai Hu
  • Xingjian Du
  • Yueqi Peng
  • Yongxin Huang
  • Shuai Fan
  • Xiaobing Li

We introduce NotaGen, a symbolic music generation model aiming to explore the potential of producing high-quality classical sheet music. Inspired by the success of Large Language Models (LLMs), NotaGen adopts pre-training, fine-tuning, and reinforcement learning paradigms (henceforth referred to as the LLM training paradigms). It is pre-trained on 1. 6M pieces of music in ABC notation, and then fine-tuned on approximately 9K high-quality classical compositions conditioned on "period-composer-instrumentation" prompts. For reinforcement learning, we propose the CLaMP-DPO method, which further enhances generation quality and controllability without requiring human annotations or predefined rewards. Our experiments demonstrate the efficacy of CLaMP-DPO in symbolic music generation models with different architectures and encoding schemes. Furthermore, subjective A/B tests show that NotaGen outperforms baseline models against human compositions, greatly advancing musical aesthetics in symbolic music generation.

JBHI Journal 2025 Journal Article

WaveSleepNet: An Interpretable Network for Expert-Like Sleep Staging

  • Yan Pei
  • Jiahui Xu
  • Feng Yu
  • Lisan Zhang
  • Wei Luo

Although deep learning algorithms have proven their efficiency in automatic sleep staging, their “black-box” nature has limited their clinical adoption. In this study, we propose WaveSleepNet, an interpretable neural network for sleep staging that reasons in a similar way to sleep clinical experts. In this network, we utilize the latent space representations generated during training to identify characteristic wave prototypes corresponding to different sleep stages. The feature representation of an input signal is segmented into patches within the latent space, each of which is compared against the learned wave prototypes. The proximity between these patches and the wave prototypes is quantified through scores, indicating the prototypes' presence and relative proportion within the signal. The scores serve as the decision-making criteria for final sleep staging. During training, an ensemble of loss functions is employed for the prototypes' diversity and robustness. Furthermore, the learned wave prototypes are visualized by analyzing occlusion sensitivity. The efficacy of WaveSleepNet is validated across three public datasets, achieving sleep staging performance that are on par with those of the state-of-the-art models. A detailed case study examining the decision-making process of WaveSleepNet demonstrates that it aligns closely with American Academy of Sleep Medicine (AASM) manual guidelines. Another case study systematically explained the misidentified reasons behind each sleep stage. WaveSleepNet's transparent process provides specialists with direct access to the physiological significance of the model's criteria, allowing for future validation, adoption and further enrichment by sleep clinical experts.

JBHI Journal 2024 Journal Article

DTP-Net: Learning to Reconstruct EEG Signals in Time-Frequency Domain by Multi-Scale Feature Reuse

  • Yan Pei
  • Jiahui Xu
  • Qianhao Chen
  • Chenhao Wang
  • Feng Yu
  • Lisan Zhang
  • Wei Luo

Electroencephalography (EEG) signals are prone to contamination by noise, such as ocular and muscle artifacts. Minimizing these artifacts is crucial for EEG-based downstream applications like disease diagnosis and brain-computer interface (BCI). This paper presents a new EEG denoising model, DTP-Net. It is a fully convolutional neural network comprising Densely-connected Temporal Pyramids (DTPs) placed between two learnable time-frequency transformations. In the time-frequency domain, DTPs facilitate efficient propagation of multi-scale features extracted from EEG signals of any length, leading to effective noise reduction. Comprehensive experiments on two public semi-simulated datasets demonstrate that the proposed DTP-Net consistently outperforms existing state-of-the-art methods on metrics including relative root mean square error (RRMSE) and signal-to-noise ratio improvement ( $\Delta$ SNR). Moreover, the proposed DTP-Net is applied to a BCI classification task, yielding an improvement of up to 5. 55% in accuracy. This confirms the potential of DTP-Net for applications in the fields of EEG-based neuroscience and neuro-engineering. An in-depth analysis further illustrates the representation learning behavior of each module in DTP-Net, demonstrating its robustness and reliability.

NeurIPS Conference 2024 Conference Paper

S-MolSearch: 3D Semi-supervised Contrastive Learning for Bioactive Molecule Search

  • Gengmo Zhou
  • Zhen Wang
  • Feng Yu
  • Guolin Ke
  • Zhewei Wei
  • Zhifeng Gao

Virtual Screening is an essential technique in the early phases of drug discovery, aimed at identifying promising drug candidates from vast molecular libraries. Recently, ligand-based virtual screening has garnered significant attention due to its efficacy in conducting extensive database screenings without relying on specific protein-binding site information. Obtaining binding affinity data for complexes is highly expensive, resulting in a limited amount of available data that covers a relatively small chemical space. Moreover, these datasets contain a significant amount of inconsistent noise. It is challenging to identify an inductive bias that consistently maintains the integrity of molecular activity during data augmentation. To tackle these challenges, we propose S-MolSearch, the first framework to our knowledge, that leverages molecular 3D information and affinity information in semi-supervised contrastive learning for ligand-based virtual screening. % S-MolSearch processes both labeled and unlabeled data, trains molecular structural encoders, and generates soft labels for unlabeled data, drawing on the principles of inverse optimal transport. Drawing on the principles of inverse optimal transport, S-MolSearch efficiently processes both labeled and unlabeled data, training molecular structural encoders while generating soft labels for the unlabeled data. This design allows S-MolSearch to adaptively utilize unlabeled data within the learning process. Empirically, S-MolSearch demonstrates superior performance on widely-used benchmarks LIT-PCBA and DUD-E. It surpasses both structure-based and ligand-based virtual screening methods for AUROC, BEDROC and EF.

NeurIPS Conference 2024 Conference Paper

SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery Detection

  • Yachao Liang
  • Min Yu
  • Gang Li
  • Jianguo Jiang
  • Boquan Li
  • Feng Yu
  • Ning Zhang
  • Xiang Meng

Detection of face forgery videos remains a formidable challenge in the field of digital forensics, especially the generalization to unseen datasets and common perturbations. In this paper, we tackle this issue by leveraging the synergy between audio and visual speech elements, embarking on a novel approach through audio-visual speech representation learning. Our work is motivated by the finding that audio signals, enriched with speech content, can provide precise information effectively reflecting facial movements. To this end, we first learn precise audio-visual speech representations on real videos via a self-supervised masked prediction task, which encodes both local and global semantic information simultaneously. Then, the derived model is directly transferred to the forgery detection task. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods in terms of cross-dataset generalization and robustness, without the participation of any fake video in model training.

TIST Journal 2023 Journal Article

Unsupervised Graph Representation Learning with Cluster-aware Self-training and Refining

  • Yanqiao Zhu
  • Yichen Xu
  • Feng Yu
  • Qiang Liu
  • Shu Wu

Unsupervised graph representation learning aims to learn low-dimensional node embeddings without supervision while preserving graph topological structures and node attributive features. Previous Graph Neural Networks (GNN) require a large number of labeled nodes, which may not be accessible in real-world applications. To this end, we present a novel unsupervised graph neural network model with Cluster-aware Self-training and Refining ( CLEAR ). Specifically, in the proposed CLEAR model, we perform clustering on the node embeddings and update the model parameters by predicting the cluster assignments. To avoid degenerate solutions of clustering, we formulate the graph clustering problem as an optimal transport problem and leverage a balanced clustering strategy. Moreover, we observe that graphs often contain inter-class edges, which mislead the GNN model to aggregate noisy information from neighborhood nodes. Therefore, we propose to refine the graph topology by strengthening intra-class edges and reducing node connections between different classes based on cluster labels, which better preserves cluster structures in the embedding space. We conduct comprehensive experiments on two benchmark tasks using real-world datasets. The results demonstrate the superior performance of the proposed model over baseline methods. Notably, our model gains over 7% improvements in terms of accuracy on node clustering over state-of-the-arts.

JMLR Journal 2022 Journal Article

ALMA: Alternating Minimization Algorithm for Clustering Mixture Multilayer Network

  • Xing Fan
  • Marianna Pensky
  • Feng Yu
  • Teng Zhang

The paper considers a Mixture Multilayer Stochastic Block Model (MMLSBM), where layers can be partitioned into groups of similar networks, and networks in each group are equipped with a distinct Stochastic Block Model. The goal is to partition the multilayer network into clusters of similar layers, and to identify communities in those layers. Jing et al. (2020) introduced the MMLSBM and developed a clustering methodology, TWIST, based on regularized tensor decomposition. The present paper proposes a different technique, an alternating minimization algorithm (ALMA), that aims at simultaneous recovery of the layer partition, together with estimation of the matrices of connection probabilities of the distinct layers. Compared to TWIST, ALMA achieves higher accuracy, both theoretically and numerically. [abs] [ pdf ][ bib ] &copy JMLR 2022. ( edit, beta )

TIST Journal 2021 Journal Article

Disentangled Item Representation for Recommender Systems

  • Zeyu Cui
  • Feng Yu
  • Shu Wu
  • Qiang Liu
  • Liang Wang

Item representations in recommendation systems are expected to reveal the properties of items. Collaborative recommender methods usually represent an item as one single latent vector. Nowadays the e-commercial platforms provide various kinds of attribute information for items (e.g., category, price, and style of clothing). Utilizing this attribute information for better item representations is popular in recent years. Some studies use the given attribute information as side information, which is concatenated with the item latent vector to augment representations. However, the mixed item representations fail to fully exploit the rich attribute information or provide explanation in recommender systems. To this end, we propose a fine-grained Disentangled Item Representation (DIR) for recommender systems in this article, where the items are represented as several separated attribute vectors instead of a single latent vector. In this way, the items are represented at the attribute level, which can provide fine-grained information of items in recommendation. We introduce a learning strategy, LearnDIR, which can allocate the corresponding attribute vectors to items. We show how DIR can be applied to two typical models, Matrix Factorization (MF) and Recurrent Neural Network (RNN). Experimental results on two real-world datasets show that the models developed under the framework of DIR are effective and efficient. Even using fewer parameters, the proposed model can outperform the state-of-the-art methods, especially in the cold-start situation. In addition, we make visualizations to show that our proposition can provide explanation for users in real-world applications.

TIST Journal 2018 Journal Article

Mining Significant Microblogs for Misinformation Identification

  • Qiang Liu
  • Feng Yu
  • Shu Wu
  • Liang Wang

With the rapid growth of social media, massive misinformation is also spreading widely on social media, e.g., Weibo and Twitter, and brings negative effects to human life. Today, automatic misinformation identification has drawn attention from academic and industrial communities. Whereas an event on social media usually consists of multiple microblogs, current methods are mainly constructed based on global statistical features. However, information on social media is full of noise, which should be alleviated. Moreover, most of the microblogs about an event have little contribution to the identification of misinformation, where useful information can be easily overwhelmed by useless information. Thus, it is important to mine significant microblogs for constructing a reliable misinformation identification method. In this article, we propose an attention-based approach for identification of misinformation (AIM). Based on the attention mechanism, AIM can select microblogs with the largest attention values for misinformation identification. The attention mechanism in AIM contains two parts: content attention and dynamic attention. Content attention is the calculated-based textual features of each microblog. Dynamic attention is related to the time interval between the posting time of a microblog and the beginning of the event. To evaluate AIM, we conduct a series of experiments on the Weibo and Twitter datasets, and the experimental results show that the proposed AIM model outperforms the state-of-the-art methods.

IJCAI Conference 2017 Conference Paper

A Convolutional Approach for Misinformation Identification

  • Feng Yu
  • Qiang Liu
  • Shu Wu
  • Liang Wang
  • Tieniu Tan

The fast expanding of social media fuels the spreading of misinformation which disrupts people's normal lives. It is urgent to achieve goals of misinformation identification and early detection in social media. In dynamic and complicated social media scenarios, some conventional methods mainly concentrate on feature engineering which fail to cover potential features in new scenarios and have difficulty in shaping elaborate high-level interactions among significant features. Moreover, a recent Recurrent Neural Network (RNN) based method suffers from deficiencies that it is not qualified for practical early detection of misinformation and poses a bias to the latest input. In this paper, we propose a novel method, Convolutional Approach for Misinformation Identification (CAMI) based on Convolutional Neural Network (CNN). CAMI can flexibly extract key features scattered among an input sequence and shape high-level interactions among significant features, which help effectively identify misinformation and achieve practical early detection. Experiment results on two large-scale datasets validate the effectiveness of CAMI model on both misinformation identification and early detection tasks.

v2026.09.13