Arrow Research search

Author name cluster

Xin Luo

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

21 papers
1 author row

Possible papers

21

EAAI Journal 2026 Journal Article

Generative artificial intelligence method for vibration data of rolling bearing with multi-domain joint loss and interpretable physical laws

  • Qiwu Zhao
  • Xiaoli Zhang
  • Xin Luo
  • Shuangxuan Liang
  • Erick Mbeka

The prediction and health management (PHM) of rolling bearings have evolved toward data-driven transformation under Industry 4. 0. Generative artificial intelligence, which generates content with logical consistency and coherence based on the multimodal large models, provides a novel data generation technology for intelligent PHM of rolling bearing that suffers from insufficient samples of incomplete measurement. However, the generated data lacks physical interpretability due to the black-box nature. The rolling bearing fault or degradation features represented by the generated data can only be recognized by intelligent models and cannot be explained manually. To address this issue, a generative artificial intelligence method for vibration data of rolling bearing with multi-domain joint loss and interpretable physical laws is proposed. Small samples measured from the fault stage of the rolling bearing are input into the multilayer perceptron (MLP) to fit the parameter of stiffness and damping ratio, the simulated data can be obtained by solving the dynamic model of rolling bearing with fitted parameters, and the physical laws of the rolling bearing working in the fault stage are described by the impulse response and fault frequency features of the simulated data. In the degradation stage of rolling bearings, the parameters' sequence of stiffness and damping ratio are fitted, respectively, by inputting a small sample of measured degradation data into MLP, and the future trends of parameter sequences are predicted by the temporal convolutional network (TCN) with single-step iterative prediction method. The simulated degradation data is produced by solving the dynamic model with fitted parameter sequences, which follow the physical laws of the rolling bearing working in the degradation stage. Experimental validation, ablation, and comparison experiments are carried out based on the Case Western Reserve University (CWRU) and XJTU-SY bearing datasets. The results show that the simulated data exhibit similar distribution characteristics to the measured data in time and frequency domains, which corresponds to physical laws and provides clear interpretability.

TMLR Journal 2026 Journal Article

Hierarchical Filtering and Refinement Classification for Few-Shot Class-Incremental Learning

  • Li-Jun Zhao
  • Zhen-Duo Chen
  • Xin Luo
  • Xin-Shun Xu

Few-shot class-incremental learning (FSCIL) aims at recognizing novel classes continually with limited novel class samples. A mainstream baseline for FSCIL is first to train the whole model in the base session, then freeze the feature extractor in the incremental sessions. Despite achieving high overall accuracy, most methods exhibit notably low accuracy on incremental classes. While some recent methods have recognized this issue, their strategies remain constrained by a unified classification objective across all samples, making it difficult to simultaneously satisfy the performance requirements of both base and incremental classes. In this paper, considering that base and incremental classes play different yet both critical roles in FSCIL, we approach FSCIL from a more structured perspective by decomposing the overall classification objective into three sub-objectives. Building on this insight, we propose a novel classification framework called Hierarchical Filtering and Refinement Classification (HFRC) to hierarchically decompose and address the classification task. Extensive experiments demonstrate that our method effectively balances the classification accuracy between base and incremental classes, and achieves superior performance compared to state-of-the-art methods.

JBHI Journal 2026 Journal Article

Knowledge-Driven Multiple Instance Learning With Hierarchical Cluster-Incorporated Aware Filtering for Larynx Pathological Grading

  • Chentao Li
  • Pan Huang
  • Jing Qin
  • Xin Luo

Pathological grading of laryngeal squamous cell carcinoma (LSCC) based on whole-slide image (WSI) is crucial for the diagnosis, treatment and prognosis. According to pathologists’ knowledge, tumor regions are highly associated with grading. However, existing multiple instance learning (MIL) methods tend to overrepresent weakly relevant non-tumor regions and irrelevant background, leading to poor grading performance and interpretability. Motivated by the above problems, we propose an end-to-end knowledge-driven MIL network with hierarchical cluster-incorporated aware filtering, i. e. HCF-MIL. Firstly, we develop the tumor-guiding cluster filtering for feature representation, which awarely filters out irrelevant instance-level information and adaptively assigns learnable weights to tumor and non-tumor instances. Secondly, conventional mean-based and max-based aggregation primarily capture the overall patterns, neglecting the contributions of the most representative individual instances. Therefore, we propose a novel enhanced filtering aggregation learning strategy to strengthen hierarchical tumor-related feature representation. Through end-to-end optimization, HCF-MIL reduces model’s entropy value and facilitates better alignment between decision-making process and diagnostic behaviors of pathologists. Experiments on larynx and multicentre datasets show that HCF-MIL significantly improves both pathological grading performance and interpretability, providing a strong foundation for reliable clinical deployment.

AAAI Conference 2026 Conference Paper

Multi-granularity Interactive Attention Framework for Residual Hierarchical Pronunciation Assessment

  • Hong Han
  • Hao-Chen Pei
  • Zhao-Zheng Nie
  • Xin Luo
  • Xin-Shun Xu

Automatic pronunciation assessment plays a crucial role in computer-assisted pronunciation training systems. Due to the ability to perform multiple pronunciation tasks simultaneously, multi-aspect multi-granularity pronunciation assessment methods are gradually receiving more attention and achieving better performance than single-level modeling tasks. However, existing methods only consider unidirectional dependencies between adjacent granularity levels, lacking bidirectional interaction among phoneme, word, and utterance levels and thus insufficiently capturing the acoustic structural correlations. To address this issue, we propose a novel residual hierarchical interactive method, HIA for short, that enables bidirectional modeling across granularities. As the core of HIA, the Interactive Attention Module leverages an attention mechanism to achieve dynamic bidirectional interaction, effectively capturing linguistic features at each granularity while integrating correlations between different granularity levels. We also propose a residual hierarchical structure to alleviate the feature forgetting problem when modeling acoustic hierarchies. In addition, we use 1-D convolutional layers to enhance the extraction of local contextual cues at each granularity. Extensive experiments on the speechocean762 dataset show that our model is comprehensively ahead of the existing state-of-the-art methods.

EAAI Journal 2025 Journal Article

A novel lightweight model integrating convolutional neural network and self-attention mechanism for corn seeds quality image recognition

  • Shuai Zhang
  • Shuqi Ma
  • Xin Luo
  • Hancheng Chai
  • Jie Zhu

In the field of agriculture, accurate identification of seeds quality based on artificial intelligence algorithms is revolutionary for enhancing crop yields and qualities. The adoption of convolutional neural networks (CNNs) with scalability and self-attention mechanisms with robustness greatly improves the accuracy of seeds quality recognition. However, CNNs rely mainly on local receptive fields, making it challenging to capture long-range spatial dependencies, while self-attention mechanisms are computationally intensive and consume significant computational resources to achieve effective training. To address these limitations, a lightweight, efficient, and high-accuracy model architecture (denoted as LightMCS) is proposed, which consists of an efficient self-attention mechanism (ESA), partial convolution (PConv), and context broadcasting (CB). The LightMCS gradually extracts and refines features hierarchically. Notably, the PConv and ESA efficiently learn local-global representations, and the CB improves the model's capacity and generalization ability. Additionally, training acceleration techniques are employed to further enhance accuracy and reduce training time. The comparable results show that the proposed LightMCS demonstrates a higher classification accuracy (84. 12 %) in the quality recognition task of corn seeds. More importantly, LightMCS has only 3. 66 million parameters and a total training time of 243 min. The LightMCS achieves a balance between lightness and efficiency while improving classification accuracy. Thus, the proposed LightMCS provides an effective solution for the accurate identification of corn seeds quality, offering a promising tool in agricultural fields where accurate sorting of defective seeds scenarios is required.

NeurIPS Conference 2025 Conference Paper

Evolving and Regularizing Meta-Environment Learner for Fine-Grained Few-Shot Class-Incremental Learning

  • Li-Jun Zhao
  • Zhen-Duo Chen
  • Yongxin Wang
  • Xin Luo
  • Xin-Shun Xu

Recently proposed Fine-Grained Few-Shot Class-Incremental Learning (FG-FSCIL) offers a practical and efficient solution for enabling models to incrementally learn new fine-grained categories under limited data conditions. However, existing methods still settle for the fine-grained feature extraction capabilities learned from the base classes. Unlike conventional datasets, fine-grained categories exhibit subtle inter-class variations, naturally fostering latent synergy among sub-categories. Meanwhile, the incremental learning framework offers an opportunity to progressively strengthen this synergy by incorporating new sub-category data over time. Motivated by this, we theoretically formulate the FSCIL problem and derive a generalization error bound within a shared fine-grained meta-category environment. Guided by our theoretical insights, we design a novel Meta-Environment Learner (MEL) for FG-FSCIL, which evolves fine-grained feature extraction to enhance meta-environment understanding and simultaneously regularizes hypothesis space complexity. Extensive experiments demonstrate that our method consistently and significantly outperforms existing approaches.

AAAI Conference 2025 Conference Paper

Exploiting Diffusion Prior for Real-World Image Dehazing with Unpaired Training

  • Yunwei Lan
  • Zhigao Cui
  • Chang Liu
  • Jialun Peng
  • Nian Wang
  • Xin Luo
  • Dong Liu

Unpaired training has been verified as one of the most effective paradigms for real scene dehazing by learning from unpaired real-world hazy and clear images. Although numerous studies have been proposed, current methods demonstrate limited generalization for various real scenes due to limited feature representation and insufficient use of real-world prior. Inspired by the strong generative capabilities of diffusion models in producing both hazy and clear images, we exploit diffusion prior for real-world image dehazing, and propose an unpaired framework named Diff-Dehazer. Specifically, we leverage diffusion prior as bijective mapping learners within the CycleGAN, a classic unpaired learning framework. Considering that physical priors contain pivotal statistics information of real-world data, we further excavate real-world knowledge by integrating physical priors into our framework. Furthermore, we introduce a new perspective for adequately leveraging the representation ability of diffusion models by removing degradation in image and text modalities, so as to improve the dehazing effect. Extensive experiments on multiple real-world datasets demonstrate the superior performance of our method.

AAAI Conference 2025 Conference Paper

Few-Shot Fine-Grained Image Classification with Progressively Feature Refinement and Continuous Relationship Modeling

  • Zhen-Xiang Ma
  • Zhen-Duo Chen
  • Tai Zheng
  • Xin Luo
  • Zixia Jia
  • Xin-Shun Xu

Recently, a number of effective methods have been proposed to tackle the challenging task of Few-Shot Fine-Grained Image Classification (FS-FGIC). However, how to fully leverage the backbone network to discover and extract detailed features to generate more discriminative class prototypes, as well as how to accurately model the similarity relationship between query samples and the class prototypes, are still issues to be further considered. Therefore, we propose a novel progreSsively featUre refInement and conTinuous rElationship moDeling method, SUITED for short, to address these two issues existing in the State-of-the-Art FS-FGIC methods. Specifically, we design the Progressive Feature Refinement Module (PFRM) to fully exploit the backbone network's progressive feature extraction capabilities, forming multi-scale feature representations to further enhance discriminative features. Then, the Continuous Relationship Modeling Module (CRMM) is proposed to capture the dependencies between query samples and the corresponding class prototypes, achieving precise optimization of the distances among corresponding sample points in the feature space. We conducted extensive experiments on five fine-grained benchmark datasets, and the experimental results demonstrate that the proposed method is comprehensively ahead of the existing State-of-the-Art methods.

JBHI Journal 2025 Journal Article

Graph-Based Prediction of miRNA-Drug Associations with Multisource Information and Metapath Enhancement Matrices

  • Ming-Yang Wu
  • Peng-Wei Hu
  • Zhu-Hong You
  • Jun Zhang
  • Lun Hu
  • Xin Luo

Recent studies have demonstrated that miRNA expression dysregulation is closely related to the occurrence of various diseases; thus, miRNA-based drug development strategies have received increasing research interest. Most existing computational methods focus on the attribute information of individual nodes and are limited to the direct associations between nodes, thereby ignoring the complex associations inherent in the network. This limitation may lead to the loss of key potential information, which impacts the prediction accuracy. To address these issues, we propose a multisource information fusion and metapath enhancement matrix based graph autoencoder (MSMP-GAE) to predict the potential associations between miRNAs and drugs. The proposed MSMP-GAE model comprises a metapath instance extraction module, a metapath feature-enhanced encoder module, a weighted feature fusion module, and a graph autoencoder. First, we construct an miRNA–drug heterogeneous network using experimentally validated miRNA–drug interactions and integrate various miRNA and drug features into an initial feature matrix to comprehensively represent their intrinsic property information. Then, we extract metapath instances from the interaction network, generate multiple metapath enhancement matrices, and fuse them with the initial feature matrix to generate high-quality node feature embeddings. Finally, we employ the graph autoencoder for fivefold cross-validation on a public dataset and test it on an independent test set. Experimental results demonstrate that the proposed MSMP-GAE model obtained an area under the curve (AUC) and AUPR values of 98. 61% and 98. 23%, respectively, which is considerably better than the several state-of-the-art methods. This highlights the importance of the higher-order complex associations between nodes in the miRNA–drug association (MDA) prediction task and provides a new method and approach to advance MDA prediction.

IJCAI Conference 2025 Conference Paper

Learning Real Facial Concepts for Independent Deepfake Detection

  • Ming-Hui Liu
  • Harry Cheng
  • Tianyi Wang
  • Xin Luo
  • Xin-Shun Xu

Deepfake detection models often struggle with generalization to unseen datasets, manifesting as misclassifying real instances as fake in target domains. This is primarily due to an overreliance on forgery artifacts and a limited understanding of real faces. To address this challenge, we propose a novel approach RealID to enhance generalization by learning a comprehensive concept of real faces while assessing the probabilities of belonging to the real and fake classes independently. RealID comprises two key modules: the Real Concept Capture Module (RealC^2) and the Independent Dual-Decision Classifier (IDC). With the assistance of a Multi-Real Memory, RealC^2 maintains various prototypes for real faces, allowing the model to capture a comprehensive concept of real class. Meanwhile, IDC redefines the classification strategy by making independent decisions based on the concept of the real class and the presence of forgery artifacts. Through the combined effect of the above modules, the influence of forgery-irrelevant patterns is alleviated, and extensive experiments on five widely used datasets demonstrate that RealID significantly outperforms existing state-of-the-art methods, achieving a 1. 74% improvement in average accuracy.

YNIMG Journal 2025 Journal Article

Neural connectivity and balance control in aging: Insights from directed cortical networks during sensory conflict

  • Guozheng Wang
  • Yi Yang
  • Xiaoxia Liu
  • Anke Hua
  • Xin Luo
  • Yiming Cai
  • Yanhua Song
  • Jian Wang

Balance control is crucial for stability during daily activities, relying on the integration of sensory inputs from the visual, vestibular, and somatosensory systems. Aging impairs the efficiency of these systems, leading to an increased risk of falls; however, the neural mechanisms underlying this decline, particularly under sensory conflict, are not fully understood. This study investigated the effects of aging on neural connectivity and sensory integration during balance tasks. Ninety-six participants (47 older adults and 49 young adults) were subjected to balance perturbation tasks under sensory-congruent and sensory-conflict conditions using a virtual reality headset and rotating platform. Behavioral measures, including postural sway and perceptual accuracy, were recorded. Electroencephalography (EEG) data were analyzed using generalized partial directed coherence (GPDC) to assess the directed functional connectivity and network efficiency. Older adults exhibited significantly greater postural sway, reduced perceptual accuracy, and a diminished ability to detect sensory conflicts than young adults, particularly under conflict conditions. As demonstrated by connectivity analysis, young adults showed adaptive shifts in connectivity from the visual to somatosensory regions during sensory conflict. In contrast, older adults demonstrated a less adaptable mode of connectivity. At the same time, global efficiency and clustering coefficients of young adults were higher, suggesting more effective and modular brain networks. Correlation analyses in older adults revealed that higher visual cortex efficiency was linked to lower postural sway specifically during sensory conflict, whereas higher motor cortex efficiency was associated with greater sway only under sensory-congruent conditions. In short, neural adaptability is vital in sensory integration and balance control. Due to decreased neural flexibility and network efficiency in older adults, their sensory reweighting was undermined and instability increased during the sensory conflict. These findings establish a foundation for development of targeted interventions to strengthen balance and lower the risks of falls in older adults.

IJCAI Conference 2025 Conference Paper

SDDiff: Boosting Radar Perception via Spatial-Doppler Diffusion

  • Shengpeng Wang
  • Xin Luo
  • Yulong Xie
  • Wei Wang

Point cloud extraction (PCE) and ego velocity estimation (EVE) are key capabilities gaining attention in 3D radar perception. However, existing work typically treats these two tasks independently, which may neglect the interplay between radar's spatial and Doppler domain features, potentially introducing additional bias. In this paper, we observe an underlying correlation between 3D points and ego velocity, which offers reciprocal benefits for PCE and EVE. To fully unlock such inspiring potential, we take the first step to design a Spatial-Doppler Diffusion (SDDiff) model for simultaneously dense PCE and accurate EVE. To seamlessly tailor it to radar perception, SDDiff improves the conventional latent diffusion process in three major aspects. First, we introduce a representation that embodies both spatial occupancy and Doppler features. Second, we design a directional diffusion with radar priors to streamline the sampling. Third, we propose Iterative Doppler Refinement to enhance the model’s adaptability to density variations and ghosting effects. Extensive evaluations show that SDDiff significantly outperforms state-of-the-art baselines by achieving 59% higher in EVE accuracy, 4X greater in valid generation density while boosting PCE effectiveness and reliability. The code and dataset will be available on https: //github. com/StellarEsti/SDDiff.

AAAI Conference 2024 Conference Paper

Cross-Layer and Cross-Sample Feature Optimization Network for Few-Shot Fine-Grained Image Classification

  • Zhen-Xiang Ma
  • Zhen-Duo Chen
  • Li-Jun Zhao
  • Zi-Chao Zhang
  • Xin Luo
  • Xin-Shun Xu

Recently, a number of Few-Shot Fine-Grained Image Classification (FS-FGIC) methods have been proposed, but they primarily focus on better fine-grained feature extraction while overlooking two important issues. The first one is how to extract discriminative features for Fine-Grained Image Classification tasks while reducing trivial and non-generalizable sample level noise introduced in this procedure, to overcome the over-fitting problem under the setting of Few-Shot Learning. The second one is how to achieve satisfying feature matching between limited support and query samples with variable spatial positions and angles. To address these issues, we propose a novel Cross-layer and Cross-sample feature optimization Network for FS-FGIC, C2-Net for short. The proposed method consists of two main modules: Cross-Layer Feature Refinement (CLFR) module and Cross-Sample Feature Adjustment (CSFA) module. The CLFR module further refines the extracted features while integrating outputs from multiple layers to suppress sample-level feature noise interference. Additionally, the CSFA module addresses the feature mismatch between query and support samples through both channel activation and position matching operations. Extensive experiments have been conducted on five fine-grained benchmark datasets, and the results show that the C2-Net outperforms other state-of-the-art methods by a significant margin in most cases. Our code is available at: https://github.com/zenith0923/C2-Net.

AAAI Conference 2024 Conference Paper

MKG-FENN: A Multimodal Knowledge Graph Fused End-to-End Neural Network for Accurate Drug–Drug Interaction Prediction

  • Di Wu
  • Wu Sun
  • Yi He
  • Zhong Chen
  • Xin Luo

Taking incompatible multiple drugs together may cause adverse interactions and side effects on the body. Accurate prediction of drug-drug interaction (DDI) events is essential for avoiding this issue. Recently, various artificial intelligence-based approaches have been proposed for predicting DDI events. However, DDI events are associated with complex relationships and mechanisms among drugs, targets, enzymes, transporters, molecular structures, etc. Existing approaches either partially or loosely consider these relationships and mechanisms by a non-end-to-end learning framework, resulting in sub-optimal feature extractions and fusions for prediction. Different from them, this paper proposes a Multimodal Knowledge Graph Fused End-to-end Neural Network (MKGFENN) that consists of two main parts: multimodal knowledge graph (MKG) and fused end-to-end neural network (FENN). First, MKG is constructed by comprehensively exploiting DDI events-associated relationships and mechanisms from four knowledge graphs of drugs-chemical entities, drug-substructures, drugs-drugs, and molecular structures. Correspondingly, a four channels graph neural network is designed to extract high-order and semantic features from MKG. Second, FENN designs a multi-layer perceptron to fuse the extracted features by end-to-end learning. With such designs, the feature extractions and fusions of DDI events are guaranteed to be comprehensive and optimal for prediction. Through extensive experiments on real drug datasets, we demonstrate that MKG-FENN exhibits high accuracy and significantly outperforms state-of-the-art models in predicting DDI events. The source code and supplementary file of this article are available on: https://github.com/wudi1989/MKG-FENN.

AIJ Journal 2024 Journal Article

Polarized message-passing in graph neural networks

  • Tiantian He
  • Yang Liu
  • Yew-Soon Ong
  • Xiaohu Wu
  • Xin Luo

In this paper, we present Polarized message-passing (PMP), a novel paradigm to revolutionize the design of message-passing graph neural networks (GNNs). In contrast to existing methods, PMP captures the power of node-node similarity and dissimilarity to acquire dual sources of messages from neighbors. The messages are then coalesced to enable GNNs to learn expressive representations from sparse but strongly correlated neighbors. Three novel GNNs based on the PMP paradigm, namely PMP graph convolutional network (PMP-GCN), PMP graph attention network (PMP-GAT), and PMP graph PageRank network (PMP-GPN) are proposed to perform various downstream tasks. Theoretical analysis is also conducted to verify the high expressiveness of the proposed PMP-based GNNs. In addition, an empirical study of five learning tasks based on 12 real-world datasets is conducted to validate the performances of PMP-GCN, PMP-GAT, and PMP-GPN. The proposed PMP-GCN, PMP-GAT, and PMP-GPN outperform numerous strong message-passing GNNs across all five learning tasks, demonstrating the effectiveness of the proposed PMP paradigm.

EAAI Journal 2024 Journal Article

Scheduling analysis of automotive glass manufacturing systems subject to sequence-independent setup time, no-idle machines, and permissive maximum total tardiness constraint

  • YunFang He
  • Yan Qiao
  • Naiqi Wu
  • Jiewu Leng
  • Xin Luo

With the increasing demand for automotive glass, improving the efficiency of automotive glass manufacturing systems can make full use of production resources and reduce the waste of natural and social resources. Therefore, this work aims to provide an efficient method for a real-world two-stage hybrid flow shop scheduling problem with small batches in an automotive glass manufacturing system. For the investigated problem, there is a significant setup time at the first stage, the second stage is served by machines that should not be interrupted, and each batch has a due date. Such constraints make this scheduling problem challenging. To solve this problem, a mixed integer linear programming model is established. Also, two properties and three theorems are given to enhance the problem-solving process. Subsequently, an efficient genetic algorithm is carefully designed to solve large-scale problems by considering the properties of the system. Meanwhile, an improvement scheme is proposed to decrease the running time of the algorithm, and experimental results show that this scheme can reduce the running time by 280 s at most from the average results of different scale problems. Finally, extensive experiments are carried out and a real-world case is solved to demonstrate the efficiency and effectiveness of the designed genetic algorithm. Also, the Taguchi method is adopted to tune the parameters for the designed algorithm.

TIST Journal 2023 Journal Article

Diagnose Like Doctors: Weakly Supervised Fine-Grained Classification of Breast Cancer

  • Jieru Tian
  • Yongxin Wang
  • Zhenduo Chen
  • Xin Luo
  • Xinshun Xu

Breast cancer is the most common type of cancers in women. Therefore, how to accurately and timely diagnose it becomes very important. Some computer-aided diagnosis models based on pathological images have been proposed for this task. However, there are still some issues that need to be further addressed. For example, most deep learning based models suffer from a lack of interpretability. In addition, some of them cannot fully exploit the information in medical data, e.g., hierarchical label structure and scattered distribution of target objects. To address these issues, we propose a weakly supervised fine-grained medical image classification method for breast cancer diagnosis, i.e., DLD-Net for short. It simulates the diagnostic procedures of pathologists by multiple attention-guided cropping and dropping operations, making it have good clinical interpretability. Moreover, it cannot only exploit the global information of a whole image, but also further mine the critical local information by generating and selecting critical regions from the image. In light of this, those subtle discriminating information hidden in scattered regions can be exploited. In addition, we also design a novel hierarchical cross-entropy loss to utilize the hierarchical label information in medical images, making the classification results more discriminative. Furthermore, DLD-Net is a weakly supervised network, which can be trained end-to-end without any additional region annotations. Extensive experimental results on three benchmark datasets demonstrate that DLD-Net is able to achieve good results and outperforms some state-of-the-art methods.

AAAI Conference 2022 Conference Paper

Online Enhanced Semantic Hashing: Towards Effective and Efficient Retrieval for Streaming Multi-Modal Data

  • Xiao-ming Wu
  • Xin Luo
  • Yu-Wei Zhan
  • Chen-Lu Ding
  • Zhen-Duo Chen
  • Xin-Shun Xu

With the vigorous development of multimedia equipments and applications, efficient retrieval of large-scale multi-modal data has become a trendy research topic. Thereinto, hashing has become a prevalent choice due to its retrieval efficiency and low storage cost. Although multi-modal hashing has drawn lots of attention in recent years, there still remain some problems. The first point is that existing methods are mainly designed in batch mode and not able to efficiently handle streaming multi-modal data. The second point is that all existing online multi-modal hashing methods fail to effectively handle unseen new classes which come continuously with streaming data chunks. In this paper, we propose a new model, termed Online enhAnced SemantIc haShing (OASIS). We design novel semantic-enhanced representation for data, which could help handle the new coming classes, and thereby construct the enhanced semantic objective function. An efficient and effective discrete online optimization algorithm is further proposed for OASIS. Extensive experiments show that our method can exceed the state-of-the-art models. For good reproducibility and benefiting the community, our code and data are already publicly available.

IJCAI Conference 2018 Conference Paper

SDMCH: Supervised Discrete Manifold-Embedded Cross-Modal Hashing

  • Xin Luo
  • Xiao-Ya Yin
  • Liqiang Nie
  • Xuemeng Song
  • Yongxin Wang
  • Xin-Shun Xu

Cross-modal hashing methods have attracted considerable attention. Most pioneer approaches only preserve the neighborhood relationship by constructing the correlations among heterogeneous modalities. However, they neglect the fact that the high-dimensional data often exists on a low-dimensional manifold embedded in the ambient space and the relative proximity between the neighbors is also important. Although some methods leverage the manifold learning to generate the hash codes, most of them fail to explicitly explore the discriminative information in the class labels and discard the binary constraints during optimization, generating large quantization errors. To address these issues, in this paper, we present a novel cross-modal hashing method, named Supervised Discrete Manifold-Embedded Cross-Modal Hashing (SDMCH). It can not only exploit the non-linear manifold structure of data and construct the correlation among heterogeneous multiple modalities, but also fully utilize the semantic information. Moreover, the hash codes can be generated discretely by an iterative optimization algorithm, which can avoid the large quantization errors. Extensive experimental results on three benchmark datasets demonstrate that SDMCH outperforms ten state-of-the-art cross-modal hashing methods.

IJCAI Conference 2017 Conference Paper

Symmetric Non-negative Latent Factor Models for Undirected Large Networks

  • Xin Luo
  • Ming-Sheng Shang

Undirected, high dimensional and sparse networks are frequently encountered in industrial applications. They contain rich knowledge regarding various useful patterns. Non-negative latent factor (NLF) models have proven to be effective and efficient in acquiring useful knowledge from asymmetric networks. However, they cannot correctly describe the symmetry of an undirected network. For addressing this issue, this work analyzes the NLF extraction processes on asymmetric and symmetric matrices respectively, thereby innovatively achieving the symmetric and non-negative latent factor (SNLF) models for undirected, high dimensional and sparse networks. The proposed SNLF models are equipped with a) high efficiency, b) non-negativity, and c) symmetry. Experimental results on real networks show that they are able to a) represent the symmetry of the target network rigorously; b) maintain the non-negativity of resulting latent factors; and c) achieve high computational efficiency when performing data analysis tasks as missing data estimation.

EAAI Journal 2012 Journal Article

A parallel matrix factorization based recommender by alternating stochastic gradient decent

  • Xin Luo
  • Huijun Liu
  • Gaopeng Gou
  • Yunni Xia
  • Qingsheng Zhu

Collaborative Filtering (CF) can be achieved by Matrix Factorization (MF) with high prediction accuracy and scalability. Most of the current MF based recommenders, however, are serial, which prevent them sharing the efficiency brought by the rapid progress in parallel programming techniques. Aiming at parallelizing the CF recommender based on Regularized Matrix Factorization (RMF), we first carry out the theoretical analysis on the parameter updating process of RMF, whereby we can figure out that the main obstacle preventing the model from parallelism is the inter-dependence between item and user features. To remove the inter-dependence among parameters, we apply the Alternating Stochastic Gradient Solver (ASGD) solver to deal with the parameter training process. On this basis, we subsequently propose the parallel RMF (P-RMF) model, of which the training process can be parallelized through simultaneously training different user/item features. Experiments on two large, real datasets illustrate that our P-RMF model can provide a faster solution to CF problem when compared to the original RMF and another parallel MF based recommender.

v2026.09.13