Arrow Research search

Author name cluster

Xin Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

187 papers
2 author rows

Possible papers

187

AAAI Conference 2026 Conference Paper

A Causal Target for Learning to Defer Under Hidden Confounding

  • Yanmin Li
  • Lihua Liu
  • Xin Wang
  • Zhilong Mao
  • Jibing Wu
  • Weidong Bao

Learning decision policies from confounded observational data is a challenging task in causal inference, as unobserved confounders can lead to biased or suboptimal actions when relying solely on machine learning models. A synergistic approach is learning to defer, which decides when to act itself and when to defer to a human expert with access to unobserved information. However, constructing the learning target, which defines the probability of choosing each action or deferral, remains a core challenge. To address this, we propose causal-target-based learning to defer (CTLD) framework, where the causal target is constructed from sharp bounds on potential outcomes. Specifically, the degree of overlap between these bounds determines the probability of deferral, while their relative positions and widths define the probabilities over actions. CTLD aligns model predictions with this causal target to make probabilistic decisions over actions and deferral. We present comprehensive theoretical guarantees for the learned policy and demonstrate the effectiveness of CTLD on synthetic and semi-synthetic datasets.

EAAI Journal 2026 Journal Article

A conditional diffusion vision transformer model via data augmentation for few-shot fault diagnosis

  • Beijia Zhao
  • Dongsheng Yang
  • Jiayue Sun
  • Yanhong Luo
  • Zhong Luo
  • Xin Wang

The scarcity of labeled training data degrades the performance of accurate fault diagnosis models, highlighting the critical need for research in few-shot fault diagnosis (FSFD). Despite being a predominant FSFD solution, current data augmentation-based methods still suffer from distribution mismatch between generated and real data, as well as insufficient hierarchical diversity, particularly in fault severities. To overcome these limitations, the conditional diffusion vision transformer model (CDViT) is proposed for FSFD. CDViT leverages a dual-constrained denoising diffusion probabilistic model to accurately model the underlying distribution of real fault data. Subsequently, a fault severity attention module is designed to effectively extract fault severity features by capturing local–global hierarchical characteristics. Additionally, a fault refinement classifier is employed to better capture fault severity, improving diagnostic performance. The effectiveness of CDViT has been validated through multiple comparative experiments conducted on three datasets, demonstrating superior performance compared to mainstream methods in FSFD.

EAAI Journal 2026 Journal Article

A framework integrating data-driven and computational fluid dynamics simulation for continuous blast furnace monitoring

  • Xin Wang
  • Xiao-Yu Tang
  • Zheng Hao
  • Kunwei Lin
  • Chunjie Yang
  • Wenhai Wang

Due to the complex and nonlinear characteristics of blast furnace (BF) systems, conventional data/mechanism-driven modeling methods have been facing challenges in continuously on-field BF internal state monitoring. Mechanism-driven computational fluid dynamics (CFD) simulations, while interpretable, have high computational costs that prevent real-time application. Conversely, data-driven methods often ignore the coupling relationships between variables and fail to provide a comprehensive understanding of the BF's internal state, which hinders precise control. In response to the aforementioned issues, this paper proposes a novel state continuous monitoring framework that incorporates data-driven with offline pre-calculated CFD simulations. It decouples the intricate and nonlinear BF operation process into a set of interpretable sub-modes, modeling each sub-mode via numerical simulations, then reconstructs the real-time BF transient by properly selecting and fusing some sub-modes. In the offline stage, the particle swarm optimization (PSO) algorithm is employed to obtain the sub-modes that best represent BF operational states. The CFD simulation is then conducted for multiple physical field states corresponding to each sub-mode. For online monitoring, a multi-mode fusion strategy (MMFS) is designed to achieve the real-time BF transient modeling. Application in a BF in South China validates the effectiveness of the proposed framework.

JBHI Journal 2026 Journal Article

A Hybrid Deep Learning Approach for Epileptic Seizure Detection in EEG signals

  • Ijaz Ahmad
  • Xin Wang
  • Danish Javeed
  • Prabhat Kumar
  • Oluwarotimi Williams Samuel
  • Shixiong Chen

Early detection and proper treatment of epilepsy is essential and meaningful to those who suffer from this disease. The adoption of deep learning (DL) techniques for automated epileptic seizure detection using electroencephalography (EEG) signals has shown great potential in making the most appropriate and fast medical decisions. However, DL algorithms have high computational complexity and suffer low accuracy with imbalanced medical data in multi seizure-classification task. Motivated from the aforementioned challenges, we present a simple and effective hybrid DL approach for epileptic seizure detection in EEG signals. Specifically, first we use a K-means Synthetic minority oversampling technique (SMOTE) to balance the sampling data. Second, we integrate a 1D convolutional neural network (CNN) with a Bidirectional Long Short-Term Memory (BiLSTM) network based on Truncated Backpropagation Through Time (TBPTT) to efficiently extract spatial and temporal sequence information while reducing computational complexity. Finally, the proposed DL architecture uses softmax and sigmoid classifiers at the classification layer to perform multi and binary seizure-classification tasks. In addition, the 10-fold cross-validation technique is performed to show the significance of the proposed DL approach. Experimental results using the publicly available UCI epileptic seizure recognition data set shows better performance in terms of precision, sensitivity, specificity, and F1-score over some baseline DL algorithms and recent state-of-the-art techniques.

AAAI Conference 2026 Conference Paper

Binary Message Passing for Generalizable Semi-Supervised Graph Anomaly Detection

  • Jingyuan Zhang
  • Xin Wang
  • Lei Yu
  • Li Yang
  • Fengjun Zhang

Graph Neural Networks (GNNs) have achieved impressive performance in semi-supervised graph anomaly detection (GAD). While many GNN variants have been developed for this task, they largely focus on advanced message aggregation schemes, leaving the message routing aspect underexplored. We argue that the commonly used broadcast-based routing can also hinder generalization, particularly in the presence of rare and structurally challenging (vertices with a high-degree) anomalies. To address this, we propose Binary Message Passing (BMP), a novel routing paradigm that models the message flow of each vertex as a binary tree (BMP tree), where vanilla graph convolution is decoupled by its left and right subtrees. Each vertex recursively gathers information from neighbors with higher anomaly probabilities within each subtree, thereby amplifying the propagation of anomaly information across the topology. The anomaly probabilities are estimated and updated by the model itself, enabling adaptive, self-supervised routing over iterations. Furthermore, combining multiple BMP trees into a BMP forest provides multi-scale structural context, enhancing the expressiveness of final vertex embeddings. Extensive experiments show that BMP improves detection performance under limited supervision while exhibiting better generalization across structurally diverse anomalies.

AAAI Conference 2026 Conference Paper

BuildingWorld: A Structured 3D Building Dataset for Urban Foundation Models

  • Shangfeng Huang
  • Ruisheng Wang
  • Xin Wang

As digital twins become central to the transformation of modern cities, accurate and structured 3D building models emerge as a key enabler of high-fidelity, updatable urban representations. These models underpin diverse applications including energy modeling, urban planning, autonomous navigation, and real-time reasoning. Despite recent advances in 3D urban modeling, most learning-based models are trained on building datasets with limited architectural diversity, which significantly undermines their generalizability across heterogeneous urban environments. To address this limitation, we present BuildingWorld, a comprehensive and structured 3D building dataset designed to bridge the gap in stylistic diversity. It encompasses buildings from geographically and architecturally diverse regions—including North America, Europe, Asia, Africa, and Oceania—offering a globally representative dataset for urban-scale foundation modeling and analysis. Specifically, BuildingWorld provides about Five million LOD2 building models collected from diverse sources, accompanied by both real and simulated airborne LiDAR point clouds. This enables comprehensive research on 3D reconstruction, building detection and segmentation, as well as roof structure segmentation. Cyber City, a virtual city model, is introduced to enable the generation of unlimited training data with customized and structurally diverse point cloud distributions. Furthermore, we provide standardized evaluation metrics tailored for building reconstruction, aiming to facilitate the training, evaluation, and comparison of large-scale vision models and foundation models in structured 3D urban environments

EAAI Journal 2026 Journal Article

Connectivity-aware three-dimensional fracture segmentation method for core computed tomography images

  • Xiangxin Zhao
  • Xin Wang
  • Liguo Niu
  • Xintao Mu
  • Xuefeng Liu

Accurately extracting the fracture structures from three-dimensional (3D) computed tomography (CT) images is essential for simulating and analyzing the physical properties of digital rocks. However, the heterogeneity within the rocks makes it difficult for threshold-based methods to identify blurred fracture boundaries. Furthermore, fractures have a complex spatial topological structure, resulting in existing slice-based segmentation methods ineffective in capturing spatial connectivity information. To address the above problems, a novel fracture segmentation method for 3D core CT images is proposed in this study. Firstly, we introduced a 3D multi-layer Transformer(3D-MLT) network to capture long-range dependence information and pixel spatial continuity features between adjacent layers. Then, we fed three axial slices into a two-dimensional (2D) multi-layer Transformer(2D-MLT) network to extract anisotropic features from multi-views. Subsequently, these features are fed into the Gradient Boosting Decision Tree (GBDT) module, which is iteratively enhanced by weaker learners to obtain preliminary segmentation probability maps. To correct the contribution of these maps to the segmentation results, we add dynamic weights to each of them and adjust it by backpropagation of the loss function. Finally, a multi-scale context-aware fusion(MSCAF) module fused spatial continuity features with these maps to obtain segmentation results. We compare it with other state-of-the-art(SOTA) methods and the experiment results demonstrate the superiority of our method in spatial structure connectivity of fracture.

AAAI Conference 2026 Conference Paper

Cross-Scale Collaboration between LLMs and Lightweight Sequential Recommenders with Domain-Specific Latent Reasoning

  • Yipeng Zhang
  • Xin Wang
  • Hong Chen
  • Junwei Pan
  • Qian Li
  • Jun Zhang
  • Jie Jiang
  • Hong Mei

Sequential recommendation aims to predict the next item based on historical interactions. To further enhance the reasoning capability in sequential recommendation, LLMs are employed to predict the next item or generate semantic IDs for item representation, given LLMs' extensive domain knowledge and reasoning ability. However, existing LLM-based methods suffer from two limitations. (i) The scarcity of recommendation data with reasoning paths makes it challenging to design suitable chain-of-thought prompting templates, and the full potential of LLMs' reasoning abilities remains underutilized. (ii) Upon obtaining semantic IDs, the LLMs and their representations are excluded from the subsequent recommendation model training, preventing downstream models from fully utilizing the rich semantic information encoded within these IDs. To address these issues, we propose a novel CoderRec framework, which is capable of fully exploiting the information encoded in semantic IDs to guide the recommendation process. Specifically, to address the problem of scarcity in reasoning path-augmented data, we introduce latent reasoning into sequential recommendation and treat the representation captured by the downstream model as domain-specific latent thought, enabling implicit logical inference without requiring explicit CoT annotations. To ensure that the downstream recommendation models are able to deeply leverage the semantic information within IDs, we propose a novel cross-scale model collaboration strategy, which employs cross-scale IDs and a two-phase approach to align LLM-derived semantics with recommendation objectives. Extensive experiments have shown the effectiveness of our proposed CoderRec framework.

AAAI Conference 2026 Conference Paper

Dual Mamba for Node-Specific Representation Learning: Tackling Over-Smoothing with Selective State Space Modeling

  • Xin He
  • Yili Wang
  • Yiwei Dai
  • Xin Wang

Over-smoothing remains a fundamental challenge in deep Graph Neural Networks (GNNs), where repeated message passing causes node representations to become indistinguishable. While existing solutions, such as residual connections and skip layers, alleviate this issue to some extent, they fail to explicitly model how node representations evolve in a node-specific and progressive manner across layers. Moreover, these methods do not take global information into account, which is also crucial for mitigating the over-smoothing problem. To address the aforementioned issues, in this work, we propose a Dual Mamba-enhanced Graph Convolutional Network (DMbaGCN), which is a novel framework that integrates Mamba into GNNs to address over-smoothing from both local and global perspectives. DMbaGCN consists of two modules: the Local State-Evolution Mamba (LSEMba) for local neighborhood aggregation and utilizing Mamba’s selective state space modeling to capture node-specific representation dynamics across layers, and the Global Context-Aware Mamba (GCAMba) that leverages Mamba’s global attention capabilities to incorporate global context for each node. By combining these components, DMbaGCN enhances node discriminability in deep GNNs, thereby mitigating over-smoothing. Extensive experiments on multiple benchmarks demonstrate the effectiveness and efficiency of our method.

JBHI Journal 2026 Journal Article

FIGNet: A Robust and Interpretable Fuzzy-Irreversible Gated Network for Auditory Brainstem Response Classification

  • Ke Zhang
  • Chunrui Zhao
  • Zenan Li
  • Caiwei Li
  • Desheng Jia
  • Yongchao Chen
  • Shang Yan
  • Xin Wang

Auditory brainstem response (ABR) is an important tool for newborn hearing screening and neurological assessment. However, its signals are often difficult to be accurately resolved due to noise interference and weak waveforms, and the need for repeated measurements under multiple sound intensity conditions results in time-consuming data acquisition. Therefore, there is an urgent need to develop an automatic classification model with high accuracy, robustness and good interpretability to achieve stable and effective recognition performance with minimal ABR data. This study presents FIGNet, a new deep learning model that combines type-2 fuzzy logic with a time-irreversible attention mechanism to address uncertainty and temporal direction in ABR signals. Fuzzy attention helps reduce the impact of noise, while the irreversible attention models the one-way nature of neural responses. Experiments on real ABR datasets show that FIGNet outperforms existing models in both binary and five-class classification tasks. It achieves 93. 72% accuracy in binary classification and 84. 42% accuracy in five-class classification. Visualization results—including confusion matrices, and accuracy curves under different noise levels—further confirm that FIGNet can focus on key waveform areas and stay reliable even in noisy conditions. These findings demonstrate that FIGNet offers fast, interpretable, and robust performance for clinical ABR analysis, achieving high classification accuracy under both clean and noisy conditions.

AAAI Conference 2026 Conference Paper

HyperD: Hybrid Periodicity Decoupling Framework for Traffic Forecasting

  • Minlan Shao
  • Zijian Zhang
  • Yili Wang
  • Yiwei Dai
  • Xu Shen
  • Xin Wang

Accurate traffic forecasting plays a vital role in intelligent transportation systems, enabling applications such as congestion control, route planning, and urban mobility optimization. However, traffic forecasting remains challenging due to two key factors: (1) complex spatial dependencies arising from dynamic interactions between road segments and traffic sensors across the network, and (2) the coexistence of multi-scale periodic patterns (e.g., daily and weekly periodic patterns driven by human routines) with irregular fluctuations caused by unpredictable events (e.g., accidents, weather, or construction). To tackle these challenges, we propose HyperD (Hybrid Periodic Decoupling), a novel framework that decouples traffic data into periodic and residual components. The periodic component is handled by the Hybrid Periodic Representation Module, which extracts fine-grained daily and weekly patterns using learnable periodic embeddings and spatial-temporal attention. The residual component, which captures non-periodic, high-frequency fluctuations, is modeled by the Frequency-Aware Residual Representation Module, leveraging complex-valued MLP in frequency domain. To enforce semantic separation between the two components, we further introduce a Dual-View Alignment Loss, which aligns low-frequency information with the periodic branch and high-frequency information with the residual branch. Extensive experiments on four real-world traffic datasets demonstrate that HyperD achieves state-of-the-art prediction accuracy, while offering superior robustness under disturbances and improved computational efficiency compared to existing methods.

EAAI Journal 2026 Journal Article

Image-plane geometric decoding for view-invariant indoor scene reconstruction

  • Mingyang Li
  • Yimeng Fan
  • Changsong Liu
  • Lixue Xu
  • Xin Wang
  • Yanyan Liu
  • Wei Zhang

Volume-based indoor scene reconstruction offers superior generalization and real-time potential. However, existing frameworks rely on weak multi-view geometric constraints, leading to quality degradation as input views decrease. In sparse-view scenarios, these methods often exhibit geometric fragmentation due to the lack of robust priors. To address this, we propose Image-Plane Geometric Decoding Reconstruction (IPDRecon) pipeline, a framework integrating geometric optical principles as inductive bias to systematically exploit single-view spatial information for view-invariant reconstruction. Our approach establishes a structured geometric constraint mechanism through three synergistic modules: the Pixel-level Confidence Encoder (PCE) leverages state–space modeling with diffuse reflection principles to extract distance and position awareness; the Affine Compensation Module (ACM) enforces rigid geometric constraints via affine invariance, enabling accurate recovery of complex structures under sparse views; and the Image-Plane Spatial Decoder (IPSD) employs a multi-source geometric prior fusion strategy to transform traditional back-projection into geometry-aware spatial encoding. Extensive experiments on benchmark datasets (ScanNet V2) demonstrate exceptional stability, achieving 79. 7% Precision and a 0. 722 harmonic mean of precision and recall (F-score). In robustness evaluations averaged on a per-scene basis across the validation set, our method shows remarkable resilience when reducing views from 100 to 60. It maintains a 99. 7% mean performance retention rate, with a per-scene coefficient of variation of 0. 24% and a maximum performance drop of only 0. 42%. These results confirm that our physics-guided approach provides a robust solution for high-fidelity reconstruction in view-limited applications.

AAAI Conference 2026 Conference Paper

Inference Scaling Law for Retrieval Augmented Generation

  • Shu Zhou
  • Yuxuan Ao
  • Yunyang Xuan
  • Xin Wang
  • Tao Fan
  • Hao Wang

Retrieval-augmented generation (RAG) has recently emerged as a powerful framework for knowledge-intensive natural language processing tasks, which leverages the strengths of both pre-trained language models and external knowledge. While significant progress has been made, the scaling behavior of these approaches during inference remains poorly understood. Towards this end, this paper presents a comprehensive study of inference scaling law for RAG models, which investigates how inference performance scales with respect to key factors including retriever model scale, generator model scale, number of retrieved documents, and context window size. Through extensive experiments on benchmark datasets, we establish empirical scaling laws that reveal power-law and sigmoid-type relationships between these factors and performance. We further build a joint inference scaling law with theoretical justification. With the proposed scaling laws, we can understand the performance tendency of RAG models under different computational resources. We believe our insights can pave the way for efficient and effective deployment of RAG models in more applications.

JBHI Journal 2026 Journal Article

Joint Learning of Confidence Fusion, Semantic Alignment and Group-Guided Reliability: A Novel Semi-Supervised Learning Framework for 3D Medical Image Segmentation

  • Xinghu Zhou
  • Guanghan Wang
  • Yuanzhi Cheng
  • Zixuan Wang
  • Xin Wang
  • Guohua Wang
  • Shinichi Tamura

Semi-supervised learning (SSL) has shown strong potential in reducing the reliance on large-scale voxel-level annotations for 3D medical image segmentation. However, existing SSL methods often suffer from unstable training and limited generalization due to unreliable pseudo-labels and insufficient structural modeling in unlabeled data. These challenges are especially evident in volumetric contexts, where anatomical structures exhibit high inter-class imbalance and complex spatial dependencies. To address these issues, we propose a semi-supervised framework built upon a single-network architecture that integrates feature learning, consistency regularization, and pseudo-label reliability modeling in a unified manner. The framework comprises three key components: 1) a Confidence-aware Multi-level Fusion Network (CMFN) for capturing robust multi-scale semantic representations; 2) a Semantic-Enhanced Center Alignment (SECA) module to align feature distributions of group-level anatomical structures and mitigate semantic drift in pseudo-labels; and 3) a Group-Guided Reliability Assessment (GGRA) module that enhances pseudo-label reliability by modeling confidence errors in a group-aware structural context. Together, these modules enhance both feature discriminability and the reliability of pseudo-labels. We evaluate our framework on three public 3D medical image segmentation benchmarks: LA, BTCV, and BraTS19. Extensive experiments demonstrate that our method consistently outperforms state-of-the-art approaches under limited annotation, achieving superior accuracy and generalization across diverse anatomical structures and segmentation tasks.

AAAI Conference 2026 Conference Paper

LUMIN: A Longitudinal Multi-modal Knowledge Decomposition Network for Predicting Breast Cancer Recurrence

  • Chunyao Lu
  • Tianyu Zhang
  • Xinglong Liang
  • Yuan Gao
  • Luyi Han
  • Xin Wang
  • Nika Rasoolzadeh
  • Tao Tan

Accurate prediction of breast cancer recurrence after treatment is essential for improving long-term outcomes. However, existing models are limited by three key challenges: (1) they typically rely on single-modal data, missing cross-modal interactions; (2) they analyze static snapshots, failing to capture disease progression over time; and (3) they often perform coarse feature fusion, lacking semantic disentanglement and interpretability. To address these issues, we propose LUMIN (Longitudinal Multi-modal Knowledge Decomposition Network), a novel framework that integrates longitudinal mammograms and electronic health records (EHRs) for recurrence prediction. LUMIN leverages a vision-language contrastive pretraining backbone to align multi-modal representations and introduces two knowledge extraction modules: (1) a Cross-Modal Disentangled Knowledge Extractor (CM-DKE) that separates shared, complementary, and modality-specific information across imaging and text; and (2) a Temporal Evolution Disentangled Knowledge Extractor (TE-DKE) that captures time-invariant, time-varying, and time-specific features to model disease dynamics. Experiments on a large-scale dataset of 3,924 patients and 19,684 exams show that LUMIN significantly outperforms state-of-the-art baselines, demonstrating its effectiveness in capturing both multi-modal semantics and temporal heterogeneity for recurrence prediction.

EAAI Journal 2026 Journal Article

Multi-objective optimization of rail welded joint grinding in railroad tracks via reinforcement learning

  • Tianci Gao
  • Yuan Wang
  • Xin Wang

Rail welded joints are prone to geometric defects that compromise operational safety and ride quality. Traditional grinding strategies often rely on fixed rules, which fail to adapt to the diverse and irregular profiles of weld defects, leading to either excessive material removal or inefficient maintenance. To address this issue, this paper develops a reinforcement learning-based multi-objective optimization framework using a dual-critic Deep Deterministic Policy Gradient (DDPG) to generate efficient grinding strategies. The proposed model simultaneously minimizes grinding volume and the number of grinding passes, addressing both material conservation and operational efficiency. A continuous-action actor network is used to predict the optimal grinding depths and locations, while two separate critic networks evaluate the trade-offs between the competing objectives. The model is trained on high-resolution field data collected from 55 real-world rail welded joints across high-speed, conventional, and subway lines. After 6000 training steps, both critic networks converged with stable critic losses and smooth policy gradients, ensuring accurate value estimations under different grinding conditions. Across 300 sampled trade-off weights, the model generates over 220 feasible grinding strategies per case, with defined Pareto-optimal fronts. Case studies demonstrate that the optimized strategies reduce grinding volume and passes, improving profile trough restoration from −0. 41 to −0. 18 mm. Therefore, the proposed method can serve as an interpretable and adaptive tool to support rail maintenance decision-making by offering context-specific grinding strategies—from aggressive single-pass interventions to multi-pass, low-depth approaches that prioritize rail longevity.

AAAI Conference 2026 Conference Paper

rMMEA: Robust Multi-Modal Entity Alignment with Missing and Noise Visual Modality

  • Lingbing Guo
  • Zhuo Chen
  • Yichi Zhang
  • Wenbin Guo
  • Haonan Yang
  • Zhao Li
  • Zirui Chen
  • Xin Wang

Recently, multi-modal embedding methods have flourished in entity alignment. As state-of-the-art approaches evolve rapidly, visual modality (i.e., images) missing emerges as a critical challenge. While visual modality typically offers the most informative signals in multi-modal entity alignment (MMEA), it is frequently unavailable for many entities. The existing methods commonly use dummy vectors to represent visual-missing embeddings, which negatively impacts both model training and inference. In this paper, we propose robust multi-modal entity alignment (rMMEA), which leverages ranking-based knowledge distillation and mutual information (MI) estimation to address missing modalities while enhancing noise robustness. Unlike conventional teacher-student distillation that requires the student to replicate teacher outputs, our rMMEA learns soft rankings from pure and complete modality sides while capturing implicit key semantics of teacher embeddings through mutual information maximization, allowing rMMEA to avoid strict point-to-point alignment. The experimental results across multiple benchmarks and settings demonstrate that rMMEA significantly outperforms the state-of-the-art anti-modality-missing methods in terms of effectiveness and efficiency.

AAAI Conference 2026 Conference Paper

Scalable Semi-supervised Community Search via Graph Transformer on Attributed Heterogeneous Information Networks

  • Linlin Ding
  • Zhaosong Zhao
  • Mo Li
  • Yishan Pan
  • Xin Wang
  • Renata Borovica-Gajic

Attributed heterogeneous information networks (AHINs) encode rich semantics through diverse node and edge types. Recent learning-based community search methods on AHINs have shown promising performance but face two major limitations: i) difficulty scaling to large graphs due to memory-intensive neighbor-based propagation (e.g., GNNs and node-level attention), and ii) reliance on explicit community-level labels, which are often unavailable or costly to obtain. To address these issues, we propose a scalable Semi-supervised Community Search framework on AHINs (SCSAH), enabling scalability and efficiency, while eliminating the need for community-level labels by leveraging readily available node classification labels. Specifically, we devise MvSF2Token to extract Multi-view Semantic Features (MvSFs) as compact subgraph-level tokens before training, significantly reducing model propagation complexity. We then design a View-Aware Semantic Graph Transformer (VASGhormer) to effectively encode MvSFs by capturing cross-view dependencies and fusing semantic features. The combination of MvSF2Token and VASGhormer ensures scalability, efficiency, and robust performance. Furthermore, we design a View-Aware Contrastive Learner to train VASGhormer without requiring community-level supervision. Extensive experiments on five real-world datasets show that SCSAH outperforms state-of-the-art methods, achieving 18.06% higher performance and 10.43 times faster training.

AAAI Conference 2026 Conference Paper

Selective Diffusion Distillation for Real-World High-Scale Image Super-Resolution

  • Wenli Zheng
  • Huiyuan Fu
  • Zekai Xu
  • Xin Wang
  • Huadong Ma

High-scale image super-resolution (SR) has become increasingly important with the rapid growth of mobile devices and high-resolution displays. However, current SR methods primarily focus on lower scales and generalize poorly to high-scale scenarios due to severe information loss and complex real-world degradations. In this paper, we propose a novel Selective Diffusion Distillation (SDD) framework for real-world high-scale SR, which distills reliable knowledge from a low-scale diffusion teacher to a high-scale student. Specifically, considering severe information loss in high-scale inputs, directly distilling from low-scale models may result in feature misalignment. To address this, we introduce a Degradation-aware Metric Learning (DML) approach to align feature distributions across different degradation levels. In addition, since the diffusion-based teacher may hallucinate artifacts in ambiguous regions, blindly imitating these unreliable outputs can degrade the student’s fidelity. To tackle this, we propose a Region-aware Selective Distillation (RSD) strategy to filter out uncertain predictions and adaptively supervise only on reliable areas. To evaluate the effectiveness of our method, we introduce Real-UltraSR, a new real-world benchmark that contains diverse high-scale LR-HR pairs, including x8, x10, x12, and x14. Extensive experiments demonstrate that our SDD framework achieves state-of-the-art performance across multiple benchmarks.

AAAI Conference 2026 Conference Paper

SMART: A Surrogate Model for Predicting Application Runtime in Dragonfly Systems

  • Xin Wang
  • Pietro Lodi Rizzini
  • Sourav Medya
  • Zhiling Lan

The Dragonfly network, with its high-radix and low-diameter structure, is a leading interconnect in high-performance computing. A major challenge is workload interference on shared network links. Parallel discrete event simulation (PDES) is commonly used to analyze workload interference. However, high-fidelity PDES is computationally expensive, making it impractical for large-scale or real-time scenarios. Hybrid simulation that incorporates data-driven surrogate models offers a promising alternative, especially for forecasting application runtime, a task complicated by the dynamic behavior of network traffic. We present SMART, a surrogate model that combines graph neural networks (GNNs) and large language models (LLMs) to capture both spatial and temporal patterns from port level router data. SMART outperforms existing statistical and machine learning baselines, enabling accurate runtime prediction and supporting efficient hybrid simulation of Dragonfly networks.

EAAI Journal 2026 Journal Article

Spatial-channel collaborative multi-scale graph interaction deep transfer learning for unsupervised rotating machinery fault diagnosis

  • Xin Wang
  • Hongkai Jiang
  • Yutong Dong
  • Mingzhe Mu

Accurate machine fault diagnosis under unlabeled scenarios remains a major challenge in the intelligent transformation driven by Industry 4. 0/5. 0. To enable cross-domain diagnosis in unlabeled target scenarios, it is both urgent and essential to extract valuable and transferable knowledge from diverse historical source domains. Graph-based multi-source transfer learning offers a promising solution. However, current methods are often constrained by inaccurate feature extraction and insufficient feature interaction, which hinder diagnostic performance. Therefore, a spatial-channel collaborative multi-scale graph interaction deep transfer learning (SCMGIDTL) is proposed. Firstly, a spatial-channel collaborative prototype extraction module is built to refine features in both spatial and channel dimensions, obtaining precise multi-domain feature prototypes to construct a high-quality graph network. Secondly, a multi-scale graph interaction transfer network is creatively established to enable multi-scale feature interaction across the multi-source domain, guiding the graph network to fuse deeper neighborhood features that benefit target graph nodes, thus enabling more accurate fault diagnosis. Finally, a category constraint loss is designed to simultaneously constrain category feature relationships from both local and global perspectives, facilitating domain alignment at the category level and further improving unsupervised fault diagnosis performance. Ablation experiments demonstrate that, starting from the graph-based transfer baseline method, the three proposed components introduce cumulative performance gains of 5. 41%, 8. 35%, and 1. 54%, respectively. The average diagnosis accuracy of multiple tasks in the two cases reaches 99. 87% and 99. 60%. These results indicate that SCMGIDTL achieves outstanding performance in unsupervised machine fault diagnosis.

EAAI Journal 2026 Journal Article

Spatiotemporal interactive multiple self-attention network for skeleton-based action recognition

  • Xin Wang
  • Long Liu
  • Siying Ren
  • Kai Lu

Graph Convolutional Networks (GCN) have shown significant advantages in skeleton-based action recognition as an effective technique for extracting action representations. However, the inherently limited receptive field of graph convolution restricts the ability of GCN-based methods to capture long-range dependencies among distant joints. Additionally, these methods normally utilize a uniform skeleton topology that models only the physically connected joints for all frames. This neglects the dependencies among non-physically connected joints and the temporal variability of joint features. To address these issues, we propose a SpatioTemporal Interactive Multiple Self-Attention (STI-MSA) network. The spatiotemporal interaction (STI) module first disentangles action features into spatiotemporal, spatial, and temporal sub-representations through convolution and multiple self-attention (MSA). Then, it performs sufficient cross-dimensional interactions to learn comprehensive and effective local and global spatiotemporal dependencies. We introduce a complementary integration of global distance encoding and the adjacency matrix as a unified prior for all representations. This enables the network to adaptively focus on the relationships between joints at varying distances, including physically and non-physically connected joints. The MSA constructs hierarchical topologies based on dimension-specific channel correlations, which are integrated with the global distance encoding and adjacency matrix to form shared local and global structures. This overcomes the limitations of fixed topologies and enhances representational capacity. Extensive experiments demonstrate the superior performance of our STI-MSA on public datasets.

EAAI Journal 2026 Journal Article

Structure-based curriculum learning ultrasound gallbladder image classification network

  • Xintao Mu
  • Shengbiao Yang
  • Jing Zhuo
  • Yang Li
  • Jia Wang
  • Cheng Peng
  • Xin Wang

Gallbladder cancer (GBC) is one of the most prevalent malignant tumors in the digestive system worldwide. Typically diagnosed at advanced stages, early detection is crucial for improving patient survival rates. Ultrasound imaging has emerged as an effective screening modality due to its noninvasive nature, real time capability, and cost-effectiveness. However, the considerable variations in lesion size, complex textural patterns, and substantial noise interference in GBC ultrasound images pose significant challenges for deep learning model training and inference. To address these challenges, this study proposes a novel two-stage deep learning framework for GBC classification. The first stage involves a comprehensive evaluation of mainstream object detection models, incorporating a newly developed gallbladder coverage ratio metric along with conventional evaluation criteria to select the optimal region-of-interest (ROI) detection network for precise gallbladder localization, thereby minimizing background noise and artifact interference. The second stage introduces an innovative curriculum learning strategy combining Relative Total Variation (RTV) and Visual Acuity (VA), which enhances classification performance in noisy environments by suppressing textural details while emphasizing structural information. Experimental results demonstrate that compared to state-of-the-art (SOTA) classification networks, our proposed method achieves minimum improvements in accuracy, specificity, and sensitivity of 1. 3%, 10. 7%, and 6. 0%, respectively.

AAAI Conference 2026 Conference Paper

U2UData+: A Scalable Swarm UAVs Autonomous Flight Dataset for Embodied Long-horizon Tasks

  • Tongtong Feng
  • Xin Wang
  • Feilin Han
  • Leping Zhang
  • Wenwu Zhu

Swarm UAV autonomous flight for Embodied Long-Horizon (ELH) tasks is crucial for advancing the low-altitude economy. However, existing methods focus only on specific basic tasks due to dataset limitations, failing in real-world deployment for ELH tasks. ELH tasks are not mere concatenations of basic tasks, requiring handling long-term dependencies, maintaining embodied persistent states, and adapting to dynamic goal shifts. This paper presents U2UData+, the first large-scale swarm UAV autonomous flight dataset for ELH tasks and the first scalable swarm UAV data online collection and algorithm closed-loop verification platform. The dataset is captured by 15 UAVs in autonomous collaborative flights for ELH tasks, comprising 12 scenes, 720 traces, 120 hours, 600 seconds per trajectory, 4.32M LiDAR frames, and 12.96M RGB frames. This dataset also includes brightness, temperature, humidity, smoke, and airflow values covering all flight routes. The platform supports the customization of simulators, UAVs, sensors, flight algorithms, formation modes, and ELH tasks. Through a visual control window, this platform allows users to collect customized datasets through one-click deployment online and to verify algorithms by closed-loop simulation. U2UData+ also introduces an ELH task for wildlife conservation and provides comprehensive benchmarks with 9 SOTA models.

JBHI Journal 2026 Journal Article

WGB-GLFI: A Novel Graph-Based Global-Local Feature Interaction Framework for Automated Seizure Detection

  • Xiang Li
  • Mingxing Zhu
  • Chuqi Yang
  • Ke Zhang
  • Xin Wang
  • Sunday Timothy Aboyeji
  • Fei Chen
  • Chen Yao

Epilepsy detection faces significant challenges due to unpredictable seizures, ranging from brief awareness lapses to severe convulsions, posing risks to patients' safety and quality of life. In recent years, deep learning has become a mainstream approach in this field, leveraging advanced computational resources and EEG datasets. However, a key challenge remains: existing methods often lack unified spatial modeling and struggle to effectively handle local detailed features, thereby limiting their accuracy and robustness. To address these issues, we propose the Weighted Graph Building Global-Local Feature Interaction (WGB-GLFI) framework, which integrates spatial connectivity and dynamic patterns through a Weighted Graph Building (WGB) module and a Global-Local Feature Interaction (GLFI) module. This approach excels by comprehensively capturing the dynamic spatial relationships during epileptic seizures and achieving seamless global-local feature integration, significantly enhancing seizure detection performance. Its effectiveness has been validated across multiple datasets, including CHB-MIT, Siena Scalp, and private datasets, demonstrating robust and reliable results. Evaluated on these datasets, our model achieves accuracy rates of 99. 28%, 99. 21%, and 99. 30%, respectively. The reliability and robustness of our framework provide epilepsy patients with faster and more reliable seizure detection, which helps to intervene in a timely manner and improve the quality of life of patients.

EAAI Journal 2025 Journal Article

A general framework for chromosomal anomaly detection based on dual constraints of nearest-neighbor and regionality

  • Yue Hao
  • Xin Wang
  • Ge Song
  • Zhiyuan Li
  • Lei Wang
  • Lingwei Li
  • Yongqi Nie
  • Peng Wang

The precise identification of structural chromosomal abnormalities (SCA) is essential for the diagnosis of genetic disorders and malignancies. Traditional karyotype analysis is labor-intensive and necessitates the expertise of cytogeneticists. We propose a dual-constraint enhanced framework that combines nearest-neighbor contrastive learning with one-class classification, facilitating automated abnormality detection without the need for anomalous data. Initially, positive sample pairs are constructed utilizing a Chromosomal Query Library (CQL). This process involves the dynamic selection of nearest neighbors, employing soft nearest neighbor selection and cosine similarity to improve feature consistency. Gaussian noise injection enhances generalization by diversifying representations, whereas a momentum update refines CQL embeddings. The Chromosome Banding module (CB module) extracts chromosomal features at multiple scales, whereas the Chromosome Batch Perception module (CBP module) emphasizes challenging samples through spatial and channel attention mechanisms. In the second stage, we present ChromosomeCutMix to create synthetic chromosomal anomalies, enhancing inter-class separation and improving anomaly detection. The proposed framework attains a classification accuracy of 97. 32% and an F1-score of 96. 69%, surpassing current methodologies in terms of sensitivity and robustness. Validated on public and clinical datasets, it offers dependable localization of biological anomalies and automated cytogenetic diagnostics, thereby enhancing the analysis of genetic disorders.

NeurIPS Conference 2025 Conference Paper

A Implies B: Circuit Analysis in LLMs for Propositional Logical Reasoning

  • Guan Zhe Hong
  • Nishanth Dikkala
  • Enming Luo
  • Cyrus Rashtchian
  • Xin Wang
  • Rina Panigrahy

Due to the size and complexity of modern large language models (LLMs), it has proven challenging to uncover the underlying mechanisms that models use to solve reasoning problems. For instance, is their reasoning for a specific problem localized to certain parts of the network? Do they break down the reasoning problem into modular components that are then executed as sequential steps as we go deeper in the model? To better understand the reasoning capability of LLMs, we study a minimal propositional logic problem that requires combining multiple facts to arrive at a solution. By studying this problem on Mistral and Gemma models, up to 27B parameters, we illuminate the core components the models use to solve such logic problems. From a mechanistic interpretability point of view, we use causal mediation analysis to uncover the pathways and components of the LLMs' reasoning processes. Then, we offer fine-grained insights into the functions of attention heads in different layers. We not only find a sparse circuit that computes the answer, but we decompose it into sub-circuits that have four distinct and modular uses. Finally, we reveal that three distinct models -- Mistral-7B, Gemma-2-9B and Gemma-2-27B -- contain analogous but not identical mechanisms.

AAAI Conference 2025 Conference Paper

Adaptive Dual Guidance Knowledge Distillation

  • Tong Li
  • Long Liu
  • Kang Liu
  • Xin Wang
  • Bo Zhou
  • Hongguang Yang
  • Kai Lu

Knowledge distillation (KD) aims to improve the performance of lightweight student networks under the guidance of pre-trained teachers. However, the large capacity gap between teachers and students limits the distillation gains. Previous methods addressing this problem have two weaknesses. First, most of them decrease the performance of pre-trained teachers, hindering students from achieving comparable performance. Second, these methods fail to dynamically adjust the transferred knowledge to be compatible with the representation ability of students, which is less effective in bridging the capacity gap. In this paper, we propose Adaptive Dual Guidance Knowledge Distillation (ADG-KD), which retains the guidance of the pre-trained teacher and uses the teacher's bidirectional optimization route guiding the student to alleviate the capacity gap problem. Specifically, ADG-KD introduces an initialized teacher, which has an identical structure to the pre-trained teacher and is optimized through the bidirectional supervision from both the pre-trained teacher and student. In this way, we construct the teacher's bidirectional optimization route to provide the students with an easy-to-hard and compatible knowledge sequence. ADG-KD trains the students under the proposed dual guidance approaches and automatically determines their importance weights, making the transferred knowledge better compatible with the representation ability of students. Extensive experiments on CIFAR-100, ImageNet, and MS-COCO demonstrate the effectiveness of our method.

EAAI Journal 2025 Journal Article

Adaptive model-agnostic meta-learning network for cross-machine fault diagnosis with limited samples

  • Mingzhe Mu
  • Hongkai Jiang
  • Xin Wang
  • Yutong Dong

Deep learning-based methods have been extensively studied in rotating machinery defect diagnosis. However, training an accurate and robust diagnostic model is still a challenge under severe domain bias and limited samples. For this reason, a new adaptive model-agnostic meta-learning (AMAML) is proposed for cross-machine fault diagnosis with limited samples. First, a novel adaptive feature encode network is built, incorporating lightweight spatial-bilateral channel attention. This enables the network to extract critical fault information in multiple dimensions adaptively within limited samples, which improves the learning efficiency of generalized diagnostic knowledge. Then, an adaptive loss computation (ALC) method is devised, which inventively realizes the interaction between loss computation and model performance. The underfitting and overfitting dilemmas under few-shot conditions are tackled by ALC. Finally, an adaptive meta-optimization strategy is proposed for dynamically adapting the update strategy of the base learner, so that the model is always optimized in the direction of strong generalizability while obtaining high performance. Six cross-machine diagnosis tasks are conducted to verify the effectiveness of AMAML. The average diagnostic accuracy of the AMAML under the 5-shot setting reached 97. 42%. Experiments confirm that AMAML is superior to other prevailing methods and is potentially promising for engineering applications.

IJCAI Conference 2025 Conference Paper

Adversarial Propensity Weighting for Debiasing in Collaborative Filtering

  • Kuiyu Zhu
  • Tao Qin
  • Pinghui Wang
  • Xin Wang

Debiased recommendation focuses on alleviating the negative impact of various biases on recommendation quality to achieve fairer personalized recommendations. Current research mainly relies on propensity score estimation or causal inference methods to alleviate selection bias; at the same time, research on prevalence bias has proposed a variety of methods based on causal graphs and contrastive learning. However, these methods have shortcomings in dealing with unstable propensity score estimates, bias interactions, and decoupling of interest and bias signals, which limits the performance improvement of recommender systems. To this end, this paper proposes APWCF, a collaborative filtering debiased method that combines dynamic propensity modeling and adversarial learning. APWCF solves the problem of high variance in propensity scores through the dynamic propensity factor, and decouples user interests and bias signals through the adversarial learning to effectively remove multiple biases. Experiments show that APWCF significantly outperforms existing methods across various benchmark datasets from different domains. Compared with the current optimal baseline PDA, Recall@10 and NDCG@10 improve by 0. 10%-5. 42% and 1. 01%-8. 60% respectively.

IROS Conference 2025 Conference Paper

An insect-scale multimodal amphibious piezoelectric robot

  • Le Wang
  • Xin Wang
  • Hanlin Wang
  • Xiqing Zuo
  • Chao Xu

Miniature amphibious robots are capable of performing various tasks in complex terrestrial and aquatic environments due to their superior adaptability. However, the mobility of existing miniature amphibious robots in such environments is limited by their complex locomotion systems and single mode of motion. This work presents a novel insect-scale amphibious robot, powered by a single piezoelectric actuator. The prototype of the robot is fabricated and preliminarily tested preliminarily. By exploiting the different vibration modes of the piezoelectric actuators, the robot achieves movement in an amphibious environment. The robot employs the acoustic flow generated by the higher-order mode to achieve rapid motion at the water surface. In addition, the robot attains forward and backward motion on the ground by means of friction force between the driving feet and the ground. The findings of this study offer significant insights into the development of amphibious robots that exhibit enhanced flexibility and adaptability. These insights lay the foundation for the future applications of such robots in narrow amphibious settings.

AAAI Conference 2025 Conference Paper

Behavior Importance-Aware Graph Neural Architecture Search for Cross-Domain Recommendation

  • Chendi Ge
  • Xin Wang
  • Ziwei Zhang
  • Yijian Qin
  • Hong Chen
  • Haiyang Wu
  • Yang Zhang
  • Yuekui Yang

Cross-domain recommendation (CDR) mitigates data sparsity and cold-start issues in recommendation systems. While recent CDR approaches using graph neural networks (GNNs) capture complex user-item interactions, they rely on manually designed architectures that are often suboptimal and labor-intensive. Additionally, extracting valuable behavioral information from source domains to improve target domain recommendations remains challenging. To address these challenges, we propose Behavior importance-aware Graph Neural Architecture Search (BiGNAS), a framework that jointly optimizes GNN architecture and data importance for CDR. BiGNAS introduces two key components: a Cross-Domain Customized Supernetwork and a Graph-Based Behavior Importance Perceptron. The supernetwork, as a one-shot, retrain-free module, automatically searches the optimal GNN architecture for each domain without the need for retraining. The perceptron uses auxiliary learning to dynamically assess the importance of source domain behaviors, thereby improving target domain recommendations. Extensive experiments on benchmark CDR datasets and a large-scale industry advertising dataset demonstrate that BiGNAS consistently outperforms state-of-the-art baselines. To the best of our knowledge, this is the first work to jointly optimize GNN architecture and behavior data importance for cross-domain recommendation.

AAAI Conference 2025 Conference Paper

ConDSeg: A General Medical Image Segmentation Framework via Contrast-Driven Feature Enhancement

  • Mengqi Lei
  • Haochen Wu
  • Xinhua Lv
  • Xin Wang

Medical image segmentation plays an important role in clinical decision making, treatment planning, and disease tracking. However, it still faces two major challenges. On the one hand, there is often a "soft boundary" between foreground and background in medical images, with poor illumination and low contrast further reducing the distinguishability of foreground and background within the image. On the other hand, co-occurrence phenomena are widespread in medical images, and learning these features is misleading to the model's judgment. To address these challenges, we propose a general framework called Contrast-Driven Medical Image Segmentation (ConDSeg). First, we develop a contrastive training strategy called Consistency Reinforcement. It is designed to improve the encoder's robustness in various illumination and contrast scenarios, enabling the model to extract high-quality features even in adverse environments. Second, we introduce a Semantic Information Decoupling module, which is able to decouple features from the encoder into foreground, background, and uncertainty regions, gradually acquiring the ability to reduce uncertainty during training. The Contrast-Driven Feature Aggregation module then contrasts the foreground and background features to guide multi-level feature fusion and key feature enhancement, further distinguishing the entities to be segmented. We also propose a Size-Aware Decoder to solve the scale singularity of the decoder. It accurately locate entities of different sizes in the image, thus avoiding erroneous learning of co-occurrence features. Extensive experiments on five datasets across three scenarios demonstrate the state-of-the-art performance of our method, proving its advanced nature and general applicability to various medical image segmentation scenarios.

EAAI Journal 2025 Journal Article

Dual-path aggregation transformer network for super-resolution with images occlusions and variability

  • Qinghui Chen
  • LunQian Wang
  • Zekai Zhang
  • XingHua Wang
  • Weilin Liu
  • Bo Xia
  • Hao Ding
  • Jinglin Zhang

While Transformer-based approaches have recently achieved notable success in super-resolution, their extensive computational requirements impede widespread practical adoption. High-resolution meteorological satellite cloud imagery is essential for weather analysis and forecasting. Enhancing image resolution through super-resolution techniques facilitates the accurate identification and localization of geographic features by meteorological systems. However, current super-resolution methods fail to restore the intricacies of cloud formations and complex regions fully. This research introduces a novel dual-path aggregation Transformer network (DPAT) tailored to enhance the super-resolution of meteorological satellite cloud images. The DPAT network adeptly captures cloud imagery's subtle details and textures, effectively addressing occlusions and the variability inherent in satellite imagery. It bolsters the model's ability to manage the complex attributes of cloud images through the introduction of the Dual-path Aggregation Self-Attention (DASA) mechanism and the Multi-scale Feature Aggregation Block (MFAB), thereby enhancing performance in processing intricate cloud features. The DASA mechanism synthesizes features across spatial, depth, and channel dimensions via a dual-path approach, thoroughly exploiting feature correlations. The MFAB, designed to supplant the multilayer perceptron, incorporates shift convolution and a multi-scale interaction block to augment feature information, compensating for the deficiency in local information absorption due to fixed receptive fields. Experimental outcomes indicate that DPAT delivers superior super-resolution outcomes. With a parameter count of only 32% of the Enhanced Deep Residual Network (EDSR) or 77% of the Image Restoration using Shift Window Transformer (SwinIR), DPAT matches SwinIR's performance on the satellite cloud dataset. Moreover, DPAT balances accuracy and parameter economy across various datasets. This technology is expected to improve image super-resolution capabilities in multiple fields such as human action recognition and industrial recognition, and indirectly improve the accuracy of image perception tasks.

IJCAI Conference 2025 Conference Paper

Dyn-D^2P: Dynamic Differentially Private Decentralized Learning with Provable Utility Guarantee

  • Zehan Zhu
  • Yan Huang
  • Xin Wang
  • Shouling Ji
  • Jinming Xu

Most existing decentralized learning methods with differential privacy (DP) guarantee rely on constant gradient clipping bounds and fixed-level DP Gaussian noises for each node throughout the training process, leading to a significant accuracy degradation compared to non-private counterparts. In this paper, we propose a new Dynamic Differentially Private Decentralized learning approach (termed Dyn-D^2P) tailored for general time-varying directed networks. Leveraging the Gaussian DP (GDP) framework for privacy accounting, Dyn-D^2P dynamically adjusts gradient clipping bounds and noise levels based on gradient convergence. This proposed dynamic noise strategy enables us to enhance model accuracy while preserving the total privacy budget. Extensive experiments on benchmark datasets demonstrate the superiority of Dyn-D^2P over its counterparts employing fixed-level noises, especially under strong privacy guarantees. Furthermore, we provide a provable utility bound for Dyn-D^2P that establishes an explicit dependency on network-related parameters, with a scaling factor of 1/sqrt{n} in terms of the number of nodes n up to a bias error term induced by gradient clipping. To our knowledge, this is the first model utility analysis for differentially private decentralized non-convex optimization with dynamic gradient clipping bounds and noise levels.

TMLR Journal 2025 Journal Article

Efficient Diffusion Models: A Survey

  • Hui Shen
  • Jingxuan Zhang
  • Boning Xiong
  • Rui Hu
  • Shoufa Chen
  • Zhongwei Wan
  • Xin Wang
  • Yu Zhang

Diffusion models have emerged as powerful generative models capable of producing high-quality contents such as images, videos, and audio, demonstrating their potential to revolutionize digital content creation. However, these capabilities come at the cost of significant computational resources and lengthy generation time, underscoring the critical need to develop efficient techniques for practical deployment. In this survey, we provide a systematic and comprehensive review of research on efficient diffusion models. We organize the literature in a taxonomy consisting of three main categories, covering distinct yet interconnected efficient diffusion model topics from algorithm-level, system-level, and framework perspective, respectively. We have also created a GitHub repository where we organize the papers featured in this survey at github.com/AIoT-MLSys-Lab/Efficient-Diffusion-Model-Survey. We hope our survey can serve as a valuable resource to help researchers and practitioners gain a systematic understanding of efficient diffusion model research and inspire them to contribute to this important and exciting field.

IJCAI Conference 2025 Conference Paper

Enhancing Counterfactual Estimation: A Focus on Temporal Treatments

  • Xin Wang
  • Shengfei Lyu
  • Kangyang Luo
  • Lishan Yang
  • Huanhuan Chen
  • Chunyan Miao

In the medical field, treatment sequences significantly influence future outcomes through complex temporal interactions. Therefore, highlighting the role of temporal treatments within the model is crucial for accurate counterfactual estimation, which is often overlooked in current methods. To address this, we employ Koopman theory, known for its capability to model complex dynamic systems, and introduce a novel model named the Counterfactual Temporal Dynamics Network via Neural Koopman Operators (CTD-NKO). This model utilizes Koopman operators to encapsulate sequential treatment data, aiming to capture the causal dynamics within the system induced by temporal interactions between treatments. Moreover, CTD-NKO implements a weighting strategy that aligns joint and marginal distributions of the system state and the current treatment to mitigate time-varying confounding bias. This deviates from the balanced representation strategy employed by existing methods, as we demonstrate that such a strategy may suffer from the potential information loss of historical treatments. These designs allow CTD-NKO to exploit treatment information more thoroughly and effectively, resulting in superior performance on both synthetic and real-world datasets.

JBHI Journal 2025 Journal Article

Ensemble Feature Selection for Microarray Data Classification

  • Xiaojian Ding
  • Pengcheng Shi
  • Xin Wang
  • Kaixiang Wang

Microarray data classification is challenged by high dimensionality and small sample sizes, causing feature selection instability. Traditional ensemble feature selection methods struggle to balance diversity and quality effectively. We propose a novel Ensemble Feature Selection Method (EFSM) that introduces a feature mapping diversity metric to generate a robust candidate pool. EFSM first generates a diverse candidate pool of feature selectors by leveraging randomized neural networks to create multiple non-linear feature mappings (views) of the original data. Its core innovation is an ensemble pruning technique formulated as an optimization problem that jointly maximizes both the predictive accuracy of individual selectors and their pairwise diversity. We simplify this NP-hard problem by converting it into a Semi-Definite Programming (SDP) problem and deriving a novel bound for efficient solution. Finally, the rankings from the pruned ensemble are aggregated using the Borda count method. Extensive experiments on 15 biological datasets demonstrate that EFSM outperforms nine state-of-the-art feature selection methods across popular classifiers, achieving superior and stable performance for high-dimensional data analysis.

IJCAI Conference 2025 Conference Paper

FedSaaS: Class-Consistency Federated Semantic Segmentation via Global Prototype Supervision and Local Adversarial Harmonization

  • Xiaoyang Yu
  • Xiaoming Wu
  • Xin Wang
  • Dongrun Li
  • Ming Yang
  • Peng Cheng

Federated semantic segmentation enables pixel-level classification in images through collaborative learning while maintaining data privacy. However, existing research commonly overlooks the fine-grained class relationships within the semantic space when addressing heterogeneous problems, particularly domain shift. This oversight results in ambiguities between class representation. To overcome this challenge, we propose a novel federated segmentation framework that strikes class consistency, termed FedSaaS. Specifically, we introduce class exemplars as a criterion for both local- and global-level class representations. On the server side, the uploaded class exemplars are leveraged to model class prototypes, which supervise global branch of clients, ensuring alignment with global-level representation. On the client side, we incorporate an adversarial mechanism to harmonize contributions of global and local branches, leading to consistent output. Moreover, multilevel contrastive losses are employed on both sides to enforce consistency between two-level representations in the same semantic space. Extensive experiments on five driving scene segmentation datasets demonstrate that our framework outperforms state-of-the-art methods, significantly improving average segmentation accuracy and effectively addressing the class-consistency representation problem.

NeurIPS Conference 2025 Conference Paper

GMM-based VAE model with Normalising Flow for effective stochastic segmentation

  • Conghui Li
  • Chern Hong Lim
  • Xin Wang

While deep neural networks possess the capability to perform semantic segmentation, producing a single deterministic output limits reliability in safety-critical applications, caused by uncertainty and annotation variability. To address this, stochastic segmentation models using Conditional Variational Autoencoders (CVAE), Bayesian networks, and diffusion have been explored. However, existing approaches suffer from limited latent expressiveness and interpretability. Furthermore, our experiments showed that models like Probabilistic U-Net rely excessively on high latent variance, leading to posterior collapse. This work propose a novel framework by integrating Gaussian Mixture Model (GMM) with Normalizing Flow (NF) in CVAE for stochastic segmentation. GMM structures the latent space into meaningful semantic clusters, while NF captures feature deformations with quantified uncertainty. Our method stabilizes latent distributions through constrained variance and mean ranges. Experiments on LIDC, Crack500, and Cityscapes datasets show that our approach outperformed state-of-the-art in curvilinear structure and medical image segmentation.

NeurIPS Conference 2025 Conference Paper

GraphChain: Large Language Models for Large-scale Graph Analysis via Tool Chaining

  • Chunyu Wei
  • Wenji Hu
  • Xingjia Hao
  • Xin Wang
  • Yifan Yang
  • Yunhai Wang
  • Yang Tian
  • Yueguo Chen

Large Language Models (LLMs) face significant limitations when applied to large-scale graphs, struggling with context constraints and inflexible reasoning. We introduce GraphChain, a novel framework enabling LLMs to analyze large graphs by orchestrating dynamic sequences of specialized tools, mimicking human exploratory processes. GraphChain incorporates two core technical contributions: (1) Progressive Graph Distillation, a reinforcement learning approach that learns to generate tool sequences balancing task relevance and intermediate state compression, thereby overcoming LLM context limitations. (2) Structure-aware Test-Time Adaptation (STTA), a mechanism using a lightweight, self-supervised adapter conditioned on graph spectral properties to efficiently adapt a frozen LLM policy to diverse graph structures via soft prompts without retraining. Experiments show GraphChain significantly outperforms prior methods, enabling scalable and adaptive LLM-driven graph analysis.

NeurIPS Conference 2025 Conference Paper

GRIT: Teaching MLLMs to Think with Images

  • Yue Fan
  • Xuehai He
  • Diji Yang
  • Kaizhi Zheng
  • Ching-Chen Kuo
  • Yuting Zheng
  • Xinze Guan
  • Xin Wang

Recent studies have demonstrated the efficacy of using Reinforcement Learning (RL) in building reasoning models that articulate chains of thoughts prior to producing final answers. However, despite ongoing advances that aim at enabling reasoning for vision-language tasks, existing open-source visual reasoning models typically generate reasoning content with pure natural language, lacking explicit integration of visual information. This limits their ability to produce clearly articulated and visually grounded reasoning chains. To this end, we propose Grounded Reasoning with Images and Texts (GRIT), a novel method for training MLLMs to think with images. GRIT introduces a grounded reasoning paradigm, in which models generate reasoning chains that interleave natural language and explicit bounding box coordinates. These coordinates point to regions of the input image that the model consults during its reasoning process. Additionally, GRIT is equipped with a reinforcement learning approach, GRPO-GR, built upon the GRPO algorithm. GRPO-GR employs robust rewards focused on the final answer accuracy and format of the grounded reasoning output, which eliminates the need for data with reasoning chain annotations or explicit bounding box labels. As a result, GRIT achieves exceptional data efficiency, requiring as few as 20 image-question-answer triplets from existing datasets. Comprehensive evaluations demonstrate that GRIT effectively trains MLLMs to produce coherent and visually grounded reasoning chains, showing a successful unification of reasoning and grounding abilities. All code, data, and checkpoints will be released.

NeurIPS Conference 2025 Conference Paper

GTR-Loc: Geospatial Text Regularization Assisted Outdoor LiDAR Localization

  • Shangshu Yu
  • Wen Li
  • Xiaotian Sun
  • Zhimin Yuan
  • Xin Wang
  • Sijie Wang
  • Rui She
  • Cheng Wang

Prevailing scene coordinate regression methods for LiDAR localization suffer from localization ambiguities, as distinct locations can exhibit similar geometric signatures — a challenge that current geometry-based regression approaches have yet to solve. Recent vision–language models show that textual descriptions can enrich scene understanding, supplying potential localization cues missing from point cloud geometries. In this paper, we propose GTR-Loc, a novel text-assisted LiDAR localization framework that effectively generates and integrates geospatial text regularization to enhance localization accuracy. We propose two novel designs: a Geospatial Text Generator that produces discrete pose-aware text descriptions, and a LiDAR-Anchored Text Embedding Refinement module that dynamically constructs view-specific embeddings conditioned on current LiDAR features. The geospatial text embeddings act as regularization to effectively reduce localization ambiguities. Furthermore, we introduce a Modality Reduction Distillation strategy to transfer textual knowledge. It enables high-performance LiDAR-only localization during inference, without requiring runtime text generation. Extensive experiments on challenging large-scale outdoor datasets, including QEOxford, Oxford Radar RobotCar, and NCLT, demonstrate the effectiveness of GTR-Loc. Our method significantly outperforms state-of-the-art approaches, notably achieving a 9. 64%/8. 04% improvement in position/orientation accuracy on QEOxford. Our code is available at https: //github. com/PSYZ1234/GTR-Loc.

AAAI Conference 2025 Conference Paper

Identity-Text Video Corpus Grounding

  • Bin Huang
  • Xin Wang
  • Hong Chen
  • Houlun Chen
  • Yaofei Wu
  • Wenwu Zhu

Video corpus grounding (VCG), which aims to retrieve relevant video moments from a video corpus, has attracted significant attention in the multimedia research community. However, the existing VCG setting primarily focuses on matching textual descriptions with videos and ignores the distinct visual identities in the videos, thus resulting in inaccurate understanding of video content and deteriorated retrieval performances. To address this limitation, we introduce a novel task, Identity-Text Video Corpus Grounding (ITVCG), which simultaneously utilize textual descriptions and visual identities as queries. As such, ITVCG benefits in enabling more accurate video corpus grounding with visual identities, as well as providing users with more flexible options to locate relevant frames based on either textual descriptions or textual descriptions and visual identities. To conduct evaluations regarding the novel ITVCG task, we propose the TVR-IT dataset, comprising 463 identity images from 6 TV shows, with 68,840 out of 72,840 queries containing at least one identity image. Furthermore, we propose Video-Locator, the first model designed for the ITVCG task. Our proposed Video-Locator integrates video-identity-text alignment and multi-modal fine-grained fusion components, facilitating a video large language model (Video LLM) to jointly understand textual descriptions, visual identities, as well as videos. Experimental results demonstrate the effectiveness of the proposed Video-Locator model and highlight the importance of identity-generalization capability for ITVCG.

YNIMG Journal 2025 Journal Article

Image-based meta- and mega-analysis (IBMMA): A unified framework for large-scale, multi-site, neuroimaging data analysis

  • Nick Steele
  • Ashley A. Huggins
  • Rajendra A. Morey
  • Ahmed Hussain
  • Courtney Russell
  • Benjamin Suarez-Jimenez
  • Elena Pozzi
  • Hadis Jameei

The increasing scale and complexity of neuroimaging datasets aggregated from multiple study sites present substantial analytic challenges, as existing statistical analysis tools struggle to handle missing voxel-data, suffer from limited computational speed and inefficient memory allocation, and are restricted in the types of statistical designs they are able to model. We introduce Image-Based Meta- & Mega-Analysis (IBMMA), a novel software package implemented in R and Python that provides a unified framework for analyzing diverse neuroimaging features, efficiently handles large-scale datasets through parallel processing, offers flexible statistical modeling options, and properly manages missing voxel-data commonly encountered in multi-site studies. IBMMA successfully analyzed a large-n dataset of several thousand participants and revealed findings in brain regions that some traditional software overlooked due to missing voxel-data resulting in gaps in brain coverage. IBMMA has the potential to accelerate discoveries in neuroscience and enhance the clinical utility of neuroimaging findings.

ICML Conference 2025 Conference Paper

Implicit degree bias in the link prediction task

  • Rachith Aiyappa
  • Xin Wang
  • Munjung Kim
  • Ozgur Can Seckin
  • Yong-Yeol Ahn
  • Sadamori Kojaku

Link prediction—a task of distinguishing actual hidden edges from random unconnected node pairs—is one of the quintessential tasks in graph machine learning. Despite being widely accepted as a universal benchmark and a downstream task for representation learning, the link prediction benchmark’s validity has rarely been questioned. Here, we show that the common edge sampling procedure in the link prediction task has an implicit bias toward high-degree nodes. This produces a highly skewed evaluation that favors methods overly dependent on node degree. In fact a “null” link prediction method based solely on node degree can yield nearly optimal performance in this setting. We propose a degree-corrected link prediction benchmark that offers a more reasonable assessment and better aligns with the performance on the recommendation task. Finally, we demonstrate that the degree-corrected benchmark can more effectively train graph machine-learning models by reducing overfitting to node degrees and facilitating the learning of relevant structures in graphs.

AAAI Conference 2025 Conference Paper

Improving Generalization for AI-Synthesized Voice Detection

  • Hainan Ren
  • Li Lin
  • Chun-Hao Liu
  • Xin Wang
  • Shu Hu

AI-synthesized voice technology has the potential to create realistic human voices for beneficial applications, but it can also be misused for malicious purposes. While existing AI-synthesized voice detection models excel in intra-domain evaluation, they face challenges in generalizing across different domains, potentially becoming obsolete as new voice generators emerge. Current solutions use diverse data and advanced machine learning techniques (e.g., domain-invariant representation, self-supervised learning), but are limited by predefined vocoders and sensitivity to factors like background noise and speaker identity. In this work, we introduce an innovative disentanglement framework aimed at extracting domain-agnostic artifact features related to vocoders. Utilizing these features, we enhance model learning in a flat loss landscape, enabling escape from suboptimal solutions and improving generalization. Extensive experiments on benchmarks show our approach outperforms state-of-the-art methods, achieving up to 5.12% improvement in the equal error rate metric in intra-domain and 7.59% in cross-domain evaluations.

AAAI Conference 2025 Conference Paper

JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-Exploration

  • Mingzi Wang
  • Yuan Meng
  • Chen Tang
  • Weixiang Zhang
  • Yijian Qin
  • Yang Yao
  • Yingxin Li
  • Tongtong Feng

The co-design of neural network architectures, quantization precisions, and hardware accelerators offers a promising approach to achieving an optimal balance between performance and efficiency, particularly for model deployment on resource-constrained edge devices. In this work, we propose the JAQ Framework, which jointly optimizes the three critical dimensions. However, effectively automating the design process across the vast search space of those three dimensions poses significant challenges, especially when pursuing extremely low-bit quantization. Specifical, the primary challenges include: (1) Memory overhead in software-side: Low-precision quantization-aware training can lead to significant memory usage due to storing large intermediate features and latent weights for backpropagation, potentially causing memory exhaustion. (2) Search time-consuming in hardware-side: The discrete nature of hardware parameters and the complex interplay between compiler optimizations and individual operators make the accelerator search time-consuming. To address these issues, JAQ mitigates the memory overhead through a channel-wise sparse quantization (CSQ) scheme, selectively applying quantization to the most sensitive components of the model during optimization. Additionally, JAQ designs BatchTile, which employs a hardware generation network to encode all possible tiling modes, thereby speeding up the search for the optimal compiler mapping strategy. Extensive experiments demonstrate the effectiveness of JAQ, achieving approximately 7% higher Top-1 accuracy on ImageNet compared to previous methods and reducing the hardware search time per iteration to 0.15 seconds.

IJCAI Conference 2025 Conference Paper

Latte: Transfering LLMs' Latent-level Knowledge for Few-shot Tabular Learning

  • Ruxue Shi
  • Hengrui Gu
  • Hangting Ye
  • Yiwei Dai
  • Xu Shen
  • Xin Wang

Few-shot tabular learning, in which machine learning models are trained with a limited amount of labeled data, provides a cost-effective approach to addressing real-world challenges. The advent of Large Language Models (LLMs) has sparked interest in leveraging their pre-trained knowledge for few-shot tabular learning. Despite promising results, existing approaches either rely on test-time knowledge extraction, which introduces undesirable latency, or text-level knowledge, which leads to unreliable feature engineering. To overcome these limitations, we propose Latte, a training-time knowledge extraction framework that transfers the latent prior knowledge within LLMs to optimize a more generalized downstream model. Latte enables general knowledge-guided downstream tabular learning, facilitating the weighted fusion of information across different feature values while reducing the risk of overfitting to limited labeled data. Furthermore, Latte is compatible with existing unsupervised pre-training paradigms and effectively utilizes available unlabeled samples to overcome the performance limitations imposed by an extremely small labeled dataset. Extensive experiments on various few-shot tabular learning benchmarks demonstrate the superior performance of Latte, establishing it as a state-of-the-art approach in this domain. Our code is available at https: //github. com/ruxueshi/Latte. git.

IJCAI Conference 2025 Conference Paper

Mamba-Based Graph Convolutional Networks: Tackling Over-smoothing with Selective State Space

  • Xin He
  • Yili Wang
  • Wenqi Fan
  • Xu Shen
  • Xin Juan
  • Rui Miao
  • Xin Wang

Graph Neural Networks (GNNs) have shown great success in various graph-based learning tasks. However, it often faces the issue of over-smoothing as the model depth increases, which causes all node representations to converge to a single value and become indistinguishable. This issue stems from the inherent limitations of GNNs, which struggle to distinguish the importance of information from different neighborhoods. In this paper, we introduce MbaGCN, a novel graph convolutional architecture that draws inspiration from the Mamba paradigm—originally designed for sequence modeling. MbaGCN presents a new backbone for GNNs, consisting of three key components: the Message Aggregation Layer, the Selective State Space Transition Layer, and the Node State Prediction Layer. These components work in tandem to adaptively aggregate neighborhood information, providing greater flexibility and scalability for deep GNN models. While MbaGCN may not consistently outperform all existing methods on each dataset, it provides a foundational framework that demonstrates the effective integration of the Mamba paradigm into graph representation learning. Through extensive experiments on benchmark datasets, we demonstrate that MbaGCN paves the way for future advancements in graph neural network research. Our code is in https: //github. com/hexin5515/MbaGCN.

AAAI Conference 2025 Conference Paper

Modular-Cam: Modular Dynamic Camera-view Video Generation with LLM

  • Zirui Pan
  • Xin Wang
  • Yipeng Zhang
  • Hong Chen
  • Kwan Man Cheng
  • Yaofei Wu
  • Wenwu Zhu

Text-to-Video generation, which utilizes the provided text prompt to generate high-quality videos, has drawn increasing attention and achieved great success due to the development of diffusion models recently. Existing methods mainly rely on a pre-trained text encoder to capture the semantic information and perform cross attention with the encoded text prompt to guide the generation of video. However, when it comes to complex prompts that contain dynamic scenes and multiple camera-view transformations, these methods can not decompose the overall information into separate scenes, as well as fail to smoothly change scenes based on the corresponding camera-views. To solve these problems, we propose a novel method, i.e., Modular-Cam. Specifically, to better understand a given complex prompt, we utilize a large language model to analyze user instructions and decouple them into multiple scenes together with transition actions. To generate a video containing dynamic scenes that match the given camera-views, we incorporate the widely-used temporal transformer into the diffusion model to ensure continuity within a single scene and propose CamOperator, a modular network based module that well controls the camera movements. Moreover, we propose AdaControlNet, which utilizes ControlNet to ensure consistency across scenes and adaptively adjusts the color tone of the generated video. Extensive qualitative and quantitative experiments prove our proposed Modular-Cam's strong capability of generating multi-scene videos together with its ability to achieve fine-grained control of camera movements. Generated results are available at https://modular-cam.github.io.

NeurIPS Conference 2025 Conference Paper

More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models

  • Zhongxing Xu
  • Chengzhi Liu
  • Qingyue Wei
  • Juncheng Wu
  • James Zou
  • Xin Wang
  • Yuyin Zhou
  • Sheng Liu

Test-time compute has empowered multimodal large language models to generate extended reasoning chains, yielding strong performance on tasks such as multimodal math reasoning. However, we observe that this improved reasoning ability often comes with increased hallucination: as generations become longer, models tend to drift away from image-grounded content and rely more on language priors. Attention analysis reveals that longer reasoning chains reduce focus on visual inputs, contributing to hallucination. To systematically study this phenomenon, we introduce RH-AUC, a metric that quantifies how a model's perception accuracy changes with reasoning length, enabling evaluation of whether the model preserves visual grounding while reasoning. We also release RH-Bench, a diagnostic benchmark covering diverse multimodal tasks, designed to jointly assess the balance of reasoning ability and hallucination. We find that (i) larger models generally exhibit a better balance between reasoning and perception; (ii) reasoning and perception balance depends more on the types and domains of the training data than its volume. Our findings highlight the need for evaluation frameworks that account for both reasoning quality and perceptual reliability.

JBHI Journal 2025 Journal Article

Multi-Modal Longitudinal Representation Learning for Predicting Neoadjuvant Therapy Response in Breast Cancer Treatment

  • Yuan Gao
  • Tao Tan
  • Xin Wang
  • Regina Beets-Tan
  • Tianyu Zhang
  • Luyi Han
  • Antonio Portaluri
  • Chunyao Lu

Longitudinal medical imaging is crucial for monitoring neoadjuvant therapy (NAT) response in clinical practice. However, mainstream artificial intelligence (AI) methods for disease monitoring commonly rely on extensive segmentation labels to evaluate lesion progression. While self-supervised vision-language (VL) learning efficiently captures medical knowledge from radiology reports, existing methods focus on single time points, missing opportunities to leverage temporal self-supervision for disease progression tracking. In addition, extracting dynamic progression from longitudinal unannotated images with corresponding textual data poses challenges. In this work, we explicitly account for longitudinal NAT examinations and accompanying reports, encompassing scans before NAT and follow-up scans during mid-/post-NAT. We introduce the multi-modal longitudinal representation learning pipeline (MLRL), a temporal foundation model, that employs multi-scale self-supervision scheme, including single-time scale vision-text alignment (VTA) learning and multi-time scale visual/textual progress (TVP/TTP) learning to extract temporal representations from each modality, thereby facilitates the downstream evaluation of tumor progress. Our method is evaluated against several state-of-the-art self-supervised longitudinal learning and multi-modal VL methods. Results from internal and external datasets demonstrate that our approach not only enhances label efficiency across the zero-, few- and full-shot regime experiments but also significantly improves tumor response prediction in diverse treatment scenarios. Furthermore, MLRL enables interpretable visual tracking of progressive areas in temporal examinations, offering insights into longitudinal VL foundation tools and potentially facilitating the temporal clinical decision-making process.

AIIM Journal 2025 Journal Article

Online continuous learning of users suicidal risk on social media

  • Lei Cao
  • Ling Feng
  • Yang Ding
  • Huijun Zhang
  • Xin Wang
  • Kaisheng Zeng
  • Yi Dai

Suicide is a tragedy for family and society. With social media becoming an integral part of people’s life nowadays, assessing suicidal risk based on one’s social media behavior has drawn increasing research attentions. The majority of the works trained a machine learning model to classify user’s suicidal risk severity level in a batch learning setting on the entire training data. This is not a timely and scalable solution in the context of social media where new data arrives sequentially in a stream form. In this study, we formulate and address the continuous suicidal risk assessment problem through a three-layered joint memory network, consisting of a short-term personal memory and long-term personal and global memories. Unlike existing methods that rely on static classification, our model supports real-time, continuous learning from users’ emotional and behavioral dynamics without the need for full retraining. This allows for personalized and adaptive risk tracking over time. We also present a way to continuously capture users’ personal features and integrate them in suicidal risk assessment. The performance on the constructed dataset containing 95 suicidal and 95 non-suicidal social media users shows that 96% of accuracy can be achieved with the proposed method.

NeurIPS Conference 2025 Conference Paper

Optimization Inspired Few-Shot Adaptation for Large Language Models

  • Boyan Gao
  • Xin Wang
  • Yibo Yang
  • David Clifton

Large Language Models (LLMs) have demonstrated remarkable performance in real-world applications. However, adapting LLMs to novel tasks via fine-tuning often requires substantial training data and computational resources that are impractical in few-shot scenarios. Existing approaches, such as In-context learning and Parameter-Efficient Fine-Tuning (PEFT), face key limitations: In-context learning introduces additional inference computational overhead with limited performance gains, while PEFT models are prone to overfitting on the few demonstration examples. In this work, we reinterpret the forward pass of LLMs as an optimization process, a sequence of preconditioned gradient descent steps refining internal representations. Based on this connection, we propose Optimization-Inspired Few-Shot Adaptation (OFA), integrating a parameterization that learns preconditioners without introducing additional trainable parameters, and an objective that improves optimization efficiency by learning preconditioners based on a convergence bound, while simultaneously steering the optimization path toward the flat local minimum. Our method overcomes both issues of ICL-based and PEFT-based methods, and demonstrates superior performance over the existing methods on a variety of few-shot adaptation tasks in experiments.

NeurIPS Conference 2025 Conference Paper

Out-of-Distribution Generalized Graph Anomaly Detection with Homophily-aware Environment Mixup

  • Sibo Tian
  • Xin Wang
  • Zeyang Zhang
  • Haibo Chen
  • Wenwu Zhu

Graph anomaly detection (GAD) is widely prevalent in scenarios such as financial fraud detection, anti-money laundering, and social bot detection. However, structural distribution shifts are commonly observed in real-world GAD data due to selection bias, resulting in reduced homophily. Existing GAD methods tend to rely on homophilic shortcuts when trained on high-homophily structures, limiting their ability to generalize well to data with low homophily under structural distribution shifts. In this study, we propose to handle structural distribution shifts by generating novel environments characterized by diverse homophilic structures and utilizing invariant patterns, i. e. , features and structures with the capability of stable prediction across structural distribution shifts, which face two challenges: (1) How to discover invariant patterns from entangled features and structures, as structures are sensitive to varying homophilic distributions. (2) How to systematically construct new environments with diverse homophilic structures. To address these challenges, we propose the Ego-Neighborhood Disentangled Encoder with Homophily-aware Environment Mixup (HEM), which effectively handles structural distribution shifts in GAD by discovering invariant patterns. Specifically, we first propose an ego-neighborhood disentangled encoder to decouple the learning of feature embeddings and structural embeddings, which facilitates subsequent improvements in the invariance of structural embeddings for prediction. Next, we introduce a homophily-aware environment mixup that dynamically adjusts edge weights through adversarial learning, effectively generating environments with diverse structural distributions. Finally, we iteratively train the classifier and environment mixup via adversarial training, simultaneously improving the diversity of constructed environments and discovering invariant patterns under structural distribution shifts. Extensive experiments on real-world datasets demonstrate that our method outperforms existing baselines and achieves state-of-the-art performance under structural distribution shift conditions.

NeurIPS Conference 2025 Conference Paper

Pattern-Guided Adaptive Prior for Structure Learning

  • Lyuzhou Chen
  • Yijia Sun
  • Yanze Gao
  • Xiangyu Wang
  • Derui Lyu
  • Taiyu Ban
  • Xin Wang
  • Xiren Zhou

Learning the causality between variables, known as DAG structure learning, is critical yet challenging due to issues such as insufficient data and noise. While prior knowledge can improve the learning process and refine the DAG structure, incorporating prior knowledge is not without pitfalls. In particular, we find that the gap between the imprecise prior knowledge and the exact weights modeled by existing methods may result in deviation in edge weights. Such deviation can subsequently cause significant inaccuracies when learning the DAG structure. This paper addresses this challenge by providing a theoretical analysis of the impact of deviation in edge weights during the optimization process of structure learning. We identify two special graph patterns that arise due to the deviation and show that their occurrence increases as the degree of deviation grows. Building on this analysis, we propose the Pattern-Guided Adaptive Prior (PGAP) framework. PGAP detects these patterns as structural signals during optimization and adaptively adjusts the structure learning process to counteract the identified weight deviation, thereby improving the integration of prior knowledge. Experiments verify the effectiveness and robustness of the proposed method.

EAAI Journal 2025 Journal Article

Polarization-based Camouflaged Object Detection with high-resolution adaptive fusion Network

  • Xin Wang
  • Junfeng Xu
  • Jiajia Ding

In comparison to traditional object detection or segmentation tasks, Camouflaged Object Detection (COD) poses greater challenges, as humans are often perplexed or deceived by the inherent similarities between foreground objects and their background surroundings. Polarization information serves as a valuable asset for discerning the attributes of objects with varied characteristics and surface texture. Taking inspiration from the polarization vision systems observed in animals, this study presents the High-Resolution Intensity & Polarization Fusion (HIPF) Net, a high-efficiency cross-modal fusion network that leverages trichromatic intensity and linear orthogonal polarization cues to produce a scene representation that is rich in texture and edge details. Specifically, the Early Adaptive Stokes Fusion (EASF) module maximizes the utilization of information from linear orthogonal polarization images. Subsequently, the Mix-Attention Feature Interaction Module (MAI) is introduced to facilitate complementary interaction among low-level features. Additionally, the Attentional Receptive Field Block (ARFB) enables the model to uncover concealed cues effectively, capturing objects of various sizes. Finally, the Weighted Cross-Level Decoder(WCFD) is designed to dynamically fuse and assign weights to cross-level contextual information for robust detection. Training and extensive validation of our model are performed on the polarization-based dataset as well as non-polarization-based datasets, with experimental results demonstrating that HIPFNet consistently outperforms state-of-the-art methods. Source codes are available at https: //github. com/CVhfut/HIPFNet.

ICML Conference 2025 Conference Paper

Predictive Performance of Deep Quantum Data Re-uploading Models

  • Xin Wang
  • Hanxiao Tao
  • Rebing Wu

Quantum machine learning models incorporating data re-uploading circuits have garnered significant attention due to their exceptional expressivity and trainability. However, their ability to generate accurate predictions on unseen data, referred to as the predictive performance, remains insufficiently investigated. This study reveals a fundamental limitation in predictive performance when deep encoding layers are employed within the data re-uploading model. Concretely, we theoretically demonstrate that when processing high-dimensional data with limited-qubit data re-uploading models, their predictive performance progressively degenerates to near random-guessing levels as the number of encoding layers increases. In this context, the repeated data uploading cannot mitigate the performance degradation. These findings are validated through experiments on both synthetic linearly separable datasets and real-world datasets. Our results demonstrate that when processing high-dimensional data, the quantum data re-uploading models should be designed with wider circuit architectures rather than deeper and narrower ones.

ICML Conference 2025 Conference Paper

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

  • Utkarsh Saxena
  • Sayeh Sharify
  • Kaushik Roy 0001
  • Xin Wang

Post-training quantization (PTQ) of large language models (LLMs) holds the promise in reducing the prohibitive computational cost at inference time. Quantization of all weight, activation and key-value (KV) cache tensors to 4-bit without significantly degrading generalizability is challenging, due to the high quantization error caused by extreme outliers in activations. To tackle this problem, we propose ResQ, a PTQ method that pushes further the state-of-the-art. By means of principal component analysis (PCA), it identifies a low-rank subspace (in practice 1/8 of the hidden dimension) in which activation variances are highest, and keep the coefficients within this subspace in high precision, e. g. 8-bit, while quantizing the rest to 4-bit. Within each subspace, invariant random rotation is applied to further suppress outliers. We show that this is a provably optimal mixed precision quantization scheme that minimizes error. With the Llama and Qwen2. 5 families of models, we demonstrate that ResQ outperforms recent uniform and mixed precision PTQ methods on a variety of benchmarks, achieving up to 33% lower perplexity on Wikitext than the next best method SpinQuant, and upto 3X speedup over 16-bit baseline. Anonymous code repository available at https: //anonymous. 4open. science/r/project-resq-2142.

NeurIPS Conference 2025 Conference Paper

Restricted Global-Aware Graph Filters Bridging GNNs and Transformer for Node Classification

  • Jingyuan Zhang
  • Xin Wang
  • Lei Yu
  • Zhirong Huang
  • Li Yang
  • Fengjun Zhang

Transformers have been widely regarded as a promising direction for breaking through the performance bottlenecks of Graph Neural Networks (GNNs), primarily due to their global receptive fields. However, a recent empirical study suggests that tuned classical GNNs can match or even outperform state-of-the-art Graph Transformers (GTs) on standard node classification benchmarks. Motivated by this fact, we deconstruct several representative GTs to examine how global attention components influence node representations. We find that the global attention module does not provide significant performance gains and may even exacerbate test error oscillations. Consequently, we consider that the Transformer is barely able to learn connectivity patterns that meaningfully complement the original graph topology. Interestingly, we further observe that mitigating such oscillations enables the Transformer to improve generalization in GNNs. In a nutshell, we reinterpret the Transformer through the lens of graph spectrum and reformulate it as a global-aware graph filter with band-pass characteristics and linear complexity. This unique perspective introduces multi-channel filtering constraints that effectively suppress test error oscillations. Extensive experiments (17 homophilous, heterophilous graphs) provide comprehensive empirical evidence for our perspective. This work clarifies the role of Transformers in GNNs and suggests that advancing modern GNN research may still require a return to the graph itself.

IJCAI Conference 2025 Conference Paper

RLMiniStyler: Light-weight RL Style Agent for Arbitrary Sequential Neural Style Generation

  • Jing Hu
  • Chengming Feng
  • Shu Hu
  • Ming-Ching Chang
  • Xin Li
  • Xi Wu
  • Xin Wang

Arbitrary style transfer aims to apply the style of any given artistic image to another content image. Still, existing deep learning-based methods often require significant computational costs to generate diverse stylized results. Motivated by this, we propose a novel reinforcement learning-based framework for arbitrary style transfer RLMiniStyler. This framework leverages a unified reinforcement learning policy to iteratively guide the style transfer process by exploring and exploiting stylization feedback, generating smooth sequences of stylized results while achieving model lightweight. Furthermore, we introduce an uncertainty-aware multi-task learning strategy that automatically adjusts loss weights to adapt to the content and style balance requirements at different training stages, thereby accelerating model convergence. Through a series of experiments across image various resolutions, we have validated the advantages of RLMiniStyler over other state-of-the-art methods in generating high-quality, diverse artistic image sequences at a lower cost. Codes are available at https: //github. com/fengxiaoming520/RLMiniStyler.

NeurIPS Conference 2025 Conference Paper

SafeVid: Toward Safety Aligned Video Large Multimodal Models

  • Yixu Wang
  • Jiaxin Song
  • Yifeng Gao
  • Xin Wang
  • Yang Yao
  • Yan Teng
  • Xingjun Ma
  • Yingchun Wang

As Video Large Multimodal Models (VLMMs) rapidly advance, their inherent complexity introduces significant safety challenges, particularly the issue of mismatched generalization where static safety alignments fail to transfer to dynamic video contexts. We introduce SafeVid, a framework designed to instill video-specific safety principles in VLMMs. SafeVid uniquely transfers robust textual safety alignment capabilities to the video domain by employing detailed textual video descriptions as an interpretive bridge, facilitating LLM-based rule-driven safety reasoning. This is achieved through a closed-loop system comprising: 1) generation of SafeVid-350K, a novel 350, 000-pair video-specific safety preference dataset; 2) targeted alignment of VLMMs using Direct Preference Optimization (DPO); and 3) comprehensive evaluation via our new SafeVidBench benchmark. Alignment with SafeVid-350K significantly enhances VLMM safety, with models like LLaVA-NeXT-Video demonstrating substantial improvements (e. g. , up to 42. 39%) on SafeVidBench. SafeVid provides critical resources and a structured approach, demonstrating that leveraging textual descriptions as a conduit for safety reasoning markedly improves the safety alignment of VLMMs in complex multimodal scenarios.

AAAI Conference 2025 Conference Paper

SCALM: Detecting Bad Practices in Smart Contracts Through LLMs

  • Zongwei Li
  • Xiaoqi Li
  • Wenkai Li
  • Xin Wang

As the Ethereum platform continues to mature and gain widespread usage, it is crucial to maintain high standards of smart contract writing practices. While bad practices in smart contracts may not directly lead to security issues, they do elevate the risk of encountering problems. Therefore, to understand and avoid these bad practices, this paper introduces the first systematic study of bad practices in smart contracts, delving into over 35 specific issues. Specifically, we propose a large language models (LLMs)-based framework, SCALM. It combines Step-Back Prompting and Retrieval-Augmented Generation (RAG) to effectively identify and address various bad practices. Our extensive experiments using multiple LLMs and datasets have shown that SCALM outperforms existing tools in detecting bad practices in smart contracts.

AAAI Conference 2025 Conference Paper

Set-Valued Sensitivity Analysis of Deep Neural Networks

  • Xin Wang
  • Feilong Wang
  • Xuegang (Jeff) Ban

This paper proposes a sensitivity analysis framework based on set-valued mapping for deep neural networks (DNN) to understand and compute how the solutions (model weights) of DNN respond to perturbations in the training data. As a DNN may not exhibit a unique solution (minima) and the algorithm of solving a DNN may lead to different solutions with minor perturbations to input data, we focus on the sensitivity of the solution set of DNN, instead of studying a single solution. In particular, we are interested in the expansion and contraction of the solution set in response to data perturbations. If the change of solution set can be bounded by the extent of the data perturbation, the model is said to exhibit the Lipschitz-like property. This 'set-to-set' analysis approach provides a deeper understanding of the robustness and reliability of DNNs during training. Our framework incorporates both isolated and non-isolated minima, and critically, does not require the assumption that the Hessian of loss function is non-singular. By developing set-level metrics such as distance between sets, convergence of sets, derivatives of set-valued mapping, and stability across the solution set, we prove that the solution set of the Fully Connected Neural Network holds Lipschitz-like properties. For general neural networks (e.g. Resnet), we introduce a graphical-derivative-based method to estimate the new solution set following data perturbation without retraining.

NeurIPS Conference 2025 Conference Paper

Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space

  • Zhen Zhang
  • Xuehai He
  • Weixiang Yan
  • Ao Shen
  • Chenyang Zhao
  • Xin Wang

Human cognition typically involves thinking through abstract, fluid concepts rather than strictly using discrete linguistic tokens. Current Large Language Models (LLMs), however, are constrained to reasoning within the boundaries of human language, processing discrete token embeddings that represent fixed points in semantic space. This discrete constraint restricts the expressive power and upper potential of such reasoning models, often causing incomplete exploration of reasoning paths, as standard Chain-of-Thought (CoT) methods rely on sampling one token per step. In this work, we introduce Soft Thinking, a training-free method that emulates human-like ``soft'' reasoning by generating abstract concept tokens in a continuous concept space. These concept tokens are created by the probability-weighted mixture of token embeddings, which span the continuous concept space, enabling smooth transitions and richer representations that transcend traditional discrete boundaries. In essence, each generated concept token encapsulates multiple meanings from related discrete tokens, implicitly exploring various reasoning paths to converge effectively toward the correct answer. Empirical evaluations on diverse mathematical and coding benchmarks consistently demonstrate the effectiveness and efficiency of Soft Thinking, improving pass@1 accuracy by up to 2. 48 points while simultaneously reducing token usage by up to 22. 4\% compared to standard CoT. Qualitative analysis further reveals that Soft Thinking outputs remain highly interpretable and readable, highlighting the potential of Soft Thinking to break the inherent limits of discrete language-based reasoning.

YNIMG Journal 2025 Journal Article

The power of pain: The temporal-spatial dynamics of empathy induced by body gestures and facial expressions

  • Xin Wang
  • Benjamin Becker
  • Shelley Xiuli Tong

Two non-verbal pain representations, body gestures and facial expressions, can communicate pain to others and elicit our own empathic responses. However, the specific impact of these representations on neural responses of empathy, particularly in terms of temporal and spatial neural mechanisms, remains unclear. To address this issue, the present study developed a kinetic pain empathy paradigm comprising short animated videos depicting a protagonist's "real life" pain and no-pain experiences through body gestures and facial expressions. Electroencephalographic (EEG) recordings were conducted on 52 neurotypical adults; while they viewed the animations. Results from multivariate pattern, event-related potential, event-related spectrum perturbation, and source localization analyses revealed that pain expressed through facial expressions, but not body gestures, elicited increased N200 and P200 responses and activated various brain regions, i.e., the anterior cingulate cortex, insula, thalamus, ventromedial prefrontal cortex, temporal gyrus, cerebellum, and right supramarginal gyrus. Enhanced theta power with distinct spatial distributions were observed during early affective arousal and late cognitive reappraisal stages of the pain event. Multiple regression analyses showed a negative correlation between the N200 amplitude and pain catastrophizing, and a positive correlation between the P200 amplitude and autism traits. These findings demonstrate the temporal evolution of empathy evoked by dynamic pain display, highlighting the significant impact of facial expression and its association with individuals' unique psychological traits.

ECAI Conference 2025 Conference Paper

Towards Mitigation of Hallucination for LLM-Empowered Agents: Progressive Generalization Bound Exploration and Watchdog Monitor

  • Siyuan Liu
  • Wenjing Liu
  • Zhiwei Xu
  • Xin Wang
  • Bo Chen
  • Tao Li

Empowered by large language models (LLMs), intelligent agents have become a popular paradigm for interacting with open environments to facilitate AI deployment. However, hallucinations generated by LLMs—where outputs are inconsistent with facts—pose a significant challenge, undermining the credibility of intelligent agents. Only if hallucinations can be mitigated, the intelligent agents can be used in real-world without any catastrophic risk. Therefore, effective detection and mitigation of hallucinations are crucial to ensure the dependability of agents. Unfortunately, the related approaches either depend on white-box access to LLMs or fail to accurately identify hallucinations. To address the challenge posed by hallucinations of intelligent agents, we present HalMit, a novel black-box watchdog framework that models the generalization bound of LLM-empowered agents and thus detect hallucinations without requiring internal knowledge of the LLM’s architecture. Specifically, a probabilistic fractal sampling technique is proposed to generate a sufficient number of queries to trigger the incredible responses in parallel, efficiently identifying the generalization bound of the target agent. Experimental evaluations demonstrate that HalMit significantly outperforms existing approaches in hallucination monitoring. Its black-box nature and superior performance make HalMit a promising solution for enhancing the dependability of LLM-powered systems.

ICRA Conference 2025 Conference Paper

Universal Online Temporal Calibration for Optimization-Based Visual-Inertial Navigation Systems

  • Yunfei Fan 0001
  • Tianyu Zhao
  • Linan Guo
  • Chen Chen 0129
  • Xin Wang
  • Fengyi Zhou

6-Degree of Freedom (6DoF) motion estimation with a combination of visual and inertial sensors is a growing area with numerous real-world applications. However, precise calibration of the time offset between these two sensor types is a prerequisite for accurate and robust tracking. To address this, we propose a universal online temporal calibration strategy for optimization-based visual-inertial navigation systems. Technically, we incorporate the time offset $t_{d}$ as a state parameter in the optimization residual model to align the IMU state to the corresponding image timestamp using $t_{d}$, angular velocity and translational velocity. This allows the temporal misalignment $t_{d}$ to be optimized alongside other tracking states during the process. As our method only modifies the structure of the residual model, it can be applied to various optimization-based frameworks with different tracking frontends. We evaluate our calibration method with both EuRoC [1] and simulation data and extensive experiments demonstrate that our approach provides more accurate time offset estimation and faster convergence, particularly in the presence of noisy sensor data. The experimental code is available at https://github.com/bytedance/Ts_Online_Optimization.

JBHI Journal 2024 Journal Article

A SwinTransformer-Based Segmentation Framework With Self-Supervised Strategy for Post-Operative Prostate Cancer Radiotherapy

  • Dong Miao
  • Jielang Li
  • Meng Dou
  • Linjie Fu
  • Yu Yao
  • Xin Wang
  • Feng Wen
  • Yali Shen

Radical prostatectomy (prostate removal) is a standard treatment for clinically localized prostate cancer and is often followed by postoperative radiotherapy. Postoperative radiotherapy requires accurate delineation of the clinical target volume (CTV) and lymph node drainage area (LNA) on computed tomography (CT) images. However, the CTV contour cannot be determined by the simple prostate expansion after resection of the prostate in the CT image. Constrained by this factor, the manual delineation process in postoperative radiotherapy is more time-consuming and challenging than in radical radiotherapy. In addition, CTV and LNA have no boundaries that can be distinguished by pixel values in CT images, and existing automatic segmentation models cannot get satisfactory results. Radiation oncologists generally determine CTV and LNA profiles according to clinical consensus and guidelines regarding surrounding organs at risk (OARs). In this work, we design a cascade segmentation block to explicitly establish correlations between CTV, LNA, and OARs, leveraging OARs features to guide CTV and LNA segmentation. Furthermore, inspired by the success of the self-attention mechanism and self-supervised learning, we adopt SwinTransformer as our backbone and propose a pure SwinTransformer-based segmentation network with self-supervised learning strategies. We performed extensive quantitative and qualitative evaluations of the proposed method. Compared to other competitive segmentation models, our model shows higher dice scores with minor standard deviations, and the detailed visualization results are more consistent with the ground truth. We believe this work can provide a feasible solution to this problem, making the postoperative radiotherapy process more efficient.

NeurIPS Conference 2024 Conference Paper

Causal language modeling can elicit search and reasoning capabilities on logic puzzles

  • Kulin Shah
  • Nishanth Dikkala
  • Xin Wang
  • Rina Panigrahy

Causal language modeling using the Transformer architecture has yielded remarkable capabilities in Large Language Models (LLMs) over the last few years. However, the extent to which fundamental search and reasoning capabilities emerged within LLMs remains a topic of ongoing debate. In this work, we study if causal language modeling can learn a complex task such as solving Sudoku puzzles. To solve a Sudoku, the model is first required to search over all empty cells of the puzzle to decide on a cell to fill and then apply an appropriate strategy to fill the decided cell. Sometimes, the application of a strategy only results in thinning down the possible values in a cell rather than concluding the exact value of the cell. In such cases, multiple strategies are applied one after the other to fill a single cell. We observe that Transformer models trained on this synthetic task can indeed learn to solve Sudokus (our model solves $94. 21\%$ of the puzzles fully correctly) when trained on a logical sequence of steps taken by a solver. We find that training Transformers with the logical sequence of steps is necessary and without such training, they fail to learn Sudoku. We also extend our analysis to Zebra puzzles (known as Einstein puzzles) and show that the model solves $92. 04 \%$ of the puzzles fully correctly. In addition, we study the internal representations of the trained Transformer and find that through linear probing, we can decode information about the set of possible values in any given cell from them, pointing to the presence of a strong reasoning engine implicit in the Transformer weights.

EAAI Journal 2024 Journal Article

Combining optical flow and Swin Transformer for Space-Time video super-resolution

  • Xin Wang
  • Hua Wang
  • Mingli Zhang
  • Fan Zhang

Space–time video super-resolution is a task that aims to interpolate low frame rate, low resolution videos to high frame rate, high resolution ones. While existing Transformer-based methods have achieved results comparable to convolutional neural networks-based methods, the computational cost of Transformer limits its performance with constrained computational resources. Moreover, Swin Transformer may fail to fully exploit the spatio-temporal information of video frames due to the limitation of window size, impeding its effectiveness in handling large motions. To address these limitations, we propose an end-to-end space–time video super-resolution architecture based on optical flow alignment and Swin Transformer. The alignment module is introduced to extract spatio-temporal information from adjacent frames without significantly increasing the computational burden. Additionally, we design a residual convolution layer to enhance the translational invariance of the features extracted by Swin Transformer and introduces additional nonlinear transformations. Experimental results demonstrate that our proposed method achieves superior performance on various benchmark datasets compared to state-of-the-art methods. In terms of Peak Signal-to-Noise Ratio, our method outperforms the state-of-the-art methods by at least 0. 15 dB on Vimeo-Medium dataset.

AAAI Conference 2024 Conference Paper

Data-Augmented Curriculum Graph Neural Architecture Search under Distribution Shifts

  • Yang Yao
  • Xin Wang
  • Yijian Qin
  • Ziwei Zhang
  • Wenwu Zhu
  • Hong Mei

Graph neural architecture search (NAS) has achieved great success in designing architectures for graph data processing.However, distribution shifts pose great challenges for graph NAS, since the optimal searched architectures for the training graph data may fail to generalize to the unseen test graph data. The sole prior work tackles this problem by customizing architectures for each graph instance through learning graph structural information, but failed to consider data augmentation during training, which has been proven by existing works to be able to improve generalization.In this paper, we propose Data-augmented Curriculum Graph Neural Architecture Search (DCGAS), which learns an architecture customizer with good generalizability to data under distribution shifts. Specifically, we design an embedding-guided data generator, which can generate sufficient graphs for training to help the model better capture graph structural information. In addition, we design a two-factor uncertainty-based curriculum weighting strategy, which can evaluate the importance of data in enabling the model to learn key information in real-world distribution and reweight them during training. Experimental results on synthetic datasets and real datasets with distribution shifts demonstrate that our proposed method learns generalizable mappings and outperforms existing methods.

NeurIPS Conference 2024 Conference Paper

Differentiable Structure Learning with Partial Orders

  • Taiyu Ban
  • Lyuzhou Chen
  • Xiangyu Wang
  • Xin Wang
  • Derui Lyu
  • Huanhuan Chen

Differentiable structure learning is a novel line of causal discovery research that transforms the combinatorial optimization of structural models into a continuous optimization problem. However, the field has lacked feasible methods to integrate partial order constraints, a critical prior information typically used in real-world scenarios, into the differentiable structure learning framework. The main difficulty lies in adapting these constraints, typically suited for the space of total orderings, to the continuous optimization context of structure learning in the graph space. To bridge this gap, this paper formalizes a set of equivalent constraints that map partial orders onto graph spaces and introduces a plug-and-play module for their efficient application. This module preserves the equivalent effect of partial order constraints in the graph space, backed by theoretical validations of correctness and completeness. It significantly enhances the quality of recovered structures while maintaining good efficiency, which learns better structures using 90\% fewer samples than the data-based method on a real-world dataset. This result, together with a comprehensive evaluation on synthetic cases, demonstrates our method's ability to effectively improve differentiable structure learning with partial orders.

TMLR Journal 2024 Journal Article

Efficient Large Language Models: A Survey

  • Zhongwei Wan
  • Xin Wang
  • Che Liu
  • Samiul Alam
  • Yu Zheng
  • Jiachen Liu
  • Zhongnan Qu
  • Shen Yan

Large Language Models (LLMs) have demonstrated remarkable capabilities in important tasks such as natural language understanding and language generation, and thus have the potential to make a substantial impact on our society. Such capabilities, however, come with the considerable resources they demand, highlighting the strong need to develop effective techniques for addressing their efficiency challenges. In this survey, we provide a systematic and comprehensive review of efficient LLMs research. We organize the literature in a taxonomy consisting of three main categories, covering distinct yet interconnected efficient LLMs topics from model-centric, data-centric, and framework-centric perspective, respectively. We have also created a GitHub repository where we organize the papers featured in this survey at https://github.com/AIoT-MLSys-Lab/Efficient-LLMs-Survey. We will actively maintain the repository and incorporate new research as it emerges. We hope our survey can serve as a valuable resource to help researchers and practitioners gain a systematic understanding of efficient LLMs research and inspire them to contribute to this important and exciting field.

JBHI Journal 2024 Journal Article

Evolutionary Ensemble Learning for EEG-Based Cross-Subject Emotion Recognition

  • Hanzhong Zhang
  • Tienyu Zuo
  • Zhiyang Chen
  • Xin Wang
  • Poly Z.H. Sun

Electroencephalogram (EEG) has been widely utilized in emotion recognition due to its high temporal resolution and reliability. However, the individual differences and non-stationary characteristics of EEG, along with the complexity and variability of emotions, pose challenges in generalizing emotion recognition models across subjects. In this paper, an end-to-end framework is proposed to improve the performance of cross-subject emotion recognition. A novel evolutionary programming (EP)-based optimization strategy with neural network (NN) as the base classifier termed NN ensemble with EP (EPNNE) is designed for cross-subject emotion recognition. The effectiveness of the proposed method is evaluated on the publicly available DEAP, FACED, SEED, and SEED-IV datasets. Numerical results demonstrate that the proposed method is superior to state-of-the-art cross-subject emotion recognition methods. The proposed end-to-end framework for cross-subject emotion recognition aids biomedical researchers in effectively assessing individual emotional states, thereby enabling efficient treatment and interventions.

AAAI Conference 2024 Conference Paper

Exponential Hardness of Optimization from the Locality in Quantum Neural Networks

  • Hao-Kai Zhang
  • Chengkai Zhu
  • Geng Liu
  • Xin Wang

Quantum neural networks (QNNs) have become a leading paradigm for establishing near-term quantum applications in recent years. The trainability issue of QNNs has garnered extensive attention, spurring demand for a comprehensive analysis of QNNs in order to identify viable solutions. In this work, we propose a perspective that characterizes the trainability of QNNs based on their locality. We prove that the entire variation range of the loss function via adjusting any local quantum gate vanishes exponentially in the number of qubits with a high probability for a broad class of QNNs. This result reveals extra harsh constraints independent of gradients and unifies the restrictions on gradient-based and gradient-free optimizations naturally. We showcase the validity of our results with numerical simulations of representative models and examples. Our findings, as a fundamental property of random quantum circuits, deepen the understanding of the role of locality in QNNs and serve as a guideline for assessing the effectiveness of diverse training strategies for quantum neural networks.

NeurIPS Conference 2024 Conference Paper

FUG: Feature-Universal Graph Contrastive Pre-training for Graphs with Diverse Node Features

  • Jitao Zhao
  • Di Jin
  • Meng Ge
  • Lianze Shan
  • Xin Wang
  • Dongxiao He
  • Zhiyong Feng

Graph Neural Networks (GNNs), known for their effective graph encoding, are extensively used across various fields. Graph self-supervised pre-training, which trains GNN encoders without manual labels to generate high-quality graph representations, has garnered widespread attention. However, due to the inherent complex characteristics in graphs, GNNs encoders pre-trained on one dataset struggle to directly adapt to others that have different node feature shapes. This typically necessitates either model rebuilding or data alignment. The former results in non-transferability as each dataset need to rebuild a new model, while the latter brings serious knowledge loss since it forces features into a uniform shape by preprocessing such as Principal Component Analysis (PCA). To address this challenge, we propose a new Feature-Universal Graph contrastive pre-training strategy (FUG) that naturally avoids the need for model rebuilding and data reshaping. Specifically, inspired by discussions in existing work on the relationship between contrastive Learning and PCA, we conducted a theoretical analysis and discovered that PCA's optimization objective is a special case of that in contrastive Learning. We designed an encoder with contrastive constraints to emulate PCA's generation of basis transformation matrix, which is utilized to losslessly adapt features in different datasets. Furthermore, we introduced a global uniformity constraint to replace negative sampling, reducing the time complexity from $O(n^2)$ to $O(n)$, and by explicitly defining positive samples, FUG avoids the substantial memory requirements of data augmentation. In cross domain experiments, FUG has a performance close to the re-trained new models. The source code is available at: https: //github. com/hedongxiao-tju/FUG.

NeurIPS Conference 2024 Conference Paper

Gorilla: Large Language Model Connected with Massive APIs

  • Shishir G. Patil
  • Tianjun Zhang
  • Xin Wang
  • Joseph E. Gonzalez

Large Language Models (LLMs) have seen an impressive wave of advances, withmodels now excelling in a variety of tasks, such as mathematical reasoning andprogram synthesis. However, their potential to effectively use tools via API callsremains unfulfilled. This is a challenging task even for today’s state-of-the-artLLMs such as GPT-4 largely due to their unawareness of what APIs are availableand how to use them in a frequently updated tool set. We develop Gorilla, afinetuned LLaMA model that surpasses the performance of GPT-4 on writing APIcalls. Trained with the novel Retriever Aware Training (RAT), when combinedwith a document retriever, Gorilla demonstrates a strong capability to adapt totest-time document changes, allowing flexible user updates or version changes. It also substantially mitigates the issue of hallucination, commonly encounteredwhen prompting LLMs directly. To evaluate the model’s ability, we introduceAPIBench, a comprehensive dataset consisting of HuggingFace, TorchHub, andTensorHub APIs. The successful integration of the retrieval system with Gorillademonstrates the potential for LLMs to use tools more accurately, keep up withfrequently updated documentation, and consequently increase the reliability andapplicability of their outputs. Gorilla’s code, model, data, and demo are availableat: https: //gorilla. cs. berkeley. edu

ICRA Conference 2024 Conference Paper

Investigation on the multi-solution problem of the kinetostatics of cable-driven continuum manipulators

  • Yicheng Dai
  • Zuan Li
  • Xin Wang
  • Han Yuan

Cable-driven continuum manipulators have gained considerable attention due to their high dexterity and inherent structural compliance, making them a popular research topic. However, previous studies have overlooked the kinetostatics of these manipulators, which can result in a multi-solution problem. This issue is critical as having multiple equilibrium states can lead to erroneous estimations of the manipulator's profile. To address this issue, the kinetostatic model is presented and simulations based on both the interval analysis method and the commonly used floating-point optimization algorithm are conducted under the same actuating forces and external loads. Results show that there are multiple solutions to the kinetostatics of cable-driven continuum manipulators with constant cross section or variable cross section. This paper fills a gap in the current literature and offers valuable insights for researchers in the field of cable-driven continuum manipulators.

EAAI Journal 2024 Journal Article

IPNet: Polarization-based Camouflaged Object Detection via dual-flow network

  • Xin Wang
  • Jiajia Ding
  • Zhao Zhang
  • Junfeng Xu
  • Jun Gao

Camouflaged Object Detection (COD) is a critical task in a variety of domains, such as medicine and military applications. The main challenge in COD is accurately detecting and extracting the concealed object from the complex background. The similarity between the camouflaged objects and their background significantly reduces the accuracy of object extraction. Polarization information can provide valuable insights into the characteristics of objects with different material properties and surface roughness. It reflects the difference in polarization information between the object and the background, which increases the contrast between the two and improves the object detection accuracy even under complex scenes. In this paper, we propose IPNet, an efficient cross-modal fusion network that utilizes both RGB intensity and linear polarization cues to generate scene representation with high contrast. Our novel network architecture dynamically fuses RGB intensity and polarization cues using an efficient cross-modal fusion module, leveraging cross-level contextual information to achieve robust detection. For training and evaluating the proposed network, we construct a polarization-based PCOD_1200 dataset that contains 89 subclasses and 1200 samples. A comprehensive set of experiments demonstrates the effectiveness of IPNet to fuse polarization and RGB intensity information and shows that our approach outperforms state-of-the-art methods.

NeurIPS Conference 2024 Conference Paper

Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models

  • Jiayu Wang
  • Yifei Ming
  • Zhenmei Shi
  • Vibhav Vineet
  • Xin Wang
  • Yixuan Li
  • Neel Joshi

Large language models (LLMs) and vision-language models (VLMs) have demonstrated remarkable performance across a wide range of tasks and domains. Despite this promise, spatial understanding and reasoning—a fundamental component of human cognition—remains under-explored. We propose SpatialEval, a novel benchmark that covers diverse aspects of spatial reasoning such as relationship understanding, navigation, and counting. We conduct a comprehensive evaluation of competitive language and vision-language models. Our findings reveal several counter-intuitive insights that have been overlooked in the literature: (1) Spatial reasoning poses significant challenges where competitive models can fall behind random guessing; (2) Despite additional visual input, VLMs often under-perform compared to their LLM counterparts; (3) When both textual and visual information is available, multi-modal language models become less reliant on visual information if sufficient textual clues are provided. Additionally, we demonstrate that leveraging redundancy between vision and text can significantly enhance model performance. We hope our study will inform the development of multimodal models to improve spatial intelligence and further close the gap with human intelligence. Our code is available at https: //github. com/jiayuww/SpatialEval.

AIJ Journal 2024 Journal Article

Mitigating social biases of pre-trained language models via contrastive self-debiasing with double data augmentation

  • Yingji Li
  • Mengnan Du
  • Rui Song
  • Xin Wang
  • Mingchen Sun
  • Ying Wang

Pre-trained Language Models (PLMs) have been shown to inherit and even amplify the social biases contained in the training corpus, leading to undesired stereotype in real-world applications. Existing techniques for mitigating the social biases of PLMs mainly rely on data augmentation with manually designed prior knowledge or fine-tuning with abundant external corpora to debias. However, these methods are not only limited by artificial experience, but also consume a lot of resources to access all the parameters of the PLMs and are prone to introduce new external biases when fine-tuning with external corpora. In this paper, we propose a Contrastive Self-Debiasing Model with Double Data Augmentation (named CD3) for mitigating social biases of PLMs. Specifically, CD3 consists of two stages: double data augmentation and contrastive self-debiasing. First, we build on counterfactual data augmentation to perform a secondary augmentation using biased prompts that are automatically searched by maximizing the differences in PLMs' encoding across demographic groups. Double data augmentation further amplifies the biases between sample pairs to break the limitations of previous debiasing models that heavily rely on prior knowledge in data augmentation. We then leverage the augmented data for contrastive learning to train a plug-and-play adapter to mitigate the social biases in PLMs' encoding without tuning the PLMs. Extensive experimental results on BERT, ALBERT, and RoBERTa on several real-world datasets and fairness metrics show that CD3 outperforms baseline models on gender debiasing and race debiasing while retaining the language modeling capabilities of PLMs.

AAAI Conference 2024 Conference Paper

Multimodal Graph Neural Architecture Search under Distribution Shifts

  • Jie Cai
  • Xin Wang
  • Haoyang Li
  • Ziwei Zhang
  • Wenwu Zhu

Multimodal graph neural architecture search (MGNAS) has shown great success for automatically designing the optimal multimodal graph neural network (MGNN) architecture by leveraging multimodal representation, crossmodal information and graph structure in one unified framework. However, existing MGNAS fails to handle distribution shifts that naturally exist in multimodal graph data, since the searched architectures inevitably capture spurious statistical correlations under distribution shifts. To solve this problem, we propose a novel Out-of-distribution Generalized Multimodal Graph Neural Architecture Search (OMG-NAS) method which optimizes the MGNN architecture with respect to its performance on decorrelated OOD data. Specifically, we propose a multimodal graph representation decorrelation strategy, which encourages the searched MGNN model to output representations that eliminate spurious correlations through iteratively optimizing the feature weights and controller. In addition, we propose a global sample weight estimator that facilitates the sharing of optimal sample weights learned from existing architectures. This design promotes the effective estimation of the sample weights for candidate MGNN architectures to generate decorrelated multimodal graph representations, concentrating more on the truly predictive relations between invariant features and ground-truth labels. Extensive experiments on real-world multimodal graph datasets demonstrate the superiority of our proposed method over SOTA baselines.

NeurIPS Conference 2024 Conference Paper

Non-asymptotic Approximation Error Bounds of Parameterized Quantum Circuits

  • Zhan Yu
  • Qiuhao Chen
  • Yuling Jiao
  • Yinan Li
  • Xiliang Lu
  • Xin Wang
  • Jerry Z. Yang

Understanding the power of parameterized quantum circuits (PQCs) in accomplishing machine learning tasks is one of the most important questions in quantum machine learning. In this paper, we focus on the PQC expressivity for general multivariate function classes. Previously established Universal Approximation Theorems for PQCs are either nonconstructive or assisted with parameterized classical data processing, making it hard to justify whether the expressive power comes from the classical or quantum parts. We explicitly construct data re-uploading PQCs for approximating multivariate polynomials and smooth functions and establish the first non-asymptotic approximation error bounds for such functions in terms of the number of qubits, the quantum circuit depth and the number of trainable parameters of the PQCs. Notably, we show that for multivariate polynomials and multivariate smooth functions, the quantum circuit size and the number of trainable parameters of our proposed PQCs can be smaller than the deep ReLU neural networks. We further demonstrate the approximation capability of PQCs via numerical experiments. Our results pave the way for designing practical PQCs that can be implemented on near-term quantum devices with limited resources.

IJCAI Conference 2024 Conference Paper

PrivSGP-VR: Differentially Private Variance-Reduced Stochastic Gradient Push with Tight Utility Bounds

  • Zehan Zhu
  • Yan Huang
  • Xin Wang
  • Jinming Xu

In this paper, we propose a differentially private decentralized learning method (termed PrivSGP-VR) which employs stochastic gradient push with variance reduction and guarantees (epsilon, delta)-differential privacy (DP) for each node. Our theoretical analysis shows that, under DP Gaussian noise with constant variance, PrivSGP-VR achieves a sub-linear convergence rate of O(1/sqrt(nK)), where n and K are the number of nodes and iterations, respectively, which is independent of stochastic gradient variance, and achieves a linear speedup with respect to n. Leveraging the moments accountant method, we further derive an optimal K to maximize the model utility under certain privacy budget in decentralized settings. With this optimized K, PrivSGP-VR achieves a tight utility bound of O(sqrt(d*log(1/delta))/(sqrt(n)*J*epsilon)), where J and d are the number of local samples and the dimension of decision variable, respectively, which matches that of the server-client distributed counterparts, and exhibits an extra factor of 1/sqrt(n) improvement compared to that of the existing decentralized counterparts, such as A(DP)2SGD. Extensive experiments corroborate our theoretical findings, especially in terms of the maximized utility with optimized K, in fully decentralized settings.

AAAI Conference 2024 Conference Paper

Rethinking Propagation for Unsupervised Graph Domain Adaptation

  • Meihan Liu
  • Zeyu Fang
  • Zhen Zhang
  • Ming Gu
  • Sheng Zhou
  • Xin Wang
  • Jiajun Bu

Unsupervised Graph Domain Adaptation (UGDA) aims to transfer knowledge from a labelled source graph to an unlabelled target graph in order to address the distribution shifts between graph domains. Previous works have primarily focused on aligning data from the source and target graph in the representation space learned by graph neural networks (GNNs). However, the inherent generalization capability of GNNs has been largely overlooked. Motivated by our empirical analysis, we reevaluate the role of GNNs in graph domain adaptation and uncover the pivotal role of the propagation process in GNNs for adapting to different graph domains. We provide a comprehensive theoretical analysis of UGDA and derive a generalization bound for multi-layer GNNs. By formulating GNN Lipschitz for k-layer GNNs, we show that the target risk bound can be tighter by removing propagation layers in source graph and stacking multiple propagation layers in target graph. Based on the empirical and theoretical analysis mentioned above, we propose a simple yet effective approach called A2GNN for graph domain adaptation. Through extensive experiments on real-world datasets, we demonstrate the effectiveness of our proposed A2GNN framework.

JBHI Journal 2024 Journal Article

Robust Epileptic Seizure Detection Based on Biomedical Signals Using an Advanced Multi-View Deep Feature Learning Approach

  • Ijaz Ahmad
  • Zhenzhen Liu
  • Lin Li
  • Inam Ullah
  • Sunday Timothy Aboyeji
  • Xin Wang
  • Oluwarotimi Williams Samuel
  • Guanglin Li

Epilepsy is a neurological disorder characterized by abnormal neuronal discharges that manifest in life-threatening seizures. These are often monitored via EEG signals, a key aspect of biomedical signal processing (BSP). Accurate epileptic seizure (ES) detection significantly depends on the precise identification of key EEG features, which requires a deep understanding of the data's intrinsic domain. Therefore, this study presents an Advanced Multi-View Deep Feature Learning (AMV-DFL) framework based on machine learning (ML) technology to enhance the detection of relevant EEG signal features for ES. Our method initially applies a fast Fourier transform (FFT) on EEG data for traditional frequency domain feature (TFD-F) extraction and directly incorporates time domain (TD) features from the raw EEG signals, establishing a comprehensive traditional multi-view feature (TMV-F). Deep features are subsequently extracted autonomously from optimal layers of one-dimensional convolutional neural networks (1D CNN), resulting in multi-view deep features (MV-DF) integrating both time and frequency domains. A multi-view forest (MV-F) is an interpretable rule-based advanced ML classifier used to construct a robust, generalized classification. Tree-based SHAP explainable artificial intelligence (T-XAI) is incorporated for interpreting and explaining the underlying rules. Experimental results confirm our method's superiority, surpassing models using TMV-FL and single-view deep features (SV-DF) by 4% and outperforming other state-of-the-art methods by an average of 3% in classification accuracy. The AMV-DFL approach aids clinicians in identifying EEG features indicative of ES, potentially discovering novel biomarkers, and improving diagnostic capabilities in epilepsy management.

IJCAI Conference 2024 Conference Paper

Self-Supervised Learning for Enhancing Spatial Awareness in Free-Hand Sketches

  • Xin Wang
  • Tengjie Li
  • Sicong Zang
  • Shikui Tu
  • Lei Xu

Free-hand sketch, as a versatile medium of communication, can be viewed as a collection of strokes arranged in a spatial layout to convey a concept. Due to the abstract nature of the sketches, changes in stroke position may make them difficult to recognize. Recently, Graphic sketch representations are effective in representing sketches. However, existing methods overlook the significance of the spatial layout of strokes and the phenomenon of strokes being drawn in the wrong positions is common. Therefore, we developed a self-supervised task to correct stroke placement and investigate the impact of spatial layout on learning sketch representations. For this task, we propose a spatially aware method, named SketchGloc, utilizing multiple graphs for graphic sketch representations. This method utilizes grids for each stroke to describe the spatial layout with other strokes, allowing for the construction of multiple graphs. Unlike other methods that rely on a single graph, this design conveys more detailed spatial layout information and alleviates the impact of misplaced strokes. The experimental results demonstrate that our model outperforms existing methods in both our proposed task and the traditional controllable sketch synthesis task. Additionally, we found that SketchGloc can learn more robust representations under our proposed task setting. The source code is available at https: //github. com/CMACH508/SketchGloc.

YNICL Journal 2024 Journal Article

Static and dynamic interactions within the triple-network model in stroke patients with multidomain cognitive impairments

  • Yingying Wang
  • Hongxu Chen
  • Caihong Wang
  • Jingchun Liu
  • Peifang Miao
  • Ying Wei
  • Luobing Wu
  • Xin Wang

BACKGROUND: Internal capsule strokes often result in multidomain cognitive impairments across memory, attention, and executive function, typically due to disruptions in brain network connectivity. Our study examines these impairments by analyzing interactions within the triple-network model, focusing on both static and dynamic aspects. METHODS: We collected resting-state fMRI data from 62 left (CI_L) and 56 right (CI_R) internal capsule stroke patients, along with 57 healthy controls (HC). Using independent component analysis to extract the default mode (DMN), executive control (ECN), and salience networks (SAN), we conducted static and dynamic functional network connectivity analyses (DFNC) to identify differences between stroke patients and controls. For DFNC, we used k-means clustering to focus on temporal properties and multilayer network analysis to examine integration and modularity Q, where integration represents dynamic interactions between networks, and modularity Q measures how well the network is divided into distinct modules. We then calculated the correlations between SFNC/DFNC properties with significant inter-group differences and cognitive scales. RESULTS: Compared to HC, both CI_L and CI_R patients showed increased static FCs between SAN and DMN and decreased dynamic interactions between ECN and other networks. CI_R patients also had heightened static FCs between SAN and ECN and maintained a state with strongly positive FNCs across all networks in the triple-network model. Additionally, CI_R patients displayed decreased modularity Q. CONCLUSION: These findings highlight that stroke can result in the disruption of static and dynamic interactions in the triple network model, aiding our understanding of the neuropathological basis for multidomain cognitive deficits after internal capsule stroke.

YNIMG Journal 2024 Journal Article

The heritability and structural correlates of resting-state fMRI complexity

  • Yi Zhen
  • Yaqian Yang
  • Yi Zheng
  • Xin Wang
  • Longzhao Liu
  • Zhiming Zheng
  • Hongwei Zheng
  • Shaoting Tang

The complexity of fMRI signals quantifies temporal dynamics of spontaneous neural activity, which has been increasingly recognized as providing important insights into cognitive functions and psychiatric disorders. However, its heritability and structural underpinnings are not well understood. Here, we utilize multi-scale sample entropy to extract resting-state fMRI complexity in a large healthy adult sample from the Human Connectome Project. We show that fMRI complexity at multiple time scales is heritable in broad brain regions. Heritability estimates are modest and regionally variable. We relate fMRI complexity to brain structure including surface area, cortical myelination, cortical thickness, subcortical volumes, and total brain volume. We find that surface area is negatively correlated with fine-scale complexity and positively correlated with coarse-scale complexity in most cortical regions, especially the association cortex. Most of these correlations are related to common genetic and environmental effects. We also find positive correlations between cortical myelination and fMRI complexity at fine scales and negative correlations at coarse scales in the prefrontal cortex, lateral temporal lobe, precuneus, lateral parietal cortex, and cingulate cortex, with these correlations mainly attributed to common environmental effects. We detect few significant associations between fMRI complexity and cortical thickness. Despite the non-significant association with total brain volume, fMRI complexity exhibits significant correlations with subcortical volumes in the hippocampus, cerebellum, putamen, and pallidum at certain scales. Collectively, our work establishes the genetic basis and structural correlates of resting-state fMRI complexity across multiple scales, supporting its potential application as an endophenotype for psychiatric disorders.

NeurIPS Conference 2024 Conference Paper

VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding

  • Houlun Chen
  • Xin Wang
  • Hong Chen
  • Zeyang Zhang
  • Wei Feng
  • Bin Huang
  • Jia Jia
  • Wenwu Zhu

Existing Video Corpus Moment Retrieval (VCMR) is limited to coarse-grained understanding that hinders precise video moment localization when given fine-grained queries. In this paper, we propose a more challenging fine-grained VCMR benchmark requiring methods to localize the best-matched moment from the corpus with other partially matched candidates. To improve the dataset construction efficiency and guarantee high-quality data annotations, we propose VERIFIED, an automatic \underline{V}id\underline{E}o-text annotation pipeline to generate captions with \underline{R}el\underline{I}able \underline{FI}n\underline{E}-grained statics and \underline{D}ynamics. Specifically, we resort to large language models (LLM) and large multimodal models (LMM) with our proposed Statics and Dynamics Enhanced Captioning modules to generate diverse fine-grained captions for each video. To filter out the inaccurate annotations caused by the LLM hallucination, we propose a Fine-Granularity Aware Noise Evaluator where we fine-tune a video foundation model with disturbed hard-negatives augmented contrastive and matching losses. With VERIFIED, we construct a more challenging fine-grained VCMR benchmark containing Charades-FIG, DiDeMo-FIG, and ActivityNet-FIG which demonstrate a high level of annotation quality. We evaluate several state-of-the-art VCMR models on the proposed dataset, revealing that there is still significant scope for fine-grained video understanding in VCMR.

JBHI Journal 2024 Journal Article

WDFF-Net: Weighted Dual-Branch Feature Fusion Network for Polyp Segmentation With Object-Aware Attention Mechanism

  • Jie Cao
  • Xin Wang
  • Zhiwei Qu
  • Li Zhuo
  • Xiaoguang Li
  • Hui Zhang
  • Yang Yang
  • Wei Wei

Colon polyps in colonoscopy images exhibit significant differences in color, size, shape, appearance, and location, posing significant challenges to accurate polyp segmentation. In this paper, a Weighted Dual-branch Feature Fusion Network is proposed for Polyp Segmentation, named WDFF-Net, which adopts HarDNet68 as the backbone network. First, a dual-branch feature fusion network architecture is constructed, which includes a shared feature extractor and two feature fusion branches, i. e. Progressive Feature Fusion (PFF) branch and Scale-aware Feature Fusion (SFF) branch. The branches fuse the deep features of multiple layers for different purposes and with different fusion ways. The PFF branch is to address the under-segmentation or over-segmentation problems of flat polyps with low-edge contrast by iteratively fusing the features from low, medium, and high layers. The SFF branch is to tackle the the problem of drastic variations in polyp size and shape, especially the missed segmentation problem for small polyps. These two branches are complementary and play different roles, in improving segmentation accuracy. Second, an Object-aware Attention Mechanism (OAM) is proposed to enhance the features of the target regions and suppress those of the background regions, to interfere with the segmentation performance. Third, a weighted dual-branch the segmentation loss function is specifically designed, which dynamically assigns the weight factors of the loss functions for two branches to optimize their collaborative training. Experimental results on five public colon polyp datasets demonstrate that, the proposed WDFF-Net can achieve a superior segmentation performance with lower model complexity and faster inference speed, while maintaining good generalization ability.

NeurIPS Conference 2024 Conference Paper

WelQrate: Defining the Gold Standard in Small Molecule Drug Discovery Benchmarking

  • Yunchao Liu
  • Ha Dong
  • Xin Wang
  • Rocco Moretti
  • Yu Wang
  • Zhaoqian Su
  • Jiawei Gu
  • Bobby Bodenheimer

While deep learning has revolutionized computer-aided drug discovery, the AI community has predominantly focused on model innovation and placed less emphasis on establishing best benchmarking practices. We posit that without a sound model evaluation framework, the AI community's efforts cannot reach their full potential, thereby slowing the progress and transfer of innovation into real-world drug discovery. Thus, in this paper, we seek to establish a new gold standard for small molecule drug discovery benchmarking, WelQrate. Specifically, our contributions are threefold: WelQrate dataset collection - we introduce a meticulously curated collection of 9 datasets spanning 5 therapeutic target classes. Our hierarchical curation pipelines, designed by drug discovery experts, go beyond the primary high-throughput screen by leveraging additional confirmatory and counter screens along with rigorous domain-driven preprocessing, such as Pan-Assay Interference Compounds (PAINS) filtering, to ensure the high-quality data in the datasets; WelQrate Evaluation Framework - we propose a standardized model evaluation framework considering high-quality datasets, featurization, 3D conformation generation, evaluation metrics, and data splits, which provides a reliable benchmarking for drug discovery experts conducting real-world virtual screening; Benchmarking - we evaluate model performance through various research questions using the WelQrate dataset collection, exploring the effects of different models, dataset quality, featurization methods, and data splitting strategies on the results. In summary, we recommend adopting our proposed WelQrate as the gold standard in small molecule drug discovery benchmarking. The WelQrate dataset collection, along with the curation codes, and experimental scripts are all publicly available at www. WelQrate. org.

EAAI Journal 2023 Journal Article

A dynamic spectrum loss generative adversarial network for intelligent fault diagnosis with imbalanced data

  • Xin Wang
  • Hongkai Jiang
  • Yunpeng Liu
  • Shaowei Liu
  • Qiao Yang

Intelligent fault diagnosis with imbalanced data is a problem that often raises concerns. The diagnosis is more effective when the imbalanced dataset is supplemented with data augmentation methods, but there is always a gap between the real data and the generated data, especially in the frequency domain. Therefore, a dynamic spectrum loss generative adversarial network (DSLGAN) is developed for intelligent fault diagnosis. Firstly, a generative information enhancement module is built to simultaneously enhance inefficient information of the generative network from different information sources, thus creating a stable and efficient environment for the generation. Secondly, the spectrum distance is designed to find the difference in spectrum location between the real data and the generated data quantitatively by distance metric, which is used to guide the model training to generate high-quality data with similar features to the real data. Finally, the dynamic spectrum loss is proposed based on the spectrum distance to break through the synthesis of difficult frequencies in the data, by reducing the weight of easily synthesized frequencies in the spectrum while dynamically focusing on the difficult frequency components during training to achieve better generation results. In addition, experiments are conducted using several datasets, and the diagnostic accuracy of DSLGAN is 99. 63% and 99. 65%, reaching a very high level and verifying the effectiveness and superiority of DSLGAN.

NeurIPS Conference 2023 Conference Paper

Alternating Updates for Efficient Transformers

  • Cenk Baykal
  • Dylan Cutler
  • Nishanth Dikkala
  • Nikhil Ghosh
  • Rina Panigrahy
  • Xin Wang

It has been well established that increasing scale in deep transformer networks leads to improved quality and performance. However, this increase in scale often comes with prohibitive increases in compute cost and inference latency. We introduce Alternating Updates (AltUp), a simple-to-implement method to increase a model's capacity without the computational burden. AltUp enables the widening of the learned representation, i. e. , the token embedding, while only incurring a negligible increase in latency. AltUp achieves this by working on a subblock of the widened representation at each layer and using a predict-and-correct mechanism to update the inactivated blocks. We present extensions of AltUp, such as its applicability to the sequence dimension, and demonstrate how AltUp can be synergistically combined with existing approaches, such as Sparse Mixture-of-Experts models, to obtain efficient models with even higher capacity. Our experiments on benchmark transformer models and language tasks demonstrate the consistent effectiveness of AltUp on a diverse set of scenarios. Notably, on SuperGLUE and SQuAD benchmarks, AltUp enables up to $87\%$ speedup relative to the dense baselines at the same accuracy.

JBHI Journal 2023 Journal Article

Continuous Stress Detection Based on Social Media

  • Yang Ding
  • Ling Feng
  • Lei Cao
  • Yi Dai
  • Xin Wang
  • Huijun Zhang
  • Ningyun Li
  • Kaisheng Zeng

Leveraging social media for stress detection has been growing attention in recent years. Most relevant studies so far concentrated on training a stress detection model on the entire data in a closed environment, and did not continuously incorporate new information into the already established models but instead regularly reconstruct a new model from scratch. In this study, we formulate a social media based continuous stress detection task with two particular questions to be addressed: (1) when to adapt a learned stress detection model? and (2) how to adapt a learned stress detection model? We design a protocol to quantify the conditions that trigger model's adaptation, and develop a layer-inheritance based knowledge distillation method to continually adapt the learned stress detection model to incoming data, while retaining the knowledge gained previously. The experimental results on a constructed dataset containing 69 users on Tencent Weibo validate the effectiveness of the proposed adaptive layer-inheritance based knowledge distillation method, achieving 86. 32% and 91. 56% of accuracy in 3-label and 2-label continuous stress detection. Implications and further possible improvements are also discussed at the end of the article.

IJCAI Conference 2023 Conference Paper

Controlling Neural Style Transfer with Deep Reinforcement Learning

  • Chengming Feng
  • Jing Hu
  • Xin Wang
  • Shu Hu
  • Bin Zhu
  • Xi Wu
  • Hongtu Zhu
  • Siwei Lyu

Controlling the degree of stylization in the Neural Style Transfer (NST) is a little tricky since it usually needs hand-engineering on hyper-parameters. In this paper, we propose the first deep Reinforcement Learning (RL) based architecture that splits one-step style transfer into a step-wise process for the NST task. Our RL-based method tends to preserve more details and structures of the content image in early steps, and synthesize more style patterns in later steps. It is a user-easily-controlled style-transfer method. Additionally, as our RL-based model performs the stylization progressively, it is lightweight and has lower computational complexity than existing one-step Deep Learning (DL) based models. Experimental results demonstrate the effectiveness and robustness of our method.

IJCAI Conference 2023 Conference Paper

Curriculum Graph Machine Learning: A Survey

  • Haoyang Li
  • Xin Wang
  • Wenwu Zhu

Graph machine learning has been extensively studied in both academia and industry. However, in the literature, most existing graph machine learning models are designed to conduct training with data samples in a random order, which may suffer from suboptimal performance due to ignoring the importance of different graph data samples and their training orders for the model optimization status. To tackle this critical problem, curriculum graph machine learning (Graph CL), which integrates the strength of graph machine learning and curriculum learning, arises and attracts an increasing amount of attention from the research community. Therefore, in this paper, we comprehensively overview approaches on Graph CL and present a detailed survey of recent advances in this direction. Specifically, we first discuss the key challenges of Graph CL and provide its formal problem definition. Then, we categorize and summarize existing methods into three classes based on three kinds of graph machine learning tasks, i. e. , node-level, link-level, and graph-level tasks. Finally, we share our thoughts on future research directions. To the best of our knowledge, this paper is the first survey for curriculum graph machine learning.

AAAI Conference 2023 Conference Paper

Curriculum Multi-Negative Augmentation for Debiased Video Grounding

  • Xiaohan Lan
  • Yitian Yuan
  • Hong Chen
  • Xin Wang
  • Zequn Jie
  • Lin Ma
  • Zhi Wang
  • Wenwu Zhu

Video Grounding (VG) aims to locate the desired segment from a video given a sentence query. Recent studies have found that current VG models are prone to over-rely the groundtruth moment annotation distribution biases in the training set. To discourage the standard VG model's behavior of exploiting such temporal annotation biases and improve the model generalization ability, we propose multiple negative augmentations in a hierarchical way, including cross-video augmentations from clip-/video-level, and self-shuffled augmentations with masks. These augmentations can effectively diversify the data distribution so that the model can make more reasonable predictions instead of merely fitting the temporal biases. However, directly adopting such data augmentation strategy may inevitably carry some noise shown in our cases, since not all of the handcrafted augmentations are semantically irrelevant to the groundtruth video. To further denoise and improve the grounding accuracy, we design a multi-stage curriculum strategy to adaptively train the standard VG model from easy to hard negative augmentations. Experiments on newly collected Charades-CD and ActivityNet-CD datasets demonstrate our proposed strategy can improve the performance of the base model on both i.i.d and o.o.d scenarios.

AAAI Conference 2023 Conference Paper

Dynamic Heterogeneous Graph Attention Neural Architecture Search

  • Zeyang Zhang
  • Ziwei Zhang
  • Xin Wang
  • Yijian Qin
  • Zhou Qin
  • Wenwu Zhu

Dynamic heterogeneous graph neural networks (DHGNNs) have been shown to be effective in handling the ubiquitous dynamic heterogeneous graphs. However, the existing DHGNNs are hand-designed, requiring extensive human efforts and failing to adapt to diverse dynamic heterogeneous graph scenarios. In this paper, we propose to automate the design of DHGNN, which faces two major challenges: 1) how to design the search space to jointly consider the spatial-temporal dependencies and heterogeneous interactions in graphs; 2) how to design an efficient search algorithm in the potentially large and complex search space. To tackle these challenges, we propose a novel Dynamic Heterogeneous Graph Attention Search (DHGAS) method. Our proposed method can automatically discover the optimal DHGNN architecture and adapt to various dynamic heterogeneous graph scenarios without human guidance. In particular, we first propose a unified dynamic heterogeneous graph attention (DHGA) framework, which enables each node to jointly attend its heterogeneous and dynamic neighbors. Based on the framework, we design a localization space to determine where the attention should be applied and a parameterization space to determine how the attention should be parameterized. Lastly, we design a multi-stage differentiable search algorithm to efficiently explore the search space. Extensive experiments on real-world dynamic heterogeneous graph datasets demonstrate that our proposed method significantly outperforms state-of-the-art baselines for tasks including link prediction, node classification and node regression. To the best of our knowledge, DHGAS is the first dynamic heterogeneous graph neural architecture search method.

EAAI Journal 2023 Journal Article

Gaze-aware hand gesture recognition for intelligent construction

  • Xin Wang
  • Dharmaraj Veeramani
  • Zhenhua Zhu

The advances in construction robotics in recent decades has been a powerful driver of construction automation. User-friendly interfaces to support human–robot work collaboration are critical for increasing adoption of construction robots. Among different interfaces, hand gesture is an effective and reliable interaction cue in the noisy construction environment. This paper proposes a novel human gaze-aware hand gesture recognition framework as a human–robot interface for intelligent construction. The proposed framework relies on an eye tracker to visually detect and track robotic machines in the first-person view. Then, the machine-of-interest is determined based on the bounding boxes of machines and gaze points. Finally, a hand gesture recognition architecture is incorporated with the machine information for conveying messages to the machine-of-interest. This approach was tested through a framework validation test and achieved precision and recall of 93. 8% and 95. 0%, respectively. A pilot study was conducted to demonstrate interaction with a robotic excavator and a dump truck, and the results illustrated that the proposed framework could serve as an effective interface for one-to-many human–robot collaborations in construction.

NeurIPS Conference 2023 Conference Paper

Joint Data-Task Generation for Auxiliary Learning

  • Hong Chen
  • Xin Wang
  • Yuwei Zhou
  • Yijian Qin
  • Chaoyu Guan
  • Wenwu Zhu

Current auxiliary learning methods mainly adopt the methodology of reweighing losses for the manually collected auxiliary data and tasks. However, these methods heavily rely on domain knowledge during data collection, which may be hardly available in reality. Therefore, current methods will become less effective and even do harm to the primary task when unhelpful auxiliary data and tasks are employed. To tackle the problem, we propose a joint data-task generation framework for auxiliary learning (DTG-AuxL), which can bring benefits to the primary task by generating the new auxiliary data and task in a joint manner. The proposed DTG-AuxL framework contains a joint generator and a bi-level optimization strategy. Specifically, the joint generator contains a feature generator and a label generator, which are designed to be applicable and expressive for various auxiliary learning scenarios. The bi-level optimization strategy optimizes the joint generator and the task learning model, where the joint generator is effectively optimized in the upper level via the implicit gradient from the primary loss and the explicit gradient of our proposed instance regularization, while the task learning model is optimized in the lower level by the generated data and task. Extensive experiments show that our proposed DTG-AuxL framework consistently outperforms existing methods in various auxiliary learning scenarios, particularly when the manually collected auxiliary data and tasks are unhelpful.

AAAI Conference 2023 Conference Paper

JR2Net: Joint Monocular 3D Face Reconstruction and Reenactment

  • Jiaxiang Shang
  • Yu Zeng
  • Xin Qiao
  • Xin Wang
  • Runze Zhang
  • Guangyuan Sun
  • Vishal Patel
  • Hongbo Fu

Face reenactment and reconstruction benefit various applications in self-media, VR, etc. Recent face reenactment methods use 2D facial landmarks to implicitly retarget facial expressions and poses from driving videos to source images, while they suffer from pose and expression preservation issues for cross-identity scenarios, i.e., when the source and the driving subjects are different. Current self-supervised face reconstruction methods also demonstrate impressive results. However, these methods do not handle large expressions well, since their training data lacks samples of large expressions, and 2D facial attributes are inaccurate on such samples. To mitigate the above problems, we propose to explore the inner connection between the two tasks, i.e., using face reconstruction to provide sufficient 3D information for reenactment, and synthesizing videos paired with captured face model parameters through face reenactment to enhance the expression module of face reconstruction. In particular, we propose a novel cascade framework named JR2Net for Joint Face Reconstruction and Reenactment, which begins with the training of a coarse reconstruction network, followed by a 3D-aware face reenactment network based on the coarse reconstruction results. In the end, we train an expression tracking network based on our synthesized videos composed by image-face model parameter pairs. Such an expression tracking network can further enhance the coarse face reconstruction. Extensive experiments show that our JR2Net outperforms the state-of-the-art methods on several face reconstruction and reenactment benchmarks.

NeurIPS Conference 2023 Conference Paper

Multi-task Graph Neural Architecture Search with Task-aware Collaboration and Curriculum

  • Yijian Qin
  • Xin Wang
  • Ziwei Zhang
  • Hong Chen
  • Wenwu Zhu

Graph neural architecture search (GraphNAS) has shown great potential for automatically designing graph neural architectures for graph related tasks. However, multi-task GraphNAS capable of handling multiple tasks simultaneously has been largely unexplored in literature, posing great challenges to capture the complex relations and influences among different tasks. To tackle this problem, we propose a novel multi-task graph neural architecture search with task-aware collaboration and curriculum (MTGC3), which is able to simultaneously discover optimal architectures for different tasks and learn the collaborative relationships among different tasks in a joint manner. Specifically, we design the layer-wise disentangled supernet capable of managing multiple architectures in a unified framework, which combines with our proposed soft task-collaborative module to learn the transferability relationships between tasks. We further develop the task-wise curriculum training strategy to improve the architecture search procedure via reweighing the influence of different tasks based on task difficulties. Extensive experiments show that our proposed MTGC3 model achieves state-of-the-art performance against several baselines in multi-task scenarios, demonstrating its ability to discover effective architectures and capture the collaborative relationships for multiple tasks.

YNIMG Journal 2023 Journal Article

Neuroimaging-based classification of PTSD using data-driven computational approaches: A multisite big data study from the ENIGMA-PGC PTSD consortium

  • Xi Zhu
  • Yoojean Kim
  • Orren Ravid
  • Xiaofu He
  • Benjamin Suarez-Jimenez
  • Sigal Zilcha-Mano
  • Amit Lazarov
  • Seonjoo Lee

BACKGROUND: Recent advances in data-driven computational approaches have been helpful in devising tools to objectively diagnose psychiatric disorders. However, current machine learning studies limited to small homogeneous samples, different methodologies, and different imaging collection protocols, limit the ability to directly compare and generalize their results. Here we aimed to classify individuals with PTSD versus controls and assess the generalizability using a large heterogeneous brain datasets from the ENIGMA-PGC PTSD Working group. METHODS: We analyzed brain MRI data from 3,477 structural-MRI; 2,495 resting state-fMRI; and 1,952 diffusion-MRI. First, we identified the brain features that best distinguish individuals with PTSD from controls using traditional machine learning methods. Second, we assessed the utility of the denoising variational autoencoder (DVAE) and evaluated its classification performance. Third, we assessed the generalizability and reproducibility of both models using leave-one-site-out cross-validation procedure for each modality. RESULTS: We found lower performance in classifying PTSD vs. controls with data from over 20 sites (60 % test AUC for s-MRI, 59 % for rs-fMRI and 56 % for d-MRI), as compared to other studies run on single-site data. The performance increased when classifying PTSD from HC without trauma history in each modality (75 % AUC). The classification performance remained intact when applying the DVAE framework, which reduced the number of features. Finally, we found that the DVAE framework achieved better generalization to unseen datasets compared with the traditional machine learning frameworks, albeit performance was slightly above chance. CONCLUSION: These results have the potential to provide a baseline classification performance for PTSD when using large scale neuroimaging datasets. Our findings show that the control group used can heavily affect classification performance. The DVAE framework provided better generalizability for the multi-site data. This may be more significant in clinical practice since the neuroimaging-based diagnostic DVAE classification models are much less site-specific, rendering them more generalizable.

JBHI Journal 2023 Journal Article

Novel Graph Topology Learning for Spatio-Temporal Analysis of COVID-19 Spread

  • Baoling Shan
  • Xin Yuan
  • Wei Ni
  • Xin Wang
  • Ren Ping Liu
  • Eryk Dutkiewicz

This article presents a new graph-learning technique to accurately infer the graph structure of COVID-19 data, helping to reveal the correlation of pandemic dynamics among different countries and identify influential countries for pandemic response analysis. The new technique estimates the graph Laplacian of the COVID-19 data by first deriving analytically its precise eigenvectors, also known as graph Fourier transform (GFT) basis. Given the eigenvectors, the eigenvalues of the graph Laplacian are readily estimated using convex optimization. With the graph Laplacian, we analyze the confirmed cases of different COVID-19 variants among European countries based on centrality measures and identify a different set of the most influential and representative countries from the current techniques. The accuracy of the new method is validated by repurposing part of COVID-19 data to be the test data and gauging the capability of the method to recover missing test data, showing 33. 3% better in root mean squared error (RMSE) and 11. 11% better in correlation of determination than existing techniques. The set of identified influential countries by the method is anticipated to be meaningful and contribute to the study of COVID-19 spread.

NeurIPS Conference 2023 Conference Paper

Public Opinion Field Effect Fusion in Representation Learning for Trending Topics Diffusion

  • Junliang Li
  • Yang Yajun
  • Qinghua Hu
  • Xin Wang
  • Hong Gao

Trending topic diffusion and prediction analysis is an important problem and has been well studied in social networks. Representation learning is an effective way to extract node embeddings, which can help for topic propagation analysis by completing downstream tasks such as link prediction and node classification. In real world, there are often several trending topics or opinion leaders in public opinion space at the same time and they can be regarded as different centers of public opinion. A public opinion field will be formed surrounding every center. These public opinion fields compete for public's attention and it will potentially affect the development of public opinion. However, the existing methods do not consider public opinion field effect for trending topics diffusion. In this paper, we introduce three well-known observations about public opinion field effect in media and communication studies, and propose a novel and effective heterogeneous representation learning framework to incorporate public opinion field effect and social circle influence effect. To the best of our knowledge, our work is the first to consider these effects in representation learning for trending topic diffusion. Extensive experiments on real-world datasets validate the superiority of our model.

AAAI Conference 2023 Conference Paper

Reducing Sentiment Bias in Pre-trained Sentiment Classification via Adaptive Gumbel Attack

  • Jiachen Tian
  • Shizhan Chen
  • Xiaowang Zhang
  • Xin Wang
  • Zhiyong Feng

Pre-trained language models (PLMs) have recently enabled rapid progress on sentiment classification under the pre-train and fine-tune paradigm, where the fine-tuning phase aims to transfer the factual knowledge learned by PLMs to sentiment classification. However, current fine-tuning methods ignore the risk that PLMs cause the problem of sentiment bias, that is, PLMs tend to inject positive or negative sentiment from the contextual information of certain entities (or aspects) into their word embeddings, leading them to establish spurious correlations with labels. In this paper, we propose an adaptive Gumbel-attacked classifier that immunes sentiment bias from an adversarial-attack perspective. Due to the complexity and diversity of sentiment bias, we construct multiple Gumbel-attack expert networks to generate various noises from mixed Gumbel distribution constrained by mutual information minimization, and design an adaptive training framework to synthesize complex noise by confidence-guided controlling the number of expert networks. Finally, we capture these noises that effectively simulate sentiment bias based on the feedback of the classifier, and then propose a multi-channel parameter updating algorithm to strengthen the classifier to recognize these noises by fusing the parameters between the classifier and each expert network. Experimental results illustrate that our method significantly reduced sentiment bias and improved the performance of sentiment classification.

NeurIPS Conference 2023 Conference Paper

Spectral Invariant Learning for Dynamic Graphs under Distribution Shifts

  • Zeyang Zhang
  • Xin Wang
  • Ziwei Zhang
  • Zhou Qin
  • Weigao Wen
  • Hui Xue'
  • Haoyang Li
  • Wenwu Zhu

Dynamic graph neural networks (DyGNNs) currently struggle with handling distribution shifts that are inherent in dynamic graphs. Existing work on DyGNNs with out-of-distribution settings only focuses on the time domain, failing to handle cases involving distribution shifts in the spectral domain. In this paper, we discover that there exist cases with distribution shifts unobservable in the time domain while observable in the spectral domain, and propose to study distribution shifts on dynamic graphs in the spectral domain for the first time. However, this investigation poses two key challenges: i) it is non-trivial to capture different graph patterns that are driven by various frequency components entangled in the spectral domain; and ii) it remains unclear how to handle distribution shifts with the discovered spectral patterns. To address these challenges, we propose Spectral Invariant Learning for Dynamic Graphs under Distribution Shifts (SILD), which can handle distribution shifts on dynamic graphs by capturing and utilizing invariant and variant spectral patterns. Specifically, we first design a DyGNN with Fourier transform to obtain the ego-graph trajectory spectrums, allowing the mixed dynamic graph patterns to be transformed into separate frequency components. We then develop a disentangled spectrum mask to filter graph dynamics from various frequency components and discover the invariant and variant spectral patterns. Finally, we propose invariant spectral filtering, which encourages the model to rely on invariant patterns for generalization under distribution shifts. Experimental results on synthetic and real-world dynamic graph datasets demonstrate the superiority of our method for both node classification and link prediction tasks under distribution shifts.

NeurIPS Conference 2023 Conference Paper

Statistical Analysis of Quantum State Learning Process in Quantum Neural Networks

  • Hao-Kai Zhang
  • Chenghong Zhu
  • Mingrui Jing
  • Xin Wang

Quantum neural networks (QNNs) have been a promising framework in pursuing near-term quantum advantage in various fields, where many applications can be viewed as learning a quantum state that encodes useful data. As a quantum analog of probability distribution learning, quantum state learning is theoretically and practically essential in quantum machine learning. In this paper, we develop a no-go theorem for learning an unknown quantum state with QNNs even starting from a high-fidelity initial state. We prove that when the loss value is lower than a critical threshold, the probability of avoiding local minima vanishes exponentially with the qubit count, while only grows polynomially with the circuit depth. The curvature of local minima is concentrated to the quantum Fisher information times a loss-dependent constant, which characterizes the sensibility of the output state with respect to parameters in QNNs. These results hold for any circuit structures, initialization strategies, and work for both fixed ansatzes and adaptive methods. Extensive numerical simulations are performed to validate our theoretical results. Our findings place generic limits on good initial guesses and adaptive methods for improving the learnability and scalability of QNNs, and deepen the understanding of prior information's role in QNNs.

NeurIPS Conference 2023 Conference Paper

Unsupervised Graph Neural Architecture Search with Disentangled Self-Supervision

  • Zeyang Zhang
  • Xin Wang
  • Ziwei Zhang
  • Guangyao Shen
  • Shiqi Shen
  • Wenwu Zhu

The existing graph neural architecture search (GNAS) methods heavily rely on supervised labels during the search process, failing to handle ubiquitous scenarios where supervisions are not available. In this paper, we study the problem of unsupervised graph neural architecture search, which remains unexplored in the literature. The key problem is to discover the latent graph factors that drive the formation of graph data as well as the underlying relations between the factors and the optimal neural architectures. Handling this problem is challenging given that the latent graph factors together with architectures are highly entangled due to the nature of the graph and the complexity of the neural architecture search process. To address the challenge, we propose a novel Disentangled Self-supervised Graph Neural Architecture Search (DSGAS) model, which is able to discover the optimal architectures capturing various latent graph factors in a self-supervised fashion based on unlabeled graph data. Specifically, we first design a disentangled graph super-network capable of incorporating multiple architectures with factor-wise disentanglement, which are optimized simultaneously. Then, we estimate the performance of architectures under different factors by our proposed self-supervised training with joint architecture-graph disentanglement. Finally, we propose a contrastive search with architecture augmentations to discover architectures with factor-specific expertise. Extensive experiments on 11 real-world datasets demonstrate that the proposed model is able to achieve state-of-the-art performance against several baseline methods in an unsupervised manner.

YNIMG Journal 2022 Journal Article

A comparison of methods to harmonize cortical thickness measurements across scanners and sites

  • Delin Sun
  • Gopalkumar Rakesh
  • Courtney C. Haswell
  • Mark Logue
  • C. Lexi Baird
  • Erin N. O'Leary
  • Andrew S. Cotton
  • Hong Xie

Results of neuroimaging datasets aggregated from multiple sites may be biased by site-specific profiles in participants’ demographic and clinical characteristics, as well as MRI acquisition protocols and scanning platforms. We compared the impact of four different harmonization methods on results obtained from analyses of cortical thickness data: (1) linear mixed-effects model (LME) that models site-specific random intercepts (LMEINT), (2) LME that models both site-specific random intercepts and age-related random slopes (LMEINT+SLP), (3) ComBat, and (4) ComBat with a generalized additive model (ComBat-GAM). Our test case for comparing harmonization methods was cortical thickness data aggregated from 29 sites, which included 1, 340 cases with posttraumatic stress disorder (PTSD) (6. 2–81. 8 years old) and 2, 057 trauma-exposed controls without PTSD (6. 3–85. 2 years old). We found that, compared to the other data harmonization methods, data processed with ComBat-GAM was more sensitive to the detection of significant case-control differences (Χ 2(3) = 63. 704, p < 0. 001) as well as case-control differences in age-related cortical thinning (Χ 2(3) = 12. 082, p = 0. 007). Both ComBat and ComBat-GAM outperformed LME methods in detecting sex differences (Χ 2(3) = 9. 114, p = 0. 028) in regional cortical thickness. ComBat-GAM also led to stronger estimates of age-related declines in cortical thickness (corrected p-values < 0. 001), stronger estimates of case-related cortical thickness reduction (corrected p-values < 0. 001), weaker estimates of age-related declines in cortical thickness in cases than controls (corrected p-values < 0. 001), stronger estimates of cortical thickness reduction in females than males (corrected p-values < 0. 001), and stronger estimates of cortical thickness reduction in females relative to males in cases than controls (corrected p-values < 0. 001). Our results support the use of ComBat-GAM to minimize confounds and increase statistical power when harmonizing data with non-linear effects, and the use of either ComBat or ComBat-GAM for harmonizing data with linear effects.

NeurIPS Conference 2022 Conference Paper

A Theoretical View on Sparsely Activated Networks

  • Cenk Baykal
  • Nishanth Dikkala
  • Rina Panigrahy
  • Cyrus Rashtchian
  • Xin Wang

Deep and wide neural networks successfully fit very complex functions today, but dense models are starting to be prohibitively expensive for inference. To mitigate this, one promising research direction is networks that activate a sparse subgraph of the network. The subgraph is chosen by a data-dependent routing function, enforcing a fixed mapping of inputs to subnetworks (e. g. , the Mixture of Experts (MoE) paradigm in Switch Transformers). However, there is no theoretical grounding for these sparsely activated models. As our first contribution, we present a formal model of data-dependent sparse networks that captures salient aspects of popular architectures. Then, we show how to construct sparse networks that provably match the approximation power and total size of dense networks on Lipschitz functions. The sparse networks use much fewer inference operations than dense networks, leading to a faster forward pass. The key idea is to use locality sensitive hashing on the input vectors and then interpolate the function in subregions of the input space. This offers a theoretical insight into why sparse networks work well in practice. Finally, we present empirical findings that support our theory; compared to dense networks, sparse networks give a favorable trade-off between number of active units and approximation quality.

JBHI Journal 2022 Journal Article

Category-Aware Chronic Stress Detection on Microblogs

  • Lei Cao
  • Huijun Zhang
  • Ningyun Li
  • Xin Wang
  • Wisong Ri
  • Ling Feng

People today live a stressful life. Compared with acute stress, long-term chronic stress is more harmful, and may cause or exacerbate many serious health problems, including high blood pressure, heart disease, chronic pain, and mental diseases. With social media becoming an integral part of our daily lives for information sharing and self-expression, detecting category-aware long-standing chronic stress from a large volume of historic open posts made by social media users is possible. In this study, we construct a data set containing 971 chronically stressed users with totally 54 546 open posts on Sina microblog from July 5, 2018 to December 1, 2019, and design two techniques for category-aware chronic stress detection: (1) a stress-oriented word embedding on the basis of an existing pre-trained word embedding, aiming to strengthen the sensibility of stress-related expressions for linguistic post analysis; (2) a multi-attention model with three layers (i. e. , category-attention layer, posts self-attention layer, and category-specific post attention layer), aiming to capture inter-relevance from a sequence of posts and infer long-term stress categories and stress levels. The experimental results show that the proposed multi-attention model equipped with the stress-oriented word embedding can achieve 80. 65% accuracy in detecting category-aware stress levels, 86. 49% accuracy in detecting chronic stress levels only, and 93. 07% accuracy in detecting chronic stress categories only. Limitations and implications of the study are also discussed at the end of the paper.

NeurIPS Conference 2022 Conference Paper

Concentration of Data Encoding in Parameterized Quantum Circuits

  • Guangxi Li
  • Ruilin Ye
  • Xuanqiang Zhao
  • Xin Wang

Variational quantum algorithms have been acknowledged as the leading strategy to realize near-term quantum advantages in meaningful tasks, including machine learning and optimization. When applied to tasks involving classical data, such algorithms generally begin with data encoding circuits and train quantum neural networks (QNNs) to minimize target functions. Although QNNs have been widely studied to improve these algorithms' performance on practical tasks, there is a gap in systematically understanding the influence of data encoding on the eventual performance. In this paper, we make progress in filling this gap by considering the common data encoding strategies based on parameterized quantum circuits. We prove that, under reasonable assumptions, the distance between the average encoded state and the maximally mixed state could be explicitly upper-bounded with respect to the width and depth of the encoding circuit. This result in particular implies that the average encoded state will concentrate on the maximally mixed state at an exponential speed on depth. Such concentration seriously limits the capabilities of quantum classifiers, and strictly restricts the distinguishability of encoded states from a quantum information perspective. To support our findings, we numerically verify these results on both synthetic and public data sets. Our results highlight the significance of quantum data encoding and may shed light on the future design of quantum encoding strategies.

NeurIPS Conference 2022 Conference Paper

Dynamic Graph Neural Networks Under Spatio-Temporal Distribution Shift

  • Zeyang Zhang
  • Xin Wang
  • Ziwei Zhang
  • Haoyang Li
  • Zhou Qin
  • Wenwu Zhu

Dynamic graph neural networks (DyGNNs) have demonstrated powerful predictive abilities by exploiting graph structural and temporal dynamics. However, the existing DyGNNs fail to handle distribution shifts, which naturally exist in dynamic graphs, mainly because the patterns exploited by DyGNNs may be variant with respect to labels under distribution shifts. In this paper, we propose to handle spatio-temporal distribution shifts in dynamic graphs by discovering and utilizing {\it invariant patterns}, i. e. , structures and features whose predictive abilities are stable across distribution shifts, which faces two key challenges: 1) How to discover the complex variant and invariant spatio-temporal patterns in dynamic graphs, which involve both time-varying graph structures and node features. 2) How to handle spatio-temporal distribution shifts with the discovered variant and invariant patterns. To tackle these challenges, we propose the Disentangled Intervention-based Dynamic graph Attention networks (DIDA). Our proposed method can effectively handle spatio-temporal distribution shifts in dynamic graphs by discovering and fully utilizing invariant spatio-temporal patterns. Specifically, we first propose a disentangled spatio-temporal attention network to capture the variant and invariant patterns. Then, we design a spatio-temporal intervention mechanism to create multiple interventional distributions by sampling and reassembling variant patterns across neighborhoods and time stamps to eliminate the spurious impacts of variant patterns. Lastly, we propose an invariance regularization term to minimize the variance of predictions in intervened distributions so that our model can make predictions based on invariant patterns with stable predictive abilities and therefore handle distribution shifts. Experiments on three real-world datasets and one synthetic dataset demonstrate the superiority of our method over state-of-the-art baselines under distribution shifts. Our work is the first study of spatio-temporal distribution shifts in dynamic graphs, to the best of our knowledge.

JBHI Journal 2022 Journal Article

Dynamic Link Prediction for Discovery of New Impactful COVID-19 Research Approaches

  • Xiangyu Wang
  • Yuan Li
  • Taiyu Ban
  • Jiarun Zhu
  • Lyuzhou Chen
  • Muhammad Usman
  • Xin Wang
  • Huanhuan Chen

In fighting the COVID-19 pandemic, the main challenges include the lack of prior research and the urgency to find effective solutions. It is essential to accurately and rapidly summarize the relevant research work and explore potential solutions for diagnosis, treatment and prevention of COVID-19. It is a daunting task to summarize the numerous existing research works and to assess their effectiveness. This paper explores the discovery of new COVID-19 research approaches based on dynamic link prediction, which analyze the dynamic topological network of keywords to predict possible connections of research concepts. A dynamic link prediction method based on multi-granularity feature fusion is proposed. Firstly, a multi-granularity temporal feature fusion method is adopted to extract the temporal evolution of different order subgraphs. Secondly, a hierarchical feature weighting method is proposed to emphasize actively evolving nodes. Thirdly, a semantic repetition sampling mechanism is designed to avoid the negative effect of semantically equivalent medical entities on the real structure of the graph, and to capture the real topological structure features. Experiments are performed on the COVID-19 Open Research Dataset to assess the performance of the model. The results show that the proposed model performs significantly better than existing state-of-the-art models, thereby confirming the effectiveness of the proposed method for the discovery of new COVID-19 research approaches.

NeurIPS Conference 2022 Conference Paper

Generalization Bounds for Estimating Causal Effects of Continuous Treatments

  • Xin Wang
  • Shengfei Lyu
  • Xingyu Wu
  • Tianhao Wu
  • Huanhuan Chen

We focus on estimating causal effects of continuous treatments (e. g. , dosage in medicine), also known as dose-response function. Existing methods in causal inference for continuous treatments using neural networks are effective and to some extent reduce selection bias, which is introduced by non-randomized treatments among individuals and might lead to covariate imbalance and thus unreliable inference. To theoretically support the alleviation of selection bias in the setting of continuous treatments, we exploit the re-weighting schema and the Integral Probability Metric (IPM) distance to derive an upper bound on the counterfactual loss of estimating the average dose-response function (ADRF), and herein the IPM distance builds a bridge from a source (factual) domain to an infinite number of target (counterfactual) domains. We provide a discretized approximation of the IPM distance with a theoretical guarantee in the practical implementation. Based on the theoretical analyses, we also propose a novel algorithm, called Average Dose- response estiMatIon via re-weighTing schema (ADMIT). ADMIT simultaneously learns a re-weighting network, which aims to alleviate the selection bias, and an inference network, which makes factual and counterfactual estimations. In addition, the effectiveness of ADMIT is empirically demonstrated in both synthetic and semi-synthetic experiments by outperforming the existing benchmarks.

NeurIPS Conference 2022 Conference Paper

Generative Status Estimation and Information Decoupling for Image Rain Removal

  • Di Lin
  • Xin Wang
  • Jia Shen
  • Renjie Zhang
  • Ruonan Liu
  • Miaohui Wang
  • Wuyuan Xie
  • Qing Guo

Image rain removal requires the accurate separation between the pixels of the rain streaks and object textures. But the confusing appearances of rains and objects lead to the misunderstanding of pixels, thus remaining the rain streaks or missing the object details in the result. In this paper, we propose SEIDNet equipped with the generative Status Estimation and Information Decoupling for rain removal. In the status estimation, we embed the pixel-wise statuses into the status space, where each status indicates a pixel of the rain or object. The status space allows sampling multiple statuses for a pixel, thus capturing the confusing rain or object. In the information decoupling, we respect the pixel-wise statuses, decoupling the appearance information of rain and object from the pixel. Based on the decoupled information, we construct the kernel space, where multiple kernels are sampled for the pixel to remove the rain and recover the object appearance. We evaluate SEIDNet on the public datasets, achieving state-of-the-art performances of image rain removal. The experimental results also demonstrate the generalization of SEIDNet, which can be easily extended to achieve state-of-the-art performances on other image restoration tasks (e. g. , snow, haze, and shadow removal).

NeurIPS Conference 2022 Conference Paper

Learning Invariant Graph Representations for Out-of-Distribution Generalization

  • Haoyang Li
  • Ziwei Zhang
  • Xin Wang
  • Wenwu Zhu

Graph representation learning has shown effectiveness when testing and training graph data come from the same distribution, but most existing approaches fail to generalize under distribution shifts. Invariant learning, backed by the invariance principle from causality, can achieve guaranteed generalization under distribution shifts in theory and has shown great successes in practice. However, invariant learning for graphs under distribution shifts remains unexplored and challenging. To solve this problem, we propose Graph Invariant Learning (GIL) model capable of learning generalized graph representations under distribution shifts. Our proposed method can capture the invariant relationships between predictive graph structural information and labels in a mixture of latent environments through jointly optimizing three tailored modules. Specifically, we first design a GNN-based subgraph generator to identify invariant subgraphs. Then we use the variant subgraphs, i. e. , complements of invariant subgraphs, to infer the latent environment labels. We further propose an invariant learning module to learn graph representations that can generalize to unknown test graphs. Theoretical justifications for our proposed method are also provided. Extensive experiments on both synthetic and real-world datasets demonstrate the superiority of our method against state-of-the-art baselines under distribution shifts for the graph classification task.

AAAI Conference 2022 Conference Paper

Learning to Solve Travelling Salesman Problem with Hardness-Adaptive Curriculum

  • Zeyang Zhang
  • Ziwei Zhang
  • Xin Wang
  • Wenwu Zhu

Various neural network models have been proposed to tackle combinatorial optimization problems such as the travelling salesman problem (TSP). Existing learningbased TSP methods adopt a simple setting that the training and testing data are independent and identically distributed. However, the existing literature fails to solve TSP instances when training and testing data have different distributions. Concretely, we find that different training and testing distribution will result in more difficult TSP instances, i. e. , the solution obtained by the model has a large gap from the optimal solution. To tackle this problem, in this work, we study learning-based TSP methods when training and testing data have different distributions using adaptivehardness, i. e. , how difficult a TSP instance can be for a solver. This problem is challenging because it is nontrivial to (1) define hardness measurement quantitatively; (2) efficiently and continuously generate sufficiently hard TSP instances upon model training; (3) fully utilize instances with different levels of hardness to learn a more powerful TSP solver. To solve these challenges, we first propose a principled hardness measurement to quantify the hardness of TSP instances. Then, we propose a hardness-adaptive generator to generate instances with different hardness. We further propose a curriculum learner fully utilizing these instances to train the TSP solver. Experiments show that our hardness-adaptive generator can generate instances ten times harder than the existing methods, and our proposed method achieves significant improvement over state-of-the-art models in terms of the optimality gap.

NeurIPS Conference 2022 Conference Paper

Module-Aware Optimization for Auxiliary Learning

  • Hong Chen
  • Xin Wang
  • Yue Liu
  • Yuwei Zhou
  • Chaoyu Guan
  • Wenwu Zhu

Auxiliary learning is a widely adopted practice in deep learning, which aims to improve the model performance on the primary task by exploiting the beneficial information in the auxiliary loss. Existing auxiliary learning methods only focus on balancing the auxiliary loss and the primary loss, ignoring the module-level auxiliary influence, i. e. , an auxiliary loss will be beneficial for optimizing specific modules within the model but harmful to others, failing to make full use of auxiliary information. To tackle the problem, we propose a Module-Aware Optimization approach for Auxiliary Learning (MAOAL). The proposed approach considers the module-level influence through the learnable module-level auxiliary importance, i. e. , the importance of each auxiliary loss to each module. Specifically, the proposed approach jointly optimizes the module-level auxiliary importance and the model parameters in a bi-level manner. In the lower optimization, the model parameters are optimized with the importance parameterized gradient, while in the upper optimization, the module-level auxiliary importance is updated with the implicit gradient from a small developing dataset. Extensive experiments show that our proposed MAOAL method consistently outperforms state-of-the-art baselines for different auxiliary losses on various datasets, demonstrating that our method can serve as a powerful generic tool for auxiliary learning.

NeurIPS Conference 2022 Conference Paper

NAS-Bench-Graph: Benchmarking Graph Neural Architecture Search

  • Yijian Qin
  • Ziwei Zhang
  • Xin Wang
  • Zeyang Zhang
  • Wenwu Zhu

Graph neural architecture search (GraphNAS) has recently aroused considerable attention in both academia and industry. However, two key challenges seriously hinder the further research of GraphNAS. First, since there is no consensus for the experimental setting, the empirical results in different research papers are often not comparable and even not reproducible, leading to unfair comparisons. Secondly, GraphNAS often needs extensive computations, which makes it highly inefficient and inaccessible to researchers without access to large-scale computation. To solve these challenges, we propose NAS-Bench-Graph, a tailored benchmark that supports unified, reproducible, and efficient evaluations for GraphNAS. Specifically, we construct a unified, expressive yet compact search space, covering 26, 206 unique graph neural network (GNN) architectures and propose a principled evaluation protocol. To avoid unnecessary repetitive training, we have trained and evaluated all of these architectures on nine representative graph datasets, recording detailed metrics including train, validation, and test performance in each epoch, the latency, the number of parameters, etc. Based on our proposed benchmark, the performance of GNN architectures can be directly obtained by a look-up table without any further computation, which enables fair, fully reproducible, and efficient comparisons. To demonstrate its usage, we make in-depth analyses of our proposed NAS-Bench-Graph, revealing several interesting findings for GraphNAS. We also showcase how the benchmark can be easily compatible with GraphNAS open libraries such as AutoGL and NNI. To the best of our knowledge, our work is the first benchmark for graph neural architecture search.

AAAI Conference 2022 Conference Paper

Orthogonal Graph Neural Networks

  • Kai Guo
  • Kaixiong Zhou
  • Xia Hu
  • Yu Li
  • Yi Chang
  • Xin Wang

Graph neural networks (GNNs) have received tremendous attention due to their superiority in learning node representations. These models rely on message passing and feature transformation functions to encode the structural and feature information from neighbors. However, stacking more convolutional layers significantly decreases the performance of GNNs. Most recent studies attribute this limitation to the over-smoothing issue, where node embeddings converge to indistinguishable vectors. Through a number of experimental observations, we argue that the main factor degrading the performance is the unstable forward normalization and backward gradient resulted from the improper design of the feature transformation, especially for shallow GNNs where the over-smoothing has not happened. Therefore, we propose a novel orthogonal feature transformation, named Ortho- GConv, which could generally augment the existing GNN backbones to stabilize the model training and improve the model’s generalization performance. Specifically, we maintain the orthogonality of the feature transformation comprehensively from three perspectives, namely hybrid weight initialization, orthogonal transformation, and orthogonal regularization. By equipping the existing GNNs (e. g. GCN, JKNet, GCNII) with Ortho-GConv, we demonstrate the generality of the orthogonal feature transformation to enable stable training, and show its effectiveness for node and graph classification tasks.

NeurIPS Conference 2022 Conference Paper

Power and limitations of single-qubit native quantum neural networks

  • Zhan Yu
  • Hongshun Yao
  • Mujin Li
  • Xin Wang

Quantum neural networks (QNNs) have emerged as a leading strategy to establish applications in machine learning, chemistry, and optimization. While the applications of QNN have been widely investigated, its theoretical foundation remains less understood. In this paper, we formulate a theoretical framework for the expressive ability of data re-uploading quantum neural networks that consist of interleaved encoding circuit blocks and trainable circuit blocks. First, we prove that single-qubit quantum neural networks can approximate any univariate function by mapping the model to a partial Fourier series. We in particular establish the exact correlations between the parameters of the trainable gates and the Fourier coefficients, resolving an open problem on the universal approximation property of QNN. Second, we discuss the limitations of single-qubit native QNNs on approximating multivariate functions by analyzing the frequency spectrum and the flexibility of Fourier coefficients. We further demonstrate the expressivity and limitations of single-qubit native QNNs via numerical experiments. We believe these results would improve our understanding of QNNs and provide a helpful guideline for designing powerful QNNs for machine learning tasks.

NeurIPS Conference 2022 Conference Paper

Sketching based Representations for Robust Image Classification with Provable Guarantees

  • Nishanth Dikkala
  • Sankeerth Rao Karingula
  • Raghu Meka
  • Jelani Nelson
  • Rina Panigrahy
  • Xin Wang

How do we provably represent images succinctly so that their essential latent attributes are correctly captured by the representation to as high level of detail as possible? While today's deep networks (such as CNNs) produce image embeddings they do not have any provable properties and seem to work in mysterious non-interpretable ways. In this work we theoretically study synthetic images that are composed of a union or intersection of several mathematically specified shapes using thresholded polynomial functions (for e. g. ellipses, rectangles). We show how to produce a succinct sketch of such an image so that the sketch “smoothly” maps to the latent-coefficients producing the different shapes in the image. We prove several important properties such as: easy reconstruction of the image from the sketch, similarity preservation (similar shapes produce similar sketches), being able to index sketches so that other similar images and parts of other images can be retrieved, being able to store the sketches into a dictionary of concepts and shapes so parts of the same or different images that refer to the same shape can point to the same entry in this dictionary of common shape attributes.

AAAI Conference 2022 Conference Paper

Stochastic Planner-Actor-Critic for Unsupervised Deformable Image Registration

  • Ziwei Luo
  • Jing Hu
  • Xin Wang
  • Shu Hu
  • Bin Kong
  • Youbing Yin
  • Qi Song
  • Xi Wu

Large deformations of organs, caused by diverse shapes and nonlinear shape changes, pose a significant challenge for medical image registration. Traditional registration methods need to iteratively optimize an objective function via a specific deformation model along with meticulous parameter tuning, but which have limited capabilities in registering images with large deformations. While deep learning-based methods can learn the complex mapping from input images to their respective deformation field, it is regression-based and is prone to be stuck at local minima, particularly when large deformations are involved. To this end, we present Stochastic Planner-Actor-Critic (SPAC), a novel reinforcement learningbased framework that performs step-wise registration. The key notion is warping a moving image successively by each time step to finally align to a fixed image. Considering that it is challenging to handle high dimensional continuous action and state spaces in the conventional reinforcement learning (RL) framework, we introduce a new concept ‘Plan’ to the standard Actor-Critic model, which is of low dimension and can facilitate the actor to generate a tractable high dimensional action. The entire framework is based on unsupervised training and operates in an end-to-end manner. We evaluate our method on several 2D and 3D medical image datasets, some of which contain large deformations. Our empirical results highlight that our work achieves consistent, significant gains and outperforms state-of-the-art methods.

JMLR Journal 2022 Journal Article

Sum of Ranked Range Loss for Supervised Learning

  • Shu Hu
  • Yiming Ying
  • Xin Wang
  • Siwei Lyu

In forming learning objectives, one oftentimes needs to aggregate a set of individual values to a single output. Such cases occur in the aggregate loss, which combines individual losses of a learning model over each training sample, and in the individual loss for multi-label learning, which combines prediction scores over all class labels. In this work, we introduce the sum of ranked range (SoRR) as a general approach to form learning objectives. A ranked range is a consecutive sequence of sorted values of a set of real numbers. The minimization of SoRR is solved with the difference of convex algorithm (DCA). We explore two applications in machine learning of the minimization of the SoRR framework, namely the AoRR aggregate loss for binary/multi-class classification at the sample level and the TKML individual loss for multi-label/multi-class classification at the label level. A combination loss of AoRR and TKML is proposed as a new learning objective for improving the robustness of multi-label learning in the face of outliers in sample and labels alike. Our empirical results highlight the effectiveness of the proposed optimization frameworks and demonstrate the applicability of proposed losses using synthetic and real data sets. [abs] [ pdf ][ bib ] [ code ] &copy JMLR 2022. ( edit, beta )

NeurIPS Conference 2022 Conference Paper

VLMbench: A Compositional Benchmark for Vision-and-Language Manipulation

  • Kaizhi Zheng
  • Xiaotong Chen
  • Odest Chadwicke Jenkins
  • Xin Wang

Benefiting from language flexibility and compositionality, humans naturally intend to use language to command an embodied agent for complex tasks such as navigation and object manipulation. In this work, we aim to fill the blank of the last mile of embodied agents---object manipulation by following human guidance, e. g. , “move the red mug next to the box while keeping it upright. ” To this end, we introduce an Automatic Manipulation Solver (AMSolver) system and build a Vision-and-Language Manipulation benchmark (VLMbench) based on it, containing various language instructions on categorized robotic manipulation tasks. Specifically, modular rule-based task templates are created to automatically generate robot demonstrations with language instructions, consisting of diverse object shapes and appearances, action types, and motion constraints. We also develop a keypoint-based model 6D-CLIPort to deal with multi-view observations and language input and output a sequence of 6 degrees of freedom (DoF) actions. We hope the new simulator and benchmark will facilitate future research on language-guided robotic manipulation.

TCS Journal 2021 Journal Article

Approximation algorithms for the submodular edge cover problem with submodular penalties

  • Xin Wang
  • Suogang Gao
  • Bo Hou
  • Lidong Wu
  • Wen Liu

In this paper, we consider the submodular edge cover problem with submodular penalties. In this problem, we are given an undirected graph G = ( V, E ) with vertex set V and edge set E. Assume the covering cost function c: 2 E → R + and the penalty function p: 2 V → R + are both submodular with p non-decreasing, c ( ∅ ) = 0 and p ( ∅ ) = 0. The goal of the submodular edge cover problem with submodular penalties is to select an edge subset to cover some vertices and penalize the vertex subset containing uncovered vertices such that the total cost of covering and penalty is minimized. For this problem, we first give a 2Δ-approximation algorithm by using a primal-dual technique, where Δ is the maximal degree of the graph G. Then we transform this problem into a submodular set cover problem, and by applying a known result for the submodular set cover problem we conclude that there is an approximation algorithm with an approximation ratio Δ + 1.

IJCAI Conference 2021 Conference Paper

Automated Machine Learning on Graphs: A Survey

  • Ziwei Zhang
  • Xin Wang
  • Wenwu Zhu

Machine learning on graphs has been extensively studied in both academic and industry. However, as the literature on graph learning booms with a vast number of emerging methods and techniques, it becomes increasingly difficult to manually design the optimal machine learning algorithm for different graph-related tasks. To solve this critical challenge, automated machine learning (AutoML) on graphs which combines the strength of graph machine learning and AutoML together, is gaining attention from the research community. Therefore, we comprehensively survey AutoML on graphs in this paper, primarily focusing on hyper-parameter optimization (HPO) and neural architecture search (NAS) for graph machine learning. We further overview libraries related to automated graph machine learning and in-depth discuss AutoGL, the first dedicated open-source library for AutoML on graphs. In the end, we share our insights on future research directions for automated graph machine learning. This paper is the first systematic and comprehensive review of automated machine learning on graphs to the best of our knowledge.

NeurIPS Conference 2021 Conference Paper

Curriculum Disentangled Recommendation with Noisy Multi-feedback

  • Hong Chen
  • Yudong Chen
  • Xin Wang
  • Ruobing Xie
  • Rui Wang
  • Feng Xia
  • Wenwu Zhu

Learning disentangled representations for user intentions from multi-feedback (i. e. , positive and negative feedback) can enhance the accuracy and explainability of recommendation algorithms. However, learning such disentangled representations from multi-feedback data is challenging because i) multi-feedback is complex: there exist complex relations among different types of feedback (e. g. , click, unclick, and dislike, etc) as well as various user intentions, and ii) multi-feedback is noisy: there exists noisy (useless) information both in features and labels, which may deteriorate the recommendation performance. Existing works on disentangled representation learning only focus on positive feedback, failing to handle the complex relations and noise hidden in multi-feedback data. To solve this problem, in this work we propose a Curriculum Disentangled Recommendation (CDR) model that is capable of efficiently learning disentangled representations from complex and noisy multi-feedback for better recommendation. Concretely, we design a co-filtering dynamic routing mechanism that simultaneously captures the complex relations among different behavioral feedback and user intentions as well as denoise the representations in the feature level. We then present an adjustable self-evaluating curriculum that is able to evaluate sample difficulties for better model training and conduct denoising in the label level via disregarding useless information. Our extensive experiments on several real-world datasets demonstrate that the proposed CDR model can significantly outperform several state-of-the-art methods in terms of recommendation accuracy.

NeurIPS Conference 2021 Conference Paper

Disentangled Contrastive Learning on Graphs

  • Haoyang Li
  • Xin Wang
  • Ziwei Zhang
  • Zehuan Yuan
  • Hang Li
  • Wenwu Zhu

Recently, self-supervised learning for graph neural networks (GNNs) has attracted considerable attention because of their notable successes in learning the representation of graph-structure data. However, the formation of a real-world graph typically arises from the highly complex interaction of many latent factors. The existing self-supervised learning methods for GNNs are inherently holistic and neglect the entanglement of the latent factors, resulting in the learned representations suboptimal for downstream tasks and difficult to be interpreted. Learning disentangled graph representations with self-supervised learning poses great challenges and remains largely ignored by the existing literature. In this paper, we introduce the Disentangled Graph Contrastive Learning (DGCL) method, which is able to learn disentangled graph-level representations with self-supervision. In particular, we first identify the latent factors of the input graph and derive its factorized representations. Each of the factorized representations describes a latent and disentangled aspect pertinent to a specific latent factor of the graph. Then we propose a novel factor-wise discrimination objective in a contrastive learning manner, which can force the factorized representations to independently reflect the expressive information from different latent factors. Extensive experiments on both synthetic and real-world datasets demonstrate the superiority of our method against several state-of-the-art baselines.

NeurIPS Conference 2021 Conference Paper

Graph Differentiable Architecture Search with Structure Learning

  • Yijian Qin
  • Xin Wang
  • Zeyang Zhang
  • Wenwu Zhu

Discovering ideal Graph Neural Networks (GNNs) architectures for different tasks is labor intensive and time consuming. To save human efforts, Neural Architecture Search (NAS) recently has been used to automatically discover adequate GNN architectures for certain tasks in order to achieve competitive or even better performance compared with manually designed architectures. However, existing works utilizing NAS to search GNN structures fail to answer the question: how NAS is able to select the desired GNN architectures? In this paper, we investigate this question to solve the problem, for the first time. We conduct a measurement study with experiments to discover that gradient based NAS methods tend to select proper architectures based on the usefulness of different types of information with respect to the target task. Our explorations further show that gradient based NAS also suffers from noises hidden in the graph, resulting in searching suboptimal GNN architectures. Based on our findings, we propose a Graph differentiable Architecture Search model with Structure Optimization (GASSO), which allows differentiable search of the architecture with gradient descent and is able to discover graph neural architectures with better performance through employing graph structure learning as a denoising process in the search procedure. The proposed GASSO model is capable of simultaneously searching the optimal architecture and adaptively adjusting graph structure by jointly optimizing graph architecture search and graph structure denoising. Extensive experiments on real-world graph datasets demonstrate that our proposed GASSO model is able to achieve state-of-the-art performance compared with existing baselines.

EAAI Journal 2021 Journal Article

Multiscale domain adaption models and their application in fault transfer diagnosis of planetary gearboxes

  • Qunwang Yao
  • Yi Qin
  • Xin Wang
  • Quan Qian

In recent years, various deep domain adaption (DDA) models, such as deep domain confusion (DDC) and deep adaption network (DAN), are proposed. These models can adapt a trained model in the source domain to new classification tasks in the target domain. However, these classical DDA models suffer from some inherent drawbacks. For example, traditional DDA models can output only one transfer feature (TF) with high dimension and fixed scales, thus possibly losing important information while performing domain confusion operation. Multiscale domain adaption (MSDA) is proposed in this paper to remold and improve the classical DDA models to solve the aforementioned problems. MSDA is a universally applicable strategy that can be embedded into most existing classical DDA models. An MSDA block is designed and constructed on the basis of multiscale convolution networks to replace the last convolutional layer of the original DDA model. The MSDA block has four parallel pipelines consisting of multiscale convolutional and global average pooling operations. Therefore, MSDA can extract more domain-invariant features than the original feature extractor, meanwhile the four low-dimension TFs can simplify the calculation of domain confusion losses. The four TFs are concatenated into a feature vector, and then it is input into the top classifier for fault identification. MSDA can be effectively applied to five classical DDA models and enhance their abilities of domain adaption. The effectiveness and advantage of the proposed MSDA are verified through 18 fault transfer diagnosis tasks of planetary gearboxes.

NeurIPS Conference 2021 Conference Paper

Not All Low-Pass Filters are Robust in Graph Convolutional Networks

  • Heng Chang
  • Yu Rong
  • Tingyang Xu
  • Yatao Bian
  • Shiji Zhou
  • Xin Wang
  • Junzhou Huang
  • Wenwu Zhu

Graph Convolutional Networks (GCNs) are promising deep learning approaches in learning representations for graph-structured data. Despite the proliferation of such methods, it is well known that they are vulnerable to carefully crafted adversarial attacks on the graph structure. In this paper, we first conduct an adversarial vulnerability analysis based on matrix perturbation theory. We prove that the low- frequency components of the symmetric normalized Laplacian, which is usually used as the convolutional filter in GCNs, could be more robust against structural perturbations when their eigenvalues fall into a certain robust interval. Our results indicate that not all low-frequency components are robust to adversarial attacks and provide a deeper understanding of the relationship between graph spectrum and robustness of GCNs. Motivated by the theory, we present GCN-LFR, a general robust co-training paradigm for GCN-based models, that encourages transferring the robustness of low-frequency components with an auxiliary neural network. To this end, GCN-LFR could enhance the robustness of various kinds of GCN-based models against poisoning structural attacks in a plug-and-play manner. Extensive experiments across five benchmark datasets and five GCN-based models also confirm that GCN-LFR is resistant to the adversarial attacks without compromising on performance in the benign situation.

IJCAI Conference 2021 Conference Paper

Stochastic Actor-Executor-Critic for Image-to-Image Translation

  • Ziwei Luo
  • Jing Hu
  • Xin Wang
  • Siwei Lyu
  • Bin Kong
  • Youbing Yin
  • Qi Song
  • Xi Wu

Training a model-free deep reinforcement learning model to solve image-to-image translation is difficult since it involves high-dimensional continuous state and action spaces. In this paper, we draw inspiration from the recent success of the maximum entropy reinforcement learning framework designed for challenging continuous control problems to develop stochastic policies over high dimensional continuous spaces including image representation, generation, and control simultaneously. Central to this method is the Stochastic Actor-Executor-Critic (SAEC) which is an off-policy actor-critic model with an additional executor to generate realistic images. Specifically, the actor focuses on the high-level representation and control policy by a stochastic latent action, as well as explicitly directs the executor to generate low-level actions to manipulate the state. Experiments on several image-to-image translation tasks have demonstrated the effectiveness and robustness of the proposed SAEC when facing high-dimensional continuous space problems.

JBHI Journal 2021 Journal Article

Synergic Adversarial Label Learning for Grading Retinal Diseases via Knowledge Distillation and Multi-Task Learning

  • Lie Ju
  • Xin Wang
  • Xin Zhao
  • Huimin Lu
  • Dwarikanath Mahapatra
  • Paul Bonnington
  • Zongyuan Ge

The need for comprehensive and automated screening methods for retinal image classification has long been recognized. Well-qualified doctors annotated images are very expensive and only a limited amount of data is available for various retinal diseases such as diabetic retinopathy (DR) and age-related macular degeneration (AMD). Some studies show that some retinal diseases such as DR and AMD share some common features like haemorrhages and exudation but most classification algorithms only train those disease models independently when the only single label for one image is available. Inspired by multi-task learning where additional monitoring signals from various sources is beneficial to train a robust model. We propose a method called synergic adversarial label learning (SALL) which leverages relevant retinal disease labels in both semantic and feature space as additional signals and train the model in a collaborative manner using knowledge distillation. Our experiments on DR and AMD fundus image classification task demonstrate that the proposed method can significantly improve the accuracy of the model for grading diseases by 5. 91% and 3. 69% respectively. In addition, we conduct additional experiments to show the effectiveness of SALL from the aspects of reliability and interpretability in the context of medical imaging application.

NeurIPS Conference 2021 Conference Paper

Towards a Unified Game-Theoretic View of Adversarial Perturbations and Robustness

  • Jie Ren
  • Die Zhang
  • Yisen Wang
  • Lu Chen
  • Zhanpeng Zhou
  • Yiting Chen
  • Xu Cheng
  • Xin Wang

This paper provides a unified view to explain different adversarial attacks and defense methods, i. e. the view of multi-order interactions between input variables of DNNs. Based on the multi-order interaction, we discover that adversarial attacks mainly affect high-order interactions to fool the DNN. Furthermore, we find that the robustness of adversarially trained DNNs comes from category-specific low-order interactions. Our findings provide a potential method to unify adversarial perturbations and robustness, which can explain the existing robustness-boosting methods in a principle way. Besides, our findings also make a revision of previous inaccurate understanding of the shape bias of adversarially learned features. Our code is available online at https: //github. com/Jie-Ren/A-Unified-Game-Theoretic-Interpretation-of-Adversarial-Robustness.

NeurIPS Conference 2021 Conference Paper

VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation

  • Linjie Li
  • Jie Lei
  • Zhe Gan
  • Licheng Yu
  • Yen-Chun Chen
  • Rohit Pillai
  • Yu Cheng
  • Luowei Zhou

Most existing video-and-language (VidL) research focuses on a single dataset, or multiple datasets of a single task. In reality, a truly useful VidL system is expected to be easily generalizable to diverse tasks, domains, and datasets. To facilitate the evaluation of such systems, we introduce Video-And-Language Understanding Evaluation (VALUE) benchmark, an assemblage of 11 VidL datasets over 3 popular tasks: (i) text-to-video retrieval; (ii) video question answering; and (iii) video captioning. VALUE benchmark aims to cover a broad range of video genres, video lengths, data volumes, and task difficulty levels. Rather than focusing on single-channel videos with visual information only, VALUE promotes models that leverage information from both video frames and their associated subtitles, as well as models that share knowledge across multiple tasks. We evaluate various baseline methods with and without large-scale VidL pre-training, and systematically investigate the impact of video input channels, fusion methods, and different video representations. We also study the transferability between tasks, and conduct multi-task learning under different settings. The significant gap between our best model and human performance calls for future study for advanced VidL models. VALUE is available at https: //value-benchmark. github. io/.

TCS Journal 2021 Journal Article

Vertex-pancyclicity of the (n,k)-bubble-sort networks

  • Xin Wang
  • Chaoqun Ma
  • Jia Guo

The cycle embedding is an important problem of networks, which can determine the fault tolerance of the networks. A network can be viewed as a graph. Let m be an integer with m ≥ 4, G be a graph and w ∈ V ( G ) be an arbitrary vertex. The graph G is vertex-pancyclic if G has a cycle C l of length l with w ∈ V ( C l ) for every l ∈ { 3, 4, ⋯, | V ( G ) | } and G is m-weak-vertex-pancyclic if G has a cycle C l of length l with w ∈ V ( C l ) for every l ∈ { m, m + 1, ⋯, | V ( G ) | }. Let G ′ be a bipartite graph and w ′ ∈ V ( G ′ ) be an arbitrary vertex. The graph G ′ is vertex-bipancyclic if G ′ has a cycle C h of length h with w ′ ∈ V ( C h ) for any even integer h with 4 ≤ h ≤ | V ( G ′ ) |. In this paper, we study the cycle embedding in the ( n, k ) -bubble-sort network B n, k. We obtain that (1) B n, 1 is vertex-pancyclic for n ≥ 3. (2) B n, n − 1 is vertex-bipancyclic for n ≥ 4. (3) B 4, 2 and B 5, 2 are 6-weak-vertex-pancyclic and B 5, 3 is vertex-pancyclic. (4) B n, k is vertex-pancyclic for n ≥ 6 with 2 ≤ k ≤ n − 2 and every constructed cycle of B n, k contains a residual edge for n ≥ 4 with 2 ≤ k ≤ n − 2.

AAAI Conference 2021 Conference Paper

VSQL: Variational Shadow Quantum Learning for Classification

  • Guangxi Li
  • Zhixin Song
  • Xin Wang

Classification of quantum data is essential for quantum machine learning and near-term quantum technologies. In this paper, we propose a new hybrid quantum-classical framework for supervised quantum learning, which we call Variational Shadow Quantum Learning (VSQL). Our method in particular utilizes the classical shadows of quantum data, which fundamentally represent the side information of quantum data with respect to certain physical observables. Specifically, we first use variational shadow quantum circuits to extract classical features in a convolution way and then utilize a fullyconnected neural network to complete the classification task. We show that this method could sharply reduce the number of parameters and thus better facilitate quantum circuit training. Simultaneously, less noise will be introduced since fewer quantum gates are employed in such shadow circuits. Moreover, we show that the Barren Plateau issue, a significant gradient vanishing problem in quantum machine learning, could be avoided in VSQL. Finally, we demonstrate the efficiency of VSQL in quantum classification via numerical experiments on the classification of quantum states and the recognition of multi-labeled handwritten digits. In particular, our VSQL approach outperforms existing variational quantum classifiers in the test accuracy in the binary case of handwritten digit recognition and notably requires much fewer parameters.

AAAI Conference 2020 Conference Paper

Adaptive Activation Network and Functional Regularization for Efficient and Flexible Deep Multi-Task Learning

  • Yingru Liu
  • Xuewen Yang
  • Dongliang Xie
  • Xin Wang
  • Li Shen
  • Haozhi Huang
  • Niranjan Balasubramanian

Multi-task learning (MTL) is a common paradigm that seeks to improve the generalization performance of task learning by training related tasks simultaneously. However, it is still a challenging problem to search the flexible and accurate architecture that can be shared among multiple tasks. In this paper, we propose a novel deep learning model called Task Adaptive Activation Network (TAAN) that can automatically learn the optimal network architecture for MTL. The main principle of TAAN is to derive flexible activation functions for different tasks from the data with other parameters of the network fully shared. We further propose two functional regularization methods that improve the MTL performance of TAAN. The improved performance of both TAAN and the regularization methods is demonstrated by comprehensive experiments.

AAAI Conference 2020 Conference Paper

Attention-Guide Walk Model in Heterogeneous Information Network for Multi-Style Recommendation Explanation

  • Xin Wang
  • Ying Wang
  • Yunzhi Ling

Explainable Recommendation aims at not only providing the recommended items to users, but also making users aware why these items are recommended. Too many interactive factors between users and items can be used to interpret the recommendation in a heterogeneous information network. However, these interactive factors are usually massive, implicit and noisy. The existing recommendation explanation approaches only consider the single explanation style, such as aspect-level or review-level. To address these issues, we propose a framework (MSRE) of generating the multi-style recommendation explanation with the attention-guide walk model on affiliation relations and interaction relations in the heterogeneous information network. Inspired by the attention mechanism, we determine the important contexts for recommendation explanation and learn joint representation of multi-style user-item interactions for enhancing recommendation performance. Constructing extensive experiments on three real-world datasets verifies the effectiveness of our framework on both recommendation performance and recommendation explanation.

IJCAI Conference 2020 Conference Paper

Classification with Rejection: Scaling Generative Classifiers with Supervised Deep Infomax

  • Xin Wang
  • Siu Ming Yiu

Deep Infomax (DIM) is an unsupervised representation learning framework by maximizing the mutual information between the inputs and the outputs of an encoder, while probabilistic constraints are imposed on the outputs. In this paper, we propose Supervised Deep InfoMax (SDIM), which introduces supervised probabilistic constraints to the encoder outputs. The supervised probabilistic constraints are equivalent to a generative classifier on high-level data representations, where class conditional log-likelihoods of samples can be evaluated. Unlike other works building generative classifiers with conditional generative models, SDIMs scale on complex datasets, and can achieve comparable performance with discriminative counterparts. With SDIM, we could perform classification with rejection. Instead of always reporting a class label, SDIM only makes predictions when test samples' largest class conditional surpass some pre-chosen thresholds, otherwise they will be deemed as out of the data distributions, and be rejected. Our experiments show that SDIM with rejection policy can effectively reject illegal inputs, including adversarial examples and out-of-distribution samples.

AAAI Conference 2020 Conference Paper

DGE: Deep Generative Network Embedding Based on Commonality and Individuality

  • Sheng Zhou
  • Xin Wang
  • Jiajun Bu
  • Martin Ester
  • Pinggang Yu
  • Jiawei Chen
  • Qihao Shi
  • Can Wang

Network embedding plays a crucial role in network analysis to provide effective representations for a variety of learning tasks. Existing attributed network embedding methods mainly focus on preserving the observed node attributes and network topology in the latent embedding space, with the assumption that nodes connected through edges will share similar attributes. However, our empirical analysis of real-world datasets shows that there exist both commonality and individuality between node attributes and network topology. On the one hand, similar nodes are expected to share similar attributes and have edges connecting them (commonality). On the other hand, each information source may maintain individual differences as well (individuality). Simultaneously capturing commonality and individuality is very challenging due to their exclusive nature and existing work fail to do so. In this paper, we propose a deep generative embedding (DGE) framework which simultaneously captures commonality and individuality between network topology and node attributes in a generative process. Stochastic gradient variational Bayesian (SGVB) optimization is employed to infer model parameters as well as the node embeddings. Extensive experiments on four real-world datasets show the superiority of our proposed DGE framework in various tasks including node classification and link prediction.

AAAI Conference 2020 Conference Paper

Generative Adversarial Zero-Shot Relational Learning for Knowledge Graphs

  • Pengda Qin
  • Xin Wang
  • Wenhu Chen
  • Chunyun Zhang
  • Weiran Xu
  • William Yang Wang

Large-scale knowledge graphs (KGs) are shown to become more important in current information systems. To expand the coverage of KGs, previous studies on knowledge graph completion need to collect adequate training instances for newlyadded relations. In this paper, we consider a novel formulation, zero-shot learning, to free this cumbersome curation. For newly-added relations, we attempt to learn their semantic features from their text descriptions and hence recognize the facts of unseen relations with no examples being seen. For this purpose, we leverage Generative Adversarial Networks (GANs) to establish the connection between text and knowledge graph domain: The generator learns to generate the reasonable relation embeddings merely with noisy text descriptions. Under this setting, zero-shot learning is naturally converted to a traditional supervised classification task. Empirically, our method is model-agnostic that could be potentially applied to any version of KG embeddings, and consistently yields performance improvements on NELL and Wiki dataset.

NeurIPS Conference 2020 Conference Paper

Learning by Minimizing the Sum of Ranked Range

  • Shu Hu
  • Yiming Ying
  • Xin Wang
  • Siwei Lyu

In forming learning objectives, one oftentimes needs to aggregate a set of individual values to a single output. Such cases occur in the aggregate loss, which combines individual losses of a learning model over each training sample, and in the individual loss for multi-label learning, which combines prediction scores over all class labels. In this work, we introduce the sum of ranked range (SoRR) as a general approach to form learning objectives. A ranked range is a consecutive sequence of sorted values of a set of real numbers. The minimization of SoRR is solved with the difference of convex algorithm (DCA). We explore two applications in machine learning of the minimization of the SoRR framework, namely the AoRR aggregate loss for binary classification and the TKML individual loss for multi-label/multi-class classification. Our empirical results highlight the effectiveness of the proposed optimization framework and demonstrate the applicability of proposed losses using synthetic and real datasets.

AAAI Conference 2020 Conference Paper

SNEQ: Semi-Supervised Attributed Network Embedding with Attention-Based Quantisation

  • Tao He
  • Lianli Gao
  • Jingkuan Song
  • Xin Wang
  • Kejie Huang
  • Yuanfang Li

Learning accurate low-dimensional embeddings for a network is a crucial task as it facilitates many network analytics tasks. Moreover, the trained embeddings often require a significant amount of space to store, making storage and processing a challenge, especially as large-scale networks become more prevalent. In this paper, we present a novel semi-supervised network embedding and compression method, SNEQ, that is competitive with state-of-art embedding methods while being far more space- and time-efficient. SNEQ incorporates a novel quantisation method based on a self-attention layer that is trained in an end-to-end fashion, which is able to dramatically compress the size of the trained embeddings, thus reduces storage footprint and accelerates retrieval speed. Our evaluation on four real-world networks of diverse characteristics shows that SNEQ outperforms a number of state-of-the-art embedding methods in link prediction, node classification and node recommendation. Moreover, the quantised embedding shows a great advantage in terms of storage and time compared with continuous embeddings as well as hashing methods.

IJCAI Conference 2020 Conference Paper

TransRHS: A Representation Learning Method for Knowledge Graphs with Relation Hierarchical Structure

  • Fuxiang Zhang
  • Xin Wang
  • Zhao Li
  • Jianxin Li

Representation learning of knowledge graphs aims to project both entities and relations as vectors in a continuous low-dimensional space. Relation Hierarchical Structure (RHS), which is constructed by a generalization relationship named subRelationOf between relations, can improve the overall performance of knowledge representation learning. However, most of the existing methods ignore this critical information, and a straightforward way of considering RHS may have a negative effect on the embeddings and thus reduce the model performance. In this paper, we propose a novel method named TransRHS, which is able to incorporate RHS seamlessly into the embeddings. More specifically, TransRHS encodes each relation as a vector together with a relation-specific sphere in the same space. Our TransRHS employs the relative positions among the vectors and spheres to model the subRelationOf, which embodies the inherent generalization relationships among relations. We evaluate our model on two typical tasks, i. e. , link prediction and triple classification. The experimental results show that our TransRHS model significantly outperforms all baselines on both tasks, which verifies that the RHS information is significant to representation learning of knowledge graphs, and TransRHS can effectively and efficiently fuse RHS into knowledge graph embeddings.

YNIMG Journal 2019 Journal Article

Agreeableness modulates group member risky decision-making behavior and brain activity

  • Fang Wang
  • Xin Wang
  • Fenghua Wang
  • Li Gao
  • Hengyi Rao
  • Yu Pan

When facing difficult decisions, people typically believe that “two heads are better than one”. However, findings from previous studies are inconsistent regarding the advantages of decision-making in groups as compared to individual decision-making. We hypothesize that personality traits may modulate risk-taking behavior and brain activity changes during group decision-making. In this study, we used event-related potentials (ERP) with a well-validated balloon analogue risk task (BART) paradigm to examine the relationships between personality traits, decision-making behavior, and brain activity patterns when a cohort of male participants make decisions and take risks both in groups and in isolation. We found significantly increased risk-taking behavior and reduced P300 component during group decision-making as compared to individual decision-making only for participants with high Agreeableness, but not for those with low Agreeableness. Moreover, Agreeableness scores correlated with risk-taking behavior and P300 amplitude changes in group decisions. These findings suggest that Agreeableness personality modulates risk-taking behavior and brain activity when people make decisions in groups, which have implications for future group decision research and practice.

AAMAS Conference 2019 Conference Paper

Coordinated Multiagent Reinforcement Learning for Teams of Mobile Sensing Robots

  • Chao Yu
  • Xin Wang
  • Zhanbo Feng

A mobile sensing robot team (MSRT) is a typical application of multi-agent systems. This paper investigates multiagent reinforcement learning in the MSRT problem. A naive coordinated learning approach is first proposed that uses a coordination graph to model interaction relationships among robots. To further reduce the computation complexity in the context of continuously changing topology caused by robots’ movement, we then propose an on-line transfer learning method that is capable of transferring the past interaction experience and learned knowledge to a new context in a dynamic environment. Simulations verify that the method can achieve reasonable team performance by properly balancing robots’ local selfish interests and global team performance.

AAAI Conference 2019 Conference Paper

Discrete Social Recommendation

  • Chenghao Liu
  • Xin Wang
  • Tao Lu
  • Wenwu Zhu
  • Jianling Sun
  • Steven Hoi

Social recommendation, which aims at improving the performance of traditional recommender systems by considering social information, has attracted broad range of interests. As one of the most widely used methods, matrix factorization typically uses continuous vectors to represent user/item latent features. However, the large volume of user/item latent features results in expensive storage and computation cost, particularly on terminal user devices where the computation resource to operate model is very limited. Thus when taking extra social information into account, precisely extracting K most relevant items for a given user from massive candidates tends to consume even more time and memory, which imposes formidable challenges for efficient and accurate recommendations. A promising way is to simply binarize the latent features (obtained in the training phase) and then compute the relevance score through Hamming distance. However, such a two-stage hashing based learning procedure is not capable of preserving the original data geometry in the real-value space and may result in a severe quantization loss. To address these issues, this work proposes a novel discrete social recommendation (DSR) method which learns binary codes in a unified framework for users and items, considering social information. We further put the balanced and uncorrelated constraints on the objective to ensure the learned binary codes can be informative yet compact, and finally develop an efficient optimization algorithm to estimate the model parameters. Extensive experiments on three real-world datasets demonstrate that DSR runs nearly 5 times faster and consumes only with 1/37 of its real-value competitor’s memory usage at the cost of almost no loss in accuracy.

IJCAI Conference 2019 Conference Paper

Disparity-preserved Deep Cross-platform Association for Cross-platform Video Recommendation

  • Shengze Yu
  • Xin Wang
  • Wenwu Zhu
  • Peng Cui
  • Jingdong Wang

Cross-platform recommendation aims to improve recommendation accuracy through associating information from different platforms. Existing cross-platform recommendation approaches assume all cross-platform information to be consistent with each other and can be aligned. However, there remain two unsolved challenges: i) there exist inconsistencies in cross-platform association due to platform-specific disparity, and ii) data from distinct platforms may have different semantic granularities. In this paper, we propose a cross-platform association model for cross-platform video recommendation, i. e. , Disparity-preserved Deep Cross-platform Association (DCA), taking platform-specific disparity and granularity difference into consideration. The proposed DCA model employs a partially-connected multi-modal autoencoder, which is capable of explicitly capturing platform-specific information, as well as utilizing nonlinear mapping functions to handle granularity differences. We then present a cross-platform video recommendation approach based on the proposed DCA model. Extensive experiments for our cross-platform recommendation framework on real-world dataset demonstrate that the proposed DCA model significantly outperform existing cross-platform recommendation methods in terms of various evaluation metrics.

AAAI Conference 2019 Conference Paper

Dynamic Spatial-Temporal Graph Convolutional Neural Networks for Traffic Forecasting

  • Zulong Diao
  • Xin Wang
  • Dafang Zhang
  • Yingru Liu
  • Kun Xie
  • Shaoyao He

Graph convolutional neural networks (GCNN) have become an increasingly active field of research. It models the spatial dependencies of nodes in a graph with a pre-defined Laplacian matrix based on node distances. However, in many application scenarios, spatial dependencies change over time, and the use of fixed Laplacian matrix cannot capture the change. To track the spatial dependencies among traffic data, we propose a dynamic spatio-temporal GCNN for accurate traffic forecasting. The core of our deep learning framework is the finding of the change of Laplacian matrix with a dynamic Laplacian matrix estimator. To enable timely learning with a low complexity, we creatively incorporate tensor decomposition into the deep learning framework, where real-time traffic data are decomposed into a global component that is stable and depends on long-term temporal-spatial traffic relationship and a local component that captures the traffic fluctuations. We propose a novel design to estimate the dynamic Laplacian matrix of the graph with above two components based on our theoretical derivation, and introduce our design basis. The forecasting performance is evaluated with two realtime traffic datasets. Experiment results demonstrate that our network can achieve up to 25% accuracy improvement.

AAAI Conference 2019 Conference Paper

Learning to Compose Topic-Aware Mixture of Experts for Zero-Shot Video Captioning

  • Xin Wang
  • Jiawei Wu
  • Da Zhang
  • Yu Su
  • William Yang Wang

Although promising results have been achieved in video captioning, existing models are limited to the fixed inventory of activities in the training corpus, and do not generalize to open vocabulary scenarios. Here we introduce a novel task, zeroshot video captioning, that aims at describing out-of-domain videos of unseen activities. Videos of different activities usually require different captioning strategies in many aspects, i. e. word selection, semantic construction, and style expression etc, which poses a great challenge to depict novel activities without paired training data. But meanwhile, similar activities share some of those aspects in common. Therefore, we propose a principled Topic-Aware Mixture of Experts (TAMoE) model for zero-shot video captioning, which learns to compose different experts based on different topic embeddings, implicitly transferring the knowledge learned from seen activities to unseen ones. Besides, we leverage external topic-related text corpus to construct the topic embedding for each activity, which embodies the most relevant semantic vectors within the topic. Empirical results not only validate the effectiveness of our method in utilizing semantic knowledge for video captioning, but also show its strong generalization ability when describing novel activities.

ICML Conference 2019 Conference Paper

Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization

  • Hesham Mostafa
  • Xin Wang

Modern deep neural networks are typically highly overparameterized. Pruning techniques are able to remove a significant fraction of network parameters with little loss in accuracy. Recently, techniques based on dynamic reallocation of non-zero parameters have emerged, allowing direct training of sparse networks without having to pre-train a large dense model. Here we present a novel dynamic sparse reparameterization method that addresses the limitations of previous techniques such as high computational cost and the need for manual configuration of the number of free parameters allocated to each layer. We evaluate the performance of dynamic reallocation methods in training deep convolutional networks and show that our method outperforms previous static and dynamic reparameterization methods, yielding the best accuracy for a fixed parameter budget, on par with accuracies obtained by iteratively pruning a pre-trained dense model. We further investigated the mechanisms underlying the superior generalization performance of the resultant sparse networks. We found that neither the structure, nor the initialization of the non-zero parameters were sufficient to explain the superior performance. Rather, effective learning crucially depended on the continuous exploration of the sparse network structure space during training. Our work suggests that exploring structural degrees of freedom during training is more effective than adding extra parameters to the network.

AAAI Conference 2019 Conference Paper

Recursively Learning Causal Structures Using Regression-Based Conditional Independence Test

  • Hao Zhang
  • Shuigeng Zhou
  • Chuanxu Yan
  • Jihong Guan
  • Xin Wang

This paper addresses two important issues in causality inference. One is how to reduce redundant conditional independence (CI) tests, which heavily impact the efficiency and accuracy of existing constraint-based methods. Another is how to construct the true causal graph from a set of Markov equivalence classes returned by these methods. For the first issue, we design a recursive decomposition approach where the original data (a set of variables) is first decomposed into three small subsets, each of which is then recursively decomposed into three smaller subsets until none of subsets can be decomposed further. Consequently, redundant CI tests can be reduced by inferring causality from these subsets. Advantage of this decomposition scheme lies in two aspects: 1) it requires only low-order CI tests, and 2) it does not violate d-separation. Thus, the complete causality can be reconstructed by merging all the partial results of the subsets. For the second issue, we employ regression-based conditional independence test to check CIs in linear non-Gaussian additive noise cases, which can identify more causal directions by x−E(x|Z)⊥z (or y−E(y|Z)⊥z). Therefore, causal direction learning is no longer limited by the number of returned Vstructures and the consistent propagation. Extensive experiments show that the proposed method can not only substantially reduce redundant CI tests but also effectively distinguish the equivalence classes, thus is superior to the state of the art constraint-based methods in causality inference.

AAMAS Conference 2019 Conference Paper

Reinforcement Learning for Cooperative Overtaking

  • Chao Yu
  • Xin Wang
  • Jianye Hao
  • Zhanbo Feng

This paper solves the cooperative overtaking problem in autonomous driving using reinforcement learning techniques. Learning in such a situation is challenging due to vehicular mobility, which renders a continuously changing environment for each learning vehicle. Without no explicit coordination mechanisms, inefficient behaviors among vehicles might cause fatal uncoordinated outcomes. To solve this issue, we propose two basic coordination models to enable distributed learning of cooperative overtaking maneuvers in a group of vehicles. Extension mechanisms are then presented to make these models workable in more complex and realistic settings with any number of vehicles. Experiments verify that, by capturing the underlying consistency of identities or positions during vehicles’ movement, efficient coordinated behaviors can be achieved simply through vehicles’ local learning interactions.

AAAI Conference 2019 Conference Paper

RS3CIS: Robust Single-Step Spectral Clustering with Intrinsic Subspace

  • Yun Xiao
  • Pengzhen Ren
  • Zhihui Li
  • Xiaojiang Chen
  • Xin Wang
  • Dingyi Fang

Spectral clustering has been widely adopted because it can mine structures between data clusters. The clustering performance of spectral clustering depends largely on the quality of the constructed affinity graph, especially when the data has noise. Subspace learning can transform the original input features to a low-dimensional subspace and help to produce a robust method. Therefore, how to learn an intrinsic subspace and construct a pure affinity graph on a dataset with noise is a challenge in spectral clustering. In order to deal with this challenge, a new Robust Single-Step Spectral Clustering with Intrinsic Subspace (RS3 CIS) method is proposed in this paper. RS3 CIS uses a local representation method that projects the original data into a low-dimensional subspace through a row-sparse transformation matrix and uses the `2, 1-norm of the transformation matrix as a penalty term to achieve noise suppression. In addition, RS3 CIS introduces Laplacian matrix rank constraint so that it can output an affinity graph with an explicit clustering structure, which makes the final clustering result to be obtained in a single-step of constructing an affinity matrix. One synthetic dataset and six real benchmark datasets are used to verify the performance of the proposed method by performing clustering and projection experiments. Experimental results show that RS3 CIS outperforms the related methods with respect to clustering quality, robustness and dimension reduction.

YNIMG Journal 2018 Journal Article

Detectability and reproducibility of the olfactory fMRI signal under the influence of magnetic susceptibility artifacts in the primary olfactory cortex

  • Jiaming Lu
  • Xin Wang
  • Zhao Qing
  • Zhu Li
  • Wen Zhang
  • Ying Liu
  • Lihua Yuan
  • Le Cheng

For human olfactory functional MRI studies, the primary olfactory cortex (POC) suffers severe magnetic susceptibility artifacts, which adversely influences the detectability and reproducibility of the olfactory fMRI data and its clinical applications. The goal of this work is to assess the impacts of the image artifacts on the detectability and reproducibility of the olfactory activation in the POC. The severity of artifacts in the POC were classified into three levels using a Subjective Artifact score (SA_score). The mean temporal signal-to-noise ratio (tSNR) of the fMRI data acquired by a given MRI sequence and olfactory activation (β value) in POC were evaluated and compared to the concurrent activations in the primary visual cortex (Brodmann area 17, BA17) by an odor-visual association paradigm using ninety-nine normal human subjects. Our study revealed that the mean tSNR in POC was above the threshold for reliable detection of the functional activation signal, and, consequently, the mean olfactory activations in the POC were not significantly different from those in BA17. The reproducibility of the activation in the POC was assessed by a random half-split stimulation of a test-retest experiment. The overlap of the activation maps for all the trials (n = 1000) in the POC were not statistically different from that observed in BA17. These results show that the detectability and reproducibility of olfactory activation in the presence of susceptibility artifacts in the POC was at similar level of that in the visual cortex.

IJCAI Conference 2018 Conference Paper

Robust Auto-Weighted Multi-View Clustering

  • Pengzhen Ren
  • Yun Xiao
  • Pengfei Xu
  • Jun Guo
  • Xiaojiang Chen
  • Xin Wang
  • Dingyi Fang

Multi-view clustering has played a vital role in real-world applications. It aims to cluster the data points into different groups by exploring complementary information of multi-view. A major challenge of this problem is how to learn the explicit cluster structure with multiple views when there is considerable noise. To solve this challenging problem, we propose a novel Robust Auto-weighted Multi-view Clustering (RAMC), which aims to learn an optimal graph with exactly k connected components, where k is the number of clusters. ℓ1-norm is employed for robustness of the proposed algorithm. We have validated this in the later experiment. The new graph learned by the proposed model approximates the original graphs of each individual view but maintains an explicit cluster structure. With this optimal graph, we can immediately achieve the clustering results without any further post-processing. We conduct extensive experiments to confirm the superiority and robustness of the proposed algorithm.

NeurIPS Conference 2017 Conference Paper

Flexpoint: An Adaptive Numerical Format for Efficient Training of Deep Neural Networks

  • Urs Köster
  • Tristan Webb
  • Xin Wang
  • Marcel Nassar
  • Arjun Bansal
  • William Constable
  • Oguz Elibol
  • Scott Gray

Deep neural networks are commonly developed and trained in 32-bit floating point format. Significant gains in performance and energy efficiency could be realized by training and inference in numerical formats optimized for deep learning. Despite advances in limited precision inference in recent years, training of neural networks in low bit-width remains a challenging problem. Here we present the Flexpoint data format, aiming at a complete replacement of 32-bit floating point format training and inference, designed to support modern deep network topologies without modifications. Flexpoint tensors have a shared exponent that is dynamically adjusted to minimize overflows and maximize available dynamic range. We validate Flexpoint by training AlexNet, a deep residual network and a generative adversarial network, using a simulator implemented with the \emph{neon} deep learning framework. We demonstrate that 16-bit Flexpoint closely matches 32-bit floating point in training all three models, without any need for tuning of model hyperparameters. Our results suggest Flexpoint as a promising numerical format for future hardware for training and inference.

IJCAI Conference 2017 Conference Paper

Understanding Users' Budgets for Recommendation with Hierarchical Poisson Factorization

  • Yunhui Guo
  • Congfu Xu
  • Hanzhang Song
  • Xin Wang

People consume and rate products in online shopping websites. The historical purchases of customers reflect their personal consumption habits and indicate their future shopping behaviors. Traditional preference-based recommender systems try to provide recommendations by analyzing users' feedback such as ratings and clicks. But unfortunately, most of the existing recommendation algorithms ignore the budget of the users. So they cannot avoid recommending users with products that will exceed their budgets. And they also cannot understand how the users will assign their budgets to different products. In this paper, we develop a generative model named collaborative budget-aware Poisson factorization (CBPF) to connect users' ratings and budgets. The CBPF model is intuitive and highly interpretable. We compare the proposed model with several state-of-the-art budget-unaware recommendation methods on several real-world datasets. The results show the advantage of uncovering users' budgets for recommendation.

AAAI Conference 2016 Conference Paper

Co-Regularized PLSA for Multi-Modal Learning

  • Xin Wang
  • MingChing Chang
  • Yiming Ying
  • Siwei Lyu

Many learning problems in real world applications involve rich datasets comprising multiple information modalities. In this work, we study co-regularized PLSA (coPLSA) as an ef- ficient solution to probabilistic topic analysis of multi-modal data. In coPLSA, similarities between topic compositions of a data entity across different data modalities are measured with divergences between discrete probabilities, which are incorporated as a co-regularizer to augment individual PLSA models over each data modality. We derive efficient iterative learning algorithms for coPLSA with symmetric KL, 2 and 1 divergences as co-regularizers, in each case the essential optimization problem affords simple numerical solutions that entail only matrix arithmetic operations and numerical solution of 1D nonlinear equations. We evaluate the performance of the coPLSA algorithms on text/image cross-modal retrieval tasks, on which they show competitive performance with state-of-the-art methods.

IJCAI Conference 2016 Conference Paper

Constrained Preference Embedding for Item Recommendation

  • Xin Wang
  • Congfu Xu
  • Yunhui Guo
  • Hui Qian

To learn users' preference, their feedback information is commonly modeled as scalars and integrated into matrix factorization (MF) based algorithms. Based on MF techniques, the preference degree is computed by the product of user and item vectors, which is also represented by scalars. On the contrary, in this paper, we express users' feedback as constrained vectors, and call the idea constrained preference embedding (CPE); it means that we regard users, items and all users' behavior as vectors. We find that this viewpoint is more flexible and powerful than traditional MF for item recommendation. For example, by the proposed assumption, users' heterogeneous actions can be coherently mined because all entities and actions can be transferred to a space of the same dimension. In addition, CPE is able to model the feedback of uncertain preference degree. To test our assumption, we propose two models called CPE-s and CPE-ps based on CPE for item recommendation, and show that the popular pair-wise ranking model BPR-MF can be deduced by some restrictions and variations on CPE-s. In the experiments, we will test CPE and the proposed algorithms, and prove their effectiveness.

JMLR Journal 2016 Journal Article

Multiplicative Multitask Feature Learning

  • Xin Wang
  • Jinbo Bi
  • Shipeng Yu
  • Jiangwen Sun
  • Minghu Song

We investigate a general framework of multiplicative multitask feature learning which decomposes individual task's model parameters into a multiplication of two components. One of the components is used across all tasks and the other component is task-specific. Several previous methods can be proved to be special cases of our framework. We study the theoretical properties of this framework when different regularization conditions are applied to the two decomposed components. We prove that this framework is mathematically equivalent to the widely used multitask feature learning methods that are based on a joint regularization of all model parameters, but with a more general form of regularizers. Further, an analytical formula is derived for the across-task component as related to the task- specific component for all these regularizers, leading to a better understanding of the shrinkage effects of different regularizers. Study of this framework motivates new multitask learning algorithms. We propose two new learning formulations by varying the parameters in the proposed framework. An efficient blockwise coordinate descent algorithm is developed suitable for solving the entire family of formulations with rigorous convergence analysis. Simulation studies have identified the statistical properties of data that would be in favor of the new formulations. Extensive empirical studies on various classification and regression benchmark data sets have revealed the relative advantages of the two new formulations by comparing with the state of the art, which provides instructive insights into the feature learning problem with multiple tasks. [abs] [ pdf ][ bib ] &copy JMLR 2016. ( edit, beta )

AAAI Conference 2016 Conference Paper

Recommending Groups to Users Using User-Group Engagement and Time-Dependent Matrix Factorization

  • Xin Wang
  • Roger Donaldson
  • Christopher Nell
  • Peter Gorniak
  • Martin Ester
  • Jiajun Bu

Social networks often provide group features to help users with similar interests associate and consume content together. Recommending groups to users poses challenges due to their complex relationship: user-group affinity is typically measured implicitly and varies with time; similarly, group characteristics change as users join and leave. To tackle these challenges, we adapt existing matrix factorization techniques to learn user-group affinity based on two different implicit engagement metrics: (i) which group-provided content users consume; and (ii) which content users provide to groups. To capture the temporally extended nature of group engagement we implement a time-varying factorization. We test the assertion that latent preferences for groups and users are sparse in investigating elastic-net regularization. Our experiments indicate that the time-varying implicit engagement-based model provides the best top-K group recommendations, illustrating the benefit of the added model complexity.

AAAI Conference 2016 Conference Paper

Write-righter: An Academic Writing Assistant System

  • Yuanchao Liu
  • Xin Wang
  • Ming Liu
  • Xiaolong Wang

Writing academic articles in English is a challenging task for non-native speakers, as more effort has to be spent to enhance their language expressions. This paper presents an academic writing assistant system called Write-righter, which can provide real-time hint and recommendation by analyzing the input context. To achieve this goal, some novel strategies, e. g. , semantic extension based sentence retrieval and LDA based sentence structure identification have been proposed. Write-righter is expected to help people express their ideas correctly by recommending top N most possible expressions.

AAAI Conference 2015 Conference Paper

Exploring Social Context for Topic Identification in Short and Noisy Texts

  • Xin Wang
  • Ying Wang
  • Wanli Zuo
  • Guoyong Cai

With the pervasion of social media, topic identification in short texts attracts increasing attention in recent years. However, in nature the texts of social media are short and noisy, and the structures are sparse and dynamic, resulting in difficulty to identify topic categories exactly from online social media. Inspired by social science findings that preference consistency and social contagion are observed in social media, we investigate topic identification in short and noisy texts by exploring social context from the perspective of social sciences. In particular, we present a mathematical optimization formulation that incorporates the preference consistency and social contagion theories into a supervised learning method, and conduct feature selection to tackle short and noisy texts in social media, which result in a Sociological framework for Topic Identification (STI). Experimental results on real-world datasets from Twitter and Citation Network demonstrate the effectiveness of the proposed framework. Further experiments are conducted to understand the importance of social context in topic identification.

AAAI Conference 2015 Conference Paper

Modeling Status Theory in Trust Prediction

  • Ying Wang
  • Xin Wang
  • Jiliang Tang
  • Wanli Zuo
  • Guoyong Cai

With the pervasion of social media, trust has been playing more of an important role in helping online users collect reliable information. In reality, user-specified trust relations are often very sparse; hence, inferring unknown trust relations has attracted increasing attention in recent years. Social status is one of the most important concepts in trust, and status theory is developed to help us understand the important role of social status in the formation of trust relations. In this paper, we investigate how to exploit social status in trust prediction by modeling status theory. We first verify status theory in trust relations, then provide a principled way to model it mathematically, and propose a novel framework sTrust which incorporates status theory for trust prediction. Experimental results on real-world datasets demonstrate the effectiveness of the proposed framework. Further experiments are conducted to understand the importance of status theory in trust prediction.

IJCAI Conference 2015 Conference Paper

Recommendation Algorithms for Optimizing Hit Rate, User Satisfaction and Website Revenue

  • Xin Wang
  • Yunhui Guo
  • Congfu Xu

We usually use hit rate to measure the performance of item recommendation algorithms. In addition to hit rate, we consider additional two important factors which are ignored by most previous works. First, we consider whether users are satisfied with the recommended items. It is possible that a user has bought an item but dislikes it. Hence high hit rate may not reflect high customer satisfaction. Second, we consider whether the website retailers are satisfied with the recommendation results. If a customer is interested in two products and wants to buy one of them, it may be better to suggest the item which can help bring more profit. Therefore, a good recommendation algorithm should not only consider improving hit rate but also consider optimizing user satisfaction and website revenue. In this paper, we propose two algorithms for the above purposes and design two modified hit rate based metrics to measure them. Experimental results on 10 real-world datasets show that our methods can not only achieve better hit rate, but also improve user satisfaction and website revenue comparing with the state-of-the-art models.

NeurIPS Conference 2014 Conference Paper

On Multiplicative Multitask Feature Learning

  • Xin Wang
  • Jinbo Bi
  • Shipeng Yu
  • Jiangwen Sun

We investigate a general framework of multiplicative multitask feature learning which decomposes each task's model parameters into a multiplication of two components. One of the components is used across all tasks and the other component is task-specific. Several previous methods have been proposed as special cases of our framework. We study the theoretical properties of this framework when different regularization conditions are applied to the two decomposed components. We prove that this framework is mathematically equivalent to the widely used multitask feature learning methods that are based on a joint regularization of all model parameters, but with a more general form of regularizers. Further, an analytical formula is derived for the across-task component as related to the task-specific component for all these regularizers, leading to a better understanding of the shrinkage effect. Study of this framework motivates new multitask learning algorithms. We propose two new learning formulations by varying the parameters in the proposed framework. Empirical studies have revealed the relative advantages of the two new formulations by comparing with the state of the art, which provides instructive insights into the feature learning problem with multiple tasks.

YNIMG Journal 2013 Journal Article

Edited magnetic resonance spectroscopy detects an age-related decline in brain GABA levels

  • Fei Gao
  • Richard A.E. Edden
  • Muwei Li
  • Nicolaas A.J. Puts
  • Guangbin Wang
  • Cheng Liu
  • Bin Zhao
  • Huiquan Wang

Gamma-aminobutyric acid (GABA) is the primary inhibitory neurotransmitter in the brain. Although measurements of GABA levels in vivo in the human brain using edited proton magnetic resonance spectroscopy (1H-MRS) have been established for some time, it is has not been established how regional GABA levels vary with age in the normal human brain. In this study, 49 healthy men and 51 healthy women aged between 20 and 76years were recruited and J-difference edited spectra were recorded at 3T to determine the effect of age on GABA levels, and to investigate whether there are regional and gender differences in GABA in mesial frontal and parietal regions. Because the signal detected at 3. 02ppm using these experimental parameters is also expected to contain contributions from both macromolecules (MM) and homocarnosine, in this study the signal is labeled GABA+ rather than GABA. Significant negative correlations were observed between age and GABA+ in both regions studied (GABA+/Cr: frontal region, r=−0. 68, p<0. 001, parietal region, r=−0. 54, p<0. 001; GABA+/NAA: frontal region, r=−0. 58, p<0. 001, parietal region, r=−0. 49, p<0. 001). The decrease in GABA+ with age in the frontal region was more rapid in women than men. Evidence of a measureable decline in GABA is important in considering the neurochemical basis of the cognitive decline that is associated with normal aging.

NeurIPS Conference 2013 Conference Paper

On Algorithms for Sparse Multi-factor NMF

  • Siwei Lyu
  • Xin Wang

Nonnegative matrix factorization (NMF) is a popular data analysis method, the objective of which is to decompose a matrix with all nonnegative components into the product of two other nonnegative matrices. In this work, we describe a new simple and efficient algorithm for multi-factor nonnegative matrix factorization problem ({mfNMF}), which generalizes the original NMF problem to more than two factors. Furthermore, we extend the mfNMF algorithm to incorporate a regularizer based on Dirichlet distribution over normalized columns to encourage sparsity in the obtained factors. Our sparse NMF algorithm affords a closed form and an intuitive interpretation, and is more efficient in comparison with previous works that use fix point iterations. We demonstrate the effectiveness and efficiency of our algorithms on both synthetic and real data sets.

TCS Journal 2012 Journal Article

Almost optimal distributed M2M multicasting in wireless mesh networks

  • Qin Xin
  • Fredrik Manne
  • Yan Zhang
  • Xin Wang

Wireless Mesh Networking (WMN) is an emerging communication paradigm to enable resilient, cost-efficient and reliable services for the future-generation wireless networks. In this paper, we study the problem of multipoint-to-multipoint (M2M) multicasting in a WMN which aims to use the minimum number of time slots to exchange messages among a group of k mesh nodes in a multi-hop WMN with n mesh nodes. We study the M2M multicasting problem in a distributed environment where each participant only knows that there are k participants and it does not know who are other k − 1 participants among n mesh nodes. It is known that the computation of an optimal M2M multicasting schedule isNP-hard. We present a fully distributed deterministic algorithm for such an M2M multicasting problem and analyze its time complexity. We show that if the maximum hop distance between any two out of the k participants is d, then the studied M2M multicasting problem can be solved in time O ( d log 2 n + k log 3 n log k ) with a polynomial-time computation, which is an almost optimal scheme due to the lower bound Ω ( d + k log n log k ) given by Chlebus et al. (2009) [5]. Our algorithm also improves the currently best known result with running time O ( d log 2 n + k log 4 n ) by Gąsieniec et al. (2006) [13]. In this paper, we also propose a distributed deterministic algorithm which accomplishes the M2M multicasting in time O ( d + k ) with a polynomial-time computation in unit disk graphs. This is an asymptotically optimal algorithm in the sense that there exists a WMN topology, e. g. , a line, a ring, a star or a complete graph, in which the M2M multicasting cannot be completed in less than Ω ( d + k ) units of time.

YNIMG Journal 2010 Journal Article

A multiple-plane approach to measure the structural properties of functionally active regions in the human cortex

  • Xin Wang
  • Sarah N. Garfinkel
  • Anthony P. King
  • Mike Angstadt
  • Michael J. Dennis
  • Hong Xie
  • Robert C. Welsh
  • Marijo B. Tamburrino

Advanced magnetic resonance imaging (MRI) techniques provide the means of studying both the structural and the functional properties of various brain regions, allowing us to address the relationship between the structural changes in human brain regions and the activity of these regions. However, analytical approaches combining functional (fMRI) and structural (sMRI) information are still far from optimal. In order to improve the accuracy of measurement of structural properties in active regions, the current study tested a new analytical approach that repeated a surface-based analysis at multiple planes crossing different depths of cortex. Twelve subjects underwent a fear conditioning study. During these tasks, fMRI and sMRI scans were acquired. The fMRI images were carefully registered to the sMRI images with an additional correction for cortical borders. The fMRI images were then analyzed with the new multiple-plane surface-based approach as compared to the volume-based approach, and the cortical thickness and volume of an active region were measured. The results suggested (1) using an additional correction for cortical borders and an intermediate template image produced an acceptable registration of fMRI and sMRI images; (2) surface-based analysis at multiple depths of cortex revealed more activity than the same analysis at any single depth; (3) projection of active surface vertices in a ribbon fashion improved active volume estimates; and (4) correction with gray matter segmentation removed non-cortical regions from the volumetric measurement of active regions. In conclusion, the new multiple-plane surface-based analysis approaches produce improved measurement of cortical thickness and volume of active brain regions. These results support the use of novel approaches for combined analysis of functional and structural neuroimaging.

IJCAI Conference 2009 Conference Paper

  • Fei Wang
  • Xin Wang
  • Tao Li

Clustering aggregation has emerged as an important extension of the classical clustering problem. It refers to the situation in which a number of different (input) clusterings have been obtained for a particular data set and it is desired to aggregate those clustering results to get a better clustering solution. In this paper, we propose a unified framework to solve the clustering aggregation problem, where the aggregated clustering result is obtained by minimizing the (weighted) sum of the Bregman divergence between it and all the input clusterings. Moreover, under our algorithm framework, we also propose a novel cluster aggregation problem where some must-link and cannot-link constraints are given in addition to the input clusterings. Finally the experimental results on some real world data sets are presented to show the effectiveness of our method.

TCS Journal 2008 Journal Article

Online scheduling of equal-processing-time task systems

  • Yumei Huo
  • Joseph Y.-T. Leung
  • Xin Wang

We consider the problem of online scheduling a set of equal-processing-time tasks with precedence constraints so as to minimize the makespan. For arbitrary precedence constraints, it is known that any list scheduling algorithm has a competitive ratio of 2 − 1 / m, where m is the number of machines. We show that for intree precedence constraints, Hu’s algorithm yields an asymptotic competitive ratio of 3/2.

IROS Conference 2003 Conference Paper

A study on geometric algorithms for real-time grasping force optimization

  • Jijie Xu
  • Guanfeng Liu 0002
  • Xin Wang
  • Zexiang Li 0001

In this paper we propose several strategies for selecting such a step size according to the properties of each algorithm and a method for searching a valid initial point. By investigating the structure of the affine-scaling vector fields associated with the optimization problem, we give a detailed convergence analysis of these algorithms. Simulation and experimental results show the different performance of these algorithms from computation time and convergence rates.

IROS Conference 2003 Conference Paper

A study on quality functions for grasp synthesis and fixture planning

  • Jijie Xu
  • Guanfeng Liu 0002
  • Xin Wang
  • Zexiang Li 0001

Planning a proper set of contact points on a given object/workpiece so as to satisfy a certain optimality criterion is a common problem in grasp synthesis for multifingered robotic hands and in fixture planning for manufacturing automation. In this paper, we formulate the grasp planning problem as optimization problems with respect to two grasp quality functions. For real-time computation, a simplified Min-analytic-center problem is proposed. Simulation and experimental results illustrate the validity of the proposed approach for optimal grasp planning.

NeurIPS Conference 2001 Conference Paper

Batch Value Function Approximation via Support Vectors

  • Thomas Dietterich
  • Xin Wang

We present three ways of combining linear programming with the kernel trick to find value function approximations for reinforcement learning. One formulation is based on SVM regression; the second is based on the Bellman equation; and the third seeks only to ensure that good moves have an advantage over bad moves. All formu(cid: 173) lations attempt to minimize the number of support vectors while fitting the data. Experiments in a difficult, synthetic maze problem show that all three formulations give excellent performance, but the advantage formulation is much easier to train. Unlike policy gradi(cid: 173) ent methods, the kernel methods described here can easily 'adjust the complexity of the function approximator to fit the complexity of the value function.

NeurIPS Conference 2001 Conference Paper

Stabilizing Value Function Approximation with the BFBP Algorithm

  • Xin Wang
  • Thomas Dietterich

We address the problem of non-convergence of online reinforcement learning algorithms (e. g. , Q learning and SARSA(A)) by adopt(cid: 173) ing an incremental-batch approach that separates the exploration process from the function fitting process. Our BFBP (Batch Fit to Best Paths) algorithm alternates between an exploration phase (during which trajectories are generated to try to find fragments of the optimal policy) and a function fitting phase (during which a function approximator is fit to the best known paths from start states to terminal states). An advantage of this approach is that batch value-function fitting is a global process, which allows it to address the tradeoffs in function approximation that cannot be handled by local, online algorithms. This approach was pioneered by Boyan and Moore with their GROWSUPPORT and ROUT al(cid: 173) gorithms. We show how to improve upon their work by applying a better exploration process and by enriching the function fitting procedure to incorporate Bellman error and advantage error mea(cid: 173) sures into the objective function. The results show improved per(cid: 173) formance on several benchmark problems.

IROS Conference 1999 Conference Paper

Feasible online learning neural and fuzzy intelligent control for a mobile vehicle

  • Xin Wang
  • Masanori Sugisaka

The goal of the paper is to submit an idea from the practical point of view for control technologies. Our basic idea is to make a mobile vehicle (NMV) be capable of evaluating results of actions and modifying the inputs of the controller itself in order to drive optimal or near-optimal paths. The NMV system consists of a charge-coupled device (CCD) camera, a Von Neumann type computer and a neurocomputer, and image processing units etc. The hardware neurocomputer RN-2000 is the kernel part of the system. The purpose of the paper is to solve the problem of how to realize the hardware neurocomputer by backpropagation (BP) neural network learning online. The strategy presented in the paper is based on modifying the past patterns and adjusting the content of the driving patterns by a new algorithm. Learning happens during the driving procedure of the mobile vehicle. This research shows the possibility of the neurocomputer which the BP neural network is inside can learn human knowledge online by the aid of software. Some words about our past researches are also given to help understanding the new method.

NeurIPS Conference 1993 Conference Paper

Asynchronous Dynamics of Continuous Time Neural Networks

  • Xin Wang
  • Qingnan Li
  • Edward Blum

Motivated by mathematical modeling, analog implementation and distributed simulation of neural networks, we present a definition of asynchronous dynamics of general CT dynamical systems defined by ordinary differential equations, based on notions of local times and communication times. We provide some preliminary results on globally asymptotical convergence of asynchronous dynamics for contractive and monotone CT dynamical systems. When ap(cid: 173) plying the results to neural networks, we obtain some conditions that ensure additive-type neural networks to be asynchronizable.

v2026.09.13