Arrow Research search

Author name cluster

Wei Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

276 papers
2 author rows

Possible papers

276

EAAI Journal 2026 Journal Article

A causal generative model-based optimal scheduling method for blast furnace gas system considering unknown scenarios

  • Feng Jin
  • Xiaoxue Wang
  • Jun Zhao
  • Wei Wang

Blast furnace gas is a significant category of byproduct energy source produced during the ironmaking process, and its rational utilization is crucial to improving energy efficiency in steel plants. However, frequent operator interventions continually create new operating scenarios, which makes it difficult for traditional methods to maintain effective scheduling. Existing scheduling methods based on optimization and generative adversarial networks (GANs) rely excessively on historical scenarios and fail to capture the explicit causal relationships among the key factors, thus restricting their applicability for scheduling under unknown conditions. To tackle such an issue, an optimal scheduling method for BFG system based on an improved causal generative model, which is capable of generating diverse and physically consistent scenarios, is proposed in this study. Each scenario is characterized by three interpretable factors, i. e. gas tank level, generation-consumption flow difference, and consumption of adjustable units. A causal conditional Wasserstein GAN (Causal-CWGAN) is then constructed by embedding a process-informed adjacency matrix and a differentiable acyclicity constraint into the WGAN-GP framework. In addition, a correction model combined with mechanism-based rationality rules is adopted to further optimize the consumption and filters unreasonable scenarios. Subsequently, the tank-level prediction is performed to update the generated scenario set and derive practical adjustment suggestions. Experimental results on real data from a steel enterprise show that, compared with the WGAN and the MGAN methods, the proposed one yields smaller Wasserstein distances, generates more rational scenarios, and provides adjustment strategies that can stably keep the gas tank level within the safety operating range.

EAAI Journal 2026 Journal Article

A dual-branch fusion network based on Convolutional Neural Network and Mamba for outdoor fire detection

  • Jin He
  • Wei Wang
  • Zhijian Gou
  • Xiaozhao Jin
  • Gexiang Zhang
  • Mianxiong Dong

Outdoor fires (OFs) pose significant threats to human safety, property, and ecological stability. However, existing detection algorithms often suffer from performance degradation in complex real-world environments, resulting in high false alarm and missed detection rates. To address these limitations, we propose a novel Convolutional Neural Network (CNN)-Mamba Dual-branch Fusion Network (CMDFNet). The framework consists of a CNN branch for fine-grained local feature extraction and a Vision Mamba (VMamba) branch for efficient global context modeling. To further enhance representation, we introduce an eight-directional selective scanning strategy for irregular flame contour perception, a selective state space mechanism for dynamic temporal adaptation, and a CNN-VMamba fusion block to facilitate deep interaction between local and global features. Extensive experiments show that CMDFNet achieves superior performance, surpassing state-of-the-art methods by 2. 0% in detection accuracy and 2. 1% in recall, while maintaining high computational efficiency. The model has also been deployed at over 20 monitoring sites, where it successfully detected more than 1600 OF events, further confirming its effectiveness and robustness in complex, real-world environments. Code is available at: https: //github. com/hejinsome/CMDFNet.

EAAI Journal 2026 Journal Article

A novel joint path planning method for drones with hopfield neural network and Gaussian sampling

  • Wei Wang
  • Yuxin Liao
  • Zequan Xu
  • Junqi Guan
  • Guoliang Zhang
  • Zhengyang Li

How to ensure the safe flight of unmanned aerial vehicles (UAVs) in complex environments with static and dynamic obstacles can be a big challenge. Therefore, a new joint path planning method with integrated of an improved non-dominated sorting genetic algorithm (NSGA-III) and an improved artificial potential field (APF) is proposed, which exhibits high computation efficiency and excellent global optimum searching capability. Three objectives including flying efficiency, stability and obstacles avoidance are constructed to meet the strict requirements of the complex environment. Greatest novel features of this new method include three aspects, and they are: 1) Hopfield neural network is introduced to replace the random strategy to generate the initial population of NSGA-III, which can enhance the iteration efficiency significantly, 2) Gaussian sampling is designed to create a new potential solution space that is adjacent to the local optimum, which can guide the further searching for a better result and 3) virtual path points as well as tracked distances are both developed in APF to help escape from the stuck area during the dynamic obstacle avoidance. A comprehensive compared study is also carried out by use of popular traditional NSGA-III, improved A∗ and RRT∗ algorithms. The joint planning algorithm can ameliorate path length and flying efficiency by an average of 12. 6% and 67. 5%, respectively. Simulations and experiments approve the validation.

AAAI Conference 2026 Conference Paper

Adaptive Hallucination Alleviation in Multimodal Large Language Models: From Strategic Data Selection to Severity-Guided Training

  • Yuanyi Xu
  • Xiangru Zhu
  • Sihang Jiang
  • Zhixu Li
  • Bei Yang
  • Xiaoxiao Xu
  • Yanghua Xiao
  • Wei Wang

Multimodal Large Language Models (MLLMs) have recently achieved strong performance across a variety of multimodal tasks. However, they still suffer from various forms of hallucination, which hinder their practical deployment. Prior approaches often struggle to efficiently construct high-quality hallucination-related samples and to process them in a fine-grained manner, resulting in limited effectiveness in hallucination alleviation. To address this issue, we propose a data sampling strategy that selects samples better suited for hallucination-oriented training, thereby enhancing training effectiveness. In addition, we introduce a quantitative method for measuring hallucination severity and assign individualized weights to training samples accordingly. Building on this, we present Hallucination-Differentiated Direct Preference Optimization (HD-DPO), a novel preference optimization framework. During fine-tuning, HD-DPO incorporates these weights into both the formulation of customized loss functions and the modulation of localized visual attention, enabling fine-grained optimization. Experimental results demonstrate that our method outperforms existing fine-tuning strategies across multiple benchmarks and generalizes well to diverse MLLM architectures, effectively reducing hallucination rates and enhancing overall model performance.

AAAI Conference 2026 Conference Paper

Ambiguity-aware Truncated Flow Matching for Ambiguous Medical Image Segmentation

  • Fanding Li
  • Xiangyu Li
  • Xianghe Su
  • Xingyu Qiu
  • Suyu Dong
  • Wei Wang
  • Kuanquan Wang
  • Gongning Luo

A simultaneous enhancement of accuracy and diversity of predictions remains a challenge in ambiguous medical image segmentation (AMIS) due to the inherent trade-offs. While truncated diffusion probabilistic models (TDPMs) hold strong potential with a paradigm optimization, existing TDPMs suffer from entangled accuracy and diversity of predictions with insufficient fidelity and plausibility. To address the aforementioned challenges, we propose Ambiguity-aware Truncated Flow Matching (ATFM), which introduces a novel inference paradigm and dedicated model components. Firstly, we propose Data-Hierarchical Inference, a redefinition of AMIS-specific inference paradigm, which enhances accuracy and diversity at data-distribution and data-sample level, respectively, for an effective disentanglement. Secondly, Gaussian Truncation Representation (GTR) is introduced to enhance both fidelity of predictions and reliability of truncation distribution, by explicitly modeling it as a Gaussian distribution at Ttrunc instead of using sampling-based approximations. Thirdly, Segmentation Flow Matching (SFM) is proposed to enhance the plausibility of diverse predictions by extending semantic-aware flow transformation in Flow Matching (FM). Comprehensive evaluations on LIDC and ISIC3 datasets demonstrate that ATFM outperforms SOTA methods and simultaneously achieves a more efficient inference. ATFM improves GED and HM-IoU by up to 12% and 7.3% compared to advanced methods.

JBHI Journal 2026 Journal Article

CATransformer: A Cycle-Aware Transformer for High-Fidelity ECG Generation From PPG

  • Xiaoyan Yuan
  • Wei Wang
  • Xiaohe Li
  • Yuanting Zhang
  • Xiping Hu
  • M. Jamal Deen

Electrocardiography (ECG) is the gold standard for monitoring heart function and is crucial for preventing the worsening of cardiovascular diseases (CVDs). However, the inconvenience of ECG acquisition poses challenges for long-term continuous monitoring. Consequently, researchers have explored non-invasive and easily accessible photoplethysmography (PPG) as an alternative, converting it into ECG. Previous studies have focused on peaks or simple mapping to generate ECG, ignoring the inherent periodicity of cardiovascular signals. This results in an inability to accurately extract physiological information during the cycle, thus compromising the generated ECG signals' clinical utility. To this end, we introduce a novel PPG-to-ECG translation model called CATransformer, capable of adaptive modeling based on the cardiac cycle. Specifically, CATransformer automatically extracts the cycle using a cycle-aware module and creates multiple semantic views of the cardiac cycle. It leverages a transformer to capture detailed features within each cycle and the dynamics across cycles. Our method outperforms existing approaches, exhibiting the lowest RMSE across five paired PPG-ECG databases. Additionally, extensive experiments are conducted on four cardiovascular-related tasks to assess the clinical utility of the generated ECG, achieving consistent state-of-the-art performance. Experimental results confirm that CATransformer generates highly faithful ECG signals while preserving their physiological characteristics.

JBHI Journal 2026 Journal Article

Causality-Adjusted Data Augmentation for Domain Continual Medical Image Segmentation

  • Zhanshi Zhu
  • Qing Dong
  • Gongning Luo
  • Wei Wang
  • Suyu Dong
  • Kuanquan Wang
  • Ye Tian
  • Guohua Wang

In domain continual medical image segmentation, distillation-based methods mitigate catastrophic forgetting by continuously reviewing old knowledge. However, these approaches often exhibit biases towards both new and old knowledge simultaneously due to confounding factors, which can undermine segmentation performance. To address these biases, we propose the Causality-Adjusted Data Augmentation (CauAug) framework, introducing a novel causal intervention strategy called the Texture-Domain Adjustment Hybrid-Scheme (TDAHS) alongside two causality-targeted data augmentation approaches: the Cross Kernel Network (CKNet) and the Fourier Transformer Generator (FTGen). (1) TDAHS establishes a domain-continual causal model that accounts for two types of knowledge biases by identifying irrelevant local textures (L) and domain-specific features (D) as confounders. It introduces a hybrid causal intervention that combines traditional confounder elimination with a proposed replacement approach to better adapt to domain shifts, thereby promoting causal segmentation. (2) CKNet eliminates confounder L to reduce biases in new knowledge absorption. It decreases reliance on local textures in input images, forcing the model to focus on relevant anatomical structures and thus improving generalization. (3) FTGen causally intervenes on confounder D by selectively replacing it to alleviate biases that impact old knowledge retention. It restores domain-specific features in images, aiding in the comprehensive distillation of old knowledge. Our experiments show that CauAug significantly mitigates catastrophic forgetting and surpasses existing methods in various medical image segmentation tasks.

TIST Journal 2026 Journal Article

Clue and Context Fusion for Sarcasm Detection with Large Multimodal Models

  • Qiuyu Li
  • Yushan Pan
  • Ding Wang
  • Wei Wang
  • Xiaowei Huang
  • Zhijie Xu

Detecting sarcasm in social media is fundamentally different from general VLM benchmarks: it is a pragmatic contradiction problem in which the literal signal in one modality is intentionally misaligned with the intended meaning, while dominant pre-training (e.g., CLIP-style contrastive agreement) biases models toward modality alignment rather than incongruity detection. We present SCARF, a contradiction-aware framework that equips large multimodal models with explicit sarcasm cues and context-sensitive retrieval. SCARF constructs coarse scene cues and fine localized evidence via tag-constrained QA, then distills them with visual tokens into a [FUSION] control vector for the LLM; a label-contrastive retriever supplies type- and context-matched exemplars, and a local multi-view encoder surfaces micro-cues. With the same backbone and training data, SCARF attains 87.92% Acc/86.67% F1 on MMSD2.0 and 77.14% Acc/76.44% F1 zero-shot on XDMSD, outperforming a comparably fine-tuned LLaVA-1.5. Ablations show sarcasm clue fusion is the main driver of gains, and tag-constrained QA improves rationale grounding and reduces hallucinations.

JBHI Journal 2026 Journal Article

CoMIL: A Contrastive CNN-Transformer Framework with Multi-Instance Learning for Whole-Slide Pathology Image Classification

  • Bowen Liu
  • Hongbo Zhu
  • Xiaotong Wei
  • Chuan Lin
  • Wei Wang

Whole slide image (WSI) classification faces challenges due to gigapixel scale and weak supervision, often struggling to balance global context with local details. We propose CoMIL, a dual-branch framework based on symmetric mutual learning. Firstly, to resolve the dilemma where single-stream networks struggle to simultaneously capture global context and fine-grained details, we employ dual parallel pathways: a Transformer branch models long-range instance dependencies, while a CNN branch captures localized tissue morphology. Secondly, to address spatial information loss, we design a Hyper Positional Generator (HyperPG). This module integrates multi-scale adaptive mechanisms with deformable convolutions, enhancing spatial awareness with linear complexity. Finally, to improve model robustness against weak label noise, bidirectional learning between branches is achieved through KL divergence minimization. Extensive experiments show that our proposed method achieves an area under the curve of 98. 6% and an accuracy of 95. 3% on the Camelyon16 dataset, and an area under the curve of 98. 8% and an accuracy of 93. 3% on the TCGA_Kidney dataset, surpassing the performance of known advanced WSI classification methods.

AIIM Journal 2026 Journal Article

Context-aware heterogeneous graph neural network for multi-level description and invasiveness prediction in renal cell carcinoma

  • Xiaoming Jiang
  • Guoying Ji
  • Ye Yan
  • Xiongjun Ye
  • Chao Liang
  • Bao Li
  • Wei Wang
  • Shudong Zhang

The invasiveness prediction in renal cell carcinoma (RCC) is of significant importance for the decision of clinical surgical plans and the patients' prognosis. Currently, besides invasive pathological assessment, it mainly relies on observation through computed tomography (CT) imaging. However, limitations of human vision and qualitative descriptions restrict the accuracy of the diagnosis of renal sinus invasion (RSI). Recently, artificial intelligence approaches have shown promising prospects in cancer diagnosis. Due to the complex imaging characteristics of invasiveness, prediction models that only focus on tumor regions are inadequate, requiring comprehensive evaluation of intratumoral heterogeneity, peritumoral information, and the kidney in which the tumor resides. Therefore, in this study, we propose a context-aware heterogeneous graph neural network for multi-level description and invasiveness prediction in RCC. The superiority of the proposed model lies in its ability to integrate imaging features at multi-level, and to learn disturbance invariant features through a data-driven diffusion perturbation strategy. To evaluate the effectiveness and generalization of our model, we conduct extensive experiments on a multi-center dataset (including CT scan images of 437 patients) to compare our model with a series of state-of-the-art (SOTA) classification models. The experimental results show the superiority of our model for RSI classification ( AUC = 0. 88 ). Additionally, we also perform a comparative study with clinical experts, and the proposed method is significantly better than existing assessment methods and clinical experts ( p < 0. 05 ). In general, our work provides an effective assessment tool for automated diagnosis of RSI in RCC and also offers new insights for constructing more precise tumor prediction models.

AAAI Conference 2026 Conference Paper

CoordAR: One-Reference 6D Pose Estimation of Novel Objects via Autoregressive Coordinate Map Generation

  • Dexin Zuo
  • Ang Li
  • Wei Wang
  • Wenxian Yu
  • Danping Zou

Object 6D pose estimation, a crucial task for robotics and augmented reality applications, becomes particularly challenging when dealing with novel objects whose 3D models are not readily available. To reduce dependency on 3D models, recent studies have explored one-reference-based pose estimation, which requires only a single reference view instead of a complete 3D model. However, existing methods that rely on real-valued coordinate regression suffer from limited global consistency due to the local nature of convolutional architectures and face challenges in symmetric or occluded scenarios owing to a lack of uncertainty modeling. We present CoordAR, a novel autoregressive framework for one-reference 6D pose estimation of unseen objects. CoordAR formulates 3D-3D correspondences between the reference and query views as a map of discrete tokens, which is obtained in an autoregressive and probabilistic manner. To enable accurate correspondence regression, CoordAR introduces 1) a novel coordinate map tokenization that enables probabilistic prediction over discretized 3D space; 2) a modality-decoupled encoding strategy that separately encodes RGB appearance and coordinate cues; and 3) an autoregressive transformer decoder conditioned on both position-aligned query features and the partially generated token sequence. With these novel mechanisms, CoordAR significantly outperforms existing methods on multiple benchmarks and demonstrates strong robustness to symmetry, occlusion, and other challenges in real-world tests.

AAAI Conference 2026 Conference Paper

Decision-Driven Orthogonal Learning with Complementary Feature Mining for Robust Synthetic Image Detection

  • Kai Li
  • Wei Wang
  • Linchao Zhang
  • Siying Zhu
  • Wenqi Ren

The widespread and inconsistent compression applied by Online Social Networks severely degrades the performance of synthetic image detectors. We attribute this degradation to two main issues: 1) the model confuses forgery artifacts with compression artifacts, and 2) compression erodes crucial discriminative high-frequency details. Existing methods suppress compression features during training but overlook the overlap between compression features and forgery-related features, leading to the unintended removal of forgery traces. To address artifact confusion, we introduce a Decision-Driven Orthogonal Constraint, which defines a classification decision axis pointing from the real class centroid to the forged class centroid. This constraint enforces compression artifacts to be orthogonal to the decision axis, mitigating their interference with forgery detection without entirely removing them, thus preventing the suppression of forgery-related features. To mitigate the erosion of high-frequency details, we propose to mine complementary forgery cues from both low-frequency information and compressed high-frequency components. A bidirectional update strategy and an adaptive global-local modulator are proposed to facilitate the utilization of forgery cues. Extensive experiments demonstrate that our method achieves state-of-the-art generalization performance in challenging open-world detection scenarios.

AAAI Conference 2026 Conference Paper

Diffusion Reconstruction-based Data Likelihood Estimation for Core-Set Selection

  • Mingyang Chen
  • Jiawei Du
  • Bo Huang
  • Yi Wang
  • Xiaobo Zhang
  • Wei Wang

Existing core-set selection methods predominantly rely on heuristic scoring signals such as training dynamics or model uncertainty, lacking explicit modeling of data likelihood. This omission may hinder the constructed subset from capturing subtle yet critical distributional structures that underpin effective model training. In this work, we propose a novel, theoretically grounded approach that leverages diffusion models to estimate data likelihood via reconstruction deviation induced by partial reverse denoising. Specifically, we establish a formal connection between reconstruction error and data likelihood, grounded in the Evidence Lower Bound (ELBO) of Markovian diffusion processes, thereby enabling a principled, distribution-aware scoring criterion for data selection. Complementarily, we introduce an efficient information-theoretic method to identify the optimal reconstruction timestep, ensuring that the deviation provides a reliable signal indicative of underlying data likelihood. Extensive experiments on ImageNet demonstrate that reconstruction deviation offers an effective scoring criterion, consistently outperforming existing baselines across selection ratios, and closely matching full-data training using only 50% of the data. Further analysis shows that the likelihood-informed nature of our score reveals informative insights in data selection, shedding light on the interplay between data distributional characteristics and model learning preferences.

AAAI Conference 2026 Conference Paper

Emotion and Intention Guided Multi-Modal Learning for Sticker Response Selection

  • Yuxuan Hu
  • Jian Chen
  • Yuhao Wang
  • Zixuan Li
  • Jing Xiong
  • Pengyue Jia
  • Wei Wang
  • Chengming Li

Stickers are widely used in online communication to convey emotions and implicit intentions. The Sticker Response Selection (SRS) task aims to select the most contextually appropriate sticker based on the dialogue. However, existing methods typically rely on semantic matching and model emotional and intentional cues separately, which can lead to mismatches when emotions and intentions are misaligned. To address this issue, we propose Emotion and Intention Guided Multi-Modal Learning (EIGML). This framework is the first to jointly model emotion and intention, effectively reducing the bias caused by isolated modeling and significantly improving selection accuracy. Specifically, we introduce Dual-Level Contrastive Framework to perform both intra-modality and inter-modality alignment, ensuring consistent representation of emotional and intentional features within and across modalities. In addition, we design an Intention-Emotion Guided Multi-Modal Fusion module that integrates emotional and intentional information progressively through three components: Emotion-Guided Intention Knowledge Selection, Intention-Emotion Guided Attention Fusion, and Similarity-Adjusted Matching Mechanism. This design injects rich, effective information into the model and enables a deeper understanding of the dialogue, ultimately enhancing sticker selection performance. Experimental results on two public datasets show that EIGML outperforms state-of-the-art baselines, achieving higher accuracy and a better understanding of emotional and intentional features.

AAAI Conference 2026 Conference Paper

FedCure: Mitigating Participation Bias in Semi-Asynchronous Federated Learning with Non-IID Data

  • Yue Chen
  • Jianfeng Lu
  • Shuqin Cao
  • Wei Wang
  • Gang Li
  • Guanghui Wen

While semi-asynchronous federated learning (SAFL) combines the efficiency of synchronous training with the flexibility of asynchronous updates, it inherently suffers from participation bias, which is further exacerbated by non-IID data distributions. More importantly, hierarchical architecture shifts participation from individual clients to client groups, thereby further intensifying this issue. Despite notable advancements in SAFL research, most existing works still focus on conventional cloud-end architectures while largely overlooking the critical impact of non-IID data on scheduling across the cloud–edge–client hierarchy. To tackle these challenges, we propose FedCure, an innovative semiasynchronous Federated learning framework that leverages Coalition construction and participation-aware scheduling to mitigate participation bias with non-IID data. Specifically, FedCure operates through three key rules: (1) a preference rule that optimizes coalition formation by maximizing collective benefits and establishing theoretically stable partitions to reduce non-IID-induced performance degradation; (2) a scheduling rule that integrates the virtual queue technique with Bayesian-estimated coalition dynamics, mitigating efficiency loss while ensuring mean rate stability; and (3) a resource allocation rule that enhances computational efficiency by optimizing client CPU frequencies based on estimated coalition dynamics while satisfying delay requirements. Comprehensive experiments on four real-world datasets demonstrate that FedCure improves accuracy by up to 5.1x compared with four state-of-the-art baselines, while significantly enhancing efficiency with the lowest coefficient of variation 0.0223 for per-round latency and maintaining long-term balance across diverse scenarios.

EAAI Journal 2026 Journal Article

Federated learning for big data: A survey on opportunities, applications, and future directions

  • Thippa Reddy Gadekallu
  • Quoc-Viet Pham
  • Thien Huynh-The
  • Hailin Feng
  • Kai Fang
  • Sharnil Pandya
  • Madhusanka Liyanage
  • Wei Wang

In recent years, data generation has grown exponentially, and Big Data has emerged as a propelling force in the development of various machine learning advances and Internet of Things devices. In this regard, the analytical and learning tools that transport data from several sources to a central cloud for processing, training, and storage enable the realization of the potential of Big Data. Nevertheless, since the data may contain sensitive information like banking account information, government information, and personal information, these traditional approaches often raise serious privacy concerns. To overcome such challenges, Federated Learning (FL) has emerged as a sub-field of machine learning that focuses on scenarios where several entities (commonly termed as clients) work together to train a model while maintaining the decentralization of their data. Although significant research efforts have been dedicated to this area, a comprehensive review focusing on FL within the realm of Big Data services is still lacking. This paper, therefore, emphasizes the use of FL in handling Big Data and related services, which provides a comprehensive review of the potential of FL in Big Data acquisition, storage, Big Data analytics, and further privacy preservation. Subsequently, the potential of FL in Big Data applications, such as smart city, smart healthcare, smart transportation, smart grid, and social media are also explored. The paper also highlights various projects related to FL for Big Data and discusses the challenges associated with such implementations. These discussions provide a direction for further research, encouraging the development of plausible solutions.

AAAI Conference 2026 Conference Paper

Frequency-Aligned Cross-Modal Learning with Top-K Wavelet Fusion and Dynamic Expert Routing for Enhanced Retinal Disease Diagnosis

  • Yuxin Lin
  • Haoran Li
  • Haoyu Cao
  • Yongting Hu
  • Qihao Xu
  • Chengliang Liu
  • Xiaoling Luo
  • Zhihao Wu

Multimodal fusion of color fundus photography (CFP) and optical coherence tomography (OCT) B-scan images has demonstrated superior diagnostic potential for retinal diseases compared to single-modality approaches. However, existing fusion paradigms - whether through naive concatenation or attention mechanisms - treat cross-modal interactions indiscriminately, lacking adaptive modulation of modality-specific contributions under varying clinical scenarios. We propose an adaptive fusion framework that dynamically routes and refines multimodal signals for enhancing disease recognition. The framework comprises two key components: 1) Dynamic Cross-Modal Expert Routing (CMER), which selectively activates convolutional neural network (CNN) experts from one modality based on contextual guidance from the other, ensuring only the most relevant feature extractors contribute to fusion; and 2) Top-K Expert-Guided Wavelet Fusion (TEWF), which performs discrete wavelet transform (DWT) to decompose selected features into low- and high-frequency subbands. Cross-modal attention is then applied specifically to high-frequency components, where lesion-specific microstructures reside, enabling frequency-aware fusion. Finally, inverse DWT (IDWT) reconstructs the fused representation, weighted by CMER-derived importance scores to amplify informative modality cues while suppressing redundancy. Experimental validation on two multimodal retinal datasets demonstrates that our method achieves state-of-the-art performance, outperforming existing fusion strategies by significant margins in disease classification accuracy and robustness.

AAAI Conference 2026 Conference Paper

From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling

  • Zhengyu Chen
  • Yudong Wang
  • Teng Xiao
  • Ruochen Zhou
  • Xusheng Yang
  • Wei Wang
  • Zhifang Sui
  • Jingang Wang

Recent advancements in improving the reasoning capabilities of Large Language Models have underscored the efficacy of Process Reward Models (PRMs) in addressing intermediate errors through structured feedback mechanisms. This study analyzes PRMs from multiple perspectives, including training methodologies, scalability, and generalization capabilities. We investigate the interplay between pre-training and reward model training FLOPs to assess their influence on PRM efficiency and accuracy in complex reasoning tasks. Our analysis reveals a pattern of diminishing returns in performance with increasing PRM scale, highlighting the importance of balancing model size and computational cost. Furthermore, the diversity of training datasets significantly impacts PRM performance, emphasizing the importance of diverse data to enhance both accuracy and efficiency. We further examine test-time scaling strategies, identifying Monte Carlo Tree Search as the most effective method when computational resources are abundant, while Best-of-N Sampling serves as a practical alternative under resource-limited conditions. Notably, our findings indicate that PRMs trained on mathematical datasets exhibit performance comparable to those tailored for code generation, suggesting robust cross-domain generalization. Employing a gradient-based metric, we observe that PRMs exhibit a preference for selecting responses with similar underlying patterns, further informing their optimization.

AAAI Conference 2026 Conference Paper

ICLR: Inter-Chrominance and Luminance Interaction for Natural Color Restoration in Low-Light Image Enhancement

  • Xin Xu
  • Hao Liu
  • Wei Liu
  • Wei Wang
  • Jiayi Wu
  • Kui Jiang

Low-Light Image Enhancement (LLIE) task aims at improving contrast while restoring details and textures for images captured in low-light conditions. HVI color space has made significant progress in this task by enabling precise decoupling of chrominance and luminance. However, for the interaction of chrominance and luminance branches, substantial distributional differences between the two branches prevalent in natural images limit complementary feature extraction, and luminance errors are propagated to chrominance channels through the nonlinear parameter. Furthermore, for interaction between different chrominance branches, images with large homogeneous-color regions usually exhibit weak correlation between chrominance branches due to concentrated distributions. Traditional pixel-wise losses exploit strong inter-branch correlations for co-optimization, causing gradient conflicts in weakly correlated regions. Therefore, we propose an Inter-Chrominance and Luminance Interaction (ICLR) framework including a Dual-stream Interaction Enhancement Module (DIEM) and a Covariance Correction Loss (CCL). The DIEM improves the extraction of complementary information from two dimensions, fusion and enhancement, respectively. The CCL utilizes luminance residual statistics to penalize chrominance errors and balances gradient conflicts by constraining chrominance branches covariance. Experimental results on multiple datasets show that the proposed ICLR framework outperforms state-of-the-art methods.

AAAI Conference 2026 Conference Paper

KnowLCP: Knowledge Augmented Lane Change Prediction for Autonomous Driving

  • Yuhuan Lu
  • Pengpeng Xu
  • Wei Wang
  • Zhen Zhang
  • Han Liu
  • Xiping Hu

Lane change prediction, encompassing both intention recognition and trajectory forecasting, is essential for the safe operation of autonomous vehicles in mixed-traffic environments. Existing models predominantly follow a data-driven paradigm, learning directly from historical vehicle states through an end-to-end approach. Inspired by the emerging paradigm of enhancing model generalizability through domain knowledge, we propose KnowLCP to explicitly model and integrate driving knowledge into the lane change prediction task. Specifically, we incorporate three types of knowledge: traffic risk awareness to improve intention prediction, vehicle kinematics to ensure the physical feasibility of predicted trajectories, and intention intensity to refine trajectory forecasting. Furthermore, we introduce a novel knowledge injection strategy that enhances mutual information during integration and proves superior to the traditional parallel input mechanism, which simply feeds knowledge features alongside historical states. Extensive experiments on two real-world trajectory datasets demonstrate that KnowLCP achieves average improvements of 8.3-10.3% in intention prediction and 10.1-10.3% in trajectory prediction over the best-performing baselines.

AAAI Conference 2026 Conference Paper

M2I2: Learning Efficient Multi-Agent Communication via Masked State Modeling and Intention Inference

  • Chuxiong Sun
  • Peng He
  • Qirui Ji
  • Zehua Zang
  • Jiangmeng Li
  • Rui Wang
  • Wei Wang

Communication is essential in coordinating the behaviors of multiple agents. However, existing methods primarily emphasize content, timing, and partners for information sharing, often neglecting the critical aspect of integrating shared information. This gap can significantly impact agents' ability to understand and respond to complex, uncertain interactions, thus affecting overall communication efficiency. To address this issue, we introduce M2I2, a novel framework designed to enhance the agents' capabilities to assimilate and utilize received information effectively. M2I2 equips agents with advanced capabilities for masked state modeling and joint-action prediction, enriching their perception of environmental uncertainties and facilitating the anticipation of teammates' intentions. This approach ensures that agents are furnished with both comprehensive and relevant information, bolstering more informed and synergistic behaviors. Moreover, we propose a Dimensional Rational Network, innovatively trained via a meta-learning paradigm, to identify the importance of dimensional pieces of information, evaluating their contributions to decision-making and auxiliary tasks. Then, we implement an importance-based heuristic for selective information masking and sharing. This strategy optimizes the efficiency of masked state modeling and the rationale behind information sharing. We evaluate M2I2 across diverse multi-agent tasks, the results demonstrate its superior performance, efficiency, and generalization capabilities, over existing state-of-the-art methods in various complex scenarios.

AAAI Conference 2026 Conference Paper

MambaOVSR: Multiscale Fusion with Global Motion Modeling for Chinese Opera Video Super-Resolution

  • Hua Chang
  • Xin Xu
  • Wei Liu
  • Wei Wang
  • Xin Yuan
  • Kui Jiang

Chinese opera is celebrated for preserving classical art. However, early filming equipment limitations have degraded videos of last-century performances by renowned artists (e.g., low frame rates and resolution), hindering archival efforts. Although space-time video super-resolution (STVSR) has advanced significantly, applying it directly to opera videos remains challenging. The scarcity of datasets impedes the recovery of high-frequency details, and existing STVSR methods lack global modeling capabilities—compromising visual quality when handling opera’s characteristic large motions. To address these challenges, we pioneer a large-scale Chinese Opera Video Clip (COVC) dataset and propose the Mamba-based multiscale fusion network for space-time Opera Video Super-Resolution (MambaOVSR). Specifically, MambaOVSR involves three novel components: the Global Fusion Module (GFM) for motion modeling through a multiscale alternating scanning mechanism, and the Multiscale Synergistic Mamba Module (MSMM) for alignment across different sequence lengths. Additionally, our MambaVR block resolves feature artifacts and positional information loss during alignment. Experimental results on the COVC dataset show that MambaOVSR significantly outperforms the SOTA STVSR method by an average of 1.86 dB in terms of PSNR.

AAAI Conference 2026 Conference Paper

Mind the Gap: The Divergence Between Human and LLM-Generated Tasks

  • Yi-Long Lu
  • Jiajun Song
  • Chunhui Zhang
  • Wei Wang

Humans constantly generate a diverse range of tasks guided by internal motivations. While generative agents powered by large language models (LLMs) aim to simulate this complex behavior, it remains uncertain whether they operate on similar cognitive principles. To address this, we conducted a task-generation experiment comparing human responses with those of an LLM agent (GPT-4o). We find that human task generation is consistently influenced by psychological drivers, including personal values (e.g., Openness to Change) and cognitive style. Even when these psychological drivers are explicitly provided to the LLM, it fails to reflect the corresponding behavioral patterns. They produce tasks that are markedly less social, less physical, and thematically biased toward abstraction. Interestingly, while the LLM's tasks were perceived as more fun and novel, this highlights a disconnect between its linguistic proficiency and its capacity to generate human-like, embodied goals. We conclude that there is a core gap between the value-driven, embodied nature of human cognition and the statistical patterns of LLMs, highlighting the necessity of incorporating intrinsic motivation and physical grounding into the design of more human-aligned agents.

JBHI Journal 2026 Journal Article

Morphology Prior Enhanced Teeth Segmentation for High-Resolution Oral Scans

  • Yuxian Jiang
  • Xiuying Wang
  • Tao Yang
  • Changkai Ji
  • Lanshan He
  • Yusheng Liu
  • Wei Wang
  • Min Liu

Deep learning methods have been proposed for tooth segmentation on high-resolution intra-oral scans (IOS) that plays a crucial role in clinical dental practice. However, they generally segment teeth in a low-resolution data with a fixed receptive field and generate final segmentation by up-sampling interpolation, and neglect teeth’s morphology priors: their similar dental arch structures and significantly different curvatures in different parts of each tooth. They thus lack adaptability to different parts of each tooth, and show less accurate segmentation of boundary points between teeth and gums due to the up-sampling computation. Further, cluttered poses of IOS limit their generalization and usability of teeth location and geometric information. To address these limitations, a morphology prior enhanced teeth segmentation framework is proposed in this paper. Firstly, a robust preprocessing is introduced to align poses of different IOS by computing their dental arch orientations, thereby improving segmentation generalization and usability of IOS geometric information. Secondly, a decomposition-merging strategy is designed to avoid the up-sampling limitation, which decomposes an IOS into multiple low-resolution data and merges their segmentation outcomes into a high-resolution result. Thirdly, an innovative module integrating semantic and geometric features is proposed to adaptively select deformable receptive fields. It geometrically samples within a variable probability space to construct receptive fields with varied graph relationships for different points, facilitating adaptive segmentation of different parts of each tooth. Experimental results on 6238 IOS from four centers demonstrate that our method significantly outperforms 11 state-of-the-art methods, achieving a 6. 93% enhancement for cross-center testing.

AAAI Conference 2026 Conference Paper

MPA: Multimodal Prototype Augmentation for Few-Shot Learning

  • Liwen Wu
  • Wei Wang
  • Lei Zhao
  • Zhan Gao
  • Qika Lin
  • Shaowen Yao
  • Zuozhu Liu
  • Bin Pu

Recently, Few-shot Learning (FSL) has become a popular task that aims to recognize new classes from only a few labeled examples and has been widely applied in fields such as natural science, remote sensing, and medical images. However, most existing methods focus only on the visual modality and compute prototypes directly from raw support images, which lack comprehensive and rich multimodal information. To address these limitations, we propose a novel Multimodal Prototype Augmentation FSL framework called MPA, including LLM-based Multi-Variant Semantic Enhancement (LMSE), Hierarchical Multi-View Augmentation (HMA), and an Adaptive Uncertain Class Absorber (AUCA). LMSE leverages large language models to generate diverse paraphrased category descriptions, enriching the support set with additional semantic cues. HMA exploits both natural and multi-view augmentations to enhance feature diversity (e.g., changes in viewing distance, camera angles, and lighting conditions). AUCA models uncertainty by introducing uncertain classes via interpolation and Gaussian sampling, effectively absorbing uncertain samples. Extensive experiments on four single-domain and six cross-domain FSL benchmarks demonstrate that MPA achieves superior performance compared to existing state-of-the-art methods across most settings. Notably, MPA surpasses the second-best method by 12.29% and 24.56% in the single-domain and cross-domain setting, respectively, in the 5-way 1-shot setting.

EAAI Journal 2026 Journal Article

Multi-view feature learning and enhanced hypergraph neural networks for synergistic prediction of drug combination

  • Wei Wang
  • Mengyi Ma
  • Hongjun Zhang
  • Yun Zhou
  • Guangsheng Wu

Drug combination therapy demonstrates more significant efficacy than monotherapy in cancer treatment. Despite the proposal of several computational approaches aimed at effectively identifying synergistic drug combinations, challenges persist due to inadequate multi-level learning within multimodal data. Furthermore, existing models still struggle to adequately capture the complex biological network interactions between drug combinations and cell lines. To overcome these issues, we propose a novel hypergraph neural network method for synergistic drug combination prediction. This method integrates multi-view feature learning and enhanced hypergraph neural networks to improve drug combination prediction. First, multi-view learning is independently applied to the multimodal data of drugs and cell lines. This framework employs a fine-tuned ChemBERTa model enhanced by contrastive learning to effectively capture the contextual information of drug SMILES. Second, enhanced hypergraph neural networks equipped with a multi-head attention mechanism are designed to capture the complex topological information between drugs and cell lines and to address the limited ability of the hypergraph to capture global information. Third, the similarity-based multi-task supervision module further stabilizes the model. The experimental results show that our method outperforms state-of-the-art methods in various scenarios, including leave-drug-combination-out, leave-cell-out, and leave-drug-out scenarios. Specifically, in the leave-drug combination-out scenario, our method achieves a Mean Squared Error of 163. 635, a Root Mean Squared Error of 12. 792, and a Pearson Correlation Coefficient of 0. 751. Finally, a case study demonstrates the efficacy of the model in predicting novel synergistic drug combinations.

EAAI Journal 2026 Journal Article

Multimodal and multiscale learning network for efficient polyp segmentation

  • Yuyang Jie
  • Wei Wang
  • Wentao Shi
  • Zhenkun Lu

Colon cancer has become the second leading cause of cancer-related deaths worldwide, resulting in more than 900 000 deaths each year. Accurate segmentation of intestinal polyps is a key step in preventing the progression of colorectal cancer; however, the robustness of existing methods in complex scenarios still needs improvement. To address this issue, we propose an efficient analysis network that integrates multi-channel and multiscale features, termed the Multimodal and Multiscale Learning Network (MML-Net). Built on the encoder architecture of the Pyramid Vision Transformer (PVT), MML-Net employs a Multi-Channel Feature Fusion Module (MCF) and a Multiscale Parallel Attention Module (MPA) to achieve cross-level feature fusion and deep, fine-grained representation learning, and introduces a Local Attention (LA) mechanism to enhance local context modeling. Experiments show strong learning and generalization capabilities, yielding Dice coefficients of 0. 934 and 0. 925 on the internal datasets Colonoscopy Vision Clinic Database and Kvasir Dataset, and 0. 830 and 0. 813 on the challenging external datasets Colonoscopy Vision Colon Database and ETIS Larib Polyp Database, respectively. The model contains only 34. 4 million (M) parameters and requires 13. 9 giga floating-point operations (GFLOPs), achieving an excellent accuracy–efficiency balance. Implemented artificial-intelligence techniques include PVT, MCF, MPA and LA. This work demonstrates the application of artificial intelligence to medical image analysis for intestinal-polyp segmentation in colonoscopy.

TAAS Journal 2026 Journal Article

Physics-Constrained Adversarial Attack Generation and Active Defense for the Hybrid Deep Reinforcement Learning-Based Load Frequency Control

  • Zhenyong Zhang
  • Wei Wang
  • Mufeng Wang
  • Haiming Wang
  • Jichao Bi
  • Guowen Xu

With the transition to Industry 5.0, there is a growing demand to deploy highly autonomous and resilient artificial intelligence (AI) systems in critical infrastructures such as power grids. In the field of load frequency control (LFC) in power grids, a hybrid control architecture in which deep reinforcement learning (DRL) controllers coexist with traditional proportionalintegral-derivative (PID) controllers can become a typical deployment model during this technological transition. However, the inherent vulnerability of DRL controllers to adversarial attacks introduces new security challenges in such complex environments: attacks not only affect the DRL-controlled areas but may also propagate to PID-controlled areas through inter-area power exchanges, potentially causing broader system instability. To accurately assess the cascading risks under this hybrid architecture, we propose a physics-constrained adversarial attack framework to simulate realistic threats targeting DRL controllers that can propagate across areas. First, we design and implement three typical hybrid control scenarios, i.e., single-agent DRL, partial-agent DRL, and full-agent DRL. Second, we propose a key-feature selection method based on gradient saliency, and we design an attack strategy that adheres to physical constraints while maintaining stealth and efficiency. Third, to enhance the system’s resiliency, we propose a two-stage active defense strategy highly compatible with the hybrid architecture. Finally, we conduct extensive simulation experiments under three typical hybrid control scenarios to evaluate the impact of the adversarial attack and the performance of the defense strategy.

AAAI Conference 2026 Conference Paper

RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis

  • Linfeng Dong
  • Yuchen Yang
  • Hao Wu
  • Wei Wang
  • Yuenan Hou
  • Zhihang Zhong
  • Xiao Sun

We introduce RacketVision, a novel dataset and benchmark for advancing computer vision in sports analytics, covering table tennis, tennis, and badminton. The dataset is the first to provide large-scale, fine-grained annotations for racket pose alongside traditional ball positions, enabling research into complex human-object interactions. It is designed to tackle three interconnected tasks: fine-grained ball tracking, articulated racket pose estimation, and predictive ball trajectory forecasting. Our evaluation of established baselines reveals a critical insight for multi-modal fusion: while naively concatenating racket pose features degrades performance, a Cross-Attention mechanism is essential to unlock their value, leading to trajectory prediction results that surpass strong unimodal baselines. RacketVision provides a versatile resource and a strong starting point for future research in dynamic object tracking, conditional motion forecasting, and multi-modal analysis in sports.

AAAI Conference 2026 Conference Paper

RaLD: Generating High-Resolution 3D Radar Point Clouds with Latent Diffusion

  • Ruijie Zhang
  • Bixin Zeng
  • Shengpeng Wang
  • Fuhui Zhou
  • Wei Wang

Millimeter-wave radar offers a promising sensing modality for autonomous systems thanks to its robustness in adverse conditions and low cost. However, its utility is significantly limited by the sparsity and low resolution of radar point clouds, which poses challenges for tasks requiring dense and accurate 3D perception. Despite that recent efforts have shown great potential by exploring generative approaches to address this issue, they often rely on dense voxel representations that are inefficient and struggle to preserve structural detail. To fill this gap, we make the key observation that latent diffusion models (LDMs), though successful in other modalities, have not been effectively leveraged for radar-based 3D generation due to a lack of compatible representations and conditioning strategies. We introduce RaLD, a framework that bridges this gap by integrating scene-level frustum-based LiDAR autoencoding, order-invariant latent representations, and direct radar spectrum conditioning. These insights lead to a more compact and expressive generation process. Experiments show that RaLD produces dense and accurate 3D point clouds from raw radar spectrums, offering a promising solution for robust perception in challenging environments.

AAAI Conference 2026 Conference Paper

Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment Perspective

  • Deyang Kong
  • Qi Guo
  • Xiangyu Xi
  • Wei Wang
  • Jingang Wang
  • Xunliang Cai
  • Shikun Zhang
  • Wei Ye

The low sampling efficiency during the rollout phase poses a significant challenge to scaling reinforcement learning for large language model reasoning. Existing methods attempt to improve efficiency by scheduling problems based on problem difficulties. However, these approaches suffer from unstable and biased estimations of problem difficulty and fail to capture the alignment between model competence and problem difficulty in RL training, leading to suboptimal results. To address these challenges, we introduce Competence-Difficulty Alignment Sampling (CDAS). This approach allows for accurate and stable estimation of problem difficulties by aggregating historical performance discrepancies across problems. Subsequently, model competence is quantified to adaptively select problems whose difficulties align with the model's current competence using a fixed-point system. Extensive experiments in mathematical RL training show that CDAS consistently outperforms strong baselines, achieving the highest average accuracy of 45.89%. Furthermore, CDAS reduces the training step time overhead by 57.06% compared to the widely-used Dynamic Sampling strategy, verifying the efficiency of CDAS. Additional experiments on different tasks, model architectures, and model sizes demonstrate the generalization capability of CDAS.

AAAI Conference 2026 Conference Paper

SLCFormer: Spectral-Local Context Transformer with Physics-Grounded Flare Synthesis for Nighttime Flare Removal

  • Xiyu Zhu
  • Wei Wang
  • Xin Yuan
  • Xiao Wang

Lens flare is a common nighttime artifact caused by strong light sources scattering within camera lenses, leading to hazy streaks, halos, and glare that degrade visual quality. However, existing methods usually fail to effectively address nonuniform scattered flares, which severely reduces their applicability to complex real-world scenarios with diverse lighting conditions.To address this issue, we propose SLCFormer, a novel spectral-local context transformer framework for effective nighttime lens flare removal. SLCFormer integrates two key modules: the Frequency Fourier and Excitation Module (FFEM), which captures efficient global contextual representations in the frequency domain to model flare characteristics, and the Directionally-Enhanced Spatial Module (DESM) for local structural enhancement and directional features in the spatial domain for precise flare removal. Furthermore, we introduce a ZernikeVAE-based scatter flare generation pipeline to synthesize physically realistic scatter flares with spatially varying PSFs, bridging optical physics and data-driven training. Extensive experiments on the Flare7K++ dataset demonstrate that our method achieves state-of-the-art performance, outperforming existing approaches in both quantitative metrics and perceptual visual quality, and generalizing robustly to real nighttime scenes with complex flare artifacts.

YNIMG Journal 2026 Journal Article

Spinal cord stimulation improves brain connectivity and consciousness level in patients with disorders of consciousness

  • Adilijiang Aihemaitiniyazi
  • Tiemin Li
  • Huawei Zhang
  • Da Wei
  • Pu Cai
  • Wei Wang
  • Guoming Luan
  • Yong Wang

OBJECTIVE: Spinal cord stimulation (SCS) is an advanced neuromodulation technology in disorders of consciousness (DOC) field. However, research on the modulation effects and mechanisms of SCS is limited. METHOD: We proposed a study design (SCS and sham) to study the short-term effects of 20 minutes' SCS, in which resting state EEG and Coma Recovery Scale-Revised (CRS-R) were used to measure the changes in neural and behavioral activity caused by SCS. We used the Genuine Permutation Cross Mutual Information(G_PCMI) to analyze EEG data and study changes in cortical connectivity during SCS. Finally, all patients' CRS-R results were obtained after 6 months' SCS treatment. RESULTS: Short-term SCS (20 min) did not alter the patient's CRS-R score, but long-term SCS (6 months) can improve the CRS-R scores of all patients. EEG results show G_PCMI of the frontal and central brain regions significantly change before and after short-term SCS (p < 0.01) and PCMI of the F-P, F-O regions have significant differences before and after short-term SCS (p < 0.05). Besides, the G_PCMI changes in frontal, parietal, F-P and F-O regions show a significant positive correlation with CRS-R changes (r = 0.80, 0.66, 0.68 and 0.72; p < 0.05). However, the sham group showed no significant G_PCMI changes. CONCLUSION: SCS can improve the awareness level of DOC patients. SCS improves cortical short- and long-distance connectivity of DOC patients may contribute the improvement of consciousness level.

AAAI Conference 2026 Conference Paper

SSTODE: Ocean-Atmosphere Physics-Informed Neural ODEs for Sea Surface Temperature Prediction

  • Zheng Jiang
  • Wei Wang
  • Gaowei Zhang
  • Yi Wang

Sea Surface Temperature (SST) is crucial for understanding upper-ocean thermal dynamics and ocean-atmosphere interactions, which have profound economic and social impacts. While data-driven models show promise in SST prediction, their black-box nature often limits interpretability and overlooks key physical processes. Recently, physics-informed neural networks have been gaining momentum but struggle with complex ocean-atmosphere dynamics due to 1) inadequate characterization of seawater movement (e.g., coastal upwelling) and 2) insufficient integration of external SST drivers (e.g., turbulent heat fluxes). To address these challenges, we propose SSTODE, a physics-informed Neural Ordinary Differential Equations (Neural ODEs) framework for SST prediction. First, we derive ODEs from fluid transport principles, incorporating both advection and diffusion to model ocean spatiotemporal dynamics. Through variational optimization, we recover a latent velocity field that explicitly governs the temporal dynamics of SST. Building upon ODE, we introduce an Energy Exchanges Integrator (EEI)-inspired by ocean heat budget equations-to account for external forcing factors. Thus, the variations in the components of these factors provide deeper insights into SST dynamics. Extensive experiments demonstrate that SSTODE achieves state-of-the-art performances in global and regional SST forecasting benchmarks. Furthermore, SSTODE visually reveals the impact of advection dynamics, thermal diffusion patterns, and diurnal heating-cooling cycles on SST evolution. These findings demonstrate the model's interpretability and physical consistency.

JBHI Journal 2026 Journal Article

TKRL: Targeted Knowledge Rectification Learning Against Teacher-Originated Defects in Domain Continual Segmentation

  • Zhanshi Zhu
  • Wenjian Gu
  • Xiangyu Li
  • Qince Li
  • Yongfeng Yuan
  • Wei Wang
  • Kuanquan Wang
  • Suyu Dong

Knowledge distillation can mitigate catastrophic forgetting in domain continual segmentation by transferring knowledge from the older model to the newer model. However, existing distillation-based methods primarily emphasize knowledge retention while overlooking inherent defects in the older teacher models. As a result, these teacher-originated defects, such as knowledge gaps or biases, are propagated and exacerbate forgetting. To address this challenge, we propose a Targeted Knowledge Rectification Learning framework (TKRL) to probe and correct teacher-originated defects. TKRL consists of two modules: (1) Probe-augmented Class Distillation, which generates gradient-driven “probes” to uncover underrepresented features in the older model, thereby bridging knowledge gaps by distilling hidden information into the new model; (2) Variance-guided Masked Autoencoder, which selectively masks and reconstructs critical high-uncertainty patches across multi-level semantic regions, thereby correcting biases inherited from the older model. Our experimental results show that TKRL effectively rectifies knowledge gaps and biases, thereby mitigating catastrophic forgetting and enhancing performance in domain continual segmentation. The implementation code is publicly available at: https://github.com/PerceptionComputingLab/TKRL_DCMIS.

JBHI Journal 2026 Journal Article

Towards Cognitive Impairment Screening in Elderly Communities with Audio-Visual Modal Disentangled Representation Learning

  • Rui Feng
  • Hongbin Chen
  • Yihao Yao
  • Liuyu Wu
  • Tao Liang
  • Wentao Xiang
  • Jie Li
  • Chu Kiong Loo

Alzheimer's disease (AD) is pressing global health concerns, for which early diagnosis is critical to effective intervention. However, conventional approaches, including neuropsychological assessments and neuroimaging techniques, are resource-intensive and impractical for community-level screening. In contrast, artificial intelligence-driven behavioral analyses, including speech pattern and facial expression recognition, have demonstrated considerable potential for scalable and non invasive cognitive assessment. This work presents a community-oriented intelligent screening system for cognitive impairment screening in elderly populations. As a foundation, we introduce CIR-AV, the first large-scale Mandarin-based multimodal dataset for cognitive impairment recognition in Chinese older adults, encompassing 574 community-dwelling participants with comprehensive facial expression and speech data. Building upon this resource, we propose DiVA, a disentangled audio-visual fusion framework that decomposes multimodal features into shared and specific representations. A trajectory constrained mechanism enhances representation purity, while a cross-modal attention-based dynamic fusion (CMF) module adaptively balances modality contributions, ensuring robust performance under real-world conditions. Experimental results demonstrate that DiVA achieves an AUC of 78. 66% at the segment level and an accuracy of 79. 46% at the subject level, significantly outperforming state-of the-art methods. With its cost-efficient and scalable de sign, it is well-suited for large-scale community screening, providing a practical solution for early dementia detection in resource-limited settings with considerable social and economic value.

AAAI Conference 2026 Conference Paper

Trustworthy Classification for Complex Social Surveys: A Memory-Enhanced Hierarchical Framework with Calibrated Uncertainty

  • Zeqiang Wang
  • Rebecca Oldroyd
  • Yuqi Wang
  • Jiageng Wu
  • Jie Yang
  • Wei Wang
  • Nishanth R. Sastry
  • Jon Johnson

Automated classification of complex social survey questionnaires is crucial for large-scale social science research but faces significant reliability challenges due to intricate hierarchical label structures, severe class imbalance, semantic ambiguity, and incomplete data coverage. Conventional classification methods often struggle with these combined complexities, yielding results that lack trustworthiness. We introduce HOCM, a framework designed for trustworthy classification in complex, real-world taxonomies. It features two synergistic components: (1) memory-enhanced contrastive learning, tailored to learn robust representations from noisy, imbalanced data by leveraging quality-aware category memory banks; and (2) hierarchical uncertainty calibration, which enforces taxonomic consistency while providing reliable confidence estimates and identifying inputs falling outside well-represented known categories. Our evaluation on a large-scale, real-world social survey dataset—a challenging exemplar of our target problem class—demonstrates that HOCM maintains strong accuracy on known classes while effectively identifying uncertain cases, significantly boosting accuracy on confident predictions. Furthermore, it adeptly detects low-resource/unknown categories. HOCM provides a more reliable automated classification tool, enabling efficient expert review and enhancing the trustworthiness of analysis in domains with complex, hierarchical data.

ICRA Conference 2025 Conference Paper

A Data-Efficient Progressive Learning Framework for Robot Scooping Task

  • Shuai Wang 0007
  • Entang Wang
  • Bidan Huang
  • Chong Zhang
  • Wei Wang
  • Yu Zheng 0001

Robot scooping is a challenging and important task in robotic tool manipulation research due to the complex relationship between the robot, the tool, and target objects/environment. Taking into account different tools, different target objects and varying environments, the required scooping manipulation strategy usually varies greatly. Even considering a specific type of spoon, the question of how to obtain a policy model that requires less demonstration data but shows better generalization capabilities deserves further exploration. In this paper, we propose a progressive learning framework for general robot scooping tasks, which requires a limited number of demonstrations but shows promising generalization capability. We first learn a scooping policy via human demonstrations with a specific setup. We then use this as a pre-train model for reinforcement learning in a curriculum manner to achieve a scooping strategy that is generalizable to different task setups. Finally, we evaluate the capabilities of the policy with a series of experiments both in simulation and on a real robot.

JBHI Journal 2025 Journal Article

A Drug-Drug Interaction Prediction Method Based on Atomic 3D Position Encoding and Elastic Message Passing Graph Neural Network

  • Tao Luo
  • Tao Lin
  • Chun Yang
  • Lingjie Fan
  • Wei Wang

Drug-drug interaction (DDI) refers to the inhibitory or enhancing effects between different drugs. Existing DDI prediction methods primarily use graph neural networks (GNNs) to directly represent drug molecular features. However, they often ignore the 3D structures of different atoms within drug molecules and the impact of noise in GNNs on DDI prediction. Consequently, the accuracy of GNN-based DDI prediction remains unsatisfactory. To address these limitations, this study proposes a DDI prediction method based on atomic 3D position encoding and an elastic message passing graph neural network (A3DPE-EMPGNN). Firstly, we construct an atomic feature network based on an attention mechanism and a message passing neural network. This network leverages 3D position encoding based on the molecular centroid to learn the features of different atoms and their associated chemical bonds, thereby constructing a graph-based molecular representation. Secondly, we design a molecular feature network that incorporates an attention mechanism, utilizing multi-head attention to capture interaction information between different drug molecules. Thirdly, we employ an adversarial attack detection and defense strategy, integrating supervised and contrastive loss learning to optimize the model and enhance its robustness while performing DDI prediction. Lastly, we evaluate the effectiveness of A3DPE-EMPGNN on two real-world datasets. Experimental results clearly demonstrate that our method achieves over 98% accuracy across ACC, AUC, AP, and F1-score, outperforming state-of-the-art GNN-based models.

AAAI Conference 2025 Conference Paper

A Lightweight Sparse Interaction Network for Time Series Forecasting

  • Xu Zhang
  • Qitong Wang
  • Peng Wang
  • Wei Wang

Recent work shows that linear models can outperform several transformer models in long-term time-series forecasting (TSF). However, instead of explicitly performing temporal interaction through self-attention, linear models implicitly perform it based on stacked MLP structures, which may be insufficient in capturing the complex temporal dependencies and their performance still has potential for improvement. To this end, we propose a Lightweight Sparse Interaction Network (LSINet) for TSF task. Inspired by the sparsity of self-attention, we propose a Multihead Sparse Interaction Mechanism (MSIM). Different from self-attention, MSIM learns the important connections between time steps through sparsity-induced Bernoulli distribution to capture temporal dependencies for TSF. The sparsity is ensured by the proposed self-adaptive regularization loss. Moreover, we observe the shareability of temporal interactions and propose to perform Shared Interactions Learning (SIL) for MSIM to further enhance efficiency and improve convergence. LSINet is a linear model comprising only MLP structures with low overhead and equipped with explicit temporal interaction mechanisms. Extensive experiments on public datasets show that LSINet achieves both higher accuracy and better efficiency than advanced linear models and transformer models in TSF tasks.

IROS Conference 2025 Conference Paper

A novel event-based structured light system for high-precision and high-speed depth sensing

  • Gongzhe Su
  • Fulong Sun
  • Wei Wang
  • Wei Xi

This paper presents a novel event-based depth sensing system with line laser scan. Our main contribution involves both hardware and software improvements to previous state-of-the-art works. The polygon mirror scanner is designed to steer line laser with a constant velocity, which minimizes non-linearity of the projected time map to improve depth precision. A piecewise linear model is then proposed to model the behavior of the scanner, which is simple and easy to calibrate. The corresponding reconstruction pipeline achieves a high-speed depth map with an efficient plane-ray intersection-based depth calculation. Experimental results verify the approach is capable of realizing 0. 6mm precision at a distance of 500mm and 8. 3ms depth reconstruction runtime on embedded platforms.

JBHI Journal 2025 Journal Article

A Novel Multi-Perspective Framework for Molecule Pretraining: From Atom to Motif Views

  • Wei Wang
  • Dengzhen Lu
  • Suyu Dong
  • Gongning Luo
  • Kuanquan Wang
  • Shanzhuo Zhang

Predicting molecular properties is vital for drug discovery, but experimental measurement is costly and limited by scarce labeled data. Self-supervised molecular pretraining can leverage large unlabeled datasets, reducing dependence on extensive annotations. However, most methods struggle to preserve domain-specific chemical knowledge, especially clinically relevant substructures such as motifs. Random masking and generic graph augmentations often degrade critical chemical information and harm interpretability. Many approaches also work at a single scale-either atom or motif-missing opportunities for cross-scale integration. We propose A2M-Mol, a multi-perspective molecular pretraining framework that combines atom-level and motif-level views through four parallel graph constructions. This design enables cross-view alignment and multiscale fusion, explicitly encoding chemical knowledge. A2M-Mol employs a suite of self-supervised tasks, including cross-view correspondence, atomic reconstruction, global topology modeling, and property constraint enforcement, all coordinated via tailored contrastive learning. Extensive experiments across benchmarks and backbone architectures show consistent improvements over state-of-the-art methods. Ablation studies confirm strong synergies among the tasks. A2M-Mol maintains robust predictive accuracy across data scales, demonstrating effectiveness for real-world molecular property prediction and potential to accelerate drug discovery.

AAAI Conference 2025 Conference Paper

A Trusted Lesion-assessment Network for Interpretable Diagnosis of Coronary Artery Disease in Coronary CT Angiography

  • Xinghua Ma
  • Xinyan Fang
  • Mingye Zou
  • Gongning Luo
  • Wei Wang
  • Kuanquan Wang
  • Zhaowen Qiu
  • Xin Gao

Coronary Artery Disease (CAD) poses a significant threat to cardiovascular patients worldwide, underscoring the critical importance of automated CAD diagnostic technologies in clinical practice. Previous technologies for lesion assessment in Coronary CT Angiography (CCTA) images have been insufficient in terms of interpretability, resulting in solutions that lack clinical reliability in both network architecture and prediction outcomes, even when diagnoses are accurate. To address the limitation of interpretability, we introduce the Trusted Lesion-Assessment Network (TLA-Net), which provides a clinically reliable solution for multi-view CAD diagnosis: (1) The causality-informed evidence collection constructs a causal graph for the diagnostic process and implements causal interventions, preventing confounders' interference and enhancing the transparency of the network architecture. (2) The clinically-aligned uncertainty integration hierarchically combines Dirichlet distributions from various views based on clinical priors, offering confidence coefficients for prediction outcomes that align with physicians' image analysis procedures. Experimental results on a dataset of 2,618 lesions demonstrate that TLA-Net, supported by its interpretable methodological design, exhibits superior performance with outstanding generalization, domain adaptability, and robustness.

EAAI Journal 2025 Journal Article

A unified rotating machinery health management framework leveraging large language models for diverse components, conditions, and tasks

  • Haotian Peng
  • Jie Gao
  • Jiawei Liu
  • Jinsong Du
  • Wei Wang

This study introduces the Rotating Machinery Large Language Model (RotLLM), a unified framework for rotating machinery health management that integrates deep learning with large language models (LLMs) to address diverse operational conditions, components, and health management tasks. RotLLM employs a novel Spectral Folding Network (SFN) to transform vibration spectrum into a unified feature space that preserves essential health state information. A dedicated projection layer then maps these features into the semantic domain of an LLM. The framework is trained using a three-stage strategy: first, pre-training the encoder on the Large-scale Multimodal Rotating Machinery (LMR) dataset, which comprises 237, 298 vibration samples collected under hundreds of operating conditions; second, initializing the projection layer with textual health state labels; and finally, fine-tuning using parameter-efficient Low-Rank Adaptation (LoRA) with high-quality corpus for various health management tasks. Experimental evaluations demonstrate that RotLLM achieves state-of-the-art performance in fault classification, maintains strong robustness under noisy conditions, and delivers rapid multi-task inference with minimal computational overhead. The framework consistently outperforms conventional methods, enabling efficient, accurate, and context-aware health management for rotating machinery across diverse conditions and tasks. The dataset and source code are open-sourced (https: //github. com/SIA-IDE/RotLLM), fostering collaboration, reproducibility, and broader adoption in industrial prognostics research.

ICLR Conference 2025 Conference Paper

AgentRefine: Enhancing Agent Generalization through Refinement Tuning

  • Dayuan Fu
  • Keqing He 0001
  • Yejie Wang
  • Wentao Hong
  • Zhuoma Gongque
  • Weihao Zeng
  • Wei Wang
  • Jingang Wang

Large Language Model (LLM) based agents have proved their ability to perform complex tasks like humans. However, there is still a large gap between open-sourced LLMs and commercial models like the GPT series. In this paper, we focus on improving the agent generalization capabilities of LLMs via instruction tuning. We first observe that the existing agent training corpus exhibits satisfactory results on held-in evaluation sets but fails to generalize to held-out sets. These agent-tuning works face severe formatting errors and are frequently stuck in the same mistake for a long while. We analyze that the poor generalization ability comes from overfitting to several manual agent environments and a lack of adaptation to new situations. They struggle with the wrong action steps and can not learn from the experience but just memorize existing observation-action relations. Inspired by the insight, we propose a novel AgentRefine framework for agent-tuning. The core idea is to enable the model to learn to correct its mistakes via observation in the trajectory. Specifically, we propose an agent synthesis framework to encompass a diverse array of environments and tasks and prompt a strong LLM to refine its error action according to the environment feedback. AgentRefine significantly outperforms state-of-the-art agent-tuning work in terms of generalization ability on diverse agent tasks. It also has better robustness facing perturbation and can generate diversified thought in inference. Our findings establish the correlation between agent generalization and self-refinement and provide a new paradigm for future research.

AIIM Journal 2025 Journal Article

AI-based methods for diagnosing and grading diabetic retinopathy: A comprehensive review

  • Ibrahim Saleh
  • Niveen Nasr El-Den
  • Mohamed Elsharkawy
  • Ali Mahmoud
  • Ashraf Sewelam
  • Wei Wang
  • Mohammed Ghazal
  • Ayman El-Baz

Diabetic retinopathy (DR) is a leading cause of blindness worldwide, requiring early detection and accurate grading for effective intervention. Advances in artificial intelligence (AI), computer vision, machine learning, and deep learning (DL) have enabled automated detection and classification of DR through various imaging modalities. This review comprehensively evaluates 91 studies employing AI-based methods in the detection and classification of DR using fundus color photography, optical coherence tomography (OCT), OCT-angiography (OCTA), and fundus fluorescein angiography, providing a holistic understanding of their strengths, challenges, and limitations. Additionally, this review compares the characteristics of 23 public datasets for DR. Across modalities, DL approaches generally outperform traditional methods. Among the studies reviewed, 81% utilized fundus images, followed by 9% using OCT, 6% using OCTA, and 2% incorporating multiple modalities. Regarding classification tasks, 62% used AI for multi-way classification, 28% for binary classification, and 10% incorporated both. The paper concludes with future directions, including explainable AI frameworks, multimodal data integration, and suggested protocols to integrate into existing healthcare workflows.

IROS Conference 2025 Conference Paper

An Inflatable Deployable Origami Grasper for Adaptive and High-Load Grasping

  • Peng Yan
  • Guang Liang
  • Sen Wang
  • Hailin Huang
  • Wei Wang
  • Xu Li
  • Bing Li

Robotic graspers are essential for enhancing the efficiency and versatility of robots in grasping tasks. In this paper, we propose a novel inflatable deployable origami grasper with a rigid-flexible coupling structure. The proposed grasper can achieve multiple deployment configurations under a single pneumatic actuation, enabling both deployment and grasping operations while also allowing for passive self-folding during deflation. The design and fabrication of the grasper are presented. Then, the stiffness model for the inflatable deployable origami unit is developed based on the equivalent truss method. Experimental results show that the grasper successfully grasps objects of various shapes and sizes in both enveloping and fingertip grasping modes, using either two or four fingers. With its simple mechanical system and high deploy/fold ratio, the proposed grasper holds significant potential for applications in industrial automation and space exploration.

AAAI Conference 2025 Conference Paper

Arbitrary Reading Order Scene Text Spotter with Local Semantics Guidance

  • Jiahao Lyu
  • Wei Wang
  • Dongbao Yang
  • Jinwen Zhong
  • Yu Zhou

Scene text spotting has attracted the enthusiasm of relative researchers in recent years. Most existing scene text spotters follow the detection-then-recognition paradigm, where the vanilla detection module hardly determines the reading order and leads to failure recognition. After rethinking the auto-regressive scene text recognition method, we find that a well-trained recognizer can implicitly perceive the local semantics of all characters in a complete word or a sentence without a character-level detection module. Local semantic knowledge not only includes text content but also spatial information in the right reading order. Motivated by the above analysis, we propose the Local Semantics Guided scene text Spotter (LSGSpotter), which auto-regressively decodes the position and content of characters guided by the local semantics. Specifically, two effective modules are proposed in LSGSpotter. On the one hand, we design a Start Point Localization Module (SPLM) for locating text start points to determine the right reading order. On the other hand, a Multi-scale Adaptive Attention Module (MAAM) is proposed to adaptively aggregate text features in a local area. In conclusion, LSGSpotter achieves the arbitrary reading order spotting task without the limitation of sophisticated detection, while alleviating the cost of computational resources with the grid sampling strategy. Extensive experiment results show LSGSpotter achieves state-of-the-art performance on the InverseText benchmark. Moreover, our spotter demonstrates superior performance on English benchmarks for arbitrary-shaped text, achieving improvements of 0.7% and 2.5% on Total-Text and SCUT-CTW1500, respectively. These results validate our text spotter is effective for scene texts in arbitrary reading order and shape.

EAAI Journal 2025 Journal Article

Bayesian bidirectional long short-term memory-based kinematics-dynamics fusion for fault-tolerant vehicle state estimation under yaw rate sensor failures

  • Min Gao
  • Jiaqi Li
  • Wei Wang
  • Renguang Wang
  • Jin Luo
  • Jing Li

Accurate estimation of vehicle states is a fundamental component of vehicle stability control systems. To address the issue of inaccurate estimation of vehicle state parameters resulting from yaw rate sensor failures, this study proposes a three-mode collaborative fault-tolerant state estimation method based on Bayesian Bidirectional Long Short-Term Memory (BiLSTM) kinematics-dynamics fusion. First, the kinematics-based method is established using the kinematics model. Second, the dynamics-based method is designed by integrating the Unscented Kalman Filter (UKF) with the dynamics model. Subsequently, a BiLSTM network fusion model based on Bayesian optimization is presented. The model utilizes estimates from kinematic and kinetic methods as a priori inputs and combines the bidirectional information capturing capability of BiLSTM with hyperparameter tuning from Bayesian optimization. The results indicate that when the yaw rate sensor fails, the proposed method achieves an average Root Mean Square Error (RMSE) of 0. 0276 km per hour (km/h) for longitudinal speed, 0. 0008 radian (rad) for side slip angle, and 0. 0072 radian per second (rad/s) for yaw rate across all scenarios. This performance demonstrates a superiority over various maneuvers. This paper combines kinematics, dynamics, and deep learning to provide a reliable solution for fault-tolerant estimation of vehicle states.

AAAI Conference 2025 Conference Paper

BearLLM: A Prior Knowledge-Enhanced Bearing Health Management Framework with Unified Vibration Signal Representation

  • Haotian Peng
  • Jiawei Liu
  • Jinsong Du
  • Jie Gao
  • Wei Wang

We propose a bearing health management framework leveraging large language models (BearLLM), a novel multimodal model that unifies multiple bearing-related tasks by processing user prompts and vibration signals. Specifically, we introduce a prior knowledge-enhanced unified vibration signal representation to handle various working conditions across multiple datasets. This involves adaptively sampling the vibration signals based on the sampling rate of the sensor, incorporating the frequency domain to unify input dimensions, and using a fault-free reference signal as an auxiliary input. To extract features from vibration signals, we first train a fault classification network, then convert and align the extracted features into word embedding, and finally concatenate these with text embedding as input to an LLM. To evaluate the performance of the proposed method, we constructed the first large-scale multimodal bearing health management (MBHM) dataset, including paired vibration signals and textual descriptions. With our unified vibration signal representation, BearLLM using one set of pre-trained weights achieves state-of-the-art performance on nine publicly available fault diagnosis benchmarks, outperforming specific methods designed for individual datasets. We provide a dataset, our model, and code to inspire future research on building more capable industrial multimodal models.

EAAI Journal 2025 Journal Article

Behavioral decision-making of mobile robots simulating the functions of cerebellum, basal ganglia and hippocampus

  • Qi Liu
  • Wei Wang
  • Dongshu Wang

In facing unknown and complex environments, the ability of robots to make efficient and good behavioral decisions is an important prerequisite for them to accomplish various tasks. When robots perform tasks in changing environments or in a new environment, they generally have poor decision-making skills and difficulty acquiring stable and continuous learning capabilities. Therefore, investigating how to make robots have human-like decision-making and learning ability and be able to make better decisions facing unknown complex tasks is currently an important challenge for robots. The authors have attempted to achieve flexible behavioral decision-making in dynamic environments of mobile robots by integrating the cerebellar supervised learning function and the reinforcement learning function of the basal ganglia, but the proposed algorithms have limited ability to adapt to dynamic environments. Based on previous work, this article redesigns the cerebellum module and adopts the motivated developmental network 2 to simulate the working mechanism of the cerebellum. Motivated developmental network 2 realizes the precise characterization of intrinsic features of the complex environment through the lateral connection between hidden neurons, in order to enhance the robot decision-making ability. A memory consolidation module is designed according to the working principle of the hippocampus, which makes the robot not only able to make good behavioral decision making when facing an unknown dynamic environment, but also able to continuously acquire knowledge from the unknown environment. Meanwhile, a state–action prediction method is proposed to realize flexible behavioral decision-making based on cerebellar supervised learning and basal ganglia’s reinforcement learning in dynamic environment, which further improves its decision-making ability. Experimental results in dynamic and real environments verify the potential of the model proposed in this paper.

ICLR Conference 2025 Conference Paper

Bridging and Modeling Correlations in Pairwise Data for Direct Preference Optimization

  • Yuxin Jiang
  • Bo Huang
  • Yufei Wang 0005
  • Xingshan Zeng
  • Liangyou Li
  • Yasheng Wang
  • Xin Jiang 0002
  • Lifeng Shang

Direct preference optimization (DPO), a widely adopted offline preference optimization algorithm, aims to align large language models (LLMs) with human-desired behaviors using pairwise preference data. However, the generation of the winning response and the losing response within pairwise data are typically isolated, leading to weak correlations between them as well as suboptimal alignment performance. To address this issue, we propose an effective framework for Bridging and Modeling Correlations in pairwise data, named BMC. Firstly, we increase the consistency and informativeness of the pairwise preference signals through targeted modifications, synthesizing a pseudo-winning response by improving the losing response with the winning response as a reference. Secondly, we identify that DPO alone is insufficient to model these correlations and capture nuanced variations. Therefore, we propose learning token-level correlations by dynamically leveraging the policy model's confidence during training. Comprehensive experiments on QA, math, and instruction-following tasks demonstrate the effectiveness of our approach, significantly surpassing competitive baselines, including DPO. Additionally, our in-depth quantitative analysis reveals the reasons behind our method's superior performance over DPO and showcases its versatility to other DPO variants.

NeurIPS Conference 2025 Conference Paper

Can Large Language Models Master Complex Card Games?

  • Wei Wang
  • Fuqing Bie
  • Junzhe Chen
  • Dan Zhang
  • Shiyu Huang
  • Evgeny Kharlamov
  • Jie Tang

Complex games have long been an important benchmark for testing the progress of artificial intelligence algorithms. AlphaGo, AlphaZero, and MuZero have defeated top human players in Go and Chess, garnering widespread societal attention towards artificial intelligence. Concurrently, large language models (LLMs) have exhibited remarkable capabilities across various tasks, raising the question of whether LLMs can achieve similar success in complex games. In this paper, we explore the potential of LLMs in mastering complex card games. We systematically assess the learning capabilities of LLMs across eight diverse card games, evaluating the impact of fine-tuning on high-quality gameplay data, and examining the models' ability to retain general capabilities while mastering these games. Our findings indicate that: (1) LLMs can approach the performance of strong game AIs through supervised fine-tuning on high-quality data, (2) LLMs can achieve a certain level of proficiency in multiple complex card games simultaneously, with performance augmentation for games with similar rules and conflicts for dissimilar ones, and (3) LLMs experience a decline in general capabilities when mastering complex games, but this decline can be mitigated by integrating a certain amount of general instruction data. The evaluation results demonstrate strong learning ability and versatility of LLMs. The code is available at https: //github. com/THUDM/LLM4CardGame

NeurIPS Conference 2025 Conference Paper

CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding

  • Hongyong Han
  • Wei Wang
  • Gaowei Zhang
  • Mingjie Li
  • Yi Wang

Coral reefs are vital yet vulnerable ecosystems that require continuous monitoring to support conservation. While coral reef images provide essential information in coral monitoring, interpreting such images remains challenging due to the need for domain expertise. Visual Question Answering (VQA), powered by Large Vision-Language Models (LVLMs), has great potential in user-friendly interaction with coral reef images. However, applying VQA to coral imagery demands a dedicated dataset that addresses two key challenges: domain-specific annotations and multidimensional questions. In this work, we introduce CoralVQA, the first large-scale VQA dataset for coral reef analysis. It contains 12, 805 real-world coral images from 67 coral genera collected from 3 oceans, along with 277, 653 question-answer pairs that comprehensively assess ecological and health-related conditions. To construct this dataset, we develop a semi-automatic data construction pipeline in collaboration with marine biologists to ensure both scalability and professional-grade data quality. CoralVQA presents novel challenges and provides a comprehensive benchmark for studying vision-language reasoning in the context of coral reef images. By evaluating several state-of-the-art LVLMs, we reveal key limitations and opportunities. These insights form a foundation for future LVLM development, with a particular emphasis on supporting coral conservation efforts.

AAAI Conference 2025 Conference Paper

Critical Forgetting-Based Multi-Scale Disentanglement for Deepfake Detection

  • Kai Li
  • Wenqi Ren
  • Jianshu Li
  • Wei Wang
  • Xiaochun Cao

Recent face forgery detection methods based on disentangled representation learning utilize paired images for cross-reconstruction, aiming to extract forgery-relevant attributes and forgery-irrelevant content. However, there still exist the following issues that may comprise the detector performance: 1) using information-dense images as the decoupling targets increases the decoupling difficulty; 2) the extracted attribute features are reconstruction-irrelevant rather than forgery-relevant, and single-scale forgery representation decoupling cannot capture sufficient discriminative information; 3) the generalization performance of decoupled attribute features is poor as the detector focuses on learning specific artifact types in the training set. To address these issues, we propose a novel disentangled representation learning framework for deepfake detection. First, we extract features by partitioning the dense information within the image, focusing independently on texture, color, or edges. These features are then used as the decoupling targets rather than the images themselves, which could mitigate the decoupling difficulty. Second, we extend reconstruction loss from image-level to feature-level, thus extending the forgery representation decoupling from single-scale to multi-scale. Third, we propose a critical forgetting mechanism that forces the detector to forget the most salient features during training, which correspond to specific forgery artifact types in the training set. Extensive experimental results validate the efficacy of the proposed method.

NeurIPS Conference 2025 Conference Paper

DCI: Dual-Conditional Inversion for Boosting Diffusion-Based Image Editing

  • Zixiang Li
  • Haoyu Wang
  • Wei Wang
  • Chuangchuang Tan
  • Yunchao Wei
  • Yao Zhao

Diffusion models have achieved remarkable success in image generation and editing tasks. Inversion within these models aims to recover the latent noise representation for a real or generated image, enabling reconstruction, editing, and other downstream tasks. However, to date, most inversion approaches suffer from an intrinsic trade-off between reconstruction accuracy and editing flexibility. This limitation arises from the difficulty of maintaining both semantic alignment and structural consistency during the inversion process. In this work, we introduce Dual-Conditional Inversion (DCI), a novel framework that jointly conditions on the source prompt and reference image to guide the inversion process. Specifically, DCI formulates the inversion process as a dual-condition fixed-point optimization problem, minimizing both the latent noise gap and the reconstruction error under the joint guidance. This design anchors the inversion trajectory in both semantic and visual space, leading to more accurate and editable latent representations. Our novel setup brings new understanding to the inversion process. Extensive experiments demonstrate that DCI achieves state-of-the-art performance across multiple editing tasks, significantly improving both reconstruction quality and editing precision. Furthermore, we also demonstrate that our method achieves strong results in reconstruction tasks, implying a degree of robustness and generalizability approaching the ultimate goal of the inversion process. Our codes are available at: https: //github. com/Lzxhh/Dual-Conditional-Inversion

AAAI Conference 2025 Conference Paper

Deep Hierarchies and Invariant Disease-Indicative Feature Learning for Computer Aided Diagnosis of Multiple Fundus Diseases

  • Yuxin Lin
  • Wei Wang
  • Xiaoling Luo
  • Zhihao Wu
  • Chengliang Liu
  • Jie Wen
  • Yong Xu

With the advancement of computer vision, numerous models have been proposed for screening of fundus diseases. However, the recognition of multiple fundus diseases is often hampered by the simultaneous presence of multiple disease types and the confluence of lesion types in fundus images. This paper addresses these challenges by conceptualizing them as multi-level feature fusion and self-supervised disease-indicative feature learning problems. We decode fundus images at various levels of granularity to delineate scenarios wherein multiple diseases and lesions co-occur. To effectively integrate these features, we introduce a hierarchical vision transformer (HVT) that adeptly captures both inter-level and intra-level dependencies. A novel forward-attention module is proposed to enhance the integration of lower-level semantic information into higher semantic layers, thereby enriching the representation of complex features. Additionally, we introduce a novel self-supervised mask-consistent feature learner (MCFL). Unlike traditional mask-autoencoders that reconstruct original images using encoder-decoder structures, MCFL utilizes a teacher-student framework to reconstruct mask-consistent feature maps. In this setup, exponential moving averaging is employed to derive classification-guided features, serving as labels for reconstruction rather than merely reconstructing the original images. This innovative approach facilitates the extraction of disease-indicative features. Extensive experiments demonstrate that our method significantly outperforms existing state-of-the-art models.

IJCAI Conference 2025 Conference Paper

Deep Opinion-Unaware Blind Image Quality Assessment by Learning and Adapting from Multiple Annotators

  • Zhihua Wang
  • Xuelin Liu
  • Jiebin Yan
  • Jie Wen
  • Wei Wang
  • Chao Huang

Existing deep neural network (DNN)-based blind image quality assessment (BIQA) methods primarily rely on human-rated datasets for training. However, collecting human labels is extremely time-consuming and labor-intensive, posing a significant bottleneck for practical applications. To address this challenge, we propose a Deep opinion-Unaware BIQA model by learning and adapting from Multiple Annotators, termed DUBMA, thereby eliminating the need for human annotations. Specifically, we first generate a large-scale set of distorted image pairs and then assign relative quality rankings using existing full-reference IQA models. The resulting dataset is subsequently employed for training our DUBMA. Due to the inherent discrepancies between synthetic and real-world distortions, a domain shift may occur. To address this, we propose an outlier-robust unsupervised domain adaptation approach leveraging optimal transport. This strategy effectively reduces the gap between synthetic and real-world distortion domains, thereby boosting the model’s adaptability and overall performance. Extensive experiments show that DUBMA outperforms existing opinion-unaware BIQA methods in terms of prediction accuracy across multiple datasets.

IJCAI Conference 2025 Conference Paper

DERI: Cross-Modal ECG Representation Learning with Deep ECG-Report Interaction

  • Jian Chen
  • Xiaoru Dong
  • Wei Wang
  • Shaorui Zhou
  • Lequan Yu
  • Xiping Hu

Electrocardiogram (ECG) is widely used to diagnose cardiac conditions via deep learning methods. Although existing self-supervised learning (SSL) methods have achieved great performance in learning representation for ECG-based cardiac conditions classification, the clinical semantics can not be effectively captured. To overcome this limitation, we proposed to learn cross-modal ECG representations that contain more clinical semantics via a novel framework with \textbf{D}eep \textbf{E}CG-\textbf{R}eport \textbf{I}nteraction (\textbf{DERI}). Specifically, we design a novel framework combining multiple alignments and mutual feature reconstructions to learn effective representation of the ECG with the clinical report, which fuses the clinical semantics of the report. An RME-Module inspired by masked modeling is proposed to improve the ECG representation learning. Furthermore, we extend ECG representation learning to report generation with a language model, which is significant for evaluating clinical semantics in the learned representations and even clinical applications. Comprehensive experiments with various settings are conducted on various datasets to show the superior performance of our DERI. Our code is released on https: //github. com/cccccj-03/DERI.

AAAI Conference 2025 Conference Paper

Disentangle Nighttime Lens Flares: Self-supervised Generation-based Lens Flare Removal

  • Yuwen He
  • Wei Wang
  • Wanyu Wu
  • Kui Jiang

Lens flares arise from light reflection and refraction within sensor arrays, whose diverse types include glow, veiling glare, reflective flare and so on. Existing methods are specialized for one specific type only, and overlook the simultaneous occurrence of multiple typed lens flares, which is common in the real-world, e.g. coexistence of glow and displacement reflections from the same light source. These co-occurring lens flares cannot be effectively resolved by the simple combination of individual flare removal methods, since these coexisting flares originates from the same light source and are generated simultaneously within the same sensor array, exhibit a complex interdependence rather than simple additive relation. To model this interdependent flares’ relationship, our Nighttime Lens Flare Formation model is the first attempt to learn the intrinsic physical relationship between flares on the imaging plane. Building on this physical model, we introduce a solution to this joint flare removal task named Self-supervised Generation-based Lens Flare Removal Network (SGLFR-Net), which is self-supervised without pre-training. Specifically, the nighttime glow is detangled in PSF Rendering Network(PSFR-Net) based on PSF Rendering Prior, while the reflective flare is modelled in Texture Prior Based Reflection Flare Removal Network (TPRR-Net). Empirical evaluations demonstrate the effectiveness of the proposed method in both joint and individual glare removal tasks.

NeurIPS Conference 2025 Conference Paper

Don’t Forget the Enjoin: FocalLoRA for Instruction Hierarchical Alignment in Large Language Models

  • Zitong Shi
  • Frank Wan
  • Haixin Wang
  • Ruoyan Li
  • Zijie Huang
  • Wanjia Zhao
  • Yijia Xiao
  • Xiao Luo

Recent studies reveal that large language models (LLMs) often struggle to resolve conflicting instructions embedded within hierarchical prompts, resulting in decreased compliance with system-level directives and compromising the reliability of safety-critical applications. While earlier approaches attempt to improve instruction hierarchy awareness through prompt engineering or embedding-level modifications, they typically lack structural modeling and either offer limited gains or require extensive fine-tuning. In this work, we introduce $\textbf{FocalLoRA}$, a parameter-efficient and structure-aware framework that strengthens hierarchical instruction adherence by selectively optimizing structurally critical attention heads, referred to as $\textit{focal heads}$, which exhibit heightened sensitivity to instruction conflicts. Experiments across multiple models and a dedicated benchmark demonstrate that FocalLoRA markedly enhances system instruction compliance with minimal tuning cost. For instance, on Llama-8B, fine-tuning only 0. 0188\% of parameters yields a 35. 52\% $\uparrow$ in system instruction compliance.

AAAI Conference 2025 Conference Paper

Dual-View Interaction-Aware Lane Change Prediction for Autonomous Driving

  • Yuhuan Lu
  • Zhen Zhang
  • Rufan Bai
  • Han Liu
  • Wei Wang

As artificial intelligence techniques evolve, we are approaching a critical moment for the widespread deployment of autonomous vehicles. Subsequently, the emergence of mixed-autonomy traffic environments presents formidable challenges to autonomous vehicles, especially for the accurate prediction of lane change intentions of their surrounding human-driven vehicles, which is crucial for ensuring the safety of autonomous vehicles. Existing lane change prediction models mainly focus on capturing the temporal variations in the movement dynamics of individual vehicles. However, the neglect to consider inter-vehicle interactions hinders their capability in complex lane change scenarios, resulting in suboptimal prediction performance. Moreover, current interaction-aware approaches for autonomous driving fail to explicitly model future interactions between vehicles, leading to unreasonable prediction results that can cause collisions between vehicles. To address the above issues, we propose to incorporate the concept of perceived safety into future interaction modeling and design a dual-view interaction-aware lane change prediction model. We evaluate the proposed model on two real-world datasets and experimental results show that the proposed model achieves average improvements of 11.7-12.4% in classification ability and 75.6-95.7% in forecast ability over the best-performing baselines across the two datasets. The ablation study and investigation into future interaction modeling demonstrate that our model has advantages in interpreting lane change scenarios from a driving safety perspective.

IROS Conference 2025 Conference Paper

Dynamic Modeling and Efficient Data-Driven Optimal Control for Micro Autonomous Surface Vehicles

  • Zhiheng Chen
  • Wei Wang

Micro Autonomous Surface Vehicles (MicroASVs) offer significant potential for operations in confined or shallow waters and swarm robotics applications. However, achieving precise and robust control at such small scales remains highly challenging, mainly due to the complexity of modeling nonlinear hydrodynamic forces and the increased sensitivity to self-motion effects and environmental disturbances, including waves and boundary effects in confined spaces. This paper presents a physics-driven dynamics model for an over-actuated MicroASV and introduces a data-driven optimal control framework that leverages a weak formulation-based online model learning method. Our approach continuously refines the physics-driven model in real time, enabling adaptive control that adjusts to changing system parameters. Simulation results demonstrate that the proposed method substantially enhances trajectory tracking accuracy and robustness, even under unknown payloads and external disturbances. These findings highlight the potential of data-driven online learning-based optimal control to improve MicroASV performance, paving the way for more reliable and precise autonomous surface vehicle operations.

IJCAI Conference 2025 Conference Paper

ECG2TOK: ECG Pre-Training with Self-Distillation Semantic Tokenizers

  • Xiaoyan Yuan
  • Wei Wang
  • Han Liu
  • Jian Chen
  • Xiping Hu

Self-supervised learning (SSL) has garnered increasing attention in electrocardiogram (ECG) analysis for its effectiveness in resource-limited settings. Existing state-of-the-art SSL methods rely on time-frequency detail reconstruction, but due to the inherent redundancy of ECG signals and individual variability, these approaches often yield suboptimal performance. In contrast, discrete label prediction becomes a superior pre-training objective by encouraging models to efficiently abstract ECG high-level semantics. However, the continuity and significant variability of ECG signals pose a challenge in generating semantically discrete labels. To address this issue, we propose an ECG pretraining framework with a self-distillation semantic tokenizer (ECG2TOK), which maps continuous ECG signals into discrete labels for self-supervised training. Specifically, the tokenizer extracts semantically aware embeddings of ECG by self-distillation and performs online clustering to generate semantically rich discrete labels. Subsequently, the SSL model is trained in conjunction with masking strategies and discrete label prediction to facilitate the abstraction of high-level semantic representations. We evaluate ECG2TOK in six downstream tasks, demonstrating that ECG2TOK efficiently achieves state-of-the-art performance and up to a 30. 73% AUC increase in low-resource scenarios. Moreover, visualization experiments demonstrate that the discrete labels generated by ECG2TOK exhibit consistent semantics closely associated with clinical features. Our code is available on https: //github. com/YXYanova/ECG2TOK.

IJCAI Conference 2025 Conference Paper

Enhancing the Performance of Global Model by Improving the Adaptability of Local Models in Federated Learning

  • Wujun Zhou
  • Shu Ding
  • Zelin Li
  • Wei Wang

Federated learning enables the clients to collaboratively train a global model, which is aggregated from local models. Due to the heterogeneous data distributions over clients and data privacy in federated learning, it is difficult to train local models to achieve a well-performed global model. In this paper, we introduce the adaptability of local models, i. e. , the average performance of local models on data distributions over clients, and enhance the performance of the global model by improving the adaptability of local models. Since each client does not know the data distributions over other clients, the adaptability of the local model cannot be directly optimized. First, we provide the property of an appropriate local model which has good adaptability on the data distributions over clients. Then, we formalize the property into the local training objective with a constraint and propose a feasible solution to train the local model. Extensive experiments on federated learning benchmarks demonstrate that our method significantly improves the adaptability of local models and achieves a well-performed global model that consistently outperforms the baseline methods.

AAAI Conference 2025 Conference Paper

Enhancing Vision-Language Models with Morphological and Taxonomic Knowledge: Towards Coral Recognition for Ocean Health

  • Hongyong Han
  • Wei Wang
  • Gaowei Zhang
  • Mingjie Li
  • Yi Wang

Coral reefs play a crucial role in marine ecosystems, offering a nutrient-rich environment and safe shelter for numerous marine species. Automated coral image recognition aids in monitoring ocean health at a scale without experts' manual effort. Recently, large vision-language models like CLIP have greatly enhanced zero-shot and low-shot classification capabilities for various visual tasks. However, these models struggle with fine-grained coral-related tasks due to a lack of specific knowledge. To bridge this gap, we compile a fine-grained coral image dataset consisting of 16,659 images with taxonomy labels (from Kingdom to Species), accompanied by morphology-specific text descriptions for each species. Based on the dataset, we propose CORAL-Adapter, integrating two complementary kinds of coral-specific knowledge (biological taxonomy and coral morphology) with general knowledge learned by CLIP. CORAL-Adapter is a simple yet powerful extension of CLIP with only a few parameter updates and can be used as a plug-and-play module with various CLIP-based methods. We show improvements in accuracy across diverse coral recognition tasks, e.g., recognizing corals unseen during training that are prone to bleaching or originate from different oceans.

JBHI Journal 2025 Journal Article

Enhancing Weakly Supervised Semantic Segmentation With Multi-Label Contrastive Learning and LLM Features Guidance

  • Wentian Cai
  • Yijiang Li
  • Yandan Chen
  • Jing Lin
  • Zihao Huang
  • Ping Gao
  • Thippa Reddy Gadekallu
  • Wei Wang

Histopathological whole-slide images (WSIs) segmentation is essential for precise tissue characterization in medical diagnostics. However, traditional approaches require labor-intensive pixel-level annotations. To this end, we study weakly supervised semantic segmentation (WSSS) which uses patch-level classification labels, reducing annotation efforts significantly. However, the complexity of WSIs and the challenge of sparse classification labels hinder effective dense pixel predictions. Moreover, due to the multi-label nature of WSI, existing approaches of single-label contrastive learning designed for the representation of single-category, neglecting the presence of other relevant categories and thus fail to adapt to WSI tasks. This paper presents a novel multi-label contrastive learning method for WSSS by incorporating class-specific embedding extraction with LLM features guidance. Specifically, we propose to obtain class-specific embeddings by utilizing classifier weights, followed by a dot-product-based attention fusion method that leverages LLM features to enrich their semantics, facilitating contrastive learning between different classes from single image. Besides, we propose a Robust Learning approach that leverages multi-layer features to evaluate the uncertainty of pseudo-labels, thereby mitigating the impact of noisy pseudo-labels on the learning process of segmentation. Extensive experiments have been conducted on two histopathological image segmentation datasets, i. e. LUAD dataset and BCSS dataset, demonstrating the effectiveness of our methods with leading performance.

AAAI Conference 2025 Conference Paper

Federated Weakly Supervised Video Anomaly Detection with Multimodal Prompt

  • Benfeng Wang
  • Chao Huang
  • Jie Wen
  • Wei Wang
  • Yabo Liu
  • Yong Xu

Video anomaly detection (VAD) aims at locating the abnormal events in videos. Recently, the Weakly Supervised VAD has made great progress, which only requires video-level annotations when training. In practical applications, different institutions may have different types of abnormal videos. However, the abnormal videos cannot be circulated on the internet due to privacy protection. To train a more generalized anomaly detector that can identify various anomalies, it is reasonable to introduce federated learning into WSVAD. In this paper, we propose Global and Local Context-driven Federated Learning, a new paradigm for privacy protected weakly supervised video anomaly detection. Specifically, we utilize the vision-language association of CLIP to detect whether the video frame is abnormal. Instead of leveraging handcrafted text prompts for CLIP, we propose a text prompt generator. The generated prompt is simultaneously influenced by text and visual. On the one hand, the text provides global context related to anomaly, which improves the model's ability of generalization. On the other hand, the visual provides personalized local context because different clients may have videos with different types of anomalies or scenes. The generated prompt ensures global generalization while processing personalized data from different clients. Extensive experiments show that the proposed method achieves remarkable performance.

NeurIPS Conference 2025 Conference Paper

Flow Field Reconstruction with Sensor Placement Policy Learning

  • Ruoyan Li
  • Frank Wan
  • Zijie Huang
  • Zixiao Liu
  • Haixin Wang
  • Xiao Luo
  • Wei Wang
  • Yizhou Sun

Flow‐field reconstruction from sparse sensor measurements remains a central challenge in modern fluid dynamics, as the need for high‐fidelity data often conflicts with practical limits on sensor deployment. Existing deep learning–based methods have demonstrated promising results, but they typically depend on simplifying assumptions such as two‐dimensional domains, predefined governing equations, synthetic datasets derived from idealized flow physics, and unconstrained sensor placement. In this work, we address these limitations by studying flow reconstruction under realistic conditions and introducing a \emph{directional transport‐aware Graph Neural Network (GNN)} that explicitly encodes both flow directionality and information transport. We further show that conventional sensor placement strategies frequently yield suboptimal configurations. To overcome this, we propose a novel \emph{Two‐Step Constrained PPO} procedure for Proximal Policy Optimization (PPO), which jointly optimizes sensor layouts by incorporating flow variability and accounts for reconstruction model's performance disparity with respect to sensor placement. We conduct comprehensive experiments under realistic assumptions to benchmark the performance of our reconstruction model and sensor placement policy. Together, they achieve significant improvements over existing methods.

TMLR Journal 2025 Journal Article

Graph Fourier Neural ODEs: Modeling Spatial-temporal Multi-scales in Molecular Dynamics

  • Fang Sun
  • Zijie Huang
  • Haixin Wang
  • Huacong Tang
  • Xiao Luo
  • Wei Wang
  • Yizhou Sun

Accurately predicting long-horizon molecular dynamics (MD) trajectories remains a significant challenge, as existing deep learning methods often struggle to retain fidelity over extended simulations. We hypothesize that one key factor limiting accuracy is the difficulty of capturing interactions that span distinct spatial and temporal scales—ranging from high-frequency local vibrations to low-frequency global conformational changes. To address these limitations, we propose **Graph Fourier Neural ODEs (GF-NODE)**, integrating a graph Fourier transform for spatial frequency decomposition with a Neural ODE framework for continuous-time evolution. Specifically, GF-NODE first decomposes molecular configurations into multiple spatial frequency modes using the graph Laplacian, then evolves the frequency components in time via a learnable Neural ODE module that captures both local and global dynamics, and finally reconstructs the updated molecular geometry through an inverse graph Fourier transform. By explicitly modeling high- and low-frequency phenomena in this unified pipeline, GF-NODE more effectively captures long-range correlations and local fluctuations alike. We provide theoretical insight through heat equation analysis on a simplified diffusion model, demonstrating how graph Laplacian eigenvalues can determine temporal dynamics scales, and crucially validate this correspondence through comprehensive empirical analysis on real molecular dynamics trajectories showing quantitative spatial-temporal correlations across diverse molecular systems. Experimental results on challenging MD benchmarks, including MD17 and alanine dipeptide, demonstrate that GF-NODE achieves state-of-the-art accuracy while preserving essential geometrical features over extended simulations. These findings highlight the promise of bridging spectral decomposition with continuous-time modeling to improve the robustness and predictive power of MD simulations. Our implementation is publicly available at https://github.com/FrancoTSolis/GF-NODE-code.

NeurIPS Conference 2025 Conference Paper

Hierarchical Frequency Tagging Probe (HFTP): A Unified Approach to Investigate Syntactic Structure Representations in Large Language Models and the Human Brain

  • Jingmin An
  • Yilong Song
  • Ruolin Yang
  • Nai Ding
  • Lingxi Lu
  • Yuxuan Wang
  • Wei Wang
  • Chu Zhuang

Large Language Models (LLMs) demonstrate human-level or even superior language abilities, effectively modeling syntactic structures, yet the specific computational units responsible remain unclear. A key question is whether LLM behavioral capabilities stem from mechanisms akin to those in the human brain. To address these questions, we introduce the Hierarchical Frequency Tagging Probe (HFTP), a tool that utilizes frequency-domain analysis to identify neuron-wise components of LLMs (e. g. , individual Multilayer Perceptron (MLP) neurons) and cortical regions (via intracranial recordings) encoding syntactic structures. Our results show that models such as GPT-2, Gemma, Gemma 2, Llama 2, Llama 3. 1, and GLM-4 process syntax in analogous layers, while the human brain relies on distinct cortical regions for different syntactic levels. Representational similarity analysis reveals a stronger alignment between LLM representations and the left hemisphere of the brain (dominant in language processing). Notably, upgraded models exhibit divergent trends: Gemma 2 shows greater brain similarity than Gemma, while Llama 3. 1 shows less alignment with the brain compared to Llama 2. These findings offer new insights into the interpretability of LLM behavioral improvements, raising questions about whether these advancements are driven by human-like or non-human-like mechanisms, and establish HFTP as a valuable tool bridging computational linguistics and cognitive neuroscience. This project is available at https: //github. com/LilTiger/HFTP.

AIIM Journal 2025 Journal Article

Human visual perception-inspired medical image segmentation network with multi-feature compression

  • Guangju Li
  • Qinghua Huang
  • Wei Wang
  • Longzhong Liu

Medical image segmentation is crucial for computer-aided diagnosis and treatment planning, directly influencing clinical decision-making. To enhance segmentation accuracy, existing methods typically fuse local, global, and various other features. However, these methods often ignore the negative impact of noise on the results during the feature fusion process. In contrast, certain regions of the human visual system, such as the inferotemporal cortex and parietal cortex, effectively suppress irrelevant noise while integrating multiple features—a capability lacking in current methods. To address this gap, we propose MS-Net, a medical image segmentation network inspired by human visual perception. MS-Net incorporates a multi-feature compression (MFC) module that mimics the human visual system’s processing of complex images, first learning various feature types and subsequently filtering out irrelevant ones. Additionally, MS-Net features a segmentation refinement (SR) module that emulates how physicians segment lesions. This module initially performs coarse segmentation to capture the lesion’s approximate location and shape, followed by a refinement step to achieve precise boundary delineation. Experimental results demonstrate that MS-Net not only attains state-of-the-art segmentation performance across three public datasets but also significantly reduces the number of parameters compared to existing models. Code is available at https: //github. com/guangguangLi/MS-Net

IJCAI Conference 2025 Conference Paper

HyperTrans: Efficient Hypergraph-Driven Cross-Domain Pattern Transfer in Image Anomaly Detection

  • Tengyu Zhang
  • Deyu Zeng
  • Baoqiang Li
  • Wei Wang
  • Wei Liu
  • Zongze Wu

Anomaly detection plays a pivotal role in industrial quality assurance processes, with cross-domain problems, exemplified by the model upgrade from RGB to 3D, being prevalent in real-world scenarios yet remaining systematically underexplored. To address the severe challenges posed by the extreme lack of datasets in target domain, we retain the knowledge from source models and explore a novel solution for anomaly detection through cross-domain learning, introducing HyperTrans. Targeting few-shot scenarios, HyperTrans centers around hypergraphs to model the relationship of the limited patch features and employs a perturbation-rectification-scoring architecture. The domain perturbation module injects and adapts channel-level statistical perturbations, mitigating style shifts during domain transfer. Subsequently, a residual hypergraph restoration module utilizes a cross-domain hypergraph to capture higher-order correlations in patches and align them across domains. Ultimately, with feature patterns exhibiting reduced domain shifts, an inter-domain scoring module aggregates similarity information between patches and normal patterns within the multi-domain subhypergraphs to make an integrated decision, generating multi-level anomaly predictions. Extensive experiments demonstrate that HyperTrans offers significant advantages in anomaly classification and anomaly segmentation tasks, outperforming state-of-the-art non-cross-domain methods in image-wise ROCAUC by 13%, 12%, and 15% in 1-shot, 2-shot, and 5-shot settings on MVTec3D AD.

YNICL Journal 2025 Journal Article

Increased glymphatic system activity and thalamic vulnerability in drug-naive somatic depression: Evidenced by DTI-ALPS index

  • Zipeng Deng
  • Wei Wang
  • Zhaowen Nie
  • Simeng Ma
  • Enqi Zhou
  • Xinhui Xie
  • Qian Gong
  • Lihua Yao

Major depressive disorder (MDD) is a significant contributor to global disease burden, with somatic symptoms frequently complicating its diagnosis and treatment. Recent advances in neuroimaging have provided insights into the neurobiological underpinnings of MDD, yet the role of the glymphatic system remains largely unexplored. This study aimed to assess glymphatic function in drug-naïve somatic depression (SMD) patients using the diffusion tensor image analysis along the perivascular space (DTI-ALPS) index. A total of 272 participants, including somatic depression patients (SMD), pure depression (PMD), and healthy controls (HC), were enrolled. We collected T1-weighted (T1w) and DTI (diffusion tensor image) scans and clinical data of all participants. The DTI-ALPS indices were calculated and compared among three groups. Gray matter regions associated with the DTI-ALPS index were identified by voxel-based morphometry analysis (VBM), revealing a cluster located in the thalamus. Then, we performed partial correlation analyses to further investigate the relationships between the DTI-ALPS index, thalamic volume, and clinical data. The DTI-ALPS index was significantly higher in the MDD group compared to the HC group, particularly in the SMD group. Furthermore, a significant positive correlation was observed between the DTI-ALPS index and thalamic volume, with lower DTI-ALPS values associated with reduced thalamic volumes, especially in the SMD group. Our findings suggest heightened glymphatic activity in MDD patients, especially SMD patients, and a potential link between glymphatic function and thalamic vulnerability. Therefore, the thalamus' vulnerability to glymphatic system function may play a role in the pathophysiology of depression, particularly somatic depression, suggesting that both the glymphatic system and the thalamus could serve as potential therapeutic or intervention targets for future treatments.

AAAI Conference 2025 Conference Paper

Intra and Inter Parser-Prompted Transformers for Effective Image Restoration

  • Cong Wang
  • Jinshan Pan
  • Liyan Wang
  • Wei Wang

We propose Intra and Inter Parser-Prompted Transformers (PPTformer) that explore useful features from visual foundation models for image restoration. Specifically, PPTformer contains two parts: an Image Restoration Network (IRNet) for restoring images from degraded observations and a Parser-Prompted Feature Generation Network (PPFGNet) for providing IRNet with reliable parser information to boost restoration. To enhance the integration of the parser within IRNet, we propose Intra Parser-Prompted Attention (IntraPPA) and Inter Parser-Prompted Attention (InterPPA) to implicitly and explicitly learn useful parser features to facilitate restoration. The IntraPPA re-considers cross attention between parser and restoration features, enabling implicit perception of the parser from a long-range and intra-layer perspective. Conversely, the InterPPA initially fuses restoration features with those of the parser, followed by formulating these fused features within an attention mechanism to explicitly perceive parser information. Further, we propose a parser-prompted feed-forward network to guide restoration within pixel-wise gating modulation. Experimental results show that PPTformer achieves state-of-the-art performance on image deraining, defocus deblurring, desnowing, and low-light enhancement.

AAAI Conference 2025 Conference Paper

InverseCoder: Self-improving Instruction-Tuned Code LLMs with Inverse-Instruct

  • Yutong Wu
  • Di Huang
  • Wenxuan Shi
  • Wei Wang
  • Yewen Pu
  • Lingzhe Gao
  • Shihao Liu
  • Ziyuan Nan

Recent advancements in open-source code large language models (LLMs) have been driven by fine-tuning on the data generated from powerful closed-source LLMs, which are expensive to obtain. This paper explores whether it is possible to use a fine-tuned open-source model to generate additional data to augment its instruction-tuning dataset. We make two observations: (1) A code snippet can serve as the response to different instructions. (2) Instruction-tuned code LLMs perform better at translating code into instructions than the reverse. Based on these observations, we propose Inverse-Instruct, a data augmentation technique that uses a fine-tuned LLM to generate additional instructions of code responses from its own training dataset. The additional instruction-response pairs are added to the original dataset, and a stronger code LLM can be obtained by fine-tuning on the augmented dataset. We empirically validate Inverse-Instruct on a range of open-source code models (e.g. CodeLlama-Python and DeepSeek-Coder) and benchmarks (e.g., HumanEval(+), MBPP(+), DS-1000 and MultiPL-E), showing it consistently improves the base models.

NeurIPS Conference 2025 Conference Paper

Localized Data Shapley: Accelerating Valuation for Nearest Neighbor Algorithms

  • Guangyi Zhang
  • Yanhao Wang
  • Chengliang Chai
  • Qiyu Liu
  • Wei Wang

Data Shapley values provide a principled approach for quantifying the contribution of individual training examples to machine learning models. However, computing these values often requires computational complexity that is exponential in the data size, and this has led researchers to pursue efficient algorithms tailored to specific machine learning models. Building on the prior success of the Shapley valuation for $K$-nearest neighbor (KNN) models, in this paper, we introduce a localized data Shapley framework that significantly accelerates the valuation of data points. Our approach leverages the distance-based local structure in the data space to decompose the global valuation problem into smaller, localized computations. Our primary contribution is an efficient valuation algorithm for a threshold-based KNN variant and shows that it provides provable speedups over the baseline under mild assumptions. Extensive experiments on real-life datasets demonstrate that our methods achieve a substantial speedup compared to previous approaches.

JBHI Journal 2025 Journal Article

MedFILIP: Medical Fine-Grained Language-Image Pre-Training

  • Xinjie Liang
  • Xiangyu Li
  • Fanding Li
  • Jie Jiang
  • Qing Dong
  • Wei Wang
  • Kuanquan Wang
  • Suyu Dong

Medical vision-language pretraining (VLP) that leverages naturally-paired medical image-report data is crucial for medical image analysis. However, existing methods struggle to accurately characterize associations between images and diseases, leading to inaccurate or incomplete diagnostic results. In this work, we propose MedFILIP, a fine-grained VLP model, introduces medical image-specific knowledge through contrastive learning, specifically: 1) An information extractor based on a large language model is proposed to decouple comprehensive disease details from reports, which excels in extracting disease deals through flexible prompt engineering, thereby effectively reducing text complexity while retaining rich information at a tiny cost. 2) A knowledge injector is proposed to construct relationships between categories and visual attributes, which help the model to make judgments based on image features, and fosters knowledge extrapolation to unfamiliar disease categories. 3) A semantic similarity matrix based on fine-grained annotations is proposed, providing smoother, information-richer labels, thus allowing fine-grained image-text alignment. 4) We validate MedFILIP on numerous datasets, e. g. , RSNA-Pneumonia, NIH ChestX-ray14, VinBigData, and COVID-19. For single-label, multi-label, and fine-grained classification, our model achieves state-of-the-art performance, the classification accuracy has increased by a maximum of 6. 69%.

AAAI Conference 2025 Conference Paper

Memorize and Rank: Elevating Large Language Models for Clinical Diagnosis Prediction

  • Mingyu Derek Ma
  • Xiaoxuan Wang
  • Yijia Xiao
  • Anthony Cuturrufo
  • Vijay S Nori
  • Eran Halperin
  • Wei Wang

Clinical diagnosis prediction models, when provided with a patient's medical history, aim to detect potential diseases early, facilitating timely intervention and improving prognostic outcomes. However, the inherent scarcity of patient data and large disease candidate space often pose challenges in developing satisfactory models for this intricate task. The exploration of leveraging Large Language Models (LLMs) for encapsulating clinical decision processes has been limited. We introduce MERA, a clinical diagnosis prediction model that bridges pertaining natural language knowledge with medical practice. We apply hierarchical contrastive learning on a disease candidate ranking list to alleviate the large decision space issue. With concept memorization through fine-tuning, we bridge the natural language clinical knowledge with medical codes. Experimental results on MIMIC-III and IV datasets show that MERA achieves the state-of-the-art diagnosis prediction performance and dramatically elevates the diagnosis prediction capabilities of generative LMs.

ICRA Conference 2025 Conference Paper

MicroASV: An Affordable 3D-Printed Centimeter-Scale Autonomous Surface Vehicle

  • Kevin Macauley
  • Zhiheng Chen
  • Wei Wang

This paper introduces the design, fabrication, and autonomous control of MicroASV, a low-cost, centimeter-scale autonomous surface Vehicle (ASV). MicroASV has a square footprint with a side length of 85 mm. Its propulsion system consists of four custom water jets arranged in a “Diamond” shaped actuator configuration, powered by magnetically coupled brushless motors. This setup allows for complete 2D mobility, enabling forward and backward motion, lateral translation, and in-place rotation. The MicroASV is built using commercially available motors and 3D-printed components, creating a modular, appendage-free structure that is simple to assemble. An onboard camera and inertial measurement unit (IMU) are integrated to enable real-time localization, with position and heading controllers developed to provide autonomous feedback control. Preliminary experiments validate the platform's effectiveness in motion, sensing, and control, establishing MicroASV as a valuable tool for studying centimeter-scale ASV control, both individually and in collective swarm operations.

IJCAI Conference 2025 Conference Paper

MMGIA: Gradient Inversion Attack Against Multimodal Federated Learning via Intermodal Correlation

  • Lele Zheng
  • Yang Cao
  • Leo Yu Zhang
  • Wei Wang
  • Yulong Shen
  • Xiaochun Cao

Multimodal federated learning (MMFL) enables collaborative model training across multiple modalities, such as images and text, without requiring direct data sharing. However, the inherent correlations between modalities introduce new privacy vulnerabilities, making MMFL more susceptible to gradient inversion attacks. In this work, we propose MMGIA, an intermodal correlation-driven gradient inversion attack that systematically exploits multimodal correlation to enhance data reconstruction quality. MMGIA consists of a two-stage optimization framework: the first stage independently reconstructs each modality using traditional gradient inversion techniques, while the second stage refines these reconstructions through pre-trained feature extractors to align modalities in a shared latent space. To further improve reconstruction accuracy, we introduce a quality-weighted fusion strategy, which dynamically integrates multimodal embeddings into a global fused representation that serves as a guiding signal for refining each modality’s reconstruction. This ensures that high-quality reconstructions contribute more to the optimization process, preventing degradation in well-reconstructed modalities while enhancing weaker ones. We conduct extensive experiments on multiple multimodal scenarios, demonstrating that MMGIA outperforms both the only existing multimodal attack and state-of-the-art single-modal attacks, revealing the heightened privacy risks in MMFL.

AAAI Conference 2025 Conference Paper

Multi-Label Few-Shot Image Classification via Pairwise Feature Augmentation and Flexible Prompt Learning

  • Han Liu
  • Yuanyuan Wang
  • Xiaotong Zhang
  • Feng Zhang
  • Wei Wang
  • Fenglong Ma
  • Hong Yu

Multi-label few-shot image classification is a crucial and challenging task due to limited annotated data and elusive category specificity. However, research on this topic is still in the rudimentary stage and few methods are available. Existing methods either leverage data augmentation to alleviate data scarcity or utilize label features as auxiliary knowledge to eliminate the negative effect caused by irrelevant categories, but they ignore the utilization of image region features for data augmentation, and overlook to learn appropriate text feature to better match the image features of specific categories. Moreover, these methods only focus on one side and do not effectively tackle the above two issues simultaneously. In this paper, we introduce a novel prototype-based multi-label few-shot learning framework that seamlessly integrates pairwise feature augmentation and flexible prompt learning. Specifically, by pairwise feature augmentation, we leverage the region features of images in the support set to generate more image features and construct image prototypes, thus alleviating the issue of data scarcity. By flexible prompt learning, we adaptively acquire class-specific prompts to build text prototypes that highly match the image features of specific classes, thereby mitigating the impact of irrelevant classes. Finally, with adaptive learnable parameters, we merge image and text prototypes to obtain the final prototypes, achieving a more powerful classifier for multi-label few-shot image classification. Extensive experimental results demonstrate that our proposed method can push the performance to a higher level.

JBHI Journal 2025 Journal Article

Multi-source Signal Fusion with Contrastive AutoEncoder for Emotion Classification

  • Shen Zhao
  • Yuzhu Hu
  • Jian Chen
  • Wei Wang
  • Xiping Hu

Emotion recognition is of great importance for human-computer interaction. Emotion recognition technology based on physiological signals has shown great potential because of its strong objectivity and real-time capability. One of the most challenging tasks in this field is how to better fuse multi-source signals to extract information as comprehensively as possible. We propose a new framework for multi-source signal fusion and emotion recognition to address key challenges in feature alignment and representation learning. First, to reduce the distance between multi-source homogeneous signals in the feature space, we design a novel Contrastive Pairs AutoEncoder (CPAE), which is for feature alignment before aggregating the signals obtained from the Dual-LSTM. We also propose a designed cross-modal frequency module (CMF-Module), using a multi-layer perceptron (MLP) to learn the real and imaginary components of the signal's frequency representation, which integrates Resblock to achieve dual-channel time-domain and frequency-domain feature extraction. Furthermore, we incorporate the hidden ordinal relationships among emotional categories into the feature space through regression loss, and constrain the feature distribution using the Wasserstein distance. Experiments on public datasets show the best performance of our proposed method by comparing with baselines. We also conduct ablation studies to better verify the effect of the proposed method.

EAAI Journal 2025 Journal Article

Multi-To-Binary: A generalizable deepfake detection approach with multi-classification guidance

  • Fei Wang
  • Bo Wang
  • Botao Jing
  • Wei Wang
  • Fei Wei
  • Junxin Chen

Visual content forgery techniques, such as Deepfake, have rapidly advanced in recent years. Due to the potential misuse of these techniques for malicious purposes, there is increasing attention to the corresponding detection methods. Most existing methods focus on specific forgery patterns, making it difficult to detect forgeries with unknown or evolving patterns. In this work, we propose a novel forgery detection framework designed to extract comprehensive features utilizing multiple classification models. More specifically, our proposed framework consists of both binary-classification and multi-classification models working collaboratively, enhanced by innovative fusion and freezing mechanisms to improve accuracy and efficiency. We conducted extensive experiments to evaluate the performance of our approach. The results demonstrate that our approach outperforms state-of-the-art techniques in terms of generalization to new forgery patterns and robustness against various types of forgeries. This demonstrates promising effectiveness for real-world applications where forgeries can be diverse and sophisticated. Our code is available at https: //github. com/Phoebe-cap/M2B-main.

AAAI Conference 2025 Conference Paper

Multi-view Evidential Learning-based Medical Image Segmentation

  • Chao Huang
  • Yushu Shi
  • Waikeung Wong
  • Chengliang Liu
  • Wei Wang
  • Zhihua Wang
  • Jie Wen

Medical image segmentation provides useful information about the shape and size of organs, which is beneficial for improving diagnosis, analysis, and treatment. Despite traditional deep learning-based models can extract domain-specific knowledge, they face a generalization bottleneck due to the limited embedded knowledge scope. Vision foundation models have been demonstrated to be effective in extracting generalizable knowledge, but they cannot extract domain-specific knowledge without fine-tuning. In this work, we propose a novel multi-view evidential learning-based framework, which can extract both domain-specific and generalizable knowledge from multi-view features by combining the advantages of traditional and vision foundation models. Specifically, a novel multi-view state space model (MV-SSM) is designed to extract task-related knowledge while removing redundant information within multi-view features. The proposed MV-SSM utilizes Mamba, a state space model, to model cross-view contextual dependencies between domain-specific and generalizable features. Additionally, evidential learning is adopted to quantify the segmentation uncertainty of the model for boundary. In special, variational Dirichlet is introduced to characterize the distribution of the result probabilities, parameterized with collected evidence to quantify uncertainty. As a result, the model can reduce the segmentation uncertainties of boundaries by optimizing the parameters of the Dirichlet distribution. Experimental results on three datasets show that our method obtains superior segmentation performance.

TIST Journal 2025 Journal Article

Multimodal Large Language Model with LoRA Fine-Tuning for Multimodal Sentiment Analysis

  • Jie Mu
  • Wei Wang
  • Wenqi Liu
  • Tiantian Yan
  • Guanglu Wang

Multimodal sentiment analysis has become a popular research topic in recent years. However, existing methods have two unaddressed limitations: (1) they use limited supervised labels to train models, which makes it impossible for model to fully learn sentiments in different modal data; (2) they employ text and image pre-trained models trained in different unimodal tasks to extract different modal features, so that the extracted features cannot take into account the interactive information between image and text. To solve these problems, in this paper we propose a Vision-Language Contrastive Learning network (VLCLNet). First, we introduce a pre-trained Large Language Model (LLM), which is trained from vast quantities of multimodal data, has better understanding ability for image and text contents, thus being effectively applied to different tasks while requiring few amount of labelled training data. Second, we adapt a Multimodal Large Language Model (MLLM), BLIP-2 (Bootstrapping Language-Image Pre-training) network, to extract multimodal fusion feature. Such MLLM can fully consider the correlation between images and texts when extracting features. In addition, due to the discrepancy between the pre-training task and the sentiment analysis task, the pre-trained model will output the suboptimal prediction results. We use Low-Rank Adaptation (LoRA) fine-tuning strategy to update the model parameters on sentiment analysis task, which avoids the issue of inconsistent task between pre-training task and downstream task. Experiments verify that the proposed VLCLNet is superior to other strong baselines.

IJCAI Conference 2025 Conference Paper

MutationGuard: A Graph and Temporal-Spatial Neural Method for Detecting Mutation Telecommunication Fraud

  • Haitao Bai
  • Pinghui Wang
  • Ruofei Zhang
  • Ziyang Zhou
  • Juxiang Zeng
  • Yulou Su
  • Li Xing
  • Zhou Su

Telecommunication fraud refers to deceptive activities in the field of communication services. This research focuses on a category of fraud identified as ''mutation telecommunication fraud". There is currently a lack of research on mutation telecommunication fraud detection, allowing this type of fraud to persist uncaught. We identify that detecting mutation fraud requires capturing multi-source patterns, including user communication graphs and temporal-spatial Voice of Call (VOC) features. Specifically, we introduce MutationGuard, which leverages Graph Neural Networks (GNN) to capture changes in user communication graphs. For VOC records, we map call start times onto a 3D cylindrical surface, thereby representing each VOC record in spatial coordinates and utilizing proposed LFFE and TCFE modules to capture local fraud behaviors and temporal behavior changes. The proposed neural modeling approach that facilitates multi-source information fusion constitutes a significant advancement in detecting mutation fraud. Experiment results reveal a significant improvement in the AUC score by 1. 52% and the F1 score by 1. 36% on the proposed telecommunication fraud dataset. Particularly, our method shows a significant improvement of 13. 93% in accuracy on mutation fraud data. We also validate the effectiveness of our method on the publicly available Sichuan Telecommunication Fraud dataset.

NeurIPS Conference 2025 Conference Paper

Neighborhood Self-Dissimilarity Attention for Medical Image Segmentation

  • Junren Chen
  • Rui Chen
  • Wei Wang
  • Junlong Cheng
  • Gang Liang
  • Liangyin Chen

Medical image segmentation based on neural networks is pivotal in promoting digital health equity. The attention mechanism increasingly serves as a key component in modern neural networks, as it enables the network to focus on regions of interest, thus improving the segmentation accuracy in medical images. However, current attention mechanisms confront an accuracy-complexity trade-off paradox: accuracy gains demand higher computational costs, while reducing complexity sacrifices model accuracy. Such a contradiction inherently restricts the real-world deployment of attention mechanisms in resource-limited settings, thus exacerbating healthcare disparities. To overcome this dilemma, we propose a parameter-free Neighborhood Self-Dissimilarity Attention (NSDA), inspired by radiologists' diagnostic patterns of prioritizing regions exhibiting substantial differences during clinical image interpretation. Unlike pairwise-similarity-based self-attention mechanisms, NSDA constructs a size-adaptive local dissimilarity measure that quantifies element-neighborhood differences. By assigning higher attention weights to regions with larger feature differences, NSDA directs the neural network to focus on high-discrepancy regions, thus improving segmentation accuracy without adding trainable parameters directly related to computational complexity. The experimental results demonstrate the effectiveness and generalization of our method. This study presents a parameter-free attention paradigm, designed with clinical prior knowledge, to improve neural network performance for medical image analysis and contribute to digital health equity in low-resource settings. The code is available at https: //github. com/ChenJunren-Lab/Neighborhood-Self-Dissimilarity-Attention.

IJCAI Conference 2025 Conference Paper

Object-Level Backdoor Attacks in RGB-T Semantic Segmentation with Cross-Modality Trigger Optimization

  • Xianghao Jiao
  • Di Wang
  • Jiawei Liang
  • Jianjie Huang
  • Wei Wang
  • Xiaochun Cao

The escalating threat of backdoor risks in deep vision models is a pressing concern. Existing research on backdoor attacks is often confined to a single modality, neglecting the challenges posed by multi-modality scene perception. This work is a pioneer of backdoor attacks in RGB-Thermal (RGB-T) semantic segmentation. We overcome the critical limitation of current segmentation backdoor attacks that indiscriminately compromise all objects of a victim class, failing to provide fine-grained control for selectively targeting specific objects as required by adversaries. To address this, we introduce a novel Object-level Backdoor Attack pipeline, termed OBA. The OBA first employs a precise data poisoning (PDP) to lock a specific victim object. Specifically, the PDP embeds the trigger into the only victim object and modifies its label’s pixels at the corresponding positions, thus enabling object-level attacks. In addition, the domain gap between static single-modality triggers and multi-modality scenarios limits the PDP. We therefore introduce a Cross-Modality Trigger Generation (CMTG) method. Through style designs of triggers and cross-modality trigger co-optimization, the target domain semantics and multi-modality model perception patterns are encoded into triggers, achieving high effectiveness, stealth, and physical feasibility of triggers. Extensive experiments show that the proposed OBA enables precise manipulation of the designated object within the specific class.

IROS Conference 2025 Conference Paper

Online Residual Model Learning for Model Predictive Control of Autonomous Surface Vehicles in Real-World Environments

  • Arturo Gamboa-Gonzalez
  • Chunlin Li
  • Michael Wehner
  • Wei Wang

Model predictive control (MPC) relies on an accurate dynamics model to achieve precise and safe robot operation. In complex and dynamic aquatic environments, developing an accurate model that captures hydrodynamic details and accounts for environmental disturbances like waves, currents, and winds is challenging for aquatic robots. In this paper, we propose an online residual model learning framework for MPC, which leverages approximate models to learn complex unmodeled dynamics and environmental disturbances in dynamic aquatic environments. We integrate offline learning from previous simulation experience with online learning from the robot’s real-time interactions with the environments. These three components—residual modeling, offline learning, and on-line learning—enable a highly sample-efficient learning process, allowing for accurate real-time inference of model dynamics in complex and dynamic conditions. We further integrate this online learning residual model into a nonlinear model predictive controller, enabling it to actively choose the optimal control actions that optimize the control performance. Extensive simulations and real-world experiments with an autonomous surface vehicle demonstrate that our residual model learning MPC significantly outperforms conventional MPCs in dynamic field environments.

NeurIPS Conference 2025 Conference Paper

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

  • Yihe Deng
  • Hritik Bansal
  • Fan Yin
  • Nanyun Peng
  • Wei Wang
  • Kai-Wei Chang

We introduce OpenVLThinker, one of the first open-source large vision–language models (LVLMs) to exhibit sophisticated chain-of-thought reasoning, achieving notable performance gains on challenging visual reasoning tasks. While text-based reasoning models (e. g. , Deepseek R1) show promising results in text-only tasks, distilling their reasoning into LVLMs via supervised fine-tuning (SFT) often results in performance degradation due to imprecise visual grounding. Conversely, purely reinforcement learning (RL)-based methods face a large search space, hindering the emergence of reflective behaviors in smaller models (e. g. , 7B LVLMs). Surprisingly, alternating between SFT and RL ultimately results in significant performance improvements after a few iterations. Our analysis reveals that the base model rarely exhibits reasoning behaviors initially, but SFT effectively surfaces these latent actions and narrows the RL search space, accelerating the development of reasoning capabilities. Each subsequent RL stage further refines the model's reasoning skills, producing higher-quality SFT data for continued self-improvement. OpenVLThinker-7B consistently advances performance across six benchmarks demanding mathematical and general reasoning, notably improving MathVista by 3. 2\%, EMMA by 1. 4\%, and HallusionBench by 2. 7\%. Beyond demonstrating the synergy between SFT and RL for complex reasoning tasks, our findings provide early evidence towards achieving R1-style reasoning in multimodal contexts.

JBHI Journal 2025 Journal Article

Partial-Label Contrastive Representation Learning for Fine-Grained Biomarkers Prediction From Histopathology Whole Slide Images

  • Yushan Zheng
  • Kun Wu
  • Jun Li
  • Kunming Tang
  • Jun Shi
  • Haibo Wu
  • Zhiguo Jiang
  • Wei Wang

In the domain of histopathology analysis, existing representation learning methods for biomarkers prediction from whole slide images (WSIs) face challenges due to the complexity of tissue subtypes and label noise problems. This paper proposed a novel partial-label contrastive representation learning approach to enhance the discrimination of histopathology image representations for fine-grained biomarkers prediction. We designed a partial-label contrastive clustering (PLCC) module for partial-label disambiguation and a dynamic clustering algorithm to sample the most representative features of each category to the clustering queue during the contrastive learning process. We conducted comprehensive experiments on three gene mutation prediction datasets, including USTC-EGFR, BRCA-HER2, and TCGA-EGFR. The results show that our method outperforms 9 existing methods in terms of Accuracy, AUC, and F1 Score. Specifically, our method achieved an AUC of 0. 950 in EGFR mutation subtyping of TCGA-EGFR and an AUC of 0. 853 in HER2 0/1+/2+/3+ grading of BRCA-HER2, which demonstrates its superiority in fine-grained biomarkers prediction from histopathology whole slide images.

JBHI Journal 2025 Journal Article

Predicting Longitudinal Visual Field Progression With Class Imbalanced Data

  • Ling Chen
  • Chun-Hung Chen
  • Wei Wang
  • Da-Wen Lu
  • Vincent S. Tseng

Glaucoma is the leading cause of irreversible blindness worldwide. The clinical standard for glaucoma diagnosis and progression tracking remains visual field (VF) testing via standard automated perimetry. One outstanding challenge of many ophthalmic prediction tasks is the issue of class imbalance, where the majority class outnumbers the minority class(es). Although this issue has been reported in several prior studies on the prediction of VF progression or glaucoma, it has not been addressed in the context of longitudinal VF data. In this work, we proposed, VF-Transformer, a transformer-based framework for VF progression prediction based on longitudinal VF examination results. In particular, we addressed the class imbalance issue by incorporating our proposed inverted class-dependent temperature (ICDT) loss and weight normalization. The proposed framework was developed and evaluated on a public VF dataset and further validated on an external hospital dataset, using accuracy, sensitivity, specificity, and area under the receiver operating characteristic curve (AUC) as evaluation metrics. Extensive experiments and comparisons with existing state-of-the-art methods and class imbalance handling strategies confirmed the effectiveness of the proposed framework in predicting VF progression in the presence of class imbalance.

NeurIPS Conference 2025 Conference Paper

Private Online Learning against an Adaptive Adversary: Realizable and Agnostic Settings

  • Bo Li
  • Wei Wang
  • Peng Ye

We revisit the problem of private online learning, in which a learner receives a sequence of $T$ data points and has to respond at each time-step a hypothesis. It is required that the entire stream of output hypotheses should satisfy differential privacy. Prior work of Golowich and Livni [2021] established that every concept class $\mathcal{H}$ with finite Littlestone dimension $d$ is privately online learnable in the realizable setting. In particular, they proposed an algorithm that achieves an $O_{d}(\log T)$ mistake bound against an oblivious adversary. However, their approach yields a suboptimal $\tilde{O}\_{d}(\sqrt{T})$ bound against an adaptive adversary. In this work, we present a new algorithm with a mistake bound of $O_{d}(\log T)$ against an adaptive adversary, closing this gap. We further investigate the problem in the agnostic setting, which is more general than the realizable setting as it does not impose any assumptions on the data. We give an algorithm that obtains a sublinear regret of $\tilde{O}_d(\sqrt{T})$ for generic Littlestone classes, demonstrating that they are also privately online learnable in the agnostic setting.

AAAI Conference 2025 Conference Paper

Promptable Representation Distribution Learning and Data Augmentation for Gigapixel Histopathology WSI Analysis

  • Kunming Tang
  • Zhiguo Jiang
  • Jun Shi
  • Wei Wang
  • Haibo Wu
  • Yushan Zheng

Gigapixel image analysis, particularly for whole slide images (WSIs), often relies on multiple instance learning (MIL). Under the paradigm of MIL, patch image representations are extracted and then fixed during the training of the MIL classifiers for efficiency consideration. However, the invariance of representations makes it difficult to perform data augmentation for WSI-level model training, which significantly limits the performance of the downstream WSI analysis. The current data augmentation methods for gigapixel images either introduce additional computational costs or result in a loss of semantic information, which is hard to meet the requirements for efficiency and stability needed for WSI model training. In this paper, we propose a Promptable Representation Distribution Learning framework (PRDL) for both patch-level representation learning and WSI-level data augmentation. Meanwhile, we explore the use of prompts to guide data augmentation in feature space, which achieves promptable data augmentation for training robust WSI-level models. The experimental results have demonstrated that the proposed method stably outperforms state-of-the-art methods.

NeurIPS Conference 2025 Conference Paper

RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering

  • Rongyang Zhang
  • Yuqing Huang
  • Chengqiang Lu
  • Qimeng Wang
  • Yan Gao
  • Yao Hu
  • Yin Xu
  • Wei Wang

In real-world scenarios, providing user queries with visually enhanced responses can considerably benefit understanding and memory, underscoring the great value of interleaved image-text generation. Despite recent progress, like the visual autoregressive model that unifies text and image processing in a single transformer architecture, generating high-quality interleaved content remains challenging. Moreover, evaluations of these interleaved sequences largely remain underexplored, with existing benchmarks often limited by unimodal metrics that inadequately assess the intricacies of combined image-text outputs. To address these issues, we present RAG-IGBench, a thorough benchmark designed specifically to evaluate the task of Interleaved Generation based on Retrieval-Augmented Generation (RAG-IG) in open-domain question answering. RAG-IG integrates multimodal large language models (MLLMs) with retrieval mechanisms, enabling the models to access external image-text information for generating coherent multimodal content. Distinct from previous datasets, RAG-IGBench draws on the latest publicly available content from social platforms and introduces innovative evaluation metrics that measure the quality of text and images, as well as their consistency. Through extensive experiments with state-of-the-art MLLMs (both open-source and proprietary) on RAG-IGBench, we provide an in-depth analysis examining the capabilities and limitations of these models. Additionally, we validate our evaluation metrics by demonstrating their high correlation with human assessments. Models fine-tuned on RAG-IGBench's training set exhibit improved performance across multiple benchmarks, confirming both the quality and practical utility of our dataset. Our benchmark is available at https: //github. com/zry13/RAG-IGBench.

NeurIPS Conference 2025 Conference Paper

Rethinking Joint Maximum Mean Discrepancy for Visual Domain Adaptation

  • Wei Wang
  • Haifeng Xia
  • Chao Huang
  • Zhengming Ding
  • Cong Wang
  • Haojie Li
  • Xiaochun Cao

In domain adaption (DA), joint maximum mean discrepancy (JMMD), as a famous distribution-distance metric, aims to measure joint probability distribution difference between the source domain and target domain, while it is still not fully explored and especially hard to be applied into a subspace-learning framework as its empirical estimation involves a tensor-product operator whose partial derivative is difficult to obtain. To solve this issue, we deduce a concise JMMD based on the Representer theorem that avoids the tensor-product operator and obtains two essential findings. First, we reveal the uniformity of JMMD by proving that previous marginal, class conditional, and weighted class conditional probability distribution distances are three special cases of JMMD with different label reproducing kernels. Second, inspired by graph embedding, we observe that the similarity weights, which strengthen the intra-class compactness in the graph of Hilbert Schmidt independence criterion (HSIC), take opposite signs in the graph of JMMD, revealing why JMMD degrades the feature discrimination. This motivates us to propose a novel loss JMMD-HSIC by jointly considering JMMD and HSIC to promote discrimination of JMMD. Extensive experiments on several cross-domain datasets could demonstrate the validity of our revealed theoretical results and the effectiveness of our proposed JMMD-HSIC.

AAAI Conference 2025 Conference Paper

ScaleOT: Privacy-utility-scalable Offsite-tuning with Dynamic LayerReplace and Selective Rank Compression

  • Kai Yao
  • Zhaorui Tan
  • Tiandi Ye
  • Lichun Li
  • Yuan Zhao
  • Wenyan Liu
  • Wei Wang
  • Jianke Zhu

Offsite-tuning is a privacy-preserving method for tuning large language models (LLMs) by sharing a lossy compressed emulator from the LLM owners with data owners for downstream task tuning. This approach protects the privacy of both the model and data owners. However, current offsite tuning methods often suffer from adaptation degradation, high computational costs, and limited protection strength due to uniformly dropping LLM layers or relying on expensive knowledge distillation. To address these issues, we propose ScaleOT, a novel privacy-utility-scalable offsite-tuning framework that effectively balances privacy and utility. ScaleOT introduces a novel layerwise lossy compression algorithm that uses reinforcement learning to obtain the importance of each layer. It employs lightweight networks, termed harmonizers, to replace the raw LLM layers. By combining important original LLM layers and harmonizers in different ratios, ScaleOT generates emulators tailored for optimal performance with various model scales for enhanced privacy protection. Additionally, we present a rank reduction method to further compress the original LLM layers, significantly enhancing privacy with negligible impact on utility. Comprehensive experiments show that ScaleOT can achieve nearly lossless offsite tuning performance compared with full fine-tuning while obtaining better model privacy.

AAAI Conference 2025 Conference Paper

SdalsNet: Self-Distilled Attention Localization and Shift Network for Unsupervised Camouflaged Object Detection

  • Peiyao Shou
  • Yixiu Liu
  • Wei Wang
  • Yaoqi Sun
  • Zhigao Zheng
  • Shangdong Zhu
  • Chenggang Yan

Unsupervised camouflaged object detection (UCOD) poses significant challenges, primarily attributed to the absence of human labels. Existing UCOD methodologies, leveraging attention mechanisms, often struggle to achieve precise localization of camouflaged objects. To overcome this limitation, we introduce a groundbreaking fully unsupervised algorithm for attention-guided camouflaged object localization, shift, and inference, termed the self-distilled attention localization and shift network (SdalsNet). In this study, we formulate an attention localization methodology aimed at accurately identifying the central coordinate of the camouflaged object. Furthermore, we propose four distinct loss functions tailored to refine the precision of attentional positioning. These loss functions effectively constrain the distances between three types of class tokens, facilitating seamless attentional shifting across the input sample. Additionally, we design a sophisticated prediction inference technique to reconstruct the binary output of an attention map, thereby providing a comprehensive understanding of the detected camouflaged objects. Experimental results on four challenging COD benchmark datasets corroborate the effectiveness of our proposed approach, demonstrating notable superiority over state-of-the-art methods.

IJCAI Conference 2025 Conference Paper

SDDiff: Boosting Radar Perception via Spatial-Doppler Diffusion

  • Shengpeng Wang
  • Xin Luo
  • Yulong Xie
  • Wei Wang

Point cloud extraction (PCE) and ego velocity estimation (EVE) are key capabilities gaining attention in 3D radar perception. However, existing work typically treats these two tasks independently, which may neglect the interplay between radar's spatial and Doppler domain features, potentially introducing additional bias. In this paper, we observe an underlying correlation between 3D points and ego velocity, which offers reciprocal benefits for PCE and EVE. To fully unlock such inspiring potential, we take the first step to design a Spatial-Doppler Diffusion (SDDiff) model for simultaneously dense PCE and accurate EVE. To seamlessly tailor it to radar perception, SDDiff improves the conventional latent diffusion process in three major aspects. First, we introduce a representation that embodies both spatial occupancy and Doppler features. Second, we design a directional diffusion with radar priors to streamline the sampling. Third, we propose Iterative Doppler Refinement to enhance the model’s adaptability to density variations and ghosting effects. Extensive evaluations show that SDDiff significantly outperforms state-of-the-art baselines by achieving 59% higher in EVE accuracy, 4X greater in valid generation density while boosting PCE effectiveness and reliability. The code and dataset will be available on https: //github. com/StellarEsti/SDDiff.

AAAI Conference 2025 Conference Paper

Security Attacks on LLM-based Code Completion Tools

  • Wen Cheng
  • Ke Sun
  • Xinyu Zhang
  • Wei Wang

The rapid development of large language models (LLMs) has significantly advanced code completion capabilities, giving rise to a new generation of LLM-based Code Completion Tools (LCCTs). Unlike general-purpose LLMs, these tools possess unique workflows, integrating multiple information sources as input and prioritizing code suggestions over natural language interaction, which introduces distinct security challenges. Additionally, LCCTs often rely on proprietary code datasets for training, raising concerns about the potential exposure of sensitive data. This paper exploits these distinct characteristics of LCCTs to develop targeted attack methodologies on two critical security risks: jailbreaking and training data extraction attacks. Our experimental results expose significant vulnerabilities within LCCTs, including a 99.4% success rate in jailbreaking attacks on GitHub Copilot and a 46.3% success rate on Amazon Q. Furthermore, We successfully extracted sensitive user data from GitHub Copilot, including 54 real email addresses and 314 physical addresses associated with GitHub usernames. Our study also demonstrates that these code-based attack methods are effective against general-purpose LLMs, highlighting a broader security misalignment in the handling of code by modern LLMs. These findings underscore critical security challenges associated with LCCTs and suggest essential directions for strengthening their security frameworks.

EAAI Journal 2025 Journal Article

Semantic analysis-based recommender system using sequential clustering and convolutional neural network

  • Yanjun Xu
  • Chunqi Tian
  • Wei Wang
  • Lizhi Bai

Accurate prediction of user preferences and generation of personalized recommendations remain as critical challenges in intelligent recommendation systems. In this study, we propose a novel recommendation model that transforms the rating prediction problem into a single-label multiclass classification task. The model integrates three key components: (1) ordered clustering information derived from user review text similarity, (2) rating rank similarity reflecting users’ behavioral tendencies, and (3) a convolutional neural network (CNN) to extract semantic representations from user textual data. First, user review embeddings are clustered to capture high-level semantic preferences, where cluster indices are utilized as ordered categorical features. Second, rating rank similarity features are constructed by comparing the relative ranking of items rated by similar users. These features are fused and fed into a CNN model, which outputs a predicted rating class (e. g. , 1–5 stars) for each unobserved item, treated as a single-label classification target. To generate final Top-N recommendations, we further incorporate user-specific rating habits and item popularity to re-rank the classification outputs. The experimental results on public benchmark datasets indicate that our model substantially improves the prediction accuracy and recommendation quality compared with existing baselines. The proposed method offers a robust and interpretable approach to bridging textual review semantics, user behavior, and deep learning for rating-aware personalized recommendation.

NeurIPS Conference 2025 Conference Paper

Symmetry-Preserving Conformer Ensemble Networks for Molecular Representation Learning

  • Yanqiao Zhu
  • Yidan Shi
  • Yuanzhou Chen
  • Fang Sun
  • Yizhou Sun
  • Wei Wang

Molecular representation learning has emerged as a promising approach for modeling molecules with deep learning in chemistry and beyond. While 3D geometric models effectively capture molecular structure, they typically process single static conformers, overlooking the inherent flexibility and dynamics of molecules. In reality, many molecular properties depend on distributions of thermodynamically accessible conformations rather than single structures. Recent works show that learning from conformer ensembles can improve molecular representations, but existing approaches either produce unphysical structures through averaging or require restrictive molecular alignment. In this paper, we propose SymmetryPreserving Conformer Ensemble networks (SPiCE), which introduces two key innovations: (1) geometric mixture-of-experts for selective processing of scalar and vector features, and (2) hierarchical ensemble encoding that combines ensemblelevel representation with cross-conformer integration. Crucially, SPiCE ensures physically meaningful representations by maintaining joint equivariance to geometric transformations of individual conformers and conformer permutations. Extensive experiments demonstrate that SPiCE consistently outperforms existing conformer ensemble methods and state-of-the-art structural aggregation models across quantum mechanical and biological property prediction tasks.

AAAI Conference 2025 Conference Paper

Template-Driven LLM-Paraphrased Framework for Tabular Math Word Problem Generation

  • Xiaoqiang Kang
  • Zimu Wang
  • Xiaobo Jin
  • Wei Wang
  • Kaizhu Huang
  • Qiufeng Wang

Solving tabular math word problems (TMWPs) has become a critical role in evaluating the mathematical reasoning ability of large language models (LLMs), where large-scale TMWP samples are commonly required for fine-tuning. Since the collection of high-quality TMWP datasets is costly and time-consuming, recent research has concentrated on automatic TMWP generation. However, current generated samples usually suffer from issues of either correctness or diversity. In this paper, we propose a Template-driven LLM-paraphrased (TeLL) framework for generating high-quality TMWP samples with diverse backgrounds and accurate tables, questions, answers, and solutions. To this end, we first extract templates from existing real samples to generate initial problems, ensuring correctness. Then, we adopt an LLM to extend templates and paraphrase problems, obtaining diverse TMWP samples. Furthermore, we find the reasoning annotation is important for solving TMWPs. Therefore, we propose to enrich each solution with illustrative reasoning steps. Through the proposed framework, we construct a high-quality dataset TabMWP-TeLL by adhering to the question types in the TabMWP dataset, and we conduct extensive experiments on a variety of LLMs to demonstrate the effectiveness of TabMWP-TeLL in improving TMWP-solving performance.

YNIMG Journal 2025 Journal Article

The brain-gut microbiota network (BGMN) is correlated with symptom severity and neurocognition in patients with schizophrenia

  • Runlin Peng
  • Wei Wang
  • Liqin Liang
  • Rui Han
  • Yi Li
  • Haiyuan Wang
  • Yuran Wang
  • Wenhao Li

The association between the human brain and gut microbiota, known as the "brain-gut-microbiota axis", is involved in the neuropathological mechanisms of schizophrenia (SZ); however, its association patterns and correlations with symptom severity and neurocognition are still largely unknown. In this study, 43 SZ patients and 55 normal controls (NCs) were included, and resting-state functional magnetic resonance imaging (rs-fMRI) and gut microbiota data were acquired for each participant. First, the brain features of brain images and functional brain networks were computed from rs-fMRI data; the gut features of gut microbiota abundance and the gut microbiota network were computed from gut microbiota data. Second, we propose a novel methodology to construct an individual brain-gut microbiota network (BGMN) for each participant by combining the brain and gut features via multiple strategies. Third, discriminative models between SZ patients and NCs were built using the connectivity matrices of the BGMN as input features. Moreover, the correlations between the most discriminative features and the scores of symptom severity and neurocognition were analyzed in SZ patients. The results showed that the best discriminative model between SZ patients and NCs was achieved using the connectivity matrices of the BGMN when all the brain and gut features were integrated, with an accuracy of 0.90 and an area under the curve value of 0.97. The most discriminative features were related primarily to the genera Faecalibacterium and Collinsella, in which the genus Faecalibacterium was linked to the visual system and subcortical cortices and the genus Collinsella was linked to the default network and subcortical cortices. Furthermore, parts of the most discriminative features were significantly correlated with the scores of neurocognition in the SZ patients. The methodology for constructing individual BGMNs proposed in this study can help us reveal the associations between the brain and gut microbiota and understand the neuropathology of SZ.

NeurIPS Conference 2025 Conference Paper

Time-IMM: A Dataset and Benchmark for Irregular Multimodal Multivariate Time Series

  • Ching Chang
  • Jeehyun Hwang
  • Yidan Shi
  • Haixin Wang
  • Wei Wang
  • Wen-Chih Peng
  • Tien-Fu Chen

Time series data in real-world applications such as healthcare, climate modeling, and finance are often irregular, multimodal, and messy, with varying sampling rates, asynchronous modalities, and pervasive missingness. However, existing benchmarks typically assume clean, regularly sampled, unimodal data, creating a significant gap between research and real-world deployment. We introduce Time-IMM, a dataset specifically designed to capture cause-driven irregularity in multimodal multivariate time series. Time-IMM represents nine distinct types of time series irregularity, categorized into trigger-based, constraint-based, and artifact-based mechanisms. Complementing the dataset, we introduce IMM-TSF, a benchmark library for forecasting on irregular multimodal time series, enabling asynchronous integration and realistic evaluation. IMM-TSF includes specialized fusion modules, including a timestamp-to-text fusion module and a multimodality fusion module, which support both recency-aware averaging and attention-based integration strategies. Empirical results demonstrate that explicitly modeling multimodality on irregular time series data leads to substantial gains in forecasting performance. Time-IMM and IMM-TSF provide a foundation for advancing time series analysis under real-world conditions. The dataset is publicly available at \url{https: //github. com/blacksnail789521/Time-IMM}, and the benchmark library can be accessed at \url{https: //github. com/blacksnail789521/IMM-TSF}.

AAAI Conference 2025 Conference Paper

Towards Better Robustness Against Natural Corruptions in Document Tampering Localization

  • Huiru Shao
  • Kaizhu Huang
  • Wei Wang
  • Xiaowei Huang
  • Qiufeng Wang

Marvelous advances have been exhibited in recent document tampering localization (DTL) systems. However, confronted with corrupted tampered document images, their vulnerability is fatal in real-world scenarios. While robustness against adversarial attack has been extensively studied by adversarial training (AT), the robustness on natural corruptions remains under-explored for DTL. In this paper, to overcome forensic dependency, we propose the adversarial forensic regularization (AFR) based on min-max optimization to improve robustness. Specifically, we adopt mutual information (MI) to represent forensic dependency between two random variable over tampered and authentic pixels spaces, where the MI can be approximated by Jensen-Shannon-Divergence (JSD) with empirical sampling. To further enable a trade-off between predictive representations in clean tampered document pixels and robust ones in corrupted pixels, an additional regularization term is formulated with divergence between clean and perturbed pixels distribution (DDR). Following min-max optimization framework, our method can also work well against adversarial attacks. To evaluate our proposed method, we collect a dataset (i.e., TSorie-CRP) for evaluating robustness against natural corruptions in real scenarios. Extensive experiments demonstrate the effectiveness of our method against natural corruptions. Without any surprise, our method also achieves good performance against adversarial attack on DTL benchmark datasets.

TMLR Journal 2025 Journal Article

Towards LifeSpan Cognitive Systems

  • Yu Wang
  • Chi Han
  • Tongtong Wu
  • Xiaoxin He
  • Wangchunshu Zhou
  • Nafis Sadeq
  • Xiusi Chen
  • Zexue He

Building a human-like system that continuously interacts with complex environments—whether simulated digital worlds or human society—presents several key challenges. Central to this is enabling continuous, high-frequency interactions, where the interactions are termed experiences. We refer to this envisioned system as the LifeSpan Cognitive System (LSCS). A critical feature of LSCS is its ability to engage in incremental and rapid updates while retaining and accurately recalling past experiences. In this paper we focus on the domain of Large Language Models (LLMs), where we identify two major challenges: (1) Abstraction and Experience Merging, and (2) Long-term Retention with Accurate Recall. These properties are essential for storing new experiences, organizing past experiences, and responding to the environment in ways that leverage relevant historical data. Unlike language models with continual learning, which typically rely on large corpora for fine-tuning and focus on improving performance within specific domains or tasks, LSCS must rapidly and incrementally update with new information from its environment at a high frequency. Existing technologies with the potential of solving the above two major challenges can be classified into four classes based on a conceptual metric called Storage Complexity, which measures the relative space required to store past experiences. Each of these four classes of technologies has its own strengths and limitations while we argue none of them alone can achieve LSCS alone. To this end, we propose a potential instantiation for LSCS that can integrate all four classes of technologies. The new instantiation, serving as a conjecture, operates through two core processes: Absorbing Experiences and Generating Responses.

IJCAI Conference 2025 Conference Paper

Unlocking the Potential of Lightweight Quantized Models for Deepfake Detection

  • Renshuai Tao
  • Ziheng Qin
  • Yifu Ding
  • Chuangchuang Tan
  • Jiakai Wang
  • Wei Wang

Deepfake detection is increasingly crucial due to the rapid rise of AI-generated content. Existing methods achieve high performance relying on computationally intensive large models, making real-time detection on resource-constrained edge devices challenging. Given that deepfake detection is a binary classification task, there is potential for model compression and acceleration. In this paper, we propose a low-bit quantization framework for lightweight and efficient deepfake detection. The Connected Quantized Block extracts common forgery features via the quantized path and retains method-specific textures through the shortcut connections. Additionally, the Shifted Logarithmic Redistribution Quantizer mitigates information loss in near-zero domains by unfolding the unbalanced activations, enabling finer quantization granularity. Comprehensive experiments demonstrate that this new framework significantly reduces 10. 8x computational costs and 12. 4x storage requirements while maintaining high detection performance, even surpassing SOTA methods using less than 5% FLOPs, paving the way for efficient deepfake detection in resource-limited scenarios.

AAAI Conference 2025 Conference Paper

Unsupervised Region-Based Image Editing of Denoising Diffusion Models

  • Zixiang Li
  • Yue Song
  • Renshuai Tao
  • Xiaohong Jia
  • Yao Zhao
  • Wei Wang

Although diffusion models have achieved remarkable success in the field of image generation, their latent space remains under-explored. Current methods for identifying semantics within latent space often rely on external supervision, such as textual information and segmentation masks. In this paper, we propose a method to identify semantic attributes in the latent space of pre-trained diffusion models without any further training. By projecting the Jacobian of the targeted semantic region into a low-dimensional subspace which is orthogonal to the non-masked regions, our approach facilitates precise semantic discovery and control over local masked areas, eliminating the need for annotations. We conducted extensive experiments across multiple datasets and various architectures of diffusion models, achieving state-of-the-art performance. In particular, for some specific face attributes, the performance of our proposed method even surpasses that of supervised approaches, demonstrating its superior ability in editing local image properties.

NeurIPS Conference 2025 Conference Paper

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought

  • Chao Huang
  • Benfeng Wang
  • Wei Wang
  • Jie Wen
  • Chengliang Liu
  • Li Shen
  • Xiaochun Cao

Recent advancements in reasoning capability of Multimodal Large Language Models (MLLMs) demonstrate its effectiveness in tackling complex visual tasks. However, existing MLLM-based Video Anomaly Detection (VAD) methods remain limited to shallow anomaly descriptions without deep reasoning. In this paper, we propose a new task named Video Anomaly Reasoning (VAR), which aims to enable deep analysis and understanding of anomalies in the video by requiring MLLMs to think explicitly before answering. To this end, we propose Vad-R1, an end-to-end MLLM-based framework for VAR. Specifically, we design a Perception-to-Cognition Chain-of-Thought (P2C-CoT) that simulates the human process of recognizing anomalies, guiding the MLLMs to reason about anomalies step-by-step. Based on the structured P2C-CoT, we construct Vad-Reasoning, a dedicated dataset for VAR. Furthermore, we propose an improved reinforcement learning algorithm AVA-GRPO, which explicitly incentivizes the anomaly reasoning capability of MLLMs through a self-verification mechanism with limited annotations. Experimental results demonstrate that Vad-R1 achieves superior performance, outperforming both open-source and proprietary models on VAD and VAR tasks.

TAAS Journal 2025 Journal Article

Vehicle Dynamics and Interaction for Trajectory Prediction and Traffic Control

  • Jian Chen
  • Shaorui Zhou
  • Wei Wang
  • Yuzhu Hu
  • Jianqing Li
  • Ben-guo He
  • Junxin Chen
  • Marwan Omar

Trajectory prediction is a crucial challenge in autonomous vehicle motion planning and decision-making techniques. However, existing methods face limitations in accurately capturing vehicle dynamics and interactions. To address this issue, this article proposes a novel approach to extracting vehicle velocity and acceleration, enabling the learning of vehicle dynamics and encoding them as auxiliary information. The VDI-LSTM model is designed, incorporating graph convolution and attention mechanisms to capture vehicle interactions using trajectory data and dynamic information. Specifically, a dynamics encoder is designed to capture the dynamic information, a dynamic graph is employed to represent vehicle interactions, and an attention mechanism is introduced to enhance the performance of LSTM and graph convolution. To demonstrate the effectiveness of our model, extensive experiments are conducted, including comparisons with several baselines and ablation studies on real-world highway datasets. Experimental results show that VDI-LSTM outperforms other baselines compared, which obtains a 3% improvement on the average RMSE indicator over the five prediction steps.

JMLR Journal 2024 Journal Article

A flexible empirical Bayes approach to multiple linear regression and connections with penalized regression

  • Youngseok Kim
  • Wei Wang
  • Peter Carbonetto
  • Matthew Stephens

We introduce a new empirical Bayes approach for large-scale multiple linear regression. Our approach combines two key ideas: (i) the use of flexible "adaptive shrinkage" priors, which approximate the nonparametric family of scale mixture of normal distributions by a finite mixture of normal distributions; and (ii) the use of variational approximations to efficiently estimate prior hyperparameters and compute approximate posteriors. Combining these two ideas results in fast and flexible methods, with computational speed comparable to fast penalized regression methods such as the Lasso, and with competitive prediction accuracy across a wide range of scenarios. Further, we provide new results that establish conceptual connections between our empirical Bayes methods and penalized methods. Specifically, we show that the posterior mean from our method solves a penalized regression problem, with the form of the penalty function being learned from the data by directly solving an optimization problem (rather than being tuned by cross-validation). Our methods are implemented in an R package, mr.ash.alpha, available from https://github.com/stephenslab/mr.ash.alpha. [abs] [ pdf ][ bib ] [ code ] &copy JMLR 2024. ( edit, beta )

AAAI Conference 2024 Conference Paper

A Unified Environmental Network for Pedestrian Trajectory Prediction

  • Yuchao Su
  • Yuanman Li
  • Wei Wang
  • Jiantao Zhou
  • Xia Li

Accurately predicting pedestrian movements in complex environments is challenging due to social interactions, scene constraints, and pedestrians' multimodal behaviors. Sequential models like long short-term memory fail to effectively integrate scene features to make predicted trajectories comply with scene constraints due to disparate feature modalities of scene and trajectory. Though existing convolution neural network (CNN) models can extract scene features, they are ineffective in mapping these features into scene constraints for pedestrians and struggle to model pedestrian interactions due to the loss of target pedestrian information. To address these issues, we propose a unified environmental network based on CNN for pedestrian trajectory prediction. We introduce a polar-based method to reflect the distance and direction relationship between any position in the environment and the target pedestrian. This enables us to simultaneously model scene constraints and pedestrian social interactions in the form of feature maps. Additionally, we capture essential local features in the feature map, characterizing potential multimodal movements of pedestrians at each time step to prevent redundant predicted trajectories. We verify the performance of our proposed model on four trajectory prediction datasets, encompassing both short-term and long-term predictions. The experimental results demonstrate the superiority of our approach over existing methods.

AAAI Conference 2024 Conference Paper

AE-NeRF: Audio Enhanced Neural Radiance Field for Few Shot Talking Head Synthesis

  • Dongze Li
  • Kang Zhao
  • Wei Wang
  • Bo Peng
  • Yingya Zhang
  • Jing Dong
  • Tieniu Tan

Audio-driven talking head synthesis is a promising topic with wide applications in digital human, film making and virtual reality. Recent NeRF-based approaches have shown superiority in quality and fidelity compared to previous studies. However, when it comes to few-shot talking head generation, a practical scenario where only few seconds of talking video is available for one identity, two limitations emerge: 1) they either have no base model, which serves as a facial prior for fast convergence, or ignore the importance of audio when building the prior; 2) most of them overlook the degree of correlation between different face regions and audio, e.g., mouth is audio related, while ear is audio independent. In this paper, we present Audio Enhanced Neural Radiance Field (AE-NeRF) to tackle the above issues, which can generate realistic portraits of a new speaker with few-shot dataset. Specifically, we introduce an Audio Aware Aggregation module into the feature fusion stage of the reference scheme, where the weight is determined by the similarity of audio between reference and target image. Then, an Audio-Aligned Face Generation strategy is proposed to model the audio related and audio independent regions respectively, with a dual-NeRF framework. Extensive experiments have shown AE-NeRF surpasses the state-of-the-art on image fidelity, audio-lip synchronization, and generalization ability, even in limited training set or training iterations.

JBHI Journal 2024 Journal Article

An Ensemble Classification Model for Depression Based on Wearable Device Sleep Data

  • Yuzhu Hu
  • Jian Chen
  • Junxin Chen
  • Wei Wang
  • Shen Zhao
  • Xiping Hu

Depression is one of the most common mental disorders, with sleep disturbances as typical symptoms. With the popularity of wearable devices increasing in recent years, more and more people wear portable devices to track sleep quality. Based on this, we believe that depression detection through wearable sleep data is more intelligent and economical. However, the majority of wearable devices face the problem of missing data during the data collection process. Otherwise, most existing studies of depression identification focus on the utilization of complex data, making it difficult to generalize and susceptible to noise interference. To address these issues, we propose a systematic ensemble classification model for depression (ECD). For the missing data problem of wearable devices, we design an improved GAIN method to further control the generation range of interpolated values, which can achieve a more reasonable treatment of missing values. Compared with the original GAIN approach, the improved method shows a 28. 56% improvement when using MAE as the metric. For depression recognition, we use ensemble learning to construct a depression classification model which combines five classification models, including SVM, KNN, LR, CBR, and DT. Ensemble learning can improve the model's robustness and generalization. The voting mechanism is used in several places to improve noise immunity. The final classification model performed great on the dataset, with a precision of 92. 55% and a recall of 91. 89%. These results illustrate how efficient this method is in automatically detecting depression.

TIST Journal 2024 Journal Article

Biomedical Information Retrieval with Positive-Unlabeled Learning and Knowledge Graphs

  • Yuqi Wang
  • Qiuyi Chen
  • Haiyang Zhang
  • Wei Wang
  • Qiufeng Wang
  • Yushan Pan
  • Liangru Xie
  • Kaizhu Huang

The rapid growth of biomedical publications has presented significant challenges in the field of information retrieval. Most existing work focuses on document retrieval given explicit queries. However, in real applications such as curated biomedical database maintenance, explicit queries are missing. In this paper, we propose a two-step model for biomedical information retrieval in the case that only a small set of example documents is available without explicit queries. Initially, we extract keywords from the observed documents using large pre-trained language models and biomedical knowledge graphs. These keywords are then enriched with domain-specific entities. Information retrieval techniques can subsequently use the collected entities to rank the documents. Following this, we introduce an iterative Positive-Unlabeled learning method to classify all unlabeled documents. Experiments conducted on the PubMed dataset demonstrate that the proposed technique outperforms the state-of-the-art positive-unlabeled learning methods. The results underscore the effectiveness of integrating large language models and biomedical knowledge graphs in improving zero-shot information retrieval performance in the biomedical domain.

IROS Conference 2024 Conference Paper

CLAT: Convolutional Local Attention Tracker for Real-time UAV Target Tracking System with Feedback Information

  • Xiaolou Sun
  • Zhibin Quan
  • Wei Wang
  • Wufei Si
  • Chunyan Wang
  • Yuntian Li
  • Yuan Wu
  • Meng Shen

Real-time UAV vision target tracking systems encounter the intricate challenges of striking a trade-off for tracking speed and performance, and the robustness of the following control. In existing tracking systems, the global attention mechanism enhances tracking performance, but it introduces higher computational complexity, impacting target tracking speed; the local attention mechanism can reduce computational complexity but often exhibits limitations in modeling the receptive field. In this paper, we propose a new framework named Convolutional Local Attention Tracker (CLAT) to address these challenges. Firstly, we design a hierarchical convolutional local attention structure as the feature extractor for CLAT. This leverages convolutional projection before local window partitioning, facilitating connections between non-overlapping windows and expanding the receptive field. Secondly, we introduce a streamlined feature fusion network comprising the unshared-weights convolutional layer and a global attention network. The whole design can balance speed and accuracy. Furthermore, to enhance servo control robustness, we have redesigned the upper-level controller by integrating all bounding box information. To capture feedback spatiotemporal information in CLAT, a dynamic template update is implemented by incorporating an IOU head into the predictor. Extensive experiments on visual tracking benchmarks and in the real world demonstrate that CLAT achieves competitive performance. Moreover, we have developed a comprehensive tracking system demonstration capable of precisely tracking targets across various categories. The tracker code will be released on https://github.com/xiaolousun/refine-pytracking.git.

AAAI Conference 2024 Conference Paper

Correlation Matching Transformation Transformers for UHD Image Restoration

  • Cong Wang
  • Jinshan Pan
  • Wei Wang
  • Gang Fu
  • Siyuan Liang
  • Mengzhu Wang
  • Xiao-ming Wu
  • Jun Liu

This paper proposes UHDformer, a general Transformer for Ultra-High-Definition (UHD) image restoration. UHDformer contains two learning spaces: (a) learning in high-resolution space and (b) learning in low-resolution space. The former learns multi-level high-resolution features and fuses low-high features and reconstructs the residual images, while the latter explores more representative features learning from the high-resolution ones to facilitate better restoration. To better improve feature representation in low-resolution space, we propose to build feature transformation from the high-resolution space to the low-resolution one. To that end, we propose two new modules: Dual-path Correlation Matching Transformation module (DualCMT) and Adaptive Channel Modulator (ACM). The DualCMT selects top C/r (r is greater or equal to 1 which controls the squeezing level) correlation channels from the max-pooling/mean-pooling high-resolution features to replace low-resolution ones in Transformers, which can effectively squeeze useless content to improve the feature representation in low-resolution space to facilitate better recovery. The ACM is exploited to adaptively modulate multi-level high-resolution features, enabling to provide more useful features to low-resolution space for better learning. Experimental results show that our UHDformer reduces about ninety-seven percent model sizes compared with most state-of-the-art methods while significantly improving performance under different training sets on 3 UHD image restoration tasks, including low-light image enhancement, image dehazing, and image deblurring. The source codes will be made available at https://github.com/supersupercong/UHDformer.

AAAI Conference 2024 Conference Paper

Deep Variational Incomplete Multi-View Clustering: Exploring Shared Clustering Structures

  • Gehui Xu
  • Jie Wen
  • Chengliang Liu
  • Bing Hu
  • Yicheng Liu
  • Lunke Fei
  • Wei Wang

Incomplete multi-view clustering (IMVC) aims to reveal shared clustering structures within multi-view data, where only partial views of the samples are available. Existing IMVC methods primarily suffer from two issues: 1) Imputation-based methods inevitably introduce inaccurate imputations, which in turn degrade clustering performance; 2) Imputation-free methods are susceptible to unbalanced information among views and fail to fully exploit shared information. To address these issues, we propose a novel method based on variational autoencoders. Specifically, we adopt multiple view-specific encoders to extract information from each view and utilize the Product-of-Experts approach to efficiently aggregate information to obtain the common representation. To enhance the shared information in the common representation, we introduce a coherence objective to mitigate the influence of information imbalance. By incorporating the Mixture-of-Gaussians prior information into the latent representation, our proposed method is able to learn the common representation with clustering-friendly structures. Extensive experiments on four datasets show that our method achieves competitive clustering performance compared with state-of-the-art methods.

AAAI Conference 2024 Conference Paper

Depression Detection via Capsule Networks with Contrastive Learning

  • Han Liu
  • Changya Li
  • Xiaotong Zhang
  • Feng Zhang
  • Wei Wang
  • Fenglong Ma
  • Hongyang Chen
  • Hong Yu

Depression detection is a challenging and crucial task in psychological illness diagnosis. Utilizing online user posts to predict whether a user suffers from depression seems an effective and promising direction. However, existing methods suffer from either poor interpretability brought by the black-box models or underwhelming performance caused by the completely separate two-stage model structure. To alleviate these limitations, we propose a novel capsule network integrated with contrastive learning for depression detection (DeCapsNet). The highlights of DeCapsNet can be summarized as follows. First, it extracts symptom capsules from user posts by leveraging meticulously designed symptom descriptions, and then distills them into class-indicative depression capsules. The overall workflow is in an explicit hierarchical reasoning manner and can be well interpreted by the Patient Health Questionnaire-9 (PHQ9), which is one of the most widely adopted questionnaires for depression diagnosis. Second, it integrates with contrastive learning, which can facilitate the embeddings from the same class to be pulled closer, while simultaneously pushing the embeddings from different classes apart. In addition, by adopting the end-to-end training strategy, it does not necessitate additional data annotation, and mitigates the potential adverse effects from the upstream task to the downstream task. Extensive experiments on three widely-used datasets show that in both within-dataset and cross-dataset scenarios our proposed method outperforms other strong baselines significantly.

YNICL Journal 2024 Journal Article

Dysregulated cerebral blood flow, rather than gray matter Volume, exhibits stronger correlations with blood inflammatory and lipid markers in depression

  • Lijun Kang
  • Wei Wang
  • Zhaowen Nie
  • Qian Gong
  • Lihua Yao
  • Dan Xiang
  • Nan Zhang
  • Ning Tu

Arterial spin labeling (ASL) can be used to detect differences in perfusion for multiple brain regions thought to be important in major depressive disorder (MDD). However, the potential of cerebral blood flow (CBF) to predict MDD and its correlations between the blood lipid levels and immune markers, which are closely related to MDD and brain function change, remain unclear. The 451 individuals - 298 with MDD and 133 healthy controls who underwent MRI at a single time point with arterial spin labelling and a high resolution T1-weighted structural scan. A proportion of MDD also provided blood samples for analysis of lipid and immune markers. We performed CBF case-control comparisons, random forest model construction, and exploratory correlation analyses. Moreover, we investigated the relationship between gray matter volume (GMV), blood lipids, and the immune system within the same sample to assess the differences in CBF and GMV. We found that the left inferior parietal but supramarginal and angular gyrus were significantly different between the MDD patients and HCs (voxel-wise P < 0.001, cluster-wise FWE correction). And bilateral inferior temporal (ITG), right middle temporal gyrus and left precentral gyrus CBF predict MDD (the area under the receiver operating characteristic curve of the random forest model is 0.717) and that CBF is a more sensitive predictor of MDD than GMV. The left ITG showed a positive correlation trend with immunoglobulin G (r = 0.260) and CD4 counts (r = 0.283). The right ITG showed a correlation trend with Total Cholesterol (r = -0.249) and tumour necrosis factor-alpha (r = -0.295). Immunity and lipids were closely related to CBF change, with the immunity relationship potentially playing a greater role. The interactions between CBF, plasma lipids and immune index could therefore represent an MDD pathophysiological mechanism. The current findings provide evidence for targeted regulation of CBF or immune properties in MDD.

NeurIPS Conference 2024 Conference Paper

Enhancing Large Vision Language Models with Self-Training on Image Comprehension

  • Yihe Deng
  • Pan Lu
  • Fan Yin
  • Ziniu Hu
  • Sheng Shen
  • Quanquan Gu
  • James Zou
  • Kai-Wei Chang

Large vision language models (LVLMs) integrate large language models (LLMs) with pre-trained vision encoders, thereby activating the perception capability of the model to understand image inputs for different queries and conduct subsequent reasoning. Improving this capability requires high-quality vision-language data, which is costly and labor-intensive to acquire. Self-training approaches have been effective in single-modal settings to alleviate the need for labeled data by leveraging model's own generation. However, effective self-training remains a challenge regarding the unique visual perception and reasoning capability of LVLMs. To address this, we introduce S elf- T raining on I mage C omprehension ( STIC ), which emphasizes a self-training approach specifically for image comprehension. First, the model self-constructs a preference dataset for image descriptions using unlabeled images. Preferred responses are generated through a step-by-step prompt, while dis-preferred responses are generated from either corrupted images or misleading prompts. To further self-improve reasoning on the extracted visual information, we let the model reuse a small portion of existing instruction-tuning data and append its self-generated image descriptions to the prompts. We validate the effectiveness of STIC across seven different benchmarks, demonstrating substantial performance gains of 4. 0% on average while using 70% less supervised fine-tuning data than the current method. Further studies dive into various components of STIC and highlight its potential to leverage vast quantities of unlabeled images for self-training.

IJCAI Conference 2024 Conference Paper

Explore Internal and External Similarity for Single Image Deraining with Graph Neural Networks

  • Cong Wang
  • Wei Wang
  • Chengjin Yu
  • Jie Mu

Patch-level non-local self-similarity is an important property of natural images. However, most existing methods do not consider this property into neural networks for image deraining, thus affecting recovery performance. Motivated by this property, we find that there exists significant patch recurrence property of a rainy image, that is, similar patches tend to recur many times in one image and its multi-scale images and external images. To better model this property for image detaining, we develop a multi-scale graph network with exemplars, called MSGNN, that contains two branches: 1) internal data-based supervised branch is used to model the internal relations of similar patches from the rainy image itself and its multi-scale images and 2) external data-participated unsupervised branch is used to model the external relations of the similar patches in the rainy image and exemplar. Specifically, we construct a graph model by searching the k-nearest neighboring patches from both the rainy images in a multi-scale framework and the exemplar. After obtaining the corresponding k neighboring patches from the multi-scale images and exemplar, we build a graph and aggregate them in an attentional manner so that the graph can provide more information from similar patches for image deraining. We embed the proposed graph in a deep neural network and train it in an end-to-end manner. Extensive experiments demonstrate that the proposed algorithm performs favorably against eight state-of-the-art methods on five public synthetic datasets and one real-world dataset. The source codes will be available at https: //github. com/supersupercong/MSGNN.

AAAI Conference 2024 Conference Paper

Generalizing across Temporal Domains with Koopman Operators

  • Qiuhao Zeng
  • Wei Wang
  • Fan Zhou
  • Gezheng Xu
  • Ruizhi Pu
  • Changjian Shui
  • Christian Gagné
  • Shichun Yang

In the field of domain generalization, the task of constructing a predictive model capable of generalizing to a target domain without access to target data remains challenging. This problem becomes further complicated when considering evolving dynamics between domains. While various approaches have been proposed to address this issue, a comprehensive understanding of the underlying generalization theory is still lacking. In this study, we contribute novel theoretic results that aligning conditional distribution leads to the reduction of generalization bounds. Our analysis serves as a key motivation for solving the Temporal Domain Generalization (TDG) problem through the application of Koopman Neural Operators, resulting in Temporal Koopman Networks (TKNets). By employing Koopman Neural Operators, we effectively address the time-evolving distributions encountered in TDG using the principles of Koopman theory, where measurement functions are sought to establish linear transition relations between evolving domains. Through empirical evaluations conducted on synthetic and real-world datasets, we validate the effectiveness of our proposed approach.

NeurIPS Conference 2024 Conference Paper

GraphVis: Boosting LLMs with Visual Knowledge Graph Integration

  • Yihe Deng
  • Chenchen Ye
  • Zijie Huang
  • Mingyu Derek Ma
  • Yiwen Kou
  • Wei Wang

The rapid evolution of large language models (LLMs) has expanded their capabilities across various data modalities, extending from well-established image data to increasingly popular graph data. Given the limitation of LLMs in hallucinations and inaccuracies in recalling factual knowledge, Knowledge Graph (KG) has emerged as a crucial data modality to support more accurate reasoning by LLMs. However, integrating structured knowledge from KGs into LLMs remains challenging, as most current KG-enhanced LLM methods directly convert the KG into linearized text triples, which is not as expressive as the original structured data. To address this, we introduce GraphVis, which conserves the intricate graph structure through the visual modality to enhance the comprehension of KGs with the aid of Large Vision Language Models (LVLMs). Our approach incorporates a unique curriculum fine-tuning scheme which first instructs LVLMs to recognize basic graphical features from the images, and subsequently incorporates reasoning on QA tasks with the visual graphs. This cross-modal methodology not only markedly enhances performance on standard textual QA but also shows improved zero-shot VQA performance by utilizing synthetic graph images to augment the data for VQA tasks. We present comprehensive evaluations across commonsense reasoning QA benchmarks, where GraphVis provides an average improvement of 11. 1% over its base model and outperforms existing KG-enhanced LLM approaches. Across VQA benchmarks such as ScienceQA that share similar scientific diagram images, GraphVis provides a notable gain of 4. 32%.

JBHI Journal 2024 Journal Article

Guest Editorial AI-Empowered Internet of Things for Data-Driven Psychophysiological Computing and Patient Monitoring

  • Kai Fang
  • Wei Wang
  • Marcin Woźniak
  • Qingchen Zhang
  • Keping Yu
  • Junxin Chen
  • Amr Tolba
  • Leo Zhang

As The cornerstone of human health, physical and mental well-being are intricately linked, influencing both an individual's physical condition and their emotional state [1]. Chronic diseases such as hypertension and diabetes can have a significant impact on mental health, leading to anxiety and depression [2]. Similarly, psychological problems such as stress, anxiety, and depression can weaken the immune system, making individuals more susceptible to physical illnesses. In recent years, the rapid development of technology has brought exciting new possibilities to the field of physical and psychological health. The Internet of Things (IoT) and artificial intelligence (AI) have shown great potential in building a comprehensive health management system that empowers individuals to take a more proactive role in their well-being.

JBHI Journal 2024 Journal Article

Guest Editorial Special Issue on Data-driven Cognitive Computing for Smart Healthcare Systems

  • Syed Hassan Shah
  • Wei Wei
  • Wei Wang

With the rising costs of drugs, medical devices, and diagnostic development, the topic of data-driven cognitive computing is currently an emerging research area in smart healthcare construction. With the support of machine learning and artificial intelligence empowered cognitive computing, the significant insights and knowledge hidden behind medical data can be capitalized for process optimization, anomaly detection, energy management, and so on. The special issue is an effort to provide a platform for researchers to explore healthcare issues supported by data-driven cognitive computing-related technologies from both theoretical and practical perspectives.

EAAI Journal 2024 Journal Article

Health prognosis via feature optimization and convolutional neural network for lithium-ion batteries

  • Mingqiang Lin
  • Leisi Ke
  • Wei Wang
  • Jinhao Meng
  • Yajuan Guan
  • Ji Wu

With the rapid expansion of the electric vehicle market, the demand for lithium-ion batteries (LIBs) is exploding. The state of health (SOH) of LIBs is receiving more widespread attention, which is the key parameter for battery health management. This paper proposes a SOH estimation method for LIBs via feature optimization and convolutional neural network (CNN) to reduce the information redundancy of existing multiple features, aimed at leveraging multiple sources of features while optimizing their combination to minimize redundancy. Firstly, multiple features are extracted from different perspectives, including electrical, thermodynamic, and electrochemical properties, to comprehensively characterize the aging of batteries. Secondly, we construct a SOH estimator based on principal component analysis (PCA) with CNN (PCA-CNN). Finally, the dimension of features is optimized with a simulated annealing algorithm (SA) under the mean-variance objective function. Moreover, Comparative experiments are conducted on the Oxford dataset for validation. The results demonstrate the effectiveness of the proposed multi-feature description method in terms of accuracy and smoothness. Compared to traditional CNN methods and fixed-dimension PCA-CNN, this estimation approach significantly improves performance, showing more than 20% and 30% increases in the key metrics of MAE and RMSE, respectively. This study successfully optimized feature combinations to reduce redundancy within the feature set while enhancing the accuracy of SOH estimation.

IROS Conference 2024 Conference Paper

IDF-MFL: Infrastructure-free and Drift-free Magnetic Field Localization for Mobile Robot

  • Hongming Shen
  • Zhenyu Wu 0001
  • Wei Wang
  • Qiyang Lyu
  • Huiqin Zhou
  • Danwei Wang

In recent years, infrastructure-based localization methods have achieved significant progress thanks to their reliable and drift-free localization capability. However, the preinstalled infrastructures suffer from inflexibilities and high maintenance costs. This poses an interesting problem of how to develop a drift-free localization system without using the preinstalled infrastructures. In this paper, an infrastructure-free and drift-free localization system is proposed using the ambient magnetic field (MF) information, namely IDF-MFL. IDF-MFL is infrastructure-free thanks to the high distinctiveness of the ambient MF information produced by inherent ferromagnetic objects in the environment, such as steel and reinforced concrete structures of buildings, and underground pipelines. The MF-based localization problem is defined as a stochastic optimization problem with the consideration of the non-Gaussian heavy-tailed noise introduced by MF measurement outliers (caused by dynamic ferromagnetic objects), and an outlier-robust state estimation algorithm is derived to find the optimal distribution of robot state that makes the expectation of MF matching cost achieves its lower bound. The proposed method is evaluated in multiple scenarios 1, including experiments on high-fidelity simulation, and real-world environments. The results demonstrate that the proposed method can achieve high-accuracy, reliable, and real-time localization without any pre-installed infrastructures.

IJCAI Conference 2024 Conference Paper

LEAP: Optimization Hierarchical Federated Learning on Non-IID Data with Coalition Formation Game

  • Jianfeng Lu
  • Yue Chen
  • Shuqin Cao
  • Longbiao Chen
  • Wei Wang
  • Yun Xin

Although Hierarchical Federated Learning (HFL) utilizes edge servers (ESs) to alleviate communication burdens, its model performance will be degraded by non-IID data and limited communication resources. Current works often assume that data is uniformly distributed, which however contradicts the heterogeneity of IoT. Solutions involving additional model training to check the data distribution inevitably increase computational costs and the risk of privacy leakage. The challenges in solving these issues are how to reduce the impact of non-IID data without involving raw data, and how to rationalize the communication resource allocation for addressing straggler problem. To tackle these challenges, we propose a novel optimization method based on coaLition formation gamE and grAdient Projection, called LEAP. Specifically, we combine edge data distribution with coalition formation game innovatively to adjust the correlations between clients and ESs dynamically, ensuring optimal correlations. We further capture the client heterogeneity to achieve the rational bandwidth allocation from coalition perception and determine the optimal transmission power within specified delay constraints at the client level. Experimental results on four real datasets show that LEAP is able to achieve 20. 62% improvement in model accuracy compared to the state-of-the-art baselines. Moreover, LEAP effectively reduces transmission energy consumption by at least about 2. 24 times.

AAAI Conference 2024 Conference Paper

Learning Dense Correspondence for NeRF-Based Face Reenactment

  • Songlin Yang
  • Wei Wang
  • Yushi Lan
  • Xiangyu Fan
  • Bo Peng
  • Lei Yang
  • Jing Dong

Face reenactment is challenging due to the need to establish dense correspondence between various face representations for motion transfer. Recent studies have utilized Neural Radiance Field (NeRF) as fundamental representation, which further enhanced the performance of multi-view face reenactment in photo-realism and 3D consistency. However, establishing dense correspondence between different face NeRFs is non-trivial, because implicit representations lack ground-truth correspondence annotations like mesh-based 3D parametric models (e.g., 3DMM) with index-aligned vertexes. Although aligning 3DMM space with NeRF-based face representations can realize motion control, it is sub-optimal for their limited face-only modeling and low identity fidelity. Therefore, we are inspired to ask: Can we learn the dense correspondence between different NeRF-based face representations without a 3D parametric model prior? To address this challenge, we propose a novel framework, which adopts tri-planes as fundamental NeRF representation and decomposes face tri-planes into three components: canonical tri-planes, identity deformations, and motion. In terms of motion control, our key contribution is proposing a Plane Dictionary (PlaneDict) module, which efficiently maps the motion conditions to a linear weighted addition of learnable orthogonal plane bases. To the best of our knowledge, our framework is the first method that achieves one-shot multi-view face reenactment without a 3D parametric model prior. Extensive experiments demonstrate that we produce better results in fine-grained motion control and identity preservation than previous methods.

IJCAI Conference 2024 Conference Paper

Learning from Long-Tailed Noisy Data with Sample Selection and Balanced Loss

  • Lefan Zhang
  • Zhang-Hao Tian
  • Wujun Zhou
  • Wei Wang

The success of deep learning depends on large-scale and well-curated training data, while data in real-world applications are commonly long-tailed and noisy. Existing methods are usually dependent on label frequency to tackle class imbalance, while the model bias on different classes is not directly related to label frequency and the true label frequency is inaccessible under label noise. To solve this, we propose a robust method for learning from long-tailed noisy data with sample selection and balanced loss. Specifically, we separate the noisy training data into clean labeled set and unlabeled set with sample selection, and train the deep neural network in a semi-supervised manner with a balanced loss based on model bias. Extensive experiments on benchmarks demonstrate that our method outperforms existing state-of-the-art methods.

AAAI Conference 2024 Conference Paper

Liberating Seen Classes: Boosting Few-Shot and Zero-Shot Text Classification via Anchor Generation and Classification Reframing

  • Han Liu
  • Siyang Zhao
  • Xiaotong Zhang
  • Feng Zhang
  • Wei Wang
  • Fenglong Ma
  • Hongyang Chen
  • Hong Yu

Few-shot and zero-shot text classification aim to recognize samples from novel classes with limited labeled samples or no labeled samples at all. While prevailing methods have shown promising performance via transferring knowledge from seen classes to unseen classes, they are still limited by (1) Inherent dissimilarities among classes make the transformation of features learned from seen classes to unseen classes both difficult and inefficient. (2) Rare labeled novel samples usually cannot provide enough supervision signals to enable the model to adjust from the source distribution to the target distribution, especially for complicated scenarios. To alleviate the above issues, we propose a simple and effective strategy for few-shot and zero-shot text classification. We aim to liberate the model from the confines of seen classes, thereby enabling it to predict unseen categories without the necessity of training on seen classes. Specifically, for mining more related unseen category knowledge, we utilize a large pre-trained language model to generate pseudo novel samples, and select the most representative ones as category anchors. After that, we convert the multi-class classification task into a binary classification task and use the similarities of query-anchor pairs for prediction to fully leverage the limited supervision signals. Extensive experiments on six widely used public datasets show that our proposed method can outperform other strong baselines significantly in few-shot and zero-shot tasks, even without using any seen class samples.

EAAI Journal 2024 Journal Article

Lightweight railroad semantic segmentation network and distance estimation for railroad Unmanned aerial vehicle images

  • R.S. Rampriya
  • Sabari Nathan
  • R. Suganya
  • Sahaya Beni Prathiba
  • P. Shunmuga Perumal
  • Wei Wang

Derailments significantly harm railroads in terms of severity and fatality rates. Manually monitoring railway tracks is a tedious and often insufficient task to prevent derailment accidents. This necessitates the development of an automated monitoring system to oversee the condition of railway tracks. In the foreseeable future, Unmanned Aerial Vehicles (UAVs), artificial intelligence, and computer vision will play a crucial role in periodically assessing the railway environment to ensure passenger safety. This paper proposes a lightweight and efficient railroad semantic segmentation network (Lite-RSNet), which segments real-time aerial railroad images into rails, gauges, and backgrounds. The segmented regions then compute the distance between rails and obstacles, identifying potential derailment hazards in railroad applications. Lite-RSNet incorporates a multi-level feature fusion framework that includes a residual layer, a multi-head attention layer, a Concurrent Spatial and Channel Squeeze Excitation (CSSE) block, and a Convolutional Block Attention Module (CBAM). The designed approach enhances onboard processing efficiency by streamlining the model to use fewer parameters and reduce storage demands. Performance assessment of the model utilized the Rail Segmentation Dataset (RSD), which consists of high-resolution, low-altitude aerial photographs of Indian railroads. The Lite-RSNet model outperforms current leading methods, maintaining a lean structure with only 6 million parameters and achieving notable metrics of 0. 971 for the Dice Coefficient (DC) and 0. 947 for the Jaccard Index (JI) on the test dataset. Furthermore, distance estimation between an obstacle and rail has been implemented using image processing techniques as an additional work of this proposal to know about the caution of an obstacle.

NeurIPS Conference 2024 Conference Paper

LucidAction: A Hierarchical and Multi-model Dataset for Comprehensive Action Quality Assessment

  • Linfeng Dong
  • Wei Wang
  • Yu Qiao
  • Xiao Sun

Action Quality Assessment (AQA) research confronts formidable obstacles due to limited, mono-modal datasets sourced from one-shot competitions, which hinder the generalizability and comprehensiveness of AQA models. To address these limitations, we present LucidAction, the first systematically collected multi-view AQA dataset structured on curriculum learning principles. LucidAction features a three-tier hierarchical structure, encompassing eight diverse sports events with four curriculum levels, facilitating sequential skill mastery and supporting a wide range of athletic abilities. The dataset encompasses multi-modal data, including multi-view RGB video, 2D and 3D pose sequences, enhancing the richness of information available for analysis. Leveraging a high-precision multi-view Motion Capture (MoCap) system ensures precise capture of complex movements. Meticulously annotated data, incorporating detailed penalties from professional gymnasts, ensures the establishment of robust and comprehensive ground truth annotations. Experimental evaluations employing diverse contrastive regression baselines on LucidAction elucidate the dataset's complexities. Through ablation studies, we investigate the advantages conferred by multi-modal data and fine-grained annotations, offering insights into improving AQA performance. The data and code will be openly released to support advancements in the AI sports field.

AAAI Conference 2024 Conference Paper

MathAttack: Attacking Large Language Models towards Math Solving Ability

  • Zihao Zhou
  • Qiufeng Wang
  • Mingyu Jin
  • Jie Yao
  • Jianan Ye
  • Wei Liu
  • Wei Wang
  • Xiaowei Huang

With the boom of Large Language Models (LLMs), the research of solving Math Word Problem (MWP) has recently made great progress. However, there are few studies to examine the robustness of LLMs in math solving ability. Instead of attacking prompts in the use of LLMs, we propose a MathAttack model to attack MWP samples which are closer to the essence of robustness in solving math problems. Compared to traditional text adversarial attack, it is essential to preserve the mathematical logic of original MWPs during the attacking. To this end, we propose logical entity recognition to identify logical entries which are then frozen. Subsequently, the remaining text are attacked by adopting a word-level attacker. Furthermore, we propose a new dataset RobustMath to evaluate the robustness of LLMs in math solving ability. Extensive experiments on our RobustMath and two another math benchmark datasets GSM8K and MultiAirth show that MathAttack could effectively attack the math solving ability of LLMs. In the experiments, we observe that (1) Our adversarial samples from higher-accuracy LLMs are also effective for attacking LLMs with lower accuracy (e.g., transfer from larger to smaller-size LLMs, or from few-shot to zero-shot prompts); (2) Complex MWPs (such as more solving steps, longer text, more numbers) are more vulnerable to attack; (3) We can improve the robustness of LLMs by using our adversarial samples in few-shot prompts. Finally, we hope our practice and observation can serve as an important attempt towards enhancing the robustness of LLMs in math solving ability. The code and dataset is available at: https://github.com/zhouzihao501/MathAttack.

AAAI Conference 2024 System Paper

MIDDAG: Where Does Our News Go? Investigating Information Diffusion via Community-Level Information Pathways

  • Mingyu Derek Ma
  • Alexander K. Taylor
  • Nuan Wen
  • Yanchen Liu
  • Po-Nien Kung
  • Wenna Qin
  • Shicheng Wen
  • Azure Zhou

We present MIDDAG, an intuitive, interactive system that visualizes the information propagation paths on social media triggered by COVID-19-related news articles accompanied by comprehensive insights including user/community susceptibility level, as well as events and popular opinions raised by the crowd while propagating the information. Besides discovering information flow patterns among users, we construct communities among users and develop the propagation forecasting capability, enabling tracing and understanding of how information is disseminated at a higher level. A demo video and more are available at https://info-pathways.github.io.

ICRA Conference 2024 Conference Paper

MM4MM: Map Matching Framework for Multi-Session Mapping in Ambiguous and Perceptually-Degraded Environments

  • Zhenyu Wu 0001
  • Wei Wang
  • Chunyang Zhao
  • Yufeng Yue
  • Jun Zhang 0042
  • Hongming Shen
  • Danwei Wang

Multi-session mapping serves as the pre-requisite for autonomous robots to fulfill various long-term tasks (e. g. , map updating, navigation, collaboration). However, it is challenging to implement multi-session mapping in enclosed or partially enclosed ambiguous environments (e. g. , long corridors, industrial warehouses). Existing solutions either depend heavily on the matching of elementary geometric features (e. g. , points, lines, and planes), which tends to fail in environments with ambiguous geometric features; or depend on the given guess of the initial transformation matrix of multiple single-session maps, which is not always obtainable and accurate enough. The ambient magnetic field has exhibited ubiquity and high distinctiveness at different location, which makes it suitable for estimating the initial transformation matrix. Thus, this paper proposes a novel probabilistic magnetic-aware Map Matching framework for Multi-session Mapping, namely MM4MM, to estimate the relative transformation of multiple single-session maps and to build the globally consistent maps in ambiguous and perceptually-degraded environments. The key novelties of this work are the designing of the hierarchical probabilistic map matching framework and the Particle Swarm Optimization strategy to associate the magnetic data of multiple sessions. Evaluations on both simulated and real world experiments demonstrate the greatly improved utility, accuracy, and robustness of multi-session mapping over the comparative methods.

IJCAI Conference 2024 Conference Paper

Multi-Attention Based Visual-Semantic Interaction for Few-Shot Learning

  • Peng Zhao
  • Yin Wang
  • Wei Wang
  • Jie Mu
  • Huiting Liu
  • Cong Wang
  • Xiaochun Cao

Few-Shot Learning (FSL) aims to train a model that can generalize to recognize new classes, with each new class having only very limited training samples. Since extracting discriminative features for new classes with few samples is challenging, existing FSL methods leverage visual and semantic prior knowledge to guide discriminative feature learning. However, for meta-learning purposes, the semantic knowledge of the query set is unavailable, so their features lack discriminability. To address this problem, we propose a novel Multi-Attention based Visual-Semantic Interaction (MAVSI) approach for FSL. Specifically, we utilize spatial and channel attention mechanisms to effectively select discriminative visual features for the support set based on its ground-truth semantics while using all the support set semantics for each query set sample. Then, a relation module with class prototypes of the support set is employed to supervise and select discriminative visual features for the query set. To further enhance the discriminability of the support set, we introduce a visual-semantic contrastive learning module to promote the similarity between visual features and their corresponding semantic features. Extensive experiments on four benchmark datasets demonstrate that our proposed MAVSI could outperform existing state-of-the-art FSL methods.

NeurIPS Conference 2024 Conference Paper

Non-Euclidean Mixture Model for Social Network Embedding

  • Roshni G. Iyer
  • Yewen Wang
  • Wei Wang
  • Yizhou Sun

It is largely agreed that social network links are formed due to either homophily or social influence. Inspired by this, we aim at understanding the generation of links via providing a novel embedding-based graph formation model. Different from existing graph representation learning, where link generation probabilities are defined as a simple function of the corresponding node embeddings, we model the link generation as a mixture model of the two factors. In addition, we model the homophily factor in spherical space and the influence factor in hyperbolic space to accommodate the fact that (1) homophily results in cycles and (2) influence results in hierarchies in networks. We also design a special projection to align these two spaces. We call this model Non-Euclidean Mixture Model, i. e. , NMM. We further integrate NMM with our non-Euclidean graph variational autoencoder (VAE) framework, NMM-GNN. NMM-GNN learns embeddings through a unified framework which uses non-Euclidean GNN encoders, non-Euclidean Gaussian priors, a non-Euclidean decoder, and a novel space unification loss component to unify distinct non-Euclidean geometric spaces. Experiments on public datasets show NMM-GNN significantly outperforms state-of-the-art baselines on social network generation and classification tasks, demonstrating its ability to better explain how the social network is formed.

IJCAI Conference 2024 Conference Paper

Optimal Graph Learning and Nuclear Norm Maximization for Deep Cross-Domain Robust Label Propagation

  • Wei Wang
  • Hanyang Li
  • Ke Shi
  • Chao Huang
  • Yang Cao
  • Cong Wang
  • Xiaochun Cao

Domain adaptation aims to achieve label transfer from a labeled source domain to an unlabeled target domain, where the two domains exhibit different distributions. Existing methods primarily concentrate on designing a feature extractor to learn better domain-invariant features, along with developing an effective classifier for reliable predictions. In this paper, we introduce optimal graph learning to generate a cross-domain graph that effectively connects the two domains, and two domain-specific graphs to capture domain-specific structures. On the one hand, we incorporate the three graphs into the label propagation (LP) classifier to enhance its robustness to distribution difference. On the other hand, we leverage the three graphs to introduce graph embedding losses, promoting the learning of locally discriminative and domain-invariant features. Furthermore, we maximize the nuclear norm of predictions in LP to enhance class diversity, thereby improving its robustness to class imbalance problem. Correspondingly, we develop an efficient algorithm to solve the associated optimization problem. Finally, we integrate the proposed LP and graph embedding losses into a deep neural network, resulting in our proposed deep cross-domain robust LP. Extensive experiments conducted on three cross-domain benchmark datasets demonstrate that our proposed approach could outperform existing state-of-the-art domain adaptation methods.

EAAI Journal 2024 Journal Article

Path planning for dual-arm fiber patch placement with temperature loss constraints

  • Xiangli Li
  • Rui Zhou
  • Wei Wang
  • Mengde Li
  • Yi Gong
  • Miao Li

In the process of composite material molding, the role of robots has become increasingly significant. However, the laying process of composite materials is heavily influenced by factors such as temperature and pressure, which affect the quality of the layup. Currently, existing robot path planning algorithms do not account for temperature constraints, presenting challenges in achieving high-quality composite material layups. This paper proposes the RRT*-MTL (Rapidly-Exploring Random Trees* with Minimum Temperature Loss) path planning algorithm. Firstly, a neural network model for predicting temperature loss is developed. Then, temperature loss is integrated as a constraint into the RRT* (Rapidly-Exploring Random Trees*) path planning algorithm to determine the laying path with minimal temperature loss. Experimental results demonstrate that this algorithm effectively preserves the object's heat retention.

I&C Journal 2024 Journal Article

Pathwise-randomness and models of second-order arithmetic

  • George Barmpalias
  • Wei Wang

A tree is pathwise-random if all of its paths are Martin-Löf random. We show that: (a) no weakly 2-random real computes a perfect pathwise-random tree; it follows that the class of perfect pathwise-random trees is null, with respect to any computable measure; (b) there exists a positive-measure pathwise-random tree which does not compute any complete extension of Peano arithmetic; and (c) there exists a perfect pathwise-random tree which does not compute any tree of positive measure and finite randomness deficiency. We then obtain models of second-order arithmetic that separate principles below weak Königs lemma.

NeurIPS Conference 2024 Conference Paper

Physics-Informed Regularization for Domain-Agnostic Dynamical System Modeling

  • Zijie Huang
  • Wanjia Zhao
  • Jingdong Gao
  • Ziniu Hu
  • Xiao Luo
  • Yadi Cao
  • Yuanzhou Chen
  • Yizhou Sun

Learning complex physical dynamics purely from data is challenging due to the intrinsic properties of systems to be satisfied. Incorporating physics-informed priors, such as in Hamiltonian Neural Networks (HNNs), achieves high-precision modeling for energy-conservative systems. However, real-world systems often deviate from strict energy conservation and follow different physical priors. To address this, we present a framework that achieves high-precision modeling for a wide range of dynamical systems from the numerical aspect, by enforcing Time-Reversal Symmetry (TRS) via a novel regularization term. It helps preserve energies for conservative systems while serving as a strong inductive bias for non-conservative, reversible systems. While TRS is a domain-specific physical prior, we present the first theoretical proof that TRS loss can universally improve modeling accuracy by minimizing higher-order Taylor terms in ODE integration, which is numerically beneficial to various systems regardless of their properties, even for irreversible systems. By integrating the TRS loss within neural ordinary differential equation models, the proposed model TREAT demonstrates superior performance on diverse physical systems. It achieves a significant 11. 5% MSE improvement in a challenging chaotic triple-pendulum scenario, underscoring TREAT’s broad applicability and effectiveness.

EAAI Journal 2024 Journal Article

Predicting water quality in municipal water management systems using a hybrid deep learning model

  • Wenxian Luo
  • Leijun Huang
  • Jiabin Shu
  • Hailin Feng
  • Wenjie Guo
  • Kai Xia
  • Kai Fang
  • Wei Wang

Increasing municipal waste generation puts more and more municipal water resources at high risk. Accurate prediction of water quality becomes critical for effective protection of the water resources. Due to the nonlinear and non-stationary characteristics of water quality data of the municipal water resources, it is challenging to achieve high prediction accuracy, especially for medium-term and long-term predictions. To address this issue, we propose a novel hybrid deep learning model to predict water quality multiple steps ahead. The proposed model adopts the encoder–decoder structure in the form of two long short-term memory (LSTM) networks, integrated with the attention mechanism and a convolutional neural network (CNN). The model extracts the complex correlation between multiple water quality features through the CNN, and uses the two LSTM networks to transfer historical information to predictions, with an attention layer assigning different weights to the different parts of the historical information. Using three years of water quality data collected from an urban river, we experimentally show that the proposed model outperforms the baseline models by 11%–34% in root mean squared error (RMSE) when predicting dissolved oxygen multiple steps ahead, and by 1%–7% when predicting total phosphorus. Similar improvement has also been found in Nash–Sutcliffeefficiency (NSE) and mean absolute error (MAE). The proposed model is a feasible solution for multi-step medium-term water quality prediction.

YNIMG Journal 2024 Journal Article

Prediction of anxious depression using multimodal neuroimaging and machine learning

  • Enqi Zhou
  • Wei Wang
  • Simeng Ma
  • Xinhui Xie
  • Lijun Kang
  • Shuxian Xu
  • Zipeng Deng
  • Qian Gong

Anxious depression is a common subtype of major depressive disorder (MDD) associated with adverse outcomes and severely impaired social function. It is important to clarify the underlying neurobiology of anxious depression to refine the diagnosis and stratify patients for therapy. Here we explored associations between anxiety and brain structure/function in MDD patients. A total of 260 MDD patients and 127 healthy controls underwent three-dimensional T1-weighted structural scanning and resting-state functional magnetic resonance imaging. Demographic data were collected from all participants. Differences in gray matter volume (GMV), (fractional) amplitude of low-frequency fluctuation ((f)ALFF), regional homogeneity (ReHo), and seed point-based functional connectivity were compared between anxious MDD patients, non-anxious MDD patients, and healthy controls. A random forest model was used to predict anxiety in MDD patients using neuroimaging features. Anxious MDD patients showed significant differences in GMV in the left middle temporal gyrus and ReHo in the right superior parietal gyrus and the left precuneus than HCs. Compared with non-anxious MDD patients, patients with anxious MDD showed significantly different GMV in the left inferior temporal gyrus, left superior temporal gyrus, left superior frontal gyrus (orbital part), and left dorsolateral superior frontal gyrus; fALFF in the left middle temporal gyrus; ReHo in the inferior temporal gyrus and the superior frontal gyrus (orbital part); and functional connectivity between the left superior temporal gyrus(temporal pole) and left medial superior frontal gyrus. A diagnostic predictive random forest model built using imaging features and validated by 10-fold cross-validation distinguished anxious from non-anxious MDD with an AUC of 0.802. Patients with anxious depression exhibit dysregulation of brain regions associated with emotion regulation, cognition, and decision-making, and our diagnostic model paves the way for more accurate, objective clinical diagnosis of anxious depression.

NeurIPS Conference 2024 Conference Paper

Proving Olympiad Algebraic Inequalities without Human Demonstrations

  • Chenrui Wei
  • Mengzhou Sun
  • Wei Wang

Solving Olympiad-level mathematical problems represents a significant advancement in machine intelligence and automated reasoning. Current machine learning methods, however, struggle to solve Olympiad-level problems beyond Euclidean plane geometry due to a lack of large-scale, high-quality datasets. The challenge is even greater in algebraic systems, which involve infinite reasoning spaces within finite conditions. To address these issues, we propose AIPS, an Algebraic Inequality Proving System capable of autonomously generating complex inequality theorems and effectively solving Olympiad-level inequality problems without requiring human demonstrations. During proof search in a mixed reasoning manner, a value curriculum learning strategy on generated datasets is implemented to improve proving performance, demonstrating strong mathematical intuitions. On a test set of 20 International Mathematical Olympiad-level inequality problems, AIPS successfully solved 10, outperforming state-of-the-art methods. Furthermore, AIPS automatically generated a vast array of non-trivial theorems without human intervention, some of which have been evaluated by professional contestants and deemed to reach the level of the International Mathematical Olympiad. Notably, one theorem was selected as a competition problem in a major city's 2024 Mathematical Olympiad. All the materials are available at sites. google. com/view/aips2

AAAI Conference 2024 Conference Paper

SelfPromer: Self-Prompt Dehazing Transformers with Depth-Consistency

  • Cong Wang
  • Jinshan Pan
  • Wanyu Lin
  • Jiangxin Dong
  • Wei Wang
  • Xiao-ming Wu

This work presents an effective depth-consistency Self-Prompt Transformer, terms as SelfPromer, for image dehazing. It is motivated by an observation that the estimated depths of an image with haze residuals and its clear counterpart vary. Enforcing the depth consistency of dehazed images with clear ones, therefore, is essential for dehazing. For this purpose, we develop a prompt based on the features of depth differences between the hazy input images and corresponding clear counterparts that can guide dehazing models for better restoration. Specifically, we first apply deep features extracted from the input images to the depth difference features for generating the prompt that contains the haze residual information in the input. Then we propose a prompt embedding module that is designed to perceive the haze residuals, by linearly adding the prompt to the deep features. Further, we develop an effective prompt attention module to pay more attention to haze residuals for better removal. By incorporating the prompt, prompt embedding, and prompt attention into an encoder-decoder network based on VQGAN, we can achieve better perception quality. As the depths of clear images are not available at inference, and the dehazed images with one-time feed-forward execution may still contain a portion of haze residuals, we propose a new continuous self-prompt inference that can iteratively correct the dehazing model towards better haze-free image generation. Extensive experiments show that our SelfPromer performs favorably against the state-of-the-art approaches on both synthetic and real-world datasets in terms of perception metrics including NIQE, PI, and PIQE. The source codes will be made available at https://github.com/supersupercong/SelfPromer.

AAAI Conference 2024 Conference Paper

STAR: Boosting Low-Resource Information Extraction by Structure-to-Text Data Generation with Large Language Models

  • Mingyu Derek Ma
  • Xiaoxuan Wang
  • Po-Nien Kung
  • P. Jeffrey Brantingham
  • Nanyun Peng
  • Wei Wang

Information extraction tasks such as event extraction require an in-depth understanding of the output structure and sub-task dependencies. They heavily rely on task-specific training data in the form of (passage, target structure) pairs to obtain reasonable performance. However, obtaining such data through human annotation is costly, leading to a pressing need for low-resource information extraction approaches that require minimal human labeling for real-world applications. Fine-tuning supervised models with synthesized training data would be a generalizable method, but the existing data generation methods either still rely on large-scale ground-truth data or cannot be applied to complicated IE tasks due to their poor performance. To address these challenges, we propose STAR, a data generation method that leverages Large Language Models (LLMs) to synthesize data instances given limited seed demonstrations, thereby boosting low-resource information extraction performance. Our approach involves generating target structures (Y) followed by generating passages (X), all accomplished with the aid of LLMs. We design fine-grained step-by-step instructions to obtain the initial data instances. We further reduce errors and improve data quality through self-reflection error identification and self-refinement with iterative revision. Our experiments show that the data generated by STAR significantly improve the performance of low-resource event extraction and relation extraction tasks, even surpassing the effectiveness of human-curated data. Human assessment of the data quality shows STAR-generated data exhibit higher passage quality and better align with the task definitions compared with the human-curated data.

NeurIPS Conference 2024 Conference Paper

Stealth edits to large language models

  • Oliver J. Sutton
  • Qinghua Zhou
  • Wei Wang
  • Desmond J. Higham
  • Alexander N. Gorban
  • Alexander Bastounis
  • Ivan Y. Tyukin

We reveal the theoretical foundations of techniques for editing large language models, and present new methods which can do so without requiring retraining. Our theoretical insights show that a single metric (a measure of the intrinsic dimension of the model's features) can be used to assess a model's editability and reveals its previously unrecognised susceptibility to malicious stealth attacks. This metric is fundamental to predicting the success of a variety of editing approaches, and reveals new bridges between disparate families of editing methods. We collectively refer to these as stealth editing methods, because they directly update a model's weights to specify its response to specific known hallucinating prompts without affecting other model behaviour. By carefully applying our theoretical insights, we are able to introduce a new jet-pack network block which is optimised for highly selective model editing, uses only standard network operations, and can be inserted into existing networks. We also reveal the vulnerability of language models to stealth attacks: a small change to a model's weights which fixes its response to a single attacker-chosen prompt. Stealth attacks are computationally simple, do not require access to or knowledge of the model's training data, and therefore represent a potent yet previously unrecognised threat to redistributed foundation models. Extensive experimental results illustrate and support our methods and their theoretical underpinnings. Demos and source code are available at https: //github. com/qinghua-zhou/stealth-edits.

NeurIPS Conference 2024 Conference Paper

The Limits of Differential Privacy in Online Learning

  • Bo Li
  • Wei Wang
  • Peng Ye

Differential privacy (DP) is a formal notion that restricts the privacy leakage of an algorithm when running on sensitive data, in which privacy-utility trade-off is one of the central problems in private data analysis. In this work, we investigate the fundamental limits of differential privacy in online learning algorithms and present evidence that separates three types of constraints: no DP, pure DP, and approximate DP. We first describe a hypothesis class that is online learnable under approximate DP but not online learnable under pure DP under the adaptive adversarial setting. This indicates that approximate DP must be adopted when dealing with adaptive adversaries. We then prove that any private online learner must make an infinite number of mistakes for almost all hypothesis classes. This essentially generalizes previous results and shows a strong separation between private and non-private settings since a finite mistake bound is always attainable (as long as the class is online learnable) when there is no privacy requirement.

JBHI Journal 2024 Journal Article

Towards Wearable and Portable Spine Motion Analysis Through Dynamic Optimization of Smartphone Videos and IMU Data

  • Wei Wang
  • Yinghu Peng
  • Yilun Sun
  • Jun Wang
  • Guanglin Li

Background: Monitoring spine kinematics is crucial for applications like disease evaluation and ergonomics analysis. However, the small scale of vertebrae and the number of degrees of freedom present significant challenges for noninvasive and convenient spine kinematics estimation. Methods: This study developed a dynamic optimization framework for wearable spine motion tracking at the intervertebral joint level by integrating smartphone videos and Inertia Measurement Units (IMUs) with dynamic constraints from a thoracolumbar spine model. Validation involved motion data from 10 healthy males performing static standing, dynamic upright trunk rotations, and gait. This data included rotations of ten IMUs on vertebrae and virtual landmarks from three smartphone videos preprocessed by OpenCap, an application leveraging computer vision for pose estimation. The kinematic measures derived from the optimized solution were compared against simultaneously collected infrared optical marker-based measurements and in vivo literature data. Solutions only based on IMUs or videos were also compared for accuracy evaluation. Results: The proposed optimization approach closely matched the reference data in the intervertebral or segmental rotation range, demonstrating minimal angular differences across all motions and the highest correlation in 3D rotations (maximal Pearson and intraclass correlation coefficients of 0. 92 and 0. 94, respectively). Time-series changes of joint angles also aligned well with the optical-marker reference. Conclusion: Dynamic optimization of the spine simulation that integrates IMUs and computer vision outperforms the single-modality method. Significance: This markerless 3D spine motion capture method holds potential for spinal health assessment in large cohorts in real-world settings without dedicated laboratories.

AAAI Conference 2024 Conference Paper

When to Grow? A Fitting Risk-Aware Policy for Layer Growing in Deep Neural Networks

  • Haihang Wu
  • Wei Wang
  • Tamasha Malepathirana
  • Damith Senanayake
  • Denny Oetomo
  • Saman Halgamuge

Neural growth is the process of growing a small neural network to a large network and has been utilized to accelerate the training of deep neural networks. One crucial aspect of neural growth is determining the optimal growth timing. However, few studies investigate this systematically. Our study reveals that neural growth inherently exhibits a regularization effect, whose intensity is influenced by the chosen policy for growth timing. While this regularization effect may mitigate the overfitting risk of the model, it may lead to a notable accuracy drop when the model underfits. Yet, current approaches have not addressed this issue due to their lack of consideration of the regularization effect from neural growth. Motivated by these findings, we propose an under/over fitting risk-aware growth timing policy, which automatically adjusts the growth timing informed by the level of potential under/overfitting risks to address both risks. Comprehensive experiments conducted using CIFAR-10/100 and ImageNet datasets show that the proposed policy achieves accuracy improvements of up to 1.3% in models prone to underfitting while achieving similar accuracies in models suffering from overfitting compared to the existing methods.

AAAI Conference 2023 Conference Paper

A Generative Approach for Script Event Prediction via Contrastive Fine-Tuning

  • Fangqi Zhu
  • Jun Gao
  • Changlong Yu
  • Wei Wang
  • Chen Xu
  • Xin Mu
  • Min Yang
  • Ruifeng Xu

Script event prediction aims to predict the subsequent event given the context. This requires the capability to infer the correlations between events. Recent works have attempted to improve event correlation reasoning by using pretrained language models and incorporating external knowledge (e.g., discourse relations). Though promising results have been achieved, some challenges still remain. First, the pretrained language models adopted by current works ignore event-level knowledge, resulting in an inability to capture the correlations between events well. Second, modeling correlations between events with discourse relations is limited because it can only capture explicit correlations between events with discourse markers, and cannot capture many implicit correlations. To this end, we propose a novel generative approach for this task, in which a pretrained language model is fine-tuned with an event-centric pretraining objective and predicts the next event within a generative paradigm. Specifically, we first introduce a novel event-level blank infilling strategy as the learning objective to inject event-level knowledge into the pretrained language model, and then design a likelihood-based contrastive loss for fine-tuning the generative model. Instead of using an additional prediction layer, we perform prediction by using sequence likelihoods generated by the generative model. Our approach models correlations between events in a soft way without any external knowledge. The likelihood-based prediction eliminates the need to use additional networks to make predictions and is somewhat interpretable since it scores each word in the event. Experimental results on the multi-choice narrative cloze (MCNC) task demonstrate that our approach achieves better results than other state-of-the-art baselines. Our code will be available at https://github.com/zhufq00/mcnc.

ICRA Conference 2023 Conference Paper

A Moving Target Tracking System of Quadrotors with Visual-Inertial Localization

  • Ziyue Lin
  • Wenbo Xu
  • Wei Wang

This paper implements a vision-based moving target tracking system of quadrotors with visual-inertial localization in GNSS-denied indoor environments. We use the visual-inertial odometry to estimate the states of the UAV by minimizing visual and inertial residuals, and estimate the states of the target with extended Kalman Filter from visual detection. This research formulates the target tracking problem as optimization-based trajectory generation where a weighted sum cost function jointly penalizes the tracking error, the control cost of the trajectory and the trajectory length, while enforcing the safety and feasibility constraints. We present a strategy that represents the trajectory as piecewise Bézier curves using Bernstein polynomial basis. Due to the special properties of Bézier curves, the position of the entire trajectory and its derivatives can be directly bounded within the safe spaces, thus this facilitating the dynamics of the quadrotor. The proposed strategy can generate smooth and collision-free tracking trajectories and is time and space efficient. We conduct simulations and real-world experiments to validate the effectiveness of our system.

NeurIPS Conference 2023 Conference Paper

Binary Classification with Confidence Difference

  • Wei Wang
  • Lei Feng
  • Yuchen Jiang
  • Gang Niu
  • Min-Ling Zhang
  • Masashi Sugiyama

Recently, learning with soft labels has been shown to achieve better performance than learning with hard labels in terms of model generalization, calibration, and robustness. However, collecting pointwise labeling confidence for all training examples can be challenging and time-consuming in real-world scenarios. This paper delves into a novel weakly supervised binary classification problem called confidence-difference (ConfDiff) classification. Instead of pointwise labeling confidence, we are given only unlabeled data pairs with confidence difference that specifies the difference in the probabilities of being positive. We propose a risk-consistent approach to tackle this problem and show that the estimation error bound achieves the optimal convergence rate. We also introduce a risk correction approach to mitigate overfitting problems, whose consistency and convergence rate are also proven. Extensive experiments on benchmark data sets and a real-world recommender system data set validate the effectiveness of our proposed approaches in exploiting the supervision information of the confidence difference.

IJCAI Conference 2023 Conference Paper

Boosting Decision-Based Black-Box Adversarial Attack with Gradient Priors

  • Han Liu
  • Xingshuo Huang
  • Xiaotong Zhang
  • Qimai Li
  • Fenglong Ma
  • Wei Wang
  • Hongyang Chen
  • Hong Yu

Decision-based methods have shown to be effective in black-box adversarial attacks, as they can obtain satisfactory performance and only require to access the final model prediction. Gradient estimation is a critical step in black-box adversarial attacks, as it will directly affect the query efficiency. Recent works have attempted to utilize gradient priors to facilitate score-based methods to obtain better results. However, these gradient priors still suffer from the edge gradient discrepancy issue and the successive iteration gradient direction issue, thus are difficult to simply extend to decision-based methods. In this paper, we propose a novel Decision-based Black-box Attack framework with Gradient Priors (DBA-GP), which seamlessly integrates the data-dependent gradient prior and time-dependent prior into the gradient estimation procedure. First, by leveraging the joint bilateral filter to deal with each random perturbation, DBA-GP can guarantee that the generated perturbations in edge locations are hardly smoothed, i. e. , alleviating the edge gradient discrepancy, thus remaining the characteristics of the original image as much as possible. Second, by utilizing a new gradient updating strategy to automatically adjust the successive iteration gradient direction, DBA-GP can accelerate the convergence speed, thus improving the query efficiency. Extensive experiments have demonstrated that the proposed method outperforms other strong baselines significantly.

JBHI Journal 2023 Journal Article

Cardiac LGE MRI Segmentation With Cross-Modality Image Augmentation and Improved U-Net

  • Xinhua Yu
  • Junxin Chen
  • Bo Fang
  • Wei Wang
  • Li-bo Zhang
  • Zhihan Lv

Image segmentation is a challenging problem in imaging informatics, which stems from the intersection of imaging techniques, computer science and biomedicine. In particular, accurate segmentation of cardiac structures in late gadolinium enhancement (LGE) cardiac magnetic resonance (CMR) is of great clinical importance for cardiac function assessment and myocardial disease diagnosis. However, it is a well-known challenge due to its special imaging modality and the lack of labeled LGE samples. In this paper, we propose an unsupervised ventricular segmentation algorithm that can perform biventricular segmentation of LGE images in the absence of labeled LGE data. There are two primary modules, the data augmentation procedure and the segmentation network. The easily available annotated balanced-Steady State Free Precession (bSSFP) images are employed for cross-modal data augmentation by image translation, where a single bSSFP image is converted into multiple synthetic LGE images while preserving the original morphological structure. Then, the proposed segmentation network is trained with the synthetic LGE images and used for segmenting real LGE images. Validation experiments demonstrated the effectiveness and advantages of the proposed algorithm.

ICRA Conference 2023 Conference Paper

CAROM Air - Vehicle Localization and Traffic Scene Reconstruction from Aerial Videos

  • Duo Lu
  • Eric Eaton
  • Matt Weg
  • Wei Wang
  • Steven Como
  • Jeffrey Wishart
  • Hongbin Yu
  • Yezhou Yang

Road traffic scene reconstruction from videos has been desirable by road safety regulators, city planners, researchers, and autonomous driving technology developers. However, it is expensive and unnecessary to cover every mile of the road with cameras mounted on the road infrastructure. This paper presents a method that can process aerial videos to vehicle trajectory data so that a traffic scene can be automatically reconstructed and accurately re-simulated using computers. On average, the vehicle localization error is about 0. 1 m to 0. 3 m using a consumer-grade drone flying at 120 meters. This project also compiles a dataset of 50 reconstructed road traffic scenes from about 100 hours of aerial videos to enable various downstream traffic analysis applications and facilitate further road traffic related research. The dataset is available at https://github.com/duolu/CAROM.

AAAI Conference 2023 Conference Paper

Competition or Cooperation? Exploring Unlabeled Data via Challenging Minimax Game for Semi-supervised Relation Extraction

  • Yu Hong
  • Jiahang Li
  • Jianchuan Feng
  • Chenghua Huang
  • Zhixu Li
  • Jianfeng Qu
  • Yanghua Xiao
  • Wei Wang

Semi-Supervised Relation Extraction aims at learning well-performed RE models with limited labeled and large-scale unlabeled data. Existing methods mainly suffer from semantic drift and insufficient supervision, which severely limit the performance. To address these problems, recent work tends to design dual modules to work cooperatively for mutual enhancement. However, the consensus of two modules greatly restricts the model from exploring diverse relation expressions in unlabeled set, which hinders the performance as well as model generalization. To tackle this problem, in this paper, we propose a novel competition-based method AdvSRE. We set up a challenging minimax game on unlabeled data between two modules, Generator and Discriminator, and assign them with conflicting objectives. During the competition game, one module may find any possible chance to beat the other, which develops two modules' abilities until relation expressions cannot be further explored. To exploit label information, Discriminator is further asked to predict specific relation for each sentence. Experiment results on two benchmarks show new state-of-the-art performance over baselines, demonstrating the effectiveness of proposed AdvSRE.

JBHI Journal 2023 Journal Article

Dual-Channel Neural Network for Atrial Fibrillation Detection From a Single Lead ECG Wave

  • Bo Fang
  • Junxin Chen
  • Yu Liu
  • Wei Wang
  • Ke Wang
  • Amit Kumar Singh
  • Zhihan Lv

With the dramatic progress of wearable devices, continuous collection of single lead ECG wave is able to be implemented in a comfortable fashion. Data mining on single lead ECG wave is therefore attracting increasing attention, where atrial fibrillation (AF) detection is a hot topic. In this paper, we propose a dual-channel neural network for AF detection from a single lead ECG wave. Two primary phases are included, the data preprocessing part followed by a dual-channel neural network. A two-stage denoising procedure is developed for data preprocessing, so as to tackle the high noise and disturbance which generally resides in the ECG wave collected by wearable devices. Then the time-frequency spectrum and Poincare plot of the denoised ECG signal are imported into the developed dual-channel neural network for feature extraction and AF detection. On the 2017 PhysioNet/CinC Challenge database, the F1 values were 0. 83, 0. 90, and 0. 75 for AF rhythm and normal rhythm, and other rhythm, respectively. The results well validate the effectiveness of the proposed method for AF detection from a single lead ECG wave, and also indicate its performance advantages over some state-of-the-art counterparts.

AAAI Conference 2023 Conference Paper

Foresee What You Will Learn: Data Augmentation for Domain Generalization in Non-stationary Environment

  • Qiuhao Zeng
  • Wei Wang
  • Fan Zhou
  • Charles Ling
  • Boyu Wang

Existing domain generalization aims to learn a generalizable model to perform well even on unseen domains. For many real-world machine learning applications, the data distribution often shifts gradually along domain indices. For example, a self-driving car with a vision system drives from dawn to dusk, with the sky gradually darkening. Therefore, the system must be able to adapt to changes in ambient illuminations and continue to drive safely on the road. In this paper, we formulate such problems as Evolving Domain Generalization, where a model aims to generalize well on a target domain by discovering and leveraging the evolving pattern of the environment. We then propose Directional Domain Augmentation (DDA), which simulates the unseen target features by mapping source data as augmentations through a domain transformer. Specifically, we formulate DDA as a bi-level optimization problem and solve it through a novel meta-learning approach in the representation space. We evaluate the proposed method on both synthetic datasets and real-world datasets, and empirical results show that our approach can outperform other existing methods.

IJCAI Conference 2023 Conference Paper

From Generation to Suppression: Towards Effective Irregular Glow Removal for Nighttime Visibility Enhancement

  • Wanyu Wu
  • Wei Wang
  • Zheng Wang
  • Kui Jiang
  • Xin Xu

Most existing Low-Light Image Enhancement (LLIE) methods are primarily designed to improve brightness in dark regions, which suffer from severe degradation in nighttime images. However, these methods have limited exploration in another major visibility damage, the glow effects in real night scenes. Glow effects are inevitable in the presence of artificial light sources and cause further diffused blurring when directly enhanced. To settle this issue, we innovatively consider the glow suppression task as learning physical glow generation via multiple scattering estimation according to the Atmospheric Point Spread Function (APSF). In response to the challenges posed by uneven glow intensity and varying source shapes, an APSF-based Nighttime Imaging Model with Near-field Light Sources (NIM-NLS) is specifically derived to design a scalable Light-aware Blind Deconvolution Network (LBDN). The glow-suppressed result is then brightened via a Retinex-based Enhancement Module (REM). Remarkably, the proposed glow suppression method is based on zero-shot learning and does not rely on any paired or unpaired training data. Empirical evaluations demonstrate the effectiveness of the proposed method in both glow suppression and low-light enhancement tasks.

ICRA Conference 2023 Conference Paper

Global Localization in Repetitive and Ambiguous Environments

  • Zhenyu Wu 0001
  • Wei Wang
  • Jun Zhang 0042
  • Qiyang Lyu
  • Haoyuan Zhang
  • Danwei Wang

Accurate global localization is an essential ingredient for autonomous mobile robots (AMRs) operating in enclosed or partially enclosed repetitive environments (e. g. , office corridors, industrial warehouses, transportation centers). In such environments, the Global Navigation Satellite System (GNSS) signals are unreliable or severely degraded. The highly ambiguous structures in such challenging scenarios would also lead the ordinary geometric feature-based LiDAR/visual localization methods to fail. The ambient magnetic field (MF) has exhibited high distinctiveness at different location, which makes it a viable alternative for infrastructure-free AMR localization. However, few of the previous research has been focused on the orientation-dependency and similar-sequential-route limitations of MF-based localization. Thus, this paper proposes a novel probabilistic global localization system with 2-D LiDAR and rotation-invariant magnetic field for AMRs operating in challenging repetitive and ambiguous environments. The proposed localization system mainly consists of: 1) Two-step Initialization: laser distance and MF sequence based matching, and 2) MF-based Pose Tracking: recursive multi-dimensional MF sequence based matching. Extensive experimental results demonstrate the advantageous localization performances of the proposed localization system over the existing methods.

IJCAI Conference 2023 Conference Paper

Graph-based Molecular Representation Learning

  • Zhichun Guo
  • Kehan Guo
  • Bozhao Nan
  • Yijun Tian
  • Roshni G. Iyer
  • Yihong Ma
  • Olaf Wiest
  • Xiangliang Zhang

Molecular representation learning (MRL) is a key step to build the connection between machine learning and chemical science. In particular, it encodes molecules as numerical vectors preserving the molecular structures and features, on top of which the downstream tasks (e. g. , property prediction) can be performed. Recently, MRL has achieved considerable progress, especially in methods based on deep molecular graph learning. In this survey, we systematically review these graph-based molecular representation techniques, especially the methods incorporating chemical domain knowledge. Specifically, we first introduce the features of 2D and 3D molecular graphs. Then we summarize and categorize MRL methods into three groups based on their input. Furthermore, we discuss some typical chemical applications supported by MRL. To facilitate studies in this fast-developing area, we also list the benchmarks and commonly used datasets in the paper. Finally, we share our thoughts on future research directions.

EAAI Journal 2023 Journal Article

Image segmentation of adhesive ores based on MSBA-Unet and convex-hull defect detection

  • Wei Wang
  • Qing Li
  • Dezheng Zhang
  • Jiawei Fu

Ore particle size information is a crucial indicator to evaluate the crushing quality and judge whether there are oversized ores on the conveyor belt. Accurately separating each ore is a critical prerequisite for obtaining high-precision particle size measurement (PSM) results. However, the large size variance and natural adhesion between ores pose a huge challenge to this task, imposing under-segmentation. Hence, this study proposes an automatic method that combines semantic segmentation and morphological operations to measure the ore particle size. Specifically, a novel multi-scale connection and boundary-aware U-Net model (MSBA-Unet) that classifies boundary pixels between adhesive ores more accurately is developed to segment ore images. Second, the convex-hull defect detection (CDD) method that divides the adhesive ores with a deep concave shape into two pieces is adopted to process the predicted masks further. The experimental results demonstrate that the MSBA-Unet architecture design and the CDD method can significantly improve the performance of separating adhesive ores of different sizes. Therefore, the under-segmentation problem is tremendously alleviated, and the ore PSM results agree well with the ground truth.

IJCAI Conference 2023 Conference Paper

Label Specific Multi-Semantics Metric Learning for Multi-Label Classification: Global Consideration Helps

  • Jun-Xiang Mao
  • Wei Wang
  • Min-Ling Zhang

In multi-label classification, it is critical to capitalize on complicated data structures and semantic relationships. Metric learning serves as an effective strategy to provide a better measurement of distances between examples. Existing works on metric learning for multi-label classification mainly learn one single global metric that characterizes latent semantic similarity between multi-label instances. However, such single-semantics metric exploitation approaches can not capture the intrinsic properties of multi-label data possessed of rich semantics. In this paper, the first attempt towards multi-semantics metric learning for multi-label classification is investigated. Specifically, the proposed LIMIC approach simultaneously learns one global and multiple label-specific local metrics by exploiting label-specific side information. The global metric is learned to capture the commonality across all the labels and label-specific local metrics characterize the individuality of each semantic space. The combination of global metric and label-specific local metrics is utilized to construct latent semantic space for each label, in which similar intra-class instances are pushed closer and inter-class instances are pulled apart. Furthermore, metric-based label correlation regularization is constructed to maintain similarity between correlated label spaces. Extensive experiments on benchmark multi-label data sets validate the superiority of our proposed approach in learning effective distance metrics for multi-label classification.

JBHI Journal 2023 Journal Article

Multi-Level Adversarial Spatio-Temporal Learning for Footstep Pressure Based FoG Detection

  • Kun Hu
  • Shaohui Mei
  • Wei Wang
  • Kaylena A. Ehgoetz Martens
  • Liang Wang
  • Simon J. G. Lewis
  • David D. Feng
  • Zhiyong Wang

Freezing of gait (FoG) is one of the most common symptoms of Parkinson's disease, which is a neurodegenerative disorder of the central nervous system impacting millions of people around the world. To address the pressing need to improve the quality of treatment for FoG, devising a computer-aided detection and quantification tool for FoG has been increasingly important. As a non-invasive technique for collecting motion patterns, the footstep pressure sequences obtained from pressure sensitive gait mats provide a great opportunity for evaluating FoG in the clinic and potentially in the home environment. In this study, FoG detection is formulated as a sequential modelling task and a novel deep learning architecture, namely Adversarial Spatio-temporal Network (ASTN), is proposed to learn FoG patterns across multiple levels. ASTN introduces a novel adversarial training scheme with a multi-level subject discriminator to obtain subject-independent FoG representations, which helps to reduce the over-fitting risk due to the high inter-subject variance. As a result, robust FoG detection can be achieved for unseen subjects. The proposed scheme also sheds light on improving subject-level clinical studies from other scenarios as it can be integrated with many existing deep architectures. To the best of our knowledge, this is one of the first studies of footstep pressure-based FoG detection and the approach of utilizing ASTN is the first deep neural network architecture in pursuit of subject-independent representations. In our experiments on 393 trials collected from 21 subjects, the proposed ASTN achieved an AUC 0. 85, clearly outperforming conventional learning methods.

NeurIPS Conference 2023 Conference Paper

Penguin: Parallel-Packed Homomorphic Encryption for Fast Graph Convolutional Network Inference

  • Ran Ran
  • Nuo Xu
  • Tao Liu
  • Wei Wang
  • Gang Quan
  • Wujie Wen

The marriage of Graph Convolutional Network (GCN) and Homomorphic Encryption (HE) enables the inference of graph data on the cloud with significantly enhanced client data privacy. However, the tremendous computation and memory overhead associated with HE operations challenges the practicality of HE-based GCN inference. GCN inference involves a sequence of expensive matrix-matrix multiplications, and we observe that directly applying the state-of-the-art HE-based secure matrix-matrix multiplication solutions to accelerate HE-GCN inference is far less efficient as it does not exploit the unique aggregation mechanism of two-dimension graph node-features in GCN layer computation. As a result, in this paper, we propose a novel HE-based ciphertext packing technique, i. e. , Penguin, that can take advantage of the unique computation pattern during the HE-GCN inference to significantly reduce the computation and memory overhead associated with HE operations. Specifically, Penguin employs (i) an effective two-dimension parallel packing technique for feature ciphertext with optimal graph node partitioning and graph feature interleaving, and (ii) an interleaved assembly technique that can effectively make use of the blank slots to merge ciphertexts after feature reduction and significantly reduce the costly rotation operation. We provide theoretical analysis and experimental validation to demonstrate the speedup achieved by Penguin in accelerating GCN inference using popular GCN models and datasets. Our results show that Penguin can achieve up to $\sim10\times$ speedup and around $\sim79$% reduction in computational memory overhead, significantly outperforming state-of-the-art solutions. To the best of our knowledge, this is the first work that can ensure the protection of both graph structure and features when accelerating HE-GCN inference on encrypted data. Our code is publicly available at https: //github. com/ranran0523/Penguin.

AAAI Conference 2023 Conference Paper

Pixel Is All You Need: Adversarial Trajectory-Ensemble Active Learning for Salient Object Detection

  • Zhenyu Wu
  • Lin Wang
  • Wei Wang
  • Qing Xia
  • Chenglizhao Chen
  • Aimin Hao
  • Shuo Li

Although weakly-supervised techniques can reduce the labeling effort, it is unclear whether a saliency model trained with weakly-supervised data (e.g., point annotation) can achieve the equivalent performance of its fully-supervised version. This paper attempts to answer this unexplored question by proving a hypothesis: there is a point-labeled dataset where saliency models trained on it can achieve equivalent performance when trained on the densely annotated dataset. To prove this conjecture, we proposed a novel yet effective adversarial trajectory-ensemble active learning (ATAL). Our contributions are three-fold: 1) Our proposed adversarial attack triggering uncertainty can conquer the overconfidence of existing active learning methods and accurately locate these uncertain pixels. 2) Our proposed trajectory-ensemble uncertainty estimation method maintains the advantages of the ensemble networks while significantly reducing the computational cost. 3) Our proposed relationship-aware diversity sampling algorithm can conquer oversampling while boosting performance. Experimental results show that our ATAL can find such a point-labeled dataset, where a saliency model trained on it obtained 97%-99% performance of its fully-supervised version with only 10 annotated points per image.

AAAI Conference 2023 Conference Paper

Poisoning with Cerberus: Stealthy and Colluded Backdoor Attack against Federated Learning

  • Xiaoting Lyu
  • Yufei Han
  • Wei Wang
  • Jingkai Liu
  • Bin Wang
  • Jiqiang Liu
  • Xiangliang Zhang

Are Federated Learning (FL) systems free from backdoor poisoning with the arsenal of various defense strategies deployed? This is an intriguing problem with significant practical implications regarding the utility of FL services. Despite the recent flourish of poisoning-resilient FL methods, our study shows that carefully tuning the collusion between malicious participants can minimize the trigger-induced bias of the poisoned local model from the poison-free one, which plays the key role in delivering stealthy backdoor attacks and circumventing a wide spectrum of state-of-the-art defense methods in FL. In our work, we instantiate the attack strategy by proposing a distributed backdoor attack method, namely Cerberus Poisoning (CerP). It jointly tunes the backdoor trigger and controls the poisoned model changes on each malicious participant to achieve a stealthy yet successful backdoor attack against a wide spectrum of defensive mechanisms of federated learning techniques. Our extensive study on 3 large-scale benchmark datasets and 13 mainstream defensive mechanisms confirms that Cerberus Poisoning raises a significantly severe threat to the integrity and security of federated learning practices, regardless of the flourish of robust Federated Learning methods.

NeurIPS Conference 2023 Conference Paper

PromptRestorer: A Prompting Image Restoration Method with Degradation Perception

  • Cong Wang
  • Jinshan Pan
  • Wei Wang
  • Jiangxin Dong
  • Mengzhu Wang
  • Yakun Ju
  • Junyang Chen

We show that raw degradation features can effectively guide deep restoration models, providing accurate degradation priors to facilitate better restoration. While networks that do not consider them for restoration forget gradually degradation during the learning process, model capacity is severely hindered. To address this, we propose a Prompting image Restorer, termed as PromptRestorer. Specifically, PromptRestorer contains two branches: a restoration branch and a prompting branch. The former is used to restore images, while the latter perceives degradation priors to prompt the restoration branch with reliable perceived content to guide the restoration process for better recovery. To better perceive the degradation which is extracted by a pre-trained model from given degradation observations, we propose a prompting degradation perception modulator, which adequately considers the characters of the self-attention mechanism and pixel-wise modulation, to better perceive the degradation priors from global and local perspectives. To control the propagation of the perceived content for the restoration branch, we propose gated degradation perception propagation, enabling the restoration branch to adaptively learn more useful features for better recovery. Extensive experimental results show that our PromptRestorer achieves state-of-the-art results on 4 image restoration tasks, including image deraining, deblurring, dehazing, and desnowing.

I&C Journal 2023 Journal Article

Randomness below complete theories of arithmetic

  • George Barmpalias
  • Wei Wang

We show that reals z which compute complete extensions of arithmetic have the random join property: for each random x < T z there exists random y < T z such that z ≡ T x ⊕ y. The same is true for the truth-table and the weak truth-table reducibilities.

NeurIPS Conference 2023 Conference Paper

RRHF: Rank Responses to Align Language Models with Human Feedback

  • Hongyi Yuan
  • Zheng Yuan
  • Chuanqi Tan
  • Wei Wang
  • Songfang Huang
  • Fei Huang

Reinforcement Learning from Human Feedback (RLHF) facilitates the alignment of large language models with human preferences, significantly enhancing the quality of interactions between humans and models. InstructGPT implements RLHF through several stages, including Supervised Fine-Tuning (SFT), reward model training, and Proximal Policy Optimization (PPO). However, PPO is sensitive to hyperparameters and requires multiple models in its standard implementation, making it hard to train and scale up to larger parameter counts. In contrast, we propose a novel learning paradigm called RRHF, which scores sampled responses from different sources via a logarithm of conditional probabilities and learns to align these probabilities with human preferences through ranking loss. RRHF can leverage sampled responses from various sources including the model responses from itself, other large language model responses, and human expert responses to learn to rank them. RRHF only needs 1 to 2 models during tuning and can efficiently align language models with human preferences robustly without complex hyperparameter tuning. Additionally, RRHF can be considered an extension of SFT and reward model training while being simpler than PPO in terms of coding, model counts, and hyperparameters. We evaluate RRHF on the Helpful and Harmless dataset, demonstrating comparable alignment performance with PPO by reward model score and human labeling. Extensive experiments show that the performance of RRHF is highly related to sampling quality which suggests RRHF is a best-of-$n$ learner.

EAAI Journal 2023 Journal Article

Screening of retired batteries with gramian angular difference fields and ConvNeXt

  • Mingqiang Lin
  • Jian Wu
  • Jinhao Meng
  • Wei Wang
  • Ji Wu

With the rapid development of electric vehicles, the second usage of retired batteries becomes a key issue. The accuracy of existing screening methods for retired batteries is highly dependent on the feature selection from charging or discharging curves. This paper proposes a novel method of screening retired batteries, in which the constant current (CC) charging curves are converted into images by Gramian angular difference fields (GADF) and classified with a ConvNeXt network. Firstly, the CC charging voltage data is reasonably reduced by piecewise aggregation approximation. Secondly, the CC voltage curves are encoded into images by GADF to make small differences more distinguishable. Then, a ConvNeXt network is used for screening the retired batteries because of its excellent performance on accuracy and scalability. Finally, validation experiments are carried out on 143 retired high-power lithium-ion batteries, and the results show that the proposed screening method has a classification detection accuracy of 93. 71%.

NeurIPS Conference 2023 Conference Paper

Semi-Supervised Domain Generalization with Known and Unknown Classes

  • Lei Zhang
  • Ji-Fu Li
  • Wei Wang

Semi-Supervised Domain Generalization (SSDG) aims to learn a model that is generalizable to an unseen target domain with only a few labels, and most existing SSDG methods assume that unlabeled training and testing samples are all known classes. However, a more realistic scenario is that known classes may be mixed with some unknown classes in unlabeled training and testing data. To deal with such a scenario, we propose the Class-Wise Adaptive Exploration and Exploitation (CWAEE) method. In particular, we explore unlabeled training data by using one-vs-rest classifiers and class-wise adaptive thresholds to detect known and unknown classes, and exploit them by adopting consistency regularization on augmented samples based on Fourier Transformation to improve the unseen domain generalization. The experiments conducted on real-world datasets verify the effectiveness and superiority of our method.

ICML Conference 2023 Conference Paper

SpENCNN: Orchestrating Encoding and Sparsity for Fast Homomorphically Encrypted Neural Network Inference

  • Ran Ran
  • Xinwei Luo
  • Wei Wang
  • Tao Liu 0023
  • Gang Quan
  • Xiaolin Xu 0001
  • Caiwen Ding
  • Wujie Wen

Homomorphic Encryption (HE) is a promising technology to protect clients’ data privacy for Machine Learning as a Service (MLaaS) on public clouds. However, HE operations can be orders of magnitude slower than their counterparts for plaintexts and thus result in prohibitively high inference latency, seriously hindering the practicality of HE. In this paper, we propose a HE-based fast neural network (NN) inference framework–SpENCNN built upon the co-design of HE operation-aware model sparsity and the single-instruction-multiple-data (SIMD)-friendly data packing, to improve NN inference latency. In particular, we first develop an encryption-aware HE-group convolution technique that can partition channels among different groups based on the data size and ciphertext size, and then encode them into the same ciphertext by novel group-interleaved encoding, so as to dramatically reduce the number of bottlenecked operations in HE convolution. We further tailor a HE-friendly sub-block weight pruning to reduce the costly HE-based convolution operation. Our experiments show that SpENCNN can achieve overall speedups of 8. 37$\times$, 12. 11$\times$, 19. 26$\times$, and 1. 87$\times$ for LeNet, VGG-5, HEFNet, and ResNet-20 respectively, with negligible accuracy loss. Our code is publicly available at https: //github. com/ranran0523/SPECNN.

JBHI Journal 2023 Journal Article

Syn_SegNet: A Joint Deep Neural Network for Ultrahigh-Field 7T MRI Synthesis and Hippocampal Subfield Segmentation in Routine 3T MRI

  • Xinwei Li
  • Linjin Wang
  • Hong Liu
  • Baoqiang Ma
  • Lei Chu
  • Xiaoxi Dong
  • Debin Zeng
  • Tongtong Che

Precise delineation of hippocampus subfields is crucial for the identification and management of various neurological and psychiatric disorders. However, segmenting these subfields automatically in routine 3T MRI is challenging due to their complex morphology and small size, as well as the limited signal contrast and resolution of the 3T images. This research proposes Syn_SegNet, an end-to-end, multitask joint deep neural network that leverages ultrahigh-field 7T MRI synthesis to improve hippocampal subfield segmentation in 3T MRI. Our approach involves two key components. First, we employ a modified Pix2PixGAN as the synthesis model, incorporating self-attention modules, image and feature matching loss, and ROI loss to generate high-quality 7T-like MRI around the hippocampal region. Second, we utilize a variant of 3D-U-Net with multiscale deep supervision as the segmentation subnetwork, incorporating an anatomic weighted cross-entropy loss that capitalizes on prior anatomical knowledge. We evaluate our method on hippocampal subfield segmentation in paired 3T MRI and 7T MRI with seven different anatomical structures. The experimental findings demonstrate that Syn_SegNet's segmentation performance benefits from integrating synthetic 7T data in an online manner and is superior to competing methods. Furthermore, we assess the generalizability of the proposed approach using a publicly accessible 3T MRI dataset. The developed method would be an efficient tool for segmenting hippocampal subfields in routine clinical 3T MRI.

JBHI Journal 2023 Journal Article

Trajectory-Aware Adaptive Imaging Clue Analysis for Guidewire Artifact Removal in Intravascular Optical Coherence Tomography

  • Gongning Luo
  • Xinghua Ma
  • Jinwen Guo
  • Mingye Zou
  • Wei Wang
  • Yang Cao
  • Kuanquan Wang
  • Shuo Li

Guidewire Artifact Removal (GAR) involves restoring missing imaging signals in areas of IntraVascular Optical Coherence Tomography (IVOCT) videos affected by guidewire artifacts. GAR helps overcome imaging defects and minimizes the impact of missing signals on the diagnosis of CardioVascular Diseases (CVDs). To restore the actual vascular and lesion information within the artifact area, we propose a reliable Trajectory-aware Adaptive imaging Clue analysis Network (TAC-Net) that includes two innovative designs: (i) Adaptive clue aggregation, which considers both texture-focused original (ORI) videos and structure-focused relative total variation (RTV) videos, and suppresses texture-structure imbalance with an active weight-adaptation mechanism; (ii) Trajectory-aware Transformer, which uses a novel attention calculation to perceive the attention distribution of artifact trajectories and avoid the interference of irregular and non-uniform artifacts. We provide a detailed formulation for the procedure and evaluation of the GAR task and conduct comprehensive quantitative and qualitative experiments. The experimental results demonstrate that TAC-Net reliably restores the texture and structure of guidewire artifact areas as expected by experienced physicians ( e. g. , SSIM: 97. 23%). We also discuss the value and potential of the GAR task for clinical applications and computer-aided diagnosis of CVDs.

NeurIPS Conference 2023 Conference Paper

Universality and Limitations of Prompt Tuning

  • Yihan Wang
  • Jatin Chauhan
  • Wei Wang
  • Cho-Jui Hsieh

Despite the demonstrated empirical efficacy of prompt tuning to adapt a pretrained language model for a new task, the theoretical underpinnings of the difference between "tuning parameters before the input" against "the tuning of model weights" are limited. We thus take one of the first steps to understand the role of soft-prompt tuning for transformer-based architectures. By considering a general purpose architecture, we analyze prompt tuning from the lens of both: universal approximation and limitations with finite-depth fixed-weight pretrained transformers for continuous-valued functions. Our universality result guarantees the existence of a strong transformer with a prompt to approximate any sequence-to-sequence function in the set of Lipschitz functions. The limitations of prompt tuning for limited-depth transformers are first proved by constructing a set of datasets, that cannot be memorized by a prompt of any length for a given single encoder layer. We also provide a lower bound on the required number of tunable prompt parameters and compare the result with the number of parameters required for a low-rank update (based on LoRA) for a single-layer setting. We finally extend our analysis to multi-layer settings by providing sufficient conditions under which the transformer can at best learn datasets from invertible functions only. Our theoretical claims are also corroborated by empirical results.

JBHI Journal 2022 Journal Article

A Multi-Modal Gait Analysis-Based Detection System of the Risk of Depression

  • WEI SHAO
  • Zhiyang You
  • Lesheng Liang
  • Xiping Hu
  • Chengming Li
  • Wei Wang
  • Bin Hu

Currently, depression has become a common mental disorder, especially among postgraduates. It is reported that postgraduates have a higher risk of depression than the general public, and they are more sensitive to contact with others. Thus, a non-contact and effective method for detecting people at risk of depression becomes an urgent demand. In order to make the recognition of depression more reliable and convenient, we propose a multi-modal gait analysis-based depression detection method that combines skeleton modality and silhouette modality. Firstly, we propose a skeleton feature set to describe depression and train a Long Short-Term Memory (LSTM) model to conduct sequence strategy. Secondly, we generate Gait Energy Image (GEI) as silhouette features from RGB videos, and design two Convolutional Neural Network (CNN) models with a new loss function to extract silhouette features from front and side perspectives. Then, we construct a multi-modal fusion model consisting of fusing silhouettes from the front and side views at the feature level and the classification results of different modalities at the decision level. The proposed multi-modal model achieved accuracy at 85. 45% in the dataset consisting of 200 postgraduate students (including 86 depressive ones), 5. 17% higher than the best single-mode model. The multi-modal method also shows improved generalization by reducing the gender differences. Furthermore, we design a vivid 3D visualization of the gait skeletons, and our results imply that gait is a potent biometric for depression detection.

EAAI Journal 2022 Journal Article

Boosting the prediction of molten steel temperature in ladle furnace with a dynamic outlier ensemble

  • Biao Wang
  • Wenjing Wang
  • Guanglei Meng
  • Zhihua Qiao
  • Yuming Guo
  • Na Wang
  • Wei Wang
  • Zhizhong Mao

Molten steel temperature prediction is a critical step in the development of level-two control systems for ladle furnace. Many machine learning algorithms have been employed to complete such a work. Whereas data-driven predictors often deteriorate due to the presence of outliers in practical applications. This paper proposes to boost the predictive performance via outlier detection. Specifically, a dynamic outlier ensemble is developed inspired by the superiority of dynamic classifier selection in classification. Clustering analysis is used to determine the region of competence, on which base detectors are selected with the dedicated measure. The reason for the usage of clustering analysis lies in its efficiency during online detection. One attribute weighting algorithm is used to enhance the capability of clustering in outlier detection. The information behind regression is used to facilitate the measure of competence, results of which can promote the performance of predictors. Such a strategy can achieve double-win from the perspective of regression and outlier detection. Extensive experiments on real-world data sets show that results of all 4 predictive models with respect to accuracy and hit rate can be improved. Moreover, the detection performance in terms of G-mean and F1 score of our detector has also been confirmed via the comparison with 8 competitors.

NeurIPS Conference 2022 Conference Paper

Collaborative Learning by Detecting Collaboration Partners

  • Shu Ding
  • Wei Wang

Massive amounts of data are naturally dispersed over different clients in many real-world applications, collaborative learning has been a promising paradigm that allows to learn models through collaboration among the clients. However, leveraging these dispersed data to learn good models is still challenging since data over different clients are heterogeneous. Previous works mainly focus on learning the centralized model for all clients or learning a personalized model for each client. When there are numerous clients, the centralized model performs badly on some clients, while learning a personalized model for each client costs unaffordable computational resources. In this paper, we propose the collaborative learning method to detect collaboration partners and adaptively learn $K$ models for numerous heterogeneous clients. We theoretically prove that the model learned for each client is a good approximation of its personalized model. Experimental results on real-world datasets verify the effectiveness of our method.

AAAI Conference 2022 Conference Paper

Content-Variant Reference Image Quality Assessment via Knowledge Distillation

  • Guanghao Yin
  • Wei Wang
  • Zehuan Yuan
  • Chuchu Han
  • Wei Ji
  • Shouqian Sun
  • Changhu Wang

Generally, humans are more skilled at perceiving differences between high-quality (HQ) and low-quality (LQ) images than directly judging the quality of a single LQ image. This situation also applies to image quality assessment (IQA). Although recent no-reference (NR-IQA) methods have made great progress to predict image quality free from the reference image, they still have the potential to achieve better performance since HQ image information is not fully exploited. In contrast, full-reference (FR-IQA) methods tend to provide more reliable quality evaluation, but its practicability is affected by the requirement for pixel-level aligned reference images. To address this, we firstly propose the content-variant reference method via knowledge distillation (CVRKD-IQA). Specifically, we use non-aligned reference (NAR) images to introduce various prior distributions of high-quality images. The comparisons of distribution differences between HQ and LQ images can help our model better assess the image quality. Further, the knowledge distillation transfers more HQ-LQ distribution difference information from the FR-teacher to the NAR-student and stabilizing CVRKD-IQA performance. Moreover, to fully mine the local-global combined information, while achieving faster inference speed, our model directly processes multiple image patches from the input with the MLP-mixer. Cross-dataset experiments verify that our model can outperform all NAR/NR-IQA SOTAs, even reach comparable performance with FR-IQA methods on some occasions. Since the content-variant and non-aligned reference HQ images are easy to obtain, our model can support more IQA applications with its relative robustness to content variations. Our code and more detail elaborations of supplement are available: https: //github. com/guanghaoyin/CVRKD-IQA.

NeurIPS Conference 2022 Conference Paper

CryptoGCN: Fast and Scalable Homomorphically Encrypted Graph Convolutional Network Inference

  • Ran Ran
  • Wei Wang
  • Quan Gang
  • Jieming Yin
  • Nuo Xu
  • Wujie Wen

Recently cloud-based graph convolutional network (GCN) has demonstrated great success and potential in many privacy-sensitive applications such as personal healthcare and financial systems. Despite its high inference accuracy and performance on the cloud, maintaining data privacy in GCN inference, which is of paramount importance to these practical applications, remains largely unexplored. In this paper, we take an initial attempt towards this and develop CryptoGCN--a homomorphic encryption (HE) based GCN inference framework. A key to the success of our approach is to reduce the tremendous computational overhead for HE operations, which can be orders of magnitude higher than its counterparts in the plaintext space. To this end, we develop a solution that can effectively take advantage of the sparsity of matrix operations in GCN inference to significantly reduce the encrypted computational overhead. Specifically, we propose a novel Adjacency Matrix-Aware (AMA) data formatting method along with the AMA assisted patterned sparse matrix partitioning, to exploit the complex graph structure and perform efficient matrix-matrix multiplication in HE computation. In this way, the number of HE operations can be significantly reduced. We also develop a co-optimization framework that can explore the trade-offs among the accuracy, security level, and computational overhead by judicious pruning and polynomial approximation of activation modules in GCNs. Based on the NTU-XVIEW skeleton joint dataset, i. e. , the largest dataset evaluated homomorphically by far as we are aware of, our experimental results demonstrate that CryptoGCN outperforms state-of-the-art solutions in terms of the latency and number of homomorphic operations, i. e. , achieving as much as a 3. 10$\times$ speedup on latency and reduces the total Homomorphic Operation Count (HOC) by 77. 4\% with a small accuracy loss of 1-1. 5$\%$. Our code is publicly available at https: //github. com/ranran0523/CryptoGCN.

AAAI Conference 2022 Short Paper

Deep Representation Debiasing via Mutual Information Minimization and Maximization (Student Abstract)

  • Ruijiang Han
  • Wei Wang
  • Yuxi Long
  • Jiajie Peng

Deep representation learning has succeeded in several fields. However, pre-trained deep representations are usually biased and make downstream models sensitive to different attributes. In this work, we propose a post-processing unsupervised deep representation debiasing algorithm, DeepMin- Max, which can obtain unbiased representations directly from pre-trained representations without re-training or fine-tuning the entire model. The experimental results on synthetic and real-world datasets indicate that DeepMinMax outperforms the existing state-of-the-art algorithms on downstream tasks.

YNICL Journal 2022 Journal Article

Effects of acute high intraocular pressure on red-green and blue-yellow cortical color responses in non-human primates

  • Mengwei Li
  • Xiaoxiao Chen
  • Nini Yuan
  • Yiliang Lu
  • Ye Liu
  • Hongliang Gong
  • Liling Qian
  • Ian Max Andolina

Glaucoma is a leading cause of irreversible blindness worldwide, and intraocular pressure (IOP) is an established and modifiable risk factor for both chronic and acute glaucoma. The relationship between color vision deficits and chronic glaucoma has been described previously. However, the effects of acute glaucoma or acute primary angle closure, which has high prevalence in China, on color vision remains unclear. To address the above question, red-green or blue-yellow color responses in V1, V2, and V4 of seven rhesus macaques were monitored using intrinsic-signal optical imaging while monocular anterior chamber perfusions were performed to reversibly elevate IOP acutely over a clinically observed range of 30 to 90 mmHg. We found that the cortical population responses to both red-green and blue-yellow grating stimuli, systematically decreased as IOP increased from 30 to 90 mmHg. Although a similar decrement in magnitude was noted in V1, V2, and V4, blue-yellow responses were consistently more impaired than red-green responses at all levels of acute IOP elevation and in all monitored visual areas. This physiological study in non-human primates demonstrates that acute IOP elevations substantially depress the ability of the visual cortex to register color information. This effect is more severe for blue-yellow responses than for red-green responses, suggesting selective impairment of the koniocellular pathways compared with the parvocellular pathways. Together, we infer that blue-yellow color vision might be the most vulnerable visual function in acute glaucoma patients.

AAAI Conference 2022 Conference Paper

Exploiting Mixed Unlabeled Data for Detecting Samples of Seen and Unseen Out-of-Distribution Classes

  • Yi-Xuan Sun
  • Wei Wang

Out-of-Distribution (OOD) detection is essential in realworld applications, which has attracted increasing attention in recent years. However, most existing OOD detection methods require many labeled In-Distribution (ID) data, causing a heavy labeling cost. In this paper, we focus on the more realistic scenario, where limited labeled data and abundant unlabeled data are available, and these unlabeled data are mixed with ID and OOD samples. We propose the Adaptive In-Outaware Learning (AIOL) method, in which we employ the appropriate temperature to adaptively select potential ID and OOD samples from the mixed unlabeled data and consider the entropy over them for OOD detection. Moreover, since the test data in realistic applications may contain OOD samples whose classes are not in the mixed unlabeled data (we call them unseen OOD classes), data augmentation techniques are brought into the method to further improve the performance. The experiments are conducted on various benchmark datasets, which demonstrate the superiority of our method.

JBHI Journal 2022 Journal Article

Hematoma Expansion Context Guided Intracranial Hemorrhage Segmentation and Uncertainty Estimation

  • Xiangyu Li
  • Gongning Luo
  • Wei Wang
  • Kuanquan Wang
  • Yue Gao
  • Shuo Li

Accurate segmentation of the Intracranial Hemorrhage (ICH) in non-contrast CT images is significant for computer-aided diagnosis. Although existing methods have achieved remarkable 1 1 The code will be available from https://github.com/JohnleeHIT/SLEX-Net.results, none of them incorporated ICH’s prior information in their methods. In this work, for the first time, we proposed a novel SLice EXpansion Network (SLEX-Net), which incorporated hematoma expansion in the segmentation architecture by directly modeling the hematoma variation among adjacent slices. Firstly, a new module named Slice Expansion Module (SEM) was built, which can effectively transfer contextual information between two adjacent slices by mapping predictions from one slice to another. Secondly, to perceive contextual information from both upper and lower slices, we designed two information transmission paths: forward and backward slice expansion, and aggregated results from those paths with a novel weighing strategy. By further exploiting intra-slice and inter-slice context with the information paths, the network significantly improved the accuracy and continuity of segmentation results. Moreover, the proposed SLEX-Net enables us to conduct an uncertainty estimation with one-time inference, which is much more efficient than existing methods. We evaluated the proposed SLEX-Net and compared it with some state-of-the-art methods. Experimental results demonstrate that our method makes significant improvements in all metrics on segmentation performance and outperforms other existing uncertainty estimation methods in terms of several metrics.

IROS Conference 2022 Conference Paper

Learning Feature Decomposition for Domain Adaptive Monocular Depth Estimation

  • Shao-Yuan Lo
  • Wei Wang
  • Jim Thomas 0001
  • Jingjing Zheng
  • Vishal M. Patel
  • Cheng-Hao Kuo

Monocular depth estimation (MDE) has attracted intense study due to its low cost and critical functions for robotic tasks such as localization, mapping and obstacle detection. Supervised approaches have led to great success with the advance of deep learning, but they rely on large quantities of ground-truth depth annotations that are expensive to acquire. Unsupervised domain adaptation (UDA) transfers knowledge from labeled source data to unlabeled target data, so as to relax the constraint of supervised learning. However, existing UDA approaches may not completely align the domain gap across different datasets because of the domain shift problem. We believe better domain alignment can be achieved via well-designed feature decomposition. In this paper, we propose a novel UDA method for MDE, referred to as Learning Feature Decomposition for Adaptation (LFDA), which learns to decompose the feature space into content and style components. LFDA only attempts to align the content component since it has a smaller domain gap. Meanwhile, it excludes the style component which is specific to the source domain from training the primary task. Furthermore, LFDA uses separate feature distribution estimations to further bridge the domain gap. Extensive experiments on three domain adaptative MDE scenarios show that the proposed method achieves superior accuracy and lower computational cost compared to the state-of-the-art approaches.

IS Journal 2022 Journal Article

Metaverses and DeMetaverses: From Digital Twins in CPS to Parallel Intelligence in CPSS

  • Xiao Wang
  • Jing Yang
  • Jinpeng Han
  • Wei Wang
  • Fei-Yue Wang

A total of 12 years have been passed since this Department was created in 2010 as the first academic forum dedicated to cyber-physical-social systems (CPSS), with the first CPSS research article on the field: “The Emergence of Intelligent Enterprises: From CPS to CPSS. ” What has happened and changed during the past decade? A brief reflection and review are presented here with a focus on digital twins in CPS versus parallel intelligence in CPSS, and their relationship to blockchain intelligence, smart contracts, metaverses, DAO, Web3, and decentralized science. The concept of DeMetaverses is thus introduced and interpreted as a DAO-based decentralized autonomous metaverse. The characteristics, mechanism, and impact of DeMetaverses are discussed with a vision for achieving an integrated human, artificial, natural, and organizational intelligence that would transform our world into “6S” societies.

NeurIPS Conference 2022 Conference Paper

RankFeat: Rank-1 Feature Removal for Out-of-distribution Detection

  • Yue Song
  • Nicu Sebe
  • Wei Wang

The task of out-of-distribution (OOD) detection is crucial for deploying machine learning models in real-world settings. In this paper, we observe that the singular value distributions of the in-distribution (ID) and OOD features are quite different: the OOD feature matrix tends to have a larger dominant singular value than the ID feature, and the class predictions of OOD samples are largely determined by it. This observation motivates us to propose RankFeat, a simple yet effective post hoc approach for OOD detection by removing the rank-1 matrix composed of the largest singular value and the associated singular vectors from the high-level feature. RankFeat achieves state-of-the-art performance and reduces the average false positive rate (FPR95) by 17. 90% compared with the previous best method. Extensive ablation studies and comprehensive theoretical analyses are presented to support the empirical results.

JBHI Journal 2022 Journal Article

Skeleton-Based Abnormal Behavior Detection Using Secure Partitioned Convolutional Neural Network Model

  • Jiefan Qiu
  • Xinlei Yan
  • Wei Wang
  • Wei Wei
  • Kai Fang

Theabnormal behavior detection is the vital for evaluation of daily-life health status of the patient with cognitive impairment. Previous studies about abnormal behavior detection indicate that convolution neural network (CNN)-based computer vision owns the high robustness and accuracy for detection. However, executing CNN model on the cloud possible incurs a privacy disclosure problem during data transmission, and the high computation overhead makes difficult to execute the model on edge-end IoT devices with a well real-time performance. In this paper, we realize a skeleton-based abnormal behavior detection, and propose a secure partitioned CNN model (SP-CNN) to extract human skeleton keypoints and achieve safely collaborative computing by deploying different CNN model layers on the cloud and the IoT device. Because, the data outputted from the IoT device are processed by the several CNN layers instead of transmitting the sensitive video data, objectively it reduces the risk of privacy disclosure. Moreover, we also design an encryption method based on channel state information (CSI) to guarantee the sensitive data security. At last, we apply SP-CNN in abnormal behavior detection to evaluate its effectiveness. The experiment results illustrate that the efficiency of the abnormal behavior detection based on SP-CNN is at least 33. 2% higher than the state-of-the-art methods, and its detection accuracy arrives to 97. 54%.

AAAI Conference 2022 Conference Paper

Towards Fine-Grained Reasoning for Fake News Detection

  • Yiqiao Jin
  • Xiting Wang
  • Ruichao Yang
  • Yizhou Sun
  • Wei Wang
  • Hao Liao
  • Xing Xie

The detection of fake news often requires sophisticated reasoning skills, such as logically combining information by considering word-level subtle clues. In this paper, we move towards fine-grained reasoning for fake news detection by better reflecting the logical processes of human thinking and enabling the modeling of subtle clues. In particular, we propose a fine-grained reasoning framework by following the human’s information-processing model, introduce a mutualreinforcement-based method for incorporating human knowledge about which evidence is more important, and design a prior-aware bi-channel kernel graph network to model subtle differences between pieces of evidence. Extensive experiments show that our model outperforms the state-of-the-art methods and demonstrate the explainability of our approach.

AAAI Conference 2021 Conference Paper

A Unified Pretraining Framework for Passage Ranking and Expansion

  • Ming Yan
  • Chenliang Li
  • Bin Bi
  • Wei Wang
  • Songfang Huang

Pretrained language models have recently advanced a wide range of natural language processing tasks. Nowadays, the application of pretrained language models to IR tasks has also achieved impressive results. Typical methods either directly apply a pretrained model to improve the re-ranking stage, or use it to conduct passage expansion and term weighting for first-stage retrieval. We observe that the passage ranking and passage expansion tasks share certain inherent relations, and can benefit from each other. Therefore, in this paper, we propose a general pretraining framework to enhance both tasks with Unified Encoder-Decoder networks (UED). The overall ranking framework consists of two parts in a cascade manner: (1) passage expansion with a pretraining-based query generation method; (2) re-ranking of passage candidates from a traditional retrieval method with a pretrained transformer encoder. Both the two parts are based on the same pretrained UED model, where we jointly train the passage ranking and query generation tasks for further improving the full ranking pipeline. An extensive set of experiments have been conducted on two large-scale passage retrieval datasets to demonstrate the state-of-the-art results of the proposed framework in both the first-stage retrieval and the final re-ranking. In addition, we successfully deploy the framework to our online production system, which can stably serve industrial applications with a request volume of up to 100 QPS in less than 300ms.

AAAI Conference 2021 Conference Paper

Adversarial Defence by Diversified Simultaneous Training of Deep Ensembles

  • Bo Huang
  • Zhiwei Ke
  • Yi Wang
  • Wei Wang
  • Linlin Shen
  • Feng Liu

Learning-based classifiers are susceptible to adversarial examples. Existing defence methods are mostly devised on individual classifiers. Recent studies showed that it is viable to increase adversarial robustness by promoting diversity over an ensemble of models. In this paper, we propose adversarial defence by encouraging ensemble diversity on learning high-level feature representations and gradient dispersion in simultaneous training of deep ensemble networks. We perform extensive evaluations under white-box and blackbox attacks including transferred examples and adaptive attacks. Our approach achieves a significant gain of up to 52% in adversarial robustness, compared with the baseline and the state-of-the-art method on image benchmarks with complex data scenes. The proposed approach complements the defence paradigm of adversarial training, and can further boost the performance. The source code is available at https: //github. com/ALIS-Lab/AAAI2021-PDD.

YNIMG Journal 2021 Journal Article

Characterizing the seizure onset zone and epileptic network using EEG-fMRI in a rat seizure model

  • Junling Wang
  • Bin Jing
  • Ru Liu
  • Donghong Li
  • Wei Wang
  • Jiaoyang Wang
  • Jianfeng Lei
  • Yue Xing

Accurate epileptogenic zone (EZ) or seizure onset zone (SOZ) localization is crucial for epilepsy surgery optimization. Previous animal and human studies on epilepsy have reported that changes in blood oxygen level-dependent (BOLD) signals induced by epileptic events could be used as diagnostic markers for EZ or SOZ localization. Simultaneous electroencephalography and functional magnetic resonance imaging (EEG-fMRI) recording is gaining interest as a non-invasive tool for preoperative epilepsy evaluation. However, EEG-fMRI studies have reported inconsistent and ambiguous findings. Therefore, it remains unclear whether BOLD responses can be used for accurate EZ or SOZ localization. In this study, we used simultaneous EEG-fMRI recording in a rat model of 4-aminopyridine-induced acute focal seizures to assess the spatial concordance between individual BOLD responses and the SOZ. This was to determine the optimal use of simultaneous EEG-fMRI recording in the SOZ localization. We observed a high spatial consistency between BOLD responses and the SOZ. Further, dynamic BOLD responses were consistent with the regions where the seizures were propagated. These results suggested that simultaneous EEG-fMRI recording could be used as a noninvasive clinical diagnostic technique for localizing the EZ or SOZ and could be an effective tool for mapping epileptic networks.

AAAI Conference 2021 Conference Paper

Clinical Temporal Relation Extraction with Probabilistic Soft Logic Regularization and Global Inference

  • Yichao Zhou
  • Yu Yan
  • Rujun Han
  • J. Harry Caufield
  • Kai-Wei Chang
  • Yizhou Sun
  • Peipei Ping
  • Wei Wang

There has been a steady need in the medical community to precisely extract the temporal relations between clinical events. In particular, temporal information can facilitate a variety of downstream applications such as case report retrieval and medical question answering. However, existing methods either require expensive feature engineering or are incapable of modeling the global relational dependencies among the events. In this paper, we propose Clinical Temporal ReLation Exaction with Probabilistic Soft Logic Regularization and Global Inference (CTRL-PG), a novel method to tackle the problem at the document level. Extensive experiments on two benchmark datasets, I2B2-2012 and TB-Dense, demonstrate that CTRL-PG significantly outperforms baseline methods for temporal relation extraction.

JMLR Journal 2021 Journal Article

Empirical Bayes Matrix Factorization

  • Wei Wang
  • Matthew Stephens

Matrix factorization methods, which include Factor analysis (FA) and Principal Components Analysis (PCA), are widely used for inferring and summarizing structure in multivariate data. Many such methods use a penalty or prior distribution to achieve sparse representations (“Sparse FA/PCA"), and a key question is how much sparsity to induce. Here we introduce a general Empirical Bayes approach to matrix factorization (EBMF), whose key feature is that it estimates the appropriate amount of sparsity by estimating prior distributions from the observed data. The approach is very flexible: it allows for a wide range of different prior families and allows that each component of the matrix factorization may exhibit a different amount of sparsity. The key to this flexibility is the use of a variational approximation, which we show effectively reduces fitting the EBMF model to solving a simpler problem, the so-called “normal means" problem. We demonstrate the benefits of EBMF with sparse priors through both numerical comparisons with competing methods and through analysis of data from the GTEx (Genotype Tissue Expression) project on genetic associations across 44 human tissues. In numerical comparisons EBMF often provides more accurate inferences than other methods. In the GTEx data, EBMF identifies interpretable structure that agrees with known relationships among human tissues. Software implementing our approach is available at https://github.com/stephenslab/flashr. [abs] [ pdf ][ bib ] [ code ] &copy JMLR 2021. ( edit, beta )

IJCAI Conference 2021 Conference Paper

Few-Shot Learning with Part Discovery and Augmentation from Unlabeled Images

  • Wentao Chen
  • Chenyang Si
  • Wei Wang
  • Liang Wang
  • Zilei Wang
  • Tieniu Tan

Few-shot learning is a challenging task since only few instances are given for recognizing an unseen class. One way to alleviate this problem is to acquire a strong inductive bias via meta-learning on similar tasks. In this paper, we show that such inductive bias can be learned from a flat collection of unlabeled images, and instantiated as transferable representations among seen and unseen classes. Specifically, we propose a novel part-based self-supervised representation learning scheme to learn transferable representations by maximizing the similarity of an image to its discriminative part. To mitigate the overfitting in few-shot classification caused by data scarcity, we further propose a part augmentation strategy by retrieving extra images from a base dataset. We conduct systematic studies on miniImageNet and tieredImageNet benchmarks. Remarkably, our method yields impressive results, outperforming the previous best unsupervised methods by 7. 74% and 9. 24% under 5-way 1-shot and 5-way 5-shot settings, which are comparable with state-of-the-art supervised methods.

AAAI Conference 2021 Conference Paper

Generating Diversified Comments via Reader-Aware Topic Modeling and Saliency Detection

  • Wei Wang
  • Piji Li
  • Hai-Tao Zheng

Automatic comment generation is a special and challenging task to verify the model ability on news content comprehension and language generation. Comments not only convey salient and interesting information in news articles, but also imply various and different reader characteristics which we treat as the essential clues for diversity. However, most of the comment generation approaches only focus on saliency information extraction, while the reader-aware factors implied by comments are neglected. To address this issue, we propose a unified reader-aware topic modeling and saliency information detection framework to enhance the quality of generated comments. For reader-aware topic modeling, we design a variational generative clustering algorithm for latent semantic learning and topic mining from reader comments. For saliency information detection, we introduce Bernoulli distribution estimating on news content to select saliency information. The obtained topic representations as well as the selected saliency information are incorporated into the decoder to generate diversified and informative comments. Experimental results on three datasets show that our framework outperforms existing baseline methods in terms of both automatic metrics and human evaluation. The potential ethical issues are also discussed in detail.

JBHI Journal 2021 Journal Article

Guest Editorial AI and 5G Empowered Internet of Medical Things

  • Syed Hassan Ahmed
  • Victor Hugo de Albuquerque
  • Wei Wei
  • Wei Wang

The papers in this special section focus on artificial intelligence (AI) and 5G Internet of Medical Things. The recent developments in biomedical sensors, wireless communication systems, and information networks are transforming the conventional healthcare systems. The transformed healthcare systems are enabling distributed healthcare services to patients who may not be co-located with the healthcare providers, providing early diagnoses, and reducing the cost in the healthcare section. The Internet of Medical Things (IoMT), which includes medical devices, wearable devices, sensors and apps, is a critical piece of the digital transformation of healthcare, as it allows new business models to emerge and enables changes in work processes, productivity improvements, cost containment and enhanced customer experiences. IoMT can help monitor, inform and notify not only care-givers, but provide healthcare providers with actual data to identify issues bef

AAAI Conference 2021 Conference Paper

Improving the Efficiency and Effectiveness for BERT-based Entity Resolution

  • Bing Li
  • Yukai Miao
  • Yaoshu Wang
  • Yifang Sun
  • Wei Wang

BERT has set a new state-of-the-art performance on entity resolution (ER) task, largely owed to fine-tuning pretrained language models and the deep pair-wise interaction. Albeit being remarkably effective, it comes with a steep increase in computational cost, as the deep-interaction requires to exhaustively compute every tuple pair to search for coreferences. For ER task, it is often prohibitively expensive due to the large cardinality to be matched. To tackle this, we introduce a siamese network structure that independently encodes tuples using BERT but delays the pair-wise interaction via an enhanced alignment network. This siamese structure enables a dedicated blocking module to quickly filter out obviously dissimilar tuple pairs, and thus drastically reduces the cardinality of fine-grained matching. Further, the blocking and entity matching are integrated into a multi-task learning framework for facilitating both tasks. Extensive experiments on multiple datasets demonstrate that our model significantly outperforms state-of-the-art models (including BERT) in both efficiency and effectiveness.

AAAI Conference 2021 Short Paper

Is Each Layer Non-trivial in CNN? (Student Abstract)

  • Wei Wang
  • Yanjie Zhu
  • Zhuoxu Cui
  • Dong Liang

Convolutional neural network (CNN) models have achieved great success in many fields. With the advent of ResNet, networks used in practice are getting deeper and wider. However, is each layer non-trivial in networks? To answer this question, we trained a network on the training set, then we replace the network convolution kernels with zeros and test the result models on the test set. We compared experimental results with baseline and showed that we can reach similar or even the same performances. Although convolution kernels are the cores of networks, we demonstrate that some of them are trivial and regular in ResNet.

NeurIPS Conference 2021 Conference Paper

KS-GNN: Keywords Search over Incomplete Graphs via Graphs Neural Network

  • Yu Hao
  • Xin Cao
  • Yufan Sheng
  • Yixiang Fang
  • Wei Wang

Keyword search is a fundamental task to retrieve information that is the most relevant to the query keywords. Keyword search over graphs aims to find subtrees or subgraphs containing all query keywords ranked according to some criteria. Existing studies all assume that the graphs have complete information. However, real-world graphs may contain some missing information (such as edges or keywords), thus making the problem much more challenging. To solve the problem of keyword search over incomplete graphs, we propose a novel model named KS-GNN based on the graph neural network and the auto-encoder. By considering the latent relationships and the frequency of different keywords, the proposed KS-GNN aims to alleviate the effect of missing information and is able to learn low-dimensional representative node embeddings that preserve both graph structure and keyword features. Our model can effectively answer keyword search queries with linear time complexity over incomplete graphs. The experiments on four real-world datasets show that our model consistently achieves better performance than state-of-the-art baseline methods in graphs having missing information.

AAAI Conference 2021 Conference Paper

Learning to Copy Coherent Knowledge for Response Generation

  • Jiaqi Bai
  • Ze Yang
  • Xinnian Liang
  • Wei Wang
  • Zhoujun Li

Knowledge-driven dialog has shown remarkable performance to alleviate the problem of generating uninformative responses in the dialog system. However, incorporating knowledge coherently and accurately into response generation is still far from being solved. Previous works dropped into the paradigm of non-goal-oriented knowledge-driven dialog, they are prone to ignore the effect of dialog goal, which has potential impacts on knowledge exploitation and response generation. To address this problem, this paper proposes a Goal-Oriented Knowledge Copy network, GOKC. Specifically, a goal-oriented knowledge discernment mechanism is designed to help the model discern the knowledge facts that are highly correlated to the dialog goal and the dialog context. Besides, a context manager is devised to copy facts not only from the discerned knowledge but also from the dialog goal and the dialog context, which allows the model to accurately restate the facts in the generated response. The empirical studies are conducted on two benchmarks of goal-oriented knowledge-driven dialog generation. The results show that our model can significantly outperform several state-of-theart models in terms of both automatic evaluation and human judgments.

AAAI Conference 2021 Conference Paper

Multi-View Representation Learning with Manifold Smoothness

  • Shu Li
  • Wei Wang
  • Wen-Tao Li
  • Pan Chen

Multi-view representation learning attempts to learn a representation from multiple views and most existing methods are unsupervised. However, representation learned only from unlabeled data may not be discriminative enough for further applications (e. g. , clustering and classification). For this reason, semi-supervised methods which could use unlabeled data along with the labeled data for multi-view representation learning need to be developed. Manifold information plays an important role in semi-supervised learning, but it has not been considered for multi-view representation learning. In this paper, we introduce the manifold smoothness into multiview representation learning and propose MvDGAT which learns the representation and the intrinsic manifold simultaneously with graph attention network. Experiments conducted on real-world datasets reveal that our MvDGAT can achieve better performance than state-of-the-art methods.

AAAI Conference 2021 Conference Paper

Open Domain Dialogue Generation with Latent Images

  • Ze Yang
  • Wei Wu
  • Huang Hu
  • Can Xu
  • Wei Wang
  • Zhoujun Li

We consider grounding open domain dialogues with images. Existing work assumes that both an image and a textual context are available, but image-grounded dialogues by nature are more difficult to obtain than textual dialogues. Thus, we propose learning a response generation model with both image-grounded dialogues and textual dialogues by assuming that the visual scene information at the time of a conversation can be represented by an image, and trying to recover the latent images of the textual dialogues through text-to-image generation techniques. The likelihood of the two types of dialogues is then formulated by a response generator and an image reconstructor that are learned within a conditional variational auto-encoding framework. Empirical studies are conducted in both image-grounded conversation and text-based conversation. In the first scenario, image-grounded dialogues, especially under a low-resource setting, can be effectively augmented by textual dialogues with latent images; while in the second scenario, latent images can enrich the content of responses and at the same time keep them relevant to contexts.

AAAI Conference 2021 Conference Paper

Precise Yet Efficient Semantic Calibration and Refinement in ConvNets for Real-time Polyp Segmentation from Colonoscopy Videos

  • Huisi Wu
  • Jiafu Zhong
  • Wei Wang
  • Zhenkun Wen
  • Jing Qin

We propose a novel convolutional neural network (ConvNet) equipped with two new semantic calibration and refinement approaches for automatic polyp segmentation from colonoscopy videos. While ConvNets set state-of-the-are performance for this task, it is still difficult to achieve satisfactory results in a real-time manner, which is a necessity in clinical practice. The main obstacle is the huge semantic gap between high-level features and low-level features, making it difficult to take full advantage of complementary semantic information contained in these hierarchical features. Compared with existing solutions, which either directly aggregate these features without considering the semantic gap or employ sophisticated non-local modeling techniques to refine semantic information by introduce many extra computational costs, the proposed ConvNet is able to more precisely yet efficiently calibrate and refine semantic information for better segmentation performance without increasing model complexity; we call the proposed ConvNet as SCR-Net, which has two key modules. We first propose a semantic calibration module (SCM) to effectively transmit the semantic information from high-level layers to low-level layers by learning the semantic-spatial relations during the training procedure. We then propose a semantic refinement module (SRM) to, based on the features calibrated by SCM, enhance the discrimination capability of the features for targeting objects. Extensive experiments on the Kvasir-SEG dataset demonstrate that the proposed SCR-Net is capable of achieving better segmentation accuracy than state-of-the-art approaches with a faster speed. The proposed techniques are general enough to be applied to similar applications where precise and efficient multi-level feature fusion is critical. The code is available at https: //github. com/jiafuz/SCR-Net.

AAAI Conference 2021 Conference Paper

Region-aware Global Context Modeling for Automatic Nerve Segmentation from Ultrasound Images

  • Huisi Wu
  • Jiasheng Liu
  • Wei Wang
  • Zhenkun Wen
  • Jing Qin

We present a novel deep learning model equipped with a new region-aware global context modeling technique for automatic nerve segmentation from ultrasound images, which is a challenging task due to (1) the large variation and blurred boundaries of targets, (2) the large amount of speckle noise in ultrasound images, and (3) the inherent real-time requirement of this task. It is essential to efficiently capture long-range dependencies by global context modeling for a segmentation network to overcome these challenges. Traditional global context modeling techniques usually explore pixel-aware correlations to establish long-range dependencies, which are usually computation-intensive and greatly degrade time performance. In addition, in this application, pixel-aware modeling may inevitably introduce much speckle noise in the computation and potentially degrade segmentation performance. In this paper, we propose a novel region-aware modeling technique to establish long-range dependencies based on different regions to improve segmentation accuracy while maintaining real-time performance; we call it region-aware pyramid aggregation (RPA) module. In order to adaptively divide the feature maps into a set of semantic-independent regions, we develop an attention mechanism and integrate it into the spatial pyramid network to evaluate the semantic similarity of different regions. We further develop an adaptive pyramid fusion (APF) module to dynamically fuse the multi-level features generated from the decoder to refining the segmentation results. We conducted extensive experiments on a famous public ultrasound nerve image segmentation dataset. Experimental results demonstrate that our method consistently outperforms our rivals in terms of segmentation accuracy. The code is available at https: //github. com/jsonliu-szu/RAGCM.

ICML Conference 2021 Conference Paper

Self-supervised and Supervised Joint Training for Resource-rich Machine Translation

  • Yong Cheng
  • Wei Wang
  • Lu Jiang 0004
  • Wolfgang Macherey

Self-supervised pre-training of text representations has been successfully applied to low-resource Neural Machine Translation (NMT). However, it usually fails to achieve notable gains on resource-rich NMT. In this paper, we propose a joint training approach, F2-XEnDec, to combine self-supervised and supervised learning to optimize NMT models. To exploit complementary self-supervised signals for supervised learning, NMT models are trained on examples that are interbred from monolingual and parallel sentences through a new process called crossover encoder-decoder. Experiments on two resource-rich translation benchmarks, WMT’14 English-German and WMT’14 English-French, demonstrate that our approach achieves substantial improvements over several strong baseline methods and obtains a new state of the art of 46. 19 BLEU on English-French when incorporating back translation. Results also show that our approach is capable of improving model robustness to input perturbations such as code-switching noise which frequently appears on the social media.

IROS Conference 2021 Conference Paper

Towards Autonomous Parking using Vision-only Sensors

  • Yi Yang 0009
  • Miaoxin Pan
  • Sitan Jiang
  • Jianhang Wang
  • Wei Wang
  • Junbo Wang
  • Meiling Wang 0002

Existing autonomous parking solutions usually require special signs, pre-built maps or accurate ranging sensors to achieve reliable perception of the parking environment, but these methods are difficult to popularize because they either require preconditions or are expensive for production cars. In this paper, we propose a vision-only autonomous parking solution based on only six cameras. Through the appropriate depth estimation algorithms, our method obtains the pixel level depth of the image, and constructs a dense point cloud, so as to realize the fine perception of the parking environment. An improved Radon transform based parking space detection method are applied for better parking space detection method. Our proposed method achieves processing speed of above 5 Hz on a intermediate level computing platform. Furthermore, we demonstrate the practicability of the proposed system in real-world parking lots.

IJCAI Conference 2021 Conference Paper

Towards Understanding Deep Learning from Noisy Labels with Small-Loss Criterion

  • Xian-Jin Gui
  • Wei Wang
  • Zhang-Hao Tian

Deep neural networks need large amounts of labeled data to achieve good performance. In real-world applications, labels are usually collected from non-experts such as crowdsourcing to save cost and thus are noisy. In the past few years, deep learning methods for dealing with noisy labels have been developed, many of which are based on the small-loss criterion. However, there are few theoretical analyses to explain why these methods could learn well from noisy labels. In this paper, we theoretically explain why the widely-used small-loss criterion works. Based on the explanation, we reformalize the vanilla small-loss criterion to better tackle noisy labels. The experimental results verify our theoretical explanation and also demonstrate the effectiveness of the reformalization.

AAAI Conference 2021 Conference Paper

Two-Stream Convolution Augmented Transformer for Human Activity Recognition

  • Bing Li
  • Wei Cui
  • Wei Wang
  • Le Zhang
  • Zhenghua Chen
  • Min Wu

Recognition of human activities is an important task due to its far-reaching applications such as healthcare system, context-aware applications, and security monitoring. Recently, WiFi based human activity recognition (HAR) is becoming ubiquitous due to its non-invasiveness. Existing WiFibased HAR methods regard WiFi signals as a temporal sequence of channel state information (CSI), and employ deep sequential models (e. g. , RNN, LSTM) to automatically capture channel-over-time features. Although being remarkably effective, they suffer from two major drawbacks. Firstly, the granularity of a single temporal point is blindly elementary for representing meaningful CSI patterns. Secondly, the timeover-channel features are also important, and could be a natural data augmentation. To address the drawbacks, we propose a novel Two-stream Convolution Augmented Human Activity Transformer (THAT) model. Our model proposes to utilize a two-stream structure to capture both time-over-channel and channel-over-time features, and use the multi-scale convolution augmented transformer to capture range-based patterns. Extensive experiments on four real experiment datasets demonstrate that our model outperforms state-of-the-art models in terms of both effectiveness and efficiency 1.

AAAI Conference 2020 Conference Paper

A Recurrent Model for Collective Entity Linking with Adaptive Features

  • Xiaoling Zhou
  • Yukai Miao
  • Wei Wang
  • Jianbin Qin

The vast amount of web data enables us to build knowledge bases with unprecedented quality and coverage. Named Entity Disambiguation (NED) is an important task that automatically resolves ambiguous mentions in free text to correct target entries in the knowledge base. Traditional machine learning based methods for NED were outperformed and made obsolete by the state-of-the-art deep learning based models. However, deep learning models are more complex, requiring large amount of training data and lengthy training and parameter tuning time. In this paper, we revisit traditional machine learning techniques and propose a light-weight, tuneable and time-efficient method without using deep learning or deep learning generated features. We propose novel adaptive features that focus on extracting discriminative features to better model similarities between candidate entities and the mention’s context. We learn a local ranking model based on traditional and the new adaptive features based on the learning-to-rank framework. While arriving at linking decisions individually via the local model, our method also takes into consideration the correlation between decisions by running multiple recurrent global models, which can be deemed as a learned local search method. Our method attains performances comparable to the state-of-the-art deep learning-based methods on NED benchmark datasets while being significantly faster to train.

AAAI Conference 2020 Conference Paper

Co-GCN for Multi-View Semi-Supervised Learning

  • Shu Li
  • Wen-Tao Li
  • Wei Wang

In many real-world applications, the data have several disjoint sets of features and each set is called as a view. Researchers have developed many multi-view learning methods in the past decade. In this paper, we bring Graph Convolutional Network (GCN) into multi-view learning and propose a novel multi-view semi-supervised learning method Co-GCN by adaptively exploiting the graph information from the multiple views with combined Laplacians. Experimental results on real-world data sets verify that Co-GCN can achieve better performance compared with state-of-the-art multi-view semisupervised methods.

JBHI Journal 2020 Journal Article

Differential Diagnosis of Atypical Hepatocellular Carcinoma in Contrast-Enhanced Ultrasound Using Spatio-Temporal Diagnostic Semantics

  • Qinghua Huang
  • Fengxin Pan
  • Wei Li
  • Feiniu Yuan
  • Hangtong Hu
  • Jinhua Huang
  • Jie Yu
  • Wei Wang

Atypical Hepatocellular Carcinoma (HCC) is very hard to distinguish from Focal Nodular Hyperplasia (FNH) in routine imaging. However little attention was paid to this problem. This paper proposes a novel liver tumor Computer-Aided Diagnostic (CAD) approach extracting spatio-temporal semantics for atypical HCC. With respect to useful diagnostic semantics, our model automatically calculates three types of semantic feature with equally down-sampled frames based on Contrast-Enhanced Ultrasound (CEUS). Thereafter, a Support Vector Machine (SVM) classifier is trained to make the final diagnosis. Compared with traditional methods for diagnosing HCC, the proposed model has the advantage of less computational complexity and being able to handle the atypical HCC cases. The experimental results show that our method obtained a pretty considerable performance and outperformed two traditional methods. According to the results, the average accuracy reaches 94. 40%, recall rate 94. 76%, F1-score value 94. 62%, specificity 93. 62% and sensitivity 94. 76%, indicating good merit for automatically diagnosing atypical HCC cases.

AAAI Conference 2020 Conference Paper

Dynamic Malware Analysis with Feature Engineering and Feature Learning

  • Zhaoqi Zhang
  • Panpan Qi
  • Wei Wang

Dynamic malware analysis executes the program in an isolated environment and monitors its run-time behaviour (e. g. system API calls) for malware detection. This technique has been proven to be effective against various code obfuscation techniques and newly released (“zero-day”) malware. However, existing works typically only consider the API name while ignoring the arguments, or require complex feature engineering operations and expert knowledge to process the arguments. In this paper, we propose a novel and low-cost feature extraction approach, and an effective deep neural network architecture for accurate and fast malware detection. Specifically, the feature representation approach utilizes a feature hashing trick to encode the API call arguments associated with the API name. The deep neural network architecture applies multiple Gated-CNNs (convolutional neural networks) to transform the extracted features of each API call. The outputs are further processed through bidirectional LSTM (long-short term memory networks) to learn the sequential correlation among API calls. Experiments show that our solution outperforms baselines significantly on a large real dataset. Valuable insights about feature engineering and architecture design are derived from the ablation study.

AAAI Conference 2020 Conference Paper

ECGadv: Generating Adversarial Electrocardiogram to Misguide Arrhythmia Classification System

  • Huangxun Chen
  • Chenyu Huang
  • Qianyi Huang
  • Qian Zhang
  • Wei Wang

Deep neural networks (DNNs)-powered Electrocardiogram (ECG) diagnosis systems recently achieve promising progress to take over tedious examinations by cardiologists. However, their vulnerability to adversarial attacks still lack comprehensive investigation. The existing attacks in image domain could not be directly applicable due to the distinct properties of ECGs in visualization and dynamic properties. Thus, this paper takes a step to thoroughly explore adversarial attacks on the DNN-powered ECG diagnosis system. We analyze the properties of ECGs to design effective attacks schemes under two attacks models respectively. Our results demonstrate the blind spots of DNN-powered diagnosis systems under adversarial attacks, which calls attention to adequate countermeasures.

AAAI Conference 2020 Conference Paper

Fine-Grained Named Entity Typing over Distantly Supervised Data Based on Refined Representations

  • Muhammad Asif Ali
  • Yifang Sun
  • Bing Li
  • Wei Wang

Fine-Grained Named Entity Typing (FG-NET) is a key component in Natural Language Processing (NLP). It aims at classifying an entity mention into a wide range of entity types. Due to a large number of entity types, distant supervision is used to collect training data for this task, which noisily assigns type labels to entity mentions irrespective of the context. In order to alleviate the noisy labels, existing approaches on FG-NET analyze the entity mentions entirely independent of each other and assign type labels solely based on mention’s sentence-specific context. This is inadequate for highly overlapping and/or noisy type labels as it hinders information passing across sentence boundaries. For this, we propose an edge-weighted attentive graph convolution network that refines the noisy mention representations by attending over corpus-level contextual clues prior to the end classification. Experimental evaluation shows that the proposed model outperforms the existing research by a relative score of upto 10. 2% and 8. 3% for macro-f1 and micro-f1 respectively.

AAAI Conference 2020 Conference Paper

Generating Well-Formed Answers by Machine Reading with Stochastic Selector Networks

  • Bin Bi
  • Chen Wu
  • Ming Yan
  • Wei Wang
  • Jiangnan Xia
  • Chenliang Li

Question answering (QA) based on machine reading comprehension has been a recent surge in popularity, yet most work has focused on extractive methods. We instead address a more challenging QA problem of generating a well-formed answer by reading and summarizing the paragraph for a given question. For the generative QA task, we introduce a new neural architecture, LatentQA, in which a novel stochastic selector network composes a well-formed answer with words selected from the question, the paragraph and the global vocabulary, based on a sequence of discrete latent variables. Bayesian inference for the latent variables is performed to train the LatentQA model. The experiments on public datasets of natural answer generation confirm the effectiveness of LatentQA in generating high-quality well-formed answers.

AAAI Conference 2020 Conference Paper

GraphER: Token-Centric Entity Resolution with Graph Convolutional Neural Networks

  • Bing Li
  • Wei Wang
  • Yifang Sun
  • Linhan Zhang
  • Muhammad Asif Ali
  • Yi Wang

Entity resolution (ER) aims to identify entity records that refer to the same real-world entity, which is a critical problem in data cleaning and integration. Most of the existing models are attribute-centric, that is, matching entity pairs by comparing similarities of pre-aligned attributes, which require the schemas of records to be identical and are too coarse-grained to capture subtle key information within a single attribute. In this paper, we propose a novel graph-based ER model GraphER. Our model is token-centric: the final matching results are generated by directly aggregating token-level comparison features, in which both the semantic and structural information has been softly embedded into token embeddings by training an Entity Record Graph Convolutional Network (ER-GCN). To the best of our knowledge, our work is the first effort to do token-centric entity resolution with the help of GCN in entity resolution task. Extensive experiments on two real-world datasets demonstrate that our model stably outperforms state-of-the-art models.

AAAI Conference 2020 Conference Paper

HAMNER: Headword Amplified Multi-Span Distantly Supervised Method for Domain Specific Named Entity Recognition

  • Shifeng Liu
  • Yifang Sun
  • Bing Li
  • Wei Wang
  • Xiang Zhao

To tackle Named Entity Recognition (NER) tasks, supervised methods need to obtain sufficient cleanly annotated data, which is labor and time consuming. On the contrary, distantly supervised methods acquire automatically annotated data using dictionaries to alleviate this requirement. Unfortunately, dictionaries hinder the effectiveness of distantly supervised methods for NER due to its limited coverage, especially in specific domains. In this paper, we aim at the limitations of the dictionary usage and mention boundary detection. We generalize the distant supervision by extending the dictionary with headword based non-exact matching. We apply a function to better weight the matched entity mentions. We propose a span-level model, which classifies all the possible spans then infers the selected spans with a proposed dynamic programming algorithm. Experiments on all three benchmark datasets demonstrate that our method outperforms previous state-of-the-art distantly supervised methods.

AAAI Conference 2020 Conference Paper

Integrating Linguistic Knowledge to Sentence Paraphrase Generation

  • Zibo Lin
  • Ziran Li
  • Ning Ding
  • Hai-Tao Zheng
  • Ying Shen
  • Wei Wang
  • Cong-Zhi Zhao

Paraphrase generation aims to rewrite a text with different words while keeping the same meaning. Previous work performs the task based solely on the given dataset while ignoring the availability of external linguistic knowledge. However, it is intuitive that a model can generate more expressive and diverse paraphrase with the help of such knowledge. To fill this gap, we propose Knowledge-Enhanced Paraphrase Network (KEPN), a transformer-based framework that can leverage external linguistic knowledge to facilitate paraphrase generation. (1) The model integrates synonym information from the external linguistic knowledge into the paraphrase generator, which is used to guide the decision on whether to generate a new word or replace it with a synonym. (2) To locate the synonym pairs more accurately, we adopt an incremental encoding scheme to incorporate position information of each synonym. Besides, a multi-task architecture is designed to help the framework jointly learn the selection of synonym pairs and the generation of expressive paraphrase. Experimental results on both English and Chinese datasets show that our method significantly outperforms the state-ofthe-art approaches in terms of both automatic and human evaluation.

NeurIPS Conference 2020 Conference Paper

Learning Continuous System Dynamics from Irregularly-Sampled Partial Observations

  • Zijie Huang
  • Yizhou Sun
  • Wei Wang

Many real-world systems, such as moving planets, can be considered as multi-agent dynamic systems, where objects interact with each other and co-evolve along with the time. Such dynamics is usually difficult to capture, and understanding and predicting the dynamics based on observed trajectories of objects become a critical research problem in many domains. Most existing algorithms, however, assume the observations are regularly sampled and all the objects can be fully observed at each sampling time, which is impractical for many applications. In this paper, we pro-pose to learn system dynamics from irregularly-sampled and partial observations with underlying graph structure for the first time. To tackle the above challenge, we present LG-ODE, a latent ordinary differential equation generative model for modeling multi-agent dynamic system with known graph structure. It can simultaneously learn the embedding of high dimensional trajectories and infer continuous latent system dynamics. Our model employs a novel encoder parameterized by a graph neural network that can infer initial states in an unsupervised way from irregularly-sampled partial observations of structural objects and utilizes neuralODE to infer arbitrarily complex continuous-time latent dynamics. Experiments on motion capture, spring system, and charged particle datasets demonstrate the effectiveness of our approach.

AAAI Conference 2020 Conference Paper

Learning-Based Efficient Graph Similarity Computation via Multi-Scale Convolutional Set Matching

  • Yunsheng Bai
  • Hao Ding
  • Ken Gu
  • Yizhou Sun
  • Wei Wang

Graph similarity computation is one of the core operations in many graph-based applications, such as graph similarity search, graph database analysis, graph clustering, etc. Since computing the exact distance/similarity between two graphs is typically NP-hard, a series of approximate methods have been proposed with a trade-off between accuracy and speed. Recently, several data-driven approaches based on neural networks have been proposed, most of which model the graphgraph similarity as the inner product of their graph-level representations, with different techniques proposed for generating one embedding per graph. However, using one fixeddimensional embedding per graph may fail to fully capture graphs in varying sizes and link structures—a limitation that is especially problematic for the task of graph similarity computation, where the goal is to find the fine-grained difference between two graphs. In this paper, we address the problem of graph similarity computation from another perspective, by directly matching two sets of node embeddings without the need to use fixed-dimensional vectors to represent whole graphs for their similarity computation. The model, GRAPH- SIM, achieves the state-of-the-art performance on four realworld graph datasets under six out of eight settings (here we count a specific dataset and metric combination as one setting), compared to existing popular methods for approximate Graph Edit Distance (GED) and Maximum Common Subgraph (MCS) computation.

AAAI Conference 2020 Conference Paper

One-Shot Image Classification by Learning to Restore Prototypes

  • Wanqi Xue
  • Wei Wang

One-shot image classification aims to train image classifiers over the dataset with only one image per category. It is challenging for modern deep neural networks that typically require hundreds or thousands of images per class. In this paper, we adopt metric learning for this problem, which has been applied for few- and many-shot image classification by comparing the distance between the test image and the center of each class in the feature space. However, for one-shot learning, the existing metric learning approaches would suffer poor performance because the single training image may not be representative of the class. For example, if the image is far away from the class center in the feature space, the metriclearning based algorithms are unlikely to make correct predictions for the test images because the decision boundary is shifted by this noisy image. To address this issue, we propose a simple yet effective regression model, denoted by RestoreNet, which learns a class agnostic transformation on the image feature to move the image closer to the class center in the feature space. Experiments demonstrate that RestoreNet obtains superior performance over the state-of-the-art methods on a broad range of datasets. Moreover, RestoreNet can be easily combined with other methods to achieve further improvement.

AAAI Conference 2020 Conference Paper

Pose-Guided Multi-Granularity Attention Network for Text-Based Person Search

  • Ya Jing
  • Chenyang Si
  • Junbo Wang
  • Wei Wang
  • Liang Wang
  • Tieniu Tan

Text-based person search aims to retrieve the corresponding person images in an image database by virtue of a describing sentence about the person, which poses great potential for various applications such as video surveillance. Extracting visual contents corresponding to the human description is the key to this cross-modal matching problem. Moreover, correlated images and descriptions involve different granularities of semantic relevance, which is usually ignored in previous methods. To exploit the multilevel corresponding visual contents, we propose a pose-guided multi-granularity attention network (PMA). Firstly, we propose a coarse alignment network (CA) to select the related image regions to the global description by a similarity-based attention. To further capture the phrase-related visual body part, a fine-grained alignment network (FA) is proposed, which employs pose information to learn latent semantic alignment between visual body part and textual noun phrase. To verify the effectiveness of our model, we perform extensive experiments on the CUHK Person Description Dataset (CUHK-PEDES) which is currently the only available dataset for text-based person search. Experimental results show that our approach outperforms the state-of-the-art methods by 15 % in terms of the top-1 metric.

AAAI Conference 2020 Conference Paper

Recursively Binary Modification Model for Nested Named Entity Recognition

  • Bing Li
  • Shifeng Liu
  • Yifang Sun
  • Wei Wang
  • Xiang Zhao

Recently, there has been an increasing interest in identifying named entities with nested structures. Existing models only make independent typing decisions on the entire entity span while ignoring strong modification relations between subentity types. In this paper, we present a novel Recursively Binary Modification model for nested named entity recognition. Our model utilizes the modification relations among sub-entities types to infer the head component on top of a Bayesian framework and uses entity head as a strong evidence to determine the type of the entity span. The process is recursive, allowing lower-level entities to help better model those on the outer-level. To the best of our knowledge, our work is the first effort that uses modification relation in nested NER task. Extensive experiments on four benchmark datasets demonstrate that our model outperforms state-of-the-art models in nested NER tasks, and delivers competitive results with state-of-the-art models in flat NER task, without relying on any extra annotations or NLP tools.

AAAI Conference 2020 Conference Paper

Revisiting Probability Distribution Assumptions for Information Theoretic Feature Selection

  • Yuan Sun
  • Wei Wang
  • Michael Kirley
  • Xiaodong Li
  • Jeffrey Chan

Feature selection has been shown to be beneficial for many data mining and machine learning tasks, especially for big data analytics. Mutual Information (MI) is a well-known information-theoretic approach used to evaluate the relevance of feature subsets and class labels. However, estimating highdimensional MI poses significant challenges. Consequently, a great deal of research has focused on using low-order MI approximations or computing a lower bound on MI called Variational Information (VI). These methods often require certain assumptions made on the probability distributions of features such that these distributions are realistic yet tractable to compute. In this paper, we reveal two sets of distribution assumptions underlying many MI and VI based methods: Feature Independence Distribution and Geometric Mean Distribution. We systematically analyze their strengths and weaknesses and propose a logical extension called Arithmetic Mean Distribution, which leads to an unbiased and normalised estimation of probability densities. We conduct detailed empirical studies across a suite of 29 real-world classification problems and illustrate improved prediction accuracy of our methods based on the identification of more informative features, thus providing support for our theoretical findings.

AAAI Conference 2020 Conference Paper

RTN: Reparameterized Ternary Network

  • Yuhang Li
  • Xin Dong
  • Sai Qian Zhang
  • Haoli Bai
  • Yuanpeng Chen
  • Wei Wang

To deploy deep neural networks on resource-limited devices, quantization has been widely explored. In this work, we study the extremely low-bit networks which have tremendous speed-up, memory saving with quantized activation and weights. We first bring up three omitted issues in extremely low-bit networks: the squashing range of quantized values; the gradient vanishing during backpropagation and the unexploited hardware acceleration of ternary networks. By reparameterizing quantized activation and weights vector with full precision scale and offset for fixed ternary vector, we decouple the range and magnitude from direction to extenuate above problems. Learnable scale and offset can automatically adjust the range of quantized values and sparsity without gradient vanishing. A novel encoding and computation pattern are designed to support efficient computing for our reparameterized ternary network (RTN). Experiments on ResNet- 18 for ImageNet demonstrate that the proposed RTN finds a much better efficiency between bitwidth and accuracy and achieves up to 26. 76% relative accuracy improvement compared with state-of-the-art methods. Moreover, we validate the proposed computation pattern on Field Programmable Gate Arrays (FPGA), and it brings 46. 46× and 89. 17× savings on power and area compared with the full precision convolution.

NeurIPS Conference 2020 Conference Paper

Semi-Supervised Partial Label Learning via Confidence-Rated Margin Maximization

  • Wei Wang
  • Min-Ling Zhang

Partial label learning assumes inaccurate supervision where each training example is associated with a set of candidate labels, among which only one is valid. In many real-world scenarios, however, it is costly and time-consuming to assign candidate label sets to all the training examples. To circumvent this difficulty, the problem of semi-supervised partial label learning is investigated in this paper, where unlabeled data is utilized to facilitate model induction along with partial label training examples. Specifically, label propagation is adopted to instantiate the labeling confidence of partial label examples. After that, maximum margin formulation is introduced to jointly enable the induction of predictive model and the estimation of labeling confidence over unlabeled data. The derived formulation enforces confidence-rated margin maximization and confidence manifold preservation over partial label examples and unlabeled data. We show that the predictive model and labeling confidence can be solved via alternating optimization which admits QP solutions in either alternating step. Extensive experiments on synthetic as well as real-world data sets clearly validate the effectiveness of the proposed semi-supervised partial label learning approach.

AAAI Conference 2020 Conference Paper

Time2Graph: Revisiting Time Series Modeling with Dynamic Shapelets

  • Ziqiang Cheng
  • Yang Yang
  • Wei Wang
  • Wenjie Hu
  • Yueting Zhuang
  • Guojie Song

Time series modeling has attracted extensive research efforts; however, achieving both reliable efficiency and interpretability from a unified model still remains a challenging problem. Among the literature, shapelets offer interpretable and explanatory insights in the classification tasks, while most existing works ignore the differing representative power at different time slices, as well as (more importantly) the evolution pattern of shapelets. In this paper, we propose to extract time-aware shapelets by designing a two-level timing factor. Moreover, we define and construct the shapelet evolution graph, which captures how shapelets evolve over time and can be incorporated into the time series embeddings by graph embedding algorithms. To validate whether the representations obtained in this way can be applied effectively in various scenarios, we conduct experiments based on three public time series datasets, and two real-world datasets from different domains. Experimental results clearly show the improvements achieved by our approach compared with 16 state-of-the-art baselines.

AAAI Conference 2019 Conference Paper

A Deep Cascade Model for Multi-Document Reading Comprehension

  • Ming Yan
  • Jiangnan Xia
  • Chen Wu
  • Bin Bi
  • Zhongzhou Zhao
  • Ji Zhang
  • Luo Si
  • Rui Wang

A fundamental trade-off between effectiveness and efficiency needs to be balanced when designing an online question answering system. Effectiveness comes from sophisticated functions such as extractive machine reading comprehension (MRC), while efficiency is obtained from improvements in preliminary retrieval components such as candidate document selection and paragraph ranking. Given the complexity of the real-world multi-document MRC scenario, it is difficult to jointly optimize both in an end-to-end system. To address this problem, we develop a novel deep cascade learning model, which progressively evolves from the documentlevel and paragraph-level ranking of candidate texts to more precise answer extraction with machine reading comprehension. Specifically, irrelevant documents and paragraphs are first filtered out with simple functions for efficiency consideration. Then we jointly train three modules on the remaining texts for better tracking the answer: the document extraction, the paragraph extraction and the answer extraction. Experiment results show that the proposed method outperforms the previous state-of-the-art methods on two large-scale multidocument benchmark datasets, i. e. , TriviaQA and DuReader. In addition, our online system can stably serve typical scenarios with millions of daily requests in less than 50ms.

TIST Journal 2019 Journal Article

A Survey of Zero-Shot Learning

  • Wei Wang
  • Vincent W. Zheng
  • Han Yu
  • Chunyan Miao

Most machine-learning methods focus on classifying instances whose classes have already been seen in training. In practice, many applications require classifying instances whose classes have not been seen previously. Zero-shot learning is a powerful and promising learning paradigm, in which the classes covered by training instances and the classes we aim to classify are disjoint. In this paper, we provide a comprehensive survey of zero-shot learning. First of all, we provide an overview of zero-shot learning. According to the data utilized in model optimization, we classify zero-shot learning into three learning settings. Second, we describe different semantic spaces adopted in existing zero-shot learning works. Third, we categorize existing zero-shot learning methods and introduce representative methods under each category. Fourth, we discuss different applications of zero-shot learning. Finally, we highlight promising future research directions of zero-shot learning.

AAAI Conference 2019 Conference Paper

Adversarial Training for Community Question Answer Selection Based on Multi-Scale Matching

  • Xiao Yang
  • Madian Khabsa
  • Miaosen Wang
  • Wei Wang
  • Ahmed Hassan Awadallah
  • Daniel Kifer
  • C. Lee Giles

Community-based question answering (CQA) websites represent an important source of information. As a result, the problem of matching the most valuable answers to their corresponding questions has become an increasingly popular research topic. We frame this task as a binary (relevant/irrelevant) classification problem, and present an adversarial training framework to alleviate label imbalance issue. We employ a generative model to iteratively sample a subset of challenging negative samples to fool our classification model. Both models are alternatively optimized using REIN- FORCE algorithm. The proposed method is completely different from previous ones, where negative samples in training set are directly used or uniformly down-sampled. Further, we propose using Multi-scale Matching which explicitly inspects the correlation between words and ngrams of different levels of granularity. We evaluate the proposed method on SemEval 2016 and SemEval 2017 datasets and achieves state-of-the-art or similar performance.

AAAI Conference 2019 Conference Paper

Antonym-Synonym Classification Based on New Sub-Space Embeddings

  • Muhammad Asif Ali
  • Yifang Sun
  • Xiaoling Zhou
  • Wei Wang
  • Xiang Zhao

Distinguishing antonyms from synonyms is a key challenge for many NLP applications focused on the lexical-semantic relation extraction. Existing solutions relying on large-scale corpora yield low performance because of huge contextual overlap of antonym and synonym pairs. We propose a novel approach entirely based on pre-trained embeddings. We hypothesize that the pre-trained embeddings comprehend a blend of lexical-semantic information and we may distill the task-specific information using Distiller, a model proposed in this paper. Later, a classifier is trained based on features constructed from the distilled sub-spaces along with some word level features to distinguish antonyms from synonyms. Experimental results show that the proposed model outperforms existing research on antonym synonym distinction in both speed and performance.

NeurIPS Conference 2019 Conference Paper

Backpropagation-Friendly Eigendecomposition

  • Wei Wang
  • Zheng Dang
  • Yinlin Hu
  • Pascal Fua
  • Mathieu Salzmann

Eigendecomposition (ED) is widely used in deep networks. However, the backpropagation of its results tends to be numerically unstable, whether using ED directly or approximating it with the Power Iteration method, particularly when dealing with large matrices. While this can be mitigated by partitioning the data in small and arbitrary groups, doing so has no theoretical basis and makes its impossible to exploit the power of ED to the full. In this paper, we introduce a numerically stable and differentiable approach to leveraging eigenvectors in deep networks. It can handle large matrices without requiring to split them. We demonstrate the better robustness of our approach over standard ED and PI for ZCA whitening, an alternative to batch normalization, and for PCA denoising, which we introduce as a new normalization strategy for deep networks, aiming to further denoise the network's features.

YNICL Journal 2019 Journal Article

Changes in default mode network connectivity in different glucose metabolism status and diabetes duration

  • Huanghui Liu
  • Jun Liu
  • Limin Peng
  • Zhichao Feng
  • Lu Cao
  • Huasheng Liu
  • Hui Shen
  • Dewen Hu

AIMS/HYPOTHESES: It is now generally accepted that diabetes increases the risk for cognitive impairment, but the precise mechanisms are poorly understood. In recent years, resting-state functional magnetic resonance imaging (rs-fMRI) is increasingly used to investigate the neural basis of cognitive dysfunction in type 2 diabetes (T2D) patients. Alterations in brain functional connectivity may underlie diabetes-related cognitive dysfunction and brain damage. The aim of this study was to investigate the changes in default mode network (DMN) connectivity in different glucose metabolism status and diabetes duration. METHODS: We used a seed-based fMRI analysis to investigate positive and negative DMN connectivity in four groups (39 subjects with normal glucose metabolism [NGM], 23 subjects with impaired glucose metabolism [IGM; i.e., prediabetes], 59 T2D patients with a diabetes duration of <10 years, and 24 T2D patients with a diabetes duration of ≥10 years). RESULTS: Negative DMN connectivity increased and then regressed with deteriorating glucose metabolism status and extending diabetes duration. DMN connectivity showed a significant correlation with diabetes duration. CONCLUSION/INTERPRETATION: This study suggests that DMN connectivity may exhibit distinct patterns in different glucose metabolism status and diabetes duration, providing some potential neuroimaging evidence for early diagnosis and further understanding of the pathophysiological mechanisms of diabetic brain damage.

IROS Conference 2019 Conference Paper

Concept and Validation of a Large-scale Human-machine Safety System Based on Real-time UWB Indoor Localization *

  • Wei Wang
  • Zhuoqi Zeng
  • Wan Ding
  • Huajun Yu
  • Hannes Rose

In production line, the conventional industrial robots and automatic machines require machinery safety protection to guarantee the safety of human operators. A scalable and easy-to-conFigure safety system concept called “Real-time Safety Virtual Positioning” (RSVP) is proposed, which could act as potentially key enabler for agile production systems by eliminating fixed safety installation and thus increasing productivity and flexibility. The RSVP provides easy access to robots of automatic assembly lines in plants, e. g. , automotive OEM, and supports virtualization and transparent to fully automatic and semi-automatic assembly line by knowing the position of persons and relevant objects (e. g. , tools, finished and/or semi-finished goods, and materials). The focus of the paper will discuss the functional safety certification realization (concept approved by TÜV (Technical Inspection Association)) and validation details of indoor localization-based safety system developments. The detailed strategy of the functional safety requirements, danger diagnosis and reaction approach, communication among safety controller, robot and machine, and safe failure reaction are listed and analyzed. The physical hardware framework and software architecture of the safety system are built and developed with the 1oo2-architecture according to the Performance Levels (PL d) of ISO 13849 and the Safety Integrity Levels (SIL 2) of IEC 61508. The implicated algorithms and data process of the UWB-based (Ultra-Wide Band) indoor localization system are introduced. The safety system concept is validated and verified in an ABS (Anti-lock Braking System) production line with human-robot co-existence environment.

AAAI Conference 2019 Conference Paper

Coreset Stochastic Variance-Reduced Gradient with Application to Optimal Margin Distribution Machine

  • Zhi-Hao Tan
  • Teng Zhang
  • Wei Wang

A major problem for kernel-based predictors is the prohibitive computational complexity, which limits their application in large-scale datasets. Coreset, an approximation method which tries to cover the given examples with a small set of points, can be used to remain the prominent information and accelerate the kernel method. In this paper, we provide perhaps the first coreset-based kernel-accelerating optimization method that has a linear convergence rate, which is much faster than existing approaches. Our method can be used to train kernel SVM-style problems and obtain sparse solutions efficiently. Specifically, the method uses SVRG as the framework, and utilizes the core points to approximate the gradients, so it can significantly reduce the complexity of the kernel method. Furthermore, we apply the method to train ODM, a kernel machine enjoying better statistical property than SVM, so that we can reduce the risk of compromising the performance while encouraging the sparsity. We conduct extensive experiments on several large-scale datasets and the results verify that our method outperforms the state-of-the-art coreset approximation method in both efficiency and generalization, while simultaneously achieving significant speed-up compared to non-approximation baselines.

AAAI Conference 2019 Conference Paper

Jointly Extracting Multiple Triplets with Multilayer Translation Constraints

  • Zhen Tan
  • Xiang Zhao
  • Wei Wang
  • Weidong Xiao

Triplets extraction is an essential and pivotal step in automatic knowledge base construction, which captures structural information from unstructured text corpus. Conventional extraction models use a pipeline of named entity recognition and relation classification to extract entities and relations, respectively, which ignore the connection between the two tasks. Recently, several neural network-based models were proposed to tackle the problem, and achieved state-of-the-art performance. However, most of them are unable to extract multiple triplets from a single sentence, which are yet commonly seen in real-life scenarios. To close the gap, we propose in this paper a joint neural extraction model for multitriplets, namely, TME, which is capable of adaptively discovering multiple triplets simultaneously in a sentence via ranking with translation mechanism. In experiment, TME exhibits superior performance and achieves an improvement of 37. 6% on F1 score over state-of-the-art competitors.

IJCAI Conference 2019 Conference Paper

Learn Smart with Less: Building Better Online Decision Trees with Fewer Training Examples

  • Ariyam Das
  • Jin Wang
  • Sahil M. Gandhi
  • Jae Lee
  • Wei Wang
  • Carlo Zaniolo

Online decision tree models are extensively used in many industrial machine learning applications for real-time classification tasks. These models are highly accurate, scalable and easy to use in practice. The Very Fast Decision Tree (VFDT) is the classic online decision tree induction model that has been widely adopted due to its theoretical guarantees as well as competitive performance. However, VFDT and its variants solely rely on conservative statistical measures like Hoeffding bound to incrementally grow the tree. This makes these models extremely circumspect and limits their ability to learn fast. In this paper, we efficiently employ statistical resampling techniques to build an online tree faster using fewer examples. We first theoretically show that a naive implementation of resampling techniques like non-parametric bootstrap does not scale due to large memory and computational overheads. We mitigate this by proposing a robust memory-efficient bootstrap simulation heuristic (Mem-ES) that successfully expedites the learning process. Experimental results on both synthetic data and large-scale real world datasets demonstrate the efficiency and effectiveness of our proposed technique.

IJCAI Conference 2019 Conference Paper

Model-Agnostic Adversarial Detection by Random Perturbations

  • Bo Huang
  • Yi Wang
  • Wei Wang

Adversarial examples induce model classification errors on purpose, which has raised concerns on the security aspect of machine learning techniques. Many existing countermeasures are compromised by adaptive adversaries and transferred examples. We propose a model-agnostic approach to resolve the problem by analysing the model responses to an input under random perturbations, and study the robustness of detecting norm-bounded adversarial distortions in a theoretical framework. Extensive evaluations are performed on the MNIST, CIFAR-10 and ImageNet datasets. The results demonstrate that our detection method is effective and resilient against various attacks including black-box attacks and the powerful CW attack with four adversarial adaptations.

AAAI Conference 2019 Conference Paper

Personalized Question Routing via Heterogeneous Network Embedding

  • Zeyu Li
  • Jyun-Yu Jiang
  • Yizhou Sun
  • Wei Wang

Question Routing (QR) on Community-based Question Answering (CQA) websites aims at recommending answerers that have high probabilities of providing the “accepted answers” to new questions. The existing question routing algorithms simply predict the ranking of users based on query content. As a consequence, the question raiser information is ignored. On the other hand, they lack learnable scoring functions to explicitly compute ranking scores. To tackle these challenges, we propose NeRank that (1) jointly learns representations of question content, question raiser, and question answerers by a heterogeneous information network embedding algorithm and a long short-term memory (LSTM) model. The embeddings of the three types of entities are unified in the same latent space, and (2) conducts question routing for personalized queries, i. e. , queries with two entities (question content, question raiser), by a convolutional scoring function taking the learned embeddings of all three types of entities as input. Using the scores, NeRank routes new questions to high-ranking answerers that are skillfulness in the question domain and have similar backgrounds to the question raiser. Experimental results show that NeRank significantly outperforms competitive baseline question routing models that ignore the raiser information in three ranking metrics. In addition, NeRank is convergeable in several thousand iterations and insensitive to parameter changes, which prove its effectiveness, scalability, and robustness.

IJCAI Conference 2019 Conference Paper

Relation Extraction Using Supervision from Topic Knowledge of Relation Labels

  • Haiyun Jiang
  • Li Cui
  • Zhe Xu
  • Deqing Yang
  • Jindong Chen
  • Chenguang Li
  • Jingping Liu
  • Jiaqing Liang

Explicitly exploring the semantics of a relation is significant for high-accuracy relation extraction, which is, however, not fully studied in previous work. In this paper, we mine the topic knowledge of a relation to explicitly represent the semantics of this relation, and model relation extraction as a matching problem. That is, the matching score between a sentence and a candidate relation is predicted for an entity pair. To this end, we propose a deep matching network to precisely model the semantic similarity between a sentence-relation pair. Besides, the topic knowledge also allows us to derive the importance information of samples as well as two knowledge-guided negative sampling strategies in the training process. We conduct extensive experiments to evaluate the proposed framework and observe improvements in AUC of 11. 5% and max F1 of 5. 4% over the baselines with state-of-the-art performance.

IROS Conference 2019 Conference Paper

TerrainFusion: Real-time Digital Surface Model Reconstruction based on Monocular SLAM

  • Wei Wang
  • Yong Zhao
  • Pengcheng Han
  • Pengcheng Zhao
  • Shuhui Bu

This paper presents an algorithm which can generate live digtial surface model (DSM) during the flight based on simultaneous localization and mapping (SLAM). We process the keyframe which is output by a monocular SLAM system to generate a local DSM, and fuse the local DSM to the global tiled DSM incrementally. During the local DSM generation, a local digital elevation model (DEM) is estimated by projecting the filtered 2D Delaunay mesh to a 3D mesh, and a local orthomosaic is obtained by projecting triangle image patches onto a 2D mesh. During the DSM fusion, both the local DEM and orthomosaic are split into tiles and fused to the global tiled DEM and orthomosaic respectively with multiband algorithm. Both the efficient DSM generation and fusion algorithms contribute to achieving a real-time reconstruction. Qualitative and quantitative experiments on a public aerial image dataset with different scenarios are performed to validate the effectiveness of the proposed method. Compared with traditional structure from motion (SfM) based approaches, the presented system is able to output both large-scale high-quality DEM and orthomosaic in real-time with low computational cost.

IJCAI Conference 2019 Conference Paper

Unsupervised Inductive Graph-Level Representation Learning via Graph-Graph Proximity

  • Yunsheng Bai
  • Hao Ding
  • Yang Qiao
  • Agustin Marinovic
  • Ken Gu
  • Ting Chen
  • Yizhou Sun
  • Wei Wang

We introduce a novel approach to graph-level representation learning, which is to embed an entire graph into a vector space where the embeddings of two graphs preserve their graph-graph proximity. Our approach, UGraphEmb, is a general framework that provides a novel means to performing graph-level embedding in a completely unsupervised and inductive manner. The learned neural network can be considered as a function that receives any graph as input, either seen or unseen in the training set, and transforms it into an embedding. A novel graph-level embedding generation mechanism called Multi-Scale Node Attention (MSNA), is proposed. Experiments on five real graph datasets show that UGraphEmb achieves competitive accuracy in the tasks of graph classification, similarity ranking, and graph visualization.

TIST Journal 2019 Journal Article

Using Sparse Representation to Detect Anomalies in Complex WSNs

  • Xiaoming Li
  • Guangquan Xu
  • Xi Zheng
  • Kaitai Liang
  • Emmanouil Panaousis
  • Tao Li
  • Wei Wang
  • Chao Shen

In recent years, wireless sensor networks (WSNs) have become an active area of research for monitoring physical and environmental conditions. Due to the interdependence of sensors, a functional anomaly in one sensor can cause a functional anomaly in another sensor, which can further lead to the malfunctioning of the entire sensor network. Existing research work has analysed faulty sensor anomalies but fails to show the effectiveness throughout the entire interdependent network system. In this article, a dictionary learning algorithm based on a non-negative constraint is developed, and a sparse representation anomaly node detection method for sensor networks is proposed based on the dictionary learning. Through experiment on a specific thermal power plant in China, we verify the robustness of our proposed method in detecting abnormal nodes against four state of the art approaches and proved our method is more robust. Furthermore, the experiments are conducted on the obtained abnormal nodes to prove the interdependence of multi-layer sensor networks and reveal the conditions and causes of a system crash.

IJCAI Conference 2018 Conference Paper

An Adaptive Hierarchical Compositional Model for Phrase Embedding

  • Bing Li
  • Xiaochun Yang
  • Bin Wang
  • Wei Wang
  • Wei Cui
  • Xianchao Zhang

Phrase embedding aims at representing phrases in a vector space and it is important for the performance of many NLP tasks. Existing models only regard a phrase as either full-compositional or non-compositional, while ignoring the hybrid-compositionality that widely exists, especially in long phrases. This drawback prevents them from having a deeper insight into the semantic structure for long phrases and as a consequence, weakens the accuracy of the embeddings. In this paper, we present a novel method for jointly learning compositionality and phrase embedding by adaptively weighting different compositions using an implicit hierarchical structure. Our model has the ability of adaptively adjusting among different compositions without entailing too much model complexity and time cost. To the best of our knowledge, our work is the first effort that considers hybrid-compositionality in phrase embedding. The experimental evaluation demonstrates that our model outperforms state-of-the-art methods in both similarity tasks and analogy tasks.

JBHI Journal 2018 Journal Article

Assessment of Gait Characteristics in Total Knee Arthroplasty Patients Using a Hierarchical Partial Least Squares Method

  • Wei Wang
  • David C. Ackland
  • Jodie A. McClelland
  • Kate E. Webster
  • Saman Halgamuge

Quantitative gait analysis is an important tool in objective assessment and management of total knee arthroplasty (TKA) patients. Studies evaluating gait patterns in TKA patients have tended to focus on discrete data such as spatiotemporal information, joint range of motion and peak values of kinematics and kinetics, or consider selected principal components of gait waveforms for analysis. These strategies may not have the capacity to capture small variations in gait patterns associated with each joint across an entire gait cycle, and may ultimately limit the accuracy of gait classification. The aim of this study was to develop an automatic feature extraction method to analyse patterns from high-dimensional autocorrelated gait waveforms. A general linear feature extraction framework was proposed and a hierarchical partial least squares method derived for discriminant analysis of multiple gait waveforms. The effectiveness of this strategy was verified using a dataset of joint angle and ground reaction force waveforms from 43 patients after TKA surgery and 31 healthy control subjects. Compared with principal component analysis and partial least squares methods, the hierarchical partial least squares method achieved generally better classification performance on all possible combinations of waveforms, with the highest classification accuracy 85. 14%. The novel hierarchical partial least squares method proposed is capable of capturing virtually all significant differences between TKA patients and the controls, and provides new insights into data visualization. The proposed framework presents a foundation for more rigorous classification of gait, and may ultimately be used to evaluate the effects of interventions such as surgery and rehabilitation.

AAAI Conference 2018 Conference Paper

Information-Theoretic Domain Adaptation Under Severe Noise Conditions

  • Wei Wang
  • Hao Wang
  • Zhi-Yong Ran
  • Ran He

Cross-domain data reconstruction methods derive a shared transformation across source and target domains. These methods usually make a specific assumption on noise, which exhibits limited ability when the target data are contaminated by different kinds of complex noise in practice. To enhance the robustness of domain adaptation under severe noise conditions, this paper proposes a novel reconstruction based algorithm in an information-theoretic setting. Specifically, benefiting from the theoretical property of correntropy, the proposed algorithm is distinguished with: detecting the contaminated target samples without making any specific assumption on noise; greatly suppressing the negative influence of noise on cross-domain transformation. Moreover, a relative entropy based regularization of the transformation is incorporated to avoid trivial solutions with the reaped theoretic advantages, i. e. , non-negativity and scale-invariance. For optimization, a half-quadratic technique is developed to minimize the nonconvex information-theoretic objectives with explicitly guaranteed convergence. Experiments on two real-world domain adaptation tasks demonstrate the superiority of our method.

IJCAI Conference 2018 Conference Paper

Tri-net for Semi-Supervised Deep Learning

  • Dong-Dong Chen
  • Wei Wang
  • Wei Gao
  • Zhi-Hua Zhou

Deep neural networks have witnessed great successes in various real applications, but it requires a large number of labeled data for training. In this paper, we propose tri-net, a deep neural network which is able to use massive unlabeled data to help learning with limited labeled data. We consider model initialization, diversity augmentation and pseudo-label editing simultaneously. In our work, we utilize output smearing to initialize modules, use fine-tuning on labeled data to augment diversity and eliminate unstable pseudo-labels to alleviate the influence of suspicious pseudo-labeled data. Experiments show that our method achieves the best performance in comparison with state-of-the-art semi-supervised deep learning methods. In particular, it achieves 8. 30% error rate on CIFAR-10 by using only 4000 labeled examples.

IJCAI Conference 2017 Conference Paper

Entity Suggestion with Conceptual Expanation

  • Yi Zhang
  • Yanghua Xiao
  • Seung-won Hwang
  • Haixun Wang
  • X. Sean Wang
  • Wei Wang

Entity Suggestion with Conceptual Explanation (ESC) refers to a type of entity acquisition query in which a user provides a set of example entities as the query and obtains in return not only some related entities but also concepts which can best explain the query and the result. ESC is useful in many applications such as related-entity recommendation and query expansion. Many example based entity suggestion solutions are available in existing literatures. However, they are generally not aware of the concepts of query entities thus cannot be used for conceptual explanation. In this paper, we propose two probabilistic entity suggestion models and their computation solutions. Our models and solutions fully take advantage of the large scale taxonomies which consist of isA relations between entities and concepts. With our models and solutions, we can not only find the best entities to suggest but also derive the best concepts to explain the suggestion. Extensive evaluations on real data sets justify the accuracy of our models and the efficiency of our solutions.

AAAI Conference 2017 Conference Paper

Fredholm Multiple Kernel Learning for Semi-Supervised Domain Adaptation

  • Wei Wang
  • Hao Wang
  • Chen Zhang
  • Yang Gao

As a fundamental constituent of machine learning, domain adaptation generalizes a learning model from a source domain to a different (but related) target domain. In this paper, we focus on semi-supervised domain adaptation and explicitly extend the applied range of unlabeled target samples into the combination of distribution alignment and adaptive classifier learning. Specifically, our extension formulates the following aspects in a single optimization: 1) learning a crossdomain predictive model by developing the Fredholm integral based kernel prediction framework; 2) reducing the distribution difference between two domains; 3) exploring multiple kernels to induce an optimal learning space. Correspondingly, such an extension is distinguished with allowing for noise resiliency, facilitating knowledge transfer and analyzing diverse data characteristics. It is emphasized that we prove the differentiability of our formulation and present an effective optimization procedure based on the reduced gradient, guaranteeing rapid convergence. Comprehensive empirical studies verify the effectiveness of the proposed method.

IJCAI Conference 2017 Conference Paper

Link Prediction with Spatial and Temporal Consistency in Dynamic Networks

  • Wenchao Yu
  • Wei Cheng
  • Charu C Aggarwal
  • Haifeng Chen
  • Wei Wang

Dynamic networks are ubiquitous. Link prediction in dynamic networks has attracted tremendous research interests. Many models have been developed to predict links that may emerge in the immediate future from the past evolution of the networks. There are two key factors: 1) a node is more likely to form a link in the near future with another node within its close proximity, rather than with a random node; 2) a dynamic network usually evolves smoothly. Existing approaches seldom unify these two factors to strive for the spatial and temporal consistency in a dynamic network. To address this limitation, in this paper, we propose a link prediction model with spatial and temporal consistency (LIST), to predict links in a sequence of networks over time. LIST characterizes the network dynamics as a function of time, which integrates the spatial topology of network at each timestamp and the temporal network evolution. Comparing to existing approaches, LIST has two advantages: 1) LIST uses a generic model to express the network structure as a function of time, which makes it also suitable for a wide variety of temporal network analysis problems beyond the focus of this paper; 2) by retaining the spatial and temporal consistency, LIST yields better prediction performance. Extensive experiments on four real datasets demonstrate the effectiveness of the LIST model.

IJCAI Conference 2017 Conference Paper

Modeling Trajectories with Recurrent Neural Networks

  • Hao Wu
  • Ziyang Chen
  • Weiwei Sun
  • Baihua Zheng
  • Wei Wang

Modeling trajectory data is a building block for many smart-mobility initiatives. Existing approaches apply shallow models such as Markov chain and inverse reinforcement learning to model trajectories, which cannot capture the long-term dependencies. On the other hand, deep models such as Recurrent Neural Network (RNN) have demonstrated their strength of modeling variable length sequences. However, directly adopting RNN to model trajectories is not appropriate because of the unique topological constraints faced by trajectories. Motivated by these findings, we design two RNN-based models which can make full advantage of the strength of RNN to capture variable length sequence and meanwhile to address the constraints of topological structure on trajectory modeling. Our experimental study based on real taxi trajectory datasets shows that both of our approaches largely outperform the existing approaches.

IJCAI Conference 2017 Conference Paper

Obtaining High-Quality Label by Distinguishing between Easy and Hard Items in Crowdsourcing

  • Wei Wang
  • Xiang-Yu Guo
  • Shao-Yuan Li
  • Yuan Jiang
  • Zhi-Hua Zhou

Crowdsourcing systems make it possible to hire voluntary workers to label large-scale data by offering them small monetary payments. Usually, the taskmaster requires to collect high-quality labels, while the quality of labels obtained from the crowd may not satisfy this requirement. In this paper, we study the problem of obtaining high-quality labels from the crowd and present an approach of learning the difficulty of items in crowdsourcing, in which we construct a small training set of items with estimated difficulty and then learn a model to predict the difficulty of future items. With the predicted difficulty, we can distinguish between easy and hard items to obtain high-quality labels. For easy items, the quality of their labels inferred from the crowd could be high enough to satisfy the requirement; while for hard items, the crowd could not provide high-quality labels, it is better to choose a more knowledgable crowd or employ specialized workers to label them. The experimental results demonstrate that the proposed approach by learning to distinguish between easy and hard items can significantly improve the label quality.

AAAI Conference 2017 Conference Paper

On the Transitivity of Hypernym-Hyponym Relations in Data-Driven Lexical Taxonomies

  • Jiaqing Liang
  • Yi Zhang
  • Yanghua Xiao
  • Haixun Wang
  • Wei Wang
  • Pinpin Zhu

Taxonomy is indispensable in understanding natural language. A variety of large scale, usage-based, data-driven lexical taxonomies have been constructed in recent years. Hypernym-hyponym relationship, which is considered as the backbone of lexical taxonomies can not only be used to categorize the data but also enables generalization. In particular, we focus on one of the most prominent properties of the hypernym-hyponym relationship, namely, transitivity, which has a significant implication for many applications. We show that, unlike human crafted ontologies and taxonomies, transitivity does not always hold in data-driven lexical taxonomies. We introduce a supervised approach to detect whether transitivity holds for any given pair of hypernym-hyponym relationships. Besides solving the inferencing problem, we also use the transitivity to derive new hypernym-hyponym relationships for data-driven lexical taxonomies. We conduct extensive experiments to show the effectiveness of our approach.

TIST Journal 2017 Journal Article

UMCR

  • Hao Yin
  • Wei Wang
  • Xu Zhang
  • Yongqiang Lyu
  • Geyong Min
  • Dongchao Guo

Although mobile application ecosystems have experienced tremendous growth in recent years, retrieving content of mobile applications that serves a key to mobile content search engines still faces grand challenges. Compared to web content retrieval, it is much more difficult to capture content in mobile applications due to the diversity of applications and the lack of Uniform Resource Locator indices. In this study, we propose and implement a <underline>u</underline>ser interaction-driven <underline>m</underline>obile <underline>c</underline>ontent <underline>r</underline>etrieval (UMCR) system to address such issues, which is the first mobile content crawler in the current literature. UMCR is a distributed system that contains many measurement nodes, each of which combines the user interaction path traversing (UIPT) and Deep Package Inspection (DPI) together to obtain mobile content. UIPT determines the events of user interactions in various applications to capture the static content such as text and images, in which a traversal depth termination scheme and an optional cut-off component are adopted to balance the content coverage and traversing efficiency. Meanwhile, the analysis based on DPI is responsible for extracting the videos as well as digging the infrastructural information and performance metrics. In addition, a distributed traversal scheduling method is designed for UIPT tasks to improve the throughput and scalability in large-scale content retrieval. Experiments on retrieving content of 64 real mobile applications demonstrate that UMCR can handle diverse mobile applications efficiently. The scheduler can improve throughput by 3 times compared to the legacy arbitrary task assignment strategy.

JBHI Journal 2016 Journal Article

An Adaptive Filter for the Removal of Drifting Sinusoidal Noise Without a Reference

  • John W. Kelly
  • Daniel P. Siewiorek
  • Asim Smailagic
  • Wei Wang

This paper presents a method for filtering sinusoidal noise with a variable bandwidth filter that is capable of tracking a sinusoid's drifting frequency. The method, which is based on the adaptive noise canceling (ANC) technique, will be referred to here as the adaptive sinusoid canceler (ASC). The ASC eliminates sinusoidal contamination by tracking its frequency and achieving a narrower bandwidth than typical notch filters. The detected frequency is used to digitally generate an internal reference instead of relying on an external one as ANC filters typically do. The filter's bandwidth adjusts to achieve faster and more accurate convergence. In this paper, the focus of the discussion and the data is physiological signals, specifically electrocorticographic (ECoG) neural data contaminated with power line noise, but the presented technique could be applicable to other recordings as well. On simulated data, the ASC was able to reliably track the noise's frequency, properly adjust its bandwidth, and outperform comparative methods including standard notch filters and an adaptive line enhancer. These results were reinforced by visual results obtained from real ECoG data. The ASC showed that it could be an effective method for increasing signal to noise ratio in the presence of drifting sinusoidal noise, which is of significant interest for biomedical applications.

IJCAI Conference 2016 Conference Paper

KBQA: An Online Template Based Question Answering System over Freebase

  • Wanyun Cui
  • Yanghua Xiao
  • Wei Wang

Question answering (QA) has become a popular way for humans to access billion-scale knowledge bases. QA systems over knowledge bases produce accurate and concise answers. The key of QA over knowledge bases is to map the question to a certain substructure in the knowledge base. To do this, KBQA (Question Answering over Knowledge Bases) uses a new kind of question representation: templates, learned from a million scale QA corpora. For example, for questions about a city's population, KBQA learns templates such as What's the population of $city? , How many people are there in $city? . It learns overall 1171303 templates for 4690 relations. Based on these templates, KBQA effectively and efficiently supports binary factoid questions or complex questions.

IJCAI Conference 2016 Conference Paper

Learning Defining Features for Categories

  • Bo Xu
  • Chenhao Xie
  • Yi Zhang
  • Yanghua Xiao
  • Haixun Wang
  • Wei Wang

Categories play a fundamental role in human cognition. Defining features (short for DFs) are the key elements to define a category, which enables machines to categorize objects. Categories enriched with their DFs significantly improve the machine's ability of categorization and benefit many applications built upon categorization. However, defining features can rarely be found for categories in current knowledge bases. Traditional efforts such as manual construction by domain experts are not practical to find defining features for millions of categories. In this paper, we make the first attempt to automatically find defining features for millions of categories in the real world. We formalize the defining feature learning problem and propose a bootstrapping solution to learn defining features from the features of entities belonging to a category. Experimental results show the effectiveness and efficiency of our method. Finally, we find defining features for overall 60, 247 categories with acceptable accuracy.

IJCAI Conference 2016 Conference Paper

Makeup Like a Superstar: Deep Localized Makeup Transfer Network

  • Si Liu
  • Xinyu Ou
  • Ruihe Qian
  • Wei Wang
  • Xiaochun Cao

In this paper, we propose a novel Deep Localized Makeup Transfer Network to automatically recommend the most suitable makeup for a female and synthesis the makeup on her face. Given a before-makeup face, her most suitable makeup is determined automatically. Then, both the before makeup and the reference faces are fed into the proposed Deep Transfer Network to generate the after-makeup face. Our end-to-end makeup transfer network have several nice properties including: (1) with complete functions: including foundation, lip gloss, and eye shadow transfer; (2) cosmetic specific: different cosmetics are transferred in different manners; (3) localized: different cosmetics are applied on different facial regions; (4) producing naturally looking results without obvious artifacts; (5) controllable makeup lightness: various results from light makeup to heavy makeup can be generated. Qualitative and quantitative experiments show that our network performs much better than the methods of [Guo and Sim, 2009] and two variants of NerualStyle [Gatys et al. , 2015a].

AAAI Conference 2016 Conference Paper

Verb Pattern: A Probabilistic Semantic Representation on Verbs

  • Wanyun Cui
  • Xiyou Zhou
  • Hangyu Lin
  • Yanghua Xiao
  • Haixun Wang
  • Seung-won Hwang
  • Wei Wang

Verbs are important in semantic understanding of natural language. Traditional verb representations, such as FrameNet, PropBank, VerbNet, focus on verbs’ roles. These roles are too coarse to represent verbs’ semantics. In this paper, we introduce verb patterns to represent verbs’ semantics, such that each pattern corresponds to a single semantic of the verb. First we analyze the principles for verb patterns: generality and specificity. Then we propose a nonparametric model based on description length. Experimental results prove the high effectiveness of verb patterns. We further apply verb patterns to context-aware conceptualization, to show that verb patterns are helpful in semantic-related tasks.

UAI Conference 2015 Conference Paper

A Smart-Dumb/Dumb-Smart Algorithm for Efficient Split-Merge MCMC

  • Wei Wang
  • Stuart Russell 0001

Split-merge moves are a standard component of MCMC algorithms for tasks such as multitarget tracking and fitting mixture models with unknown numbers of components. Achieving rapid mixing for split-merge MCMC has been notoriously difficult, and state-of-the-art methods do not scale well. We explore the reasons for this and propose a new split-merge kernel consisting of two sub-kernels: one combines a “smart” split move that proposes plausible splits of heterogeneous clusters with a “dumb” merge move that proposes merging random pairs of clusters; the other combines a dumb split move with a smart merge move. We show that the resulting smart-dumb/dumb-smart (SDDS) algorithm outperforms previous methods. Experiments with entity-mention models and Dirichlet process mixture models demonstrate much faster convergence and better scaling to large data sets.

NeurIPS Conference 2015 Conference Paper

Bidirectional Recurrent Convolutional Networks for Multi-Frame Super-Resolution

  • Yan Huang
  • Wei Wang
  • Liang Wang

Super resolving a low-resolution video is usually handled by either single-image super-resolution (SR) or multi-frame SR. Single-Image SR deals with each video frame independently, and ignores intrinsic temporal dependency of video frames which actually plays a very important role in video super-resolution. Multi-Frame SR generally extracts motion information, e. g. optical flow, to model the temporal dependency, which often shows high computational cost. Considering that recurrent neural network (RNN) can model long-term contextual information of temporal sequences well, we propose a bidirectional recurrent convolutional network for efficient multi-frame SR. Different from vanilla RNN, 1) the commonly-used recurrent full connections are replaced with weight-sharing convolutional connections and 2) conditional convolutional connections from previous input layers to current hidden layer are added for enhancing visual-temporal dependency modelling. With the powerful temporal dependency modelling, our model can super resolve videos with complex motions and achieve state-of-the-art performance. Due to the cheap convolution operations, our model has a low computational complexity and runs orders of magnitude faster than other multi-frame methods.

IJCAI Conference 2015 Conference Paper

On Conceptual Labeling of a Bag of Words

  • Xiangyan Sun
  • Yanghua Xiao
  • Haixun Wang
  • Wei Wang

In natural language processing and information retrieval, the bag of words representation is used to implicitly represent the meaning of the text. Implicit semantics, however, are insufficient in supporting text or natural language based interfaces, which are adopted by an increasing number of applications. Indeed, in applications ranging from automatic ontology construction to question answering, explicit representation of semantics is starting to play a more prominent role. In this paper, we introduce the task of conceptual labeling (CL), which aims at generating a minimum set of conceptual labels that best summarize a bag of words. We draw the labels from a data driven semantic network that contains millions of highly connected concepts. The semantic network provides meaning to the concepts, and in turn, it provides meaning to the bag of words through the conceptual labels we generate. To achieve our goal, we use an information theoretic approach to trade-off the semantic coverage of a bag of words against the minimality of the output labels. Specifically, we use Minimum Description Length (MDL) as the criteria in selecting the best concepts. Our extensive experimental results demonstrate the effectiveness of our approach in representing the explicit semantics of a bag of words.

AAAI Conference 2015 Conference Paper

Transfer Feature Representation via Multiple Kernel Learning

  • Wei Wang
  • Hao Wang
  • Chen Zhang
  • Fanjiang Xu

Learning an appropriate feature representation across source and target domains is one of the most effective solutions to domain adaptation problems. Conventional cross-domain feature learning methods rely on the Reproducing Kernel Hilbert Space (RKHS) induced by a single kernel. Recently, Multiple Kernel Learning (MKL), which bases classifiers on combinations of kernels, has shown improved performance in the tasks without distribution difference between domains. In this paper, we generalize the framework of MKL for cross-domain feature learning and propose a novel Transfer Feature Representation (TFR) algorithm. TFR learns a convex combination of multiple kernels and a linear transformation in a single optimization which integrates the minimization of distribution difference with the preservation of discriminating power across domains. As a result, standard machine learning models trained in the source domain can be reused for the target domain data. After rewritten into a differentiable formulation, TFR can be optimized by a reduced gradient method and reaches the convergence. Experiments in two real-world applications verify the effectiveness of our proposed method.

AAAI Conference 2014 Conference Paper

Cross-Domain Metric Learning Based on Information Theory

  • Hao Wang
  • Wei Wang
  • Chen Zhang
  • Fanjiang Xu

Supervised metric learning plays a substantial role in statistical classification. Conventional metric learning algorithms have limited utility when the training data and testing data are drawn from related but different domains (i. e. , source domain and target domain). Although this issue has got some progress in feature-based transfer learning, most of the work in this area suffers from non-trivial optimization and pays little attention to preserving the discriminating information. In this paper, we propose a novel metric learning algorithm to transfer knowledge from the source domain to the target domain in an information-theoretic setting, where a shared Mahalanobis distance across two domains is learnt by combining three goals together: 1) reducing the distribution difference between different domains; 2) preserving the geometry of target domain data; 3) aligning the geometry of source domain data with its label information. Based on this combination, the learnt Mahalanobis distance effectively transfers the discriminating power and propagates standard classifiers across these two domains. More importantly, our proposed method has closed-form solution and can be efficiently optimized. Experiments in two real-world applications demonstrate the effectiveness of our proposed method.

JBHI Journal 2014 Journal Article

Resource Optimized TTSH-URA for Multimedia Stream Authentication in Swallowable-Capsule-Based Wireless Body Sensor Networks

  • Wei Wang
  • Chunqiu Wang
  • Min Zhao

To ease the burdens on the hospitalization capacity, an emerging swallowable-capsule technology has evolved to serve as a remote gastrointestinal (GI) disease examination technique with the aid of the wireless body sensor network (WBSN). Secure multimedia transmission in such a swallowable-capsule-based WBSN faces critical challenges including energy efficiency and content quality guarantee. In this paper, we propose a joint resource allocation and stream authentication scheme to maintain the best possible video quality while ensuring security and energy efficiency in GI-WBSNs. The contribution of this research is twofold. First, we establish a unique signature-hash (S-H) diversity approach in the authentication domain to optimize video authentication robustness and the authentication bit rate overhead over a wireless channel. Based on the full exploration of S-H authentication diversity, we propose a new two-tier signature-hash (TTSH) stream authentication scheme to improve the video quality by reducing authentication dependence overhead while protecting its integrity. Second, we propose to combine this authentication scheme with a unique S-H oriented unequal resource allocation (URA) scheme to improve the energy-distortion-authentication performance of wireless video delivery in GI-WBSN. Our analysis and simulation results demonstrate that the proposed TTSH with URA scheme achieves considerable gain in both authenticated video quality and energy efficiency.

AAAI Conference 2013 Conference Paper

Exploring the Contribution of Unlabeled Data in Financial Sentiment Analysis

  • Jimmy Ren
  • Wei Wang
  • Jiawei Wang
  • Stephen Liao

With the proliferation of its applications in various industries, sentiment analysis by using publicly available web data has become an active research area in text classification during these years. It is argued by researchers that semi-supervised learning is an effective approach to this problem since it is capable to mitigate the manual labeling effort which is usually expensive and timeconsuming. However, there was a long-term debate on the effectiveness of unlabeled data in text classification. This was partially caused by the fact that many assumptions in theoretic analysis often do not hold in practice. We argue that this problem may be further understood by adding an additional dimension in the experiment. This allows us to address this problem in the perspective of bias and variance in a broader view. We show that the well-known performance degradation issue caused by unlabeled data can be reproduced as a subset of the whole scenario. We argue that if the bias-variance tradeoff is to be better balanced by a more effective feature selection method unlabeled data is very likely to boost the classification performance. We then propose a feature selection framework in which labeled and unlabeled training samples are both considered. We discuss its potential in achieving such a balance. Besides, the application in financial sentiment analysis is chosen because it not only exemplifies an important application, the data possesses better illustrative power as well. The implications of this study in text classification and financial sentiment analysis are both discussed.

NeurIPS Conference 2010 Conference Paper

Multi-View Active Learning in the Non-Realizable Case

  • Wei Wang
  • Zhi-Hua Zhou

The sample complexity of active learning under the realizability assumption has been well-studied. The realizability assumption, however, rarely holds in practice. In this paper, we theoretically characterize the sample complexity of active learning in the non-realizable case under multi-view setting. We prove that, with unbounded Tsybakov noise, the sample complexity of multi-view active learning can be $\widetilde{O}(\log \frac{1}{\epsilon})$, contrasting to single-view setting where the polynomial improvement is the best possible achievement. We also prove that in general multi-view setting the sample complexity of active learning with unbounded Tsybakov noise is $\widetilde{O}(\frac{1}{\epsilon})$, where the order of $1/\epsilon$ is independent of the parameter in Tsybakov noise, contrasting to previous polynomial bounds where the order of $1/\epsilon$ is related to the parameter in Tsybakov noise.

YNIMG Journal 2008 Journal Article

Independent components of the haemodynamic response in intrinsic optical imaging

  • Ingo Schiessl
  • Wei Wang
  • Niall McLoughlin

Functional brain imaging methods are prone to contamination from global vascular artefacts. A variety of methods have been proposed to help segment functional from non-specific changes. Here we quantify the improvement in the signal to noise ratio (SNR) of functional maps, derived from intrinsic optical imaging studies of macaque visual cortex, through the application of Extended Spatial Decorrelation (ESD). The resulting independent component maps and their corresponding time courses reveal for the first time a fast vascular component in the haemodynamic response. ESD is a blind source separation algorithm that utilises spatial statistical features in brain images to separate the recorded mixed sources into independent components. We have investigated differential and single condition experiments using a variety of visual stimuli. To calculate the improvement of the SNR in decibel (dB) we back project separated components onto the original single trial data and analyse the corresponding Fourier spectrum. The application of ESD improved SNR in the functional brain maps from 0. 52 to 16. 88 dB on differential imaging data and from 1. 69 to 12. 83 dB in the case of single condition experiments. Analysing the independent components further we found that they can separate different functional compartments of the cortical vasculature. Some of the components, classified as arterial through slit spectroscopy, revealed a strong fast response to the stimulus onset/offset starting ∼0. 2 s after the change of the stimulus and reaching a peak after ∼0. 4 s. This fast haemodynamic response raises new questions concerning the spatial specificity of the so-called “initial dip”.

v2026.09.13