Arrow Research search

Author name cluster

Lei Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

175 papers
2 author rows

Possible papers

175

EAAI Journal 2026 Journal Article

A Bayesian attention-based Transformer with epistemic and aleatoric uncertainty quantification for trustworthy remaining useful life prediction

  • Lei Wang
  • Zaigang Chen
  • Zhiwen Liu

Transformers have emerged as a state-of-the-art method for remaining useful life (RUL) prediction due to their powerful self-attention mechanisms. However, traditional Transformers fail to consider prediction uncertainties, often resulting in overconfident results. To address this limitation, this paper proposes a Bayesian attention-based Transformer (BATformer) for uncertainty-aware and trustworthy RUL prediction. BATformer explicitly models two fundamental uncertainties: (1) Epistemic uncertainty, accounting for model uncertainty; (2) Aleatoric uncertainty, representing data uncertainty. A Bayesian attention mechanism, where attention weights are treated as latent random variables rather than deterministic values, and a quantile regressor are incorporated into a dual-uncertainty quantification framework to capture the epistemic and aleatoric uncertainties respectively. This dual-uncertainty quantification framework not only facilitates reliable uncertainty quantification but also enhances prediction accuracy. The effectiveness and superiority of BATformer is validated by experiments on a turbofan engine dataset and a tool wear dataset. The results show that BATformer improves RUL prediction performance on the turbofan engine dataset by 5. 37 % and 15. 70 %, and tool wear dataset by 8. 29 % and 15. 75 %, in terms of root mean square error and score function, respectively.

AAAI Conference 2026 Conference Paper

AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection

  • Bin-Bin Gao
  • Yue Zhou
  • Jiangtao Yan
  • Yuezhi Cai
  • Weixi Zhang
  • Meng Wang
  • Jun Liu
  • Yong Liu

Universal visual anomaly detection aims to identify anomalies from novel or unseen vision domains without additional fine-tuning, which is critical in open scenarios. Recent studies have demonstrated that pre-trained vision-language models like CLIP exhibit strong generalization with just zero or a few normal images. However, existing methods struggle to design prompt templates, handle complex token interactions, or require fine-tuning on target domains, resulting in limited flexibility. In this work, we present a simple yet effective AdaptCLIP based on two key insights. First, adaptive visual and textual representations should be learned alternately rather than jointly. Second, comparative learning between query and normal image prompt should incorporate both contextual and aligned residual features, rather than relying solely on residual features. AdaptCLIP treats CLIP models as a foundational service, adding only three simple adapters, visual adapter, textual adapter, and prompt-query adapter, at its input or output ends. AdaptCLIP supports zero-/few-shot generalization across domains and provides a training-free approach on target domains once trained on a base dataset. AdaptCLIP achieves state-of-the-art performance on 12 anomaly detection benchmarks from industrial and medical domains, significantly outperforming existing competitive methods.

AAAI Conference 2026 Conference Paper

Beyond Tokens: Dynamic Latent Reasoning via Semantic Residual Refinement

  • Fangrui Lv
  • Lei Wang
  • Ruixin Hong
  • Yong Du
  • Xiangyu Wu
  • Tingting Gao
  • Guorui Zhou
  • Changshui Zhang

Chain-of-Thought prompting has remarkably advanced LLM reasoning by generating explicit step-by-step tokens, yet its discrete nature inherently limits expressiveness and efficiency, struggling with abstract, ambiguous, or semantically divergent cognition beyond linguistic tokens. Latent reasoning offers a promising alternative by operating in the model’s internal continuous space for richer cognitive representations. However, existing methods typically rely on finetuning or token interpolation to bridge latent and input spaces, introducing training difficulty or semantic degradation. To this end, we propose Dynamic Latent Reasoning (DyLaR), a training-free framework that preserves semantic fidelity to latent space. DyLaR introduces a Semantic Residual Refinement module that progressively refines latent inputs by integrating semantic residuals from prior hidden states, thus capturing expressive semantic hierarchies that closely approximate continuous latent representations. To enhance flexibility, DyLaR further incorporates a dynamic switching policy that allows LLMs to alternate between discrete and latent reasoning based on model uncertainty, favoring explicit reasoning when confident and latent exploration under ambiguity. Empirical experiments across knowledge- and reasoning-intensive tasks demonstrate that DyLaR consistently outperforms strong baselines in both effectiveness and token efficiency. Qualitative analyses further illustrate its interpretability and flexibility in navigating complex reasoning scenarios.

AAAI Conference 2026 Conference Paper

CometNet: Contextual Motif-guided Long-term Time Series Forecasting

  • Weixu Wang
  • Xiaobo Zhou
  • Xin Qiao
  • Lei Wang
  • Tie Qiu

Long-term Time Series Forecasting is crucial across numerous critical domains, yet its accuracy remains fundamentally constrained by the receptive field bottleneck in existing models. Mainstream Transformer- and Multi-layer Perceptron (MLP)-based methods mainly rely on finite look-back windows, limiting their ability to model long-term dependencies and hurting forecasting performance. Naively extending the look-back window proves ineffective, as it not only introduces prohibitive computational complexity, but also drowns vital long-term dependencies in historical noise. To address these challenges, we propose CometNet, a novel Contextual Motif-guided Long-term Time Series Forecasting framework. CometNet first introduces a Contextual Motif Extraction module that identifies recurrent, dominant contextual motifs from complex historical sequences, providing extensive temporal dependencies far exceeding limited look-back windows; Subsequently, a Motif-guided Forecasting module is proposed, which integrates the extracted dominant motifs into forecasting. By dynamically mapping the look-back window to its relevant motifs, CometNet effectively harnesses their contextual information to strengthen long-term forecasting capability. Extensive experimental results on eight real-world datasets have demonstrated that CometNet significantly outperforms current state-of-the-art (SOTA) methods, particularly on extended forecast horizons.

YNIMG Journal 2026 Journal Article

Cortical reorganization and listening effort in sound localization with left- vs. right-sided deafness

  • Yimeng Liu
  • Xuexin Tian
  • Aifang Fu
  • Zengzhi Guo
  • Lei Wang
  • Jie Tang
  • Fei Chen
  • Hongzheng Zhang

Sound source localization is a fundamental ability for living organisms and essential for human daily functioning. Sound localization heavily relies on binaural auditory input. Unilateral auditory impairment not only disrupts sound localization but also alters cortical activation patterns. However, current understanding of these adaptive cortical changes remains limited, with a lack of multidimensional and systematic research. Here, we employed multimodal neuroimaging-functional near-infrared spectroscopy (fNIRS) and pupillometry-to investigate the physiological mechanisms of sound localization from complementary perspectives. We recruited 33 normal-hearing participants, 12 with right single-sided deafness (RSSD), and 8 with left single-sided deafness (LSSD). All participants completed a sound localization task with seven azimuths (±90°, ±60°, ±30°, 0°) in the frontal hemifield. Our results revealed symmetric activation patterns in higher-order auditory cortices within the auditory dorsal stream and distinct cognitive processing for left vs. right hemispheres during sound localization. Following unilateral auditory deprivation, we observed distinct patterns of cortical reorganization that differed by deprivation side. Specifically, RSSD was associated with maintained spatial gradient encoding but required prefrontal compensation, whereas LSSD showed sacrificed spatial resolution for hemispheric processing efficiency. These findings offer novel insights into brain adaptation to unilateral auditory deprivation and provide a foundation for future research and clinical applications.

AAAI Conference 2026 Conference Paper

DanceHA: A Multi-Agent Framework for Document-Level Aspect-Based Sentiment Analysis

  • Lei Wang
  • Min Huang
  • Eduard Dragut

Aspect-Based Sentiment Intensity Analysis (ABSIA) has garnered increasing attention, though research largely focuses on domain-specific, sentence-level settings. In contrast, document-level ABSIA--particularly in addressing complex tasks like extracting Aspect-Category-Opinion-Sentiment-Intensity (ACOSI) tuples--remains underexplored. In this work, we introduce DanceHA, a multi-agent framework designed for open-ended, document-level ABSIA with informal writing styles. DanceHA has two main components: Dance, which employs a divide-and-conquer strategy to decompose the long-context ABSIA task into smaller, manageable sub-tasks for collaboration among specialized agents; and HA, Human-AI collaboration for annotation. We release Inf-ABSIA, a multi-domain document-level ABSIA dataset featuring fine-grained and high-accuracy labels from DanceHA. Extensive experiments demonstrate the effectiveness of our agentic framework and show that the multi-agent knowledge in DanceHA can be effectively transferred into student models. Our results highlight the importance of the overlooked informal styles in ABSIA, as they often intensify opinions tied to specific aspects.

AAAI Conference 2026 Conference Paper

Dual-Channel Learning Framework for Zero-Shot CircRNA-miRNA Interaction Prediction via State Space Modeling

  • Mengmeng Wei
  • Lei Wang
  • Zhu-Hong You
  • Pengwei Hu
  • Bowei Zhao
  • Zhi-An Huang
  • Yu-An Huang
  • Haicheng Yi

CircRNA-miRNA interaction (CMI) plays a pivotal role in disease therapeutics and drug discovery. However, existing methods face several challenges in modeling complex biological networks and zero-shot learning scenarios. Biological networks encapsulate rich biological information, yet current approaches often fail to fully exploit this depth. Moreover, zero-shot prediction requires models to identify new interactions without relying on previously observed samples, imposing stringent requirements on generalization capabilities. To address these limitations, we propose a dual-channel learning framework leveraging State space modeling for Zero-shot CMI prediction (ZeroStem). ZeroStem first enhances the biological relevance of node using prior knowledge, and employs a graph Transformer to extract macro-topological representations. Subsequently, it generates semantic subgraphs based on meta-paths to focus on specific biological relationships, utilizing the Mamba to extract micro-semantic representations via state space modeling. Finally, macro-topological and micro-semantic representations are seamlessly integrated through linear transformation and residual connections, enabling high-precision zero-shot CMI prediction. Extensive experiments on multiple benchmark datasets demonstrate that ZeroStem significantly outperforms existing methods, validating its efficiency and robust generalization in CMI prediction. Case studies further illustrate that ZeroStem offers novel insights into the molecular mechanisms underlying intricate disease-associated networks.

AAAI Conference 2026 Conference Paper

Enhancing All-to-X Backdoor Attacks with Optimized Target Class Mapping

  • Lei Wang
  • Yulong Tian
  • Hao Han
  • Fengyuan Xu

Backdoor attacks pose severe threats to machine learning systems, prompting extensive research in this area. However, most existing work focuses on single-target All-to-One (A2O) attacks, overlooking the more complex All-to-X (A2X) attacks with multiple target classes, which are often assumed to have low attack success rates. In this paper, we first demonstrate that A2X attacks are robust against state-of-the-art defenses. We then propose a novel attack strategy that enhances the success rate of A2X attacks while maintaining robustness by optimizing grouping and target class assignment mechanisms. Our method improves the attack success rate by up to 28%, with average improvements of 6.7%, 16.4%, 14.1% on CIFAR10, CIFAR100, and Tiny-ImageNet, respectively. We anticipate that this study will raise awareness of A2X attacks and stimulate further research in this underexplored area.

JBHI Journal 2026 Journal Article

Extraction of Seafarers’ Occupational Plasticity Brain Network Based on Effective Connectivity Lateralization

  • Lei Wang
  • Weiming Zeng
  • Baolong Li
  • Weifang Nie
  • Hua Zhang
  • Hongyu Chen
  • Yueyang Li
  • Yuhu Shi

Lateralization is an effective model for exploring changes in brain activity and is widely used to assess brain function. Seafarers, as an occupation working in marine environments, are subjected to long-term specialized occupational demands and experiences, which inevitably impact brain function. By utilizing lateralization, the influence of occupational experience on brain activity can be further explored. A novel Effective Connectivity Lateralization Analysis (ECLA) framework is proposed, which incorporates a Transformer-based Granger causality model (Transformer-GC) to analyze the effects of seafaring on brain plasticity. The Transformer-GC model constructs effective connectivity (EC) matrices, and lateralization indices are derived to investigate occupational influences on brain activity. Two control groups of non-seafarers are included to identify seafarers’ unique occupational plasticity brain networks. Results show that Transformer-GC achieves an accuracy improvement of nearly 16% and 19. 4% over the GRU-based and MVGC model, respectively, and a 5% gain over Pearson-based functional connectivity, confirming its superior performance. Moreover, the results of the ECLA showed significant differences in VentralAttention, Somatomotor, DorsalAttention in the seafarer, demonstrating that these brain networks are affected by the long-term work of seafarers. The findings demonstrate the effectiveness of ECLA in revealing the impact of long-term maritime work on brain plasticity, particularly in identifying the brain network of seafarers’ occupational plasticity. It is shown that occupational experience can reshape the lateralization of brain functional activity, offering new insights into neural plasticity across different professions.

AAAI Conference 2026 Conference Paper

Feature Hallucination for Self-supervised Action Recognition (Abstract Reprint)

  • Lei Wang
  • Piotr Koniusz

Understanding human actions in videos requires more than raw pixel analysis; it relies on high-level semantic reasoning and effective integration of multimodal features. We propose a deep translational action recognition framework that enhances recognition accuracy by jointly predicting action concepts and auxiliary features from RGB video frames. At test time, hallucination streams infer missing cues, enriching feature representations without increasing computational overhead. To focus on action-relevant regions beyond raw pixels, we introduce two novel domain-specific descriptors. Object Detection Features (ODF) aggregate outputs from multiple object detectors to capture contextual cues, while Saliency Detection Features (SDF) highlight spatial and intensity patterns crucial for action recognition. Our framework seamlessly integrates these descriptors with auxiliary modalities such as optical flow, Improved Dense Trajectories, skeleton data, and audio cues. It remains compatible with state-of-the-art architectures, including I3D, AssembleNet, Video Transformer Network, FASTER, and recent models like VideoMAE V2 and InternVideo2. To handle uncertainty in auxiliary features, we incorporate aleatoric uncertainty modeling in the hallucination step and introduce a robust loss function to mitigate feature noise. Our multimodal self-supervised action recognition framework achieves state-of-the-art performance on multiple benchmarks, including Kinetics-400, Kinetics-600, and Something-Something V2, demonstrating its effectiveness in capturing fine-grained action dynamics.

AAAI Conference 2026 Conference Paper

FedALT: Federated Fine-Tuning Through Adaptive Local Training with Rest-of-World LoRA

  • Jieming Bian
  • Lei Wang
  • Letian Zhang
  • Jie Xu

Fine-tuning large language models (LLMs) in federated settings enables privacy-preserving adaptation but suffers from cross-client interference due to model aggregation. Existing federated LoRA fine-tuning methods, primarily based on FedAvg, struggle with data heterogeneity, leading to harmful cross-client interference and suboptimal personalization. In this work, we propose FedALT, a novel personalized federated LoRA fine-tuning algorithm that fundamentally departs from FedAvg. Instead of using an aggregated model to initialize local training, each client continues training its individual LoRA while incorporating shared knowledge through a separate Rest-of-World (RoW) LoRA component. To effectively balance local adaptation and global information, FedALT introduces an adaptive mixer that dynamically learns input-specific weightings between the individual and RoW LoRA components, drawing conceptual foundations from the Mixture-of-Experts (MoE) paradigm. Through extensive experiments on NLP benchmarks, we demonstrate that FedALT significantly outperforms state-of-the-art personalized federated LoRA fine-tuning methods, achieving superior local adaptation without sacrificing computational efficiency.

AAAI Conference 2026 Conference Paper

HSA-Net: Hierarchical and Structure-Aware Framework for Efficient and Scalable Molecular Language Modeling

  • Zihang Shao
  • Wentao Lei
  • Lei Wang
  • Wen-Cai Ye
  • Li Liu

Molecular representation learning, a cornerstone for downstream tasks like molecular captioning and molecular property prediction, heavily relies on Graph Neural Networks (GNN). However, GNN suffers from the over-smoothing problem, where node-level features collapse in deep GNN layers. While existing feature projection methods with cross-attention have been introduced to mitigate this issue, they still perform poorly in deep features. This motivated our exploration of using Mamba as an alternative projector for its ability to handle complex sequences. However, we observe that while Mamba excels at preserving global topological information from deep layers, it neglects fine-grained details in shallow layers. The capabilities of Mamba and cross-attention exhibit a global-local trade-off. To resolve this critical global-local trade-off, we propose Hierarchical and Structure-Aware Network (HSA-Net), a novel framework with two modules that enables a hierarchical feature projection and fusion. Firstly, a Hierarchical Adaptive Projector (HAP) module is introduced to process features from different graph layers. It learns to dynamically switch between a cross-attention projector for shallow layers and a structure-aware Graph-Mamba projector for deep layers, producing high-quality, multi-level features. Secondly, to adaptively merge these multi-level features, we design a Source-Aware Fusion (SAF) module, which flexibly selects fusion experts based on the characteristics of the aggregation features, ensuring a precise and effective final representation fusion. Extensive experiments demonstrate that our HSA-Net framework quantitatively and qualitatively outperforms current state-of-the-art (SOTA) methods.

AAAI Conference 2026 Conference Paper

Learning Time in Static Classifiers

  • Xi Ding
  • Lei Wang
  • Piotr Koniusz
  • Yongsheng Gao

Real-world visual data rarely presents as isolated, static instances. Instead, it often evolves gradually over time through variations in pose, lighting, object state, or scene context. However, conventional classifiers are typically trained under the assumption of temporal independence, limiting their ability to capture such dynamics. We propose a simple yet effective framework that equips standard feedforward classifiers with temporal reasoning, all without modifying model architectures or introducing recurrent modules. At the heart of our approach is a novel Support-Exemplar-Query (SEQ) learning paradigm, which structures training data into temporally coherent trajectories. These trajectories enable the model to learn class-specific temporal prototypes and align prediction sequences via a differentiable soft-DTW loss. A multi-term objective further promotes semantic consistency and temporal smoothness. By interpreting input sequences as evolving feature trajectories, our method introduces a strong temporal inductive bias through loss design alone. This proves highly effective in both static and temporal tasks: it enhances performance on fine-grained and ultra-fine-grained image classification, and delivers precise, temporally consistent predictions in video anomaly detection. Despite its simplicity, our approach bridges static and temporal learning in a modular and data-efficient manner, requiring only a simple classifier on top of pre-extracted features.

JBHI Journal 2026 Journal Article

LMSCDA: A Secondary Structure Enhanced Language Model for Predicting CircRNA and Disease Associations

  • Mian-Shuo Lu
  • Lei Wang
  • Meng-Meng Wei
  • Xiao-Rui Su
  • Bo-Wei Zhao
  • Zhu-Hong You
  • De-Shuang Huang

Circular RNA (circRNA) is a kind of non-coding RNA widely present in cells. CircRNA plays a critical role in the occurrence and treatment of diseases. Unraveling the relationships between circRNAs and diseases has become a focus for diagnosis. While computational methods for predicting circRNA-disease associations (CDA) exist, they often oversimplify the representation of circRNA structures. To address this gap, we propose a novel method LMSCDA, which focuses on enhancing circRNA and disease representation by language model to predict CDAs. Specifically, we first calculate circRNA secondary structure by the chemistry principle. Then we employ a hierarchical feature extraction model to extract the circRNA structure and semantic features and amplify features by attention mechanism. Concurrently disease semantic features encoded utilize the biomedical language model. While behavioral features of circRNA and disease captured from circRNA-miRNA and circRNA-disease networks. We integrate them into comprehensive representation to predict CDAs. LMSCDA achieves an AUC of 0. 9877 and an AUPR of 0. 9881 in 5-fold cross-validation on the CircR2Disease dataset. Our approach yields demonstrably competitive results when evaluated against prominent existing models. Our case study on breast cancer first validated predictive accuracy of LMSCDA, with 19 of the top 20 circRNA-Breast cancer associations being confirmed by literature evidence. An analysis on independent clinical transcriptomic dataset identified highly differentially expressed circRNA by LMSCDA, pinpointing candidates for future investigation.

AAAI Conference 2026 Conference Paper

M3UCD: A Multi-task Multimodal Metaphor Understanding Challenge Dataset for LLMs

  • Tianlong Zheng
  • Yating Yang
  • Rui Dong
  • Bo Ma
  • Lei Wang
  • Xi Zhou
  • Siru Miao
  • Turghun Osman

Understanding multimodal metaphors represents a crucial pathway for machines to comprehend human cognition. However, current research remains constrained by superficial dataset annotations, insufficient systematic evaluation of large language models, and fragmented task frameworks. To bridge these gaps, the paper proposes a systematic solution featuring: (I) We present the largest fine-grained Multi-task Multimodal Metaphor Understanding Challenge Dataset (M3UCD) built via multi-perspective collaborative annotation. It contains 15,345 samples, each annotated with 12 manual attribute labels. (II) Systematic benchmarking of LLMs' capacity boundaries in metaphor understanding. Evaluation results reveal the persistent challenges LLMs face in this domain while validating M3UCD's effectiveness and potential. (III) A concise and unified multi-task baseline framework was developed and demonstrated its effectiveness in enhancing the metaphor understanding capabilities of MLLMs.

AAAI Conference 2026 Conference Paper

RAPTOR: Real-Time High-Resolution UAV Video Prediction with Efficient Video Attention

  • Zhan Chen
  • Zile Guo
  • Enze Zhu
  • Peirong Zhang
  • Xiaoxuan Liu
  • Lei Wang
  • Yidan Zhang

Video prediction is plagued by a fundamental trilemma: achieving high-resolution and perceptual quality typically comes at the cost of real-time speed, hindering its use in latency-critical applications. This challenge is most acute for autonomous UAVs in dense urban environments, where foreseeing events from high-resolution imagery is non-negotiable for safety. Existing methods, reliant on iterative generation (diffusion, autoregressive models) or quadratic-complexity attention, fail to meet these stringent demands on edge hardware. To break this long-standing trade-off, we introduce RAPTOR, a video prediction architecture that achieves real-time, high-resolution performance. RAPTOR’s single-pass design avoids the error accumulation and latency of iterative approaches. Its core innovation is Efficient Video Attention (EVA), a novel translator module that factorizes spatiotemporal modeling. Instead of processing flattened spacetime tokens with O((ST)^2) or O(ST) complexity, EVA alternates operations along the spatial (S) and temporal (T) axes. This factorization reduces the time complexity to O(S + T) and memory complexity to O(max(S, T)), enabling global context modeling at 512^2 resolution and beyond, operating directly on dense feature maps with a patch-free design. Complementing this architecture is a 3-stage training curriculum that progressively refines predictions from coarse structure to sharp, temporally coherent details. Experiments show RAPTOR is the first predictor to exceed 30 FPS on a Jetson AGX Orin for 512^2 video, setting a new state-of-the-art on UAVid, KTH, and a custom high-resolution dataset in PSNR, SSIM, and LPIPS. Critically, RAPTOR boosts the mission success rate in a real-world UAV navigation task by 18%, paving the way for safer and more anticipatory embodied agents.

AAAI Conference 2026 Conference Paper

ReFINE: A Reward-Based Framework for Interpretable and Nuanced Evaluation of Radiology Report Generation

  • Yunyi Liu
  • Yingshu Li
  • Zhanyu Wang
  • Xinyu Liang
  • Lingqiao Liu
  • Lei Wang
  • Luping Zhou

Automated radiology report generation (R2Gen) has advanced significantly, yet evaluation remains challenging due to the complexity of assessing report quality. Traditional metrics often misalign with human judgments, failing to identify specific deficiencies. To address this, we introduce ReFINE, a framework for training an Evaluation Model using a novel margin-based reward enforcement loss. This approach decomposes report quality into fine-grained sub-scores across user-defined criteria, improving interpretability. Leveraging GPT-4, we generate diverse training data with paired accepted and rejected reports to train our model under a reward-based system. The trained ReFINE Score provides both granular sub-scores and an aggregated quality assessment, enabling criterion-specific evaluation. Experimental results demonstrate ReFINE's superior alignment with human judgments, outperforming traditional metrics in model selection. Its robustness is validated across three expert-annotated datasets—including chest X-rays and multimodal reports covering 9 imaging modalities—and under two distinct scoring systems.

EAAI Journal 2026 Journal Article

Semi-supervised image segmentation via selective self-ensembling and boundary uncertainty suppression

  • Xiaoguo Yang
  • Yabo Wu
  • Shouxiang Ni
  • Hongmei He
  • Hao Zhang
  • Yu Chen
  • Ke Yan
  • Chuoying Tan

Image segmentation plays a key role in many image-guided clinical applications and deep learning technology has proven effective for this task when sufficient labeled images are available. However, it is very time-consuming and labor-intensive to obtain adequate pixel-level image labels. To alleviate the scarcity of labeled images, we propose a novel semi-supervised segmentation method based on the available uncertainty-aware mean teacher (UAMT) framework by introducing two different strategies, i. e. , selective self-ensembling (SSE) and boundary uncertainty suppression (BUS). The SSE dynamically selects multiple best student models across different training steps to update the teacher model's weights, while the BUS reduces boundary segmentation errors and improves the quantitative potential of loss functions through a unique uncertainty estimation function. With the two strategies, our proposed method was able to obtain promising segmentation performance with limited labeled images and abundant unlabeled ones. We trained and validated our proposed method by segmenting multiple objects from three public datasets (i. e. , PROMISE, REFUGE, and RETA). Extensive experiments showed that our proposed method achieved better segmentation performance than the UAMT, along with the average Dice score (DSC) of 0. 7990 for three different objects, and can compete with several existing semi-supervised methods (i. e. , HCMT, SASSNet, and DTC).

AAAI Conference 2026 Conference Paper

TimeMosaic: Temporal Heterogeneity Guided Time Series Forecasting via Adaptive Granularity Patch and Segment-wise Decoding

  • Kuiye Ding
  • Fanda Fan
  • Chunyi Hou
  • Zheya Wang
  • Lei Wang
  • Zhengxin Yang
  • Jianfeng Zhan

Multivariate time series forecasting is essential in domains such as finance, transportation, climate, and energy. However, existing patch-based methods typically adopt fixed-length segmentation, overlooking the heterogeneity of local temporal dynamics and the decoding heterogeneity of forecasting. Such designs lose details in information-dense regions, introduce redundancy in stable segments, and fail to capture the distinct complexities of short-term and long-term horizons. We propose TimeMosaic, a forecasting framework that aims to address temporal heterogeneity. TimeMosaic employs adaptive patch embedding to dynamically adjust granularity according to local information density, balancing motif reuse with structural clarity while preserving temporal continuity. In addition, it introduces segment-wise decoding that treats each prediction horizon as a related subtask and adapts to horizon-specific difficulty and information requirements, rather than applying a single uniform decoder. Extensive evaluations on benchmark datasets demonstrate that TimeMosaic delivers consistent improvements over existing methods, and our model trained on the large-scale corpus with 321 billion observations achieves performance competitive with state-of-the-art TSFMs.

EAAI Journal 2026 Journal Article

Uncertainty- and hardness-weighted loss functions for medical image segmentation

  • Yanyan Zheng
  • Yabo Wu
  • Jie Chen
  • Xiaoguo Yang
  • Hao Zhang
  • Quanyong Yi
  • Jiantao Pu
  • Lei Wang

Accurate segmentation of medical images is essential for various image processing tasks and is now predominantly achieved using deep learning techniques. However, existing approaches often employ loss functions that fail to account for pixel-level differences in prediction uncertainty or hardness. This limitation frequently results in relatively large segmentation errors, particularly in object boundary regions. To address the limitation, we developed a novel class of uncertainty-/hardness-weighted loss functions by introducing two distinct pixel-wise weighting schemes: probability-guided uncertainty (PGU) and region-enhanced hardness (REH) weights. These weights, derived from the differences between network predictions and their corresponding ground truths, were designed to emphasize challenging pixels while reducing segmentation uncertainties. We validated these loss functions by integrating them with two classical neural networks, i. e. , Swin Transformer based U-shape network (Swin-Unet) and V-shape network (V-Net) to segment two- and three-dimensional target objects across four different images datasets, including Retinal Fundus Glaucoma Challenge (REFUGE) dataset, Retinal Vascular Tree Analysis (RETA) dataset, optical coherence tomography (OCT) dataset, and Atria Segmentation Challenge (ASC) dataset. Extensive experiments demonstrated that our developed loss functions outperformed classical losses, such as cross-entropy (CE) and Dice losses, along with their variants, highlighting the effectiveness and generalization of the introduced weighting schemes. The source code is available at https: //github. com/wmuLei/uhLoss.

JMLR Journal 2025 Journal Article

A Decentralized Proximal Gradient Tracking Algorithm for Composite Optimization on Riemannian Manifolds

  • Lei Wang
  • Le Bao
  • Xin Liu

This paper focuses on minimizing a smooth function combined with a nonsmooth regularization term on a compact Riemannian submanifold embedded in the Euclidean space under a decentralized setting. Typically, there are two types of approaches at present for tackling such composite optimization problems. The first, subgradient-based approaches, rely on subgradient information of the objective function to update variables, achieving an iteration complexity of $O(\epsilon^{-4}\log^2(\epsilon^{-2}))$. The second, smoothing approaches, involve constructing a smooth approximation of the nonsmooth regularization term, resulting in an iteration complexity of $O(\epsilon^{-4})$. This paper proposes a proximal gradient type algorithm that fully exploits the composite structure. The global convergence to a stationary point is established with a significantly improved iteration complexity of $O(\epsilon^{-2})$. To validate the effectiveness and efficiency of our proposed method, we present numerical results from real-world applications, showcasing its superior performance compared to existing approaches. [abs] [ pdf ][ bib ] &copy JMLR 2025. ( edit, beta )

EAAI Journal 2025 Journal Article

A general framework for chromosomal anomaly detection based on dual constraints of nearest-neighbor and regionality

  • Yue Hao
  • Xin Wang
  • Ge Song
  • Zhiyuan Li
  • Lei Wang
  • Lingwei Li
  • Yongqi Nie
  • Peng Wang

The precise identification of structural chromosomal abnormalities (SCA) is essential for the diagnosis of genetic disorders and malignancies. Traditional karyotype analysis is labor-intensive and necessitates the expertise of cytogeneticists. We propose a dual-constraint enhanced framework that combines nearest-neighbor contrastive learning with one-class classification, facilitating automated abnormality detection without the need for anomalous data. Initially, positive sample pairs are constructed utilizing a Chromosomal Query Library (CQL). This process involves the dynamic selection of nearest neighbors, employing soft nearest neighbor selection and cosine similarity to improve feature consistency. Gaussian noise injection enhances generalization by diversifying representations, whereas a momentum update refines CQL embeddings. The Chromosome Banding module (CB module) extracts chromosomal features at multiple scales, whereas the Chromosome Batch Perception module (CBP module) emphasizes challenging samples through spatial and channel attention mechanisms. In the second stage, we present ChromosomeCutMix to create synthetic chromosomal anomalies, enhancing inter-class separation and improving anomaly detection. The proposed framework attains a classification accuracy of 97. 32% and an F1-score of 96. 69%, surpassing current methodologies in terms of sensitivity and robustness. Validated on public and clinical datasets, it offers dependable localization of biological anomalies and automated cytogenetic diagnostics, thereby enhancing the analysis of genetic disorders.

EAAI Journal 2025 Journal Article

A lightweight image segmentation network leveraging inception and squeeze-excitation modules for efficient skin lesion analysis

  • Woei-Hwa Tarn
  • Chi Hou Chong
  • Lei Wang
  • Chang-Fu Kuo
  • Jenhui Chen

The U-shaped network (U-Net) and its derivatives are widely regarded as the cornerstone of medical image segmentation, with performance often improved by increasing model depth and complexity. However, this results in a greater computational burden and slower inference, limiting practical deployment. To address these issues, we propose a lightweight image segmentation based on the convolutional multilayer perceptron (MLP)-based network with U-Net (IS-UNeXt) model, a lightweight segmentation model based on an MLP framework that incorporates Inception-inspired multi-scale fusion blocks and squeeze-and-excitation (SE) modules to mitigate key limitations of existing models, such as high computational complexity, excessive parameter size, and high inference time. Evaluated on the international skin imaging collaboration 2018 (ISIC2018) and the dermoscopic image database acquired at the dermatology service of Hospital Pedro Hispano, Portugal (PH2) datasets, IS-UNeXt reduces inference time by 58. 7%, parameters by 37. 7%, and computational complexity by 48. 4% compared to the convolutional MLP-based network with U-Net (UNeXt), while reaching an intersection over union (IoU) of 81. 1% and a dice coefficient (DC) of 88. 9% on ISIC2018 and IoU of 90. 34% and DC of 94. 42% on PH2. These results demonstrate IS-UNeXt’s effectiveness and efficiency in skin lesion segmentation, rendering it highly suitable for real-time medical applications on resource-constrained devices.

IJCAI Conference 2025 Conference Paper

Adaptive Gradient Learning for Spiking Neural Networks by Exploiting Membrane Potential Dynamics

  • Jiaqiang Jiang
  • Lei Wang
  • Runhao Jiang
  • Jing Fan
  • Rui Yan

Recent advancements have focused on directly training high-performance spiking neural networks (SNNs) by estimating the approximate gradients of spiking activity through a continuous function with constant sharpness, known as surrogate gradient (SG) learning. However, as spikes propagate within neurons and among layers, the distribution of membrane potential dynamics (MPD) will deviate from the gradient-available interval of fixed SG, hindering SNNs from searching the optimal solution space. To maintain the stability of gradient flows, SG needs to align with evolving MPD. Here, we propose a novel adaptive gradient learning for SNNs by exploiting MPD, namely MPD-AGL. It fully accounts for the underlying factors contributing to membrane potential shifts and establishes a dynamic association between SG and MPD at different timesteps to relax gradient estimation, which provides a new degree of freedom for SG learning. Experimental results demonstrate that our method achieves excellent performance at low latency. Moreover, it increases the proportion of neurons that fall into the gradient-available interval compared to fixed SG, effectively mitigating the gradient vanishing problem. Code is available at https: //github. com/jqjiang1999/MPD-AGL.

EAAI Journal 2025 Journal Article

Adaptive human in the loop system for identifying non-optimal states in natural product manufacturing process

  • Qilong Xue
  • Yang Yu
  • Shixin Cen
  • Yequan Yan
  • Jiping Pang
  • Ping Li
  • Yehan Hou
  • Lei Wang

In the extraction of natural products, the identification of non-optimal production states is pivotal for ensuring consistent product quality. Currently, there is a deficiency in online, automated detection methods. This study introduces an online machine vision strategy in a real industrial setting to maintain optimal production state. Specifically, the strategy incorporates an adaptive human in the loop deep learning approach to select high-value samples. This method achieves over 90 % accuracy with fewer training samples, effectively addressing the challenges posed by the low-value density characteristic of industrial data. Additionally, a convolutional neural networks-transformer framework is employed as a classifier for video data to meet the demands of time-series data. To enhance the efficiency of processing multiple video streams, we have implemented a knowledge distillation technique to lighten the model. Finally, this model has been deployed in an actual industrial environment for online monitoring of three extraction devices. The system encapsulates the expertise of engineers to standardize the criteria for assessing production states. This integration of innovative technologies ensures a more reliable and efficient extraction process, meeting the industry's need for consistent product quality.

NeurIPS Conference 2025 Conference Paper

Adaptive LoRA Experts Allocation and Selection for Federated Fine-Tuning

  • Lei Wang
  • Jieming Bian
  • Letian Zhang
  • Jie Xu

Large Language Models (LLMs) have demonstrated impressive capabilities across various tasks, but fine-tuning them for domain-specific applications often requires substantial domain-specific data that may be distributed across multiple organizations. Federated Learning (FL) offers a privacy-preserving solution, but faces challenges with computational constraints when applied to LLMs. Low-Rank Adaptation (LoRA) has emerged as a parameter-efficient fine-tuning approach, though a single LoRA module often struggles with heterogeneous data across diverse domains. This paper addresses two critical challenges in federated LoRA fine-tuning: 1. determining the optimal number and allocation of LoRA experts across heterogeneous clients, and 2. enabling clients to selectively utilize these experts based on their specific data characteristics. We propose FedLEASE (Federated adaptive LoRA Expert Allocation and SElection), a novel framework that adaptively clusters clients based on representation similarity to allocate and train domain-specific LoRA experts. It also introduces an adaptive top-$M$ Mixture-of-Experts mechanism that allows each client to select the optimal number of utilized experts. Our extensive experiments on diverse benchmark datasets demonstrate that FedLEASE significantly outperforms existing federated fine-tuning approaches in heterogeneous client settings while maintaining communication efficiency.

AAAI Conference 2025 Conference Paper

Attention-Driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models Without Fine-Tuning

  • Hai-Ming Xu
  • Qi Chen
  • Lei Wang
  • Lingqiao Liu

Recent advancements in Multimodal Large Language Models (MLLMs) have generated significant interest in their ability to autonomously interact with and interpret Graphical User Interfaces (GUIs). A major challenge in these systems is grounding—accurately identifying critical GUI components such as text or icons based on a GUI image and a corresponding text query. Traditionally, this task has relied on fine-tuning MLLMs with specialized training data to predict component locations directly. However, in this paper, we propose a novel Tuning-free Attention-driven Grounding (TAG) method that leverages the inherent attention patterns in pretrained MLLMs to accomplish this task without the need for additional fine-tuning. Our method involves identifying and aggregating attention maps from specific tokens within a carefully constructed query prompt. Applied to MiniCPM-Llama3-V 2.5, a state-of-the-art MLLM, our tuning-free approach achieves performance comparable to tuning-based methods, with notable success in text localization. Additionally, we demonstrate that our attention map-based grounding technique significantly outperforms direct localization predictions from MiniCPM-Llama3-V 2.5, highlighting the potential of using attention maps from pretrained MLLMs and paving the way for future innovations in this domain.

JMLR Journal 2025 Journal Article

BitNet: 1-bit Pre-training for Large Language Models

  • Hongyu Wang
  • Shuming Ma
  • Lingxiao Ma
  • Lei Wang
  • Wenhui Wang
  • Li Dong
  • Shaohan Huang
  • Huaijie Wang

The increasing size of large language models (LLMs) has posed challenges for deployment and raised concerns about environmental impact due to high energy consumption. Previous research typically applies quantization after pre-training. While these methods avoid the need for model retraining, they often cause notable accuracy loss at extremely low bit-widths. In this work, we explore the feasibility and scalability of 1-bit pre-training. We introduce BitNet b1 and BitNet b1.58, the scalable and stable 1-bit Transformer architecture designed for LLMs. Specifically, we introduce BitLinear as a drop-in replacement of the nn.Linear layer in order to train 1-bit weights from scratch. Experimental results show that BitNet b1 achieves competitive performance, compared to state-of-the-art 8-bit quantization methods and FP16 Transformer baselines. With the ternary weight, BitNet b1.58 matches the half-precision Transformer LLM with the same model size and training tokens in terms of both perplexity and end-task performance, while being significantly more cost-effective in terms of latency, memory, throughput, and energy consumption. More profoundly, BitNet defines a new scaling law and recipe for training new generations of LLMs that are both high-performance and cost-effective. It enables a new computation paradigm and opens the door for designing specific hardware optimized for 1-bit LLMs. [abs] [ pdf ][ bib ] &copy JMLR 2025. ( edit, beta )

EAAI Journal 2025 Journal Article

Detecting worker loss of balance events from point cloud sequence using unsupervised motion-pose learning

  • Mingyu Zhang
  • Lei Wang
  • Yinong Hu
  • Shuai Han
  • Jiawen Zhang
  • Heng Li

Workers' loss of balance (LB), such as slip and trip, may lead to severe injuries and even fatalities. Existing methods for detecting LB typically rely on wearable sensors and focus on specific body parts. This study introduces a novel, non-contact approach utilizing light detection and ranging (LiDAR) technology to detect LB events. By capturing full-body point cloud data, the proposed method extracts both static pose and dynamic motion features across multiple body sections and detects LB events through unsupervised learning. The high-dimensional point cloud sequence is transformed into interpretable gait features, enabling effective unsupervised learning through sequence reconstruction. A two-stream network and fusion strategy are also developed to combine pose and motion features for final LB detection. Experiments with various LB events demonstrate the method's effectiveness, achieving an F1 score of 0. 98 and a recall of 0. 98. Our analysis reveals that integrating features from multiple body parts and the fusion of pose and motion information significantly enhances detection performance. This study offers a promising alternative to traditional methods, providing effective, non-intrusive monitoring of worker safety in dynamic construction environments.

EAAI Journal 2025 Journal Article

Diabetes risk assessment model based on unbalanced public health examination data

  • Liangjun Jiang
  • Jing Wang
  • Jie Xie
  • Zhenhua Xia
  • Juan Li
  • Haimei Gong
  • Lei Wang

Timely risk assessment is crucial for the prevention, treatment, and management of diabetes. In this study, we propose an intelligent risk assessment model framework for diabetes based on public health examination data, focusing on both clinical examination and personal lifestyle dimensions. First, to address the multi-feature issue in examination data, a progressive correlation-based feature selection method was established to select risk features. Fourteen and six key features were selected for the clinical examination and personal lifestyle dimensions, respectively, with the features being concise, transparent, and highly interpretable. Second, to address the issue of sample imbalance and the medical focus on minority class samples, we designed an under-sampling ensemble classification iterative boosting method using light gradient boosting machine as the base classifier. This method combines the advantages of ensemble learning and undersampling, and progressively improves model performance through adaptive sampling mechanisms and multiple rounds of iterative learning. Compared to other balancing methods, our approach demonstrated superior overall performance, achieving accuracies of 89. 02% and 87. 53% in the clinical examination and personal lifestyle dimensions, respectively. Finally, to enhance the practicality of predictions, we designed a web-based visual risk grading scorecard. On an independent test set, the accuracies for the clinical examination and personal lifestyle dimensions reached 85. 74% and 85. 48%, respectively, indicating that the information loss after binning was relatively low and that the features effectively captured factors related to diabetes risk. The proposed diabetes risk assessment model framework demonstrates good practicality and lays a solid foundation for diabetes risk warning and decision support.

NeurIPS Conference 2025 Conference Paper

FedEL: Federated Elastic Learning for Heterogeneous Devices

  • Letian Zhang
  • Bo Chen
  • Jieming Bian
  • Lei Wang
  • Jie Xu

Federated learning (FL) enables distributed devices to collaboratively train machine learning (ML) models while maintaining data privacy. However, the heterogeneous hardware capabilities of participating devices often result in significant training delays, as straggler clients with limited resources prolong the aggregation process. Existing solutions such as client selection, asynchronous FL, and partial training partially address these challenges but encounter issues such as reduced accuracy, stale updates, and compromised model performance due to inconsistent training contributions. To overcome these limitations, we propose FedEL, a federated elastic learning framework that enhances training efficiency while maintaining model accuracy. FedEL introduces a novel window-based training process, sliding the window to locate the training part of the model and dynamically selecting important tensors for training within a coordinated runtime budget. This approach ensures progressive and balanced training across all clients, including stragglers. Additionally, FedEL employs a tensor importance adjustment module, harmonizing local and global tensor importance to mitigate biases caused by data heterogeneity. The experiment results shows that FedEL achieves up to 3. 87× improvement in time-to-accuracy compared to baselines while maintaining or exceeding final test accuracy.

NeurIPS Conference 2025 Conference Paper

Graph Your Own Prompt

  • Xi Ding
  • Lei Wang
  • Piotr Koniusz
  • Yongsheng Gao

We propose Graph Consistency Regularization (GCR), a novel framework that injects relational graph structures, derived from model predictions, into the learning process to promote class-aware, semantically meaningful feature representations. Functioning as a form of self-prompting, GCR enables the model to refine its internal structure using its own outputs. While deep networks learn rich representations, these often capture noisy inter-class similarities that contradict the model's predicted semantics. GCR addresses this issue by introducing parameter-free Graph Consistency Layers (GCLs) at arbitrary depths. Each GCL builds a batch-level feature similarity graph and aligns it with a global, class-aware masked prediction graph, derived by modulating softmax prediction similarities with intra-class indicators. This alignment enforces that feature-level relationships reflect class-consistent prediction behavior, acting as a semantic regularizer throughout the network. Unlike prior work, GCR introduces a multi-layer, cross-space graph alignment mechanism with adaptive weighting, where layer importance is learned from graph discrepancy magnitudes. This allows the model to prioritize semantically reliable layers and suppress noisy ones, enhancing feature quality without modifying the architecture or training procedure. GCR is model-agnostic, lightweight, and improves semantic structure across various networks and datasets. Experiments show that GCR promotes cleaner feature structure, stronger intra-class cohesion, and improved generalization, offering a new perspective on learning from prediction structure.

TAAS Journal 2025 Journal Article

HAG-MTF: Higher-Order Adaptive Generative Graph for Massive Traffic Forecasting in Industry 5.0

  • Lei Wang
  • Huaming Wu
  • Fan Zhang
  • Keqiu Li
  • Wei Yu
  • Shuo Chen

With the evolution of urban smart transportation, the complexity of urban traffic networks escalates, emphasizing the importance of large-scale traffic data prediction in traffic management and urban planning. Traditional spatiotemporal graph models, such as Graph-WaveNet and MTGCN, face exponentially increasing computational complexity as the spatial dimensions expand. To address this challenge, we propose a novel Higher-order Adaptive Generative graph for Massive Traffic Forecasting (HAG-MTF) approach, which utilizes generative AI and high-order graph structures to model the intricate spatial dependencies in large-scale traffic data. The HAG-MTF incorporates a high-order dimensionality reduction module to optimize traffic node processing, utilizing prior graph relationships to generate a fusion graph that dynamically incorporates neighborhood information for efficient, localized graph convolution. The model further incorporates the high-order spatiotemporal relationship extraction module (H-net), enhancing the capacity and speed of traffic data processing while boosting prediction accuracy for complex spatial structures. Furthermore, HAG-MTF introduces a fusion loss function that hierarchically balances multiple objectives, ensuring both precision and computational efficiency. HAG-MTF adaptively handles large-scale real-world traffic data, meeting the needs of traffic controllers and urban planners for predicting massive datasets in practical settings. It supports efficient, flexible interactions via parameter tuning and model outputs, ultimately integrating human insights into traffic analysis and decision-making. This dynamic human-machine collaboration differs from non-Industry 5.0 approaches, which rely on purely automated systems without human input. Those lead to inflexible, brittle conclusions and recommendations, neglecting shifts in traffic patterns driven by human behavior. Extensive experiments on real-world traffic datasets demonstrate that HAG-MTF significantly improves processing efficiency for high-complexity spatial data while delivering precise, human-informed predictions through generative AI-driven operations.

EAAI Journal 2025 Journal Article

Information guided attention network for bearing remaining useful life prediction adaptive to working conditions and fault modes

  • Lei Wang
  • Hongrui Cao
  • Xuefeng Chen

Bearing remaining useful life (RUL) prediction is a major concern in prognostics and health management for industrial system. Recently, deep learning models have significantly advanced the development of RUL prediction technologies. However, the accuracy of deep learning-based RUL prediction methods is easily affected by variable working conditions and multiple fault modes. In this paper, an information guided attention network (IGAN) is developed for bearing RUL prediction adaptive to working conditions and fault modes. First, the proposed IGAN builds the multiscale convolutional layers, which are multiple dilated convolutions with different dilation rates, to fully learn multiscale representations. This ensures that no degradation information of bearings under different working conditions and different fault modes is missed. Second, a novel plug-and-play attention block called information guided attention mechanism (IGAM) is designed to adaptively highlight the informative convolutional channels informed by working conditions and fault modes. Finally, a temporal attention mechanism is integrated into the IGAN to adaptively emphasize the degradation features at different temporal locations to further enhance the feature representation. Two case studies on a bearing dataset across different working conditions and a bearing dataset under time-varying working conditions are conducted to validate the effectiveness and superiority of the proposed method.

JBHI Journal 2025 Journal Article

Integrating Transformer and Graph Attention Network for circRNA-miRNA Interaction Prediction

  • Meng-Meng Wei
  • Lei Wang
  • Bo-Wei Zhao
  • Xiao-Rui Su
  • Zhu-Hong You
  • De-Shuang Huang

CircRNA-miRNA interaction (CMI) plays a crucial role in the gene regulatory network of the cell. Numerous experiments have shown that abnormalities in CMI can impact molecular functions and physiological processes, leading to the occurrence of specific diseases. Current computational models for predicting CMI typically focus on local molecular entity relationships, thereby neglecting inherent molecular attributes and global structural information. To address these limitations, we propose a multi-feature fusion prediction model based on the transformer and graph attention network, named EGATCMI. Specifically, EGATCMI combines the transformer architecture with Word2vec to pre-train the sequence of circRNA and miRNA, capturing their sequence feature representation and sequence similarity. By leveraging the self-attention mechanism, EGATCMI extracts global structural feature from the CMI network. EGATCMI effectively integrates the obtained multi-feature for prediction, achieving AUC values of 0. 9106 and 0. 9470 on the CMI-9905 and CircBank datasets, respectively, outperforming existing methods. In case studies that the prediction of interactions between three miRNAs that are closely related to diseases and circRNAs, 8 out of 10 pairs were accurately predicted and validated. Extensive experimental results demonstrate the potential of EGATCMI as a reliable tool for candidate screening in biological investigations.

ICML Conference 2025 Conference Paper

Kona: An Efficient Privacy-Preservation Framework for KNN Classification by Communication Optimization

  • Guopeng Lin
  • Ruisheng Zhou
  • Shuyu Chen
  • Weili Han
  • Jin Tan
  • Wenjing Fang
  • Lei Wang
  • Tao Wei

K-nearest neighbors (KNN) classification plays a significant role in various applications due to its interpretability. The accuracy of KNN classification relies heavily on large amounts of high-quality data, which are often distributed among different parties and contain sensitive information. Dozens of privacy-preserving frameworks have been proposed for performing KNN classification with data from different parties while preserving data privacy. However, existing privacy-preserving frameworks for KNN classification demonstrate communication inefficiency in the online phase due to two main issues: (1) They suffer from huge communication size for secure Euclidean square distance computations. (2) They require numerous communication rounds to select the $k$ nearest neighbors. In this paper, we present $\texttt{Kona}$, an efficient privacy-preserving framework for KNN classification. We resolve the above communication issues by (1) designing novel Euclidean triples, which eliminate the online communication for secure Euclidean square distance computations, (2) proposing a divide-and-conquer bubble protocol, which significantly reduces communication rounds for selecting the $k$ nearest neighbors. Experimental results on eight real-world datasets demonstrate that $\texttt{Kona}$ significantly outperforms the state-of-the-art framework by $1. 1\times \sim 3121. 2\times$ in communication size, $16. 7\times \sim 5783. 2\times$ in communication rounds, and $1. 1\times \sim 232. 6\times$ in runtime.

NeurIPS Conference 2025 Conference Paper

MPCache: MPC-Friendly KV Cache Eviction for Efficient Private LLM Inference

  • Wenxuan Zeng
  • Ye Dong
  • Jinjin Zhou
  • Jin Tan
  • Lei Wang
  • Tao Wei
  • Runsheng Wang
  • Meng Li

Private large language model (LLM) inference based on secure multi-party computation (MPC) achieves formal data privacy protection but suffers from significant latency overhead, especially for long input sequences. While key-value (KV) cache eviction and sparse attention algorithms have been proposed for efficient LLM inference in plaintext, they are not designed for MPC and cannot benefit private LLM inference directly. In this paper, we propose an accurate and MPC-friendly KV cache eviction framework, dubbed MPCache, building on the observation that historical tokens in a long sequence may have different effects on the downstream decoding. Hence, MPCache combines a look-once static eviction algorithm to discard unimportant KV cache and a query-aware dynamic selection algorithm to activate only a small subset of KV cache for attention computation. MPCache further incorporates a series of optimizations for efficient dynamic KV cache selection, including MPC-friendly similarity approximation, hierarchical KV cache clustering, and cross-layer index-sharing strategy. Extensive experiments demonstrate that MPCache consistently outperforms prior-art KV cache eviction baselines across different generation tasks and achieves 1. 8 ~ 2. 01x and 3. 39 ~ 8. 37x decoding latency and communication reduction on different sequence lengths, respectively.

JMLR Journal 2025 Journal Article

Optimal subsampling for high-dimensional partially linear models via machine learning methods

  • Yujing Shao
  • Lei Wang
  • Heng Lian
  • Haiying Wang

In this paper, we explore optimal subsampling strategies for estimating the parametric regression coefficients in partially linear models with unknown nuisance functions involving high-dimensional and potentially endogenous covariates. To address model misspecifications and the curse of dimensionality, we leverage flexible machine learning (ML) techniques to estimate the unknown nuisance functions. By constructing an unbiased subsampling Neyman-orthogonal score function, we eliminate regularization bias. A two-step algorithm is then used to obtain appropriate ML estimators of the nuisance functions, mitigating the risk of over-fitting. Using martingale techniques, we establish the unconditional consistency and asymptotic normality of the subsample estimators. Furthermore, we derive optimal subsampling probabilities, including A-optimal and L-optimal probabilities as special cases. The proposed optimal subsampling approach is extended to partially linear instrumental variable models to account for potential endogeneity through instrumental variables. Simulation studies and an empirical analysis of the Physicochemical Properties of Protein Tertiary Structure dataset demonstrate the superior performance of our subsample estimators. [abs] [ pdf ][ bib ] &copy JMLR 2025. ( edit, beta )

NeurIPS Conference 2025 Conference Paper

Puzzles: Unbounded Video-Depth Augmentation for Scalable End-to-End 3D Reconstruction

  • Jiahao Ma
  • Lei Wang
  • Miaomiao Liu
  • David Ahmedt-Aristizabal
  • Chuong Nguyen

Multi-view 3D reconstruction remains a core challenge in computer vision. Recent methods, such as DUSt3R and its successors, directly regress pointmaps from image pairs without relying on known scene geometry or camera parameters. However, the performance of these models is constrained by the diversity and scale of available training data. In this work, we introduce Puzzles, a data augmentation strategy that synthesizes an unbounded volume of high-quality, posed video-depth data from just a single image or video clip. By simulating diverse camera trajectories and realistic scene geometry through targeted image transformations, Puzzles significantly enhances data variety. Extensive experiments show that integrating Puzzles into existing video‑based 3D reconstruction pipelines consistently boosts performance, all without modifying the underlying network architecture. Notably, models trained on only 10% of the original data, augmented with Puzzles, achieve accuracy comparable to those trained on the full dataset.

NeurIPS Conference 2025 Conference Paper

Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think

  • Ge Wu
  • Shen Zhang
  • Ruijing Shi
  • Shanghua Gao
  • Zhenyuan Chen
  • Lei Wang
  • Zhaowei Chen
  • Hongcheng Gao

REPA and its variants effectively mitigate training challenges in diffusion models by incorporating external visual representations from pretrained models, through alignment between the noisy hidden projections of denoising networks and foundational clean image representations. We argue that the external alignment, which is absent during the entire denoising inference process, falls short of fully harnessing the potential of discriminative representations. In this work, we propose a straightforward method called $\textit{$\textbf{R}$epresentation $\textbf{E}$ntanglement for $\textbf{G}$eneration}$ ($\textbf{REG}$), which entangles low-level image latents with a single high-level class token from pretrained foundation models for denoising. REG acquires the capability to produce coherent image-class pairs directly from pure noise, substantially improving both generation quality and training efficiency. This is accomplished with negligible additional inference overhead, requiring only one single additional token for denoising (<0. 5\% increase in FLOPs and latency). The inference process concurrently reconstructs both image latents and their corresponding global semantics, where the acquired semantic knowledge actively guides and enhances the image generation process. On ImageNet 256$\times$256, SiT-XL/2 + REG demonstrates remarkable convergence acceleration, achieving $\textbf{63}\times$ and $\textbf{23}\times$ faster training than SiT-XL/2 and SiT-XL/2 + REPA, respectively. More impressively, SiT-L/2 + REG trained for merely 400K iterations outperforms SiT-XL/2 + REPA trained for 4M iterations ($\textbf{10}\times$ longer). Code is available at: https: //github. com/Martinser/REG.

EAAI Journal 2025 Journal Article

Research on global path planning algorithm based on indoor map partition preprocessing

  • Jifan Yang
  • Xiaoling Li
  • Xiaoyang Liu
  • Xunding Pan
  • Lei Wang

Mobile robots utilize Simultaneous Localization and Mapping (SLAM) technology to generate environmental maps and determine their locations within these environments. Subsequently, they employ path planning algorithms to complete navigation tasks. Although existing path planning algorithms are relatively mature, they still exhibit inefficiencies in complex indoor environments. To address this issue, this paper introduces an Indoor Map Partitioning Preprocessing (IMPP) algorithm, which identifies and segments irregularly shaped, complex rooms to accelerate the path planning process. The method initially utilizes the Robot Operating System (ROS) to construct an indoor map dataset and subsequently applies an image segmentation model to identify and enclose rooms. By combining image processing techniques with path planning algorithms, this method can obtain room index information and successfully exclude irrelevant areas from the path planning process. Ultimately, the IMPP algorithm is integrated with a variety of global path planning algorithms. Experimental results demonstrate that in complex indoor environments, this method significantly surpasses existing partitioning methods in terms of room recognition accuracy. Moreover, it decreases the number of expansion points in global path planning algorithms, significantly enhancing processing speed and efficiency.

EAAI Journal 2025 Journal Article

Research on mobile robot path planning based on improved Q-evaluation ant colony optimization algorithm

  • Dongdong Li
  • Lei Wang

This paper proposes an improved Q-evaluation ant colony optimization (IQACO) algorithm to address the shortcomings of traditional ant colony optimization (ACO) algorithm in solving path planning problems, such as the difficulty of effectively reflecting the quality of nodes through the concentration of pheromones, and the problem of poor search stability caused by excessive hyperparameters at different map sizes. Firstly, by analyzing the optimization process of the Q-Learning algorithm, the feasibility of evaluating nodes with Q-value was demonstrated. Secondly, an update method from path transformation to node Q-value was studied, and its superiority over pheromone concentration was demonstrated. Finally, a node selection strategy based on Q-evaluation was proposed to reduce the hyperparameters of the algorithm and improve its stability in optimizing environments of different sizes. To validate the effectiveness of the proposed improvement strategy, both theoretical analysis and repetitive experimental results demonstrate a significant reduction in the time complexity of the proposed algorithm compared to the traditional Ant Colony Optimization (ACO) algorithm. Additionally, simulations conducted on four different map sizes show that the reduction in hyperparameters enhances the algorithm's stability across various map scales. In addition, comparative experiments with other Q-learning-enhanced ACO variants further highlight the advantages of the proposed IQACO algorithm in terms of solution quality, convergence stability, and real-time performance. Furthermore, through comparisons with related work and recent studies, the proposed algorithm's superiority in terms of search capability and stability is clearly demonstrated.

JBHI Journal 2025 Journal Article

Self-Supervised Contrastive Learning on Attribute and Topology Graphs for Predicting Relationships Among lncRNAs, miRNAs and Diseases

  • Lan Huang
  • Nan Sheng
  • Ling Gao
  • Lei Wang
  • Wenju Hou
  • Jie Hong
  • Yan Wang

Exploring associations between long non-coding RNAs (lncRNAs), microRNAs (miRNAs) and diseases is crucial for disease prevention, diagnosis and treatment. While determining these relationships experimentally is resource-intensive and time-consuming, computational methods have emerged as an attractive way. However, existing computational methods tend to focus on single tasks, neglecting the benefits of leveraging multiple biomolecular interactions and domain-specific knowledge for multi-task prediction. Furthermore, the scarcity of labeled data for lncRNA-disease associations (LDAs), miRNA-disease associations (MDAs) and lncRNA-miRNA interactions (LMIs) poses challenges for comprehensive node embedding learning. This paper proposes a multi-task prediction model (called SSCLMD) that employs self-supervised contrastive learning on attribute and topology graphs to identify potential LDAs, MDAs and LMIs. Firstly, domain knowledge of lncRNAs, miRNAs and diseases as well as their interactions are exploited to construct attribute graph and topology graph, respectively. Then, the nodes are encoded in the attribute and topology spaces to extract the specific and common feature. Meanwhile, the attention mechanism is performed to adaptively fuse the embedding from different views. SSCLMD incorporates contrastive self-supervised learning as a regularize to guide node embedding learning in both attribute and topology space without relying on labels. Severing as a regularize in multi-task learning paradigm, it to improves the model. s generalization capabilities. Extensive experiments on 2 manually curated datasets demonstrate that SSCLMD significantly outperforms baseline methods in LDA, MDA and LMI prediction tasks. Case studies on both old and new datasets further supported SSCLMD's ability to uncover novel disease-related lncRNAs and miRNAs.

JBHI Journal 2025 Journal Article

SeqNovo: De Novo Peptide Sequencing Prediction in IoMT via Seq2Seq

  • Ke Wang
  • Mingjia Zhu
  • Wadii Boulila
  • Maha Driss
  • Thippa Reddy Gadekallu
  • Chien-Ming Chen
  • Lei Wang
  • Saru Kumari

In the Internet of Medical Things (IoMT), de novo peptide sequencing prediction is one of the most important techniques for the fields of disease prediction, diagnosis, and treatment. Recently, deep-learning-based peptide sequencing prediction has been a new trend. However, most popular deep learning models for peptide sequencing prediction suffer from poor interpretability and poor ability to capture long-range dependencies. To solve these issues, we propose a model named SeqNovo, which has the encoding-decoding structure of sequence to sequence (Seq2Seq), the highly nonlinear properties of multilayer perceptron (MLP), and the ability of the attention mechanism to capture long-range dependencies. SeqNovo use MLP to improve the feature extraction and utilize the attention mechanism to discover key information. A series of experiments have been conducted to show that the SeqNovo is superior to the Seq2Seq benchmark model, DeepNovo. SeqNovo improves both the accuracy and interpretability of the predictions, which will be expected to support more related research.

NeurIPS Conference 2025 Conference Paper

Time-o1: Time-Series Forecasting Needs Transformed Label Alignment

  • Hao Wang
  • Licheng Pan
  • Zhichao Chen
  • Xu Chen
  • Qingyang Dai
  • Lei Wang
  • Haoxuan Li
  • Zhouchen Lin

Training time-series forecasting models poses unique challenges in loss function design. Most existing approaches adopt temporal mean squared error, but this study reveals two critical limitations: (1) it ignores the presence of label autocorrelation, which biases it from the true label sequence likelihood; (2) it involves excessive number of tasks, which complicates optimization, especially for long-term forecasting. To address these issues, we introduce Time-o1, a transform-enhanced loss function for time-series forecasting. The central idea is to transform the label sequence into decorrelated components with discriminated significance. Models are then trained to align the most significant components, thereby effectively mitigating label autocorrelation and reducing task amount. Experiments demonstrate that Time-o1 achieves state-of-the-art performance and is compatible with various forecast models. Code is available at https: //github. com/Master-PLC/Time-o1.

NeurIPS Conference 2025 Conference Paper

Unlocking Multimodal Mathematical Reasoning via Process Reward Model

  • Ruilin Luo
  • Zhuofan Zheng
  • Lei Wang
  • Yifan Wang
  • Xinzhe Ni
  • Zicheng Lin
  • Songtao Jiang
  • Yiyao Yu

Process Reward Models (PRMs) have shown promise in enhancing the mathematical reasoning capabilities of Large Language Models (LLMs) through Test-Time Scaling (TTS). However, their integration into multimodal reasoning remains largely unexplored. In this work, we take the first step toward unlocking the potential of PRMs in multimodal mathematical reasoning. We identify three key challenges: (i) the scarcity of high-quality reasoning data constrains the capabilities of foundation Multimodal Large Language Models (MLLMs), which imposes further limitations on the upper bounds of TTS and reinforcement learning (RL); (ii) a lack of automated methods for process labeling within multimodal contexts persists; (iii) the employment of process rewards in unimodal RL faces issues like reward hacking, which may extend to multimodal scenarios. To address these issues, we introduce URSA, a three-stage Unfolding multimodal pRocess-Supervision Aided training framework. We first construct MMathCoT-1M, a high-quality large-scale multimodal Chain-of-Thought (CoT) reasoning dataset, to build a stronger math reasoning foundation MLLM, URSA-8B. Subsequently, we go through an automatic process to synthesize process supervision data, which emphasizes both logical correctness and perceptual consistency. We introduce DualMath-1. 1M to facilitate the training of URSA-8B-RM. Finally, we propose Process-Supervised Group-Relative-Policy-Optimization (PS-GRPO), pioneering a multimodal PRM-aided online RL method that outperforms vanilla GRPO. With PS-GRPO application, URSA-8B-PS-GRPO outperforms Gemma3-12B and GPT-4o by 8. 4% and 2. 7% on average across 6 benchmarks.

TAAS Journal 2025 Journal Article

VoI-based Situation-Aware Routing Protocol for Non-linear Underwater Communication Networks

  • Kiran Saleem
  • Lei Wang
  • Rana Zeeshan Ahmed
  • Thippa Reddy Gadekallu
  • Ahmad Almadhor
  • Yang Li

One of the main challenges for underwater applications, such as environmental monitoring and disaster management, is achieving efficient data transmission in environments where conditions change rapidly, and resources need for data transport are scarce. The capability of evaluating the Value of information (VoI) enables us to assess these problems by proposing a Value of Information-based Situation-Aware Non-Linear Routing (VoI SANLR/VoI SANL) method. It aims to deal with critical event scenarios using BDI (Belief-Desire-Intention) logic criteria and prioritizing the timely uploading of data-driven information towards the destination. SANLR of VoI is developed to reduce energy consumption, end-to-end latency, jitter, and improve Packet Delivery Ratio (PDR) in underwater communication networks. VoI SANLR introduces principles of priority-based methods and intends to address challenges in terms of underwater environment such as varying channel conditions, lack energy resources, and real-time decision requirements by using SANLR. Energy optimization analysis reveals consistent outperformance, achieving a remarkable 95% reduction in energy consumption compared to other techniques. Low latency is maintained, ranging from 2.5 to 0.5 seconds, showcasing enhanced efficiency and scalability. VoI SANLR demonstrates exceptional performance in both throughput and jitter. It achieves the highest data transfer rates, ranging from 100 kbps to 110 kbps, indicating outstanding efficiency. Additionally, the jitter remains consistently low, between 1.8 ms and 2 ms, ensuring minimal delay variability and improved communication stability. PDR consistently surpasses other techniques, reaching a maximum of 99%. Additionally, network lifetime analysis demonstrates VoI SANLR's superiority, exhibiting the highest network lifetime at each node and a significant 31.25% improvement at Node 100 compared to other methods.

EAAI Journal 2024 Journal Article

A survey of causal discovery based on functional causal model

  • Lei Wang
  • Shanshan Huang
  • Shu Wang
  • Jun Liao
  • Tingpeng Li
  • Li Liu

Causal discovery finds widespread applications, ranging from estimating treatment effectiveness in medicine, analyzing policy impacts in economics, to constructing predictive models in machine learning—all of which rely on the study and discovery of causal relationships. In recent years, as causal learning has progressed, causal discovery has been classified into different categories depending on assumptions and learning strategies. In this paper, we undertake an exploration of causal discovery methods based on the functional causal model (FCM). We commence by introducing essential terminology associated with causal discovery and laying out the foundational assumptions underpinning FCM-based methods. Following this, we conduct a comprehensive exploration of classical FCM algorithms that have gained prominence in recent years. Furthermore, we scrutinize the performance of these FCM methods across a selection of benchmark datasets. Finally, we deliberate on unresolved issues within this category of methodologies and outline potential avenues for future research.

NeurIPS Conference 2024 Conference Paper

Advancing Video Anomaly Detection: A Concise Review and a New Dataset

  • Liyun Zhu
  • Lei Wang
  • Arjun Raj
  • Tom Gedeon
  • Chen Chen

Video Anomaly Detection (VAD) finds widespread applications in security surveillance, traffic monitoring, industrial monitoring, and healthcare. Despite extensive research efforts, there remains a lack of concise reviews that provide insightful guidance for researchers. Such reviews would serve as quick references to grasp current challenges, research trends, and future directions. In this paper, we present such a review, examining models and datasets from various perspectives. We emphasize the critical relationship between model and dataset, where the quality and diversity of datasets profoundly influence model performance, and dataset development adapts to the evolving needs of emerging approaches. Our review identifies practical issues, including the absence of comprehensive datasets with diverse scenarios. To address this, we introduce a new dataset, Multi-Scenario Anomaly Detection (MSAD), comprising 14 distinct scenarios captured from various camera views. Our dataset has diverse motion patterns and challenging variations, such as different lighting and weather conditions, providing a robust foundation for training superior models. We conduct an in-depth analysis of recent representative models using MSAD and highlight its potential in addressing the challenges of detecting anomalies across diverse and evolving surveillance scenarios.

ICML Conference 2024 Conference Paper

Ditto: Quantization-aware Secure Inference of Transformers upon MPC

  • Haoqi Wu
  • Wenjing Fang
  • Yancheng Zheng
  • Junming Ma
  • Jin Tan
  • Lei Wang

Due to the rising privacy concerns on sensitive client data and trained models like Transformers, secure multi-party computation (MPC) techniques are employed to enable secure inference despite attendant overhead. Existing works attempt to reduce the overhead using more MPC-friendly non-linear function approximations. However, the integration of quantization widely used in plaintext inference into the MPC domain remains unclear. To bridge this gap, we propose the framework named Ditto to enable more efficient quantization-aware secure Transformer inference. Concretely, we first incorporate an MPC-friendly quantization into Transformer inference and employ a quantization-aware distillation procedure to maintain the model utility. Then, we propose novel MPC primitives to support the type conversions that are essential in quantization and implement the quantization-aware MPC execution of secure quantized inference. This approach significantly decreases both computation and communication overhead, leading to improvements in overall efficiency. We conduct extensive experiments on Bert and GPT2 models to evaluate the performance of Ditto. The results demonstrate that Ditto is about $3. 14\sim 4. 40\times$ faster than MPCFormer (ICLR 2023) and $1. 44\sim 2. 35\times$ faster than the state-of-the-art work PUMA with negligible utility degradation.

JBHI Journal 2024 Journal Article

GSLCDA: An Unsupervised Deep Graph Structure Learning Method for Predicting CircRNA-Disease Association

  • Lei Wang
  • Zheng-Wei Li
  • Zhu-Hong You
  • De-Shuang Huang
  • Leon Wong

Growing studies reveal that Circular RNAs (circRNAs) are broadly engaged in physiological processes of cell proliferation, differentiation, aging, apoptosis, and are closely associated with the pathogenesis of numerous diseases. Clarification of the correlation among diseases and circRNAs is of great clinical importance to provide new therapeutic strategies for complex diseases. However, previous circRNA-disease association prediction methods rely excessively on the graph network, and the model performance is dramatically reduced when noisy connections occur in the graph structure. To address this problem, this paper proposes an unsupervised deep graph structure learning method GSLCDA to predict potential CDAs. Concretely, we first integrate circRNA and disease multi-source data to constitute the CDA heterogeneous network. Then the network topology is learned using the graph structure, and the original graph is enhanced in an unsupervised manner by maximize the inter information of the learned and original graphs to uncover their essential features. Finally, graph space sensitive k-nearest neighbor (KNN) algorithm is employed to search for latent CDAs. In the benchmark dataset, GSLCDA obtained 92. 67% accuracy with 0. 9279 AUC. GSLCDA also exhibits exceptional performance on independent datasets. Furthermore, 14, 12 and 14 of the top 16 circRNAs with the most points GSLCDA prediction scores were confirmed in the relevant literature in the breast cancer, colorectal cancer and lung cancer case studies, respectively. Such results demonstrated that GSLCDA can validly reveal underlying CDA and offer new perspectives for the diagnosis and therapy of complex human diseases.

ICLR Conference 2024 Conference Paper

In-context Autoencoder for Context Compression in a Large Language Model

  • Tao Ge 0001
  • Jing Hu 0001
  • Lei Wang
  • Xun Wang 0012
  • Siqing Chen
  • Furu Wei

We propose the In-context Autoencoder (ICAE), leveraging the power of a large language model (LLM) to compress a long context into short compact memory slots that can be directly conditioned on by the LLM for various purposes. ICAE is first pretrained using both autoencoding and language modeling objectives on massive text data, enabling it to generate memory slots that accurately and comprehensively represent the original context. Then, it is fine-tuned on instruction data for producing desirable responses to various prompts. Experiments demonstrate that our lightweight ICAE, introducing about 1% additional parameters, effectively achieves $4\times$ context compression based on Llama, offering advantages in both improved latency and GPU memory cost during inference, and showing an interesting insight in memorization as well as potential for scalability. These promising results imply a novel perspective on the connection between working memory in cognitive science and representation learning in LLMs, revealing ICAE's significant implications in addressing the long context problem and suggesting further research in LLM context management. Our data, code and models are available at https://github.com/getao/icae.

EAAI Journal 2024 Journal Article

Learning discriminative context for salient object detection

  • Ge Zhu
  • Lei Wang
  • Jinping Tang

Context understanding is important for salient object detection (SOD) in complex scenes. To alleviate visual confusion, we propose a context aware network that full use of the same type contextual information for SOD. Specifically, we introduce the pixel relationships into decoder, which can extract the explicit contextual information to alleviate visual confusion by learning the similarities and differences between pixels. Furthermore, we explore the optimal embedding position of the pixel relationship in network to maximize its benefits, thereby reducing the introduction of noise. Compared with 20 counterparts, experimental results on five datasets show that our approach has better generalization, which averagely improves 1. 3% and 2. 8% over the advanced CNN-based model and Transformer-based model on three evaluation metrics. The improvements demonstrate effectiveness of the proposed context learning strategies, which are helpful for dealing with various complex scenes. Moreover, our model is more efficient with an inference speed of 56. 7 FPS on a single NVIDIA 2080TI GPU. Codes are available at https: //github. com/lesonly/CANet.

JBHI Journal 2024 Journal Article

MAGCDA: A Multi-Hop Attention Graph Neural Networks Method for CircRNA-Disease Association Prediction

  • Lei Wang
  • Zheng-Wei Li
  • Zhu-Hong You
  • De-Shuang Huang
  • Leon Wong

With a growing body of evidence establishing circular RNAs (circRNAs) are widely exploited in eukaryotic cells and have a significant contribution in the occurrence and development of many complex human diseases. Disease-associated circRNAs can serve as clinical diagnostic biomarkers and therapeutic targets, providing novel ideas for biopharmaceutical research. However, available computation methods for predicting circRNA-disease associations (CDAs) do not sufficiently consider the contextual information of biological network nodes, making their performance limited. In this work, we propose a multi-hop attention graph neural network-based approach MAGCDA to infer potential CDAs. Specifically, we first construct a multi-source attribute heterogeneous network of circRNAs and diseases, then use a multi-hop strategy of graph nodes to deeply aggregate node context information through attention diffusion, thus enhancing topological structure information and mining data hidden features, and finally use random forest to accurately infer potential CDAs. In the four gold standard data sets, MAGCDA achieved prediction accuracy of 92. 58%, 91. 42%, 83. 46% and 91. 12%, respectively. MAGCDA has also presented prominent achievements in ablation experiments and in comparisons with other models. Additionally, 18 and 17 potential circRNAs in top 20 predicted scores for MAGCDA prediction scores were confirmed in case studies of the complex diseases breast cancer and Almozheimer's disease, respectively. These results suggest that MAGCDA can be a practical tool to explore potential disease-associated circRNAs and provide a theoretical basis for disease diagnosis and treatment.

JBHI Journal 2024 Journal Article

Predicting Protein Functions Based on Heterogeneous Graph Attention Technique

  • Yingwen Zhao
  • Zhihao Yang
  • Lei Wang
  • Yin Zhang
  • Hongfei Lin
  • Jian Wang

In bioinformatics, protein function prediction stands as a fundamental area of research and plays a crucial role in addressing various biological challenges, such as the identification of potential targets for drug discovery and the elucidation of disease mechanisms. However, known functional annotation databases usually provide positive experimental annotations that proteins carry out a given function, and rarely record negative experimental annotations that proteins do not carry out a given function. Therefore, existing computational methods based on deep learning models focus on these positive annotations for prediction and ignore these scarce but informative negative annotations, leading to an underestimation of precision. To address this issue, we introduce a deep learning method that utilizes a heterogeneous graph attention technique. The method first constructs a heterogeneous graph that covers the protein-protein interaction network, ontology structure, and positive and negative annotation information. Then, it learns embedding representations of proteins and ontology terms by using the heterogeneous graph attention technique. Finally, it leverages these learned representations to reconstruct the positive protein-term associations and score unobserved functional annotations. It can enhance the predictive performance by incorporating these known limited negative annotations into the constructed heterogeneous graph. Experimental results on three species (i. e. , Human, Mouse, and Arabidopsis) demonstrate that our method can achieve better performance in predicting new protein annotations than state-of-the-art methods.

EAAI Journal 2024 Journal Article

Rapid detection method for insulation performance of vacuum glass based on ensemble learning

  • Xiaoling Li
  • Shunyu Liu
  • Yuanqi Wang
  • Fuquan Zhou
  • Lei Wang

For a long time, the use of steady state method to detect the thermal insulation performance of vacuum glass caused some problems such as long detection period, many influencing factors, inaccurate detection, etc. In order to improve the efficiency of vacuum glass insulation performance detection and reduce the cost of vacuum glass industrialization, the ensemble learning method for rapid detection of vacuum glass insulation performance is studied. Firstly, the correlation between variables and the distribution of variables are analyzed based on unsteady state method. The temperature-related variables and heat transfer coefficient are used as the input variables and target variables of the model. Then, three models and one model are selected as the first and second layers of stacking model based on five-fold cross-validation and Spearman correlation analysis. Finally, the heat transfer coefficient characterizing the thermal insulation performance of vacuum glass is predicted by the designed 3 + 1 stacking model. We involve 10 single models and other 11 ensemble models to verify the effectiveness of the method. The experimental results show that the 3 + 1 stacking model based on five-fold cross-validation and Spearman correlation analysis has the best prediction effect, which outperforms single model and other ensemble models. It improves the generalization ability of prediction model.

NeurIPS Conference 2024 Conference Paper

Reflective Multi-Agent Collaboration based on Large Language Models

  • Xiaohe Bo
  • Zeyu Zhang
  • Quanyu Dai
  • Xueyang Feng
  • Lei Wang
  • Rui Li
  • Xu Chen
  • Ji-Rong Wen

Benefiting from the powerful language expression and planning capabilities of Large Language Models (LLMs), LLM-based autonomous agents have achieved promising performance in various downstream tasks. Recently, based on the development of single-agent systems, researchers propose to construct LLM-based multi-agent systems to tackle more complicated tasks. In this paper, we propose a novel framework, named COPPER, to enhance the collaborative capabilities of LLM-based agents with the self-reflection mechanism. To improve the quality of reflections, we propose to fine-tune a shared reflector, which automatically tunes the prompts of actor models using our counterfactual PPO mechanism. On the one hand, we propose counterfactual rewards to assess the contribution of a single agent’s reflection within the system, alleviating the credit assignment problem. On the other hand, we propose to train a shared reflector, which enables the reflector to generate personalized reflections according to agent roles, while reducing the computational resource requirements and improving training stability. We conduct experiments on three datasets to evaluate the performance of our model in multi-hop question answering, mathematics, and chess scenarios. Experimental results show that COPPER possesses stronger reflection capabilities and exhibits excellent generalization performance across different actor models.

EAAI Journal 2024 Journal Article

Research on vacuum glass insulation performance prediction based on unsteady state multivariate data screening and multi-model fusion self-optimization

  • Xiaoling Li
  • Yuanqi Wang
  • Fuquan Zhou
  • Lei Wang

A self-optimization regression method based on multivariate data screening and multi-model fusion is proposed for the regression prediction of vacuum glass insulation performance. Firstly, we analyzed the positive correlation between the temperature change rate and heat transfer coefficient. Secondly, we combined the technique of multi-distribution overall diffusion to generate virtual sample data with multiple variables corresponding to the small neighborhood of the target variable, addressing the issue of insufficient sample size. We conducted quantity and threshold screening on the generated data to enhance the effectiveness of the virtual sample data. The screened data and original training set data are integrated to form new training samples. Then, the limited-memory Broyden Fletcher Goldfarb Shanno with boundary constraints (L-BFGS-B) algorithm is introduced to optimize the parameters of multiple models to improve the accuracy of model regression prediction. Finally, through the self-searching optimization framework structure, the virtual sample data screening and optimization of various model parameters are updated. Eventually, the optimal model for predicting the heat transfer coefficient of vacuum glass insulation performance is determined from multiple models. In order to validate the effectiveness of self-optimization regression method, we conducted insulation performance heat transfer coefficient predictions under 30 different ways for vacuum glass. Experimental results demonstrate that the self-optimization regression method based on multivariate data screening and multi-model fusion can obtain the most effective regression prediction model, achieving accurate, rapid and intelligent prediction of vacuum glass insulation performance.

AAAI Conference 2024 Conference Paper

Roll with the Punches: Expansion and Shrinkage of Soft Label Selection for Semi-supervised Fine-Grained Learning

  • Yue Duan
  • Zhen Zhao
  • Lei Qi
  • Luping Zhou
  • Lei Wang
  • Yinghuan Shi

While semi-supervised learning (SSL) has yielded promising results, the more realistic SSL scenario remains to be explored, in which the unlabeled data exhibits extremely high recognition difficulty, e.g., fine-grained visual classification in the context of SSL (SS-FGVC). The increased recognition difficulty on fine-grained unlabeled data spells disaster for pseudo-labeling accuracy, resulting in poor performance of the SSL model. To tackle this challenge, we propose Soft Label Selection with Confidence-Aware Clustering based on Class Transition Tracking (SoC) by reconstructing the pseudo-label selection process by jointly optimizing Expansion Objective and Shrinkage Objective, which is based on a soft label manner. Respectively, the former objective encourages soft labels to absorb more candidate classes to ensure the attendance of ground-truth class, while the latter encourages soft labels to reject more noisy classes, which is theoretically proved to be equivalent to entropy minimization. In comparisons with various state-of-the-art methods, our approach demonstrates its superior performance in SS-FGVC. Checkpoints and source code are available at https://github.com/NJUyued/SoC4SS-FGVC.

AAAI Conference 2024 Conference Paper

S2WAT: Image Style Transfer via Hierarchical Vision Transformer Using Strips Window Attention

  • Chiyu Zhang
  • Xiaogang Xu
  • Lei Wang
  • Zaiyan Dai
  • Jun Yang

Transformer's recent integration into style transfer leverages its proficiency in establishing long-range dependencies, albeit at the expense of attenuated local modeling. This paper introduces Strips Window Attention Transformer (S2WAT), a novel hierarchical vision transformer designed for style transfer. S2WAT employs attention computation in diverse window shapes to capture both short- and long-range dependencies. The merged dependencies utilize the "Attn Merge" strategy, which adaptively determines spatial weights based on their relevance to the target. Extensive experiments on representative datasets show the proposed method's effectiveness compared to state-of-the-art (SOTA) transformer-based and other approaches. The code and pre-trained models are available at https://github.com/AlienZhang1996/S2WAT.

EAAI Journal 2024 Journal Article

SDG: A global large-scale airport perception disparity cognition modeling method based on deep learning and geographic knowledge

  • Ning Li
  • Liang Cheng
  • Hui Chen
  • Yalu Zhang
  • Lei Wang
  • Chen Ji
  • Manchun Li

Global airport perception levels vary due to natural geographical factors and economic development disparities. Understanding these differences is crucial for assessing regional airport development and its correlation with geographical patterns. However, there are limited methods available to effectively comprehend these disparities. To address this issue, this paper proposes a Salience, Disturbance, and Geographic-knowledge (SDG) approach for the cognitive analysis of global large-scale airport perception differences. Salience is assessed using a two-class deep learning model to evaluate the prominence of known airports. Disturbance is evaluated using an object detection model to measure background interference in large-scale airport perception. Geographic-knowledge analysis considers the correlation between regional airports and their surrounding geographic environment. The results rank perception difficulties for 17 regions worldwide, with Tajikistan exhibiting the highest difficulty at 0. 922, while the Jiangsu–Zhejiang–Shanghai region in China has the lowest at 0. 102. We also performed correlation analyses to validate the effectiveness of our model. To our knowledge, this paper pioneers the cognitive analysis of target perception difficulty differences across multiple global regions.

AIIM Journal 2024 Journal Article

Semi-supervised image segmentation using a residual-driven mean teacher and an exponential Dice loss

  • Chenyang Mei
  • Xiaoguo Yang
  • Mi Zhou
  • Shaodan Zhang
  • Hao Chen
  • Xiaokai Yang
  • Lei Wang

Semi-supervised segmentation plays an important role in computer vision and medical image analysis and can alleviate the burden of acquiring abundant expert-annotated images. In this paper, we developed a residual-driven semi-supervised segmentation method (termed RDMT) based on the classical mean teacher (MT) framework by introducing a novel model-level residual perturbation and an exponential Dice (eDice) loss. The introduced perturbation was integrated into the exponential moving average (EMA) scheme to enhance the performance of the MT, while the eDice loss was used to improve the detection sensitivity of a given network to object boundaries. We validated the developed method by applying it to segment 3D Left Atrium (LA) and 2D optic cup (OC) from the public LASC and REFUGE datasets based on the V-Net and U-Net, respectively. Extensive experiments demonstrated that the developed method achieved the average Dice score of 0. 8776 and 0. 7751, when trained on 10% and 20% labeled images, respectively for the LA and OC regions depicted on the LASC and REFUGE datasets. It significantly outperformed the MT and can compete with several existing semi-supervised segmentation methods (i. e. , HCMT, UAMT, DTC and SASS).

AAAI Conference 2024 Conference Paper

T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Large Language Model Signals for Science Question Answering

  • Lei Wang
  • Yi Hu
  • Jiabang He
  • Xing Xu
  • Ning Liu
  • Hui Liu
  • Heng Tao Shen

Large Language Models (LLMs) have recently demonstrated exceptional performance in various Natural Language Processing (NLP) tasks. They have also shown the ability to perform chain-of-thought (CoT) reasoning to solve complex problems. Recent studies have explored CoT reasoning in complex multimodal scenarios, such as the science question answering task, by fine-tuning multimodal models with high-quality human-annotated CoT rationales. However, collecting high-quality COT rationales is usually time-consuming and costly. Besides, the annotated rationales are hardly accurate due to the external essential information missed. To address these issues, we propose a novel method termed T-SciQ that aims at teaching science question answering with LLM signals. The T-SciQ approach generates high-quality CoT rationales as teaching signals and is advanced to train much smaller models to perform CoT reasoning in complex modalities. Additionally, we introduce a novel data mixing strategy to produce more effective teaching data samples for simple and complex science question answer problems. Extensive experimental results show that our T-SciQ method achieves a new state-of-the-art performance on the ScienceQA benchmark, with an accuracy of 96.18%. Moreover, our approach outperforms the most powerful fine-tuned baseline by 4.5%. The code is publicly available at https://github.com/T-SciQ/T-SciQ.

NeurIPS Conference 2024 Conference Paper

Taming Cross-Domain Representation Variance in Federated Prototype Learning with Heterogeneous Data Domains

  • Lei Wang
  • Jieming Bian
  • Letian Zhang
  • Chen Chen
  • Jie Xu

Federated learning (FL) allows collaborative machine learning training without sharing private data. While most FL methods assume identical data domains across clients, real-world scenarios often involve heterogeneous data domains. Federated Prototype Learning (FedPL) addresses this issue, using mean feature vectors as prototypes to enhance model generalization. However, existing FedPL methods create the same number of prototypes for each client, leading to cross-domain performance gaps and disparities for clients with varied data distributions. To mitigate cross-domain feature representation variance, we introduce FedPLVM, which establishes variance-aware dual-level prototypes clustering and employs a novel $\alpha$-sparsity prototype loss. The dual-level prototypes clustering strategy creates local clustered prototypes based on private data features, then performs global prototypes clustering to reduce communication complexity and preserve local data privacy. The $\alpha$-sparsity prototype loss aligns samples from underrepresented domains, enhancing intra-class similarity and reducing inter-class similarity. Evaluations on Digit-5, Office-10, and DomainNet datasets demonstrate our method's superiority over existing approaches.

IJCAI Conference 2024 Conference Paper

Trade When Opportunity Comes: Price Movement Forecasting via Locality-Aware Attention and Iterative Refinement Labeling

  • Liang Zeng
  • Lei Wang
  • Hui Niu
  • Ruchen Zhang
  • Ling Wang
  • Jian Li

Price movement forecasting, aimed at predicting financial asset trends based on current market information, has achieved promising advancements through machine learning (ML) methods. Most existing ML methods, however, struggle with the extremely low signal-to-noise ratio and stochastic nature of financial data, often mistaking noises for real trading signals without careful selection of potentially profitable samples. To address this issue, we propose LARA, a novel price movement forecasting framework with two main components: Locality-Aware Attention (LA-Attention) and Iterative Refinement Labeling (RA-Labeling). (1) LA-Attention, enhanced by metric learning techniques, automatically extracts the potentially profitable samples through masked attention scheme and task-specific distance metrics. (2) RA-Labeling further iteratively refines the noisy labels of potentially profitable samples, and combines the learned predictors robust to the unseen and noisy samples. In a set of experiments on three real-world financial markets: stocks, cryptocurrencies, and ETFs, LARA significantly outperforms several machine learning based methods on the Qlib quantitative investment platform. Extensive ablation studies confirm LARA's superior ability in capturing more reliable trading opportunities.

AAAI Conference 2024 Conference Paper

Would You Like Your Data to Be Trained? A User Controllable Recommendation Framework

  • Lei Wang
  • Xu Chen
  • Zhenhua Dong
  • Quanyu Dai

Recommender systems have a significant impact on various real-world applications, shaping people's daily lives and enhancing productivity. Traditional recommender models aim to collect extensive user information to accurately estimate user preferences. However, in practical scenarios, users may not want all their behaviors to be included in the model training process. This paper introduces a novel recommendation paradigm that allows users to indicate their ``willingness'' regarding which data should contribute to model training. The models are then optimized to maximize utility, which considers the trade-off between recommendation performance and respecting user preferences. The recommendation problem is formulated as a multiplayer game, with each user acting as a player and using a selection vector to indicate their willingness to include specific interacted items in training. To efficiently solve this game, an influence function-based model is proposed to approximate recommendation performances for different actions without re-optimizing the model. Furthermore, an enhanced model leveraging multiple anchor actions for the influence function is introduced to improve performance approximation accuracy. The convergence rate of the algorithm is theoretically analyzed, and the advantages of incorporating multiple anchor actions are demonstrated. Extensive experiments on both simulated and real-world datasets validate the effectiveness of the proposed models in balancing recommendation quality and user willingness. To promote this research direction, we have released our project at https://paitesanshi.github.io/IFRQE/.

JBHI Journal 2023 Journal Article

A High-Rate Hybrid BCI System Based on High-Frequency SSVEP and sEMG

  • Hongyan Cui
  • Xinyi Chi
  • Lei Wang
  • Xiaogang Chen

Recently, various biosignals have been combined with electroencephalography (EEG) to build hybrid brain-computer interface (BCI) systems to improve system performance. Since steady-state visual evoked potential (SSVEP) and surface electromyography (sEMG) are easy-to-use, non-invasive techniques, and have high signal-to-noise ratio (SNR), hybrid BCI systems combining SSVEP and sEMG have received much attention in the BCI literature. However, most existing studies regarding hybrid BCIs based on SSVEP and sEMG adopt low-frequency visual stimuli to induce SSVEPs. The comfort of these systems needs further improvement to meet the practical application requirements. The present study realized a novel hybrid BCI combining high-frequency SSVEP and sEMG signals for spelling applications. EEG and sEMG were obtained simultaneously from the scalp and skin surface of subjects, respectively. These two types of signals were analyzed independently and then combined to determine the target stimulus. Our online results demonstrated that the developed hybrid BCI yielded a mean accuracy of 88. 07 ± 1. 43% and ITR of 159. 12 ± 4. 31 bits/min. These results exhibited the feasibility and effectiveness of fusing high-frequency SSVEP and sEMG towards improving the total BCI system performance.

EAAI Journal 2023 Journal Article

A new data fusion driven-sparse representation learning method for bearing intelligent diagnosis in small and unbalanced samples

  • Yike Zhao
  • Xin Zhang
  • Jiaxu Wang
  • Lei Wu
  • Zhiwen Liu
  • Lei Wang

Dictionary learning has made enormous achievements for its powerful feature representation capabilities. For the bearing fault diagnosis, the lack of failure samples is always a problem demanding prompt solutions. Since failure samples are far less than the normal samples, the data is inevitably to be unbalanced. To solve these problems, a new data fusion driven-sparse representation learning (DFDSRL) framework is proposed for fault diagnosis in small and unbalanced samples. The proposed method first constructs different dictionaries severally from the data samples grouped according to the signal modes (for example, the signals measured from the bearing under different working conditions). To induce the dictionary discriminability, discriminative sparse codes errors, reconstruction errors and classification errors are integrated as optimization objectives. Then, a new data fusion strategy is developed to fuse the dictionaries from all signal modes in the same fault class, and a fused dictionary for each fault class is generated for the final fault identification. The fusion is performed at the feature level instead of the conventional data level, and the discriminability of the fused dictionaries are further enhanced with the fusion strategy by eliminating the insignificant features shared by the atoms in each fault class dictionaries. Experimental results show that the DFDSRL achieves high fault identification accuracy for the problems of small and unbalanced samples in comparison with several advanced methods, benefitting from its excellent capability of fusing more fault features. By transforming the data unbalanced problem into the balanced small sample problem, DFDSRL further improves the fault identification accuracy.

EAAI Journal 2023 Journal Article

A sketch semantic segmentation method based on point-segment level interaction

  • Shihui Zhang
  • Lei Wang
  • Xueqiang Han
  • Shi Wang

Sketch semantic segmentation is a basic computer vision task, which poses great challenges due to the abstraction of sketches and different drawing styles. In this paper, we propose a sketch semantic segmentation method based on point-segment level interaction. Specifically, an enhanced local feature aggregation (ELFA) module is developed based on two kinds of distance information between the neighboring points/segments (NPs/NSs) and the corresponding center points/segments (CPs/CSs). The ELFA module not only extracts local features adequately, but also takes into account the different effects of the NPs/NSs on the corresponding CPs/CSs. Based on the ELFA module, a point-level branch and a segment-level branch are established to encode the semantics of sketches from a point level and a segment level respectively, which makes the features extracted from the point-level branch complementary to those extracted from the segment-level branch. Further, a point-segment level interaction (PSLI) module is designed to interchange the information of the two levels and to reduce the losing of some important semantic details caused by feature selection of multiple stages. The PSLI module can be placed in several stages, which is beneficial to retain and utilize the useful details. Finally, point-level features and segment-level features are fused to obtain the semantic segmentation result. Extensive experiments on SPG and SketchSeg-150K show that the proposed method achieves state-of-the-art performance.

AAAI Conference 2023 Conference Paper

AEC-GAN: Adversarial Error Correction GANs for Auto-Regressive Long Time-Series Generation

  • Lei Wang
  • Liang Zeng
  • Jian Li

Large-scale high-quality data is critical for training modern deep neural networks. However, data acquisition can be costly or time-consuming for many time-series applications, thus researchers turn to generative models for generating synthetic time-series data. In particular, recent generative adversarial networks (GANs) have achieved remarkable success in time-series generation. Despite their success, existing GAN models typically generate the sequences in an auto-regressive manner, and we empirically observe that they suffer from severe distribution shifts and bias amplification, especially when generating long sequences. To resolve this problem, we propose Adversarial Error Correction GAN (AEC-GAN), which is capable of dynamically correcting the bias in the past generated data to alleviate the risk of distribution shifts and thus can generate high-quality long sequences. AEC-GAN contains two main innovations: (1) We develop an error correction module to mitigate the bias. In the training phase, we adversarially perturb the realistic time-series data and then optimize this module to reconstruct the original data. In the generation phase, this module can act as an efficient regulator to detect and mitigate the bias. (2) We propose an augmentation method to facilitate GAN's training by introducing adversarial examples. Thus, AEC-GAN can generate high-quality sequences of arbitrary lengths, and the synthetic data can be readily applied to downstream tasks to boost their performance. We conduct extensive experiments on six widely used datasets and three state-of-the-art time-series forecasting models to evaluate the quality of our synthetic time-series data in different lengths and downstream tasks. Both the qualitative and quantitative experimental results demonstrate the superior performance of AEC-GAN over other deep generative models for time-series generation.

AAAI Conference 2023 Conference Paper

Alignment-Enriched Tuning for Patch-Level Pre-trained Document Image Models

  • Lei Wang
  • Jiabang He
  • Xing Xu
  • Ning Liu
  • Hui Liu

Alignment between image and text has shown promising improvements on patch-level pre-trained document image models. However, investigating more effective or finer-grained alignment techniques during pre-training requires a large amount of computation cost and time. Thus, a question naturally arises: Could we fine-tune the pre-trained models adaptive to downstream tasks with alignment objectives and achieve comparable or better performance? In this paper, we propose a new model architecture with alignment-enriched tuning (dubbed AETNet) upon pre-trained document image models, to adapt downstream tasks with the joint task-specific supervised and alignment-aware contrastive objective. Specifically, we introduce an extra visual transformer as the alignment-ware image encoder and an extra text transformer as the alignment-ware text encoder before multimodal fusion. We consider alignment in the following three aspects: 1) document-level alignment by leveraging the cross-modal and intra-modal contrastive loss; 2) global-local alignment for modeling localized and structural information in document images; and 3) local-level alignment for more accurate patch-level information. Experiments on various downstream tasks show that AETNet can achieve state-of-the-art performance on various downstream tasks. Notably, AETNet consistently outperforms state-of-the-art pre-trained models, such as LayoutLMv3 with fine-tuning techniques, on three different downstream tasks. Code is available at https://github.com/MAEHCM/AET.

EAAI Journal 2023 Journal Article

Autonomous dispatch trajectory planning on flight deck: A search-resampling-optimization framework

  • Xinwei Wang
  • Bai Li
  • Xichao Su
  • Haijun Peng
  • Lei Wang
  • Chen Lu
  • Chao Wang

There is a growing expectation to realize the autonomous dispatch on flight deck, where dispatch trajectory planning is seen as the key technique. Optimal-control based method has shown great advantages in high degree of constraint satisfaction over its counterparts in the last decade. However, it suffers from low computational efficiency even numerical divergence under scenarios with complicated obstacles. To deal with such an issue, a search-resampling-optimization (SRO) framework is proposed in this paper. A hybrid A* algorithm is employed to generate a coarse path according to the boundary conditions in the search stage. Then a resampling process is implemented to pave a series of safe dispatch corridors (SDCs) along the coarse path. Finally, by replacing the common one-to-one collision-avoidance with the constructed within-SDC constraints, an optimal control problem whose scale is totally independent of the number of obstacles can be formulated. The resampled result is further fed into the optimization stage to facilitate the numerical solution. Dispatch trajectory planning for taxiing aircraft and tractor can be treated uniformly under this framework. And numerical simulations demonstrate that the SRO framework is efficient and robust even with narrow accessible tunnels. The SRO is inherently flexible and can be easily extended to the trajectory planning problem in other fields. A video of the main idea and numerical simulations in this paper is available at www. bilibili. com/video/BV1tP4y1d7xy/.

EAAI Journal 2023 Journal Article

Conditional probability based multi-objective cooperative task assignment for heterogeneous UAVs

  • Xiaohua Gao
  • Lei Wang
  • Xinyong Yu
  • Xichao Su
  • Yu Ding
  • Chen Lu
  • Haijun Peng
  • Xinwei Wang

In actual air combat, there is an inevitable risk that an unmanned aerial vehicle (UAV) will be destroyed. However, this risk is rarely considered in the mission planning phase. In this paper, we focus on cooperative mission assignment for heterogeneous UAVs. We develop a multi-objective optimization model to find a balance between mission gains and UAV losses. The objective function is expressed using conditional probability theory by introducing the probabilities of mission success and UAV loss. Munitions loading capacity, time constraints, and priority constraints are modeled as constraints. To solve this combinatorial problem, an improved multi-objective genetic algorithm, which incorporates a natural chromosome encoding format and specially designed genetic operators, is developed. An efficient unlocking method is constructed to address the unavoidable dead-lock phenomenon meanwhile maintaining the population randomness. Numerical simulations for different problem sizes and ammunition stocks are performed, and the proposed algorithm is compared with the Multi-objective Particle Swarm Optimization and the Multi-objective Grey Wolf Optimization, respectively, using different unlocking approaches. The simulation and comparison results demonstrate the practical value and effectiveness of the developed model and the proposed algorithm.

YNIMG Journal 2023 Journal Article

Exploring the neural mechanisms underlying achalasia: A study of functional connectivity and regional brain activity

  • Nina Zhang
  • Binyu Teng
  • Xinyi Lu
  • Liangliang Shi
  • Li Liu
  • Fan Zhou
  • Ni Jiang
  • Xin Zhang

BACKGROUND AND AIMS: The pathophysiology of achalasia, which involves central nuclei abnormalities, remains unknown. We investigated the resting-state functional MRI (rs-fMRI) features of patients with achalasia. METHODS: We applied resting-state functional MRI (rs-fMRI) to investigate the brain features in patients with achalasia (n = 27), compared to healthy controls (n = 29). Focusing on three regions of interest (ROIs): the dorsal motor nucleus of the vagus (DMV), the nucleus ambiguus (NA), and the nucleus of the solitary tract (NTS), we analyzed variations in resting-state functional connectivity (rs-FC), fractional amplitude of low-frequency fluctuations (fALFF), and regional homogeneity (ReHo). RESULTS: Achalasia patients demonstrated stronger functional connectivity between the NA and the right precentral gyrus, left postcentral gyrus, and left insula. No significant changes were found in the DMV or NTS. The fMRI analysis showed higher rs-FC values for NA-DMV and NA-NTS connections in achalasia patients. Achalasia patients exhibited decreased fALFF values in the NA, DMV, and NTS regions, as well as increased ReHo values in the NA and DMV regions. A positive correlation was observed between fALFF values in all six ROIs and the width of the barium meal. The NTS fALFF value and NA ReHo value displayed a positive correlation with integrated relaxation pressure (IRP), while the ReHo value in the right precentral gyrus showed an inverse correlation with the height of the barium meal. CONCLUSIONS: Abnormal rs-FC and regional brain activity was found in patients with achalasia. Our study provides new insights into the pathophysiology of achalasia and highlights the potential of rs-fMRI in improving the diagnosis and treatment of this condition.

AAAI Conference 2023 Conference Paper

Generalizing Math Word Problem Solvers via Solution Diversification

  • Zhenwen Liang
  • Jipeng Zhang
  • Lei Wang
  • Yan Wang
  • Jie Shao
  • Xiangliang Zhang

Current math word problem (MWP) solvers are usually Seq2Seq models trained by the (one-problem; one-solution) pairs, each of which is made of a problem description and a solution showing reasoning flow to get the correct answer. However, one MWP problem naturally has multiple solution equations. The training of an MWP solver with (one-problem; one-solution) pairs excludes other correct solutions, and thus limits the generalizability of the MWP solver. One feasible solution to this limitation is to augment multiple solutions to a given problem. However, it is difficult to collect diverse and accurate augment solutions through human efforts. In this paper, we design a new training framework for an MWP solver by introducing a solution buffer and a solution discriminator. The buffer includes solutions generated by an MWP solver to encourage the training data diversity. The discriminator controls the quality of buffered solutions to participate in training. Our framework is flexibly applicable to a wide setting of fully, semi-weakly and weakly supervised training for all Seq2Seq MWP solvers. We conduct extensive experiments on a benchmark dataset Math23k and a new dataset named Weak12k, and show that our framework improves the performance of various MWP solvers under different settings by generating correct and diverse solutions.

AAAI Conference 2023 Conference Paper

High-Level Semantic Feature Matters Few-Shot Unsupervised Domain Adaptation

  • Lei Yu
  • Wanqi Yang
  • Shengqi Huang
  • Lei Wang
  • Ming Yang

In few-shot unsupervised domain adaptation (FS-UDA), most existing methods followed the few-shot learning (FSL) methods to leverage the low-level local features (learned from conventional convolutional models, e.g., ResNet) for classification. However, the goal of FS-UDA and FSL are relevant yet distinct, since FS-UDA aims to classify the samples in target domain rather than source domain. We found that the local features are insufficient to FS-UDA, which could introduce noise or bias against classification, and not be used to effectively align the domains. To address the above issues, we aim to refine the local features to be more discriminative and relevant to classification. Thus, we propose a novel task-specific semantic feature learning method (TSECS) for FS-UDA. TSECS learns high-level semantic features for image-to-class similarity measurement. Based on the high-level features, we design a cross-domain self-training strategy to leverage the few labeled samples in source domain to build the classifier in target domain. In addition, we minimize the KL divergence of the high-level feature distributions between source and target domains to shorten the distance of the samples between the two domains. Extensive experiments on DomainNet show that the proposed method significantly outperforms SOTA methods in FS-UDA by a large margin (i.e., ~10%).

JBHI Journal 2023 Journal Article

Improving Protein Function Prediction by Adaptively Fusing Information From Protein Sequences and Biomedical Literature

  • Yingwen Zhao
  • Zhihao Yang
  • Yongkai Hong
  • Lei Wang
  • Yin Zhang
  • Hongfei Lin
  • Jian Wang

Proteins are the main undertakers of life activities, and accurately predicting their biological functions can help human better understand life mechanism and promote the development of themselves. With the rapid development of high-throughput technologies, an abundance of proteins are discovered. However, the gap between proteins and function annotations is still huge. To accelerate the process of protein function prediction, some computational methods taking advantage of multiple data have been proposed. Among these methods, the deep-learning-based methods are currently the most popular for their capability of learning information automatically from raw data. However, due to the diversity and scale difference between data, it is challenging for existing deep learning methods to capture related information from different data effectively. In this paper, we introduce a deep learning method that can adaptively learn information from protein sequences and biomedical literature, namely DeepAF. DeepAF first extracts the two kinds of information by using different extractors, which are built based on pre-trained language models and can capture rudimentary biological knowledge. Then, to integrate those information, it performs an adaptive fusion layer based on a Cross-attention mechanism that considers the knowledge of mutual interactions between two information. Finally, based on the mixed information, DeepAF utilizes logistic regression to obtain prediction scores. The experimental results on the datasets of two species (i. e. , Human and Yeast ) show that DeepAF outperforms other state-of-the-art approaches.

AAAI Conference 2023 Conference Paper

Learning to Generate an Unbiased Scene Graph by Using Attribute-Guided Predicate Features

  • Lei Wang
  • Zejian Yuan
  • Badong Chen

Scene Graph Generation (SGG) aims to capture the semantic information in an image and build a structured representation, which facilitates downstream tasks. The current challenge in SGG is to tackle the biased predictions caused by the long-tailed distribution of predicates. Since multiple predicates in SGG are coupled in an image, existing data re-balancing methods cannot completely balance the head and tail predicates. In this work, a decoupled learning framework is proposed for unbiased scene graph generation by using attribute-guided predicate features to construct a balanced training set. Specifically, the predicate recognition is decoupled into Predicate Feature Representation Learning (PFRL) and predicate classifier training with a class-balanced predicate feature set, which is constructed by our proposed Attribute-guided Predicate Feature Generation (A-PFG) model. In the A-PFG model, we first define the class labels of and corresponding visual feature as attributes to describe a predicate. Then the predicate feature and the attribute embedding are mapped into a shared hidden space by a dual Variational Auto-encoder (VAE), and finally the synthetic predicate features are forced to learn the contextual information in the attributes via cross reconstruction and distribution alignment. To demonstrate the effectiveness of our proposed method, our decoupled learning framework and A-PFG model are applied to various SGG models. The empirical results show that our method is substantially improved on all benchmarks and achieves new state-of-the-art performance for unbiased scene graph generation. Our code is available at https://github.com/wanglei0618/A-PFG.

ICRA Conference 2023 Conference Paper

MonoPGC: Monocular 3D Object Detection with Pixel Geometry Contexts

  • Zizhang Wu
  • Yuanzhu Gan
  • Lei Wang
  • Guilian Chen
  • Jian Pu

Monocular 3D object detection reveals an economical but challenging task in autonomous driving. Recently center-based monocular methods have developed rapidly with a great trade-off between speed and accuracy, where they usually depend on the object center's depth estimation via 2D features. However, the visual semantic features without sufficient pixel geometry information, may affect the performance of clues for spatial 3D detection tasks. To alleviate this, we propose MonoPGC, a novel end-to-end Monocular 3D object detection framework with rich Pixel Geometry Contexts. We introduce the pixel depth estimation as our auxiliary task and design depth cross-attention pyramid module (DCPM) to inject local and global depth geometry knowledge into visual features. In addition, we present the depth-space-aware transformer (DSAT) to integrate 3D space position and depth-aware features efficiently. Besides, we design a novel depth-gradient positional encoding (DGPE) to bring more distinct pixel geometry contexts into the transformer for better object detection. Extensive experiments demonstrate that our method achieves the state-of-the-art performance on the KITTI dataset.

ICRA Conference 2023 Conference Paper

MVFusion: Multi-View 3D Object Detection with Semantic-aligned Radar and Camera Fusion

  • Zizhang Wu
  • Guilian Chen
  • Yuanzhu Gan
  • Lei Wang
  • Jian Pu

Multi-view radar-camera fused 3D object detection provides a farther detection range and more helpful features for autonomous driving, especially under adverse weather. The current radar-camera fusion methods deliver kinds of designs to fuse radar information with camera data. However, these fusion approaches usually adopt the straightforward concatenation operation between multi-modal features, which ignores the semantic alignment with radar features and sufficient correlations across modals. In this paper, we present MVFusion, a novel Multi-View radar-camera Fusion method to achieve semantic-aligned radar features and enhance the cross-modal information interaction. To achieve so, we inject the semantic alignment into the radar features via the semantic-aligned radar encoder (SARE) to produce image-guided radar features. Then, we propose the radar-guided fusion transformer (RGFT) to fuse our radar and image features to strengthen the two modals' correlation from the global scope via the cross-attention mechanism. Extensive experiments show that MVFusion achieves state-of-the-art performance (51. 7% NDS and 45. 3% mAP) on the nuScenes dataset. We shall release our code and trained networks upon publication.

NeurIPS Conference 2023 Conference Paper

Online PCA in Converging Self-consistent Field Equations

  • Xihan Li
  • Xiang Chen
  • Rasul Tutunov
  • Haitham Bou Ammar
  • Lei Wang
  • Jun Wang

Self-consistent Field (SCF) equation is a type of nonlinear eigenvalue problem in which the matrix to be eigen-decomposed is a function of its own eigenvectors. It is of great significance in computational science for its connection to the Schrödinger equation. Traditional fixed-point iteration methods for solving such equations suffer from non-convergence issues. In this work, we present a novel perspective on such SCF equations as a principal component analysis (PCA) for non-stationary time series, in which a distribution and its own top principal components are mutually updated over time, and the equilibrium state of the model corresponds to the solution of the SCF equations. By the new perspective, online PCA techniques are able to engage in so as to enhance the convergence of the model towards the equilibrium state, acting as a new set of tools for converging the SCF equations. With several numerical adaptations, we then develop a new algorithm for converging the SCF equation, and demonstrated its high convergence capacity with experiments on both synthesized and real electronic structure scenarios.

EAAI Journal 2023 Journal Article

Optimization of high-performance concrete mix ratio design using machine learning

  • Bin Chen
  • Lei Wang
  • Zongbao Feng
  • Yang Liu
  • Xianguo Wu
  • Yawei Qin
  • Lingyu Xia

High-durability concrete is required in extremely cold or ocean environments, making the design of concrete mixes highly important and complicated. In this study, a hybrid intelligent framework for multi-objective optimization based on random forest (RF) and the non-dominated sorting genetic algorithm version II (NSGA-II) is developed to efficiently predict concrete durability and optimize the concrete mix ratio. The relative dynamic elastic modulus of concrete after 300 freeze–thaw cycles and the chloride ion permeability coefficient at 28 days are defined as the standard measures of durability. The concrete mix ratio is taken as the influencing parameter, and orthogonal test data and engineering practice data are collected as the datasets. The proposed framework is applied to a realistic expressway project in a cold region of China. The results demonstrate that (1) a hybrid intelligent framework based on RF-NSGA-II can effectively predict concrete durability and optimize the mix ratio. (2) The developed RF model has an excellent regression learning ability, while the goodness of fit (R2) of concrete durability reaches 0. 9503 and 0. 9551, respectively, with root mean square error (RMSE) values of only 0. 096 and 0. 043, the mean absolute percentage error (MAPE) values of 2. 54% and 2. 17%. (3) After optimization, the concrete durability reaches a high standard, with a frost resistance of >95% and a chloride ion permeability coefficient of <3*10 − 8 cm2/s, at a unit volume cost of only 376. 77 yuan. Hence, the proposed framework can be used to effectively optimize the concrete mix design and provide guidance for similar projects.

JBHI Journal 2023 Journal Article

PPAEDTI: Personalized Propagation Auto-Encoder Model for Predicting Drug-Target Interactions

  • Yue-Chao Li
  • Zhu-Hong You
  • Chang-Qing Yu
  • Lei Wang
  • Leon Wong
  • Lun Hu
  • Peng-Wei Hu
  • Yu-An Huang

Identifying protein targets for drugs establishes an indispensable knowledge foundation for drug repurposing and drug development. Though expensive and time-consuming, vitro trials are widely employed to discover drug targets, and the existing relevant computational algorithms still cannot satisfy the demand for real application in drug R&D with regards to the prediction accuracy and performance efficiency, which are urgently needed to be improved. To this end, we propose here the PPAEDTI model, which uses the graph personalized propagation technique to predict drug-target interactions from the known interaction network. To evaluate the prediction performance, six benchmark datasets were used for testing with some state-of-the-art methods compared. As a result, using the 5-fold cross-validation, the proposed PPAEDTI model achieves average AUCs>90% on 5 collected datasets. We also manually checked the top-20 prediction list for 2 proteins (hsa: 775 and hsa: 779) and a kind of drug (D00618), and successfully confirmed 18, 17, and 20 items from the public datasets, respectively. The experimental results indicate that, given known drug-target interactions, the PPAEDTI model can provide accurate predictions for the new ones, which is anticipated to serve as a useful tool for pharmacology research. Using the proposed model that was trained with the collected datasets, we have built a computational platform that is accessible at http://120.77.11.78/PPAEDTI/and corresponding codes and datasets are also released.

NeurIPS Conference 2023 Conference Paper

REASONER: An Explainable Recommendation Dataset with Comprehensive Labeling Ground Truths

  • Xu Chen
  • Jingsen Zhang
  • Lei Wang
  • Quanyu Dai
  • Zhenhua Dong
  • Ruiming Tang
  • Rui Zhang
  • Li Chen

Explainable recommendation has attracted much attention from the industry and academic communities. It has shown great potential to improve the recommendation persuasiveness, informativeness and user satisfaction. In the past few years, while a lot of promising explainable recommender models have been proposed, the datasets used to evaluate them still suffer from several limitations, for example, the explanation ground truths are not labeled by the real users, the explanations are mostly single-modal and around only one aspect. To bridge these gaps, in this paper, we build a new explainable recommendation dataset, which, to our knowledge, is the first contribution that provides a large amount of real user labeled multi-modal and multi-aspect explaination ground truths. In specific, we firstly develop a video recommendation platform, where a series of questions around the recommendation explainability are carefully designed. Then, we recruit about 3000 high-quality labelers with different backgrounds to use the system, and collect their behaviors and feedback to our questions. In this paper, we detail the construction process of our dataset and also provide extensive analysis on its characteristics. In addition, we develop a library, where ten well-known explainable recommender models are implemented in a unified framework. Based on this library, we build several benchmarks for different explainable recommendation tasks. At last, we present many new opportunities brought by our dataset, which are expected to promote the field of explainable recommendation. Our dataset, library and the related documents have been released at https: //reasoner2023. github. io/.

AAAI Conference 2023 Conference Paper

Reject Decoding via Language-Vision Models for Text-to-Image Synthesis

  • Fuxiang Wu
  • Liu Liu
  • Fusheng Hao
  • Fengxiang He
  • Lei Wang
  • Jun Cheng

Transformer-based text-to-image synthesis generates images from abstractive textual conditions and achieves prompt results. Since transformer-based models predict visual tokens step by step in testing, where the early error is hard to be corrected and would be propagated. To alleviate this issue, the common practice is drawing multi-paths from the transformer-based models and re-ranking the multi-images decoded from multi-paths to find the best one and filter out others. Therefore, the computing procedure of excluding images may be inefficient. To improve the effectiveness and efficiency of decoding, we exploit a reject decoding algorithm with tiny multi-modal models to enlarge the searching space and exclude the useless paths as early as possible. Specifically, we build tiny multi-modal models to evaluate the similarities between the partial paths and the caption at multi scales. Then, we propose a reject decoding algorithm to exclude some lowest quality partial paths at the inner steps. Thus, under the same computing load as the original decoding, we could search across more multi-paths to improve the decoding efficiency and synthesizing quality. The experiments conducted on the MS-COCO dataset and large-scale datasets show that the proposed reject decoding algorithm can exclude the useless paths and enlarge the searching paths to improve the synthesizing quality by consuming less time.

EAAI Journal 2023 Journal Article

Temporal transaction information-aware Ponzi scheme detection for ethereum smart contracts

  • Lei Wang
  • Hao Cheng
  • Zibin Zheng
  • Aijun Yang
  • Ming Xu

In recent years, the frenetic advances of blockchain techniques have promoted the large-scale application of cryptocurrency and attracted significant attention in the mushrooming applications of decentralized finance (DeFi). To guarantee the health of a DeFi ecosystem, it is critical to reduce the transaction risks in a DeFi system. In particular, as a representative DeFi ecosystem platform, Ethereum’s transaction process is mainly carried out with the help of smart contracts. Due to (pseudo)anonymity, the transaction process of Ethereum users is challenged by severe fraud threats. Ponzi scheme is the typical one. Previous studies have used machine learning methods to build Ponzi scheme detection models based on learning from the identified static smart contract samples feature data. However, in the early stage of smart contract deployment, the Ponzi scheme is difficult to detect. With the progress of transactions, Ponzi scheme will gradually show its characteristics. The existing methods are still falling short in capturing the temporal features of smart contracts for detecting Ponzi schemes in the big data environment. The recognition rate of the current approaches needs to be further improved. In this paper, we propose TTPS, a Long Short-Term Memory (LSTM) Ponzi scheme detection method considering time series transaction information of smart contracts. TTPS considers both temporal account features and code features of smart contracts. Adaptive synthetic sampling (ADASYN) is employed to effectively extend the feature data of minority class Ponzi scheme small samples. LSTM is utilized to learn from the temporal feature data of Ponzi scheme samples for TTPS model training. Experimental results verify and demonstrate the effectiveness and efficiency of TTPS.

TMLR Journal 2023 Journal Article

Trip-ROMA: Self-Supervised Learning with Triplets and Random Mappings

  • Wenbin Li
  • Xuesong Yang
  • Meihao Kong
  • Lei Wang
  • Jing Huo
  • Yang Gao
  • Jiebo Luo

Contrastive self-supervised learning (SSL) methods, such as MoCo and SimCLR, have achieved great success in unsupervised visual representation learning. They rely on a large number of negative pairs and thus require either large memory banks or large batches. Some recent non-contrastive SSL methods, such as BYOL and SimSiam, attempt to discard negative pairs and have also shown remarkable performance. To avoid collapsed solutions caused by not using negative pairs, these methods require non-trivial asymmetry designs. However, in small data regimes, we can not obtain a sufficient number of negative pairs or effectively avoid the over-fitting problem when negatives are not used at all. To address this situation, we argue that negative pairs are still important but one is generally sufficient for each positive pair. We show that a simple Triplet-based loss (Trip) can achieve surprisingly good performance without requiring large batches or asymmetry designs. Moreover, to alleviate the over-fitting problem in small data regimes and further enhance the effect of Trip, we propose a simple plug-and-play RandOm MApping (ROMA) strategy by randomly mapping samples into other spaces and requiring these randomly projected samples to satisfy the same relationship indicated by the triplets. Integrating the triplet-based loss with random mapping, we obtain the proposed method Trip-ROMA. Extensive experiments, including unsupervised representation learning and unsupervised few-shot learning, have been conducted on ImageNet-1K and seven small datasets. They successfully demonstrate the effectiveness of Trip-ROMA and consistently show that ROMA can further effectively boost other SSL methods. Code is available at https://github.com/WenbinLee/Trip-ROMA.

YNICL Journal 2022 Journal Article

Detection of emerging neurodegeneration using Bayesian linear mixed-effect modeling

  • Yann Cobigo
  • Matthew S. Goh
  • Amy Wolf
  • Adam M. Staffaroni
  • John Kornak
  • Bruce L. Miller
  • Gil D. Rabinovici
  • William W. Seeley

Early detection of neurodegeneration, and prediction of when neurodegenerative diseases will lead to symptoms, are critical for developing and initiating disease modifying treatments for these disorders. While each neurodegenerative disease has a typical pattern of early changes in the brain, these disorders are heterogeneous, and early manifestations can vary greatly across people. Methods for detecting emerging neurodegeneration in any part of the brain are therefore needed. Prior publications have described the use of Bayesian linear mixed-effects (BLME) modeling for characterizing the trajectory of change across the brain in healthy controls and patients with neurodegenerative disease. Here, we use an extension of such a model to detect emerging neurodegeneration in cognitively healthy individuals at risk for dementia. We use BLME to quantify individualized rates of volume loss across the cerebral cortex from the first two MRIs in each person and then extend the BLME model to predict future values for each voxel. We then compare observed values at subsequent time points with the values that were expected from the initial rates of change and identify voxels that are lower than the expected values, indicating accelerated volume loss and neurodegeneration. We apply the model to longitudinal imaging data from cognitively normal participants in the Alzheimer's Disease Neuroimaging Initiative (ADNI), some of whom subsequently developed dementia, and two cognitively normal cases who developed pathology-proven frontotemporal lobar degeneration (FTLD). These analyses identified regions of accelerated volume loss prior to or accompanying the earliest symptoms, and expanding across the brain over time, in all cases. The changes were detected in regions that are typical for the likely diseases affecting each patient, including medial temporal regions in patients at risk for Alzheimer's disease, and insular, frontal, and/or anterior/inferior temporal regions in patients with likely or proven FTLD. In the cases where detailed histories were available, the first regions identified were consistent with early symptoms. Furthermore, survival analysis in the ADNI cases demonstrated that the rate of spread of accelerated volume loss across the brain was a statistically significant predictor of time to conversion to dementia. This method for detection of neurodegeneration is a potentially promising approach for identifying early changes due to a variety of diseases, without prior assumptions about what regions are most likely to be affected first in an individual.

NeurIPS Conference 2022 Conference Paper

Distributed Online Convex Optimization with Compressed Communication

  • Zhipeng Tu
  • Xi Wang
  • Yiguang Hong
  • Lei Wang
  • Deming Yuan
  • Guodong Shi

We consider a distributed online convex optimization problem when streaming data are distributed among computing agents over a connected communication network. Since the data are high-dimensional or the network is large-scale, communication load can be a bottleneck for the efficiency of distributed algorithms. To tackle this bottleneck, we apply the state-of-art data compression scheme to the fundamental GD-based distributed online algorithms. Three algorithms with difference-compressed communication are proposed for full information feedback (DC-DOGD), one-point bandit feedback (DC-DOBD), and two-point bandit feedback (DC-DO2BD), respectively. We obtain regret bounds explicitly in terms of time horizon, compression ratio, decision dimension, agent number, and network parameters. Our algorithms are proved to be no-regret and match the same regret bounds, w. r. t. time horizon, with their uncompressed versions for both convex and strongly convex losses. Numerical experiments are given to validate the theoretical findings and illustrate that the proposed algorithms can effectively reduce the total transmitted bits for distributed online training compared with the uncompressed baseline.

NeurIPS Conference 2022 Conference Paper

Domain Generalization by Learning and Removing Domain-specific Features

  • Yu Ding
  • Lei Wang
  • Bin Liang
  • Shuming Liang
  • Yang Wang
  • Fang Chen

Deep Neural Networks (DNNs) suffer from domain shift when the test dataset follows a distribution different from the training dataset. Domain generalization aims to tackle this issue by learning a model that can generalize to unseen domains. In this paper, we propose a new approach that aims to explicitly remove domain-specific features for domain generalization. Following this approach, we propose a novel framework called Learning and Removing Domain-specific features for Generalization (LRDG) that learns a domain-invariant model by tactically removing domain-specific features from the input images. Specifically, we design a classifier to effectively learn the domain-specific features for each source domain, respectively. We then develop an encoder-decoder network to map each input image into a new image space where the learned domain-specific features are removed. With the images output by the encoder-decoder network, another classifier is designed to learn the domain-invariant features to conduct image classification. Extensive experiments demonstrate that our framework achieves superior performance compared with state-of-the-art methods.

AAAI Conference 2022 Conference Paper

FactorVAE: A Probabilistic Dynamic Factor Model Based on Variational Autoencoder for Predicting Cross-Sectional Stock Returns

  • Yitong Duan
  • Lei Wang
  • Qizhong Zhang
  • Jian Li

As an asset pricing model in economics and finance, factor model has been widely used in quantitative investment. Towards building more effective factor models, recent years have witnessed the paradigm shift from linear models to more flexible nonlinear data-driven machine learning models. However, due to low signal-to-noise ratio of the financial data, it is quite challenging to learn effective factor models. In this paper, we propose a novel factor model, FactorVAE, as a probabilistic model with inherent randomness for noise modeling. Essentially, our model integrates the dynamic factor model (DFM) with the variational autoencoder (VAE) in machine learning, and we propose a prior-posterior learning method based on VAE, which can effectively guide the learning of model by approximating an optimal posterior factor model with future information. Particularly, considering that risk modeling is important for the noisy stock data, Factor- VAE can estimate the variances from the distribution over the latent space of VAE, in addition to predicting returns. The experiments on the real stock market data demonstrate the effectiveness of FactorVAE, which outperforms various baseline methods.

EAAI Journal 2022 Journal Article

Hierarchical pyramid attentive network with spatial separable convolution for crowd counting

  • Shihui Zhang
  • Xiaoxiao Zhang
  • He Li
  • Huan He
  • Dandan Song
  • Lei Wang

To tackle the challenging scale variation issue of the crowd counting task so as to improve the counting accuracy, we present a novel method based on Hierarchical Pyramid Attentive Network (HPANet) for crowd counting. Specifically, a Scale-aware Pyramid Attentive (SPA) block, extracting the rich multi-scale context, is designed elaborately as using the two-branch spatial separable convolution as its core component to replace the conventional pure convolution with larger kernel size to reduce the computation, as well as adopting a self-attention operation for the spatial feature aggregation. In order to further learn the scale-aware feature representation well from the input image, we stack the designed SPA block in a hierarchical way and fuse their features flexibly as another crucial module of the proposed HPANet, the Hierarchical Feature Fusion (HFF) module. Combining the designed SPA block and HFF module, the developed HPANet could remedy the scale variation issue and thus improve the counting performance with the mighty scale-aware feature representation. The performance of the HPANet is evaluated on four public available benchmark datasets in this paper, including ShanghaiTech, Mall, Beijing BRT and UCF-QNRF. Extensive experimental results on benchmarks demonstrate that the proposed HPANet could have an effective performance for crowd counting and the ablation experimental results validate the effectiveness of the components of HPANet on the counting task. The designed HPANet could realize a preferable counting performance in view of alleviating the scale variation issue, without the cost of introducing too much additional parameters for the multi-column structure.

NeurIPS Conference 2022 Conference Paper

Improving Barely Supervised Learning by Discriminating Unlabeled Samples with Super-Class

  • Guan Gui
  • Zhen Zhao
  • Lei Qi
  • Luping Zhou
  • Lei Wang
  • Yinghuan Shi

In semi-supervised learning (SSL), a common practice is to learn consistent information from unlabeled data and discriminative information from labeled data to ensure both the immutability and the separability of the classification model. Existing SSL methods suffer from failures in barely-supervised learning (BSL), where only one or two labels per class are available, as the insufficient labels cause the discriminative information being difficult or even infeasible to learn. To bridge this gap, we investigate a simple yet effective way to leverage unlabeled samples for discriminative learning, and propose a novel discriminative information learning module to benefit model training. Specifically, we formulate the learning objective of discriminative information at the super-class level and dynamically assign different classes into different super-classes based on model performance improvement. On top of this on-the-fly process, we further propose a distribution-based loss to learn discriminative information by utilizing the similarity relationship between samples and super-classes. It encourages the unlabeled samples to stay closer to the distribution of their corresponding super-class than those of others. Such a constraint is softer than the direct assignment of pseudo labels, while the latter could be very noisy in BSL. We compare our method with state-of-the-art SSL and BSL methods through extensive experiments on standard SSL benchmarks. Our method can achieve superior results, \eg, an average accuracy of 76. 76\% on CIFAR-10 with merely 1 label per class.

YNIMG Journal 2022 Journal Article

Investigating the temporal pattern of neuroimaging-based brain age estimation as a biomarker for Alzheimer's Disease related neurodegeneration

  • Alexei Taylor
  • Fengqing Zhang
  • Xin Niu
  • Ashley Heywood
  • Jane Stocks
  • Gangyi Feng
  • Karteek Popuri
  • Mirza Faisal Beg

Neuroimaging-based brain-age estimation via machine learning has emerged as an important new approach for studying brain aging. The difference between one's estimated brain age and chronological age, the brain age gap (BAG), has been proposed as an Alzheimer's Disease (AD) biomarker. However, most past studies on the BAG have been cross-sectional. Quantifying longitudinal changes in an individual's BAG temporal pattern would likely improve prediction of AD progression and clinical outcome based on neurophysiological changes. To fill this gap, our study conducted predictive modeling using a large neuroimaging dataset with up to 8 years of follow-up to examine the temporal patterns of the BAG's trajectory and how it varies by subject-level characteristics (sex, APOEɛ4 carriership) and disease status. Specifically, we explored the pattern and rate of change in BAG over time in individuals who remain stable with normal cognition or mild cognitive impairment (MCI), as well as individuals who progress to clinical AD. Combining multimodal imaging data in a support vector regression model to estimate brain age yielded improved performance over single modality. Multilevel modeling results showed the BAG followed a linear increasing trajectory with a significantly faster rate in individuals with MCI who progressed to AD compared to cognitively normal or MCI individuals who did not progress. The dynamic changes in the BAG during AD progression were further moderated by sex and APOEɛ4 carriership. Our findings demonstrate the BAG as a potential biomarker for understanding individual specific temporal patterns related to AD progression.

AAAI Conference 2022 Conference Paper

LaSSL: Label-Guided Self-Training for Semi-supervised Learning

  • Zhen Zhao
  • Luping Zhou
  • Lei Wang
  • Yinghuan Shi
  • Yang Gao

The key to semi-supervised learning (SSL) is to explore adequate information to leverage the unlabeled data. Current dominant approaches aim to generate pseudolabels on weakly augmented instances and train models on their corresponding strongly augmented variants with high-confidence results. However, such methods are limited in excluding samples with low-confidence pseudo-labels and under-utilization of the label information. In this paper, we emphasize the cruciality of the label information and propose a Label-guided Self-training approach to Semi-supervised Learning (LaSSL), which improves pseudo-label generations from two mutually boosted strategies. First, with the ground-truth labels and iteratively-polished pseudolabels, we explore instance relations among all samples and then minimize a class-aware contrastive loss to learn discriminative feature representations that make same-class samples gathered and different-class samples scattered. Second, on top of improved feature representations, we propagate the label information to the unlabeled samples across the potential data manifold at the feature-embedding level, which can further improve the labelling of samples with reference to their neighbours. These two strategies are seamlessly integrated and mutually promoted across the whole training process. We evaluate LaSSL on several classification benchmarks under partially labeled settings and demonstrate its superiority over the state-of-the-art approaches.

JBHI Journal 2022 Journal Article

MDADP: A Webserver Integrating Database and Prediction Tools for Microbe-Disease Associations

  • Lei Wang
  • Hao Li
  • Yuqi Wang
  • Yihong Tan
  • Zhiping Chen
  • Tingrui Pei
  • Quan Zou

More and more evidence has demonstrated that microbiota play important roles in the life processes of the human body. In recent years, various computational methods have been proposed for identifying potentially disease-associated microbes to save costs in traditional biological experiments. However, prediction performances of these methods are generally limited by outdated and incomplete datasets. And moreover, until now, there are limited studies that can provide visual predictive tools for inferring possible microbe-disease associations (MDAs) as well. Hence, in this manuscript, a novel webserver called MDADP will be proposed to identify latent MDAs, in which, a new MDA database together with interactive prediction tools for MDAs studies will be designed simultaneously. Especially, in the newly constructed MDA database, 2019 known MDAs between 58 diseases and 703 microbes have been manually collected first. And then, through adopting the average ranking method and the co-confidence method respectively, eight representative computational models have been integrated together to identify potential disease-related microbes. As a result, MDADP can provide not only interactive features for users to access and capture MDAs entities, but alsoeffective tools for users to identify candidate microbes for different diseases. To our knowledge, MDADP is the first online platform that incorporates a new MDA database with comprehensive MDA prediction tools. Therefore, we believe that it will be a valuable source of information for researches in microbiology and disease-related fields. MDADP can be accessed at http://mdadp.leelab2997.cn.

AAAI Conference 2022 System Paper

MWPToolkit: An Open-Source Framework for Deep Learning-Based Math Word Problem Solvers

  • Yihuai Lan
  • Lei Wang
  • Qiyuan Zhang
  • Yunshi Lan
  • Bing Tian Dai
  • Yan Wang
  • Dongxiang Zhang
  • Ee-Peng Lim

While Math Word Problem (MWP) solving has emerged as a popular field of study and made great progress in recent years, most existing methods are benchmarked solely on one or two datasets and implemented with different configurations. In this paper, we introduce the first open-source library for solving MWPs called MWPToolkit, which provides a unified, comprehensive, and extensible framework for the research purpose. Specifically, we deploy 17 deep learningbased MWP solvers and 6 MWP datasets in our toolkit. These MWP solvers are advanced models for MWP solving, covering the categories of Seq2seq, Seq2Tree, Graph2Tree, and Pre-trained Language Models. And these MWP datasets are popular datasets that are commonly used as benchmarks in existing work. Our toolkit is featured with highly modularized and reusable components, which can help researchers quickly get started and develop their own models. We have released the code and documentation of MWPToolkit in https: //github. com/LYH-YF/MWPToolkit.

JBHI Journal 2022 Journal Article

NSECDA: Natural Semantic Enhancement for CircRNA-Disease Association Prediction

  • Lei Wang
  • Leon Wong
  • Zhu-Hong You
  • De-Shuang Huang
  • Xiao-Rui Su
  • Bo-Wei Zhao

Increasing evidence suggest that circRNA, as one of the most promising emerging biomarkers, has a very close relationship with diseases. Exploring the relationship between circRNA and diseases can provide novel perspective for diseases diagnosis and pathogenesis. The existing circRNA-disease association (CDA) prediction models, however, generally treat the data attributes equally, do not pay special attention to the attributes with more significant influence, and do not make full use of the correlation and symbiosis between attributes to dig into the latent semantic information of the data. Therefore, in response to the above problems, this paper proposes a natural semantic enhancement method NSECDA to predict CDA. In practical terms, we first recognize the circRNA sequence as a biological language, and analyze its natural semantic properties through the natural language understanding theory; then integrate it with disease attributes, circRNA and disease Gaussian Interaction Profile (GIP) kernel attributes, and use Graph Attention Network (GAT) to focus on the influential attributes, so as to mine the deeply hidden features; finally, the Rotation Forest (RoF) classifier was used to accurately determine CDA. In the gold standard data set CircR2Disease, NSECDA achieved 92. 49% accuracy with 0. 9225 AUC score. In comparison with the non-natural semantic enhancement model and other classifier models, NSECDA also shows competitive performance. Additionally, 25 of the CDA pairs with unknown associations in the top 30 prediction scores of NSECDA have been proven by newly reported studies. These achievements suggest that NSECDA is an effective model to predict CDA, which can provide credible candidate for subsequent wet experiments, thus significantly reducing the scope of investigations.

YNICL Journal 2021 Journal Article

FDG-PET in presymptomatic C9orf72 mutation carriers

  • Karteek Popuri
  • Mirza Faisal Beg
  • Hyunwoo Lee
  • Rakesh Balachandar
  • Lei Wang
  • Vesna Sossi
  • Claudia Jacova
  • Matt Baker

OBJECTIVE: Our aim is to investigate patterns of brain glucose metabolism using fluorodeoxyglucose positron emission tomography (FDG-PET) in presymptomatic carriers of the C9orf72 repeat expansion to better understand the early preclinical stages of frontotemporal dementia (FTD). METHODS: Structural MRI and FDG-PET were performed on clinically asymptomatic members of families with FTD caused by the C9orf72 repeat expansion (15 presymptomatic mutation carriers, C9orf72+; 20 non-carriers, C9orf72-). Regional glucose metabolism in cerebral and cerebellar gray matter was compared between groups. RESULTS: The mean age of the C9orf72+ and C9orf72- groups were 45.3 ± 10.6 and 56.0 ± 11.0 years respectively, and the mean age of FTD onset in their families was 56 ± 7 years. Compared to non-carrier controls, the C9orf72+ group exhibited regional hypometabolism, primarily involving the cingulate gyrus, frontal and temporal neocortices (left > right) and bilateral thalami. CONCLUSIONS: The C9orf72 repeat expansion is associated with changes in brain glucose metabolism that are demonstrable up to 10 years prior to symptom onset and before changes in gray matter volume become significant. These findings indicate that FDG-PET may be a particularly sensitive and useful method for investigating and monitoring the earliest stages of FTD in individuals with this underlying genetic basis.

JBHI Journal 2021 Journal Article

IMU-Based Gait Normalcy Index Calculation for Clinical Evaluation of Impaired Gait

  • Lei Wang
  • Yun Sun
  • Qingguo Li
  • Tao Liu
  • Jingang Yi

Inertial measurement units (IMU) have been used for gait analysis in many clinical studies, as a more convenient, low cost and less restricted alternative to the laboratory-based motion capture systems or instrumented walkways. Spatial-temporal gait parameters such as gait cycle duration and stride length calculated from the IMUs were often used in these studies for evaluating the impaired gait. However, the spatial-temporal information provided by IMUs is limited, and sometime suffers incomplete and less effective evaluation. In this study, we develop a novel IMU-based method for clinical gait evaluation. Nine gait variables including three spatial-temporal parameters and six kinematic parameters are extracted from two shank-mounted IMUs for quantifying patient's gait deviations. Based on those parameters, an IMU-based gait normalcy index (INI) is derived to evaluate the overall gait performance. Eight inpatient subjects with gait impairments caused by n-hexane neuropathy and ten healthy subjects were recruited. The proposed gait variables and INI were examined on the inpatients at three to five time instants during the rehabilitation process until being discharged. A comparison with healthy subjects and statistical analysis for the changes of gait variables and INI demonstrated that the proposed new set of gait variables and INI can provide adequate and effective information for quantifying gait abnormalities, and help understanding the progress of gait and effectiveness of therapy during rehabilitation process.

AIIM Journal 2021 Journal Article

Interactive medical image segmentation via a point-based interaction

  • Jian Zhang
  • Yinghuan Shi
  • Jinquan Sun
  • Lei Wang
  • Luping Zhou
  • Yang Gao
  • Dinggang Shen

Due to low tissue contrast, irregular shape, and large location variance, segmenting the objects from different medical imaging modalities (e. g. , CT, MR) is considered as an important yet challenging task. In this paper, a novel method is presented for interactive medical image segmentation with the following merits. (1) Its design is fundamentally different from previous pure patch-based and image-based segmentation methods. It is observed that during delineation, the physician repeatedly check the intensity from area inside-object to outside-object to determine the boundary, which indicates that comparison in an inside-out manner is extremely important. Thus, the method innovatively models the segmentation task as learning the representation of bi-directional sequential patches, starting from (or ending in) the given central point of the object. This can be realized by the proposed ConvRNN network embedded with a gated memory propagation unit. (2) Unlike previous interactive methods (requiring bounding box or seed points), the proposed method only asks the physician to merely click on the rough central point of the object before segmentation, which could simultaneously enhance the performance and reduce the segmentation time. (3) The method is utilized in a multi-level framework for better performance. It has been systematically evaluated in three different segmentation tasks, including CT kidney tumor, MR prostate, and PROMISE12 challenge, showing promising results compared with state-of-the-art methods.

JBHI Journal 2021 Journal Article

Lexicon Knowledge Boosted Interaction Graph Network for Adverse Drug Reaction Recognition From Social Media

  • Zhiheng Li
  • Zhihao Yang
  • Lei Wang
  • Yin Zhang
  • Hongfei Lin
  • Jian Wang

The World Health Organization underlines the significance of adverse drug reaction (ADR) reports for patients' safety. Actually, many potential ADRs tend to be under-reported in post-market ADR surveillance. Recognizing ADRs from social media is indispensably important and could complement post-market ADR surveillance for more effective pharmacovigilance studies. However, previous approaches pose two challenges: 1) ADRs show high expression variability in social media, and thus, many potential ADRs are out-of-lexicon ones, which are difficult to be recognized, and 2) most phrasal ADRs are non-standard mentions and their boundaries are difficult to identify accurately. To tackle these challenges, we design three interaction graphs and propose a neural network approach, i. e. , Interaction Graph Network (IGN). Specifically, to recognize more out-of-lexicon ADRs, besides the mentions in ADR lexicon, noun phrases in the input sentence are regarded as candidate phrases and their features are taken into considerations. Moreover, in an attempt to accurately identify ADR boundaries, three word-phrase interaction graphs are designed to represent lexicon knowledge and are encoded using graph attention networks (GATs) to directly integrate various boundary and contextual information of candidate phrases into ADR recognition. Experimental results on two benchmark datasets show that IGN can recognize ADR accurately and consistently outperforms other state-of-the-art approaches.

JBHI Journal 2021 Journal Article

Non-invasive Monitoring of Three Glucose Ranges Based On ECG By Using DBSCAN-CNN

  • Jingzhen Li
  • Igbe Tobore
  • Yuhang Liu
  • Abhishek Kandwal
  • Lei Wang
  • Zedong Nie

Autonomic nervous system (ANS) can maintain homeostasis through the coordination of different organs including heart. The change of blood glucose (BG) level can stimulate the ANS, which will lead to the variation of Electrocardiogram (ECG). Considering that the monitoring of different BG ranges is significant for diabetes care, in this paper, an ECG-based technique was proposed to achieve non-invasive monitoring with three BG ranges: low glucose level, moderate glucose level, and high glucose level. For this purpose, multiple experiments that included fasting tests and oral glucose tolerance tests were conducted, and the ECG signals from 21 adults were recorded continuously. Furthermore, an approach of fusing density-based spatial clustering of applications with noise and convolution neural networks (DBSCAN-CNN) was presented for ECG preprocessing of outliers and classification of BG ranges based ECG. Also, ECG's important information, which was related to different BG ranges, was graphically visualized. The result showed that the percentages of accurate classification were 87. 94% in low glucose level, 69. 36% in moderate glucose level, and 86. 39% in high glucose level. Moreover, the visualization results revealed that the highlights of ECG for the different BG ranges were different. In addition, the sensitivity of prediabetes/diabetes screening based on ECG was up to 98. 48%, and the specificity was 76. 75%. Therefore, we conclude that the proposed approach for BG range monitoring and prediabetes/diabetes screening has potentials in practical applications.

IJCAI Conference 2020 Conference Paper

Asymmetric Distribution Measure for Few-shot Learning

  • Wenbin Li
  • Lei Wang
  • Jing Huo
  • Yinghuan Shi
  • Yang Gao
  • Jiebo Luo

The core idea of metric-based few-shot image classification is to directly measure the relations between query images and support classes to learn transferable feature embeddings. Previous work mainly focuses on image-level feature representations, which actually cannot effectively estimate a class's distribution due to the scarcity of samples. Some recent work shows that local descriptor based representations can achieve richer representations than image-level based representations. However, such works are still based on a less effective instance-level metric, especially a symmetric metric, to measure the relation between a query image and a support class. Given the natural asymmetric relation between a query image and a support class, we argue that an asymmetric measure is more suitable for metric-based few-shot learning. To that end, we propose a novel Asymmetric Distribution Measure (ADM) network for few-shot learning by calculating a joint local and global asymmetric measure between two multivariate local distributions of a query and a class. Moreover, a task-aware Contrastive Measure Strategy (CMS) is proposed to further enhance the measure function. On popular miniImageNet and tieredImageNet, ADM can achieve the state-of-the-art results, validating our innovative design of asymmetric distribution measures for few-shot learning. The source code can be downloaded from https: //github. com/WenbinLee/ADM. git.

YNICL Journal 2020 Journal Article

Brain morphometric differences in youth with and without perinatally-acquired HIV: A cross-sectional study

  • C. Paula Lewis-de los Angeles
  • Paige L. Williams
  • Lisanne M. Jenkins
  • Yanling Huo
  • Kathleen Malee
  • Kathryn I. Alpert
  • Kristina A. Uban
  • Megan M. Herting

Youth with perinatally-acquired HIV (PHIV) experience specific and global cognitive deficits at increased rates compared to typically-developing HIV-uninfected youth. In youth with PHIV, HIV infects the brain early in development. Neuroimaging studies have demonstrated altered grey matter morphometry in youth with PHIV compared to typically-developing youth. This study examined cortical thickness, surface area, and gyrification of grey matter in youth (age 11-20 years old) with PHIV (n = 40) from the Pediatric HIV/AIDS Cohort Study (PHACS) compared to typically-developing presumed HIV uninfected and unexposed youth (n = 80) from the Pediatric Imaging, Neurocognition and Genetics Study (PING) using structural magnetic resonance imaging. This study also examined the relationship between grey matter morphometry and age. Youth with PHIV had reduced cortical thickness, surface area, and gyrification compared to typically-developing youth. In addition, an inverse relationship between age and grey matter volume was found in typically-developing youth, but was not observed in youth with PHIV. Longitudinal studies are necessary to understand the neurodevelopmental trajectory of youth with PHIV.

AAAI Conference 2020 Conference Paper

Differentiable Meta-Learning Model for Few-Shot Semantic Segmentation

  • Pinzhuo Tian
  • Zhangkai Wu
  • Lei Qi
  • Lei Wang
  • Yinghuan Shi
  • Yang Gao

To address the annotation scarcity issue in some cases of semantic segmentation, there have been a few attempts to develop the segmentation model in the few-shot learning paradigm. However, most existing methods only focus on the traditional 1-way segmentation setting (i. e. , one image only contains a single object). This is far away from practical semantic segmentation tasks where the K-way setting (K >1) is usually required by performing the accurate multi-object segmentation. To deal with this issue, we formulate the fewshot semantic segmentation task as a learning-based pixel classification problem, and propose a novel framework called MetaSegNet based on meta-learning. In MetaSegNet, an architecture of embedding module consisting of the global and local feature branches is developed to extract the appropriate meta-knowledge for the few-shot segmentation. Moreover, we incorporate a linear model into MetaSegNet as a base learner to directly predict the label of each pixel for the multiobject segmentation. Furthermore, our MetaSegNet can be trained by the episodic training mechanism in an end-to-end manner from scratch. Experiments on two popular semantic segmentation datasets, i. e. , PASCAL VOC and COCO, reveal the effectiveness of the proposed MetaSegNet in the K-way few-shot semantic segmentation task.

IJCAI Conference 2020 Conference Paper

Financial Thought Experiment: A GAN-based Approach to Vast Robust Portfolio Selection

  • Chi Seng Pun
  • Lei Wang
  • Hoi Ying Wong

Modern day trading practice resembles a thought experiment, where investors imagine various possibilities of future stock market and invest accordingly. Generative adversarial network (GAN) is highly relevant to this trading practice in two ways. First, GAN generates synthetic data by a neural network that is technically indistinguishable from the reality, which guarantees the reasonableness of the experiment. Second, GAN generates multitudes of fake data, which implements half of the experiment. In this paper, we present a new architecture of GAN and adapt it to portfolio risk minimization problem by adding a regression network to GAN (implementing the second half of the experiment). The new architecture is termed GANr. Battling against two distinctive networks: discriminator and regressor, GANr's generator aims to simulate a stock market that is close to the reality while allow for all possible scenarios. The resulting portfolio resembles a robust portfolio with data-driven ambiguity. Our empirical studies show that GANr portfolio is more resilient to bleak financial scenarios than CLSGAN and LASSO portfolios.

YNICL Journal 2020 Journal Article

Outward subcortical curvature associated with sub-clinical depression symptoms in adolescents

  • Lisanne M. Jenkins
  • Jessica J. Chiang
  • Katherine Vause
  • Lauren Hoffer
  • Kathryn Alpert
  • Todd B. Parrish
  • Gregory E. Miller
  • Lei Wang

OBJECTIVE: Subclinical or subthreshold depressive symptoms (StD) are frequent in adolescence and are related to suicidality and onset of depression in adulthood, however, their neurobiology is poorly understood. We examined the relationship between StD and subcortical grey matter structures in unmedicated adolescents with no history of axis I diagnosis. METHODS: 277 youths from Chicago aged 14 years participated, undergoing a structural MRI scan and completing the Revised Children's Anxiety and Depression Scale (RCADS). Blood samples provided a composite of five pro-inflammatory cytokines. Regions of interest (ROI) for vertex-based surface analysis were the left and right amygdala, hippocampus, thalamus, caudate, nucleus accumbens, pallidum and putamen. Covariates were age, pubertal status, socioeconomic disadvantage and intracranial volume. Males and females were analysed separately. RESULTS: StD had positive associations (outward shape) with subcortical morphology in the right amygdala and left hippocampus in females, and the bilateral putamen and the left caudate, hippocampus and thalamus in males. However, we also found negative associations with StD (inward contractions) in the hippocampus in females and the caudate in males. Pro-inflammatory cytokines did not mediate the relationship between StD and outward morphology or volume. CONCLUSION: This is one of the first studies to examine subcortical morphology of basal ganglia and thalamic regions related to StD in adolescents, and the first study to report mostly positive associations between StD, volume and outward morphology in youths. These findings could reflect intact neurogenesis or resilience to depression, however longitudinal research is needed to further understand the neurobiology of StD in adolescents.

IJCAI Conference 2020 Conference Paper

Teacher-Student Networks with Multiple Decoders for Solving Math Word Problem

  • Jipeng Zhang
  • Roy Ka-Wei Lee
  • Ee-Peng Lim
  • Wei Qin
  • Lei Wang
  • Jie Shao
  • Qianru Sun

Math word problem (MWP) is challenging due to the limitation in training data where only one “standard” solution is available. MWP models often simply fit this solution rather than truly understand or solve the problem. The generalization of models (to diverse word scenarios) is thus limited. To address this problem, this paper proposes a novel approach, TSN-MD, by leveraging the teacher network to integrate the knowledge of equivalent solution expressions and then to regularize the learning behavior of the student network. In addition, we introduce the multiple-decoder student network to generate multiple candidate solution expressions by which the final answer is voted. In experiments, we conduct extensive comparisons and ablative studies on two large-scale MWP benchmarks, and show that using TSN-MD can surpass the state-of-the-art works by a large margin. More intriguingly, the visualization results demonstrate that TSN-MD not only produces correct final answers but also generates diverse equivalent expressions of the solution.

TIST Journal 2019 Journal Article

A Visual Analysis Approach for Understanding Durability Test Data of Automotive Products

  • Ying Zhao
  • Lei Wang
  • Shijie Li
  • Fangfang Zhou
  • Xiaoru Lin
  • Qiang Lu
  • Lei Ren

People face data-rich manufacturing environments in Industry 4.0. As an important technology for explaining and understanding complex data, visual analytics has been increasingly introduced into industrial data analysis scenarios. With the durability test of automotive starters as background, this study proposes a visual analysis approach for understanding large-scale and long-term durability test data. Guided by detailed scenario and requirement analyses, we first propose a migration-adapted clustering algorithm that utilizes a segmentation strategy and a group of matching-updating operations to achieve an efficient and accurate clustering analysis of the data for starting mode identification and abnormal test detection. We then design and implement a visual analysis system that provides a set of user-friendly visual designs and lightweight interactions to help people gain data insights into the test process overview, test data patterns, and durability performance dynamics. Finally, we conduct a quantitative algorithm evaluation, case study, and user interview by using real-world starter durability test datasets. The results demonstrate the effectiveness of the approach and its possible inspiration for the durability test data analysis of other similar industrial products.

IJCAI Conference 2019 Conference Paper

Coarse-to-Fine Image Inpainting via Region-wise Convolutions and Non-Local Correlation

  • Yuqing Ma
  • Xianglong Liu
  • Shihao Bai
  • Lei Wang
  • Dailan He
  • Aishan Liu

Recently deep neural networks have achieved promising performance for filling large missing regions in image inpainting tasks. They usually adopted the standard convolutional architecture over the corrupted image, where the same convolution filters try to restore the diverse information on both existing and missing regions, and meanwhile ignores the long-distance correlation among the regions. Only relying on the surrounding areas inevitably leads to meaningless contents and artifacts, such as color discrepancy and blur. To address these problems, we first propose region-wise convolutions to locally deal with the different types of regions, which can help exactly reconstruct existing regions and roughly infer the missing ones from existing regions at the same time. Then, a non-local operation is introduced to globally model the correlation among different regions, promising visual consistency between missing and existing regions. Finally, we integrate the region-wise convolutions and non-local correlation in a coarse-to-fine framework to restore semantically reasonable and visually realistic images. Extensive experiments on three widely-used datasets for image inpainting tasks have been conducted, and both qualitative and quantitative experimental results demonstrate that the proposed model significantly outperforms the state-of-the-art approaches, especially for the large irregular missing regions.

AAAI Conference 2019 Conference Paper

Coupled CycleGAN: Unsupervised Hashing Network for Cross-Modal Retrieval

  • Chao Li
  • Cheng Deng
  • Lei Wang
  • De Xie
  • Xianglong Liu

In recent years, hashing has attracted more and more attention owing to its superior capacity of low storage cost and high query efficiency in large-scale cross-modal retrieval. Benefiting from deep leaning, continuously compelling results in cross-modal retrieval community have been achieved. However, existing deep cross-modal hashing methods either rely on amounts of labeled information or have no ability to learn an accuracy correlation between different modalities. In this paper, we proposed Unsupervised coupled Cycle generative adversarial Hashing networks (UCH), for cross-modal retrieval, where outer-cycle network is used to learn powerful common representation, and inner-cycle network is explained to generate reliable hash codes. Specifically, our proposed UCH seamlessly couples these two networks with generative adversarial mechanism, which can be optimized simultaneously to learn representation and hash codes. Extensive experiments on three popular benchmark datasets show that the proposed UCH outperforms the state-of-the-art unsupervised cross-modal hashing methods.

AAAI Conference 2019 Conference Paper

Distribution Consistency Based Covariance Metric Networks for Few-Shot Learning

  • Wenbin Li
  • Jinglin Xu
  • Jing Huo
  • Lei Wang
  • Yang Gao
  • Jiebo Luo

Few-shot learning aims to recognize new concepts from very few examples. However, most of the existing few-shot learning methods mainly concentrate on the first-order statistic of concept representation or a fixed metric on the relation between a sample and a concept. In this work, we propose a novel end-to-end deep architecture, named Covariance Metric Networks (CovaMNet). The CovaMNet is designed to exploit both the covariance representation and covariance metric based on the distribution consistency for the few-shot classification tasks. Specifically, we construct an embedded local covariance representation to extract the second-order statistic information of each concept and describe the underlying distribution of this concept. Upon the covariance representation, we further define a new deep covariance metric to measure the consistency of distributions between query samples and new concepts. Furthermore, we employ the episodic training mechanism to train the entire network in an end-to-end manner from scratch. Extensive experiments in two tasks, generic few-shot image classification and fine-grained fewshot image classification, demonstrate the superiority of the proposed CovaMNet. The source code can be available from https: //github. com/WenbinLee/CovaMNet. git.

AAAI Conference 2019 Conference Paper

Template-Based Math Word Problem Solvers with Recursive Neural Networks

  • Lei Wang
  • Dongxiang Zhang
  • Jipeng Zhang
  • Xing Xu
  • Lianli Gao
  • Bing Tian Dai
  • Heng Tao Shen

The design of automatic solvers to arithmetic math word problems has attracted considerable attention in recent years and a large number of datasets and methods have been published. Among them, Math23K is the largest data corpus that is very helpful to evaluate the generality and robustness of a proposed solution. The best performer in Math23K is a seq2seq model based on LSTM to generate the math expression. However, the model suffers from performance degradation in large space of target expressions. In this paper, we propose a template-based solution based on recursive neural network for math expression construction. More specifically, we first apply a seq2seq model to predict a tree-structure template, with inferred numbers as leaf nodes and unknown operators as inner nodes. Then, we design a recursive neural network to encode the quantity with Bi-LSTM and self attention, and infer the unknown operator nodes in a bottom-up manner. The experimental results clearly establish the superiority of our new framework as we improve the accuracy by a wide margin in two of the largest datasets, i. e. , from 58. 1% to 66. 9% in Math23K and from 62. 8% to 66. 8% in MAWPS.

YNIMG Journal 2018 Journal Article

3D conditional generative adversarial networks for high-quality PET image estimation at low dose

  • Yan Wang
  • Biting Yu
  • Lei Wang
  • Chen Zu
  • David S. Lalush
  • Weili Lin
  • Xi Wu
  • Jiliu Zhou

Positron emission tomography (PET) is a widely used imaging modality, providing insight into both the biochemical and physiological processes of human body. Usually, a full dose radioactive tracer is required to obtain high-quality PET images for clinical needs. This inevitably raises concerns about potential health hazards. On the other hand, dose reduction may cause the increased noise in the reconstructed PET images, which impacts the image quality to a certain extent. In this paper, in order to reduce the radiation exposure while maintaining the high quality of PET images, we propose a novel method based on 3D conditional generative adversarial networks (3D c-GANs) to estimate the high-quality full-dose PET images from low-dose ones. Generative adversarial networks (GANs) include a generator network and a discriminator network which are trained simultaneously with the goal of one beating the other. Similar to GANs, in the proposed 3D c-GANs, we condition the model on an input low-dose PET image and generate a corresponding output full-dose PET image. Specifically, to render the same underlying information between the low-dose and full-dose PET images, a 3D U-net-like deep architecture which can combine hierarchical features by using skip connection is designed as the generator network to synthesize the full-dose image. In order to guarantee the synthesized PET image to be close to the real one, we take into account of the estimation error loss in addition to the discriminator feedback to train the generator network. Furthermore, a concatenated 3D c-GANs based progressive refinement scheme is also proposed to further improve the quality of estimated images. Validation was done on a real human brain dataset including both the normal subjects and the subjects diagnosed as mild cognitive impairment (MCI). Experimental results show that our proposed 3D c-GANs method outperforms the benchmark methods and achieves much better performance than the state-of-the-art methods in both qualitative and quantitative measures.

AAAI Conference 2018 Short Paper

A Novel Embedding Method for News Diffusion Prediction

  • Ruoran Liu
  • Qiudan Li
  • Can Wang
  • Lei Wang
  • Daniel Zeng

News diffusion prediction aims to predict a sequence of news sites which will quote a particular piece of news. Most of previous propagation models make efforts to estimate propagation probabilities along observed links and ignore the characteristics of news diffusion processes, and they fail to capture the implicit relationships between news sites. In this paper, we propose an algorithm to model the news diffusion processes in a continuous space and take the attributes of news into account. Experiments performed on a real-world news dataset show that our model can take advantage of news’ attributes and predict news diffusion accurately.

YNICL Journal 2018 Journal Article

Development and validation of a novel dementia of Alzheimer's type (DAT) score based on metabolism FDG-PET imaging

  • Karteek Popuri
  • Rakesh Balachandar
  • Kathryn Alpert
  • Donghuan Lu
  • Mahadev Bhalla
  • Ian R. Mackenzie
  • Robin Ging-Yuek Hsiung
  • Lei Wang

Fluorodeoxyglucose positron emission tomography (FDG-PET) imaging based 3D topographic brain glucose metabolism patterns from normal controls (NC) and individuals with dementia of Alzheimer's type (DAT) are used to train a novel multi-scale ensemble classification model. This ensemble model outputs a FDG-PET DAT score (FPDS) between 0 and 1 denoting the probability of a subject to be clinically diagnosed with DAT based on their metabolism profile. A novel 7 group image stratification scheme is devised that groups images not only based on their associated clinical diagnosis but also on past and future trajectories of the clinical diagnoses, yielding a more continuous representation of the different stages of DAT spectrum that mimics a real-world clinical setting. The potential for using FPDS as a DAT biomarker was validated on a large number of FDG-PET images (N=2984) obtained from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database taken across the proposed stratification, and a good classification AUC (area under the curve) of 0. 78 was achieved in distinguishing between images belonging to subjects on a DAT trajectory and those images taken from subjects not progressing to a DAT diagnosis. Further, the FPDS biomarker achieved state-of-the-art performance on the mild cognitive impairment (MCI) to DAT conversion prediction task with an AUC of 0. 81, 0. 80, 0. 77 for the 2, 3, 5 years to conversion windows respectively.

IJCAI Conference 2018 Conference Paper

Evaluating Brush Movements for Chinese Calligraphy: A Computer Vision Based Approach

  • Pengfei Xu
  • Lei Wang
  • Ziyu Guan
  • Xia Zheng
  • Xiaojiang Chen
  • Zhanyong Tang
  • Dingyi Fang
  • Xiaoqing Gong

Chinese calligraphy is a popular, highly esteemed art form in the Chinese cultural sphere and worldwide. Ink brushes are the traditional writing tool for Chinese calligraphy and the subtle nuances of brush movements have a great impact on the aesthetics of the written characters. However, mastering the brush movement is a challenging task for many calligraphy learners as it requires many years’ practice and expert supervision. This paper presents a novel approach to help Chinese calligraphy learners to quantify the quality of brush movements without expert involvement. Our approach extracts the brush trajectories from a video stream; it then compares them with example templates of reputed calligraphers to produce a score for the writing quality. We achieve this by first developing a novel neural network to extract the spatial and temporal movement features from the video stream. We then employ methods developed in the computer vision and signal processing domains to track the brush movement trajectory and calculate the score. We conducted extensive experiments and user studies to evaluate our approach. Experimental results show that our approach is highly accurate in identifying brush movements, yielding an average accuracy of 90%, and the generated score is within 3% of errors when compared to the one given by human experts.

YNICL Journal 2018 Journal Article

Hippocampal functional connectivity is related to self-reported cognitive concerns in breast cancer patients undergoing adjuvant therapy

  • Alexandra C. Apple
  • Matthew P. Schroeder
  • Anthony J. Ryals
  • Lynne I. Wagner
  • David Cella
  • Pei-An Shih
  • James Reilly
  • Frank J. Penedo

Nearly three out of four survivors experience Cancer-Related Cognitive Impairment (CRCI) for months or years following treatment. Both clinical and animal studies point to the hippocampus as a likely brain region affected in CRCI, however no previous study has investigated the functional connectivity of the hippocampus in CRCI. We compared hippocampal connectivity in cancer survivors and healthy controls and tested the relationship between functional connectivity differences and measures of objective and subjective cognition. Exploratory analysis of inflammatory markers was conducted in a small subset of participants as well. FMRI data were acquired during a memory task from 16 breast cancer survivors and 17 controls. The NIH Toolbox was used to assess cognitive performance and Neuro-QoL was used to measure self-reported cognitive concerns. Whole-brain group-level comparisons identified clusters with different connectivity to the hippocampus in survivors versus controls during task. Average connectivity was extracted from clusters of significant difference between the groups and correlated with cognitive performance and subjective report. Survivors performed worse on a test of episodic memory and reported greater cognitive concern than controls. Exploratory analysis found higher IL6 in cancer survivors compared to controls. Cancer survivors demonstrated higher connectivity of hippocampus with left cuneus, left lingual, left precuneus, and right middle prefrontal gyrus compared with controls. In survivors, higher task-related hippocampal-cortical connectivity was related to worse subjective measures of cognitive concern. Of the four significant clusters, higher connectivity of the precuneus with hippocampus was significantly associated with worse cognitive concern in survivors. The observed greater hippocampal-cortical connectivity in survivors compared to controls is the first reported fMRI biomarker of subjective concern, and may represent a compensatory response to cancer and its treatments. This compensation could explain, in part, the subjective feelings of cognitive impairment that were reported by survivors.

JBHI Journal 2018 Journal Article

Left Atrial Appendage Segmentation Using Fully Convolutional Neural Networks and Modified Three-Dimensional Conditional Random Fields

  • Cheng Jin
  • Jianjiang Feng
  • Lei Wang
  • Heng Yu
  • Jiang Liu
  • Jiwen Lu
  • Jie Zhou

Thrombosis has become a global disease threatening human health. The left atrial appendage (LAA) is a major source of thrombosis in patients with atrial fibrillation (AF). Positive correlation exists between LAA volume and AF risk. LAA morphology has been suggested to influence thromboembolic risk in AF patients and to help predict thromboembolic events in low-risk patient groups. Automatic segmentation of LAA can greatly help physicians diagnose AF. In consideration of the large anatomical variations of the LAA, we proposed a robust method for automatic LAA segmentation on computed tomographic angiography (CTA) data using fully convolutional neural networks with three-dimensional (3–D) conditional random fields (CRFs). After manual localization of ROI of LAA, we adopted the FCN in natural image segmentation and transferred their learned models by fine-tuning the networks to segment each 2–D LAA slice. Subsequently, we used a modified dense 3–D CRF that accounts for the 3–D spatial information and larger contextual information to refine the segmentations of all slices. Our method was evaluated on 150 sets of CTA data using five-fold cross validation. Compared with manual annotation, we obtained a mean dice overlap of $\text{94. 76}\%$ and a mean volume overlap of $\text{91. 10}\%$ with a computation time of less than 40 s per volume. Experimental results demonstrated the robustness of our method in dealing with large anatomical variations and computational efficiency for adoption in a daily clinical routine.)

AAAI Conference 2018 Conference Paper

MathDQN: Solving Arithmetic Word Problems via Deep Reinforcement Learning

  • Lei Wang
  • Dongxiang Zhang
  • Lianli Gao
  • Jingkuan Song
  • Long Guo
  • Heng Tao Shen

Designing an automatic solver for math word problems has been considered as a crucial step towards general AI, with the ability of natural language understanding and logical inference. The state-of-the-art performance was achieved by enumerating all the possible expressions from the quantities in the text and customizing a scoring function to identify the one with the maximum probability. However, it incurs exponential search space with the number of quantities and beam search has to be applied to trade accuracy for efficiency. In this paper, we make the first attempt of applying deep reinforcement learning to solve arithmetic word problems. The motivation is that deep Q-network has witnessed success in solving various problems with big search space and achieves promising performance in terms of both accuracy and running time. To fit the math problem scenario, we propose our MathDQN that is customized from the general deep reinforcement learning framework. Technically, we design the states, actions, reward function, together with a feed-forward neural network as the deep Q-network. Extensive experimental results validate our superiority over state-ofthe-art methods. Our MathDQN yields remarkable improvement on most of datasets and boosts the average precision among all the benchmark datasets by 15%.

YNICL Journal 2018 Journal Article

Systematic comparison of different techniques to measure hippocampal subfield volumes in ADNI2

  • Susanne G. Mueller
  • Paul A. Yushkevich
  • Sandhitsu Das
  • Lei Wang
  • Koen Van Leemput
  • Juan Eugenio Iglesias
  • Kate Alpert
  • Adam Mezher

Objective: Subfield-specific measurements provide superior information in the early stages of neurodegenerative diseases compared to global hippocampal measurements. The overall goal was to systematically compare the performance of five representative manual and automated T1 and T2 based subfield labeling techniques in a sub-set of the ADNI2 population. Methods: The high resolution T2 weighted hippocampal images (T2-HighRes) and the corresponding T1 images from 106 ADNI2 subjects (41 controls, 57 MCI, 8 AD) were processed as follows. A. T1-based: 1. Freesurfer + Large-Diffeomorphic-Metric-Mapping in combination with shape analysis. 2. FreeSurfer 5.1 subfields using in-vivo atlas. B. T2-HighRes: 1. Model-based subfield segmentation using ex-vivo atlas (FreeSurfer 6.0). 2. T2-based automated multi-atlas segmentation combined with similarity-weighted voting (ASHS). 3. Manual subfield parcellation. Multiple regression analyses were used to calculate effect sizes (ES) for group, amyloid positivity in controls, and associations with cognitive/memory performance for each approach. Results: Subfield volumetry was better than whole hippocampal volumetry for the detection of the mild atrophy differences between controls and MCI (ES: 0.27 vs 0.11). T2-HighRes approaches outperformed T1 approaches for the detection of early stage atrophy (ES: 0.27 vs.0.10), amyloid positivity (ES: 0.11 vs 0.04), and cognitive associations (ES: 0.22 vs 0.19). Conclusions: T2-HighRes subfield approaches outperformed whole hippocampus and T1 subfield approaches. None of the different T2-HghRes methods tested had a clear advantage over the other methods. Each has strengths and weaknesses that need to be taken into account when deciding which one to use to get the best results from subfield volumetry.

TCS Journal 2017 Journal Article

A space efficient algorithm for the longest common subsequence in k-length substrings

  • Daxin Zhu
  • Lei Wang
  • Tinran Wang
  • Xiaodong Wang

Two space efficient algorithms to solve the L C S k problem and L C S ≥ k problem are presented in this paper. The algorithms improve the time and space complexities of the algorithms of Benson et al. [4]. The space cost of the first algorithm to solve the L C S k problem is reduced from O ( n 2 ) to O ( k n ), if the size of the two input sequences are both n. The time and space costs of the second algorithm to solve the L C S ≥ k problem are both improved. The time cost is reduced from O ( k n 2 ) to O ( n 2 ), and the space cost is reduced from O ( n 2 ) to O ( k n ). In the case of k = O ( 1 ), the two algorithms are both linear space algorithms.

JBHI Journal 2017 Journal Article

HEp-2 Cell Image Classification With Deep Convolutional Neural Networks

  • Zhimin Gao
  • Lei Wang
  • Luping Zhou
  • Jianjia Zhang

Efficient Human Epithelial-2 cell image classification can facilitate the diagnosis of many autoimmune diseases. This paper proposes an automatic framework for this classification task, by utilizing the deep convolutional neural networks (CNNs) which have recently attracted intensive attention in visual recognition. In addition to describing the proposed classification framework, this paper elaborates several interesting observations and findings obtained by our investigation. They include the important factors that impact network design and training, the role of rotation-based data augmentation for cell images, the effectiveness of cell image masks for classification, and the adaptability of the CNN-based classification system across different datasets. Extensive experimental study is conducted to verify the above findings and compares the proposed framework with the well-established image classification models in the literature. The results on benchmark datasets demonstrate that 1) the proposed framework can effectively outperform existing models by properly applying data augmentation, 2) our CNN-based framework has excellent adaptability across different datasets, which is highly desirable for cell image classification under varying laboratory settings. Our system is ranked high in the cell image classification competition hosted by ICPR 2014.

AAAI Conference 2017 Conference Paper

Multiple Kernel k-Means with Incomplete Kernels

  • Xinwang Liu
  • Miaomiao Li
  • Lei Wang
  • Yong Dou
  • Jianping Yin
  • En Zhu

Multiple kernel clustering (MKC) algorithms optimally combine a group of pre-specified base kernels to improve clustering performance. However, existing MKC algorithms cannot efficiently address the situation where some rows and columns of base kernels are absent. This paper proposes a simple while effective algorithm to address this issue. Different from existing approaches where incomplete kernels are firstly imputed and a standard MKC algorithm is applied to the imputed kernels, our algorithm integrates imputation and clustering into a unified learning procedure. Specifically, we perform multiple kernel clustering directly with the presence of incomplete kernels, which are treated as auxiliary variables to be jointly optimized. Our algorithm does not require that there be at least one complete base kernel over all the samples. Also, it adaptively imputes incomplete kernels and combines them to best serve clustering. A three-step iterative algorithm with proved convergence is designed to solve the resultant optimization problem. Extensive experiments are conducted on four benchmark data sets to compare the proposed algorithm with existing imputation-based methods. Our algorithm consistently achieves superior performance and the improvement becomes more significant with increasing missing ratio, verifying the effectiveness and advantages of the proposed joint imputation and clustering.

YNICL Journal 2017 Journal Article

Subtle hippocampal deformities in breast cancer survivors with reduced episodic memory and self-reported cognitive concerns

  • Alexandra C. Apple
  • Anthony J. Ryals
  • Kathryn I. Alpert
  • Lynne I. Wagner
  • Pei-An Shih
  • Mehmet Dokucu
  • David Cella
  • Frank J. Penedo

Cancer survivors have lingering cognitive problems, however the anatomical basis for these problems has yet to be fully elucidated. Clinical studies as well as animal models of chemotherapy have pinpointed cell and volume loss to the hippocampus, however, few studies have performed shape analysis of the hippocampus on cancer survivors. This study used high-dimensional deformation mapping analysis to test whether localized hippocampal deformation differs in breast cancer survivors who received adjuvant chemotherapy coupled with hormone blockade therapy, and if deformation was related to subjective self-reported concerns and cognitive performance. 3 T MRI images were acquired from 16 pre-menopausal breast cancer survivors and 18 healthy controls without a history of cancer. Breast cancer survivors had undergone chemotherapy within the eighteen months prior to the study, and were receiving estrogen-blockade therapy at the time of the study. Automated high-dimensional deformation mapping was used to compare localized hippocampal deformation differences between groups. Self-reported subjective concerns were assessed using Neuro-QOL Cognitive Function assessment, whereas cognitive performance was evaluated using the NIH Toolbox Cognition Battery. Relative to healthy controls, cancer survivors showed significantly more inward hippocampal deformation, worse self-reported cognitive functioning, and inferior episodic memory test score. This study is the first of its kind to examine the relationship between hippocampal deformity and cognitive impairment in cancer survivors.

YNIMG Journal 2017 Journal Article

The quantification of blood-brain barrier disruption using dynamic contrast-enhanced magnetic resonance imaging in aging rhesus monkeys with spontaneous type 2 diabetes mellitus

  • Ziqian Xu
  • Wen Zeng
  • Jiayu Sun
  • Wei Chen
  • Ruzhi Zhang
  • Zunyuan Yang
  • Zunwei Yao
  • Lei Wang

Microvascular lesions of the body are one of the most serious complications that can affect patients with type 2 diabetes mellitus. The blood-brain barrier (BBB) is a highly selective permeable barrier around the microvessels of the brain. This study investigated BBB disruption in diabetic rhesus monkeys using dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI). Multi-slice DCE-MRI was used to quantify BBB permeability. Five diabetic monkeys and six control monkeys underwent magnetic resonance brain imaging in 3 Tesla MRI system. Regions of the frontal cortex, the temporal cortex, the basal ganglia, the thalamus, and the hippocampus in the two groups were selected as regions of interest to calculate the value of the transport coefficient Ktrans using the extended Tofts model. Permeability in the diabetic monkeys was significantly increased as compared with permeability in the normal control monkeys. Histopathologically, zonula occludens protein-1 decreased, immunoglobulin G leaked out of the blood, and nuclear factor E2–related factor translocated from the cytoplasm to the nuclei. It is likely that diabetes contributed to the increased BBB permeability.

YNIMG Journal 2016 Journal Article

Activity-induced manganese-dependent MRI (AIM-MRI) and functional MRI in awake rabbits during somatosensory stimulation

  • Matthew P. Schroeder
  • Craig Weiss
  • Daniel Procissi
  • Lei Wang
  • John F. Disterhoft

Activity-induced manganese-dependent MRI (AIM-MRI) is a powerful tool to track system-wide neural activity using high resolution, quantitative T1-weighted MRI in animal models and has significant advantages for investigating neural activity over other modalities including BOLD fMRI. With AIM-MRI, Mn2+ ions enter neurons via voltage-gated calcium channels preferentially active during the time of experimental exposure. A broad range of AIM-MRI studies using different species studying different phenomena have been performed, but few of these studies provide a systematic evaluation of the factors influencing the detection of Mn2+ such as dosage and the temporal characteristics of Mn2+ uptake. We identified an optimal dose of Mn2+ (25mg/kg, s. c.) in order to characterize the time-course of Mn2+ accumulation in active neural regions in the rabbit. T1-weighted MRI and functional MRI were collected 0–3, 6–9, and 24–27h post-Mn2+ injection while the vibrissae on the right side were vibrated. Significant BOLD activation in the left somatosensory (SS) cortex and left ventral posteromedial (VPM) thalamic nucleus was detected during whisker vibration. T1-weighted signal intensities were extracted from these regions, their corresponding contralateral regions and the visual cortex (to serve as controls). A significant elevation in T1-weighted signal intensity in the left SS cortex (relative to right) was evident 6–9 and 24–27h post-Mn2+ injection while the left VPM thalamus showed a significant enhancement (relative to the right) only during the 24–27h session. Visual cortex showed no hemispheric difference at any timepoint. Our results suggest that studies employing AIM-MRI would benefit by conducting experimental manipulations 6–24h after subcutaneous MnCl2 injections to optimize the concentration of contrast agent in the regions active during the exposure.

YNIMG Journal 2016 Journal Article

Intrinsic connectivity of neural networks in the awake rabbit

  • Matthew P. Schroeder
  • Craig Weiss
  • Daniel Procissi
  • John F. Disterhoft
  • Lei Wang

The way in which the brain is functionally connected into different networks has emerged as an important research topic in order to understand normal neural processing and signaling. Since some experimental manipulations are difficult or unethical to perform in humans, animal models are better suited to investigate this topic. Rabbits are a species that can undergo MRI scanning in an awake and conscious state with minimal preparation and habituation. In this study, we characterized the intrinsic functional networks of the resting New Zealand White rabbit brain using BOLD fMRI data. Group independent component analysis revealed seven networks similar to those previously found in humans, non-human primates and/or rodents including the hippocampus, default mode, cerebellum, thalamus, and visual, somatosensory, and parietal cortices. For the first time, the intrinsic functional networks of the resting rabbit brain have been elucidated demonstrating the rabbit's applicability as a translational animal model. Without the confounding effects of anesthetics or sedatives, future experiments may employ rabbits to understand changes in neural connectivity and brain functioning as a result of experimental manipulation (e. g. , temporary or permanent network disruption, learning-related changes, and drug administration).

IJCAI Conference 2016 Conference Paper

Multiple Kernel Clustering with Local Kernel Alignment Maximization

  • Miaomiao Li
  • Xinwang Liu
  • Lei Wang
  • Yong Dou
  • Jianping Yin
  • En Zhu

Kernel alignment has recently been employed for multiple kernel clustering (MKC). However, we find that most of existing works implement this alignment in a global manner, which: i) indiscriminately forces all sample pairs to be equally aligned with the same ideal similarity; and ii) is inconsistent with a well-established concept that the similarity evaluated for two farther samples in a high dimensional space is less reliable. To address these issues, this paper proposes a novel MKC algorithm with a "local" kernel alignment, which only requires that the similarity of a sample to its k-nearest neighbours be aligned with the ideal similarity matrix. Such an alignment helps the clustering algorithm to focus on closer sample pairs that shall stay together and avoids involving unreliable similarity evaluation for farther sample pairs. We derive a new optimization problem to implement this idea, and design a two-step algorithm to efficiently solve it. As experimentally demonstrated on six challenging multiple kernel learning benchmark data sets, our algorithm significantly outperforms the state-of-the-art comparable methods in the recent literature, verifying the effectiveness and superiority of maximizing local kernel alignment.

AAAI Conference 2016 Conference Paper

Multiple Kernel k -Means Clustering with Matrix-Induced Regularization

  • Xinwang Liu
  • Yong Dou
  • Jianping Yin
  • Lei Wang
  • En Zhu

Multiple kernel k-means (MKKM) clustering aims to optimally combine a group of pre-specified kernels to improve clustering performance. However, we observe that existing MKKM algorithms do not sufficiently consider the correlation among these kernels. This could result in selecting mutually redundant kernels and affect the diversity of information sources utilized for clustering, which finally hurts the clustering performance. To address this issue, this paper proposes an MKKM clustering with a novel, effective matrix-induced regularization to reduce such redundancy and enhance the diversity of the selected kernels. We theoretically justify this matrix-induced regularization by revealing its connection with the commonly used kernel alignment criterion. Furthermore, this justification shows that maximizing the kernel alignment for clustering can be viewed as a special case of our approach and indicates the extendability of the proposed matrix-induced regularization for designing better clustering algorithms. As experimentally demonstrated on five challenging MKL benchmark data sets, our algorithm significantly improves existing MKKM and consistently outperforms the state-of-the-art ones in the literature, verifying the effectiveness and advantages of incorporating the proposed matrix-induced regularization.

YNIMG Journal 2016 Journal Article

Northwestern University schizophrenia data sharing for SchizConnect: A longitudinal dataset for large-scale integration

  • Alex Kogan
  • Kathryn Alpert
  • Jose Luis Ambite
  • Daniel S. Marcus
  • Lei Wang

In this paper, we describe an instance of the Northwestern University Schizophrenia Data and Software Tool (NUSDAST), a schizophrenia-related dataset hosted at XNAT Central, and the SchizConnect data portal used for accessing and sharing the dataset. NUSDAST was built and extended upon existing, standard schemas available for data sharing on XNAT Central (http: //central. xnat. org/). With the creation of SchizConnect, we were able to link NUSDAST to other neuroimaging data sources and create a powerful, federated neuroimaging resource.

YNIMG Journal 2016 Journal Article

SchizConnect: Mediating neuroimaging databases on schizophrenia and related disorders for large-scale integration

  • Lei Wang
  • Kathryn I. Alpert
  • Vince D. Calhoun
  • Derin J. Cobia
  • David B. Keator
  • Margaret D. King
  • Alexandr Kogan
  • Drew Landis

SchizConnect (www. schizconnect. org) is built to address the issues of multiple data repositories in schizophrenia neuroimaging studies. It includes a level of mediation—translating across data sources—so that the user can place one query, e. g. for diffusion images from male individuals with schizophrenia, and find out from across participating data sources how many datasets there are, as well as downloading the imaging and related data. The current version handles the Data Usage Agreements across different studies, as well as interpreting database-specific terminologies into a common framework. New data repositories can also be mediated to bring immediate access to existing datasets. Compared with centralized, upload data sharing models, SchizConnect is a unique, virtual database with a focus on schizophrenia and related disorders that can mediate live data as information is being updated at each data source. It is our hope that SchizConnect can facilitate testing new hypotheses through aggregated datasets, promoting discovery related to the mechanisms underlying schizophrenic dysfunction.

YNICL Journal 2016 Journal Article

Subcortical neuromorphometry in schizophrenia spectrum and bipolar disorders

  • Daniel Mamah
  • Kathryn I. Alpert
  • Deanna M. Barch
  • John G. Csernansky
  • Lei Wang

BACKGROUND: Disorders within the schizophrenia spectrum genetically overlap with bipolar disorder, yet questions remain about shared biological phenotypes. Investigation of brain structure in disease has been enhanced by developments in shape analysis methods that can identify subtle regional surface deformations. Our study aimed to identify brain structure surface deformations that were common across related psychiatric disorders, and characterize differences. METHODS: Using the automated FreeSurfer-initiated Large Deformation Diffeomorphic Metric Mapping, we examined volumes and shapes of seven brain structures: hippocampus, amygdala, caudate, nucleus accumbens, putamen, globus pallidus and thalamus. We compared findings in controls (CON; n = 40), and those with schizophrenia (SCZ; n = 52), schizotypal personality disorder (STP; n = 12), psychotic bipolar disorder (P-BP; n = 49) and nonpsychotic bipolar disorder (N-BP; n = 24), aged 15-35. Relationships between morphometric measures and positive, disorganized and negative symptoms were also investigated. RESULTS: Inward deformation was present in the posterior thalamus in SCZ, P-BP and N-BP; and in the subiculum of the hippocampus in SCZ and STP. Most brain structures however showed unique shape deformations across groups. Correcting for intracranial size resulted in volumetric group differences for caudate (p < 0.001), putamen (p < 0.01) and globus pallidus (p < 0.001). Shape analysis showed dispersed patterns of expansion on the basal ganglia in SCZ. Significant clinical relationships with hippocampal, amygdalar and thalamic volumes were observed. CONCLUSIONS: Few similarities in surface deformation patterns were seen across groups, which may reflect differing neuropathologies. Posterior thalamic contraction in SCZ and BP suggest common genetic or environmental antecedents. Surface deformities in SCZ basal ganglia may have been due to antipsychotic drug effects.

YNIMG Journal 2016 Journal Article

The Northwestern University Neuroimaging Data Archive (NUNDA)

  • Kathryn Alpert
  • Alexandr Kogan
  • Todd Parrish
  • Daniel Marcus
  • Lei Wang

The Northwestern University Neuroimaging Data Archive (NUNDA), an XNAT-powered data archiving system, aims to facilitate secure data storage; centralized data management; automated, standardized data processing; and simple, intuitive data sharing. NUNDA is a federated data archive, wherein individual project owners regulate access to their data. NUNDA supports multiple methods of data import, enabling data collection in a central repository. Data in NUNDA are available by project to any authorized user, allowing coordinated data management and review across sites. With NUNDA pipelines, users capitalize on existing procedures or standardize custom routines for consistent, automated data processing. NUNDA can be integrated with other research databases to simplify data exploration and discovery. And data on NUNDA can be confidently shared for secure collaboration.

AAAI Conference 2015 Conference Paper

Absent Multiple Kernel Learning

  • Xinwang Liu
  • Lei Wang
  • Jianping Yin
  • Yong Dou
  • Jian Zhang

Multiple kernel learning (MKL) optimally combines the multiple channels of each sample to improve classification performance. However, existing MKL algorithms cannot effectively handle the situation where some channels are missing, which is common in practical applications. This paper proposes an absent MKL (AMKL) algorithm to address this issue. Different from existing approaches where missing channels are firstly imputed and then a standard MKL algorithm is deployed on the imputed data, our algorithm directly classifies each sample with its observed channels. In specific, we define a margin for each sample in its own relevant space, which corresponds to the observed channels of that sample. The proposed AMKL algorithm then maximizes the minimum of all sample-based margins, and this leads to a difficult optimization problem. We show that this problem can be reformulated as a convex one by applying the representer theorem. This makes it readily be solved via existing convex optimization packages. Extensive experiments are conducted on five MKL benchmark data sets to compare the proposed algorithm with existing imputation-based methods. As observed, our algorithm achieves superior performance and the improvement is more significant with the increasing missing ratio.

NeurIPS Conference 2015 Conference Paper

Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question

  • Haoyuan Gao
  • Junhua Mao
  • Jie Zhou
  • Zhiheng Huang
  • Lei Wang
  • Wei Xu

In this paper, we present the mQA model, which is able to answer questions about the content of an image. The answer can be a sentence, a phrase or a single word. Our model contains four components: a Long Short-Term Memory (LSTM) to extract the question representation, a Convolutional Neural Network (CNN) to extract the visual representation, an LSTM for storing the linguistic context in an answer, and a fusing component to combine the information from the first three components and generate the answer. We construct a Freestyle Multilingual Image Question Answering (FM-IQA) dataset to train and evaluate our mQA model. It contains over 150, 000 images and 310, 000 freestyle Chinese question-answer pairs and their English translations. The quality of the generated answers of our mQA model on this dataset is evaluated by human judges through a Turing Test. Specifically, we mix the answers provided by humans and our model. The human judges need to distinguish our model from the human. They will also provide a score (i. e. 0, 1, 2, the larger the better) indicating the quality of the answer. We propose strategies to monitor the quality of this evaluation process. The experiments show that in 64. 7% of cases, the human judges cannot distinguish our model from humans. The average score is 1. 454 (1. 918 for human). The details of this work, including the FM-IQA dataset, can be found on the project page: \url{http: //idl. baidu. com/FM-IQA. html}.

YNIMG Journal 2015 Journal Article

Quantitative comparison of 21 protocols for labeling hippocampal subfields and parahippocampal subregions in in vivo MRI: Towards a harmonized segmentation protocol

  • Paul A. Yushkevich
  • Robert S.C. Amaral
  • Jean C. Augustinack
  • Andrew R. Bender
  • Jeffrey D. Bernstein
  • Marina Boccardi
  • Martina Bocchetta
  • Alison C. Burggren

Objective An increasing number of human in vivo magnetic resonance imaging (MRI) studies have focused on examining the structure and function of the subfields of the hippocampal formation (the dentate gyrus, CA fields 1−3, and the subiculum) and subregions of the parahippocampal gyrus (entorhinal, perirhinal, and parahippocampal cortices). The ability to interpret the results of such studies and to relate them to each other would be improved if a common standard existed for labeling hippocampal subfields and parahippocampal subregions. Currently, research groups label different subsets of structures and use different rules, landmarks, and cues to define their anatomical extents. This paper characterizes, both qualitatively and quantitatively, the variability in the existing manual segmentation protocols for labeling hippocampal and parahippocampal substructures in MRI, with the goal of guiding subsequent work on developing a harmonized substructure segmentation protocol. Method MRI scans of a single healthy adult human subject were acquired both at 3T and 7T. Representatives from 21 research groups applied their respective manual segmentation protocols to the MRI modalities of their choice. The resulting set of 21 segmentations was analyzed in a common anatomical space to quantify similarity and identify areas of agreement. Results The differences between the 21 protocols include the region within which segmentation is performed, the set of anatomical labels used, and the extents of specific anatomical labels. The greatest overall disagreement among the protocols is at the CA1/subiculum boundary, and disagreement across all structures is greatest in the anterior portion of the hippocampal formation relative to the body and tail. Conclusions The combined examination of the 21 protocols in the same dataset suggests possible strategies towards developing a harmonized subfield segmentation protocol and facilitates comparison between published studies.

ICRA Conference 2014 Conference Paper

3D path planning of a laser manipulation robotic system for tooth preparing

  • Lei Ma
  • Dangxiao Wang
  • Yuru Zhang
  • Lei Wang
  • Pei-jun Lv
  • Yuchun Sun

In this paper, we proposed a 3D path planning method for a miniature robotic system, which can manipulate a beam of ultra-short pulse laser to cut a decayed tooth to formulate an expected 3D shape. Using high resolution STereo Lithography (STL) models of the original decayed tooth and the target preparing shape as the input, our method consists of a fast slicing algorithm and an optimized path generating algorithm, which realized a high efficient layer-by-layer cutting for laser ablation. Theoretical analysis on the geometric distortion and surface roughness was carried out to model the influence of the path planning algorithms on the accuracy of the prepared tooth. Experimental results on a real tooth indicate that the path planning method can maintain the accuracy for the laser ablation process.

EAAI Journal 2014 Journal Article

A review of opposition-based learning from 2005 to 2012

  • Qingzheng Xu
  • Lei Wang
  • Na Wang
  • Xinhong Hei
  • Li Zhao

Diverse forms of opposition are already existent virtually everywhere around us, and utilizing opposite numbers to accelerate an optimization method is a new idea. Since 2005, opposition-based learning is a fast growing research field in which a variety of new theoretical models and technical methods have been studied for dealing with complex and significant problems. As a result, an increasing number of works have thus proposed. This paper provides a survey on the state-of-the-art of research, reported in the specialized literature to date, related to this framework. This overview covers basic concepts, theoretical foundation, combinations with intelligent algorithms, and typical application fields. A number of challenges that can be undertaken to help move the field forward are discussed according to the current state of the opposition-based learning.

NeurIPS Conference 2014 Conference Paper

Encoding High Dimensional Local Features by Sparse Coding Based Fisher Vectors

  • Lingqiao Liu
  • Chunhua Shen
  • Lei Wang
  • Anton van den Hengel
  • Chao Wang

Deriving from the gradient vector of a generative model of local features, Fisher vector coding (FVC) has been identified as an effective coding method for image classification. Most, if not all, FVC implementations employ the Gaussian mixture model (GMM) to characterize the generation process of local features. This choice has shown to be sufficient for traditional low dimensional local features, e. g. , SIFT; and typically, good performance can be achieved with only a few hundred Gaussian distributions. However, the same number of Gaussians is insufficient to model the feature space spanned by higher dimensional local features, which have become popular recently. In order to improve the modeling capacity for high dimensional features, it turns out to be inefficient and computationally impractical to simply increase the number of Gaussians. In this paper, we propose a model in which each local feature is drawn from a Gaussian distribution whose mean vector is sampled from a subspace. With certain approximation, this model can be converted to a sparse coding procedure and the learning/inference problems can be readily solved by standard sparse coding methods. By calculating the gradient vector of the proposed model, we derive a new fisher vector encoding strategy, termed Sparse Coding based Fisher Vector Coding (SCFVC). Moreover, we adopt the recently developed Deep Convolutional Neural Network (CNN) descriptor as a high dimensional local feature and implement image classification with the proposed SCFVC. Our experimental evaluations demonstrate that our method not only significantly outperforms the traditional GMM based Fisher vector encoding but also achieves the state-of-the-art performance in generic object recognition, indoor scene, and fine-grained image classification problems.

YNICL Journal 2014 Journal Article

Novel ThickNet features for the discrimination of amnestic MCI subtypes

  • Pradeep Reddy Raamana
  • Wei Wen
  • Nicole A. Kochan
  • Henry Brodaty
  • Perminder S. Sachdev
  • Lei Wang
  • Mirza Faisal Beg

BACKGROUND: Amnestic mild cognitive impairment (aMCI) is considered to be a transitional stage between healthy aging and Alzheimer's disease (AD), and consists of two subtypes: single-domain aMCI (sd-aMCI) and multi-domain aMCI (md-aMCI). Individuals with md-aMCI are found to exhibit higher risk of conversion to AD. Accurate discrimination among aMCI subtypes (sd- or md-aMCI) and controls could assist in predicting future decline. METHODS: We apply our novel thickness network (ThickNet) features to discriminate md-aMCI from healthy controls (NC). ThickNet features are extracted from the properties of a graph constructed from inter-regional co-variation of cortical thickness. We fuse these ThickNet features using multiple kernel learning to form a composite classifier. We apply the proposed ThickNet classifier to discriminate between md-aMCI and NC, sd-aMCI and NC and; and also between sd-aMCI and md-aMCI, using baseline T1 MR scans from the Sydney Memory and Ageing Study. RESULTS: ThickNet classifier achieved an area under curve (AUC) of 0.74, with 70% sensitivity and 69% specificity in discriminating md-aMCI from healthy controls. The same classifier resulted in AUC = 0.67 and 0.67 for sd-aMCI/NC and sd-aMCI/md-aMCI classification experiments respectively. CONCLUSIONS: The proposed ThickNet classifier demonstrated potential for discriminating md-aMCI from controls, and in discriminating sd-aMCI from md-aMCI, using cortical features from baseline MRI scan alone. Use of the proposed novel ThickNet features demonstrates significant improvements over previous experiments using cortical thickness alone. This result may offer the possibility of early detection of Alzheimer's disease via improved discrimination of aMCI subtypes.

TCS Journal 2014 Journal Article

On the equality constraints tolerance of Constrained Optimization Problems

  • Chengyong Si
  • Jing An
  • Tian Lan
  • Thomas Ußmüller
  • Lei Wang
  • Qidi Wu

The tolerance value plays an important role when converting equality constraints into inequality constraints in solving Constrained Optimization Problems. Many researchers use a fixed or dynamic setting directly based on trial or experiments without systematic study. As a well-known constraint handling technique, Deb's feasibility-based rule is widely adopted, but it has one drawback as the ranking is not consistent with the actual ranking after introducing the tolerance value. After carefully analyzing how the tolerance value influences the ranking difference, a novel strategy named Ranking Adjustment Strategy (RAS) is proposed, which can be considered as a complement of Deb's feasibility-based rule. The experiment has verified the effectiveness of the proposed strategy. This is the first time to analyze the inner mechanism of the tolerance value for equality constraints systematically, which can give some guide for future research.

AAAI Conference 2014 Conference Paper

Sample-adaptive Multiple Kernel Learning

  • Xinwang Liu
  • Lei Wang
  • Jian Zhang
  • Jianping Yin

Existing multiple kernel learning (MKL) algorithms indiscriminately apply a same set of kernel combination weights to all samples. However, the utility of base kernels could vary across samples and a base kernel useful for one sample could become noisy for another. In this case, rigidly applying a same set of kernel combination weights could adversely affect the learning performance. To improve this situation, we propose a sample-adaptive MKL algorithm, in which base kernels are allowed to be adaptively switched on/off with respect to each sample. We achieve this goal by assigning a latent binary variable to each base kernel when it is applied to a sample. The kernel combination weights and the latent variables are jointly optimized via margin maximization principle. As demonstrated on five benchmark data sets, the proposed algorithm consistently outperforms the comparable ones in the literature.

JBHI Journal 2014 Journal Article

The Sensitive and Efficient Detection of Quadriceps Muscle Thickness Changes in Cross-Sectional Plane Using Ultrasonography: A Feasibility Investigation

  • Jizhou Li
  • Yongjin Zhou
  • Yi Lu
  • Guangquan Zhou
  • Lei Wang
  • Yong-Ping Zheng

As a direct determinant parameter to quantify muscle activity, the muscle thickness (MT) has been investigated in many aspects and for various purposes. Ultrasonography (US) is a promising modality to detect muscle morphological changes during contractions since it is portable, noninvasive, and real time. However, there are few reports on sensitive and efficient estimation of changes of MT in a cross-sectional plane. In this feasibility investigation, we proposed a coarse-to-fine method based on a compressive-tracking algorithm for estimation of MT changes during an example task of isometric knee extension using ultrasound images. The sensitivity and efficiency are evaluated with 1920 US images from quadriceps muscle (QM) in eight subjects. The detection results were compared with those obtained from both traditional manual measurement and the well known normalized cross-correlation method, and the effect of the size of tracking window on detection performance was evaluated as well. It is demonstrated that the proposed method agrees well with the manual measurement. Meanwhile, it is not only sensitive to relatively small changes of MT but also computationally efficient.

JBHI Journal 2013 Journal Article

Automatic Tracking of Aponeuroses and Estimation of Muscle Thickness in Ultrasonography: A Feasibility Study

  • Shan Ling
  • Yongjin Zhou
  • Ye Chen
  • Yu-Qian Zhao
  • Lei Wang
  • Yong-Ping Zheng

Muscle thickness measurement in ultrasonography was traditionally conducted by a trained operator, and the manual detecting process is time consuming and subjective. In this paper, we proposed an automatic tracking strategy to achieve the continuous and quantitative measurement for gastrocnemius muscle thickness in ultrasound images. The method involved three steps: tracking of seed points, contours extraction of aponeuroses, and muscle thickness estimation. In an ultrasound image sequence, we first selected two seed points in the first frame manually for the superficial and deep aponeuroses, respectively. Seed points in all following frames were then tracked by registering to their respective previous frames. Second, we adopted the local and global intensity fitting model to extract the contours of aponeuroses. At last, the muscle thickness was achieved by calculating the distance between the contours of superficial and deep aponeuroses. The performance of the algorithm was evaluated using 500 frames of ultrasound images. It was demonstrated in the experiments that the proposed methods could be used for objective tracking of aponeuroses and estimation of muscle thickness in musculoskeletal ultrasound images.

EAAI Journal 2013 Journal Article

Multi-objective optimization using teaching-learning-based optimization algorithm

  • Feng Zou
  • Lei Wang
  • Xinhong Hei
  • Debao Chen
  • Bin Wang

Two major goals in multi-objective optimization are to obtain a set of nondominated solutions as closely as possible to the true Pareto front (PF) and maintain a well-distributed solution set along the Pareto front. In this paper, we propose a teaching-learning-based optimization (TLBO) algorithm for multi-objective optimization problems (MOPs). In our algorithm, we adopt the nondominated sorting concept and the mechanism of crowding distance computation. The teacher of the learners is selected from among current nondominated solutions with the highest crowding distance values and the centroid of the nondominated solutions from current archive is selected as the Mean of the learners. The performance of proposed algorithm is investigated on a set of some benchmark problems and real life application problems and the results show that the proposed algorithm is a challenging method for multi-objective algorithms.

IROS Conference 2013 Conference Paper

Preliminary experiments of a miniature robotic system for tooth ablation using ultra-short pulsed lasers

  • Lei Wang
  • Dangxiao Wang
  • Lei Ma
  • Yuru Zhang
  • Fusong Yuan
  • Yuchun Sun
  • Pei-jun Lv

As a preliminary step to achieve a long-term goal of developing an automatic dental preparation system for clinical operations, we design and build a miniature robotic system which can manipulate a laser beam to move in three dimensional spaces to remove hard tissue from a target tooth. The dental preparation requires the robotic system to own high accuracy, high ablation speed and small size. A 2D galvanometer scanners module is integrated to meet the requirement of a high moving speed of the laser focus. A closed-loop system based on a miniature-sized voice-coil motor and a grating ruler are developed to realize the accurate control of the focus. The overall size of the developed prototype is 108mm×56mm×43mm, which is small enough to be used in close proximity to a patient's mouth. The prototype has been tested by using two different kinds of laser generators, i. e. , a nanosecond laser and a picosecond laser. The experiment results show that the robotic system can provide high moving speed of 1000mm/s with good shape accuracy. From the results, we found that nanosecond laser beam can be controlled to ablate zirconia and aluminum, but not suitable to ablate tooth because of tissue carbonization. By selecting suitable parameters of the picosecond laser generator, a target tooth could be ablated to produce a cylinder shape without carbonization. Limitations of the prototype are identified according to the experiment results.

YNIMG Journal 2012 Journal Article

Automated detection of amnestic mild cognitive impairment in community-dwelling elderly adults: A combined spatial atrophy and white matter alteration approach

  • Yue Cui
  • Wei Wen
  • Darren M. Lipnicki
  • Mirza Faisal Beg
  • Jesse S. Jin
  • Suhuai Luo
  • Wanlin Zhu
  • Nicole A. Kochan

Amnestic mild cognitive impairment (aMCI) is a syndrome widely considered to be prodromal Alzheimer's disease. Accurate diagnosis of aMCI would enable earlier treatment, and could thus help minimize the prevalence of Alzheimer's disease. The aim of the present study was to evaluate a magnetic resonance imaging-based automated classification schema for identifying aMCI. This was carried out in a sample of community-dwelling adults aged 70–90years old: 79 with a clinical diagnosis of aMCI and 204 who were cognitively normal. Our schema was novel in using measures of both spatial atrophy, derived from T1-weighted images, and white matter alterations, assessed with diffusion tensor imaging (DTI) tract-based spatial statistics (TBSS). Subcortical volumetric features were extracted using a FreeSurfer-initialized Large Deformation Diffeomorphic Metric Mapping (FS+LDDMM) segmentation approach, and fractional anisotropy (FA) values obtained for white matter regions of interest. Features were ranked by their ability to discriminate between aMCI and normal cognition, and a support vector machine (SVM) selected an optimal feature subset that was used to train SVM classifiers. As evaluated via 10-fold cross-validation, the classification performance characteristics achieved by our schema were: accuracy, 71. 09%; sensitivity, 51. 96%; specificity, 78. 40%; and area under the curve, 0. 7003. Additionally, we identified numerous socio-demographic, lifestyle, health and other factors potentially implicated in the misclassification of individuals by our schema and those previously used by others. Given its high level of performance, our classification schema could facilitate the early detection of aMCI in community-dwelling elderly adults.

YNIMG Journal 2012 Journal Article

Cognitively normal individuals with AD parents may be at risk for developing aging-related cortical thinning patterns characteristic of AD

  • Katherine Reiter
  • Kathryn I. Alpert
  • Derin J. Cobia
  • Mary J. Kwasny
  • John C. Morris
  • John C. Csernansky
  • Lei Wang

Children of Alzheimer's disease (AD) patients are at heightened risk of developing AD due to genetic influences, including the apolipoprotein E4 (ApoE4) allele. In this study, we assessed the earliest cortical changes associated with AD in 71 cognitively healthy, adult children of AD patients (AD offspring) as compared with 69 with no family history of AD (non-AD offspring). Cortical thickness measures were obtained using FreeSurfer from 1. 5T magnetic resonance (MR) scans. ApoE genotyping was obtained. Primary analyses examined family history and ApoeE4 effects on cortical thickness. Secondary analyses examined age effects within groups. All comparisons were adjusted using False Discovery Rate at a significance threshold of p <0. 05. There were no statistically significant differences between family history and ApoE4 groups. Within AD offspring, increasing age was related to reduced cortical thickness (atrophy) over large areas of the precuneus, superior frontal and superior temporal gyri, starting at around age 60. Further, these patterns existed within female and maternal AD offspring, but were absent in male and paternal AD offspring. Within non-AD offspring, negative correlations existed over small regions of the superior temporal, insula and lingual cortices. These results suggest that as AD offspring age, cortical atrophy is more prominent, particularly if the parent with AD is mother or if the AD offspring is female.

EAAI Journal 2012 Journal Article

Hoeffding bound based evolutionary algorithm for symbolic regression

  • Li Zhao
  • Lei Wang
  • Du-wu Cui

In symbolic regression area, it is difficult for evolutionary algorithms to construct a regression model when the number of sample points is very large. Much time will be spent in calculating the fitness of the individuals and in selecting the best individuals within the population. Hoeffding bound is a probability bound for sums of independent random variables. As a statistical result, it can be used to exactly decide how many samples are necessary for choosing i individuals from a population in evolutionary algorithms without calculating the fitness completely. This paper presents a Hoeffding bound based evolutionary algorithm (HEA) for regression or approximation problems when the number of the given learning samples is very large. In HEA, the original fitness function is used in every k generations to update the approximate fitness obtained by Hoeffding bound. The parameter 1−δ is the probability of correctly selecting i best individuals from population P, which can be tuned to avoid an unstable evolution process caused by a large discrepancy between the approximate model and the original fitness function. The major advantage of the proposed HEA algorithm is that it can guarantee that the solution discovered has performance matching what would be discovered with a traditional genetic programming (GP) selection operator with a determinate probability and the running time can be reduced largely. We examine the performance of the proposed algorithm with several regression problems and the results indicate that with the similar accuracy, the HEA algorithm can find the solution more efficiently than tradition EA. It is very useful for regression problems with large number of training samples.

JMLR Journal 2012 Journal Article

Positive Semidefinite Metric Learning Using Boosting-like Algorithms

  • Chunhua Shen
  • Junae Kim
  • Lei Wang
  • Anton van den Hengel

The success of many machine learning and pattern recognition methods relies heavily upon the identification of an appropriate distance metric on the input data. It is often beneficial to learn such a metric from the input training data, instead of using a default one such as the Euclidean distance. In this work, we propose a boosting-based technique, termed BOOSTMETRIC, for learning a quadratic Mahalanobis distance metric. Learning a valid Mahalanobis distance metric requires enforcing the constraint that the matrix parameter to the metric remains positive semidefinite. Semidefinite programming is often used to enforce this constraint, but does not scale well and is not easy to implement. BOOSTMETRIC is instead based on the observation that any positive semidefinite matrix can be decomposed into a linear combination of trace-one rank-one matrices. BOOSTMETRIC thus uses rank-one positive semidefinite matrices as weak learners within an efficient and scalable boosting-based learning process. The resulting methods are easy to implement, efficient, and can accommodate various types of constraints. We extend traditional boosting algorithms in that its weak learner is a positive semidefinite matrix with trace and rank being one rather than a classifier or regressor. Experiments on various data sets demonstrate that the proposed algorithms compare favorably to those state-of-the-art methods in terms of classification accuracy and running time. [abs] [ pdf ][ bib ] &copy JMLR 2012. ( edit, beta )

AAAI Conference 2010 Conference Paper

Efficient Spectral Feature Selection with Minimum Redundancy

  • Zheng Zhao
  • Lei Wang
  • Huan Liu

Spectral feature selection identifies relevant features by measuring their capability of preserving sample similarity. It provides a powerful framework for both supervised and unsupervised feature selection, and has been proven to be effective in many real-world applications. One common drawback associated with most existing spectral feature selection algorithms is that they evaluate features individually and cannot identify redundant features. Since redundant features can have significant adverse effect on learning performance, it is necessary to address this limitation for spectral feature selection. To this end, we propose a novel spectral feature selection algorithm to handle feature redundancy, adopting an embedded model. The algorithm is derived from a formulation based on a sparse multi-output regression with a L2, 1-norm constraint. We conduct theoretical analysis on the properties of its optimal solutions, paving the way for designing an efficient path-following solver. Extensive experiments show that the proposed algorithm can do well in both selecting relevant features and removing redundancy.

EAAI Journal 2010 Journal Article

Predication based immune network for multimodal function optimization

  • Qingzheng Xu
  • Lei Wang
  • Jing Si

For the problem of indeterminate direction of local search, lacking of efficient regulation mechanism between local search and global search and regenerating new antibodies randomly in the original optimization version of artificial immune network (opt-aiNet), this paper puts forward a novel predication based immune network (PiNet) to solve multimodal function optimization more efficiently, accurately and reliably. The algorithm mimics natural phenomenon in immune system such as clonal selection, affinity maturation, immune network, immune memory and immune predication. The proposed algorithm includes two main features with opt-aiNet. The information of antibodies in continuous generations is utilized to point out the direction of local search and to adjust the balance between local and global search. PiNet also employs memory cells to generate new antibodies with high affinities. Theory analysis and experiments on 10 widely used benchmark problems show that when compared with opt-aiNet method, PiNet algorithm is capable of improving search performance significantly in successful rate, convergence speed, search ability, solution quality and algorithm stability.

YNIMG Journal 2009 Journal Article

Morphometric abnormalities and hyperanxiety in genetically epileptic rats: A model of psychiatric comorbidity?

  • Viviane Bouilleret
  • R. Edward Hogan
  • Dennis Velakoulis
  • Michael R. Salzberg
  • Lei Wang
  • Gary F. Egan
  • Terence J. O'Brien
  • Nigel C. Jones

Background Imaging studies of epilepsy patients with comorbid affective disturbance demonstrate morphometric changes in limbic brain regions implicated in psychiatric disease. Genetic Absence Epilepsy Rats from Strasbourg (GAERS), specifically bred for their epilepsy phenotype, also exhibit elevated anxiety-like behaviors suggesting a common causality. Here we examined whether relevant cerebral morphological alterations exist in this rat strain using volumetric measurements and large deformation high dimensional mapping (HDM-LD), a tool recently validated to produce accurate three-dimensional surface representations of the hippocampus. Methods Volumetric MRI and the Open Field test of anxiety were performed in adult female GAERS (n =12) and Non-Epileptic Controls (NEC; n =11). The volumes of selected brain regions, including cortex, hippocampus, amygdala, thalamus, hypothalamus and lateral ventricles, were measured using Region-Of-Interest analysis from the MRI data and total volumes compared between the two strains. Results GAERS had increased amygdala (right: p =0. 003; left p <0. 001), cortices (right: p =0. 006; left p =0. 012) and ventricular volumes (p =0. 002) when compared with NEC rats. Further, HDM-LD showed GAERS to have hippocampal volume loss in two regions: the medial hippocampal surface immediately caudal to the hippocampal commissure, and the lateral hippocampal surface over the mid-portion of the septotemporal axis. GAERS exhibited increased anxiety in the Open Field compared with NEC rats: reduced distance traveled (p <0. 001) and reduced time in the centre area (p =0. 042). Conclusions Morphometric brain changes in GAERS could be relevant to their hyperanxious and epileptic phenotypes. This model may be useful in illuminating the pathogenesis of affective disorders generally, as well as modeling psychiatric comorbidities of epilepsy.

YNIMG Journal 2009 Journal Article

Neuroanatomical asymmetry patterns in individuals with schizophrenia and their non-psychotic siblings

  • Anqi Qiu
  • Lei Wang
  • Laurent Younes
  • Michael P. Harms
  • J. Tilak Ratnanather
  • Michael I. Miller
  • John G. Csernansky

Neuroanatomical endophenotypes may reveal insights into the processes by which genetic factors increase the risk of developing schizophrenia. To determine whether patterns of neuroanatomical asymmetries may be useful as schizophrenia-related endophenotypes, we compared patterns of structural asymmetries in patients with schizophrenia, healthy controls, and their respective siblings. The surfaces of the left and right amygdala, hippocampus, thalamus, caudate nucleus, putamen, globus pallidus, and nucleus accumbens were assessed in 40 pairs of healthy comparison controls (CON) and their siblings (CON-SIB) and 25 pairs of patients with schizophrenia (SCZ) and their siblings (SCZ-SIB) in magnetic resonance (MR) images using large deformation diffeomorphic metric mapping (LDDMM) and parallel transport techniques. The within-subject asymmetry deformation of each structure was first measured via LDDMM, and then translated to a global template via parallel transport for evaluation of the patterns of asymmetry both within and across siblings. Our results revealed that asymmetries observed in CON subjects occurred in the amygdala and the anterior segment of the hippocampus with more pronounced expansion deformation in the right-sided structures (R>L asymmetry) but not in the basal ganglia and thalamus. Disturbance in this pattern of asymmetries was observed in both SCZ and SCZ-SIB subjects. More specifically, exaggerations and reductions in the normative pattern of asymmetries were observed in the amygdala–hippocampus formation, basal ganglia, and thalamus. These altered patterns of asymmetries are present in subjects with schizophrenia and their siblings, and therefore may represent a schizophrenia-related endophenotype.

NeurIPS Conference 2009 Conference Paper

Positive Semidefinite Metric Learning with Boosting

  • Chunhua Shen
  • Junae Kim
  • Lei Wang
  • Anton Hengel

The learning of appropriate distance metrics is a critical problem in classification. In this work, we propose a boosting-based technique, termed BoostMetric, for learning a Mahalanobis distance metric. One of the primary difficulties in learning such a metric is to ensure that the Mahalanobis matrix remains positive semidefinite. Semidefinite programming is sometimes used to enforce this constraint, but does not scale well. BoostMetric is instead based on a key observation that any positive semidefinite matrix can be decomposed into a linear positive combination of trace-one rank-one matrices. BoostMetric thus uses rank-one positive semidefinite matrices as weak learners within an efficient and scalable boosting-based learning process. The resulting method is easy to implement, does not require tuning, and can accommodate various types of constraints. Experiments on various datasets show that the proposed algorithm compares favorably to those state-of-the-art methods in terms of classification accuracy and running time.

EAAI Journal 2008 Journal Article

AdaBoost with SVM-based component classifiers

  • Xuchun Li
  • Lei Wang
  • Eric Sung

The use of SVM (Support Vector Machine) as component classifier in AdaBoost may seem like going against the grain of the Boosting principle since SVM is not an easy classifier to train. Moreover, Wickramaratna et al. [2001. Performance degradation in boosting. In: Proceedings of the Second International Workshop on Multiple Classifier Systems, pp. 11–21] show that AdaBoost with strong component classifiers is not viable. In this paper, we shall show that AdaBoost incorporating properly designed RBFSVM (SVM with the RBF kernel) component classifiers, which we call AdaBoostSVM, can perform as well as SVM. Furthermore, the proposed AdaBoostSVM demonstrates better generalization performance than SVM on imbalanced classification problems. The key idea of AdaBoostSVM is that for the sequence of trained RBFSVM component classifiers, starting with large σ values (implying weak learning), the σ values are reduced progressively as the Boosting iteration proceeds. This effectively produces a set of RBFSVM component classifiers whose model parameters are adaptively different manifesting in better generalization as compared to AdaBoost approach with SVM component classifiers using a fixed (optimal) σ value. From benchmark data sets, we show that our AdaBoostSVM approach outperforms other AdaBoost approaches using component classifiers such as Decision Trees and Neural Networks. AdaBoostSVM can be seen as a proof of concept of the idea proposed in Valentini and Dietterich [2004. Bias-variance analysis of support vector machines for the development of SVM-based ensemble methods. Journal of Machine Learning Research 5, 725–775] that Adaboost with heterogeneous SVMs could work well. Moreover, we extend AdaBoostSVM to the Diverse AdaBoostSVM to address the reported accuracy/diversity dilemma of the original Adaboost. By designing parameter adjusting strategies, the distributions of accuracy and diversity over RBFSVM component classifiers are tuned to maintain a good balance between them and promising results have been obtained on benchmark data sets.

YNIMG Journal 2008 Journal Article

FreeSurfer-initiated fully-automated subcortical brain segmentation in MRI using Large Deformation Diffeomorphic Metric Mapping

  • Ali R. Khan
  • Lei Wang
  • Mirza Faisal Beg

Fully-automated brain segmentation methods have not been widely adopted for clinical use because of issues related to reliability, accuracy, and limitations of delineation protocol. By combining the probabilistic-based FreeSurfer (FS) method with the Large Deformation Diffeomorphic Metric Mapping (LDDMM)-based label-propagation method, we are able to increase reliability and accuracy, and allow for flexibility in template choice. Our method uses the automated FreeSurfer subcortical labeling to provide a coarse-to-fine introduction of information in the LDDMM template-based segmentation resulting in a fully-automated subcortical brain segmentation method (FS+LDDMM). One major advantage of the FS+LDDMM-based approach is that the automatically generated segmentations generated are inherently smooth, thus subsequent steps in shape analysis can directly follow without manual post-processing or loss of detail. We have evaluated our new FS+LDDMM method on several databases containing a total of 50 subjects with different pathologies, scan sequences and manual delineation protocols for labeling the basal ganglia, thalamus, and hippocampus. In healthy controls we report Dice overlap measures of 0. 81, 0. 83, 0. 74, 0. 86 and 0. 75 for the right caudate nucleus, putamen, pallidum, thalamus and hippocampus respectively. We also find statistically significant improvement of accuracy in FS+LDDMM over FreeSurfer for the caudate nucleus and putamen of Huntington's disease and Tourette's syndrome subjects, and the right hippocampus of Schizophrenia subjects.

NeurIPS Conference 2008 Conference Paper

PSDBoost: Matrix-Generation Linear Programming for Positive Semidefinite Matrices Learning

  • Chunhua Shen
  • Alan Welsh
  • Lei Wang

In this work, we consider the problem of learning a positive semidefinite matrix. The critical issue is how to preserve positive semidefiniteness during the course of learning. Our algorithm is mainly inspired by LPBoost [1] and the general greedy convex optimization framework of Zhang [2]. We demonstrate the essence of the algorithm, termed PSDBoost (positive semidefinite Boosting), by focusing on a few different applications in machine learning. The proposed PSDBoost algorithm extends traditional Boosting algorithms in that its parameter is a positive semidefinite matrix with trace being one instead of a classifier. PSDBoost is based on the observation that any trace-one positive semidefinitematrix can be decomposed into linear convex combinations of trace-one rank-one matrices, which serve as base learners of PSDBoost. Numerical experiments are presented.

YNIMG Journal 2008 Journal Article

Symmetric abnormalities in sulcal patterning in schizophrenia

  • John G. Csernansky
  • Sarah K. Gillespie
  • Donna L. Dierker
  • Alan Anticevic
  • Lei Wang
  • Deanna M. Barch
  • David C. Van Essen

To compare the morphology of the cerebral cortex and its characteristic pattern of gyri and sulci in individuals with and without schizophrenia, T1-weighted magnetic resonance scans were collected, along with clinical and cognitive information, from 33 individuals with schizophrenia and 30 healthy individuals group-matched for age, gender, race and parental socioeconomic status. Sulcal depth was measured across the entire cerebral cortex by reconstructing surfaces of cortical mid-thickness (layer 4) in each hemisphere and registering them to the human PALS cortical atlas. Group differences in sulcal depth were tested using methods for cluster size analysis and interhemispheric symmetry analysis. A significant group difference was found bilaterally in the parietal operculum, where the average sulcal depth was shallower in individuals with schizophrenia. In addition, group differences in sulcal depth showed significant bilateral symmetry across much of the occipital, parietal, and temporal cortices. In individuals with schizophrenia, sulcal depth in the left hemisphere was correlated with the severity of impaired performance on tests of working memory and executive function.

YNIMG Journal 2007 Journal Article

Combining anatomical manifold information via diffeomorphic metric mappings for studying cortical thinning of the cingulate gyrus in schizophrenia

  • Anqi Qiu
  • Laurent Younes
  • Lei Wang
  • J. Tilak Ratnanather
  • Sarah K. Gillepsie
  • Gillian Kaplan
  • John Csernansky
  • Michael I. Miller

Spatial normalization is a crucial step in assessing patterns of neuroanatomical structure and function associated with health and disease. Errors that occur during spatial normalization can influence hypothesis testing due to the dimensionalities of mapping algorithms and anatomical manifolds (landmarks, curves, surfaces, volumes) used to drive the mapping algorithms. The primary aim of this paper is to improve statistical inference using multiple anatomical manifolds and large deformation diffeomorphic metric mapping (LDDMM) algorithms. We propose that combining information generated by the various manifolds and algorithms improves the reliability of hypothesis testing. We used this unified approach to assess variation in the thickness of the cingulate gyrus in subjects with schizophrenia and healthy comparison subjects. Three different LDDMM algorithms for mapping landmarks, curves and triangulated meshes were used to transform thickness maps of the cingulate surfaces into an atlas coordinate system. We then tested for group differences by combining the information from the three types of anatomical manifolds and LDDMM mapping algorithms. The unified approach provided reliable statistical results and eliminated ambiguous results due to surface mismatches. Subjects with schizophrenia had non-uniform cortical thinning over the left and right cingulate gyri, especially in the anterior portion, as compared to healthy comparison subjects.

YNIMG Journal 2006 Journal Article

Abnormalities of hippocampal surface structure in very mild dementia of the Alzheimer type

  • Lei Wang
  • J. Philp Miller
  • Mokhtar H. Gado
  • Daniel W. McKeel
  • Marcus Rothermich
  • Michael I. Miller
  • John C. Morris
  • John G. Csernansky

To better define the pattern of hippocampal deformity early in the course of Alzheimer's disease, we compared the pattern of hippocampal surface variation in subjects with very mild dementia of the Alzheimer type (DAT) and nondemented subjects. The surface of the hippocampus was divided a priori on a neuroanatomical template into three zones approximating the locations of underlying subfields [Csernansky, J. G. , Wang, L. , Swank, J. , Miller, J. P. , Gado, M. , McKeel, D. , Miller, M. I. , Morris, J. C. , 2005. Preclinical detection of Alzheimer's disease: hippocampal shape and volume predict dementia onset in the elderly. NeuroImage 25, 783–792]; i. e. , a lateral zone (LZ) approximating the CA1 subfield, a superior zone (SZ) approximating the combined CA2, CA3, CA4 subfields and the gyrus dentatus (GD), and an inferior-medial zone (IMZ) approximating the subiculum. Large-deformation high-dimensional brain mapping (HDBM-LD) was used to generate the hippocampal surfaces of all subjects and to register the surface zones across subjects. Average variations within each zone were calculated for the subjects with very mild DAT as compared to the average of the nondemented subjects. After correcting for multiple comparisons, the very mild DAT subjects showed significant inward variation in the left and right LZ, the left and right IMZ, but not in the left and right SZ as compared to nondemented subjects. In logistic regression analyses, inward variation of the left and right LZ or IMZ by 0. 1 mm relative to the average of the nondemented subjects increased the odds of the subject being a very mild DAT subject (range—1. 18 to 1. 57) rather than being a nondemented subject. The odds ratios for the left and right SZ were not significant. These results represent a replication of our previous findings [Csernansky, J. G. , Wang, L. , Joshi, S. , Miller, J. P. , Gado, M. , Kido, D. , McKeel, D. , Morris, J. C. , Miller, M. I. , 2000. Early DAT is distinguished from aging by high-dimensional mapping of the hippocampus. Neurology 55, 1636–1643. ] and suggest that inward deformities of the hippocampal surface in proximity to the CA1 subfield and subiculum can be used to distinguish subjects with very mild DAT from nondemented subjects.

YNIMG Journal 2006 Journal Article

MRI detects white matter reorganization after neural progenitor cell treatment of stroke

  • Quan Jiang
  • Zheng Gang Zhang
  • Guang Liang Ding
  • Brian Silver
  • Li Zhang
  • He Meng
  • Mei Lu
  • Siamak Pourabdillah-Nejed-D.

We evaluated the effects of neural progenitor cell treatment of stroke on white matter reorganization using MRI. Male Wistar rats (n = 26) were subjected to 3 h of middle cerebral artery occlusion and were treated with neural progenitor cells (n = 17) or without treatment (n = 9) and were sacrificed at 5–7 weeks thereafter. MRI measurements revealed that grafted neural progenitor cells selectively migrated towards the ischemic boundary regions. White matter reorganization, confirmed histologically, was coincident with increases of fractional anisotropy (FA, P < 0. 01) after stroke in the ischemic recovery regions compared to that in the ischemic core region in both treated and control groups. Immunoreactive staining showed axonal projections emanating from neurons and extruding from the corpus callosum into the ipsilateral striatum bounding the lesion areas after stroke. Fiber tracking (FT) maps derived from diffusion tensor imaging revealed similar orientation patterns to the immunohistological results. Complementary measurements in stroke patients indicated that FT maps exhibit an overall orientation parallel to the lesion boundary. Our data demonstrate that FA and FT identify and characterize cerebral tissue undergoing white matter reorganization after stroke and treatment with neural progenitor cells.

YNIMG Journal 2005 Journal Article

Investigation of neural progenitor cell induced angiogenesis after embolic stroke in rat using MRI

  • Quan Jiang
  • Zheng Gang Zhang
  • Guang Liang Ding
  • Li Zhang
  • James R. Ewing
  • Lei Wang
  • Ruilan Zhang
  • Lian Li

Using MRI, we investigated dynamic changes of brain angiogenesis after neural progenitor cell transplantation in the living adult rat subjected to embolic stroke. Neural progenitor cells isolated from the subventricular zone (SVZ) of the adult rat were labeled by superparamagnetic particles and intracisternally transplanted into the adult rat 48 h after stroke (n = 8). Before and after the transplantation, an array of MRI parameters were measured, including high resolution 3D MRI and quantitative T 1, T 1sat (T 1 in the presence of an off-resonance irradiation of the macromolecules of brain), T 2, the inverse of the apparent forward transfer rate for magnetization transfer (k inv), cerebral blood flow (CBF), cerebral blood volume (CBV), and blood-to-brain transfer constant (K i) of Gd-DTPA. The von Willerbrand factor (vWF) immunoreactive images of coronal sections obtained at 6 weeks after cell transplantation were used to analyze vWF immunoreactive vessels. MRI measurements revealed that grafted neural progenitor cells selectively migrated towards the ischemic boundary regions. In the ischemic boundary regions, angiogenesis confirmed by an increase in vascular density and the appearance of large thin wall mother vessels was coincident with increases of CBF and CBV (CBF, P < 0. 01; CBV, P < 0. 01) at 6 weeks after treatment, and coincident with transient increases of K i with a peak at 2 to 3 weeks after cell therapy. Relative T 1, T 1sat, T 2, and k inv decreased in the ischemic boundary regions with angiogenesis compared to that in the non-angiogenic ischemic region (T 1, P < 0. 01 at 6 weeks; T 1sat, P < 0. 05 at 2 to 6 weeks; T 2, P < 0. 05 at 3 to 6 weeks; k inv P < 0. 05 at 6 weeks). Of these methods, K i appear to be the most useful MR measurements which identify and predict the location and area of angiogenesis. CBF, CBV, T 1sat, T 1, T 2, and k inv provide complementary information to characterize ischemic tissue with and without angiogenesis. Our data suggest that select MRI parameters can identify the cerebral tissue destined to undergo angiogenesis after treatment of embolic stroke with cell therapy.

ICRA Conference 2004 Conference Paper

A Learning Market based Layered Multi-robot Architecture

  • Liu Lin
  • Lei Wang
  • Zhiqiang Zheng
  • Zengqi Sun

This work presents a novel learning market based layered architecture for distributed multi-robot systems. For the market system, a reward function is defined and adapted by some kind of learning algorithm in the dynamical system. A layered architecture combining the market based system with the typical layered robot architecture is also proposed to make the robot team execute smoothly in the dynamic environment. Results illustrate the relationships between task time, task amount and robot amount. The task/robot rate is defined to deeply study their relationships.

YNIMG Journal 2004 Journal Article

Computational anatomy and neuropsychiatric disease: probabilistic assessment of variation and statistical inference of group difference, hemispheric asymmetry, and time-dependent change

  • John G. Csernansky
  • Lei Wang
  • Sarang C. Joshi
  • J. Tilak Ratnanather
  • Michael I. Miller

Three components of computational anatomy (CA) are reviewed in this paper: (i) the computation of large-deformation maps, that is, for any given coordinate system representations of two anatomies, computing the diffeomorphic transformation from one to the other; (ii) the computation of empirical probability laws of anatomical variation between anatomies; and (iii) the construction of inferences regarding neuropsychiatric disease states. CA utilizes spatial–temporal vector field information obtained from large-deformation maps to assess anatomical variabilities and facilitate the detection and quantification of abnormalities of brain structure in subjects with neuropsychiatric disorders. Neuroanatomical structures are divided into two types: subcortical structures—gray matter (GM) volumes enclosed by a single surface—and cortical mantle structures—anatomically distinct portions of the cerebral cortical mantle layered between the white matter (WM) and cerebrospinal fluid (CSF). Because of fundamental differences in the geometry of these two types of structures, image-based large-deformation high-dimensional brain mapping (HDBM-LD) and large-deformation diffeomorphic metric matching (LDDMM) were developed for the study of subcortical structures and labeled cortical mantle distance mapping (LCMDM) was developed for the study of cortical mantle structures. Studies of neuropsychiatric disorders using CA usually require the testing of hypothesized group differences with relatively small numbers of subjects per group. Approaches that increase the power for testing such hypotheses include methods to quantify the shapes of individual structures, relationships between the shapes of related structures (e. g. , asymmetry), and changes of shapes over time. Promising preliminary studies employing these approaches to studies of subjects with schizophrenia and very mild to mild Alzheimer's disease (AD) are presented.

YNIMG Journal 2004 Journal Article

In vivo magnetic resonance imaging tracks adult neural progenitor cell targeting of brain tumor

  • Zhenggang Zhang
  • Quan Jiang
  • Feng Jiang
  • Gaungliang Ding
  • Ruilan Zhang
  • Lei Wang
  • Li Zhang
  • Adam M. Robin

Using magnetic resonance imaging (MRI), we described a method for noninvasively tracking grafted neural progenitor cells and bone marrow stromal cells (MSCs) in brain tumor of the rat. Neural progenitor cells and MSCs were labeled with lipophilic dye-coated superparamagnetic particles. The labeled neural progenitor cells and MSCs were transplanted to rats via the cisterna magna and a tail vein, respectively, 1 week after 9L-gliosarcoma cell implantation. Three-dimensional (3D) gradient echo and contrast agent images revealed dynamic migration of adult neural progenitor cells and MSCs detected by loss of MRI signals towards tumor mass and infiltrated tumor cells. Prussian blue staining and fluorescent microscope analysis showed that grafted cells targeted tumor cells and areas with grafted cells corresponded to areas with loss of MRI signals. These results demonstrate that the MRI technique provides a sensitive method for in vivo assessment of grafted cells targeting tumor mass and infiltrated tumor cells and that adult neural progenitor cells and MSCs can target tumor aggregates in the brain.

YNIMG Journal 2003 Journal Article

Changes in hippocampal volume and shape across time distinguish dementia of the Alzheimer type from healthy aging☆

  • Lei Wang
  • Jeffrey S. Swank
  • Irena E. Glick
  • Mokhtar H. Gado
  • Michael I. Miller
  • John C. Morris
  • John G. Csernansky

Rates of hippocampal volume loss have been shown to distinguish subjects with dementia of the Alzheimer type (DAT) from nondemented controls (Jack et al. , 2000). In this study, we obtained magnetic resonance scans in 18 subjects with very mild DAT (CDR 0. 5) and 26 age-matched nondemented controls (CDR 0) 2 years apart. Large-deformation high-dimensional brain mapping was used to quantify and compare changes in hippocampal shape as well as volume in the two groups of subjects. Hippocampal volume loss over time was significantly greater in the CDR 0. 5 subjects (left = 8. 3%, right = 10. 2%) than in the CDR 0 subjects (left = 4. 0%, right = 5. 5%) (ANOVA, F = 7. 81, P = 0. 0078). We used singular-value decomposition and logistic regression models to quantify hippocampal shape change across time within individuals, and this shape change in the CDR 0. 5 and CDR 0 subjects was found to be significantly different (Wilks's λ, P = 0. 014). Further, at baseline, CDR 0. 5 subjects, in comparison to CDR 0 subjects, showed inward deformation over 38% of the hippocampal surface; after 2 years this difference grew to 47%. Also, within the CDR 0 subjects, shape change between baseline and follow-up was largely confined to the head of the hippocampus and subiculum, while in the CDR 0. 5 subjects, shape change involved the lateral body of the hippocampus as well as the head region and subiculum. These results suggest that different patterns of hippocampal shape change in time as well as different rates of hippocampal volume loss distinguish very mild DAT from healthy aging.

YNIMG Journal 2001 Journal Article

Statistical Analysis of Hippocampal Asymmetry in Schizophrenia

  • Lei Wang
  • Sarang C. Joshi
  • Michael I. Miller
  • John G. Csernansky

The asymmetry of brain structures has been studied in schizophrenia to better understand its underlying neurobiology. Brain regions of interest have previously been characterized by volumes, cross-sectional and surface areas, and lengths. Using high-dimensional brain mapping, we have developed a statistical method for analyzing patterns of left–right asymmetry of the human hippocampus taken from high-resolution MR scans. We introduce asymmetry measures that capture differences in the patterns of high-dimensional vector fields between the left and right hippocampus surfaces. In 15 pairs of subjects previously studied (J. G. Csernansky et al. , 1998, Proc. Natl. Acad. Sci. USA 95, 11406–11411). we define the difference in hippocampal asymmetry patterns between the groups. Volume analysis indicated a large normative asymmetry between left and right hippocampus (R > L), and shape analysis allowed us to visualize the normative asymmetry pattern of the hippocampal surfaces. We observed that the right hippocampus was wider along its lateral side in both schizophrenia and control subjects. Also, while patterns of hippocampal asymmetry were generally similar in the schizophrenia and control groups, a principal component analysis based on left–right asymmetry vector fields detected a statistically significant difference between the two groups, specifically related to the subiculum.

YNIMG Journal 1999 Journal Article

Brain Segmentation and the Generation of Cortical Surfaces

  • Mukta Joshi
  • Jing Cui
  • Keith Doolittle
  • Sarang Joshi
  • David Van Essen
  • Lei Wang
  • Michael I. Miller

This paper describes methods for white matter segmentation in brain images and the generation of cortical surfaces from the segmentations. We have developed a system that allows a user to start with a brain volume, obtained by modalities such as MRI or cryosection, and constructs a complete digital representation of the cortical surface. The methodology consists of three basic components: local parametric modeling and Bayesian segmentation; surface generation and local quadratic coordinate fitting; and surface editing. Segmentations are computed by parametrically fitting known density functions to the histogram of the image using the expectation maximization algorithm [DLR77]. The parametric fits are obtained locally rather than globally over the whole volume to overcome local variations in gray levels. To represent the boundary of the gray and white matter we use triangulated meshes generated using isosurface generation algorithms [GH95]. A complete system of local parametric quadratic charts [JWM+95] is superimposed on the triangulated graph to facilitate smoothing and geodesic curve tracking. Algorithms for surface editing include extraction of the largest closed surface. Results for several macaque brains are presented comparing automated and hand surface generation.

v2026.09.13