Arrow Research search

Author name cluster

Ying Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

80 papers
2 author rows

Possible papers

80

EAAI Journal 2026 Journal Article

An artificial intelligence-powered design strategy for offshore wind turbine monopiles

  • Bence Kato
  • Ruizhi Huang
  • Gang Wang
  • Ying Wang

Monopiles support 60 % of existing offshore wind turbines (OWTs). Their effective design and modeling remains expensive and challenging due to complex nonlinear soil-structure interaction (SSI) under varying loads. This study aims to address these issues via an artificial intelligence (AI) model, named pAIle as the combination of “pile” and “AI”, using the long short-term memory (LSTM) model, trained on over 100 experimentally validated, high-fidelity finite element simulations. A new data structuring method has been introduced for LSTM training, where long-sequence data is temporally stacked via feature enrichment. This approach speeds up training by 10-fold while enhancing prediction accuracy, improving the overall R2 from 0. 983 to 0. 995. Testing results showed that pAIle efficiently predicts pile head displacements and rotations both at small strains and in the post-failure flow state by reproducing nonlinear SSI, such as damping and cyclic accumulation of plastic strains. Feature importance analysis showed that pAIle correctly understands which physical parameters govern pile head deformations. Exceptional extrapolation performance, evident in an order of magnitude lower normalized mean squared error compared to similar AI models, underscores pAIle's generalization capacity. Comparative studies with other popular AI architectures further demonstrated the effectiveness of LSTM model. Finally, the paper illustrates how this approach can be integrated into existing engineering workflows by enabling rapid monopile size optimization and post-storm integrity assessments. The procedures can complete in less than 2 s on a personal computer, requiring only readily available soil or pile parameters, showing that it is a feasible novel design strategy for OWT monopiles.

AAAI Conference 2026 Conference Paper

Beyond Counting: Evaluating Abstract and Emotional Reasoning in Vision-Language Models

  • Yuan Zhou
  • Yan Zhang
  • Jianlong Chang
  • Xin Gu
  • Ying Wang
  • Kun Ding
  • Guangwen Yang
  • Shiming Xiang

Despite the rapid progress of Vision Language Models (VLMs), existing benchmarks still concentrate on coarse-grained object recognition or simple relational reasoning, leaving the fine-grained and higher-order reasoning abilities of these systems largely unexamined. To bridge this critical evaluation gap, we introduce EmojiGrid, a novel diagnostic benchmark specifically designed to probe these fine-grained and higher-order skills. Leveraging the universal and semantically rich nature of emojis, we synthesize a grid‑based visual dataset paired with 29,000+ QA pairs. Each pair is explicitly anchored in a three-level cognitive taxonomy comprising (i) Perception and Information Extraction, (ii) Relational and Structural Reasoning, and (iii) Abstraction and Advanced Cognition. These dimensions further decompose into nine categories covering a broad range of cognitive skills, including counting, spatial relations, compositional logic, semantic sentiment, and related higher-order reasoning tasks. Our extensive evaluation of 25 state-of-the-art open-source and proprietary VLMs reveals a significant performance gap between foundational perceptual tasks and higher-level cognitive abilities, particularly in abstraction and advanced emotional reasoning. Notably, all models struggle with compositional logic, spatial consistency, and especially emotional and semantic understanding. EmojiGrid provides a quantifiable, fine-grained benchmark to diagnose VLM limitations and guides future progress toward models that can truly perceive, reason about, and interpret complex, symbol-rich visual scenes.

JBHI Journal 2026 Journal Article

BLADE: Breast Lesion Analysis with Domain Expertise for DCE-MRI Diagnosis

  • Zhitao Wei
  • Yi Dai
  • Yanting Liang
  • Chinting Wong
  • Yanfen Cui
  • Xiaobo Chen
  • Zhihe Zhao
  • Xiaodong Zheng

Dynamic Contrast-Enhanced Magnetic Reso nance Imaging (DCE-MRI) is pivotal in breast cancer diag nosis, yet radiologists face challenges in interpreting its complex data due to the lack of robust automated tools. Current lesion diagnosis systems struggle with limited datasets and insufficient integration of domain knowledge. To overcome these limitations, we propose Breast Lesion Analysis with DomainExpertise(BLADE), anoveldiagnosis framework that synergizes deep learning with clinical ex pertise. BLADE leverages a pre-trained vertical foundation model (optimized via Momentum Contrast on 2. 1 million MRI slices) as its encoder, ensuring robust feature extraction. Crucially, the system incorporates prior multi-phasic hemodynamic knowledge to emulate radiologists' diagnos tic reasoning and introduces a Breast Imaging Reporting and Data System (BI-RADS)-based constraint during training to align predictions with clinical standards. Extensive experiments demonstrate that BLADE outperforms state of-the-art methods, achieving an Area Under the Curve (AUC) of 0. 9228 and 0. 9553 on two external test datasets, respectively. Notably, BLADE significantly enhances clin ical workflow; when used as an assistive tool, BLADE improves diagnostic accuracy by 14. 31%, surpassing stan daloneperformanceofclinicians. This workbridgesthegap between AI-driven analysis and clinical practice in breast MRI interpretation. The source code is available at https://github.com/GDPHMediaLab/BLADE.

AAAI Conference 2026 Conference Paper

Centralized Group Equitability and Individual Envy-Freeness in the Allocation of Indivisible Items

  • Ying Wang
  • Jiaqian Li
  • Tianze Wei
  • Hau Chan
  • Minming Li

We study the fair allocation of indivisible items to groups of agents from the perspectives of both the agents and a centralized allocator. In our setting, the centralized allocator aims to ensure that the allocation is fair both among the groups and between individual agents. This setting applies to many real-world scenarios, such as when a school administrator allocates resources (e.g., office spaces and supplies) to staff members within departments or when a city council allocates limited housing units to families in need across different communities. To ensure fairness between agents, we consider the classical notion of envy-freeness (EF). To ensure fairness among groups, we introduce the notion of centralized group equitability (CGEQ), which captures fairness for groups from the centralized allocator’s perspective. Because an EF or CGEQ allocation does not always exist in general, we consider their natural relaxations: envy-freeness to one item (EF1) and centralized group equitability up to one item (CGEQ1). For different classes of valuation functions of the agents and the centralized allocator, we show that allocations satisfying both EF1 and CGEQ1 always exist, and we design efficient algorithms to compute such allocations. We also consider the centralized group maximin share (CGMMS) from the centralized allocator's perspective as a group-level fairness objective with EF1 for agents, and present several results.

AAAI Conference 2026 Conference Paper

Cyto-SSL: A Self-Supervised Pretraining Framework for Cytology Foundation Model

  • Yiming Zhang
  • Rui Yan
  • Xiaohua Wan
  • Yifan Zhao
  • Shuang Feng
  • Zhetao Xu
  • Ying Wang
  • Fa Zhang

Cytological images originate from exfoliated cells, collected via liquid-based slides and digitized into whole slide images (WSIs). Unlike histological WSIs that exhibit continuous and well-structured tissue, cytological WSIs are sparse in spatial distribution and unstructured in cellular relationships. Typically, the nucleus serves as the primary diagnostic feature, while surrounding cytoplasmic information plays a supportive role. These unique characteristics limit the development of effective foundation models and hinder the transferability of histology-based models for cytopathology. To address this, we propose **Cyto-SSL**, the first self-supervised pretraining framework for cytological images. It introduces **Nuclei-Centered Perturbation**, which highlights individual nuclei by perturbing non-nuclear regions. We also design an SR-Transformer module, which complements this by using sparse attention to concentrate on diagnostically relevant scattered cells, while iRPE helps model to capture local spatial relationships and avoids unnecessary attention to irrelevant global structures. Experimental results show that **Cyto-SSL** enhances performance across diverse cytological datasets and Multiple Instance Learning (MIL) methods. On a WSI-level dataset, it achieved 95.67% accuracy and outperformed ImageNet-pretrained ResNet-50 by 11.33%, demonstrating superior feature representation for cytological analysis. Additionally, **Cyto-SSL** modules are plug-and-play, easily integrated into other pretraining frameworks, yielding a 2.6% accuracy gain across different SSL methods.

JBHI Journal 2026 Journal Article

DBGT-PLA: Dual-Branch Graph–Transformer Fusion for Interpretable Protein– Ligand Affinity Prediction

  • Ying Wang
  • Jing Hu
  • Junlin Xu
  • Bo Li

Protein-ligand binding affinity prediction is critical for drug discovery, yet existing methods struggle to jointly model local atomic interactions and global contextual dependencies. To address this, we propose the Interpretable Dual-Branch Graph–Transformer framework for Protein–Ligand Affinity prediction (DBGT-PLA), a novel dual-branch architecture that integrates graph neural network (GNN) with a stability-enhanced Transformer equipped with learnable positional embeddings and a NaN-filtering mechanism that handles potential Not-a-Number (NaN) values arising from numerical instability or data preprocessing. We design a Gated Residual Learning (GRL) Fusion module that performs dimension-wise adaptive integration between local graph topology and global Transformer context. This mechanism enables multi-level feature coordination through a residual path, achieving biophysically consistent alignment between atomic-level interactions and global conformational dependencies. Furthermore, we introduce an edge-level Shapley attribution framework tailored to protein–ligand interaction graphs, quantifying contributions of chemical bonds (e. g. , hydrophobic contacts) and non-covalent interactions. Experiments show DBGT-PLA reduces RMSE by 18. 3% (from 1. 522 to 1. 244 on the Holdout Set 2019), outperforming state-of-the-art models. Crucially, our explainability module reveals that the ligand edges dominate affinity predictions, accounting for nearly 70%. This work not only advances predictive accuracy but also offers unprecedented, quantitative insights into interaction determinants, which can guide rational drug optimization. The code of DBGT-PLA is publicly available at https://github.com/wangwying/DBGT-PLA

YNICL Journal 2026 Journal Article

Distinct neurologic state in patients with traumatic brain injury and hemorrhagic stroke during the stage of acute disorders of consciousness and the correlation with the neurological prognosis: A multi-modal PET/rs-fMRI study

  • Danjing Yu
  • Kemeng Gao
  • Xiefeng Wang
  • Lin Zhao
  • Yi Sun
  • Zhiyan Shen
  • Yu Wang
  • Ying Wang

PURPOSE: The exact mechanisms underlying the distinct neurological outcomes between Traumatic Brain Injury (TBI) and Hemorrhagic Stroke (HS) remain unclear. Our objective is to assess distinct features of neurologic state between comatose patients with TBI and HS during the stage of acute disorder of consciousness (aDoC) and to identify the correlation of neurologic features with prognosis. METHODS: Data were analyzed from TBI and HS patients examined by positron emission tomography (PET) and resting-state functional magnetic resonance imaging (rs-fMRI) simultaneously. Primary clinical outcomes consisted of the state of consciousness and neurological prognosis. The regional neural activity was assessed by the amplitude of fractional low-frequency fluctuation (fALFF) and regional homogeneity (ReHo) on rs-fMRI scans. The standardized uptake value (SUV) on PET scans quantified neural metabolism. Functional connectivity (FC) and graph theoretic approach (GTA) were employed to compare the FC patterns between TBI and HS. Correlations of PET/rs-fMRI indicators with the prognosis of HS and TBI were identified. RESULTS: Muti-modal PET/rs-fMRI analysis showed more active local neurological state in TBI patients than HS patients, specifically in the right precentral gyrus (PreCG.R), right postcentral gyrus (PoCG.R), right superior temporal gyrus (STG.R) and right middle temporal gyrus (MTG.R). TBI patients demonstrated significantly higher clustering coefficient and nodal efficiency of the sensorimotor network (SMN) along with lower connectivity and network efficiency in the default network (DMN) compared to HS patients. PET/rs-fMRI indicators significantly correlated with the neurological prognosis of TBI and HS. CONCLUSIONS: This study elucidated the underlying mechanisms contributing to the distinct neurologic prognosis between comatose TBI and HS patients, and may contribute to the development of early targeted intervention strategies for specific diseases.

EAAI Journal 2026 Journal Article

Dynamic physics-Weighted Gaussian process regression for robust thermal error prediction under non-stationary conditions

  • Zheng Yan
  • Ying Wang
  • Zhijie Xia
  • Zengtao Chen

To address the trade-off between generalization and interpretability in thermal error compensation for Vertical Machining Centers (VMCs), this paper proposes a Dynamic Physics-Weighted Gaussian Process Regression (DPW-GPR) framework. In terms of Artificial Intelligence contribution, a unified parametric physical model is embedded as a Bayesian prior to impose manifold constraints. To further enhance residual learning capabilities, a physics-guided composite kernel function is meticulously designed to capture multi-scale thermal fluctuations. Crucially, a novel adaptive weighting mechanism driven by Kullback-Leibler (KL) divergence is introduced to quantify the real-time discrepancy between physical priors and data posteriors, acting as a probabilistic switch to seamlessly transition between steady-state physical consistency and transient data-driven learning. Regarding the engineering application, the method is validated on a VH800 VMC under complex non-stationary conditions to predict multi-directional deformations, including spindle elongation, headstock torsion, and column bending. Unlike traditional deterministic models, the proposed framework outputs 95% confidence intervals to quantify epistemic uncertainty, enabling risk-aware decision-making. Experimental results demonstrate that DPW-GPR reduces the Root Mean Square Error (RMSE) of the z-axis prediction to 3. 44 μm, outperforming Support Vector Regression (SVR) and Back Propagation Neural Networks (BPNN) by over 50%. Notably, the model exhibits superior data efficiency, achieving high accuracy with only 10% sparse training samples, significantly surpassing SVR and BPNN trained on full datasets. The proposed approach provides a robust, risk-aware, and data-efficient solution for intelligent manufacturing.

AIJ Journal 2026 Journal Article

Environment promoted invariant information learning for graph out-of-distribution generalization

  • Shuo Wang
  • Mingchen Sun
  • Qiang Huang
  • Ying Wang

Graph out-of-distribution generalization is an important task in graph data mining, which has received extensive attention in many practical applications. In recent years, an increasing number of studies have focused on applying invariant learning and causal learning to enhance the model’s cross environment generalization capability. However, existing methods often neglect subgraph estimation bias during invariant information extraction, which impacts generalization performance. Therefore, to address this issue, we construct the Causality Inspired Environment Promoted Graph Generalization Framework (CEPG), which dynamically corrects subgraph estimation biases and learns the target invariant distribution through integrating multiple constraints. Specifically, we first leverage a subgraph generation module to explicitly obtain invariant and environmental subgraphs by evaluating the link reliability. Then, we design the specific environmental information extraction module to prevent bias propagation from environmental subgraphs and capture domain-specific knowledge. Finally, we construct the environment promoted invariant information learning module. This module can align estimated invariant distribution with the target distribution through the environment promoted and reflection mechanism guidance constraints. Extensive experiments demonstrate that our approach effectively enhances generalization across various types of distribution shifts and outperforms state-of-the-art methods on both synthetic and real-world graph OOD generalization benchmarks.

JBHI Journal 2026 Journal Article

Hierarchical Coarse-to-Fine cGAN for Subtype-Specific Freezing of Gait Signal Generation

  • Xinyue Yu
  • Helena Cockx
  • Ying Wang
  • Richard van Wezel
  • Kaylena Ehgoetz Martens
  • Arash Arami

Freezing of gait (FOG), a debilitating symptom of Parkinson's disease, can manifest in three sub-types: shuffling, trembling, and akinesia, with occurrence and frequency varying across patients. While deep learning (DL) models show promise in FOG detection, their robustness and generalization across subtypes are limited by data scarcity and imbalances between FOG/non-FOG classes and among subtypes. To address this, we propose a subtype-aware FOG augmentation technique enabling training of DL models to perform consistently across subtypes. Specifically, we introduce Hierarchical Coarse-to-Fine conditional Generative Adversary Network (Hi-CF cGAN), a two-stage model that generates subtype-conditioned FOG-like ankle accelerations that are realistic and diverse, as verified through visualization, UMAPs, and Maximum Mean Discrepancy comparison against real signals. We evaluate its effective-ness by training CNNs for FOG detection with both general (subtype-stratified) and personalized (subtype-variant, based on patient-specific subtype composition) augmentation via Hi-CF cGAN, benchmarking against classical augmentations and baseline (no augmentation). Compared to baseline, general augmentation with Hi-CF cGAN effectively improves average detection rates of FOG, trembling FOG, and especially the previously overlooked minor subtypes, shuffling FOG (from 66. 8% to 81. 6%) and akinesia FOG (from 58. 7% to 77. 9%). These improvements exceed those of classical augmentations, demonstrating superior real-ism, richness, and adaptability of Hi-CF cGAN-generated data in addressing FOG/non-FOG and subtype imbalances. Personalized augmentation further enhances accuracy on targeted subtype(s) compared to general augmentation, highlighting its potential for tailored model optimization.

AAAI Conference 2026 Conference Paper

LookFlow: Training-Free and Efficient High-Resolution Image Synthesis via Dynamic Lookahead Guidance Flow

  • Yuan Zhou
  • Yan Zhang
  • Jianlong Chang
  • Xin Gu
  • Ying Wang
  • Kun Ding
  • Guangwen Yang
  • Shiming Xiang

Rectification flow Transformers (RFTs) have shown promising performance in diffusion-based image synthesis but are typically confined to lower-resolution scenarios, limiting their ability to generate high-resolution images. Existing resolution extrapolation approaches often suffer from excessive computational overhead, resulting in prolonged inference times. We propose LookFlow, a training-free high-resolution synthesis framework that accelerates inference while preserving visual quality. Building on pretrained text-to-image RFTs, LookFlow employs a dynamic lookahead guidance flow mechanism to refine high-resolution velocity predictions by leveraging multi-timestep lookahead information extracted from a low-resolution flow. Additionally, reusing temporally similar features across consecutive timesteps drastically reduces computation and significantly decreases inference time overhead. Extensive experiments on COCO demonstrate that LookFlow robustly scales resolutions from 4× to 25×, achieving up to a maximum speedup of 2.01× while maintaining competitive visual fidelity.

AAAI Conference 2026 Conference Paper

Monocular Vehicle Pose and Shape Reconstruction via Dynamic Context Adaptation and Progressive Geometry Refinement

  • Wei Li
  • Long Ji
  • Ying Wang
  • Xiao Wu
  • Zhaoquan Yuan
  • Penglin Dai

Accurate reconstruction of 3D vehicle pose and shape from monocular images is challenging, particularly for distant objects in autonomous driving. Existing methods often suffer from geometric ambiguity in depth estimation and structural hollowness in shape recovery, primarily due to inadequate multi-scale feature aggregation and unflexible prior modeling. To overcome these limitations, MonoVPR is proposed, a novel framework integrating dynamic context adaptation and progressive geometry refinement. Specifically, a Hierarchical Dual-Context Attention (HDCA) module is introduced to resolve scale-dependent degradation through gated cross-attention across multi-resolution feature maps, dynamically fusing object-centric geometric cues with scene-centric semantics. For shape refinement, the Bounded Iterative Mesh Refiner (BIMR) progressively optimizes template-guided deformations via multi-head attention and a tanh-bounded correction loop, ensuring physically plausible reconstructions.Extensive experiments on the ApolloCar3D benchmark demonstrate MonoVPR achieves state-of-the-art performance, showing exceptional capability in reconstructing geometrically consistent shapes and precise poses for challenging long-range scenarios.

AAAI Conference 2026 Conference Paper

PASA: Progressive-Adaptive Spectral Augmentation for Automated Auscultation in Data-Scarce Environments

  • Ying Wang
  • Guoheng Huang
  • Xueyuan Gong
  • Xinxin Wang
  • Xiaochen Yuan

Automated auscultation advances the detection of respiratory diseases, especially in areas with limited resources where traditional diagnostic methods are unavailable. On the other hand, the scarcity of auscultation datasets limits the automation performance, prompting the needs for data augmentation methods. However, most of the existing methods neglect the difference in acoustic sounds that requires personalized augmentation strategies. To address this, we propose a Progressive-Adaptive Spectral Augmentation (PASA), which is one of the first paradigms to adaptively select the best augmentation strategy for each sample. The PASA innovatively treats augmentation selection problem as a Markov Decision Process (MDP), creating an alternating loop between the diagnostic model and the augmentation selection. The agent selects the optimal augmentation operations and magnitudes via a task-specific design, including state construction, action sampling, Hybrid Batch-Sample (HBS) strategy execution, and reward guidance. The HBS strategy initially applies uniform augmentation across mini-batches while collecting sample-specific performance statistics. When model performance stabilizes, it transits to sample-level augmentation based on accumulated difficulty assessments. This two-phase design balances computational complexity with personalization. Extensive experiments across three benchmark datasets demonstrate that the PASA outperforms the state-of-the-art methods, pioneering a transformative paradigm for adaptive data augmentation in automated auscultation.

AAAI Conference 2026 Conference Paper

Radar-APLANC: Unsupervised Radar-based Heartbeat Sensing via Augmented Pseudo-Label and Noise Contrast

  • Ying Wang
  • Zhaodong Sun
  • Xu Cheng
  • Zuxian He
  • Xiaobai Li

Frequency Modulated Continuous Wave (FMCW) radars can measure subtle chest wall oscillations to enable non-contact heartbeat sensing. However, traditional radar-based heartbeat sensing methods face performance degradation due to noise. Learning-based radar methods achieve better noise robustness but require costly labeled signals for supervised training. To overcome these limitations, we propose the first unsupervised framework for radar-based heartbeat sensing via Augmented Pseudo-Label and Noise Contrast (Radar-APLANC). We propose to use both the heartbeat range and noise range within the radar range matrix to construct the positive and negative samples, respectively, for improved noise robustness. Our Noise-Contrastive Triplet (NCT) loss only utilizes positive samples, negative samples, and pseudo-label signals generated by the traditional radar method, thereby avoiding dependence on expensive ground-truth physiological signals. We further design a pseudo-label augmentation approach featuring adaptive noise-aware label selection to improve pseudo-label signal quality. Extensive experiments on the Equipleth dataset and our collected radar dataset demonstrate that our unsupervised method achieves performance comparable to state-of-the-art supervised methods.

ICRA Conference 2025 Conference Paper

3D Dense Captioning via Prototypical Momentum Distillation

  • Jinpeng Mi
  • Ying Wang
  • Shaofei Jin
  • Shiming Zhang
  • Xian Wei
  • Jianwei Zhang

3D dense captioning aims to describe the crucial regions in 3D visual scenes in the form of natural language. Recent prevailing approaches achieve promising results by leveraging complicated structures incorporated with large-scale models, which necessitate abundant parameters and pose challenges regarding its practical applications. Besides, with limited training data, 3D dense captioners are often susceptible to overfitting, directly degrading caption generation performance. Drawing inspiration from the recent advancements in knowledge distillation, we propose a novel approach termed Prototypical Momentum Distillation (PMD) to prompt the model to generate more detailed captions. PMD incorporates Momentum Distillation (MD) with an Uncertainty-aware Prototype-anchored Clustering (UPC) strategy to transfer knowledge by considering the uncertainty of the teacher knowledge. Specifically, we employ the original captioner as the student model and maintain an Exponential Moving Average (EMA) copy of the captioner as the teacher model to impart knowledge as the auxiliary supervision of the student. To abate the misleading caused by uncertain knowledge, we present an Uncertainty-aware Prototype-anchored Clustering (UPC) strategy to cluster the distilled knowledge according to its confidence. We then transfer the rearranged knowledge from the teacher to guide the training route of the student. We conduct extensive experiments and ablation studies on two widely used benchmark datasets, ScanRefer and Nr3D. Experimental results demonstrate that PMD outperforms all state-of-the-art approaches on the benchmarks with MLE training, highlighting its effectiveness.

EAAI Journal 2025 Journal Article

Adaptive Deformable Convolutional Neural Network Framework for depression-related behavioral analysis in mice

  • Jian Li
  • Ziyi Li
  • Peng Shan
  • Xiaoyong Lyu
  • Yu Tian
  • Chen Du
  • Ying Wang
  • Yuliang Zhao

The use of approximately 1 billion laboratory animals annually in research highlights the urgent need for advanced methods to analyze behavioral dynamics, particularly in mice. Capturing subtle and prolonged behavioral changes, such as those observed in long-term depression studies, poses a significant challenge. To address this, we propose an Adaptive Deformable Convolutional Neural Network Framework for depression-related behavioral analysis in mice. By integrating DeepLabCut (DLC) with deformable convolutional networks (DCN) and convolutional block attention module (CBAM), the framework captures subtle and prolonged behavioral changes with high precision. Adaptive image deformation encodes joint movements into image representations, enabling robust analysis of spatial and temporal patterns. In depression modeling experiment, the framework achieved over 80% classification accuracy, demonstrating its scalability and efficiency. This non-invasive, automated solution represents a transformative advancement in behavioral analysis, offering a reliable tool for long-term studies in animal models.

JBHI Journal 2025 Journal Article

AI-Driven Quantitative Analysis of Pathological Images for Membranous Nephropathy Across Macro and Micro Modalities

  • Guangze Shi
  • Ying Wang
  • Yongfei Wu
  • Xueyu Liu
  • Jia Shen
  • Hao Meng
  • Yexin Lai
  • Weixia Han

The diagnosis of membranous nephropathy (MN) has been reliant on the identification of glomerular basement membrane (GBM) variations and lesions at both macro and micro levels. At the macro level, light microscopy (LM) has been used to reveal spike- like projections that indicate pathological changes, whereas at the micro level, transmission electron microscopy (TEM) has been employed to identify GBM thickening. However, qualitative diagnosis has been limited by inter-pathologist variability, creating the need for deep learning approaches capable of quantifying pathological changes and predicting MN progression. In this study, an AI-driven framework based on the Mamba model has been proposed, in which the area and proportion of spike- like projections are quantified at the macro level, and GBM thickness is segmented and measured at the micro level. Classical machine learning models are then applied to predict MN progression based on pathological indicators extracted through factor analysis. Unlike prior approaches, the framework has been designed to emulate the diagnostic workflow of pathologists by integrating LM and TEM images for joint analysis. Experiments on an external dataset of 109 cases demonstrated strong performance in glomeruli classification, GBM segmentation, and MN progression prediction. These findings highlight the potential of multi-scale integrated quantification to provide objective, reproducible, and clinically interpretable assessment of MN progression.

AIJ Journal 2025 Journal Article

BATED: Learning fair representation for Pre-trained Language Models via biased teacher-guided disentanglement

  • Yingji Li
  • Mengnan Du
  • Rui Song
  • Mu Liu
  • Ying Wang

With the rapid development of Pre-trained Language Models (PLMs) and their widespread deployment in various real-world applications, social biases of PLMs have attracted increasing attention, especially the fairness of downstream tasks, which potentially affects the development and stability of society. Among existing debiasing methods, intrinsic debiasing methods are not necessarily effective when applied to downstream tasks, and the downstream fine-tuning process may introduce new biases or catastrophic forgetting. Most extrinsic debiasing methods rely on sensitive attribute words as prior knowledge to supervise debiasing training. However, it is difficult to collect sensitive attribute information of real data due to privacy and regulation. Moreover, limited sensitive attribute words may lead to inadequate debiasing training. To this end, this paper proposes a debiasing method to learn fair representation for PLMs via BiAsed TEacher-guided Disentanglement (called BATED). Specific to downstream tasks, BATED performs debiasing training under the guidance of a biased teacher model rather than relying on sensitive attribute information of the training data. First, we leverage causal contrastive learning to train a task-agnostic general biased teacher model. We then employ Variational Auto-Encoder (VAE) to disentangle the PLM-encoded representation into the fair representation and the biased representation. The Biased representation is further decoupled via biased teacher-guided disentanglement, while the fair representation learn downstream tasks. Therefore, BATED guarantees the performance of downstream tasks while improving the fairness. Experimental results on seven PLMs testing three downstream tasks demonstrate that BATED outperforms the state-of-the-art overall in terms of fairness and performance on downstream tasks.

JBHI Journal 2025 Journal Article

DCTP-Net: Dual-Branch CLIP-Enhance Textual Prompt-Aware Network for Acute Ischemic Stroke Lesion Segmentation From CT Image

  • Jiahao Liu
  • Hongqing Zhu
  • Ziying Wang
  • Ning Chen
  • Tong Hou
  • Bingcang Huang
  • Weiping Lu
  • Ying Wang

Detecting early ischemic lesions (EIL) in computed tomography (CT) images is crucial for reducing diagnostic time and minimizing neuron loss due to oxygen deprivation. This paper introduces DCTP-Net, a dual-branch network for segmenting acute ischemic stroke lesions in CT images, consisting of a segmentation branch and a prompt-aware branch. The segmentation branch uses an encoder-decoder network as the backbone to identify lesions, where the encoder fuses CT image features with prompt features from the prompt-aware branch. To enhance semantic feature extraction and reduce the impact of cerebral structural details, we introduce a cross-collaboration dynamic connection (CCDC) module to link the encoder and decoder. The prompt-aware branch includes a learnable prompt (LP) block to incorporate cerebral prior knowledge, and the prompt-aware encoder (PAE) combines the LP block with multi-level features from the segmentation branch for more precise representation. Additionally, we propose a CLIP-enhance textual prompt (CETP) module that utilizes the CLIP text encoder to generate specialized convolutional parameters for the segmentation head. These parameters are tailored to the unique characteristics of each input image, improving segmentation performance. Qualitative and quantitative studies reveal that DCTP-Net outperforms the current state-of-the-art, IS-Net, with Dice score increases of 3. 9% on AISD and 3. 8% on ISLES2018, demonstrating its superiority in EIL segmentation.

JAIR Journal 2025 Journal Article

Graph Collaborative Filtering Model Combining Time Factor and Attention Mechanism

  • Xianglin Zuo
  • Xin He
  • Tianhao Jia
  • Ying Wang

Recently, with the triumph of deep learning, attention mechanism, and graph convolutional networks in their respective fields, using new representation learning techniques or introducing auxiliary information to improve the representation ability of embedding has become the core content of the recommendation algorithm research. Generally, most existing GNN-based recommendation methods recursively propagate embedding information on the graph structure and capture collaborative signals by exploring the high-level connectivity between users and items. Despite the great success, those methods do not consider the influence of temporal context on user preferences embedding information propagation, nor do they distinguish the contribution of different neighbor node information to the target node. In order to address the two problems, we propose a graph collaborative filtering model TAGCF combing time factors and attention based on the existing method. The model uses the time factor to integrate temporal information into the process of embedding information propagation and uses the attention mechanism to distinguish the influence of embedding information from different neighbors. The effectiveness of TAGCF, time information, and attention mechanism are verified through comparative experiments with multiple baseline methods on the two recommendation system datasets, MovieLens and Amazon-books.

IJCAI Conference 2025 Conference Paper

Graph OOD Detection via Plug-and-Play Energy-based Evaluation and Propagation

  • Yunxia Zhang
  • Mingchen Sun
  • Yutong Zhang
  • Funing Yang
  • Ying Wang

Existing graph neural network (GNN) methods are typically built upon the i. i. d. assumption, emphasizing the enhancement of the test performance for in-distribution (ID) data. However, there has been limited exploration of their adaptability to scenarios involving unknown distribution data. On the one hand, in real-world application scenarios, graph data often expands continuously with the acquisition of external knowledge, which means that new nodes with unknown categories may be added to the graph data. The gap between the new node distribution and the original node distribution can make existing GNN methods less effective. On the other hand, existing out-of-distribution (OOD) detection methods often rely on the softmax confidence score, which makes the OOD data suffer from overconfident posterior distributions. To address the above issues, we propose an Energy Propagation-based Graph Neural Network (EPGNN), which improves the OOD generalization ability by endowing GNN with the capacity to detect the OOD nodes in the graph. Specifically, we first construct GNN encoder to obtain node embedding that incorporates neighborhood structural information. Then, we design a plug-and-play energy-based OOD evaluator by assigning corresponding energy values to different nodes. Finally, we construct a plug-and-play structure-aware energy propagation module and joint alignment regularization, which make the node energy more flexible during the training process. Extensive experiments on benchmark datasets demonstrate the superiority of our method.

AAMAS Conference 2025 Conference Paper

Group-fair Facility Location Games with Externalities

  • Minming Li
  • Cheng Peng
  • Ying Wang
  • Houyu Zhou

We study facility location games with externalities where agents are located on a real line and divided into groups. The cost of an agent is affected by the facility location and their group members. The goal is to design mechanisms to locate a facility to approximately optimize group-fair objectives while eliciting the agents’ locations truthfully. We consider two types of group interactions: competitive and collaborative, and two group-fair objectives, minimizing the maximum total group cost and minimizing the maximum average group cost. For each scenario, we analyze classic mechanisms, presenting their approximation ratios, and introduce new mechanisms that achieve improved approximation ratios. Additionally, we establish tight lower bounds for each setting, demonstrating that our mechanisms are the best possible.

IJCAI Conference 2025 Conference Paper

Hierarchy Knowledge Graph for Parameter-Efficient Entity Embedding

  • Hepeng Gao
  • Funing Yang
  • Yongjian Yang
  • Ying Wang

Traditional knowledge graphs (KGs) provide each entity with a unique embedding as a representation, which contains a lot of redundant information. Meanwhile, the space complexities of the KGs are positively related to the number of entities. In this work, we propose a hierarchical representation learning method, namely HRL, which is a parameter-efficient model where the number of model parameters is independent of dataset scales. Specifically, we propose a hierarchical model comprising a Meta Encoder and a Context Encoder to generate the representation of entities and relations. The Meta Encoder captures the common representations shared across entities, while the Context Encoder learns entity-specific representations. We further provide a theoretical analysis of model design by constructing a structural causal model (SCM) when completing a knowledge graph. The SCM outlines the relationships between nodes, where entity embeddings are conditioned on both common and entity-specific representations. Note that our model is designed to reduce model scale while maintaining competitive performance. We evaluate HRL on the knowledge graph completion task using three real-world datasets. The results demonstrate that HRL significantly outperforms existing parameter-efficient baselines, as well as traditional state-of-the-art baselines of similar scale.

NeurIPS Conference 2025 Conference Paper

Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering

  • yuyang Hong
  • Jiaqi Gu
  • Yang Qi
  • Lubin Fan
  • Yue Wu
  • Ying Wang
  • Kun Ding
  • Shiming Xiang

The task of Knowlegde-Based Visual Question Answering (KB-VQA) requires the model to understand visual features and retrieve external knowledge. Retrieval-Augmented Generation (RAG) have been employed to address this problem through knowledge base querying. However, existing work demonstrate two limitations: insufficient interactivity during knowledge retrieval and ineffective organization of retrieved information for Visual-Language Model (VLM). To address these challenges, we propose a three-stage visual language model with Process, Retrieve and Filter (VLM-PRF) framework. For interactive retrieval, VLM-PRF uses reinforcement learning (RL) to guide the model to strategically process information via tool-driven operations. For knowledge filtering, our method trains the VLM to transform the raw retrieved information into into task-specific knowledge. With a dual reward as supervisory signals, VLM-PRF successfully enable model to optimize retrieval strategies and answer generation capabilities simultaneously. Experiments on two datasets demonstrate the effectiveness of our framework.

NeurIPS Conference 2025 Conference Paper

MMOT: The First Challenging Benchmark for Drone-based Multispectral Multi-Object Tracking

  • Tianhao Li
  • Tingfa Xu
  • Ying Wang
  • Haolin Qin
  • Xu Lin
  • Jianan Li

Drone-based multi-object tracking is essential yet highly challenging due to small targets, severe occlusions, and cluttered backgrounds. Existing RGB-based multi-object tracking algorithms heavily depend on spatial appearance cues such as color and texture, which often degrade in aerial views, compromising tracking reliability. Multispectral imagery, capturing pixel-level spectral reflectance, provides crucial spectral cues that significantly enhance object discriminability under degraded spatial conditions. However, the lack of dedicated multispectral UAV datasets has hindered progress in this domain. To bridge this gap, we introduce MMOT, the first challenging benchmark for drone-based multispectral multi-object tracking dataset. It features three key characteristics: (i) Large Scale — 125 video sequences with over 488. 8K annotations across eight object categories; (ii) Comprehensive Challenges — covering diverse real-world challenges such as extreme small targets, high-density scenarios, severe occlusions and complex platform motion; and (iii) Precise Oriented Annotations — enabling accurate localization and reduced object ambiguity under aerial perspectives. To better extract spectral features and leverage oriented annotations, we further present a multispectral and orientation-aware MOT scheme adapting existing MOT methods, featuring: (i) a lightweight Spectral 3D-Stem integrating spectral features while preserving compatibility with RGB pretraining; (ii) a orientation-aware Kalman filter for precise state estimation; and (iii) an end-to-end orientation-adaptive transformer architecture. Extensive experiments across representative trackers consistently show that multispectral input markedly improves tracking performance over RGB baselines, particularly for small and densely packed objects. We believe our work will benefit the community for advancing drone-based multispectral multi-object tracking research. Our MMOT, code and benchmarks are publicly available at https: //github. com/Annzstbl/MMOT.

AAAI Conference 2025 Conference Paper

Qsco: A Quantum Scoring Module for Open-Set Supervised Anomaly Detection

  • Yifeng Peng
  • Xinyi Li
  • Zhiding Liang
  • Ying Wang

Open set anomaly detection (OSAD) is a crucial task that aims to identify abnormal patterns or behaviors in data sets, especially when the anomalies observed during training do not represent all possible classes of anomalies. The recent advances in quantum computing in handling complex data structures and improving machine learning models herald a paradigm shift in anomaly detection methodologies. This study proposes a Quantum Scoring Module (Qsco), embedding quantum variational circuits into neural networks to enhance the model's processing capabilities in handling uncertainty and unlabeled data. Extensive experiments conducted across eight real-world anomaly detection datasets demonstrate our model's superior performance in detecting anomalies across varied settings and reveal that integrating quantum simulators does not result in prohibitive time complexities. At the same time, the experimental results under different noise models also prove that Qsco is a noise-resilient algorithm. Our study validates the feasibility of quantum-enhanced anomaly detection methods in practical applications.

JAIR Journal 2025 Journal Article

Rumor Detection with Adaptive Data Augmentation and Adversarial Training

  • Ying Wang
  • Fuyuan Ma
  • Zhaoqi Yang
  • Yaodi Zhu
  • Bo Yang
  • Pengfei Shen
  • Lei Yun

Rumors are widely spread on social media, which has a negative impact on social stability. To address this problem, many rumor detection methods have been proposed. However, most existing methods overlook the potential impact of noise and adversarial attacks on their detection performance, which could compromise their effectiveness when applied in an unknown environment. To overcome these challenges and improve the framework robustness to noise and adversarial attacks, we propose a novel rumor detection framework with Adaptive Data Augmentation and Adversarial Training, named ADAAT. Our framework utilizes the adaptive data augmentation module to calculate the importance of edges and features and adaptively modify the less important among them with a greater probability. In addition, it contains a hard sample generation module which generates adversarial representations through adversarial training. These adversarial representations are treated as hard samples, which are utilized in contrastive learning to learn essential features, thereby improving the robustness of the framework. Our framework proves superiority in rumor detection tasks, increasing the accuracy by an average of 3.6%, 4.5% and 2.5% over the state-of-the-art methods on Twitter15, Twitter16 and PHEME, respectively. When the ADAAT framework is applied to attacked test data, the detection accuracy decreases by only 1.3%, 1.4%, and 1.2%. This paper appears in the AI & Society Track.

NeurIPS Conference 2025 Conference Paper

TITAN: A Trajectory-Informed Technique for Adaptive Parameter Freezing in Large-Scale VQE

  • Yifeng Peng
  • Xinyi Li
  • Samuel Yen-Chi Chen
  • Kaining Zhang
  • Zhiding Liang
  • Ying Wang
  • Yuxuan Du

Variational quantum Eigensolver (VQE) is a leading candidate for harnessing quantum computers to advance quantum chemistry and materials simulations, yet its training efficiency deteriorates rapidly for large Hamiltonians. Two issues underlie this bottleneck: (i) the no-cloning theorem imposes a linear growth in circuit evaluations with the number of parameters per gradient step; and (ii) deeper circuits encounter barren plateaus (BPs), leading to exponentially increasing measurement overheads. To address these challenges, here we propose a deep learning framework, dubbed Titan, which identifies and freezes inactive parameters of a given ansätze at initialization for a specific class of Hamiltonians, reducing the optimization overhead without sacrificing accuracy. The motivation of Titan starts with our empirical findings that a subset of parameters consistently has negligible influence on training dynamics. Its design combines a theoretically grounded data construction strategy, ensuring each training example is informative and BP-resilient, with an adaptive neural architecture that generalizes across ansätze of varying sizes. Across benchmark transverse-field Ising models, Heisenberg models, and multiple molecule systems up to $30$ qubits, Titan achieves up to $3\times$ faster convergence and $40$–$60\%$ fewer circuit evaluations than state-of-the-art baselines, while matching or surpassing their estimation accuracy. By proactively trimming parameter space, Titan lowers hardware demands and offers a scalable path toward utilizing VQE to advance practical quantum chemistry and materials science.

AIIM Journal 2025 Journal Article

Utilizing semantically enhanced self-supervised graph convolution and multi-head attention fusion for herb recommendation

  • Xianlun Tang
  • Yuze Tang
  • Xinran Liu
  • Haochuan Zhang
  • Xiaoyuan Dang
  • Ying Wang
  • Zihui Xu

Traditional Chinese herbal medicine has long been recognized as an effective natural therapy. Recently, the development of recommendation systems for herbs has garnered widespread academic attention, as these systems significantly impact the application of traditional Chinese medicine. However, existing herb recommendation systems are limited by data sparsity, insufficient correlation between prescriptions, and inadequate representation of symptoms and herb characteristics. To address these issues, this paper introduces an approach to herb recommendation based on semantically enhanced self-supervised graph convolution and multi-head attention fusion (BSGAM). This method involves efficient embedding of entities following fine-tuning of BERT; leveraging the attributes of herbs to optimize feature representation through a residual graph convolution network and self-supervised learning; and ultimately employing a multi-head attention mechanism for feature integration and recommendation. Experiments conducted on a publicly available traditional Chinese medicine prescription dataset demonstrate that our method achieves improvements of 6. 80%, 7. 46%, and 6. 60% in F1-Score@5, F1-Score@10, and F1-Score@20, respectively, compared to baseline methods. These results confirm the effectiveness of our approach in enhancing the accuracy of herb recommendations.

JBHI Journal 2024 Journal Article

Automated Prediction of Infant Cognitive Development Risk by Video: A Pilot Study

  • Shengjie Ji
  • Dan Ma
  • Lunxin Pan
  • Wenan Wang
  • Xiaohang Peng
  • Joan Toluwani Amos
  • Honorine Niyigena Ingabire
  • Min Li

Objective: Cognition is an essential human function, and its development in infancy is crucial. Traditionally, pediatricians used clinical observation or medical imaging to assess infants’ current cognitive development (CD) status. The object of pediatricians’ greater concern is however their future outcomes, because high-risk infants can be identified early in life for intervention. However, this opportunity has not yet been realized. Fortunately, some recent studies have shown that the general movement (GM) performance of infants around 3–4 months after birth might reflect their future CD status, which gives us an opportunity to achieve this goal by cameras and artificial intelligence. Methods: First, infants’ GM videos were recorded by cameras, from which a series of features reflecting their bilateral movement symmetry (BMS) were extracted. Then, after at least eight months of natural growth, the infants’ CD status was evaluated by the Bayley Infant Development Scale, and they were divided into high-risk and low-risk groups. Finally, the BMS features extracted from the early recorded GM videos were fed into the classifiers, using late infant CD risk assessment as the prediction target. Results: The area under the curve, recall and precision values reached 0. 830, 0. 832, and 0. 823 for two-group classification, respectively. Conclusion: This pilot study demonstrates that it is possible to automatically predict the CD of infants around the age of one year based on their GMs recorded early in life. Significance: This study not only helps clinicians better understand infant CD mechanisms, but also provides an economical, portable and non-invasive way to screen infants at high-risk early to facilitate their recovery.

YNIMG Journal 2024 Journal Article

Cerebellar representation during phonetic processing in tonal and non-tonal language speakers: An ALE meta-analysis

  • Xiaotong Zhang
  • Zhaowen Zhou
  • Ying Wang
  • Jinyi Long
  • Zhuoming Chen

The role of the cerebellum in phonetic processing has been discovered and widely discussed for decades. However, with the idea that the cerebral representation of phonetic processing is different in tonal language and non-tonal language speakers, whether the cerebellar representation of phonetic processing differs based on language background remains unknown. In the present study, we conducted an activation likelihood estimation (ALE) analysis among 33 functional neuroimaging studies involving 541 healthy adults (213 tonal language speakers and 328 non-tonal language speakers). The aim was to explore the cerebellar representation of phonetic perception and phonetic production in these two language backgrounds. Our results demonstrated the involvement of cerebellum left Crus I, right Crus II, lobules VI, and VIIb in phonetic perception among tonal language speakers, whereas only one focal cluster (right Crus I and Crus II) was demonstrated in non-tonal language speakers. Conjunction analysis revealed overlapping regions located in the right Crus II both in tonal and non-tonal language speakers during phonetic perception. During phonetic production, no significant cluster was detected among tonal language speakers, whereas one focal cluster (within right lobule VI) was detected in non-tonal language speakers. These results highlight the specific cerebellar representation of phonetic processing in tonal and non-tonal languages. Overall, this ALE analysis provides a profound view of the neural mechanism of phonetic processing.

JBHI Journal 2024 Journal Article

Collaborative Transfer Network for Multi-Classification of Breast Cancer Histopathological Images

  • Liangliang Liu
  • Ying Wang
  • Pei Zhang
  • Hongbo Qiao
  • Tong Sun
  • Hui Zhang
  • Xue Xu
  • Hongcai Shang

The incidence of breast cancer is increasing rapidly around the world. Accurate classification of the breast cancer subtype from hematoxylin and eosin images is the key to improve the precision of treatment. However, the high consistency of disease subtypes and uneven distribution of cancer cells seriously affect the performance of multi-classification methods. Furthermore, it is difficult to apply existing classification methods to multiple datasets. In this article, we propose a collaborative transfer network (CTransNet) for multi-classification of breast cancer histopathological images. CTransNet consists of a transfer learning backbone branch, a residual collaborative branch, and a feature fusion module. The transfer learning branch adopts the pre-trained DenseNet structure to extract image features from ImageNet. The residual branch extracts target features from pathological images in a collaborative manner. The feature fusion strategy of optimizing these two branches is used to train and fine-tune CTransNet. Experiments show that CTransNet achieves 98. 29% classification accuracy on the public BreaKHis breast cancer dataset, exceeding the performance of state-of-the-art methods. Visual analysis is carried out under the guidance of oncologists. Based on the training parameters of the BreaKHis dataset, CTransNet achieves superior performance on other two public breast cancer datasets (breast-cancer-grade-ICT and ICIAR2018_BACH_Challenge), indicating that CTransNet has good generalization performance.

JBHI Journal 2024 Journal Article

Combination of Channel Reordering Strategy and Dual CNN-LSTM for Epileptic Seizure Prediction Using Three iEEG Datasets

  • Xiaoshuang Wang
  • Ziheng Gao
  • Meiyan Zhang
  • Ying Wang
  • Lin Yang
  • Jianwen Lin
  • Tommi Kärkkäinen
  • Fengyu Cong

Objective: Intracranial electroencephalogram (iEEG) signals are generally recorded using multiple channels, and channel selection is therefore a significant means in studying iEEG-based seizure prediction. For n channels, $2^{\text{n}}{-1}$ channel cases can be generated for selection. However, by this means, an increase in n can cause an exponential increase in computational consumption, which may result in a failure of channel selection when n is too large. Hence, it is necessary to explore reasonable channel selection strategies under the premise of controlling computational consumption and ensuring high classification accuracy. Given this, we propose a novel method of channel reordering strategy combined with dual CNN-LSTM for effectively predicting seizures. Method: First, for each patient with n channels, interictal and preictal iEEG samples from each single channel are input into the CNN-LSTM model for classification. Then, the F1-score of each single channel is calculated, and the channels are reordered in descending order according to the size of F1-scores ( channel reordering strategy ). Next, iEEG signals with an increasing number of channels are successively fed into the CNN-LSTM model for classification again. Finally, according to the classification results from n channel cases, the channel case with the highest classification rate is selected. Results: Our method is evaluated on the three iEEG datasets: the Freiburg, the SWEC-ETHZ and the American Epilepsy Society Seizure Prediction Challenge (AES-SPC). At the event-based level, the sensitivities of 100%, 100% and 90. 5%, and the false prediction rates (FPRs) of 0. 10/h, 0/h and 0. 47/h, are achieved for the three datasets, respectively. Moreover, compared to an unspecific random predictor, our method also shows a better performance for all patients and dogs from the three datasets. At the segment-based level, the sensitivities-specificities-accuracies-AUCs of 88. 1%–94. 0%–93. 5%–0. 9101, 99. 1%–99. 7%–99. 6%–0. 9935, and 69. 2%–79. 9%–78. 2%–0. 7373, are attained for the three datasets, respectively. Conclusion: Our method can effectively predict seizures and address the challenge of an excessive number of channels during channel selection.

NeurIPS Conference 2024 Conference Paper

Instance-adaptive Zero-shot Chain-of-Thought Prompting

  • Xiaosong Yuan
  • Chen Shen
  • Shaotian Yan
  • Xiaofeng Zhang
  • Liang Xie
  • Wenxiao Wang
  • Renchu Guan
  • Ying Wang

Zero-shot Chain-of-Thought (CoT) prompting emerges as a simple and effective strategy for enhancing the performance of large language models (LLMs) in real-world reasoning tasks. Nonetheless, the efficacy of a singular, task-level prompt uniformly applied across the whole of instances is inherently limited since one prompt cannot be a good partner for all, a more appropriate approach should consider the interaction between the prompt and each instance meticulously. This work introduces an instance-adaptive prompting algorithm as an alternative zero-shot CoT reasoning scheme by adaptively differentiating good and bad prompts. Concretely, we first employ analysis on LLMs through the lens of information flow to detect the mechanism under zero-shot CoT reasoning, in which we discover that information flows from question to prompt and question to rationale jointly influence the reasoning results most. We notice that a better zero-shot CoT reasoning needs the prompt to obtain semantic information from the question then the rationale aggregates sufficient information from the question directly and via the prompt indirectly. On the contrary, lacking any of those would probably lead to a bad one. Stem from that, we further propose an instance-adaptive prompting strategy (IAP) for zero-shot CoT reasoning. Experiments conducted with LLaMA-2, LLaMA-3, and Qwen on math, logic, and commonsense reasoning tasks (e. g. , GSM8K, MMLU, Causal Judgement) obtain consistent improvement, demonstrating that the instance-adaptive zero-shot CoT prompting performs better than other task-level methods with some curated prompts or sophisticated procedures, showing the significance of our findings in the zero-shot CoT reasoning mechanism.

JBHI Journal 2024 Journal Article

KFDAE: CircRNA-Disease Associations Prediction Based on Kernel Fusion and Deep Auto-Encoder

  • Wen-Yue Kang
  • Ying-Lian Gao
  • Ying Wang
  • Feng Li
  • Jin-Xing Liu

CircRNA has been proved to play an important role in the diseases diagnosis and treatment. Considering that the wet-lab is time-consuming and expensive, computational methods are viable alternative in these years. However, the number of circRNA-disease associations (CDAs) that can be verified is relatively few, and some methods do not take full advantage of dependencies between attributes. To solve these problems, this paper proposes a novel method based on Kernel Fusion and Deep Auto-encoder (KFDAE) to predict the potential associations between circRNAs and diseases. Firstly, KFDAE uses a non-linear method to fuse the circRNA similarity kernels and disease similarity kernels. Then the vectors are connected to make the positive and negative sample sets, and these data are send to deep auto-encoder to reduce dimension and extract features. Finally, three-layer deep feedforward neural network is used to learn features and gain the prediction score. The experimental results show that compared with existing methods, KFDAE achieves the best performance. In addition, the results of case studies prove the effectiveness and practical significance of KFDAE, which means KFDAE is able to capture more comprehensive information and generate credible candidate for subsequent wet-lab.

AIJ Journal 2024 Journal Article

Mitigating social biases of pre-trained language models via contrastive self-debiasing with double data augmentation

  • Yingji Li
  • Mengnan Du
  • Rui Song
  • Xin Wang
  • Mingchen Sun
  • Ying Wang

Pre-trained Language Models (PLMs) have been shown to inherit and even amplify the social biases contained in the training corpus, leading to undesired stereotype in real-world applications. Existing techniques for mitigating the social biases of PLMs mainly rely on data augmentation with manually designed prior knowledge or fine-tuning with abundant external corpora to debias. However, these methods are not only limited by artificial experience, but also consume a lot of resources to access all the parameters of the PLMs and are prone to introduce new external biases when fine-tuning with external corpora. In this paper, we propose a Contrastive Self-Debiasing Model with Double Data Augmentation (named CD3) for mitigating social biases of PLMs. Specifically, CD3 consists of two stages: double data augmentation and contrastive self-debiasing. First, we build on counterfactual data augmentation to perform a secondary augmentation using biased prompts that are automatically searched by maximizing the differences in PLMs' encoding across demographic groups. Double data augmentation further amplifies the biases between sample pairs to break the limitations of previous debiasing models that heavily rely on prior knowledge in data augmentation. We then leverage the augmented data for contrastive learning to train a plug-and-play adapter to mitigate the social biases in PLMs' encoding without tuning the PLMs. Extensive experimental results on BERT, ALBERT, and RoBERTa on several real-world datasets and fairness metrics show that CD3 outperforms baseline models on gender debiasing and race debiasing while retaining the language modeling capabilities of PLMs.

AAMAS Conference 2024 Conference Paper

Positive Intra-Group Externalities in Facility Location

  • Ying Wang
  • Houyu Zhou
  • Minming Li

We study facility location games with multiple groups in one dimension where an agent’s utility is not only decided by the distance from the facility but also by their group members. The positive effect of the interactions within a group is captured by positive intra-group externalities. Our goal is to design a mechanism that is non-manipulable and respects unanimity while (approximately) optimizing an objective function. We consider three types of manipulation, misreporting only the location, misreporting only the group membership, and misreporting both, under two social objectives, the social utility and the minimum utility. For both objectives, we achieve nearly tight bounds by either designing new mechanisms or extending the existing mechanisms in terms of the first two types of manipulation. As to the negative result, we show that strategyproofness and unanimity are incompatible when each agent can misreport both the location and the group membership, which is independent of the objective functions.

YNIMG Journal 2024 Journal Article

Profiling cortical morphometric similarity in perinatal brains: Insights from development, sex difference, and inter-individual variation

  • Ying Wang
  • Dalin Zhu
  • Leilei Zhao
  • Xiaomin Wang
  • Zhe Zhang
  • Bin Hu
  • Dan Wu
  • Weihao Zheng

The topological organization of the macroscopic cortical networks important for the development of complex brain functions. However, how the cortical morphometric organization develops during the third trimester and whether it demonstrates sexual and individual differences at this particular stage remain unclear. Here, we constructed the morphometric similarity network (MSN) based on morphological and microstructural features derived from multimodal MRI of two independent cohorts (cross-sectional and longitudinal) scanned at 30-44 postmenstrual weeks (PMW). Sex difference and inter-individual variations of the MSN were also examined on these cohorts. The cross-sectional analysis revealed that both network integration and segregation changed in a nonlinear biphasic trajectory, which was supported by the results obtained from longitudinal analysis. The community structure showed remarkable consistency between bilateral hemispheres and maintained stability across PMWs. Connectivity within the primary cortex strengthened faster than that within high-order communities. Compared to females, male neonates showed a significant reduction in the participation coefficient within prefrontal and parietal cortices, while their overall network organization and community architecture remained comparable. Furthermore, by using the morphometric similarity as features, we achieved over 65 % accuracy in identifying an individual at term-equivalent age from images acquired after birth, and vice versa. These findings provide comprehensive insights into the development of morphometric similarity throughout the perinatal cortex, enhancing our understanding of the establishment of neuroanatomical organization during early life.

YNIMG Journal 2024 Journal Article

The enhanced connectivity between the frontoparietal, somatomotor network and thalamus as the most significant network changes of chronic low back pain

  • Kun Zhu
  • Jianchao Chang
  • Siya Zhang
  • Yan Li
  • Junxun Zuo
  • Haoyu Ni
  • Bingyong Xie
  • Jiyuan Yao

The prolonged duration of chronic low back pain (cLBP) inevitably leads to changes in the cognitive, attentional, sensory and emotional processing brain regions. Currently, it remains unclear how these alterations are manifested in the interplay between brain functional and structural networks. This study aimed to predict the Oswestry Disability Index (ODI) in cLBP patients using multimodal brain magnetic resonance imaging (MRI) data and identified the most significant features within the multimodal networks to aid in distinguishing patients from healthy controls (HCs). We constructed dynamic functional connectivity (dFC) and structural connectivity (SC) networks for all participants (n = 112) and employed the Connectome-based Predictive Modeling (CPM) approach to predict ODI scores, utilizing various feature selection thresholds to identify the most significant network change features in dFC and SC outcomes. Subsequently, we utilized these significant features for optimal classifier selection and the integration of multimodal features. The results revealed enhanced connectivity among the frontoparietal network (FPN), somatomotor network (SMN) and thalamus in cLBP patients compared to HCs. The thalamus transmits pain-related sensations and emotions to the cortical areas through the dorsolateral prefrontal cortex (dlPFC) and primary somatosensory cortex (SI), leading to alterations in whole-brain network functionality and structure. Regarding the model selection for the classifier, we found that Support Vector Machine (SVM) best fit these significant network features. The combined model based on dFC and SC features significantly improved classification performance between cLBP patients and HCs (AUC=0.9772). Finally, the results from an external validation set support our hypotheses and provide insights into the potential applicability of the model in real-world scenarios. Our discovery of enhanced connectivity between the thalamus and both the dlPFC (FPN) and SI (SMN) provides a valuable supplement to prior research on cLBP.

AAAI Conference 2024 Conference Paper

Weak Distribution Detectors Lead to Stronger Generalizability of Vision-Language Prompt Tuning

  • Kun Ding
  • Haojian Zhang
  • Qiang Yu
  • Ying Wang
  • Shiming Xiang
  • Chunhong Pan

We propose a generalized method for boosting the generalization ability of pre-trained vision-language models (VLMs) while fine-tuning on downstream few-shot tasks. The idea is realized by exploiting out-of-distribution (OOD) detection to predict whether a sample belongs to a base distribution or a novel distribution and then using the score generated by a dedicated competition based scoring function to fuse the zero-shot and few-shot classifier. The fused classifier is dynamic, which will bias towards the zero-shot classifier if a sample is more likely from the distribution pre-trained on, leading to improved base-to-novel generalization ability. Our method is performed only in test stage, which is applicable to boost existing methods without time-consuming re-training. Extensive experiments show that even weak distribution detectors can still improve VLMs' generalization ability. Specifically, with the help of OOD detectors, the harmonic mean of CoOp and ProGrad increase by 2.6 and 1.5 percentage points over 11 recognition datasets in the base-to-novel setting.

YNIMG Journal 2023 Journal Article

Accurate and Efficient Simulation of Very High-Dimensional Neural Mass Models with Distributed-Delay Connectome Tensors

  • Anisleidy González Mitjans
  • Deirel Paz Linares
  • Carlos López Naranjo
  • Ariosky Areces Gonzalez
  • Min Li
  • Ying Wang
  • Ronaldo Garcia Reyes
  • Maria L. Bringas-Vega

This paper introduces methods and a novel toolbox that efficiently integrates high-dimensional Neural Mass Models (NMMs) specified by two essential components. The first is the set of nonlinear Random Differential Equations (RDEs) of the dynamics of each neural mass. The second is the highly sparse three-dimensional Connectome Tensor (CT) that encodes the strength of the connections and the delays of information transfer along the axons of each connection. To date, simplistic assumptions prevail about delays in the CT, often assumed to be Dirac-delta functions. In reality, delays are distributed due to heterogeneous conduction velocities of the axons connecting neural masses. These distributed-delay CTs are challenging to model. Our approach implements these models by leveraging several innovations. Semi-analytical integration of RDEs is done with the Local Linearization (LL) scheme for each neural mass, ensuring dynamical fidelity to the original continuous-time nonlinear dynamic. This semi-analytic LL integration is highly computationally-efficient. In addition, a tensor representation of the CT facilitates parallel computation. It also seamlessly allows modeling distributed delays CT with any level of complexity or realism. This ease of implementation includes models with distributed-delay CTs. Consequently, our algorithm scales linearly with the number of neural masses and the number of equations they are represented with, contrasting with more traditional methods that scale quadratically at best. To illustrate the toolbox's usefulness, we simulate a single Zetterberg-Jansen and Rit (ZJR) cortical column, a single thalmo-cortical unit, and a toy example comprising 1000 interconnected ZJR columns. These simulations demonstrate the consequences of modifying the CT, especially by introducing distributed delays. The examples illustrate the complexity of explaining EEG oscillations, e.g., split alpha peaks, since they only appear for distinct neural masses. We provide an open-source Script for the toolbox.

JBHI Journal 2023 Journal Article

BertNDA: A Model Based on Graph-Bert and Multi-Scale Information Fusion for ncRNA-Disease Association Prediction

  • Zhiwei Ning
  • Jinyang Wu
  • Yidong Ding
  • Ying Wang
  • Qinke Peng
  • Laiyi Fu

Non-coding RNAs (ncRNAs) are a class of RNA molecules that lack the ability to encode proteins in human cells, but play crucial roles in various biological process. Understanding the interactions between different ncRNAs and their impact on diseases can significantly contribute to diagnosis, prevention, and treatment of diseases. However, predicting tertiary interactions between ncRNAs and diseases based on structural information in multiple scales remains a challenging task. To address this challenge, we propose a method called BertNDA, aiming to predict potential relationships between miRNAs, lncRNAs, and diseases. The framework identifies the local information through connectionless subgraph, which aggregate neighbor nodes’ feature. And global information is extracted by leveraging Laplace transform of graph structures and WL (Weisfeiler-Lehman) absolute role coding. Additionally, an EMLP (Element-wise MLP) structure is designed to fuse pairwise global information. The transformer-encoder is employed as the backbone of our approach, followed by a prediction-layer to output the final correlation score. Extensive experiments demonstrate that BertNDA outperforms state-of-the-art methods in prediction assignment and exhibits significant potential for various biological applications. Moreover, we develop an online prediction platform that incorporates the prediction model, providing users with an intuitive and interactive experience. Overall, our model offers an efficient, accurate, and comprehensive tool for predicting tertiary associations between ncRNAs and diseases.

YNIMG Journal 2023 Journal Article

Cortical encoding of rhythmic kinematic structures in biological motion

  • Li Shen
  • Xiqian Lu
  • Xiangyong Yuan
  • Ruichen Hu
  • Ying Wang
  • Yi Jiang

Biological motion (BM) perception is of great survival value to human beings. The critical characteristics of BM information lie in kinematic cues containing rhythmic structures. However, how rhythmic kinematic structures of BM are dynamically represented in the brain and contribute to visual BM processing remains largely unknown. Here, we probed this issue in three experiments using electroencephalogram (EEG). We found that neural oscillations of observers entrained to the hierarchical kinematic structures of the BM sequences (i.e., step-cycle and gait-cycle for point-light walkers). Notably, only the cortical tracking of the higher-level rhythmic structure (i.e., gait-cycle) exhibited a BM processing specificity, manifested by enhanced neural responses to upright over inverted BM stimuli. This effect could be extended to different motion types and tasks, with its strength positively correlated with the perceptual sensitivity to BM stimuli at the right temporal brain region dedicated to visual BM processing. Modeling results further suggest that the neural encoding of spatiotemporally integrative kinematic cues, in particular the opponent motions of bilateral limbs, drives the selective cortical tracking of BM information. These findings underscore the existence of a cortical mechanism that encodes periodic kinematic features of body movements, which underlies the dynamic construction of visual BM perception.

YNIMG Journal 2023 Journal Article

Diagnosing Parkinson's disease by combining neuromelanin and iron imaging features using an automated midbrain template approach

  • Mojtaba Jokar
  • Zhijia Jin
  • Pei Huang
  • Ying Wang
  • Youmin Zhang
  • Yan Li
  • Zenghui Cheng
  • Yu Liu

BACKGROUND AND PURPOSE: Early diagnosis of Parkinson's disease (PD) is still a clinical challenge. Most previous studies using manual or semi-automated methods for segmenting the substantia nigra (SN) are time-consuming and, despite raters being well-trained, individual variation can be significant. In this study, we used a template-based, automatic, SN subregion segmentation pipeline to detect the neuromelanin (NM) and iron features in the SN and SN pars compacta (SNpc) derived from a single 3D magnetization transfer contrast (MTC) gradient echo (GRE) sequence in an attempt to develop a comprehensive imaging biomarker that could be used to diagnose PD. MATERIALS AND METHODS: volume, SNpc volume and iron content with a variety of thresholds as well as the N1 sign in diagnosing PD. Correlation analyses were performed to study the relationship between these imaging measures and the clinical scales in PD. RESULTS: = 0.04, p = 0.013) in PD patients. CONCLUSION: volume, SNpc volume and iron content) resulted in an AUC of 0.947 and provided a comprehensive set of imaging biomarkers that, potentially, could be used to diagnose PD clinically.

JBHI Journal 2023 Journal Article

LncDLSM: Identification of Long Non-Coding RNAs With Deep Learning-Based Sequence Model

  • Ying Wang
  • Pengfei Zhao
  • Hongkai Du
  • Yingxin Cao
  • Qinke Peng
  • Laiyi Fu

Long non-coding RNAs (LncRNAs) serve a vital role in regulating gene expressions and other biological processes. Differentiation of lncRNAs from protein-coding transcripts helps researchers dig into the mechanism of lncRNA formation and its downstream regulations related to various diseases. Previous works have been proposed to identify lncRNAs, including traditional bio-sequencing and machine learning approaches. Considering the tedious work of biological characteristic-based feature extraction procedures and inevitable artifacts during bio-sequencing processes, those lncRNA detection methods are not always satisfactory. Hence, in this work, we presented lncDLSM, a deep learning-based framework differentiating lncRNA from other protein-coding transcripts without dependencies on prior biological knowledge. lncDLSM is a helpful tool for identifying lncRNAs compared with other biological feature-based machine learning methods and can be applied to other species by transfer learning achieving satisfactory results. Further experiments showed that different species display distinct boundaries among distributions corresponding to the homology and the specificity among species, respectively.

JBHI Journal 2023 Journal Article

MSGCA: Drug-Disease Associations Prediction Based on Multi-Similarities Graph Convolutional Autoencoder

  • Ying Wang
  • Ying-Lian Gao
  • Juan Wang
  • Feng Li
  • Jin-Xing Liu

Identifying drug-disease associations (DDAs) is critical to the development of drugs. Traditional methods to determine DDAs are expensive and inefficient. Therefore, it is imperative to develop more accurate and effective methods for DDAs prediction. Most current DDAs prediction methods utilize original DDAs matrix directly. However, the original DDAs matrix is sparse, which greatly affects the prediction consequences. Hence, a prediction method based on multi-similarities graph convolutional autoencoder (MSGCA) is proposed for DDAs prediction. First, MSGCA integrates multiple drug similarities and disease similarities using centered kernel alignment-based multiple kernel learning (CKA-MKL) algorithm to form new drug similarity and disease similarity, respectively. Second, the new drug and disease similarities are improved by linear neighborhood, and the DDAs matrix is reconstructed by weighted K nearest neighbor profiles. Next, the reconstructed DDAs and the improved drug and disease similarities are integrated into a heterogeneous network. Finally, the graph convolutional autoencoder with attention mechanism is utilized to predict DDAs. Compared with extant methods, MSGCA shows superior results on three datasets. Furthermore, case studies further demonstrate the reliability of MSGCA.

AAMAS Conference 2023 Conference Paper

Sybil-Proof Diffusion Auction in Social Networks

  • Hongyin Chen
  • Xiaotie Deng
  • Ying Wang
  • Yue Wu
  • Dengji Zhao

A diffusion auction is a market to sell commodities over a social network, where the challenge is to incentivize existing buyers to invite their neighbors in the network to join the market. Existing mechanisms have been designed to solve the challenge in various settings, aiming at desirable properties such as non-deficiency, incentive compatibility and social welfare maximization. Since the mechanisms are employed in dynamic networks with ever-changing structures, buyers could easily generate fake nodes in the network to manipulate the mechanisms for their own benefits, which is commonly known as the Sybil attack. We observe that strategic agents may gain an unfair advantage in existing mechanisms through such attacks. To resist this potential attack, we propose two diffusion auction mechanisms, the Sybil tax mechanism (STM) and the Sybil cluster mechanism (SCM), to achieve both Sybil-proofness and incentive compatibility in the single-item setting. Our proposal provides the first mechanisms to protect the interests of buyers against Sybil attacks with a mild sacrifice of social welfare and revenue.

YNIMG Journal 2023 Journal Article

The iron burden of cerebral microbleeds contributes to brain atrophy through the mediating effect of white matter hyperintensity

  • Ke Lv
  • Yanzhen Liu
  • Yongsheng Chen
  • Sagar Buch
  • Ying Wang
  • Zhuo Yu
  • Huiying Wang
  • Chenxi Zhao

The goal of this work was to explore the total iron burden of cerebral microbleeds (CMBs) using a semi-automatic quantitative susceptibility mapping and to establish its effect on brain atrophy through the mediating effect of white matter hyperintensities (WMH). A total of 95 community-dwelling people were enrolled. Quantitative susceptibility mapping (QSM) combined with a dynamic programming algorithm (DPA) was used to measure the characteristics of 1309 CMBs. WMH were evaluated according to the Fazekas scale, and brain atrophy was assessed using a 2D linear measurement method. Histogram analysis was used to explore the distribution of CMBs susceptibility, volume, and total iron burden, while a correlation analysis was used to explore the relationship between volume and susceptibility. Stepwise regression analysis was used to analyze the risk factors for CMBs and their contribution to brain atrophy. Mediation analysis was used to explore the interrelationship between CMBs and brain atrophy. We found that the frequency distribution of susceptibility of the CMBs was Gaussian in nature with a mean of 201 ppb and a standard deviation of 84 ppb; however, the volume and total iron burden of CMBs were more Rician in nature. A weak but significant correlation between the susceptibility and volume of CMBs was found (r = -0.113, P < 0.001). The periventricular WMH (PVWMH) was a risk factor for the presence of CMBs (number: β = 0.251, P = 0.014; volume: β = 0.237, P = 0.042; total iron burden: β = 0.238, P = 0.020) and was a risk factor for brain atrophy (third ventricle width: β = 0.325, P = 0.001; Evans's index: β = 0.323, P = 0.001). PVWMH had a significant mediating effect on the correlation between CMBs and brain atrophy. In conclusion, QSM along with the DPA can measure the total iron burden of CMBs. PVWMH might be a risk factor for CMBs and may mediate the effect of CMBs on brain atrophy.

NeurIPS Conference 2023 Conference Paper

Visual Explanations of Image-Text Representations via Multi-Modal Information Bottleneck Attribution

  • Ying Wang
  • Tim G. J. Rudner
  • Andrew G. Wilson

Vision-language pretrained models have seen remarkable success, but their application to safety-critical settings is limited by their lack of interpretability. To improve the interpretability of vision-language models such as CLIP, we propose a multi-modal information bottleneck (M2IB) approach that learns latent representations that compress irrelevant information while preserving relevant visual and textual features. We demonstrate how M2IB can be applied to attribution analysis of vision-language pretrained models, increasing attribution accuracy and improving the interpretability of such models when applied to safety-critical domains such as healthcare. Crucially, unlike commonly used unimodal attribution methods, M2IB does not require ground truth labels, making it possible to audit representations of vision-language pretrained models when multiple modalities but no ground-truth data is available. Using CLIP as an example, we demonstrate the effectiveness of M2IB attribution and show that it outperforms gradient-based, perturbation-based, and attention-based attribution methods both qualitatively and quantitatively.

AAAI Conference 2022 Conference Paper

ACDNet: Adaptively Combined Dilated Convolution for Monocular Panorama Depth Estimation

  • Chuanqing Zhuang
  • Zhengda Lu
  • Yiqun Wang
  • Jun Xiao
  • Ying Wang

Depth estimation is a crucial step for 3D reconstruction with panorama images in recent years. Panorama images maintain the complete spatial information but introduce distortion with equirectangular projection. In this paper, we propose an ACDNet based on the adaptively combined dilated convolution to predict the dense depth map for a monocular panoramic image. Specifically, we combine the convolution kernels with different dilations to extend the receptive field in the equirectangular projection. Meanwhile, we introduce an adaptive channel-wise fusion module to summarize the feature maps and get diverse attention areas in the receptive field along the channels. Due to the utilization of channel-wise attention in constructing the adaptive channel-wise fusion module, the network can capture and leverage the cross-channel contextual information efficiently. Finally, we conduct depth estimation experiments on three datasets (both virtual and real-world) and the experimental results demonstrate that our proposed ACDNet substantially outperforms the current state-of-the-art (SOTA) methods. Our codes and model parameters are accessed in https: //github. com/zcq15/ACDNet.

AAAI Conference 2022 Conference Paper

Delving into Sample Loss Curve to Embrace Noisy and Imbalanced Data

  • Shenwang Jiang
  • Jianan Li
  • Ying Wang
  • Bo Huang
  • Zhang Zhang
  • Tingfa Xu

Corrupted labels and class imbalance are commonly encountered in practically collected training data, which easily leads to over-fitting of deep neural networks (DNNs). Existing approaches alleviate these issues by adopting a sample re-weighting strategy, which is to re-weight sample by designing weighting function. However, it is only applicable for training data containing only either one type of data biases. In practice, however, biased samples with corrupted labels and of tailed classes commonly co-exist in training data. How to handle them simultaneously is a key but under-explored problem. In this paper, we find that these two types of biased samples, though have similar transient loss, have distinguishable trend and characteristics in loss curves, which could provide valuable priors for sample weight assignment. Motivated by this, we delve into the loss curves and propose a novel probe-and-allocate training strategy: In the probing stage, we train the network on the whole biased training data without intervention, and record the loss curve of each sample as an additional attribute; In the allocating stage, we feed the resulting attribute to a newly designed curve-perception network, named CurveNet, to learn to identify the bias type of each sample and assign proper weights through meta-learning adaptively. The training speed of meta learning also blocks its application. To solve it, we propose a method named skip layer meta optimization (SLMO) to accelerate training speed by skipping the bottom layers. Extensive synthetic and real experiments well validate the proposed method, which achieves state-of-the-art performance on multiple challenging benchmarks.

YNIMG Journal 2022 Journal Article

Early protein energy malnutrition impacts life-long developmental trajectories of the sources of EEG rhythmic activity

  • Jorge Bosch-Bayard
  • Fuleah Abdul Razzaq
  • Carlos Lopez-Naranjo
  • Ying Wang
  • Min Li
  • Lidice Galan-Garcia
  • Ana Calzada-Reyes
  • Trinidad Virues-Alba

Protein Energy Malnutrition (PEM) has lifelong consequences on brain development and cognitive function. We studied the lifelong developmental trajectories of resting-state EEG source activity in 66 individuals with histories of Protein Energy Malnutrition (PEM) limited to the first year of life and in 83 matched classmate controls (CON) who are all participants of the 49 years longitudinal Barbados Nutrition Study (BNS). qEEGt source z-spectra measured deviation from normative values of EEG rhythmic activity sources at 5-11 years of age and 40 years later at 45-51 years of age. The PEM group showed qEEGt abnormalities in childhood, including a developmental delay in alpha rhythm maturation and an insufficient decrease in beta activity. These profiles may be correlated with accelerated cognitive decline.

YNIMG Journal 2022 Journal Article

Harmonized-Multinational qEEG norms (HarMNqEEG)

  • Min Li
  • Ying Wang
  • Carlos Lopez-Naranjo
  • Shiang Hu
  • Ronaldo César García Reyes
  • Deirel Paz-Linares
  • Ariosky Areces-Gonzalez
  • Aini Ismafairus Abd Hamid

This paper extends frequency domain quantitative electroencephalography (qEEG) methods pursuing higher sensitivity to detect Brain Developmental Disorders. Prior qEEG work lacked integration of cross-spectral information omitting important functional connectivity descriptors. Lack of geographical diversity precluded accounting for site-specific variance, increasing qEEG nuisance variance. We ameliorate these weaknesses. (i) Create lifespan Riemannian multinational qEEG norms for cross-spectral tensors. These norms result from the HarMNqEEG project fostered by the Global Brain Consortium. We calculate the norms with data from 9 countries, 12 devices, and 14 studies, including 1564 subjects. Instead of raw data, only anonymized metadata and EEG cross-spectral tensors were shared. After visual and automatic quality control, developmental equations for the mean and standard deviation of qEEG traditional and Riemannian DPs were calculated using additive mixed-effects models. We demonstrate qEEG "batch effects" and provide methods to calculate harmonized z-scores. (ii) We also show that harmonized Riemannian norms produce z-scores with increased diagnostic accuracy predicting brain dysfunction produced by malnutrition in the first year of life and detecting COVID induced brain dysfunction. (iii) We offer open code and data to calculate different individual z-scores from the HarMNqEEG dataset. These results contribute to developing bias-free, low-cost neuroimaging technologies applicable in various health settings.

NeurIPS Conference 2022 Conference Paper

MsSVT: Mixed-scale Sparse Voxel Transformer for 3D Object Detection on Point Clouds

  • Shaocong Dong
  • Lihe Ding
  • Haiyang Wang
  • Tingfa Xu
  • Xinli Xu
  • Jie Wang
  • Ziyang Bian
  • Ying Wang

3D object detection from the LiDAR point cloud is fundamental to autonomous driving. Large-scale outdoor scenes usually feature significant variance in instance scales, thus requiring features rich in long-range and fine-grained information to support accurate detection. Recent detectors leverage the power of window-based transformers to model long-range dependencies but tend to blur out fine-grained details. To mitigate this gap, we present a novel Mixed-scale Sparse Voxel Transformer, named MsSVT, which can well capture both types of information simultaneously by the divide-and-conquer philosophy. Specifically, MsSVT explicitly divides attention heads into multiple groups, each in charge of attending to information within a particular range. All groups' output is merged to obtain the final mixed-scale features. Moreover, we provide a novel chessboard sampling strategy to reduce the computational complexity of applying a window-based transformer in 3D voxel space. To improve efficiency, we also implement the voxel sampling and gathering operations sparsely with a hash map. Endowed by the powerful capability and high efficiency of modeling mixed-scale information, our single-stage detector built on top of MsSVT surprisingly outperforms state-of-the-art two-stage detectors on Waymo. Our project page: https: //github. com/dscdyc/MsSVT.

YNIMG Journal 2022 Journal Article

No smoking signs with strong smoking symbols induce weak cravings: an fMRI and EEG study

  • Wanwan Lü
  • Qichao Wu
  • Ying Liu
  • Ying Wang
  • Zhengde Wei
  • Yu Li
  • Chuan Fan
  • An-Li Wang

No smoking signs (NSSs) that combine smoking symbols (SSs) and prohibition symbols (PSs) represent common examples of reward and prohibition competition. To evaluate how SSs within NSSs influence their effectiveness in guiding reward vs. prohibition, we studied 93 male smokers. We collected self-reported craving ratings (N=30), cue reactivity under fMRI/EEG (N=33), and smoking-behavior anticipation for paired NSSs and SSs (N=30). We found that NSS-induced cravings were negatively correlated with SS-induced cravings and PS-induced inhibition. fMRI indicated that both correlations were mediated by activation of the inferior frontal gyrus and precuneus, suggesting that the effects of SSs and PSs interact with each other. EEG revealed that the prohibition response occurs after the cigarette response, indicating that the cigarette response might be precluded by the prohibition, supporting the effect of SSs in discouraging smoking. Moreover, stronger SSs induced stronger slow positive waves and late positive potentials, and the stronger the late positive potentials, the stronger the late positive potentials. Both the amplitudes of late positive potentials and slow positive waves were positively correlated with the amplitude of N2, which was positively correlated with the attention grabbed score by the NSS. In addition, the weaker the NSS-induced craving, the greater the smoking behavior anticipation reduction, indicating the capability of NSSs to decrease smoking behavior. Our study provides empirical evidence for selecting the most effective NSSs: those combining strong SS and PS, offering insights about competition between cigarette reward and prohibition and providing neural evidence on how cigarette reward and prohibition interact.

EAAI Journal 2021 Journal Article

Feature-based evidential reasoning for probabilistic risk analysis and prediction

  • Ying Wang
  • Limao Zhang

Risk analysis plays an important role in quality control in engineering projects for the consideration of time, cost, safety, and the environment. This study proposes a feature-based evidential reasoning approach for probabilistic risk analysis and prediction, incorporating the learning process of belief degrees and estimation of the judgment quality. Firstly, classifiers are trained to estimate the probabilistic risk from sub-groups of factors. Secondly, the judgment from each classifier is evaluated according to the classifier’s performance which is characterized by the importance weight and reliability. Finally, the judgments from classifiers are fused via evidential reasoning to give the overall probabilistic risk classification result. The proposed approach displays superior performance on the dataset from Wuhan Metro with a 16% increase in precision, a 6% increase in recall, and an 8% increase in F1-score, compared to the direct model without information fusion. The fused model achieves a classification accuracy of 0. 86 on the testing samples, which is better than the direct model. Besides, the model shows good error tolerance for wrongly classified results from classifiers without information fusion. The model has an acceptable performance even when the dataset is challenging to conduct classification tasks due to high overlapping areas in the attribute space.

YNIMG Journal 2021 Journal Article

Imaging iron and neuromelanin simultaneously using a single 3D gradient echo magnetization transfer sequence: Combining neuromelanin, iron and the nigrosome-1 sign as complementary imaging biomarkers in early stage Parkinson's disease

  • Naying He
  • Kiarash Ghassaban
  • Pei Huang
  • Mojtaba Jokar
  • Ying Wang
  • Zenghui Cheng
  • Zhijia Jin
  • Yan Li

Diagnosing early stage Parkinson's disease (PD) is still a clinical challenge. Previous studies using iron, neuromelanin (NM) or the Nigrosome-1 (N1) sign in the substantia nigra (SN) by themselves have been unable to provide sufficiently high diagnostic performance for these methods to be adopted clinically. Our goal in this study was to extract the NM complex volume, iron content and volume representing the entire SN, and the N1 sign as potential complementary imaging biomarkers using a single 3D magnetization transfer contrast (MTC) gradient echo sequence and to evaluate their diagnostic performance and clinical correlations in early stage PD. A total of 40 early stage idiopathic PD subjects and 40 age- and sex-matched healthy controls (HCs) were imaged at 3T. NM boundaries (representing the SN pars compacta (SNpc) and parabrachial pigmented nucleus) and iron boundaries representing the total SN (SNpc and SN pars reticulata) were determined semi-automatically using a dynamic programming (DP) boundary detection algorithm. Receiver operating characteristic analyses were performed to evaluate the utility of these imaging biomarkers in diagnosing early stage PD. A correlation analysis was used to study the relationship between these imaging measures and the clinical scales. We also introduced the concept of NM and total iron overlap volumes to demonstrate the loss of NM relative to the iron containing SN. Furthermore, all 80 cases were evaluated for the N1 sign independently. The NM and SN volumes were lower while the iron content was higher in the SN for PD subjects compared to HCs. Interestingly, the PD subjects with bilateral loss of the N1 sign had the highest iron content. The area under the curve (AUC) values for the average of both hemispheres for single measures were: .960 for NM complex volume; .788 for total SN volume; .740 for SN iron content and. 891 for the N1 sign. Combining NM complex volume with each of the following measures through binary logistic regression led to AUC values for the averaged right and left sides of: .976 for total iron content; .969 for total SN volume, .965 for overlap volume and. 983 for the N1 sign. We found a negative correlation between SN volume and UPDRS-III (R2 =. 22, p =. 002). While the N1 sign performed well, it does not contain any information about iron content or NM quantitatively, therefore, marrying this sign with the NM and iron measures provides a better physiological explanation of what is happening when the N1 sign disappears in PD subjects. In summary, the combination of NM complex volume, SN volume, iron content and the N1 sign as derived from a single MTC sequence provides complementary information for understanding and diagnosing early stage PD.

AIIM Journal 2021 Journal Article

Tumor saliency estimation for breast ultrasound images via breast anatomy modeling

  • Fei Xu
  • Yingtao Zhang
  • H.D. Cheng
  • Boyu Zhang
  • Jianrui Ding
  • Chunping Ning
  • Ying Wang

Tumor saliency estimation aims to localize tumors by modeling the visual stimuli in medical images. However, it is a challenging task for breast ultrasound (BUS) image due to the complicated anatomic structure of the breast and poor image quality; and existing saliency estimation approaches only model the generic visual stimuli, e. g. , local and global contrast, location, and feature correlation, and achieve poor performance for tumor saliency estimation. In this paper, we propose a novel optimization model to estimate tumor saliency by utilizing breast anatomy. First, we model breast anatomy and decompose breast ultrasound image into layers using Neutro-Connectedness; then utilize the layers to generate the foreground and background maps; and finally propose a novel objective function to estimate the tumor saliency by integrating the foreground map, background map, adaptive center bias, and region-based correlation cues. The extensive experiments demonstrate that the proposed approach obtains more accurate foreground and background maps with breast anatomy; especially, for the images having large or small tumors. Meanwhile, the new objective function can handle the images without tumors. The newly proposed method achieves state-of-the-art performance comparing to eight tumor saliency estimation approaches using two BUS datasets.

YNICL Journal 2020 Journal Article

Altered resting-state dynamic functional brain networks in major depressive disorder: Findings from the REST-meta-MDD consortium

  • Yicheng Long
  • Hengyi Cao
  • Chaogan Yan
  • Xiao Chen
  • Le Li
  • Francisco Xavier Castellanos
  • Tongjian Bai
  • Qijing Bo

BACKGROUND: Major depressive disorder (MDD) is known to be characterized by altered brain functional connectivity (FC) patterns. However, whether and how the features of dynamic FC would change in patients with MDD are unclear. In this study, we aimed to characterize dynamic FC in MDD using a large multi-site sample and a novel dynamic network-based approach. METHODS: Resting-state functional magnetic resonance imaging (fMRI) data were acquired from a total of 460 MDD patients and 473 healthy controls, as a part of the REST-meta-MDD consortium. Resting-state dynamic functional brain networks were constructed for each subject by a sliding-window approach. Multiple spatio-temporal features of dynamic brain networks, including temporal variability, temporal clustering and temporal efficiency, were then compared between patients and healthy subjects at both global and local levels. RESULTS: ). Corresponding local changes in MDD were mainly found in the default-mode, sensorimotor and subcortical areas. Measures of temporal variability and characteristic temporal path length were significantly correlated with depression severity in patients (corrected p < 0.05). Moreover, the observed between-group differences were robustly present in both first-episode, drug-naïve (FEDN) and non-FEDN patients. CONCLUSIONS: Our findings suggest that excessive temporal variations of brain FC, reflecting abnormal communications between large-scale bran networks over time, may underlie the neuropathology of MDD.

AAAI Conference 2020 Conference Paper

Attention-Guide Walk Model in Heterogeneous Information Network for Multi-Style Recommendation Explanation

  • Xin Wang
  • Ying Wang
  • Yunzhi Ling

Explainable Recommendation aims at not only providing the recommended items to users, but also making users aware why these items are recommended. Too many interactive factors between users and items can be used to interpret the recommendation in a heterogeneous information network. However, these interactive factors are usually massive, implicit and noisy. The existing recommendation explanation approaches only consider the single explanation style, such as aspect-level or review-level. To address these issues, we propose a framework (MSRE) of generating the multi-style recommendation explanation with the attention-guide walk model on affiliation relations and interaction relations in the heterogeneous information network. Inspired by the attention mechanism, we determine the important contexts for recommendation explanation and learn joint representation of multi-style user-item interactions for enhancing recommendation performance. Constructing extensive experiments on three real-world datasets verifies the effectiveness of our framework on both recommendation performance and recommendation explanation.

NeurIPS Conference 2020 Conference Paper

Bayesian Bits: Unifying Quantization and Pruning

  • Mart Van Baalen
  • Christos Louizos
  • Markus Nagel
  • Rana Ali Amjad
  • Ying Wang
  • Tijmen Blankevoort
  • Max Welling

We introduce Bayesian Bits, a practical method for joint mixed precision quantization and pruning through gradient based optimization. Bayesian Bits employs a novel decomposition of the quantization operation, which sequentially considers doubling the bit width. At each new bit width, the residual error between the full precision value and the previously rounded value is quantized. We then decide whether or not to add this quantized residual error for a higher effective bit width and lower quantization noise. By starting with a power-of-two bit width, this decomposition will always produce hardware-friendly configurations, and through an additional 0-bit option, serves as a unified view of pruning and quantization. Bayesian Bits then introduces learnable stochastic gates, which collectively control the bit width of the given tensor. As a result, we can obtain low bit solutions by performing approximate inference over the gates, with prior distributions that encourage most of them to be switched off. We experimentally validate our proposed method on several benchmark datasets and show that we can learn pruned, mixed precision networks that provide a better trade-off between accuracy and efficiency than their static bit width equivalents.

YNICL Journal 2020 Journal Article

Biotypes of major depressive disorder: Neuroimaging evidence from resting-state default mode network patterns

  • Sugai Liang
  • Wei Deng
  • Xiaojing Li
  • Andrew J. Greenshaw
  • Qiang Wang
  • Mingli Li
  • Xiaohong Ma
  • Tong-Jian Bai

BACKGROUND: Major depressive disorder (MDD) is heterogeneous disorder associated with aberrant functional connectivity within the default mode network (DMN). This study focused on data-driven identification and validation of potential DMN-pattern-based MDD subtypes to parse heterogeneity of the disorder. METHODS: The sample comprised 1397 participants including 690 patients with MDD and 707 healthy controls (HC) registered from multiple sites based on the REST-meta-MDD Project in China. Baseline resting-state functional magnetic resonance imaging (rs-fMRI) data was recorded for each participant. Discriminative features were selected from DMN between patients and HC. Patient subgroups were defined by K-means and principle component analysis in the multi-site datasets and validated in an independent single-site dataset. Statistical significance of resultant clustering were confirmed. Demographic and clinical variables were compared between identified patient subgroups. RESULTS: Two MDD subgroups with differing functional connectivity profiles of DMN were identified in the multi-site datasets, and relatively stable in different validation samples. The predominant dysfunctional connectivity profiles were detected among superior frontal cortex, ventral medial prefrontal cortex, posterior cingulate cortex and precuneus, whereas one subgroup exhibited increases of connectivity (hyperDMN MDD) and another subgroup showed decreases of connectivity (hypoDMN MDD). The hyperDMN subgroup in the discovery dataset had age-related severity of depressive symptoms. Patient subgroups had comparable demographic and clinical symptom variables. CONCLUSIONS: Findings suggest the existence of two neural subtypes of MDD associated with different dysfunctional DMN connectivity patterns, which may provide useful evidence for parsing heterogeneity of depression and be valuable to inform the search for personalized treatment strategies.

AIIM Journal 2020 Journal Article

Classification of myocardial infarction based on hybrid feature extraction and artificial intelligence tools by adopting tunable-Q wavelet transform (TQWT), variational mode decomposition (VMD) and neural networks

  • Wei Zeng
  • Jian Yuan
  • Chengzhi Yuan
  • Qinghui Wang
  • Fenglin Liu
  • Ying Wang

Cardiovascular diseases (CVD) is the leading cause of human mortality and morbidity around the world, in which myocardial infarction (MI) is a silent condition that irreversibly damages the heart muscles. Currently, electrocardiogram (ECG) is widely used by the clinicians to diagnose MI patients due to its inexpensiveness and non-invasive nature. Pathological alterations provoked by MI cause slow conduction by increasing axial resistance on coupling between cells. This issue may cause abnormal patterns in the dynamics of the tip of the cardiac vector in the ECG signals. However, manual interpretation of the pathological alternations induced by MI is a time-consuming, tedious and subjective task. To overcome such disadvantages, computer-aided diagnosis techniques including signal processing and artificial intelligence tools have been developed. In this study we propose a novel technique for automatic detection of MI based on hybrid feature extraction and artificial intelligence tools. Tunable quality factor ( Q -factor) wavelet transform (TQWT), variational mode decomposition (VMD) and phase space reconstruction (PSR) are utilized to extract representative features to form cardiac vectors with synthesis of the standard 12-lead and Frank XYZ leads. They are combined with neural networks to model, identify and detect abnormal patterns in the dynamics of cardiac system caused by MI. First, 12-lead ECG signals are reduced to 3-dimensional VCG signals, which are synthesized with Frank XYZ leads to build a hybrid 4-dimensional cardiac vector. Second, this vector is decomposed into a set of frequency subbands with a number of decomposition levels by using the TQWT method. Third, VMD is employed to decompose the subband of the 4-dimensional cardiac vector into different intrinsic modes, in which the first intrinsic mode contains the majority of the cardiac vector's energy and is considered to be the predominant intrinsic mode. It is selected to construct the reference variable for analysis. Fourth, phase space of the reference variable is reconstructed, in which the properties associated with the nonlinear cardiac system dynamics are preserved. Three-dimensional (3D) PSR together with Euclidean distance (ED) has been utilized to derive features, which demonstrate significant difference in cardiac system dynamics between normal (healthy) and MI cardiac vector signals. Fifth, cardiac system dynamics can be modeled and identified using neural networks, which employ the ED of 3D PSR of the reference variable as the input features. The difference of cardiac system dynamics between healthy control and MI cardiac vector is computed and used for the detection of MI based on a bank of estimators. Finally, data sets, which include conventional 12-lead and Frank XYZ leads ECG signal fragments from 148 patients with MI and 52 healthy controls from PTB diagnostic ECG database, are used for evaluation. By using the 10-fold cross-validation style, the achieved average classification accuracy is reported to be 97. 98%. Currently, ST segment evaluation is one of the major and traditional ways for the MI detection. However, there exist weak or even undetectable ST segments in many ECG signals. Since the proposed method does not rely on the information of ST waves, it can serve as a complementary MI detection algorithm in the intensive care unit (ICU) of hospitals to assist the clinicians in confirming their diagnosis. Overall, our results verify that the proposed features may satisfactorily reflect cardiac system dynamics, and are complementary to the existing ECG features for automatic cardiac function analysis.

JBHI Journal 2020 Journal Article

Progressive Sub-Band Residual-Learning Network for MR Image Super Resolution

  • Xuetong Xue
  • Ying Wang
  • Jie Li
  • Zhicheng Jiao
  • Ziqi Ren
  • Xinbo Gao

High-resolution (HR) magnetic resonance images (MRI) provide more detailed information for clinical application. However, HR MRI is less available because of the longer scan time and lower signal-to-noise ratio. Spatial resolution is one of the key parameters of MRI. The image post-processing technique super-resolution (SR) is an alternative approach to improve the spatial resolution of MR images. Inspired by advanced deep learning based SR methods, we propose an MRI SR model named progressive sub-band residual learning SR network (PSR-SRN). The proposed model contains two parallel progressive learning streams, where one stream learns on missed high-frequency residuals by sub-band residual learning unit (ISRL) and the other focuses on reconstructing refined MR image. These two streams complement each other and enable to learn complex mappings between “Low-” and “High-” resolution MR images. Besides, we introduce brain-like mechanisms (in-depth supervision and local feedback mechanism) and progressive sub-band learning strategy to emphasize variant textures of MRI. Compared with traditional and deep learning MRI SR methods, our PSR-SRN model shows superior performance.

YNIMG Journal 2020 Journal Article

Subvoxel vascular imaging of the midbrain using USPIO-Enhanced MRI

  • Sagar Buch
  • Ying Wang
  • Min-Gyu Park
  • Pavan K. Jella
  • Jiani Hu
  • Yongsheng Chen
  • Kamran Shah
  • Yulin Ge

There is an urgent need for better detection and understanding of vascular abnormalities at the micro-level, where critical vascular nourishment and cellular metabolic changes occur. This is especially the case for structures such as the midbrain where both the feeding and draining vessels are quite small. Being able to monitor and diagnose vascular changes earlier will aid in better understanding the etiology of the disease and in the development of therapeutics. In this work, thirteen healthy volunteers were scanned with a dual echo susceptibility weighted imaging (SWI) sequence, with a resolution of 0. 22 ​× ​0. 44 ​× ​1 ​mm3 at 3T. Ultra-small superparamagnetic iron oxides (USPIO) were used to induce an increase in susceptibility in both arteries and veins. Although the increased vascular susceptibility enhances the visibility of small subvoxel vessels, the accompanying strong signal loss of the large vessels deteriorates the local tissue contrast. To overcome this problem, the SWI data were acquired at different time points during a gradual administration (final concentration ​= ​4 ​mg/kg) of the USPIO agent, Ferumoxytol, and the data was processed to combine the SWI data dynamically, in order to see through these blooming artifacts. The major vessels and their tributaries (such as the collicular artery, peduncular artery, peduncular vein and the lateral mesencephalic vein) were identified on the combined SWI data using arterio-venous maps. Dynamically combined SWI data was then compared with previous histological work to validate that this protocol was able to detect small vessels on the order of 50 ​μm–100 ​μm. A complex division-based phase unwrapping was also employed to improve the quality of quantitative susceptibility maps by reducing the artifacts due to aliased voxels at the vessel boundaries. The smallest detectable vessel size was then evaluated by revisiting numerical simulations, using estimated true susceptibilities for the basal vein and the posterior cerebral artery in the presence of Ferumoxytol. These simulations suggest that vessels as small as 50 ​μm should be visible with the maximum dose of 4 ​mg/kg.

IJCAI Conference 2019 Conference Paper

Learning Network Embedding with Community Structural Information

  • Yu Li
  • Ying Wang
  • Tingting Zhang
  • Jiawei Zhang
  • Yi Chang

Network embedding is an effective approach to learn the low-dimensional representations of vertices in networks, aiming to capture and preserve the structure and inherent properties of networks. The vast majority of existing network embedding methods exclusively focus on vertex proximity of networks, while ignoring the network internal community structure. However, the homophily principle indicates that vertices within the same community are more similar to each other than those from different communities, thus vertices within the same community should have similar vertex representations. Motivated by this, we propose a novel network embedding framework NECS to learn the Network Embedding with Community Structural information, which preserves the high-order proximity and incorporates the community structure in vertex representation learning. We formulate the problem into a principled optimization framework and provide an effective alternating algorithm to solve it. Extensive experimental results on several benchmark network datasets demonstrate the effectiveness of the proposed framework in various network analysis tasks including network reconstruction, link prediction and vertex classification.

YNIMG Journal 2018 Journal Article

Chronic nicotine exposure impairs uncertainty modulation on reinforcement learning in anterior cingulate cortex and serotonin system

  • Zhengde Wei
  • Long Han
  • Xiuying Zhong
  • Ying Liu
  • Rujing Zha
  • Ying Wang
  • Li-Zhuang Yang
  • Junjie Bu

Deficits in the computational processes of reinforcement learning have been suggested to underlie addiction. Additionally, environmental uncertainty, which is encoded in the anterior cingulate cortex (ACC), modulates reward prediction errors (RPEs) during reinforcement learning and exacerbates addiction. The present study tested whether and how the ACC would have an essential role in drug addiction by failing to use uncertainty to modulate the RPEs during reinforcement learning. In Experiment I, we found that the ACC/medial prefrontal cortex (MPFC) did not modulate RPE learning according to uncertainty in smokers. The effect of uncertainty × RPE in the ACC/MPFC was correlated with the learning rate of RPEs and the duration of nicotine use. Experiment II demonstrated that serotonin, but not dopamine, receptor mRNA expression significantly decreased in the ACC of the nicotine exposed compared to the control rats. Furthermore, there was a positive correlation between learning rate and serotonin receptor mRNA expression in the ACC. Therefore, all present results suggest that impairments in uncertainty modulation in the ACC disrupt reinforcement learning processes in chronic nicotine users and contribute to maladaptive decision-making. These findings support interventions for pathological decision-making in drug addiction that strongly focus on the serotonin system in ACC.

YNICL Journal 2018 Journal Article

Common and distinct abnormal frontal-limbic system structural and functional patterns in patients with major depression and bipolar disorder

  • Lixiang Chen
  • Ying Wang
  • Chen Niu
  • Shuming Zhong
  • Huiqing Hu
  • Ping Chen
  • Shufei Zhang
  • Guanmao Chen

Major depressive disorder (MDD) and bipolar disorder (BD) are common severe affective diseases. Although previous neuroimaging studies have investigated brain abnormalities in MDD or BD, the structural and functional differences between these two disorders remain unclear. In this study, we adopted a multimodal approach, combining voxel-based morphometry (VBM) and functional connectivity (FC), to study the common and distinct structural and functional alterations in unmedicated MDD and BD patients. The VBM analysis revealed that both the MDD and BD patients showed decreased gray matter volume (GMV) in the left anterior cingulate cortex (ACC_L) and right hippocampus (HIP_R) compared with the healthy controls, and the MDD patients showed decreased GMV in the left superior frontal gyrus (SFG_L) and ACC_L compared with the BD patients. Furthermore, we took these clusters as seed regions to analyze the abnormal resting-state functional connectivity (RSFC) in the patients. We found that both the MDD and BD groups had decreased RSFC between the ACC_L and the left orbitofrontal cortex (OFC_L) and that the MDD group had decreased RSFC between the SFG_L and the HIP_L, compared with the healthy controls. Our results revealed that the MDD and BD patients were more similar than different in GMV and RSFC. These findings indicate that investigating the frontal-limbic system could be useful for understanding the underlying mechanisms of these two disorders.

YNICL Journal 2018 Journal Article

Disruption of superficial white matter in the emotion regulation network in bipolar disorder

  • Shufei Zhang
  • Ying Wang
  • Feng Deng
  • Shuming Zhong
  • Lixiang Chen
  • Xiaomei Luo
  • Shaojuan Qiu
  • Ping Chen

Bipolar disorder (BD) is characterized by emotion dysregulation and involves changes in the gray matter (GM) and white matter (WM). Although previous diffusion tensor imaging (DTI) studies reported changes in the diffusion properties of the deep WM (DWM) in BD patients, the diffusion properties of the superficial WM (SWM) are rarely investigated. In this study, we tried to determine whether the diffusion parameters of the SWM were altered in BD patients compared to controls and whether the changes were associated with the disrupted emotion regulation of the BD patients. We collected DTI data from 37 BD patients and 42 gender- and age-matched healthy controls (HC). Using probabilistic tractography, we defined a population-based SWM mask based on all the subjects. After performing the tract-based spatial statistical (TBSS) analyses, we identified the SWM areas in which the BD patients differed from the controls. This study showed significantly reduced fractional anisotropy in the SWM (FA SWM) in the BD patients compared to the HC in the bilateral dorsolateral prefrontal cortex (dlPFC), ventrolateral prefrontal cortex (vlPFC), medial prefrontal cortex (mPFC), and the left parietal cortex. Moreover, compared to the controls, the BD patients showed significantly increased mean diffusivity (MD SWM) and radial diffusivity (RD SWM) in the SWM in the right frontal cortex. This study presents altered cortico-cortical connections proximal to the regions related to the emotion dysregulation of BD patients, which indicated that the SWM may serve as the brain's structural basis underlying the disrupted emotion regulation of BD patients. The disrupted FA SWM in the parietal cortex may indicate that the emotion dysregulation in BD patients is related to the cognitive control network.

YNIMG Journal 2017 Journal Article

Neural substrates of updating the prediction through prediction error during decision making

  • Ying Wang
  • Ning Ma
  • Xiaosong He
  • Nan Li
  • Zhengde Wei
  • Lizhuang Yang
  • Rujing Zha
  • Long Han

Learning of prediction error (PE), including reward PE and risk PE, is crucial for updating the prediction in reinforcement learning (RL). Neurobiological and computational models of RL have reported extensive brain activations related to PE. However, the occurrence of PE does not necessarily predict updating the prediction, e. g. , in a probability-known event. Therefore, the brain regions specifically engaged in updating the prediction remain unknown. Here, we conducted two functional magnetic resonance imaging (fMRI) experiments, the probability-unknown Iowa Gambling Task (IGT) and the probability-known risk decision task (RDT). Behavioral analyses confirmed that PEs occurred in both tasks but were only used for updating the prediction in the IGT. By comparing PE-related brain activations between the two tasks, we found that the rostral anterior cingulate cortex/ventral medial prefrontal cortex (rACC/vmPFC) and the posterior cingulate cortex (PCC) activated only during the IGT and were related to both reward and risk PE. Moreover, the responses in the rACC/vmPFC and the PCC were modulated by uncertainty and were associated with reward prediction-related brain regions. Electric brain stimulation over these regions lowered the performance in the IGT but not in the RDT. Our findings of a distributed neural circuit of PE processing suggest that the rACC/vmPFC and the PCC play a key role in updating the prediction through PE processing during decision making.

AAAI Conference 2015 Conference Paper

10,000+ Times Accelerated Robust Subset Selection

  • Feiyun Zhu
  • Bin Fan
  • Xinliang Zhu
  • Ying Wang
  • Shiming Xiang
  • Chunhong Pan

Subset selection from massive data with noised information is increasingly popular for various applications. This problem is still highly challenging as current methods are generally slow in speed and sensitive to outliers. To address the above two issues, we propose an accelerated robust subset selection (ARSS) method. Specifically in the subset selection area, this is the first attempt to employ the p (0 < p ≤ 1)-norm based measure for the representation loss, preventing large errors from dominating our objective. As a result, the robustness against outlier elements is greatly enhanced. Actually, data size is generally much larger than feature length, i. e. N L. Based on this observation, we propose a speedup solver (via ALM and equivalent derivations) to highly reduce the computational cost, theoretically from O N4 to O N2 L. Extensive experiments on ten benchmark datasets verify that our method not only outperforms state of the art methods, but also runs 10, 000+ times faster than the most related method.

AAAI Conference 2015 Conference Paper

Exploring Social Context for Topic Identification in Short and Noisy Texts

  • Xin Wang
  • Ying Wang
  • Wanli Zuo
  • Guoyong Cai

With the pervasion of social media, topic identification in short texts attracts increasing attention in recent years. However, in nature the texts of social media are short and noisy, and the structures are sparse and dynamic, resulting in difficulty to identify topic categories exactly from online social media. Inspired by social science findings that preference consistency and social contagion are observed in social media, we investigate topic identification in short and noisy texts by exploring social context from the perspective of social sciences. In particular, we present a mathematical optimization formulation that incorporates the preference consistency and social contagion theories into a supervised learning method, and conduct feature selection to tackle short and noisy texts in social media, which result in a Sociological framework for Topic Identification (STI). Experimental results on real-world datasets from Twitter and Citation Network demonstrate the effectiveness of the proposed framework. Further experiments are conducted to understand the importance of social context in topic identification.

AAAI Conference 2015 Conference Paper

Modeling Status Theory in Trust Prediction

  • Ying Wang
  • Xin Wang
  • Jiliang Tang
  • Wanli Zuo
  • Guoyong Cai

With the pervasion of social media, trust has been playing more of an important role in helping online users collect reliable information. In reality, user-specified trust relations are often very sparse; hence, inferring unknown trust relations has attracted increasing attention in recent years. Social status is one of the most important concepts in trust, and status theory is developed to help us understand the important role of social status in the formation of trust relations. In this paper, we investigate how to exploit social status in trust prediction by modeling status theory. We first verify status theory in trust relations, then provide a principled way to model it mathematically, and propose a novel framework sTrust which incorporates status theory for trust prediction. Experimental results on real-world datasets demonstrate the effectiveness of the proposed framework. Further experiments are conducted to understand the importance of status theory in trust prediction.

EAAI Journal 2011 Journal Article

Chaotic differential evolution methods for dynamic economic dispatch with valve-point effects

  • Youlin Lu
  • Jianzhong Zhou
  • Hui Qin
  • Ying Wang
  • Yongchuan Zhang

The dynamic economic dispatch (DED), with the consideration of valve-point effects, is a complicated non-linear constrained optimization problem with non-smooth and non-convex characteristics. In this paper, three chaotic differential evolution (CDE) methods are proposed based on the Tent equation to solve DED problem with valve-point effects. In the proposed methods, chaotic sequences are applied to obtain the dynamic parameter settings in DE. Meanwhile, a chaotic local search (CLS) operation for solving DED problem is designed to help DE avoiding premature convergence effectively. Finally, in order to handle the complicated constraints with efficiency, new heuristic constraints handling methods and feasibility based selection strategy are embedded into the proposed CDE methods. The feasibility and effectiveness of the proposed CDE methods are demonstrated for two test systems. The simulation results reveal that, compared with DE and those other methods reported in literatures recently, the proposed CDE methods are capable of obtaining better quality solutions with higher efficiency.

YNIMG Journal 2010 Journal Article

High-dimensional pattern regression using machine learning: From medical images to continuous clinical variables

  • Ying Wang
  • Yong Fan
  • Priyanka Bhatt
  • Christos Davatzikos

This paper presents a general methodology for high-dimensional pattern regression on medical images via machine learning techniques. Compared with pattern classification studies, pattern regression considers the problem of estimating continuous rather than categorical variables, and can be more challenging. It is also clinically important, since it can be used to estimate disease stage and predict clinical progression from images. In this work, adaptive regional feature extraction approach is used along with other common feature extraction methods, and feature selection technique is adopted to produce a small number of discriminative features for optimal regression performance. Then the Relevance Vector Machine (RVM) is used to build regression models based on selected features. To get stable regression models from limited training samples, a bagging framework is adopted to build ensemble basis regressors derived from multiple bootstrap training samples, and thus to alleviate the effects of outliers as well as facilitate the optimal model parameter selection. Finally, this regression scheme is tested on simulated data and real data via cross-validation. Experimental results demonstrate that this regression scheme achieves higher estimation accuracy and better generalizing ability than Support Vector Regression (SVR).

EAAI Journal 2008 Journal Article

A machine-learning approach to multi-robot coordination

  • Ying Wang
  • Clarence W. de Silva

This paper presents a machine-learning approach to the multi-robot coordination problem in an unknown dynamic environment. A multi-robot object transportation task is employed as the platform to assess and validate this approach. Specifically, a flexible two-layer multi-agent architecture is developed to implement multi-robot coordination. In this architecture, four software agents form a high-level coordination subsystem while two heterogeneous robots constitute the low-level control subsystem. Two types of machine learning—reinforcement learning (RL) and genetic algorithms (GAs)—are integrated to make decisions when the robots cooperatively transport an object to a goal location while avoiding obstacles. A probabilistic arbitrator is used to determine the winning output between the RL and GA algorithms. In particular, a modified RL algorithm called the sequential Q-learning algorithm is developed to deal with the issues of behavior conflict that arise in multi-robot cooperative transportation tasks. The learning-based high-level coordination subsystem sends commands to the low-level control subsystem, which is implemented with a hybrid force/position control scheme. Simulation and experimental results are presented to demonstrate the effectiveness and adaptivity of the developed approach.

IROS Conference 2005 Conference Paper

Robust internal model control with feedforward controller for a high-speed motion platform

  • Ying Wang
  • Zhenhua Xiong 0001
  • Han Ding 0001

A new control method based on a combination of robust control and internal model control has been proposed. This control system includes internal model controller for velocity loop, robust controller for position loop, and a feedforward controller. The internal model controller is designed to suppress disturbance. Stability robustness of the closed loop is provided by the robust controller. The zero phase error tracking controller is adopted to act as a feedforward controller to further improve the tracking performance. The theoretical analysis shows the validity of the proposed control scheme. Furthermore, simulations and experimental results are presented to demonstrate performance improvement of the proposed control structure.

ICRA Conference 2004 Conference Paper

Nonlinear Friction Compensation and Disturbance Observer for a High-speed Motion Platform

  • Ying Wang
  • Zhenhua Xiong 0001
  • Han Ding 0001
  • Xiangyang Zhu

Nonlinear friction and external disturbances affect the positioning accuracy of high-speed motion systems, especially impelled by linear motor. Thus, how to eliminate theses disturbance should be considered when designing a robust controller. This paper presents a controller, which includes three parts: a proportional-plus-derivative (PD) feedback controller, a friction compensator, and a disturbance observer. The friction compensator is based on LuGre model and it compensates for nonlinear friction. The disturbance observer is used to eliminate the friction compensation error and other external disturbances. Experimental results show that the controller gives high positioning accuracy and more robust performance in the presence of disturbances.

IROS Conference 2000 Conference Paper

Developing rigid motion constraints for the registration of free-form shapes

  • Yonghuai Liu
  • Marcos A. Rodrigues 0001
  • Ying Wang

We propose a novel method to deal with sphere ambiguity, occlusion, appearance and disappearance of points in image registration. We have developed a number of rigid motion constraints through analysis of geometrical properties of reflected correspondence vectors synthesised into a single coordinate frame. The properties are used as further constraints to eliminate false matches obtained by the iterative closest point criterion. A number of experiments based on both synthetic data and real images demonstrate that the proposed method is accurate, robust, and efficient for the registration of free-form shapes.

v2026.09.13