Arrow Research search

Author name cluster

Fei Gao

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

40 papers
2 author rows

Possible papers

40

AAAI Conference 2026 Conference Paper

CO²IF: Language-Bridging Hyperspectral-Multispectral Image Fusion with Coordinated and Cross-modal Optimal Transport

  • Mingjin Zhang
  • Zhongkai Yang
  • Fei Gao

Due to the difficulties of directly obtaining high-resolution hyperspectral images (HR-HSI), the fusion of low-resolution hyperspectral images (LR-HSI) and high-resolution multispectral images (HR-MSI) has emerged as an effective approach. While existing methods leverage image-level priors from HR-MSI, they often lack explicit semantic guidance for precise detail reconstruction. Recognizing that textual scene descriptions encapsulate valuable object attributes and contextual information, we introduce the first Language-Bridging framework for Hyperspectral and Multispectral image fusion (CO²IF). CO²IF leverages language semantics as prior knowledge to explicitly guide the reconstruction process. To bridge the modality gap between textual descriptions and high-dimensional hyperspectral data, we design a Cross-modal Optimal Transport (COT) module. COT establishes precise semantic correspondences between language features and the visual cues of individual spectral bands. Building upon this semantic alignment, we develop a Multimodal Coordinated State Space Model (CoMamba). CoMamba effectively integrates the language-derived priors with spatial information from HR-MSI and spectral information from LR-HSI. This language-guided reconstruction significantly enhances the extraction of crucial spatial-spectral details, leading to superior fidelity in the generated HR-HSI. In addition, this paper adds text descriptions for three widely used datasets. Both qualitative and quantitative experimental results on the public datasets confirm the superiority of the proposed method compared to the SOTA methods.

EAAI Journal 2026 Journal Article

Temporal-spatial parallel multiscale network with sparse three-channel mixed attention for wearable sensor-based human activity recognition

  • Renzhuo Wang
  • Hongji Xu
  • Yiran Li
  • Yonghui Yu
  • Yupeng Duan
  • Zhikai Xu
  • Wentao Ai
  • Xinya Li

Human activity recognition (HAR) based on wearable sensors using deep learning (DL) models has garnered significant attention in recent years. However, existing models encounter several challenges in fully exploiting the information from multi-source sensor positions. Notably, they often fail to provide adequate interpretability in terms of both single-channel attention extraction and inter-channel attention mixing. This paper proposes a novel temporal-spatial parallel multiscale network with sparse three-channel mixed attention (TSPM-STCMA) designed to independently and in parallel learn temporal-spatial features from both single-channel and inter-channel relationships derived from multi-source sensors at various body positions. Specifically, the multiscale dynamic convolution with sparse three-channel mixed attention (MDC-STCMA) module is designed to enhance the interpretability of both single-channel attention and inter-channel attention from the same sensor position. Furthermore, the MDC-STCMA module integrates three components based on an attention mechanism. The three-channel dynamic convolution based on improved squeeze-and-excitation (TCDC-ISE) enables a single channel to generate distinct attention responses. The three-channel mixed attention (TCMA) extracts inter-channel correlations. The sparse attention (SA) mitigates overfitting by retraining features with low contributions. Compared with previous models, the TSPM-STCMA achieves superior recognition accuracies of 98. 72 %, 96. 83 %, and 98. 40 %, on three publicly available datasets, i. e. , Physical Activity Monitoring for Aging People (PAMAP2), OPPORTUNITY activity recognition (OPPORTUNITY), and University of California Irvine HAR (UCI-HAR), respectively, while requiring significantly fewer parameters.

EAAI Journal 2025 Journal Article

A novel unmanned aerial vehicles task allocation approach based on the intuitionistic fuzzy multi-criteria bilateral matching-based decision-making method

  • Fei Gao

With the rapid advancement of unmanned aerial vehicles (UAVs), deploying UAV swarms for diverse missions has become standard practice. A critical challenge in these operations is task allocation, that is, assigning the right UAV to the right task. However, existing methods often overlook the concept of bilateral matching, where the mutual suitability between UAVs and tasks is considered under uncertainty. To this end, this study proposes a novel decision-making framework for UAV task allocation based on the intuitionistic fuzzy multi-criteria bilateral matching method designed to enhance mission outcomes through balanced and fair assignments. The proposed method employs intuitionistic fuzzy sets (IFSs) to effectively model the uncertainty and ambiguity inherent in evaluating UAVs and tasks across multiple criteria. To ensure objectivity, the criteria importance through intercriteria correlation (CRITIC) method is extended to the IFS environment for criteria weighting. Suitability and matching degrees are then formulated to comprehensively assess the compatibility between each UAV and task. Finally, a non-linear optimization model integrates these degrees to determine the optimal allocation scheme. A practical case study validates the approach, demonstrating its reliability and effectiveness in generating rational and robust allocation outcomes. In conclusion, this study provides a reliable strategy for UAV swarm task allocation in complex and uncertain environments.

AAAI Conference 2025 Conference Paper

Autoregressive Sequence Modeling for 3D Medical Image Representation

  • Siwen Wang
  • Churan Wang
  • Fei Gao
  • Lixian Su
  • Fandong Zhang
  • Yizhou Wang
  • Yizhou Yu

Three-dimensional (3D) medical images, such as Computed Tomography (CT) and Magnetic Resonance Imaging (MRI), are essential for clinical applications. However, the need for diverse and comprehensive representations is particularly pronounced when considering the variability across different organs, diagnostic tasks, and imaging modalities. How to effectively interpret the intricate contextual information and extract meaningful insights from these images remains an open challenge to the community. While current self-supervised learning methods have shown potential, they often consider an image as a whole thereby overlooking the extensive, complex relationships among local regions from one or multiple images. In this work, we introduce a pioneering method for learning 3D medical image representations through an autoregressive pre-training framework. Our approach sequences various 3D medical images based on spatial, contrast, and semantic correlations, treating them as interconnected visual tokens within a token sequence. By employing an autoregressive sequence modeling task, we predict the next visual token in the sequence, which allows our model to deeply understand and integrate the contextual information inherent in 3D medical images. Additionally, we implement a random startup strategy to avoid overestimating token relationships and to enhance the robustness of learning. The effectiveness of our approach is demonstrated by the superior performance over others on nine downstream tasks in public datasets.

IROS Conference 2025 Conference Paper

Building Hybrid Omnidirectional Visual-Lidar Map for Visual-Only Localization

  • Jingyang Huang
  • Hao Wei
  • Changze Li
  • Tong Qin
  • Fei Gao
  • Ming Yang

Recently, there has been growing interest in using low-cost sensor combinations, such as cameras and IMUs, to achieve accurate localization within pre-built pointcloud maps. In this paper, we propose a novel hybrid visual-Lidar mapping and visual-only re-localization framework, specifically designed for UAVs with limited computational resources operating in challenging environments. Keyframes function as a bridge in our system, associating images with pointcloud to facilitate efficient and accurate pose estimation. Besides, our system creates omnidirectional keyframes at the mapping stage, enabling effective re-localization from any orientation, which enhance the robustness and practicability of our system. Experiments show that the proposed algorithm achieves high localization accuracy on pre-built maps and is capable of running in real-time on UAVs for autonomous navigation tasks. The source code will be made publicly available soon.

NeurIPS Conference 2025 Conference Paper

DiffE2E: Rethinking End-to-End Driving with a Hybrid Diffusion-Regression-Classification Policy

  • Rui Zhao
  • Yuze Fan
  • Ziguo Chen
  • Fei Gao
  • Zhenhai Gao

End-to-end learning has emerged as a transformative paradigm for autonomous driving. However, the inherently multimodal nature of driving behaviors remains a fundamental challenge to robust deployment. We propose DiffE2E, a diffusion-based end-to-end autonomous driving framework. The architecture first performs multi-scale alignment of perception features from multiple sensors via a hierarchical bidirectional cross-attention mechanism. Subsequently, we design a hybrid diffusion-regression-classification decoder based on the Transformer architecture, adopting a collaborative training paradigm to seamlessly fuse the strengths of diffusion and explicit strategies. DiffE2E conducts structured modeling in the latent space: diffusion captures the multimodal distribution of future trajectories, while regression and classification act as explicit strategies to precisely model key control variables such as velocity, enhancing both the precision and controllability of the model. A global condition integration module further enables deep fusion of perception features with high-level goals, significantly improving the quality of trajectory generation. The subsequent cross-attention mechanism facilitates efficient interaction between integrated features and hybrid latent variables, promoting joint optimization of diffusion and explicit strategies for structured output generation and thereby yielding more robust control. Experimental results demonstrate that DiffE2E achieves state-of-the-art performance on both CARLA closed-loop benchmarks and NAVSIM evaluations. The proposed unified framework that integrates diffusion and explicit strategies provides a generalizable paradigm for hybrid action representation and shows substantial potential for extension to broader domains, including embodied intelligence.

YNIMG Journal 2025 Journal Article

Individualized brain radiomics-based network tracks distinct subtypes and abnormal patterns in prodromal Parkinson's disease

  • Lin Hua
  • Canpeng Huang
  • Xinglin Zeng
  • Fei Gao
  • Zhen Yuan

Individuals in the prodromal phase of Parkinson's disease (PD) exhibit significant heterogeneity and can be divided into distinct subtypes based on clinical symptoms, pathological mechanisms, and brain network patterns. However, little has been done regarding the valid subtyping of prodromal PD, which hinders the early diagnosis of PD. Therefore, we aimed to identify the subtypes of prodromal PD using the brain radiomics-based network and examine the unique patterns linked to the clinical presentations of each subtype. Individualized brain radiomics-based network was constructed for normal controls (NC; N = 110), prodromal PD patients (N = 262), and PD patients (N = 108). A data-driven clustering approach using the radiomics-based network was carried out to cluster prodromal PD patients into higher-/lower-risk subtypes. Then, the dissociated patterns of clinical manifestations, anatomical structure alterations, and gene expression between these two subtypes were evaluated. Clustering findings indicated that one prodromal PD subtype closely resembled the pattern of NCs (N-P; N = 159), while the other was similar to the pattern of PD (P-P; N = 103). Significant differences were observed between the subtypes in terms of multiple clinical measurements, neuroimaging for morphological changes, and gene enrichment for synaptic transmission. Identification of prodromal PD subtypes based on brain connectomes and a full understanding of heterogeneity at this phase could inform early and accurate PD diagnosis and effective neuroprotective interventions.

AAAI Conference 2025 Conference Paper

IRMamba: Pixel Difference Mamba with Layer Restoration for Infrared Small Target Detection

  • Mingjin Zhang
  • Xiaolong Li
  • Fei Gao
  • Jie Guo

Infrared small target detection (IRSTD) focuses on identifying small targets in infrared images. Despite advancements with deep learning, challenges persist due to the IR long-range imaging mechanism, where targets are small, dim, and easily lost in noise and background clutter. Current deep learning methods struggle to suppress noise and background interference while preserving fine details, leading to missed detections and false alarms. To address these issues, we propose IRMamba, an encoder-decoder architecture featuring Pixel Difference Mamba (PDMamba) and a Layer Restoration Module (LRM). Specifically, PDMamba integrates the intensity and directional information of pixel differences between scanning positions and their central neighborhoods into the state equation of the state space model (SSM). This enhances target detail representation and suppresses background interference by capturing local 2D dependencies from a global perspective. In addition, LRM incorporates the double-depth image prior into the iterative convergence algorithm, and utilizes the inter-layer interrelationships to gradually reverse the separation of the target layer, achieving noise suppression and refined reconstruction of the image mask. Experiments conducted on multiple public datasets, including NUAA-SIRST, NUDT-SIRST, and IRSTD-1K, demonstrate the significant advantages of IRMamba over SOTA methods.

IJCAI Conference 2025 Conference Paper

Logic Distillation: Learning from Code Function by Function for Decision-making Tasks

  • Dong Chen
  • Shilin Zhang
  • Fei Gao
  • Yueting Zhuang
  • Siliang Tang
  • Qidong Liu
  • Mingliang Xu

Large language models (LLMs) have garnered increasing attention owing to their powerful comprehension and generation capabilities. Generally, larger LLMs (L-LLMs) that require paid interfaces exhibit significantly superior performance compared to smaller LLMs (S-LLMs) that can be deployed on a variety of devices. Knowledge distillation (KD) aims to empower S-LLMs with the capabilities of L-LLMs, while S-LLMs merely mimic the outputs of L-LLMs, failing to get the powerful decision-making capability for new situations. Consequently, S-LLMs are helpless when it comes to continuous decision-making tasks that require logical reasoning. To tackle the identified challenges, we propose a novel framework called Logic Distillation (LD). Initially, LD employs L-LLMs to instantiate complex instructions into discrete functions and illustrates their usage to establish a function base. Subsequently, LD fine-tunes S-LLMs based on the function base to learn the logic employed by L-LLMs in decision-making. During testing, S-LLMs will yield decision-making outcomes, function by function, based on current states. Experiments demonstrate that with the assistance of LD, S-LLMs can achieve outstanding results in continuous decision-making tasks, comparable to, or even surpassing, those of L-LLMs. The code and data for the proposed method are provided for research purposes https: //github. com/Anfeather/Logic-Distillation.

AAAI Conference 2025 Conference Paper

MOCID: Motion Context and Displacement Information Learning for Moving Infrared Small Target Detection

  • Mingjin Zhang
  • Yuanjun Ouyang
  • Fei Gao
  • Jie Guo
  • Qiming Zhang
  • Jing Zhang

In the field of Moving Infrared Small Target Detection (MIRSTD), current methods typically use sequential modeling with two individual modules for spatial and temporal processing. However, such a modeling strategy lacks clear guidance on the motion and displacement difference between moving targets and background noise, thereby limiting the feature discriminability and resulting in error-prone target localization. This paper addresses this issue from clip and frame levels and proposes a novel architecture MOCID for MIRSTD. For clip-level feature fusion, we design a spatio-temporal backbone consisting of several proposed Fourier-inspired Spatio-temporal Attention (FISTA) layers. Each FISTA layer sequentially processes the features from spatial and temporal views to capture clip-level temporal motion context, where Fourier Transformation and Inverse Fourier Transformation are employed for each view. This context is then embedded into dynamic convolutional kernels for subsequent spatial feature extraction, thereby enabling clear motion difference guidance and generating comprehensive features. For frame-level feature fusion, we design a Displacement-aware Mamba Module (DAM) to capture detailed frame-to-frame displacement information. DAM utilizes an innovative Temporal Interpolation and Displacement-aware Scan technique to perform spatio-temporal difference-aware displacement modeling, introducing elaborate temporal indicators into feature extraction. Combining the above improvements, our model captures comprehensive motion and displacement contexts, significantly improving the detection of the small target. Extensive experiments demonstrate that MOCID achieves state-of-the-art detection accuracy on popular IRDST and DAUB datasets. Furthermore, MOCID offers a superior balance between throughput and performance compared to other methods. The code for this work will be made publicly available.

IJCAI Conference 2025 Conference Paper

Multimodal Prior Learning with Double Constraint Alignment for Snapshot Spectral Compressive Imaging

  • Mingjin Zhang
  • Longyi Li
  • Fei Gao
  • Qiming Zhang
  • Jie Guo

The objective of snapshot spectral compressive imaging reconstruction is to recover the 3D hyperspectral image (HSI) from a 2D measurement. Existing methods either focus on network architecture design or simply introduce image-level prior to the model. However, these methods lack guiding information for accurate reconstruction. Recognizing that textual description contain rich semantic information that can significantly enhance details, this paper introduces a novel framework, CAMM, which integrates text information into the model to improve the performance. The framework comprises two key components: Fine-grained Alignment Module (FAM) and Multimodal Fusion Mamba (MFM). Specifically, FAM is used to reduce the knowledge gap between the RGB domain obtained by the pre-trained vision-language model and the HSI domain. Through the double constraints of distribution similarity and entropy, the adaptive alignment of different complexity features is realized, which makes the encoded features more accurate. MFM aims to identify the guiding effect of RGB features and text features on HSI in space and channel dimensions. Instead of fusing features directly, it integrates prior at image-level and text-level prior into Mamba's state-space equation, so that each scanning step can be accurately guided. This kind of positive feedback adjustment ensures the authenticity of the guiding information. To our knowledge, this is the first text-guided model for compressive spectral imaging. Extensive experimental results the public datasets demonstrate the superior performance of CAMM, validating the effectiveness of our proposed method.

AAAI Conference 2025 Conference Paper

QuARF: Quality-Adaptive Receptive Fields for Degraded Image Perception

  • Fei Gao
  • Ying Zhou
  • Ziyun Li
  • Wenwang Han
  • Jiaqi Shi
  • Maoying Qiao
  • Jinlan Xu
  • Nannan Wang

Advanced Deep Neural Networks (DNNs) perform well for high-quality images, but their performance dramatically decreases for degraded images. Data augmentation is commonly used to alleviate this problem, but using too much perturbed data might seriously decrease the performance on pristine images. To tackle this challenge, we take our cue from the assumption of spatial coincidence in human visual perception, i.e. multiscale and varying receptive fields are required for understanding pristine and degraded images. Correspondingly, we propose a novel plug-and-play network architecture, dubbed Quality-Adaptive Receptive Fields (QuARF), to automatically select the optimal receptive fields based on the quality of the input image. To this end, we first design a multi-kernel convolutional block, which comprises multiscale continuous receptive fields. Afterward, we design a quality-adaptive routing network to predict the significance of each kernel, based on the quality features extracted from the input image. In this way, QuARF automatically selects the optimal inference route for each image. To further boost efficiency and effectiveness, the input feature map is split into multiple groups, with each group independently learning its quality-adaptive routing parameters. We apply QuARF to a variety of DNNs and conduct experiments in both discriminative and generation tasks, including semantic segmentation, image translation, and restoration. Thorough experimental results show that QuARF significantly and robustly improves the performance for degraded images, and outperforms data augmentation in most cases.

YNIMG Journal 2025 Journal Article

Right inferior frontal cortex and preSMA in response inhibition: An investigation based on PTC model

  • Lili Wu
  • Mengjie Jiang
  • Min Zhao
  • Xin Hu
  • Jing Wang
  • Kaihua Zhang
  • Ke Jia
  • Fuxin Ren

Response inhibition is an essential component of cognitive function. A large body of literature has used neuroimaging data to uncover the neural architecture that regulates inhibitory control in general and movement cancelation. The presupplementary motor area (preSMA) and the right inferior frontal cortex (rIFC) are the key nodes in the inhibitory control network. However, how these two regions contribute to response inhibition remains controversial. Based on the Pause-then-Cancel Model (PTC), this study employed functional magnetic resonance imaging (fMRI) to investigate the functional specificity of two regions in the stopping process. The Go/No-Go task (GNGT) and the Stop Signal Task (SST) were administered to the same group of participants. We used the GNGT to dissociate the pause process and both the GNGT and the SST to investigate the inhibition mechanism. Imaging data revealed that response inhibition produced by both tasks activated the preSMA and rIFC. Furthermore, an across-participants analysis showed that increased activation in the rIFC was associated with a delay in the go response in the GNGT. In contrast, increased activation in the preSMA was associated with good inhibition efficiency via the striatum in both GNGT and SST. These behavioral and imaging findings support the PTC model of the role of rIFC and preSMA, that the former is involved in a pause process to delay motor responses, whereas the preSMA is involved in the stopping of motor responses.

AAAI Conference 2025 Conference Paper

Semi-supervised Infrared Small Target Detection with Thermodynamic-Inspired Uneven Perturbation and Confidence Adaptation

  • Mingjin Zhang
  • Wenteng Shang
  • Fei Gao
  • Qiming Zhang
  • FengQin Lu
  • Jing Zhang

Single-frame Infrared Small Target (SIRST) detection has made significant advancements, but it still faces challenges due to limited labeled data and the foreground-background class imbalance. To address these issues, we introduce a novel Semi-Supervised SIRST Detection (S^3D) pipeline in this paper. First, drawing inspiration from thermodynamics, we propose augmenting infrared images using both chromatically and spatially uneven perturbations. This dual-stream perturbation enhances the diversity and balance of infrared samples, contributing to the robustness of detection models. Additionally, we develop a confidence-adaptive matching method to maintain weighted consistency among perturbed unlabeled samples. Second, to tackle class imbalance in labeled data, we compel the model to generate discriminative predictions for challenging, misclassified examples while down-weighting well-classified examples. We achieve this by modifying the standard cross-entropy loss to squeeze the detector and truncating the loss on well-classified examples. Our innovative Truncated Squeeze (TS) loss focuses on learning discriminative representations for difficult cases and prevents over-optimization for simpler ones. To assess the effectiveness of the perturbation techniques and loss functions, we apply them to various SIRST detectors and conduct comprehensive experiments on two benchmark datasets. Notably, our proposed methods consistently and significantly improve accuracy. Remarkably, our approach achieves over 98% performance of the state-of-the-art fully-supervised method using only 1/8 of the labeled samples.

YNIMG Journal 2025 Journal Article

The self-awareness brain network: Construction, characterization, and alterations in schizophrenia and major depressive disorder

  • Xiaoluan Xia
  • Fei Gao
  • Shiyang Xu
  • Kaixin Li
  • Qingxia Zhu
  • Yuwen He
  • Xinglin Zeng
  • Lin Hua

Self-awareness (SA) research is crucial for understanding cognition, social behavior, mental health, and education, but SA's underlying network architecture, particularly connectivity patterns, remains largely uncharted. We integrated meta-analytic findings with connectivity-behavior correlation analyses to systematically identify SA-related regions and connections in healthy adults. Edge-weighted networks capturing public, private, and composite SA dimensions were established, where weights represented correlation strengths between tractography-derived structural connectivities and SA levels quantified through behavioral assessments. Then, multilevel SA networks were extracted across a spectrum of correlation thresholds. Robust full-threshold analyses revealed their hierarchical continuum encompassing distinct lateralization patterns, topological transitions, and characteristic hourglass-like architectures. Pathological analysis demonstrated SA connectivity disruptions in schizophrenia (SZ) and major depressive disorder (MDD): approximately 40 % of SA-related connectivities were altered in SZ and 20 % in MDD, with 90 % of MDD alterations overlapping with SZ. While disease-specific and shared alterations were also observed in network-level topological properties, the core SA connectivity framework remained preserved in both disorders. Collectively, these findings significantly advanced our understanding of SA's neurobiological substrates and their pathological deviations.

EAAI Journal 2024 Journal Article

A novel Fermatean fuzzy BWM-VIKOR based multi-criteria decision-making approach for selecting health care waste treatment technology

  • Fei Gao
  • Meihong Han
  • Siyang Wang
  • Jie Gao

Medical waste management (MWM) is a challenging issue for medical facility managers owing to its potential to prevent environmental and health risks. Many treatment technologies (TTs) have been used for MWM, and selecting appropriate TTs is a complex multi-criteria decision-making (MCDM) problem. In this paper, we propose a novel integrated MCDM method based on the best-worst method (BWM) and the VlseKriterijuska Optimizacija I Komoromisno Resenje (VIKOR) method under the Fermatean fuzzy environment to evaluate, rank, and select treatment technologies for medical waste management. First, novel distance measure and entropy measure for Fermatean fuzzy sets are developed, and their properties are examined. Second, different treatment technologies are evaluated by experts using Fermatean fuzzy sets. Third, the experts’ weights are determined using a novel method based on the entropy measure, and the criteria weights are calculated by a novel hybrid criteria weight calculation method. Finally, the ranking of different treatment technologies is determined by the proposed method. To validate the proposed method, a real case study in Jinan, China is presented, where the proposed method is applied to determine the optimal health care waste (HCW) treatment technologies from five TTs with eight criteria. The results are compared to other MCDM methods, which shows the effectiveness and reliability of the proposed method. Moreover, sensitivity analysis is also carried out to show the robustness of the proposed method. From the results, it can be found that the proposed method could provide reliable and robust results for obtaining the optimal health care waste treatment technology.

EAAI Journal 2024 Journal Article

An integrated hesitant 2-tuple linguistic Pythagorean fuzzy decision-making method for single-pilot operations mechanism evaluation

  • Fei Gao
  • Ying Zhang
  • Yijia Li
  • Wenhao Bi

Single-pilot operations (SPO), i. e. , one pilot on board the flight in charge of all the operations, has received extensive attention in recent years. SPO involves several participants including the pilot, the autopilot, and the ground operator, and it is of great significance to determine the roles of these participants under the SPO mechanism. However, the evaluation of SPO mechanism has received little attention. Aiming at providing a reliable approach for SPO mechanism evaluation under uncertainty, this paper proposes a novel hesitant 2-tuple linguistic Pythagorean fuzzy decision-making method based on hesitant fuzzy linguistic term sets (HFLTS), 2-tuple linguistic information, Pythagorean fuzzy sets (PFS), and VIKOR method to handle the SPO mechanism evaluation problem. Firstly, based on the analysis of SPO, the evaluation criteria system for SPO mechanism is established. Then, considering the uncertain information in the SPO mechanism evaluation process, the HFLTS is used to represent the linguistic judgments of experts under uncertainty. Next, the HFLTSs of different experts are converted into 2-tuple linguistic information and PFSs to model the overall judgments of different criteria. Finally, the extended VIKOR method is adopted to evaluate and rank different SPO mechanisms. A case study of SPO mechanism evaluation for the ground proximity warning system (GPWS) is used to demonstrate the effectiveness and feasibility of the proposed method, and the SPO mechanism that the autopilot being responsible for GPWS is evaluated by the proposed method to be optimal. Sensitivity and comparison analyses further confirm that the proposed method could provide reliable and reasonable SPO mechanism evaluation results. In conclusion, the proposed method presents a novel and effective way to evaluate and select appropriate SPO mechanisms.

YNIMG Journal 2024 Journal Article

Brain extended and closed forms glutathione levels decrease with age and extended glutathione is associated with visuospatial memory

  • Xin Hu
  • Keyu Pan
  • Min Zhao
  • Jiali Lv
  • Jing Wang
  • Xiaofeng Zhang
  • Yuxi Liu
  • Yulu Song

During aging, the brain is subject to greater oxidative stress (OS), which is thought to play a critical role in cognitive impairment. Glutathione (GSH), as a major antioxidant in the brain, can be used to combat OS. However, how brain GSH levels vary with age and their associations with cognitive function is unclear. In this study, we combined point-resolved spectroscopy and edited spectroscopy sequences to investigate extended and closed forms GSH levels in the anterior cingulate cortex (ACC), posterior cingulate cortex (PCC), and occipital cortex (OC) of 276 healthy participants (extended form, 166 females, age range 20-70 years) and 15 healthy participants (closed form, 7 females, age range 26-56 years), and examined their relationships with age and cognitive function. The results revealed decreased extended form GSH levels with age in the PCC among 276 participants. Notably, the timecourse of extended form GSH level changes in the PCC and ACC differed between males and females. Additionally, positive correlations were observed between extended form GSH levels in the PCC and OC and visuospatial memory. Additionally, a decreased trend of closed form GSH levels with age was also observed in the PCC among 15 participants. Taken together, these findings enhance our understanding of the brain both closed and extended form GSH time course during normal aging and associations with sex and memory, which is an essential first step for understanding the neurochemical underpinnings of healthy aging.

YNIMG Journal 2024 Journal Article

Bundle-specific tractogram distribution estimation using higher-order streamline differential equation

  • Yuanjing Feng
  • Lei Xie
  • Jingqiang Wang
  • Qiyuan Tian
  • Jianzhong He
  • Qingrun Zeng
  • Fei Gao

Streamline tractography locally traces peak directions extracted from fiber orientation distribution (FOD) functions, lacking global information about the trend of the whole fiber bundle. Therefore, it is prone to producing erroneous tracks while missing true positive connections. In this work, we propose a new bundle-specific tractography (BST) method based on a bundle-specific tractogram distribution (BTD) function, which directly reconstructs the fiber trajectory from the start region to the termination region by incorporating the global information in the fiber bundle mask. A unified framework for any higher-order streamline differential equation is presented to describe the fiber bundles with disjoint streamlines defined based on the diffusion vectorial field. At the global level, the tractography process is simplified as the estimation of BTD coefficients by minimizing the energy optimization model, and is used to characterize the relations between BTD and diffusion tensor vector under the prior guidance by introducing the tractogram bundle information to provide anatomic priors. Experiments are performed on simulated Hough, Sine, Circle data, ISMRM 2015 Tractography Challenge data, FiberCup data, and in vivo data from the Human Connectome Project (HCP) for qualitative and quantitative evaluation. Results demonstrate that our approach reconstructs complex fiber geometry more accurately. BTD reduces the error deviation and accumulation at the local level and shows better results in reconstructing long-range, twisting, and large fanning tracts.

YNIMG Journal 2024 Journal Article

Concurrent behavioral modeling and multimodal neuroimaging reveals how feedback affects the performance of decision making in internet gaming disorder

  • Xinglin Zeng
  • Ying Hao Sun
  • Fei Gao
  • Lin Hua
  • Shiyang Xu
  • Zhen Yuan

Internet gaming disorder (IGD) prompts inquiry into how feedback from prior gaming rounds influences subsequent risk-taking behavior and potential neural mechanisms. Forty-two participants, including 15 with IGD and 27 health controls (HCs), underwent a sequential risk-taking task. Hierarchy Bayesian modeling was adopted to measure risky propensity, behavioral consistence, and affection by emotion ratings from last trial. Concurrent electroencephalogram and functional near-infrared spectroscopy (EEG-fNIRS) recordings were performed to demonstrate when, where and how the previous-round feedback affects the decision making to the next round. We discovered that the IGD illustrated heightened risk-taking propensity as compared to the HCs, indicating by the computational modeling (p = 0.028). EEG results also showed significant time window differences in univariate and multivariate pattern analysis between the IGD and HCs after the loss of the game. Further, reduced brain activation in the prefrontal cortex during the task was detected in IGD as compared to that of the control group. The findings underscore the importance of understanding the aberrant decision-making processes in IGD and suggest potential implications for future interventions and treatments aimed at addressing this behavioral addiction.

EAAI Journal 2024 Journal Article

Human-like mechanism deep learning model for longitudinal motion control of autonomous vehicles

  • Zhenhai Gao
  • Tong Yu
  • Fei Gao
  • Rui Zhao
  • Tianjun Sun

Artificial intelligence (AI) plays a critical role in the prediction, planning, and control of autonomous vehicle. The original motion control methods are increasing in accuracy, but their control style markedly differs from that of human drivers. This not only degrades the passenger's ride experience but also brings unfamiliarity to other drivers, leading to safety hazards. Moreover, some of the methods are a direct application of AI techniques without any specific enhancements tailored for vehicle control. This both fails to fully utilize the potential of AI and reduces interpretability. To address these issues, in this paper, deep learning methods are integrated with driver control mechanisms to propose a human-like neural network for vehicle longitudinal dynamics estimation and control (HNN). First, the process of driver estimation and control of vehicle dynamics is proposed and described in mathematical language. Subsequently, the HNN composed of three sub-networks is presented. The three sub-networks correspond to the three sub-processes of the driver's mechanism, which makes the proposed HNN method more applicable to vehicle dynamics control and more interpretable in specific engineering domains. The effectiveness of the HNN method is verified on a real-world dataset. The results demonstrate that the HNN not only enhances vehicle human-like control but also surpasses baselines in terms of control style consistency and convergence rate.

EAAI Journal 2024 Journal Article

The way to smart civil aviation: An integrated decision making approach for smart civil aviation assessment in China

  • Shuida Bao
  • Fei Gao
  • Zhaoyue Zhang
  • Qingjun Xia
  • Wenhao Bi

To address the critical need to objectively assess the progress and effectiveness of China’s pioneering smart civil aviation initiative, this paper integrates 2-dimensional linguistic intuitionistic fuzzy sets (2DLIFSs) with the VlseKriterijumska Optimizacija I Kompromisno Resenje (VIKOR) method. This integration develops a novel evaluation approach for assessing smart civil aviation performance across a cluster of airports, ultimately accelerating its implementation. The evaluation attribute system comprehensively covers smart travel, smart airport, technological innovation, and comprehensive effect. Furthermore, a novel hybrid distance metric is introduced to quantify the difference and closeness of 2DLIFSs with the ideal solution. The case study demonstrates the reliability and effectiveness of the proposed method for assessing smart civil aviation. Additional validation is provided through sensitivity analysis and comparative studies. This work offers a novel approach to evaluating smart civil aviation, enabling informed decisions and enhancing the efficiency and sustainability of the civil aviation sector.

ICRA Conference 2023 Conference Paper

Efficient View Path Planning for Autonomous Implicit Reconstruction

  • Jing Zeng
  • Yanxu Li
  • Yunlong Ran
  • Shuo Li
  • Fei Gao
  • Lincheng Li
  • Shibo He
  • Jiming Chen 0001

Implicit neural representations have shown promising potential for 3D scene reconstruction. Recent work applies it to autonomous 3D reconstruction by learning information gain for view path planning. Effective as it is, the computation of the information gain is expensive, and compared with that using volumetric representations, collision checking using the implicit representation for a 3D point is much slower. In the paper, we propose to 1) leverage a neural network as an implicit function approximator for the information gain field and 2) combine the implicit fine-grained representation with coarse volumetric representations to improve efficiency. Further with the improved efficiency, we propose a novel informative path planning based on a graph-based planner. Our method demonstrates significant improvements in the reconstruction quality and planning efficiency compared with autonomous reconstructions with implicit and explicit representations. We deploy the method on a real UAV and the results show that our method can plan informative views and reconstruct a scene with high quality.

YNIMG Journal 2023 Journal Article

Neurochemical and functional reorganization of the cognitive-ear link underlies cognitive impairment in presbycusis

  • Ning Li
  • Wen Ma
  • Fuxin Ren
  • Xiao Li
  • Fuyan Li
  • Wei Zong
  • Lili Wu
  • Zongrui Dai

Recent studies suggest that the interaction between presbycusis and cognitive impairment may be partially explained by the cognitive-ear link. However, the underlying neurophysiological mechanisms remain largely unknown. In this study, we combined magnetic resonance spectroscopy (MRS) and resting-state functional magnetic resonance imaging (fMRI) to investigate auditory gamma-aminobutyric acid (GABA) and glutamate (Glu) levels, intra- and inter-network functional connectivity, and their relationships with auditory and cognitive function in 51 presbycusis patients and 51 well-matched healthy controls. Our results confirmed reorganization of the cognitive-ear link in presbycusis, including decreased auditory GABA and Glu levels and aberrant functional connectivity involving auditory networks (AN) and cognitive-related networks, which were associated with reduced speech perception or cognitive impairment. Moreover, mediation analyses revealed that decreased auditory GABA levels and dysconnectivity between the AN and default mode network (DMN) mediated the association between hearing loss and impaired information processing speed in presbycusis. These findings highlight the importance of AN-DMN dysconnectivity in cognitive-ear link reorganization leading to cognitive impairment, and hearing loss may drive reorganization via decreased auditory GABA levels. Modulation of GABA neurotransmission may lead to new treatment strategies for cognitive impairment in presbycusis patients.

IJCAI Conference 2023 Conference Paper

Semantic-Aware Generation of Multi-View Portrait Drawings

  • Biao Ma
  • Fei Gao
  • Chang Jiang
  • Nannan Wang
  • Gang Xu

Neural radiance fields (NeRF) based methods have shown amazing performance in synthesizing 3D-consistent photographic images, but fail to generate multi-view portrait drawings. The key is that the basic assumption of these methods -- a surface point is consistent when rendered from different views -- doesn't hold for drawings. In a portrait drawing, the appearance of a facial point may changes when viewed from different angles. Besides, portrait drawings usually present little 3D information and suffer from insufficient training data. To combat this challenge, in this paper, we propose a Semantic-Aware GEnerator (SAGE) for synthesizing multi-view portrait drawings. Our motivation is that facial semantic labels are view-consistent and correlate with drawing techniques. We therefore propose to collaboratively synthesize multi-view semantic maps and the corresponding portrait drawings. To facilitate training, we design a semantic-aware domain translator, which generates portrait drawings based on features of photographic faces. In addition, use data augmentation via synthesis to mitigate collapsed results. We apply SAGE to synthesize multi-view portrait drawings in diverse artistic styles. Experimental results show that SAGE achieves significantly superior or highly competitive performance, compared to existing 3D-aware image synthesis methods. The codes are available at https: //github. com/AiArt-HDU/SAGE.

ICML Conference 2022 Conference Paper

Disentangling Disease-related Representation from Obscure for Disease Prediction

  • Churan Wang
  • Fei Gao
  • Fandong Zhang
  • Fangwei Zhong
  • Yizhou Yu
  • Yizhou Wang 0001

Disease-related representations play a crucial role in image-based disease prediction such as cancer diagnosis, due to its considerable generalization capacity. However, it is still a challenge to identify lesion characteristics in obscured images, as many lesions are obscured by other tissues. In this paper, to learn the representations for identifying obscured lesions, we propose a disentanglement learning strategy under the guidance of alpha blending generation in an encoder-decoder framework (DAB-Net). Specifically, we take mammogram mass benign/malignant classification as an example. In our framework, composite obscured mass images are generated by alpha blending and then explicitly disentangled into disease-related mass features and interference glands features. To achieve disentanglement learning, features of these two parts are decoded to reconstruct the mass and the glands with corresponding reconstruction losses, and only disease-related mass features are fed into the classifier for disease prediction. Experimental results on one public dataset DDSM and three in-house datasets demonstrate that the proposed strategy can achieve state-of-the-art performance. DAB-Net achieves substantial improvements of 3. 9%~4. 4% AUC in obscured cases. Besides, the visualization analysis shows the model can better disentangle the mass and glands in the obscured image, suggesting the effectiveness of our solution in exploring the hidden characteristics in this challenging problem.

YNIMG Journal 2021 Journal Article

Species and individual differences and connectional asymmetry of Broca's area in humans and macaques

  • Xiaoluan Xia
  • Fei Gao
  • Zhen Yuan

To reveal the connectional specialization of the Broca's area (or its homologue), voxel-wise inter-species and individual differences, and inter-hemispheric asymmetry were respectively inspected in humans and macaques at both whole-brain connectivity and single tract levels. It was discovered that the developed connectivity blueprint approach is able to localize connectionally comparable voxels between the two species in Broca's area, whereas the quantitative differences between blueprints of locationally or connectionally corresponding voxels enable us to generate inter-hemispheric, inter-subject, and inter-species connectional variabilities, respectively. More importantly, the inter-species and inter-subject variabilities exhibited positive correlation in both two primates, and relatively higher variabilities were detected in the anatomically defined pars triangularis. By contrast, negative relationship was identified between the inter-species variability and hemispheric asymmetry in human brain. In particular, relatively higher asymmetry was revealed in the anatomically defined pars opercularis. Therefore, our novel findings demonstrated that pars triangularis, as compared to pars opercularis, might be a more active area during primate evolution, in which the brain connectivity and possible functions of pars triangularis show relatively higher degree in species specialization, yet lower in hemispheric specialization. Meanwhile, brain connectivity and possible functions of pars opercularis manifested an opposite pattern. At the tract level, functional roles related to the ventral stream in speech comprehension were relatively conservative and bilaterally organized, while those related to the dorsal stream in speech production show relatively higher species and hemispheric specializations.

YNICL Journal 2020 Journal Article

AD-NET: Age-adjust neural network for improved MCI to AD conversion prediction

  • Fei Gao
  • Hyunsoo Yoon
  • Yanzhe Xu
  • Dhruman Goradia
  • Ji Luo
  • Teresa Wu
  • Yi Su

The prediction of Mild Cognitive Impairment (MCI) patients who are at higher risk converting to Alzheimer's Disease (AD) is critical for effective intervention and patient selection in clinical trials. Different biomarkers including neuroimaging have been developed to serve the purpose. With extensive methodology development efforts on neuroimaging, an emerging field is deep learning research. One great challenge facing deep learning is the limited medical imaging data available. To address the issue, researchers explore the use of transfer learning to extend the applicability of deep models on neuroimaging research for AD diagnosis and prognosis. Existing transfer learning models mostly focus on transferring the features from the pre-training into the fine-tuning stage. Recognizing the advantages of the knowledge gained during the pre-training, we propose an AD-NET (Age-adjust neural network) with the pre-training model serving two purposes: extracting and transferring features; and obtaining and transferring knowledge. Specifically, the knowledge being transferred in this research is an age-related surrogate biomarker. To evaluate the effectiveness of the proposed approach, AD-NET is compared with 8 classification models from literature using the same public neuroimaging dataset. Experimental results show that the proposed AD-NET outperforms the competing models in predicting the MCI patients at risk for conversion to the AD stage.

YNICL Journal 2020 Journal Article

Brain GABA+ changes in primary hypothyroidism patients before and after levothyroxine treatment: A longitudinal magnetic resonance spectroscopy study

  • Bo Liu
  • Zhensong Wang
  • Liangjie Lin
  • Huan Yang
  • Fei Gao
  • Tao Gong
  • Richard A.E. Edden
  • Guangbin Wang

OBJECTIVE: Increasing evidence indicates the involvement of the GABAergic system in the pathophysiology of hypothyroidism. We aimed to investigate longitudinal changes of brain GABA in primary hypothyroidism before and after levothyroxine (L-T4) treatment. MATERIAL AND METHODS: In 18 patients with hypothyroidism, we used the MEGA-PRESS (Mescher-Garwood point-resolved spectroscopy) editing sequence to measure brain GABA levels from medial prefrontal cortex (mPFC) and posterior cingulate cortex (PCC) at baseline and after 6-months of L-T4 treatment. Sex- and age-matched healthy controls (n = 18) were scanned at baseline. Thyroid function and neuropsychological tests were also performed. RESULTS: GABA signals were successfully quantified from all participants with fitting errors lower than 15%. GABA signal was labeled as GABA+ due to contamination from co-edited macromoleculars and homocarnosine. In hypothyroid patients, mean GABA+ was significantly lower in the mPFC region compared with controls (p = 0.031), and the mPFC GABA+ measurements were significantly correlated with depressive symptoms and memory function (r = -0.558, p = 0.016; r = 0.522, p = 0.026, respectively). After adequate L-T4 treatment, the mPFC GABA+ in hypothyroid patients increased to normal level, along with relieved neuropsychological impairments. CONCLUSION: The study suggested the decrease of GABA+ may be an important neurobiological factor in the pathophysiology of hypothyroidism. Treatment of L-T4 may reverse the abnormal GABA+ and hypothyroidism-induced neuropsychiatric impairments, indicating the action mode of L-T4 in adjunctive treatment of affective disorders.

JBHI Journal 2020 Journal Article

Deep Residual Inception Encoder–Decoder Network for Medical Imaging Synthesis

  • Fei Gao
  • Teresa Wu
  • Xianghua Chu
  • Hyunsoo Yoon
  • Yanzhe Xu
  • Bhavika Patel

Image synthesis is a novel solution in precision medicine for scenarios where important medical imaging is not otherwise available. The convolutional neural network (CNN) is an ideal model for this task because of its powerful learning capabilities through the large number of layers and trainable parameters. In this research, we propose a new architecture of residual inception encoder- decoder neural network (RIED-Net) to learn the nonlinear mapping between the input images and targeting output images. To evaluate the validity of the proposed approach, it is compared with two models from the literature: synthetic CT deep convolutional neural network (sCT-DCNN) and shallow CNN, using both an institutional mammogram dataset from Mayo Clinic Arizona and a public neuroimaging dataset from the Alzheimer's Disease Neuroimaging Initiative. Experimental results show that the proposed RIED-Net outperforms the two models on both datasets significantly in terms of structural similarity index, mean absolute percent error, and peak signal-to-noise ratio.

YNIMG Journal 2020 Journal Article

Deficits in ascending and descending pain modulation pathways in patients with postherpetic neuralgia

  • Hong Li
  • Xiaoyun Li
  • Yi Feng
  • Fei Gao
  • Yazhuo Kong
  • Li Hu

Postherpetic Neuralgia (PHN), develops after the resolution of the herpes zoster mucocutaneous eruption, is a debilitating chronic pain. However, there is a lack of knowledge regarding the underlying mechanisms associated with ascending and descending pain modulations in PHN patients. Here, we combined psychophysics with structural and functional magnetic resonance imaging (MRI) techniques to investigate the brain alternations in PHN patients. Psychophysical tests showed that compared with healthy controls, PHN patients had increased state and trait anxiety and depression. Structural MRI data indicated that PHN patients had significantly smaller gray matter volumes of the thalamus and amygdala than healthy controls, and the thalamus volume was negatively correlated with pain intensity (assessed using the Short-form of the McGill pain questionnaire) in PHN patients. When the thalamus and periaqueductal gray matter (PAG) were used as the seeds, resting-state functional MRI data revealed abnormal patterns of functional connectivity within ascending and descending pain pathways in PHN patients, e. g. , increased functional connectivity between the thalamus and somatosensory cortices and decreased functional connectivity between the PAG and frontal cortices. In addition, subjective ratings of both Present Pain Index (PPI) and Beck-Depression Inventory (BDI) were negatively correlated with the strength of functional connectivity between the PAG and primary somatosensory cortex (SI), and importantly, the effect of BDI on PPI was mediated by the PAG-SI functional connectivity. Overall, our results provided evidence suggesting deficits in ascending and descending pain modulation pathways, which were highly associated with the intensity of chronic pain and its emotional comorbidities in PHN patients. Therefore, our study deepened our understanding of the pathogenesis of PHN, which would be helpful in determining the optimized treatment for the patients.

YNIMG Journal 2020 Journal Article

Functional engagement of white matter in resting-state brain networks

  • Muwei Li
  • Yurui Gao
  • Fei Gao
  • Adam W. Anderson
  • Zhaohua Ding
  • John C. Gore

The topological characteristics of functional networks, derived from measurements of resting-state connectivity in gray matter (GM), are associated with individual cognitive abilities or specific dysfunctions. However, blood oxygen level-dependent (BOLD) signals in white matter (WM) are usually ignored or even regressed out as nuisance factors in the data analyses that underlie network models. Recent studies have demonstrated reliable detection of WM BOLD signals and imply these reflect associated neural activities. Here we evaluate quantitatively the contributions of individual WM voxels to the identification of functional networks, which we term their engagement (or conceptually, their importance). We quantify the engagement by measuring the reductions of connectivity, produced by ignoring the signal fluctuations within each WM voxel, with respect to both the entire network (global) or a single GM node (local). We observed highly reproducible spatial distributions of global engagement maps, as well as a trend toward increased relevance of deep WM voxels at delayed times. Local engagement maps exhibit homogeneous spatial distributions with respect to internal nodes that constitute a well-recognized sub-functional network, but inhomogeneous distributions with respect to other nodes. WM voxels show distinct distributions of engagement depending on their anatomical locations. These findings demonstrate the important role of WM in network modeling, thus supporting the need for changes of conventional views that WM signal variations represent only physiological noise.

YNICL Journal 2020 Journal Article

Protein-based amide proton transfer-weighted MR imaging of amnestic mild cognitive impairment

  • Zewen Zhang
  • Caiqing Zhang
  • Jian Yao
  • Xin Chen
  • Fei Gao
  • Shanshan Jiang
  • Weibo Chen
  • Jinyuan Zhou

Amide proton transfer-weighted (APTw) MRI is a novel molecular imaging technique that can noninvasively detect endogenous cellular proteins and peptides in tissue. Here, we demonstrate the feasibility of protein-based APTw MRI in characterizing amnestic mild cognitive impairment (aMCI). Eighteen patients with confirmed aMCI and 18 matched normal controls were scanned at 3 Tesla. The APTw, as well as conventional magnetization transfer ratio (MTR), signal differences between aMCI and normal groups were assessed by the independent samples t-test, and the receiver-operator-characteristic analysis was used to assess the diagnostic performance of APTw. When comparing the normal control group, aMCI brains typically had relatively higher APTw signals. Quantitatively, APTw intensity values were significantly higher in nine of 12 regions of interest in aMCI patients than in normal controls. The largest areas under the receiver-operator-characteristic curves were 0.88 (gray matter in occipital lobe) and 0.82 (gray matter in temporal lobe, white matter in occipital lobe) in diagnosing aMCI patients. On the contrary, MTR intensity values were significantly higher in only three of 12 regions of interest in the aMCI group. Additionally, the age dependency analyses revealed that these cross-sectional APTw/MTR signals had an increasing trend with age in most brain regions for normal controls, but a decreasing trend with age in most brain regions for aMCI patients. Our early results show the potential of the APTw signal as a new imaging biomarker for the noninvasive molecular diagnosis of aMCI.

YNIMG Journal 2019 Journal Article

Big GABA II: Water-referenced edited MR spectroscopy at 25 research sites

  • Mark Mikkelsen
  • Daniel L. Rimbault
  • Peter B. Barker
  • Pallab K. Bhattacharyya
  • Maiken K. Brix
  • Pieter F. Buur
  • Kim M. Cecil
  • Kimberly L. Chan

Accurate and reliable quantification of brain metabolites measured in vivo using 1H magnetic resonance spectroscopy (MRS) is a topic of continued interest. Aside from differences in the basic approach to quantification, the quantification of metabolite data acquired at different sites and on different platforms poses an additional methodological challenge. In this study, spectrally edited γ-aminobutyric acid (GABA) MRS data were analyzed and GABA levels were quantified relative to an internal tissue water reference. Data from 284 volunteers scanned across 25 research sites were collected using GABA+ (GABA + co-edited macromolecules (MM)) and MM-suppressed GABA editing. The unsuppressed water signal from the volume of interest was acquired for concentration referencing. Whole-brain T 1-weighted structural images were acquired and segmented to determine gray matter, white matter and cerebrospinal fluid voxel tissue fractions. Water-referenced GABA measurements were fully corrected for tissue-dependent signal relaxation and water visibility effects. The cohort-wide coefficient of variation was 17% for the GABA + data and 29% for the MM-suppressed GABA data. The mean within-site coefficient of variation was 10% for the GABA + data and 19% for the MM-suppressed GABA data. Vendor differences contributed 53% to the total variance in the GABA + data, while the remaining variance was attributed to site- (11%) and participant-level (36%) effects. For the MM-suppressed data, 54% of the variance was attributed to site differences, while the remaining 46% was attributed to participant differences. Results from an exploratory analysis suggested that the vendor differences were related to the unsuppressed water signal acquisition. Discounting the observed vendor-specific effects, water-referenced GABA measurements exhibit similar levels of variance to creatine-referenced GABA measurements. It is concluded that quantification using internal tissue water referencing is a viable and reliable method for the quantification of in vivo GABA levels.

NeurIPS Conference 2019 Conference Paper

Neural Machine Translation with Soft Prototype

  • Yiren Wang
  • Yingce Xia
  • Fei Tian
  • Fei Gao
  • Tao Qin
  • Cheng Xiang Zhai
  • Tie-Yan Liu

Neural machine translation models usually use the encoder-decoder framework and generate translation from left to right (or right to left) without fully utilizing the target-side global information. A few recent approaches seek to exploit the global information through two-pass decoding, yet have limitations in translation quality and model efficiency. In this work, we propose a new framework that introduces a soft prototype into the encoder-decoder architecture, which allows the decoder to have indirect access to both past and future information, such that each target word can be generated based on the better global understanding. We further provide an efficient and effective method to generate the prototype. Empirical studies on various neural machine translation tasks show that our approach brings significant improvement in generation quality over the baseline model, with little extra cost in storage and inference time, demonstrating the effectiveness of our proposed framework. Specially, we achieve state-of-the-art results on WMT2014, 2015 and 2017 English to German translation.

NeurIPS Conference 2019 Conference Paper

Normalization Helps Training of Quantized LSTM

  • Lu Hou
  • Jinhua Zhu
  • James Kwok
  • Fei Gao
  • Tao Qin
  • Tie-Yan Liu

The long-short-term memory (LSTM), though powerful, is memory and computa\x02tion expensive. To alleviate this problem, one approach is to compress its weights by quantization. However, existing quantization methods usually have inferior performance when used on LSTMs. In this paper, we first show theoretically that training a quantized LSTM is difficult because quantization makes the exploding gradient problem more severe, particularly when the LSTM weight matrices are large. We then show that the popularly used weight/layer/batch normalization schemes can help stabilize the gradient magnitude in training quantized LSTMs. Empirical results show that the normalized quantized LSTMs achieve significantly better results than their unnormalized counterparts. Their performance is also comparable with the full-precision LSTM, while being much smaller in size.

YNIMG Journal 2017 Journal Article

Big GABA: Edited MR spectroscopy at 24 research sites

  • Mark Mikkelsen
  • Peter B. Barker
  • Pallab K. Bhattacharyya
  • Maiken K. Brix
  • Pieter F. Buur
  • Kim M. Cecil
  • Kimberly L. Chan
  • David Y.-T. Chen

Magnetic resonance spectroscopy (MRS) is the only biomedical imaging method that can noninvasively detect endogenous signals from the neurotransmitter γ-aminobutyric acid (GABA) in the human brain. Its increasing popularity has been aided by improvements in scanner hardware and acquisition methodology, as well as by broader access to pulse sequences that can selectively detect GABA, in particular J-difference spectral editing sequences. Nevertheless, implementations of GABA-edited MRS remain diverse across research sites, making comparisons between studies challenging. This large-scale multi-vendor, multi-site study seeks to better understand the factors that impact measurement outcomes of GABA-edited MRS. An international consortium of 24 research sites was formed. Data from 272 healthy adults were acquired on scanners from the three major MRI vendors and analyzed using the Gannet processing pipeline. MRS data were acquired in the medial parietal lobe with standard GABA+ and macromolecule- (MM-) suppressed GABA editing. The coefficient of variation across the entire cohort was 12% for GABA+ measurements and 28% for MM-suppressed GABA measurements. A multilevel analysis revealed that most of the variance (72%) in the GABA+ data was accounted for by differences between participants within-site, while site-level differences accounted for comparatively more variance (20%) than vendor-level differences (8%). For MM-suppressed GABA data, the variance was distributed equally between site- (50%) and participant-level (50%) differences. The findings show that GABA+ measurements exhibit strong agreement when implemented with a standard protocol. There is, however, increased variability for MM-suppressed GABA measurements that is attributed in part to differences in site-to-site data acquisition. This study's protocol establishes a framework for future methodological standardization of GABA-edited MRS, while the results provide valuable benchmarks for the MRS community.

IJCAI Conference 2017 Conference Paper

Efficient Inexact Proximal Gradient Algorithm for Nonconvex Problems

  • Quanming Yao
  • James T. Kwok
  • Fei Gao
  • Wei Chen
  • Tie-Yan Liu

While proximal gradient algorithm is originally designed for convex optimization, several variants have been recently proposed for nonconvex problems. Among them, nmAPG [Li and Lin, 2015] is the state-of-art. However, it is inefficient when the proximal step does not have closed-form solution, or such solution exists but is expensive, as it requires more than one proximal steps to be exactly solved in each iteration. In this paper, we propose an efficient accelerate proximal gradient (niAPG) algorithm for nonconvex problems. In each iteration, it requires only one inexact (less expensive) proximal step. Convergence to a critical point is still guaranteed, and a O(1/k) convergence rate is derived. Experiments on image inpainting and matrix completion problems demonstrate that the proposed algorithm has comparable performance as the state-of-the-art, but is much faster.

YNIMG Journal 2015 Journal Article

Decreased auditory GABA+ concentrations in presbycusis demonstrated by edited magnetic resonance spectroscopy

  • Fei Gao
  • Guangbin Wang
  • Wen Ma
  • Fuxin Ren
  • Muwei Li
  • Yuling Dong
  • Cheng Liu
  • Bo Liu

Gamma-aminobutyric acid (GABA) is the main inhibitory neurotransmitter in the central auditory system. Altered GABAergic neurotransmission has been found in both the inferior colliculus and the auditory cortex in animal models of presbycusis. Edited magnetic resonance spectroscopy (MRS), using the MEGA-PRESS sequence, is the most widely used technique for detecting GABA in the human brain. However, to date there has been a paucity of studies exploring changes to the GABA concentrations in the auditory region of patients with presbycusis. In this study, sixteen patients with presbycusis (5 males/11 females, mean age 63. 1±2. 6years) and twenty healthy controls (6 males/14 females, mean age 62. 5±2. 3years) underwent audiological and MRS examinations. Pure tone audiometry from 0. 125 to 8kHz and tympanometry were used to assess the hearing abilities of all subjects. The pure tone average (PTA; the average of hearing thresholds at 0. 5, 1, 2 and 4kHz) was calculated. The MEGA-PRESS sequence was used to measure GABA+ concentrations in 4×3×3cm3 volumes centered on the left and right Heschl's gyri. GABA+ concentrations were significantly lower in the presbycusis group compared to the control group (left auditory regions: p =0. 002, right auditory regions: p =0. 008). Significant negative correlations were observed between PTA and GABA+ concentrations in the presbycusis group (r =−0. 57, p =0. 02), while a similar trend was found in the control group (r =−0. 40, p =0. 08). These results are consistent with a hypothesis of dysfunctional GABAergic neurotransmission in the central auditory system in presbycusis and suggest a potential treatment target for presbycusis.

YNIMG Journal 2013 Journal Article

Edited magnetic resonance spectroscopy detects an age-related decline in brain GABA levels

  • Fei Gao
  • Richard A.E. Edden
  • Muwei Li
  • Nicolaas A.J. Puts
  • Guangbin Wang
  • Cheng Liu
  • Bin Zhao
  • Huiquan Wang

Gamma-aminobutyric acid (GABA) is the primary inhibitory neurotransmitter in the brain. Although measurements of GABA levels in vivo in the human brain using edited proton magnetic resonance spectroscopy (1H-MRS) have been established for some time, it is has not been established how regional GABA levels vary with age in the normal human brain. In this study, 49 healthy men and 51 healthy women aged between 20 and 76years were recruited and J-difference edited spectra were recorded at 3T to determine the effect of age on GABA levels, and to investigate whether there are regional and gender differences in GABA in mesial frontal and parietal regions. Because the signal detected at 3. 02ppm using these experimental parameters is also expected to contain contributions from both macromolecules (MM) and homocarnosine, in this study the signal is labeled GABA+ rather than GABA. Significant negative correlations were observed between age and GABA+ in both regions studied (GABA+/Cr: frontal region, r=−0. 68, p<0. 001, parietal region, r=−0. 54, p<0. 001; GABA+/NAA: frontal region, r=−0. 58, p<0. 001, parietal region, r=−0. 49, p<0. 001). The decrease in GABA+ with age in the frontal region was more rapid in women than men. Evidence of a measureable decline in GABA is important in considering the neurochemical basis of the cognitive decline that is associated with normal aging.

v2026.09.13