Arrow Research search

Author name cluster

Feng Liu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

88 papers
2 author rows

Possible papers

88

EAAI Journal 2026 Journal Article

Dynamic path smooth unfolding network and learnable random smoothing strategy for magnetic resonance imaging compressed sensing

  • Ziqi Yang
  • Mingfeng Jiang
  • Chenghu Geng
  • Zhifeng Chen
  • Mengyu Jia
  • Xiaocheng Yang
  • Sumei Huang
  • Feng Liu

Deep Unfolding Networks (DUNs) have become the mainstream approach for compressed sensing Magnetic Resonance Imaging (MRI) reconstruction from highly under-sampled k-space data. In this paper, a novel Dynamic Path Smooth Unfolding Network (DPSU-Net) is proposed for compressed sensing MRI reconstruction by dynamically selecting different paths for smooth unfolding. Furthermore, a learnable random smoothing strategy is used to enhance model robustness by introducing perturbations through a noise generator during training stage. Experimental results on the FastMRI T1-weighted and T2-weighted images show that DPSU-Net achieves superior reconstruction performance across different under-sampling rates, with Peak Signal-to-Noise Ratio (PSNR)/Structural Similarity Index Measure (SSIM) of 48. 70/0. 9889 on T1-weighted images and 45. 68/0. 9715 on T2-weighted images, surpassing existing state-of-the-art networks. Ablation studies further confirm the effectiveness and robustness of the dynamic path selection and learnable random smoothing strategies, demonstrating improvements in reconstruction quality.

AAAI Conference 2026 Conference Paper

Generalising Traffic Forecasting to Regions Without Traffic Observations

  • Xinyu Su
  • Majid Sarvi
  • Feng Liu
  • Egemen Tanin
  • Jianzhong Qi

Traffic forecasting is essential for intelligent transportation systems. Accurate forecasting relies on continuous observations collected by traffic sensors. However, due to high deployment and maintenance costs, not all regions are equipped with such sensors. This paper aims to forecast for regions without traffic sensors, where the lack of historical traffic observations challenges the generalisability of existing models. We propose a model named **GenCast**, the core idea of which is to exploit external knowledge to compensate for the missing observations and to enhance generalisation. We integrate physics-informed neural networks into GenCast, enabling physical principles to regularise the learning process. We introduce an external signal learning module to explore correlations between traffic states and external signals such as weather conditions, further improving model generalisability. Additionally, we design a spatial grouping module to filter localised features that hinder model generalisability. Extensive experiments show that GenCast consistently reduces forecasting errors on multiple real-world datasets.

AAAI Conference 2026 Conference Paper

HiFi-Mamba: Dual-Stream?-Laplacian Enhanced Mamba for High-Fidelity MRI Reconstruction

  • Hongli Chen
  • Pengcheng Fang
  • Yuxia Chen
  • Yingxuan Ren
  • Jing Hao
  • Fangfang Tang
  • Xiaohao Cai
  • Shanshan Shan

Reconstructing high-fidelity MR images from undersampled k-space data remains a challenging problem in MRI. While Mamba variants for vision tasks offer promising long-range modeling capabilities with linear-time complexity, their direct application to MRI reconstruction inherits two key limitations: (1) insensitivity to high-frequency anatomical details; and (2) reliance on redundant multi-directional scanning. To address these limitations, we introduce High-Fidelity Mamba (HiFi-Mamba), a novel dual-stream Mamba-based architecture comprising stacked?-Laplacian (WL) and HiFi-Mamba blocks. Specifically, the WL block performs fidelity-preserving spectral decoupling, producing complementary low- and high-frequency streams. This separation enables the HiFi-Mamba block to focus on low-frequency structures, enhancing global feature modeling. Concurrently, the HiFi-Mamba block selectively integrates high-frequency features through adaptive state-space modulation, preserving comprehensive spectral details. To eliminate the scanning redundancy, the HiFi-Mamba block adopts a streamlined unidirectional traversal strategy that preserves long-range modeling capability with improved computational efficiency. Extensive experiments on standard MRI reconstruction benchmarks demonstrate that HiFi-Mamba consistently outperforms state-of-the-art CNN-based, Transformer-based, and other Mamba-based models in reconstruction accuracy while maintaining a compact and efficient model design.

TMLR Journal 2026 Journal Article

Learning Representations for Independence Testing

  • Nathaniel Xu
  • Feng Liu
  • Danica J. Sutherland

Many tools exist to detect dependence between random variables, a core question across a wide range of machine learning, statistical, and scientific endeavors. Although several statistical tests guarantee eventual detection of any dependence with enough samples, standard tests may require an exorbitant amount of samples for detecting subtle dependencies between high-dimensional random variables with complex distributions. In this work, we study two related ways to learn powerful independence tests. First, we show how to construct powerful statistical tests with finite-sample validity by using variational estimators of mutual information, such as the InfoNCE or NWJ estimators. Second, we establish a close connection between these variational mutual information-based tests and tests based on the Hilbert-Schmidt Independence Criterion (HSIC); in particular, learning a variational bound (typically parameterized by a deep network) for mutual information is closely related to learning a kernel for HSIC. Finally, we show how to, rather than selecting a representation to maximize the statistic itself, select a representation which can maximize the power of a test, in either setting; we term the former case a Neural Dependency Statistic (NDS). While HSIC power optimization has been recently considered in the literature, we correct some important misconceptions and expand to considering deep kernels. In our experiments, while all approaches can yield powerful tests with exact level control, optimized HSIC tests generally outperform the other approaches on difficult problems of detecting structured dependence.

TMLR Journal 2026 Journal Article

Let's Roll a BiFTA: Bi-refinement for Fine-grained Text-visual Alignment in Vision-Language Models

  • Yuhao Sun
  • Chengyi Cai
  • Jiacheng Zhang
  • Zesheng Ye
  • Xingliang Yuan
  • Feng Liu

Recent research has shown that aligning fine-grained text descriptions with localized image patches can significantly improve the zero-shot performance of pre-trained vision-language models (e.g., CLIP). However, we find that both fine-grained text descriptions and localized image patches often contain redundant information, making text-visual alignment less effective. In this paper, we tackle this issue from two perspectives: \emph{view refinement} and \emph{description refinement}, termed as \textit{\textbf{Bi}-refinement for \textbf{F}ine-grained \textbf{T}ext-visual \textbf{A}lignment} (BiFTA). \emph{View refinement} removes redundant image patches with high \emph{Intersection over Union} (IoU) ratios, resulting in more distinctive visual samples. \emph{Description refinement} removes redundant text descriptions with high pairwise cosine similarity, ensuring greater diversity in the remaining descriptions. BiFTA achieves superior zero-shot performance on 6 benchmark datasets for both ViT-based and ResNet-based CLIP, justifying the necessity to remove redundant information in visual-text alignment.

YNIMG Journal 2026 Journal Article

Mitigating inter-scanner heterogeneity in brain MRI data: Assessing its impact on association analyses and the effectiveness of ComBat harmonization in multi-site neuroimaging studies

  • Xiaoxiao Xiao
  • Jinyu Liu
  • Lining Guo
  • Kaizhong Xue
  • Sijia Wang
  • Feng Liu
  • Wen Qin
  • Chunshui Yu

Recruiting participants from multiple sites accelerates data acquisition and increases the total sample size in neuroimaging studies, thereby enhancing the validity and generalizability of statistical findings. While both meta-analysis and mega-analysis can accommodate multi-site data, the latter leverages more effectively the high statistical power offered by large sample sizes of multi-site datasets. However, multi-site datasets often present abiotic variances stemming from differences in device manufacturer, reconstruction algorithm, acquisition parameters, and other factors, collectively termed scanner effect or inter-scanner heterogeneity. This heterogeneity hinders the application of mega-analysis and, if inadequately addressed, may obscure true effects or produce false positive effects. Furthermore, such scanner effects may vary in their impact on different brain imaging metrics (BIMs). To comprehensively understand the scanner effects on diverse BIMs, we used a multi-modal brain magnetic resonance imaging (MRI) dataset comprising 995 BIMs acquired on 28 MR scanners from two traveling subjects to characterize the inter-scanner heterogeneity (quantified by intraclass correlation coefficients) for each BIM. We then assessed the efficacy of the ComBat (combatting batch effects) harmonization method in removing such inter-scanner heterogeneity. Subsequently, using a large-scale neuroimaging dataset of 7035 subjects from CHIMGEN (Chinese Imaging Genetics), we conducted association analyses between BIMs and demographic/behavioral variables (DBVs) using both meta-analysis and mega-analysis strategies, both before and after applying ComBat harmonization, with the aim to investigate the impact of inter-scanner heterogeneity on association analyses and to evaluate the effectiveness of ComBat in mitigating this impact. The results showed that all BIMs exhibited inter-scanner heterogeneity but with varying degrees - functional connectivity (FC)-related BIMs showed the highest and cortical volume and surface area showed the lowest heterogeneity. ComBat harmonization effectively corrected the heterogeneity for most BIMs, though it was less successful for certain BIMs, particularly those related to FC. The BIM-DBV association analyses indicated that mega-analysis outperformed meta-analysis in general. However, when using uncorrected data, mega-analysis yielded an excessive number of significant, but unreliable, associations, particularly when there were only sparse associations between a DBV and a BIM in the brain. Notably, ComBat harmonization effectively addressed this issue. These results provided the first comprehensive characterization of scanner effects on extensive BIMs and new insights into the effectiveness of the ComBat harmonization technique in mitigating inter-scanner heterogeneity in multi-site neuroimaging studies to ensure the validity and reliability of statistical findings.

AAAI Conference 2026 Conference Paper

RadarMP: Motion Perception for 4D mmWave Radar in Autonomous Driving

  • Ruiqi Cheng
  • Huijun Di
  • Jian Li
  • Feng Liu
  • Wei Liang

Accurate 3D scene motion perception significantly enhances the safety and reliability of an autonomous driving system. Benefiting from its all-weather operational capability and unique perceptual properties, 4D mmWave radar has emerged as an essential component in advanced autonomous driving. However, sparse and noisy radar points often lead to imprecise motion perception, leaving autonomous vehicles with limited sensing capabilities when optical sensors degrade under adverse weather conditions. In this paper, we propose RadarMP, a novel method for precise 3D scene motion perception using low-level radar echo signals from two consecutive frames. Unlike existing methods that separate radar target detection and motion estimation, RadarMP jointly models both tasks in a unified architecture, enabling consistent radar point cloud generation and pointwise 3D scene flow prediction. Tailored to radar characteristics, we design specialized self-supervised loss functions guided by Doppler shifts and echo intensity, effectively supervising spatial and motion consistency without explicit annotations. Extensive experiments on the public dataset demonstrate that RadarMP achieves reliable motion perception across diverse weather and illumination conditions, outperforming radar-based decoupled motion perception pipelines and enhancing perception capabilities for full-scenario autonomous driving systems.

TMLR Journal 2026 Journal Article

Semantic-aware Adversarial Fine-tuning for CLIP

  • Jiacheng Zhang
  • Jinhao Li
  • Hanxun Huang
  • Sarah Monazam Erfani
  • Benjamin I. P. Rubinstein
  • Feng Liu

Recent studies have shown that CLIP model's adversarial robustness in zero-shot classification tasks can be enhanced by adversarially fine-tuning its image encoder with adversarial examples (AEs), which are generated by minimizing the cosine similarity between images and a hand-crafted template (e.g., ''A photo of a {label}''). However, it has been shown that the cosine similarity between a single image and a single hand-crafted template is insufficient to measure the similarity for image-text pairs. Building on this, in this paper, we find that the AEs generated using cosine similarity may fail to fool CLIP when the similarity metric is replaced with semantically enriched alternatives, making the image encoder fine-tuned with these AEs less robust. To overcome this issue, we first propose a semantic-ensemble attack to generate semantic-aware AEs by minimizing the average similarity between the original image and an ensemble of refined textual descriptions. These descriptions are initially generated by a foundation model to capture core semantic features beyond hand-crafted templates and are then refined to reduce hallucinations. To this end, we propose Semantic-aware Adversarial Fine-Tuning (SAFT), which fine-tunes CLIP's image encoder with semantic-aware AEs. Extensive experiments show that SAFT outperforms current methods, achieving substantial improvements in zero-shot adversarial robustness across 16 datasets. Our code is available at: https://github.com/tmlr-group/SAFT.

JBHI Journal 2025 Journal Article

A Hierarchical Graph Convolutional Network With Infomax-Guided Graph Embedding for Population-Based ASD Detection

  • Xiaoke Hao
  • Mingming Ma
  • Jiaqing Tao
  • Jiahui Cao
  • Jing Qin
  • Feng Liu
  • Daoqiang Zhang
  • Dong Ming

Recently, functional magnetic resonance imaging (fMRI)-based brain networks have been shown to be an effective diagnostic tool with great potential for accurately detecting autism spectrum disorders (ASD). Meanwhile, the successful use of graph convolution networks (GCNs) methods based on fMRI information has improved the classification accuracy of ASD. However, many graph convolution-based methods do not fully utilize the topological information of the brain functional connectivity network (BFCN) or ignore the effect of non-imaging information. Therefore, we propose a hierarchical graph embedding model that leverage both the topological information of the BFCN and the non-imaging information of the subjects to improve the classification accuracy. Specifically, our model first use the Infomax Module to automatically identify embedded features in regions of interests (ROIs) in the brain. Then, these features, along with non-imaging information, is used to construct a population graph model. Finally, we design a graph convolution framework to propagate and aggregate the node features and obtain the results for ASD detection. Our model takes into account both the significance of the BFCN to individual subjects and relationships between subjects in the population graph. The model performed autism detection using the Autism Brain Imaging Data Exchange (ABIDE) dataset and obtained an average accuracy of 77. 2% and an AUC of 87. 2%. These results exceed those of the baseline approach. Through extensive experiments, we demonstrate the competitiveness, robustness and effectiveness of our model in aiding ASD diagnosis.

YNIMG Journal 2025 Journal Article

Accelerating multi-directional diffusion MRI through patch-based joint reconstruction

  • Zhongbiao Xu
  • Rongli Zhang
  • Wei Huang
  • Guanhua Deng
  • Xiaoyun Liang
  • Li Guo
  • Junying Cheng
  • Yaohui Wang

Diffusion magnetic resonance imaging (dMRI) is a valuable technique for studying tissue microstructure and connectivity in the brain. However, acquiring high-resolution dMRI data is time-consuming, limiting its clinical applicability. Traditional parallel imaging techniques can accelerate the acquisition of dMRI, but they are constrained by the geometry factor. In this study, we propose a novel patch-based multiple diffusion directions joint reconstruction method that simultaneously capitalizes on the intra- and inter-image correlation across multiple diffusion directions by grouping similar 3D image patches and then enforces the sparsity of these groups in sensitivity encoding (SENSE) reconstruction, termed PB-SENSE. The simulation and in vivo experiments demonstrated that the proposed method can achieve high-quality images comparable to those obtained from fully sampled data, even with an acceleration of 5. This suggests that the proposed method has the potential to enhance the practical application of high-resolution diffusion imaging.

NeurIPS Conference 2025 Conference Paper

Anchor-based Maximum Discrepancy for Relative Similarity Testing

  • Zhijian Zhou
  • Liuhua Peng
  • Xunye Tian
  • Feng Liu

The relative similarity testing aims to determine which of the distributions, $P$ or $Q$, is closer to an anchor distribution $U$. Existing kernel-based approaches often test the relative similarity with a fixed kernel in a manually specified alternative hypothesis, e. g. , $Q$ is closer to $U$ than $P$. Although kernel selection is known to be important to kernel-based testing methods, the manually specified hypothesis poses a significant challenge for kernel selection in relative similarity testing: Once the hypothesis is specified first, we can always find a kernel such that the hypothesis is rejected. This challenge makes relative similarity testing ill-defined when we want to select a good kernel after the hypothesis is specified. In this paper, we cope with this challenge via learning a proper hypothesis and a kernel simultaneously, instead of learning a kernel after manually specifying the hypothesis. We propose an anchor-based maximum discrepancy (AMD), which defines the relative similarity as the maximum discrepancy between the distances of $(U, P)$ and $(U, Q)$ in a space of deep kernels. Based on AMD, our testing incorporates two phases. In Phase I, we estimate the AMD over the deep kernel space and infer the potential hypothesis. In Phase II, we assess the statistical significance of the potential hypothesis, where we propose a unified testing framework to derive thresholds for tests over different possible hypotheses from Phase I. Lastly, we validate our method theoretically and demonstrate its effectiveness via extensive experiments on benchmark datasets. Codes are publicly available at: https: //github. com/tmlr-group/AMD.

ICML Conference 2025 Conference Paper

CALM: Consensus-Aware Localized Merging for Multi-Task Learning

  • Kunda Yan
  • Min Zhang 0068
  • Sen Cui
  • Zikun Qu
  • Bo Jiang 0016
  • Feng Liu
  • Changshui Zhang

Model merging aims to integrate the strengths of multiple fine-tuned models into a unified model while preserving task-specific capabilities. Existing methods, represented by task arithmetic, are typically classified into global- and local-aware methods. However, global-aware methods inevitably cause parameter interference, while local-aware methods struggle to maintain the effectiveness of task-specific details in the merged model. To address these limitations, we propose a Consensus Aware Localized Merging (CALM) method which incorporates localized information aligned with global task consensus, ensuring its effectiveness post-merging. CALM consists of three key components: (1) class-balanced entropy minimization sampling, providing a more flexible and reliable way to leverage unsupervised data; (2) an efficient-aware framework, selecting a small set of tasks for sequential merging with high scalability; (3) a consensus-aware mask optimization, aligning localized binary masks with global task consensus and merging them conflict-free. Experiments demonstrate the superiority and robustness of our CALM, significantly outperforming existing methods and achieving performance close to traditional MTL.

ICML Conference 2025 Conference Paper

Convergence of Mean-Field Langevin Stochastic Descent-Ascent for Distributional Minimax Optimization

  • Zhangyi Liu
  • Feng Liu
  • Rui Gao
  • Shuang Li

We study convergence properties of the discrete-time Mean-Field Langevin Stochastic Gradient Descent-Ascent (MFL-SGDA) algorithm for solving distributional minimax optimization. These problems arise in various applications, such as zero-sum games, generative adversarial networks and distributionally robust learning. Despite the significance of MFL-SGDA in these contexts, the discrete-time convergence rate remains underexplored. To address this gap, we establish a last-iterate convergence rate of $O(\frac{1}{\epsilon}\log\frac{1}{\epsilon})$ for MFL-SGDA. This rate is nearly optimal when compared to the complexity lower bound of its Euclidean counterpart. This rate also matches the complexity of mean-field Langevin stochastic gradient descent for distributional minimization and the outer-loop iteration complexity of an existing double-loop algorithm for distributional minimax problems. By leveraging an elementary analysis framework that avoids PDE-based techniques, we overcome previous limitations and achieve a faster convergence rate.

NeurIPS Conference 2025 Conference Paper

DUAL: Learning Diverse Kernels for Aggregated Two-sample and Independence Testing

  • Zhijian Zhou
  • Xunye Tian
  • Liuhua Peng
  • Chao Lei
  • Antonin Schrab
  • Danica J. Sutherland
  • Feng Liu

To adapt kernel two-sample and independence testing to complex structured data, aggregation of multiple kernels is frequently employed to boost testing power compared to single-kernel tests. However, we observe a phenomenon that directly maximizing multiple kernel-based statistics may result in highly similar kernels that capture highly overlapping information, limiting the effectiveness of aggregation. To address this, we propose an aggregated statistic that explicitly incorporates kernel diversity based on the covariance between different kernels. Moreover, we identify a fundamental challenge: a trade-off between the diversity among kernels and the test power of individual kernels, i. e. , the selected kernels should be both effective and diverse. This motivates a testing framework with selection inference, which leverages information from the training phase to select kernels with strong individual performance from the learned diverse kernel pool. We provide rigorous theoretical statements and proofs to show the consistency on the test power and control of Type-I error, along with asymptotic analysis of the proposed statistics. Lastly, we conducted extensive empirical experiments demonstrating the superior performance of our proposed approach across various benchmarks for both two-sample and independence testing.

IJCAI Conference 2025 Conference Paper

DualCast: A Model to Disentangle Aperiodic Events from Traffic Series

  • Xinyu Su
  • Feng Liu
  • Yanchuan Chang
  • Egemen Tanin
  • Majid Sarvi
  • Jianzhong Qi

Traffic forecasting is crucial for transportation systems optimisation. Current models minimise the mean forecasting errors, often favouring periodic events prevalent in the training data, while overlooking critical aperiodic ones like traffic incidents. To address this, we propose DualCast, a dual-branch framework that disentangles traffic signals into intrinsic spatial-temporal patterns and external environmental contexts, including aperiodic events. DualCast also employs a cross-time attention mechanism to capture high-order spatial-temporal relationships from both periodic and aperiodic patterns. DualCast is versatile. We integrate it with recent traffic forecasting models, consistently reducing their forecasting errors by up to 9. 6% on multiple real datasets.

NeurIPS Conference 2025 Conference Paper

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Videos Generation

  • Xiaofeng Wang
  • Kang Zhao
  • Feng Liu
  • Jiayu Wang
  • Guosheng Zhao
  • Xiaoyi Bao
  • Zheng Zhu
  • Yingya Zhang

Video generation has emerged as a promising tool for world simulation, leveraging visual data to replicate real-world environments. Within this context, egocentric video generation, which centers on the human perspective, holds significant potential for enhancing applications in virtual reality, augmented reality, and gaming. However, the generation of egocentric videos presents substantial challenges due to the dynamic nature of first-person viewpoints, the intricate diversity of actions, and the complex variety of scenes encountered. Existing datasets are inadequate for addressing these challenges effectively. To bridge this gap, we present EgoVid-5M, the first high-quality dataset specifically curated for egocentric video generation. EgoVid-5M encompasses over 5 million egocentric video clips and is enriched with detailed action annotations, including fine-grained kinematic control and high-level textual descriptions. To ensure the integrity and usability of the dataset, we implement a sophisticated data cleansing pipeline designed to maintain frame consistency, action coherence, and motion smoothness under egocentric conditions. Furthermore, we introduce EgoDreamer, which is capable of generating egocentric videos driven simultaneously by action descriptions and kinematic control signals. The EgoVid-5M dataset, associated action annotations, and all data cleansing metadata will be released for the advancement of research in egocentric video generation.

TMLR Journal 2025 Journal Article

Exploring Weak-to-Strong Generalization for CLIP-based Classification

  • Jinhao Li
  • Sarah Monazam Erfani
  • Lei Feng
  • James Bailey
  • Feng Liu

Aligning large-scale commercial models with user intent is crucial to preventing harmful outputs. Current methods rely on human supervision but become impractical as model complexity increases. When models surpass human knowledge, providing accurate feedback becomes challenging and inefficient. A novel solution proposed recently is using a weaker model to supervise a stronger model. This concept leverages the ability of weaker models to perform evaluations, thereby reducing the workload on human supervisors. Previous work has shown the effectiveness of weak-to-strong generalization in the context of language-only models. Extending this concept to vision-language models leverages these insights, adapting the proven benefits to a multi-modal context. In our study, we explore weak-to-strong generalization for CLIP-based classification. We propose a method, \emph{class prototype learning} (CPL), which aims to enhance the classification capabilities of the CLIP model, by learning more representative prototypes for each category. Our findings indicate that, despite using a simple loss function under weak supervision, CPL yields robust improvements in targeted scenarios, particularly when pretraining is limited. Extensive experiments demonstrate that our approach is effective under these settings, achieving a 3.67\% improvement over strong baseline methods.

AAAI Conference 2025 Conference Paper

FloNa: Floor Plan Guided Embodied Visual Navigation

  • Jiaxin Li
  • Weiqi Huang
  • Zan Wang
  • Wei Liang
  • Huijun Di
  • Feng Liu

Humans naturally rely on floor plans to navigate in unfamiliar environments, as they are readily available, reliable, and provide rich geometrical guidance. However, existing visual navigation settings overlook this valuable prior knowledge, leading to limited efficiency and accuracy. To eliminate this gap, we introduce a novel navigation task: Floor Plan Visual Navigation (FloNa), the first attempt to incorporate floor plans into embodied visual navigation. While the floor plan offers significant advantages, two key challenges emerge: (1) handling the spatial inconsistency between the floor plan and the actual scene layout for collision-free navigation, and (2) aligning observed images with the floor plan sketch despite their distinct modalities. To address these challenges, we propose FloDiff, a novel diffusion policy framework incorporating a localization module to facilitate alignment between the current observation and the floor plan. We further collect 20k navigation episodes across 117 scenes in the iGibson simulator to support the training and evaluation. Extensive experiments demonstrate the effectiveness and efficiency of our framework in unfamiliar scenes using floor plan knowledge.

IJCAI Conference 2025 Conference Paper

Flow Matching Based Sequential Recommender Model

  • Feng Liu
  • Lixin Zou
  • Xiangyu Zhao
  • Min Tang
  • Liming Dong
  • Dan Luo
  • Xiangyang Luo
  • Chenliang Li

Generative models, particularly diffusion model, have emerged as powerful tools for sequential recommendation. However, accurately modeling user preferences remains challenging due to the noise perturbations inherent in the forward and reverse processes of diffusion-based methods. Towards this end, this study introduces FMRec, a Flow Matching based model that employs a straight flow trajectory and a modified loss tailored for the recommendation task. Additionally, from the diffusion-model perspective, we integrate a reconstruction loss to improve robustness against noise perturbations, thereby retaining user preferences during the forward process. In the reverse process, we employ a deterministic reverse sampler, specifically an ODE-based updating function, to eliminate unnecessary randomness, thereby ensuring that the generated recommendations closely align with user needs. Extensive evaluations on four benchmark datasets reveal that FMRec achieves an average improvement of 6. 53% over state-of-the-art methods. The replication code is available at https: //github. com/FengLiu-1/FMRec.

NeurIPS Conference 2025 Conference Paper

Generative Model Inversion Through the Lens of the Manifold Hypothesis

  • Xiong Peng
  • Bo Han
  • Fengfei Yu
  • Tongliang Liu
  • Feng Liu
  • Mingyuan Zhou

Model inversion attacks (MIAs) aim to reconstruct class-representative samples from trained models. Recent generative MIAs utilize generative adversarial networks to learn image priors that guide the inversion process, yielding reconstructions with high visual quality and strong fidelity to the private data. To explore the reason behind their effectiveness, we begin by examining the gradients of inversion loss w. r. t. synthetic inputs, and find that these gradients are surprisingly noisy. Further analysis shows that generative model inversion approaches implicitly denoise the gradients by projecting them onto the tangent space of the generator manifold—filtering out directions that deviate from the manifold structure while preserving informative components aligned with it. Our empirical measurements show that, in models trained with standard supervision, loss gradients exhibit large angular deviations from the data manifold, indicating poor alignment with class-relevant directions. This observation motivates our central hypothesis: models become more vulnerable to MIAs when their loss gradients align more closely with the generator manifold. We validate this hypothesis by designing a novel training objective that explicitly promotes such alignment. Building on this insight, we further introduce a training-free approach to enhance gradient–manifold alignment during inversion, leading to consistent improvements over state-of-the-art generative MIAs.

NeurIPS Conference 2025 Conference Paper

Inference of Whole Brain Electrophysiological Networks Through Multimodal Integration of Simultaneous Scalp and Intracranial EEG

  • Shihao Yang
  • Feng Liu

Brain imaging research has transitioned over the past decades from identifying isolated regions of task-evoked activation to characterizing the spatiotemporal dynamics of large-scale brain networks. Electrophysiological signals are the direct manifestation of brain activity; thus, characterizing whole-brain electrophysiological networks (WBEN) can serve as a fundamental tool for neuroscience studies and clinical applications. In this work, we introduce a framework for integrating scalp EEG and intracranial EEG (iEEG) for WBEN estimation through a principled state-space modeling approach, where an Expectation-Maximization (EM) algorithm is designed to infer the state variables and brain connectivity simultaneously. We validated the proposed method on synthetic data, and the results revealed improved performance compared to traditional two-step methods using scalp EEG only, demonstrating the importance of including iEEG signals for WBEN estimation. For real data with simultaneous EEG and iEEG, we applied the developed framework to understand the information flows during encoding and maintenance phases of a working memory task. The information flows between subcortical and cortical regions are delineated, highlighting more significant information flows from cortical to subcortical regions during encoding than during maintenance. The results are consistent with previous research findings, but from a whole-brain perspective, which underscores the unique utility of the proposed framework.

NeurIPS Conference 2025 Conference Paper

Long-tailed Recognition with Model Rebalancing

  • JIAAN LUO
  • Feng Hong
  • Qiang Hu
  • Xiaofeng Cao
  • Feng Liu
  • Jiangchao Yao

Long-tailed recognition is ubiquitous and challenging in deep learning and even in the downstream finetuning of foundation models, since the skew class distribution generally prevents the model generalization to the tail classes. Despite the promise of previous methods from the perspectives of data augmentation, loss rebalancing and decoupled training etc. , consistent improvement in the broad scenarios like multi-label long-tailed recognition is difficult. In this study, we dive into the essential model capacity impact under long-tailed context, and propose a novel framework, Model Rebalancing (MORE), which mitigates imbalance by directly rebalancing the model's parameter space. Specifically, MORE introduces a low-rank parameter component to mediate the parameter space allocation guided by a tailored loss and sinusoidal reweighting schedule, but without increasing the overall model complexity or inference costs. Extensive experiments on diverse long-tailed benchmarks, spanning multi-class and multi-label tasks, demonstrate that MORE significantly improves generalization, particularly for tail classes, and effectively complements existing imbalance mitigation methods. These results highlight MORE's potential as a robust plug-and-play module in long-tailed settings.

ICLR Conference 2025 Conference Paper

Lossy Compression with Pretrained Diffusion Models

  • Jeremy Vonderfecht
  • Feng Liu

We apply Theis et al. (2022)'s DiffC algorithm to Stable Diffusion 1.5, 2.1, XL, and and Flux-dev, and demonstrate that these pretrained models are remarkably capable lossy image compressors. A principled algorithm for compression using pretrained diffusion models has been understood since at least 2020 (Ho et al.), but challenges in reverse-channel coding have prevented such algorithms from ever being fully implemented. We introduce simple workarounds that lead to the first complete implementation of DiffC, which is capable of compressing and decompressing images using Stable Diffusion in under 10 seconds. Despite requiring no additional training, our method is competitive with other state-of-the-art generative compression methods at low ultra-low bitrates.

IJCAI Conference 2025 Conference Paper

Multi-view Clustering via Multi-granularity Ensemble

  • Jie Yang
  • Wei Chen
  • Feng Liu
  • Peng Zhou
  • Zhongli Wang
  • Xinyan Liang
  • Bingbing Jiang

Multi-view clustering aims to integrate complementary information from multiple views to improve clustering performance. However, existing ensemble-based methods suffer from information loss due to their reliance on single-granularity labels, limiting the discriminative capability of learned representations. Meanwhile, representation and graph fusion-based approaches face challenges such as explicit view alignment and manual weight tuning, making them less effective for heterogeneous views with varying data distributions. To address these limitations, we propose a novel multi-view clustering framework via Multi-granularity Ensemble (MGE), fully using the multi-granularity information across diverse views for accurate and consistent clustering. Specifically, MGE first modifies the hierarchical clustering and then leverages it on each view (including the fused view) to achieve multi-granularity labels. Moreover, the cross-view and cross-granularity fusion strategy is designed to learn a robust co-association similarity matrix, which effectively preserves the fine-grained and coarse-grained structures of multi-view data and facilitates subsequent clustering. Therefore, MGE can provide a comprehensive representation of local and global patterns within data, eliminating the requirement for view alignment and weight tuning. Experiments demonstrate that MGE consistently outperforms state-of-the-art methods across multiple datasets, validating its effectiveness and superiority in handling heterogeneous views.

NeurIPS Conference 2025 Conference Paper

Physics-Driven Spatiotemporal Modeling for AI-Generated Video Detection

  • Shuhai Zhang
  • ZiHao Lian
  • Jiahao Yang
  • Daiyuan Li
  • Guoxuan Pang
  • Feng Liu
  • Bo Han
  • Shutao Li

AI-generated videos have achieved near-perfect visual realism (e. g. , Sora), urgently necessitating reliable detection mechanisms. However, detecting such videos faces significant challenges in modeling high-dimensional spatiotemporal dynamics and identifying subtle anomalies that violate physical laws. In this paper, we propose a physics-driven AI-generated video detection paradigm based on probability flow conservation principles. Specifically, we propose a statistic called Normalized Spatiotemporal Gradient (NSG), which quantifies the ratio of spatial probability gradients to temporal density changes, explicitly capturing deviations from natural video dynamics. Leveraging pre-trained diffusion models, we develop an NSG estimator through spatial gradients approximation and motion-aware temporal modeling without complex motion decomposition while preserving physical constraints. Building on this, we propose an NSG-based video detection method (NSG-VD) that computes the Maximum Mean Discrepancy (MMD) between NSG features of the test and real videos as a detection metric. Last, we derive an upper bound of NSG feature distances between real and generated videos, proving that generated videos exhibit amplified discrepancies due to distributional shifts. Extensive experiments confirm that NSG-VD outperforms state-of-the-art baselines by 16. 00\% in Recall and 10. 75\% in F1-Score, validating the superior performance of NSG-VD. The source code is available at \url{https: //github. com/ZSHsh98/NSG-VD}.

NeurIPS Conference 2025 Conference Paper

Practical Kernel Selection for Kernel-based Conditional Independence Test

  • Wenjie Wang
  • Mingming Gong
  • Biwei Huang
  • James Bailey
  • Bo Han
  • Kun Zhang
  • Feng Liu

Conditional independence (CI) testing is a fundamental yet challenging task in modern statistics and machine learning. One pivotal class of methods for assessing conditional independence encompasses kernel-based approaches, known for assessing CI by detecting general conditional dependence without imposing strict assumptions on relationships or data distributions. As with any method utilizing kernels, selecting appropriate kernels is crucial for precise identification. However, it remains underexplored in kernel-based CI methods, where the kernels are often determined manually or heuristically. In this paper, we analyze and propose a kernel parameter selection approach for the kernel-based conditional independence test (KCI). The kernel parameters are selected based on the ratio of the statistic to the asymptotic variance, which approximates the test power for the given parameters at large sample sizes. The search procedure is grid-based, allowing for parallelization with manageable additional computation time. We theoretically demonstrate the consistency of the proposed criterion and conduct extensive experiments on both synthetic and real data to show the effectiveness of our method.

AAAI Conference 2025 Conference Paper

Privacy-Preserving Low-Rank Adaptation Against Membership Inference Attacks for Latent Diffusion Models

  • Zihao Luo
  • Xilie Xu
  • Feng Liu
  • Yun Sing Koh
  • Di Wang
  • Jingfeng Zhang

Low-rank adaptation (LoRA) is an efficient strategy for adapting latent diffusion models (LDMs) on a private dataset to generate specific images by minimizing the adaptation loss. However, the LoRA-adapted LDMs are vulnerable to membership inference (MI) attacks that can judge whether a particular data point belongs to the private dataset, thus leading to the privacy leakage. To defend against MI attacks, we first propose a straightforward solution: Membership-Privacy-preserving LoRA (MP-LoRA). MP-LoRA is formulated as a min-max optimization problem where a proxy attack model is trained by maximizing its MI gain while the LDM is adapted by minimizing the sum of the adaptation loss and the MI gain of the proxy attack model. However, we empirically find that MP-LoRA has the issue of unstable optimization, and theoretically analyze that the potential reason is the unconstrained local smoothness, which impedes the privacy-preserving adaptation. To mitigate this issue, we further propose a Stable Membership-Privacy-preserving LoRA (SMP-LoRA) that adapts the LDM by minimizing the ratio of the adaptation loss to the MI gain. Besides, we theoretically prove that the local smoothness of SMP-LoRA can be constrained by the gradient norm, leading to improved convergence. Our experimental results corroborate that SMP-LoRA can indeed defend against MI attacks and generate high-quality images.

NeurIPS Conference 2025 Conference Paper

Revealing Multimodal Causality with Large Language Models

  • Jin Li
  • Shoujin Wang
  • Qi Zhang
  • Feng Liu
  • Tongliang Liu
  • Longbing Cao
  • Shui Yu
  • Fang Chen

Uncovering cause-and-effect mechanisms from data is fundamental to scientific progress. While large language models (LLMs) show promise for enhancing causal discovery (CD) from unstructured data, their application to the increasingly prevalent multimodal setting remains a critical challenge. Even with the advent of multimodal LLMs (MLLMs), their efficacy in multimodal CD is hindered by two primary limitations: (1) difficulty in exploring intra- and inter-modal interactions for comprehensive causal variable identification; and (2) insufficiency to handle structural ambiguities with purely observational data. To address these challenges, we propose MLLM-CD, a novel framework for multimodal causal discovery from unstructured data. It consists of three key components: (1) a novel contrastive factor discovery module to identify genuine multimodal factors based on the interactions explored from contrastive sample pairs; (2) a statistical causal structure discovery module to infer causal relationships among discovered factors; and (3) an iterative multimodal counterfactual reasoning module to refine the discovery outcomes iteratively by incorporating the world knowledge and reasoning capabilities of MLLMs. Extensive experiments on both synthetic and real-world datasets demonstrate the effectiveness of the proposed MLLM-CD in revealing genuine factors and causal relationships among them from multimodal unstructured data. The implementation code and data are available at https: //github. com/JinLi-i/MLLM-CD.

ICML Conference 2025 Conference Paper

Sample-specific Noise Injection for Diffusion-based Adversarial Purification

  • Yuhao Sun
  • Jiacheng Zhang
  • Zesheng Ye
  • Chaowei Xiao
  • Feng Liu

Diffusion-based purification (DBP) methods aim to remove adversarial noise from the input sample by first injecting Gaussian noise through a forward diffusion process, and then recovering the clean example through a reverse generative process. In the above process, how much Gaussian noise is injected to the input sample is key to the success of DBP methods, which is controlled by a constant noise level $t*$ for all samples in existing methods. In this paper, we discover that an optimal $t*$ for each sample indeed could be different. Intuitively, the cleaner a sample is, the less the noise it should be injected, and vice versa. Motivated by this finding, we propose a new framework, called Sample-specific Score-aware Noise Injection (SSNI). Specifically, SSNI uses a pre-trained score network to estimate how much a data point deviates from the clean data distribution (i. e. , score norms). Then, based on the magnitude of score norms, SSNI applies a reweighting function to adaptively adjust $t*$ for each sample, achieving sample-specific noise injections. Empirically, incorporating our framework with existing DBP methods results in a notable improvement in both accuracy and robustness on CIFAR-10 and ImageNet-1K, highlighting the necessity to allocate distinct noise levels to different samples in DBP methods. Our code is available at: https: //github. com/tmlr-group/SSNI.

ICLR Conference 2025 Conference Paper

Score-based free-form architectures for high-dimensional Fokker-Planck equations

  • Feng Liu
  • Faguo Wu
  • Xiao Zhang 0004

Deep learning methods incorporate PDE residuals as the loss function for solving Fokker-Planck equations, and usually impose the proper normalization condition to avoid a trivial solution. However, soft constraints require careful balancing of multi-objective loss functions, and specific network architectures may limit representation capacity under hard constraints. In this paper, we propose a novel framework: Fokker-Planck neural network (FPNN) that adopts a score PDE loss to decouple the score learning and the density normalization into two stages. Our method allows free-form network architectures to model the unnormalized density and strictly satisfy normalization constraints by post-processing. We demonstrate the effectiveness on various high-dimensional steady-state Fokker-Planck (SFP) equations, achieving superior accuracy and over a 20$\times$ speedup compared to state-of-the-art methods. Without any labeled data, FPNNs achieve the mean absolute percentage error (MAPE) of 11.36%, 13.87% and 12.72% for 4D Ring, 6D Unimodal and 6D Multi-modal problems respectively, requiring only 256, 980, and 980 parameters. Experimental results highlights the potential as a universal fast solver for handling more than 20-dimensional SFP equations, with great gains in efficiency, accuracy, memory and computational resource usage.

JBHI Journal 2025 Journal Article

Spatial Craving Patterns in Marijuana Users: Insights From fMRI Brain Connectivity Analysis With High-Order Graph Attention Neural Networks

  • Jun-En Ding
  • Shihao Yang
  • Anna Zilverstand
  • Kaustubh R. Kulkarni
  • Xiaosi Gu
  • Feng Liu

The excessive consumption of marijuana can induce substantial psychological and social consequences. In this investigation, we propose an elucidative framework termed high-order graph attention neural networks (HOGANN) for the classification of Marijuana addiction, coupled with an analysis of localized brain network communities exhibiting abnormal activities among chronic marijuana users. HOGANN integrates dynamic intrinsic functional brain networks, estimated from functional magnetic resonance imaging (fMRI), using graph attention-based long short-term memory (GAT-LSTM) to capture temporal network dynamics. We employ a high-order attention module for information fusion and message passing among neighboring nodes, enhancing the network community analysis. Our model is validated across two distinct data cohorts, yielding substantially higher classification accuracy than benchmark algorithms. Furthermore, we discern the most pertinent subnetworks and cognitive regions affected by persistent marijuana consumption, indicating adverse effects on functional brain networks, particularly within the dorsal attention and frontoparietal networks. Intriguingly, our model demonstrates superior performance in cohorts exhibiting prolonged dependence, implying that prolonged marijuana usage induces more pronounced alterations in brain networks. The model proficiently identifies craving brain maps, thereby delineating critical brain regions for analysis.

NeurIPS Conference 2025 Conference Paper

Towards Accurate Time Series Forecasting via Implicit Decoding

  • Xinyu Li
  • Yuchen Luo
  • Hao Wang
  • Haoxuan Li
  • Liuhua Peng
  • Feng Liu
  • Yandong Guo
  • Kun Zhang

Recent booming time series models have demonstrated remarkable forecasting performance. However, these methods often place greater focus on more effectively modelling the historical series, largely neglecting the forecasting phase, which generates long-term forecasts by separately predicting multiple time points. Given that real-world time series typically consist of various long short-term dynamics, independent predictions over individual time points may fail to express complex underlying patterns and can lead to a lack of global views. To address these issues, this work explores new perspectives from the forecasting phase and proposes a novel Implicit Forecaster (IF) as an additional decoding module. Inspired by decomposition forecasting, IF adopts a more nuanced approach by implicitly predicting constituent waves represented by their frequency, amplitude, and phase, thereby accurately forming the time series. Extensive experimental results from multiple real-world datasets show that IF can consistently boost mainstream time series models, achieving state-of-the-art forecasting performance. Code is available at this repository: https: //github. com/rakuyorain/Implicit-Forecaster.

NeurIPS Conference 2024 Conference Paper

Bayesian-guided Label Mapping for Visual Reprogramming

  • Chengyi Cai
  • Zesheng Ye
  • Lei Feng
  • Jianzhong Qi
  • Feng Liu

Visual reprogramming (VR) leverages the intrinsic capabilities of pretrained vision models by adapting their input or output interfaces to solve downstream tasks whose labels (i. e. , downstream labels) might be totally different from the labels associated with the pretrained models (i. e. , pretrained labels). When adapting the output interface, label mapping methods transform the pretrained labels to downstream labels by establishing a gradient-free one-to-one correspondence between the two sets of labels. However, in this paper, we reveal that one-to-one mappings may overlook the complex relationship between pretrained and downstream labels. Motivated by this observation, we propose a ** B ayesian-guided L abel M apping (BLM) method. BLM constructs an iteratively-updated probabilistic label mapping matrix, with each element quantifying a pairwise relationship between pretrained and downstream labels. The assignment of values to the constructed matrix is guided by Bayesian conditional probability, considering the joint distribution of the downstream labels and the labels predicted by the pretrained model on downstream samples. Experiments conducted on both pretrained vision models (e. g. , ResNeXt) and vision-language models (e. g. , CLIP) demonstrate the superior performance of BLM over existing label mapping methods. The success of BLM also offers a probabilistic lens through which to understand and analyze the effectiveness of VR. Our code is available at https: //github. com/tmlr-group/BayesianLM.

YNICL Journal 2024 Journal Article

Functional magnetic resonance imaging alternations in suicide attempts individuals and their association with gene expression

  • Yurong Jiang
  • Yujing Zhou
  • Yingying Xie
  • Junzi Zhou
  • Mengjing Cai
  • Jie Tang
  • Feng Liu
  • Juanwei Ma

BACKGROUND: Functional Magnetic Resonance Imaging (fMRI) has shown brain activity alterations in individuals with a history of attempted suicide (SA) who are diagnosed with depression disorder (DD) or bipolar disorder (BD). However, patterns of spontaneous brain activity and their genetic correlations need further investigation. METHODS: A voxel-based meta-analysis of 19 studies including 26 datasets, involving 742 patients with a history of SA and 978 controls (both nonsuicidal patients and healthy controls) was conducted. We examined fMRI changes in SA patients and analyzed the association between these changes and gene expression profiles using data from the Allen Human Brain Atlas by partial least squares regression analysis. RESULTS: SA patients demonstrated increased spontaneous brain activity in several brain regions including the bilateral inferior temporal gyrus, hippocampus, fusiform gyrus, and right insula, and decreased activity in areas like the bilateral paracentral lobule and inferior frontal gyrus. Additionally, 5,077 genes were identified, exhibiting expression patterns associated with SA-related fMRI alterations. Functional enrichment analyses demonstrated that these SA-related genes were enriched for biological functions including glutamatergic synapse and mitochondrial structure. Concurrently, specific expression analyses showed that these genes were specifically expressed in the brain tissue, in neurons cells, and during early developmental periods. CONCLUSION: Our findings suggest a neurobiological basis for fMRI abnormalities in SA patients with DD or BD, potentially guiding future genetic and therapeutic research.

YNICL Journal 2024 Journal Article

Genetic and vascular risk factors for ischemic stroke and cortical morphometry in individuals without a history of stroke: A UK Biobank observational cohort study

  • Jiawei Liu
  • Yingying Xie
  • Feng Liu
  • Wen Qin
  • Chunshui Yu

BACKGROUND: Stroke risk factors may contribute to cognitive decline and dementia by altering brain tissue integrity. If their effects on brain are nonnegligible, the target regions for stroke rehabilitation with brain stimulation identified by cross-sectional case-control studies may be biased due to the pre-existing brain differences caused by these risk factors. Here, we investigated the effects of stroke risk factors on cortical thickness (CT) and surface area (SA) in individuals without a history of stroke. METHODS: ), systolic blood pressure (SBP), diastolic blood pressure (DBP), glycated hemoglobin (HbA1c), triglycerides (TG), and low-density lipoprotein (LDL) on CT and SA of 62 cerebral regions. We excluded non-Caucasian participants and participants with missing data, unqualified brain images, or a history of stroke or any other brain diseases. We constructed a multivariate linear regression model for each phenotype to simultaneously test the effect of each factor and interaction between factors. The results were verified by sensitivity analyses of SDP or DBP input and adjusting for body-mass index, high-density lipoprotein cholesterol, or smoking and alcohol intake. By excluding participants with abnormal blood pressure, glucose, or lipid, we tested whether vascular risk factor within normal range also affected cortical phenotypes. To determine clinical relevance of our findings, we also investigated the effects of stroke risk factors and cortical phenotypes on cognitive decline assessed by fluid intelligence score (FIQ) and the mediation of cortical phenotype for the association between stroke risk factor and FIQ. RESULTS: and SBP with cognitive decline were mediated by CT phenotypes. CONCLUSIONS: Stroke risk factors have substantial effects on cortical morphometry and cognitive decline in middle-aged and older people, which should be considered in the prevention of dementia and in the identification of target regions for stroke rehabilitation with brain stimulation.

TMLR Journal 2024 Journal Article

HiFE: Hierarchical Feature Ensemble Framework for Few-shot Hypotheses Adaptation

  • Yongfeng Zhong
  • Haoang Chi
  • Feng Liu
  • Xiao-ming Wu
  • Bo Han

The process of transferring knowledge from a source domain to a target domain in the absence of source data constitutes a formidable obstacle within the field of source-free domain adaptation, often termed hypothesis adaptation. Conventional methodologies have been dependent on a robustly trained (strong) source hypothesis to encapsulate the knowledge pertinent to the source domain. However, this strong hypothesis is prone to overfitting the source domain, resulting in diminished generalization performance when applied to the target domain. To mitigate this issue, we advocate for the augmentation of transferable source knowledge via the integration of multiple (weak) source models that are underfitting. Furthermore, we propose a novel architectural framework, designated as the Hierarchical Feature Ensemble (HiFE) framework for Few-Shot Hypotheses Adaptation, which amalgamates features from both the strong and intentionally underfit source models. Empirical evidence from our experiments indicates that these weaker models, while not optimal within the source domain context, contribute to an enhanced generalization capacity of the resultant model for the target domain. Moreover, the HiFE framework we introduce demonstrates superior performance, surpassing other leading baselines across a spectrum of few-shot hypothesis adaptation scenarios.

YNIMG Journal 2024 Journal Article

Homotopic functional connectivity disruptions in schizophrenia and their associated gene expression

  • Mengjing Cai
  • Yuan Ji
  • Qiyu Zhao
  • Hui Xue
  • Zuhao Sun
  • He Wang
  • Yijing Zhang
  • Yayuan Chen

It has been revealed that abnormal voxel-mirrored homotopic connectivity (VMHC) is present in patients with schizophrenia, yet there are inconsistencies in the relevant findings. Moreover, little is known about their association with brain gene expression profiles. In this study, transcription-neuroimaging association analyses using gene expression data from Allen Human Brain Atlas and case-control VMHC differences from both the discovery (meta-analysis, including 9 studies with a total of 386 patients and 357 controls) and replication (separate group-level comparisons within two datasets, including a total of 258 patients and 287 controls) phases were performed to identify genes associated with VMHC alterations. Enrichment analyses were conducted to characterize the biological functions and specific expression of identified genes, and Neurosynth decoding analysis was performed to examine the correlation between cognitive-related processes and VMHC alterations in schizophrenia. In the discovery and replication phases, patients with schizophrenia exhibited consistent VMHC changes compared to controls, which were correlated with a series of cognitive-related processes; meta-regression analysis revealed that illness duration was negatively correlated with VMHC abnormalities in the cerebellum and postcentral/precentral gyrus. The abnormal VMHC patterns were stably correlated with 1287 genes enriched for fundamental biological processes like regulation of cell communication, nervous system development, and cell communication. In addition, these genes were overexpressed in astrocytes and immune cells, enriched in extensive cortical regions and wide developmental time windows. The present findings may contribute to a more comprehensive understanding of the molecular mechanisms underlying VMHC alterations in patients with schizophrenia.

NeurIPS Conference 2024 Conference Paper

In-N-Out: Lifting 2D Diffusion Prior for 3D Object Removal via Tuning-Free Latents Alignment

  • Dongting Hu
  • Huan Fu
  • Jiaxian Guo
  • Liuhua Peng
  • Tingjin Chu
  • Feng Liu
  • Tongliang Liu
  • Mingming Gong

Neural representations for 3D scenes have made substantial advancements recently, yet object removal remains a challenging yet practical issue, due to the absence of multi-view supervision over occluded areas. Diffusion Models (DMs), trained on extensive 2D images, show diverse and high-fidelity generative capabilities in the 2D domain. However, due to not being specifically trained on 3D data, their application to multi-view data often exacerbates inconsistency, hence impacting the overall quality of the 3D output. To address these issues, we introduce "In-N-Out", a novel approach that begins by inpainting a prior, i. e. , the occluded area from a single view using DMs, followed by outstretching it to create multi-view inpaintings via latents alignments. Our analysis identifies that the variability in DMs' outputs mainly arises from initially sampled latents and intermediate latents predicted in the denoising process. We explicitly align of initial latents using a Neural Radiance Field (NeRF) to establish a consistent foundational structure in the inpainted area, complemented by an implicit alignment of intermediate latents through cross-view attention during the denoising phases, enhancing appearance consistency across views. To further enhance rendering results, we apply a patch-based hybrid loss to optimize NeRF. We demonstrate that our techniques effectively mitigate the challenges posed by inconsistencies in DMs and substantially improve the fidelity and coherence of inpainted 3D representations.

ICRA Conference 2024 Conference Paper

Learning-Aided Control of Robotic Tether-Net with Maneuverable Nodes to Capture Large Space Debris

  • Achira Boonrath
  • Feng Liu
  • Eleonora M. Botta
  • Souma Chowdhury

Maneuverable tether-net systems launched from an unmanned spacecraft offer a promising solution for the active removal of large space debris. Guaranteeing the successful capture of such space debris is dependent on the ability to reliably maneuver the tether-net system – a flexible, many-DoF (thus complex) system – for a wide range of launch scenarios. Here, scenarios are defined by the relative location of the debris with respect to the chaser spacecraft. This paper represents and solves this problem as a hierarchically decentralized implementation of robotic trajectory planning and control and demonstrates the effectiveness of the approach when applied to two different tether-net systems, with 4 and 8 maneuverable units (MUs), respectively. Reinforcement learning (policy gradient) is used to design the centralized trajectory planner that, based on the relative location of the target debris at the launch of the net, computes the final aiming positions of each MU, from which their trajectory can be derived. Each MU then seeks to follow its assigned trajectory by using a decentralized PID controller that outputs the MU’s thrust vector and is informed by noisy sensor feedback (for realism) of its relative location. System performance is assessed in terms of capture success and overall fuel consumption by the MUs. Reward shaping and surrogate models are used to respectively guide and speed up the RL process. Simulation-based experiments show that this approach allows the successful capture of debris at fuel costs that are notably lower than nominal baselines, including in scenarios where the debris is significantly off-centered compared to the approaching chaser spacecraft.

NeurIPS Conference 2024 Conference Paper

Mind the Gap Between Prototypes and Images in Cross-domain Finetuning

  • Hongduan Tian
  • Feng Liu
  • Zhanke Zhou
  • Tongliang Liu
  • Chengqi Zhang
  • Bo Han

In cross-domain few-shot classification (CFC), recent works mainly focus on adapting a simple transformation head on top of a frozen pre-trained backbone with few labeled data to project embeddings into a task-specific metric space where classification can be performed by measuring similarities between image instance and prototype representations. Technically, an assumption implicitly adopted in such a framework is that the prototype and image instance embeddings share the same representation transformation. However, in this paper, we find that there naturally exists a gap, which resembles the modality gap, between the prototype and image instance embeddings extracted from the frozen pre-trained backbone, and simply applying the same transformation during the adaptation phase constrains exploring the optimal representation distributions and shrinks the gap between prototype and image representations. To solve this problem, we propose a simple yet effective method, contrastive prototype-image adaptation (CoPA), to adapt different transformations for prototypes and images similarly to CLIP by treating prototypes as text prompts. Extensive experiments on Meta-Dataset demonstrate that CoPA achieves the state-of-the-art performance more efficiently. Meanwhile, further analyses also indicate that CoPA can learn better representation clusters, enlarge the gap, and achieve the minimum validation loss at the enlarged gap.

JMLR Journal 2024 Journal Article

On the Learnability of Out-of-distribution Detection

  • Zhen Fang
  • Yixuan Li
  • Feng Liu
  • Bo Han
  • Jie Lu

Supervised learning aims to train a classifier under the assumption that training and test data are from the same distribution. To ease the above assumption, researchers have studied a more realistic setting: out-of-distribution (OOD) detection, where test data may come from classes that are unknown during training (i.e., OOD data). Due to the unavailability and diversity of OOD data, good generalization ability is crucial for effective OOD detection algorithms, and corresponding learning theory is still an open problem. To study the generalization of OOD detection, this paper investigates the probably approximately correct (PAC) learning theory of OOD detection that fits the commonly used evaluation metrics in the literature. First, we find a necessary condition for the learnability of OOD detection. Then, using this condition, we prove several impossibility theorems for the learnability of OOD detection under some scenarios. Although the impossibility theorems are frustrating, we find that some conditions of these impossibility theorems may not hold in some practical scenarios. Based on this observation, we next give several necessary and sufficient conditions to characterize the learnability of OOD detection in some practical scenarios. Lastly, we offer theoretical support for representative OOD detection works based on our OOD theory. [abs] [ pdf ][ bib ] &copy JMLR 2024. ( edit, beta )

TMLR Journal 2024 Journal Article

Predicting the Encoding Error of SIRENs

  • Jeremy Vonderfecht
  • Feng Liu

Implicit Neural Representations (INRs), which encode signals such as images, videos, and 3D shapes in the weights of neural networks, are becoming increasingly popular. Among their many applications is signal compression, for which there is great interest in achieving the highest possible fidelity to the original signal subject to constraints such as neural network size, training (encoding) and inference (decoding) time. But training INRs can be a computationally expensive process, making it challenging to determine the best possible tradeoff under such constraints. Towards this goal, we propose a novel problem: predicting the encoding error (i.e. training loss) that an INR will reach on a given training signal. We present a method which predicts the encoding error that a popular INR network (SIREN) will reach, given its network hyperparameters and the signal to encode. This method is trained on a unique dataset of 300,000 SIRENs, trained across a variety of images and hyperparameters. Our predictive method demonstrates the feasibility of this regression problem, and allows users to anticipate the encoding error that a SIREN network will reach in milliseconds instead of minutes or longer. We also provide insights into the behavior of SIREN networks, such as why narrow SIRENs can have very high random variation in encoding error, and how the performance of SIRENs relates to JPEG compression.

NeurIPS Conference 2024 Conference Paper

Pseudo-Private Data Guided Model Inversion Attacks

  • Xiong Peng
  • Bo Han
  • Feng Liu
  • Tongliang Liu
  • Mingyuan Zhou

In model inversion attacks (MIAs), adversaries attempt to recover private training data by exploiting access to a well-trained target model. Recent advancements have improved MIA performance using a two-stage generative framework. This approach first employs a generative adversarial network to learn a fixed distributional prior, which is then used to guide the inversion process during the attack. However, in this paper, we observed a phenomenon that such a fixed prior would lead to a low probability of sampling actual private data during the inversion process due to the inherent distribution gap between the prior distribution and the private data distribution, thereby constraining attack performance. To address this limitation, we propose increasing the density around high-quality pseudo-private data—recovered samples through model inversion that exhibit characteristics of the private training data—by slightly tuning the generator. This strategy effectively increases the probability of sampling actual private data that is close to these pseudo-private data during the inversion process. After integrating our method, the generative model inversion pipeline is strengthened, leading to improvements over state-of-the-art MIAs. This paves the way for new research directions in generative MIAs.

YNIMG Journal 2024 Journal Article

Quantitative susceptibility mapping through model-based deep image prior (MoDIP)

  • Zhuang Xiong
  • Yang Gao
  • Yin Liu
  • Amir Fazlollahi
  • Peter Nestor
  • Feng Liu
  • Hongfu Sun

The data-driven approach of supervised learning methods has limited applicability in solving dipole inversion in Quantitative Susceptibility Mapping (QSM) with varying scan parameters across different objects. To address this generalization issue in supervised QSM methods, we propose a novel training-free model-based unsupervised method called MoDIP (Model-based Deep Image Prior). MoDIP comprises a small, untrained network and a Data Fidelity Optimization (DFO) module. The network converges to an interim state, acting as an implicit prior for image regularization, while the optimization process enforces the physical model of QSM dipole inversion. Experimental results demonstrate MoDIP's excellent generalizability in solving QSM dipole inversion across different scan parameters. It exhibits robustness against pathological brain QSM, achieving over 32 % accuracy improvement than supervised deep learning methods. It is also 33 % more computationally efficient and runs 4 times faster than conventional DIP-based approaches, enabling 3D high-resolution image reconstruction in under 4.5 min.

EAAI Journal 2024 Journal Article

Robust semi-supervised learning with reciprocal weighted mixing distribution alignment

  • Ziyu Cheng
  • Xianmin Wang
  • Jing Li
  • Feng Liu
  • Yutong Xie
  • Haiyan Liang

Recent semi-supervised learning(SSL) methods have achieved great success owing to the impressive performances brought by the combination of pseudo-labeling and consistency regularization. These methods often use pre-defined constant thresholds or dynamical thresholds to select unlabeled samples that contribute to training. However, many correct/incorrect pseudo-labels may be ignored/selected. Especially in distribution mismatched scenario, threshold-adjusted strategy is often complex and ineffective. To alleviate this issue, we develop a simple yet powerful framework whose idea is to abandon this strategy and utilize distribution alignment to adjust the predictions generated from a biased model softly. Specifically, first, we create two classifiers to predict pseudo-label(i. e. , the sample belongs to a specific category) and complementary pseudo-label(i. e. , the sample does not belong to a specific category), respectively. Second, by maintaining the distributions of pseudo-labels, complementary pseudo-labels and their reverse versions from past iterations, we enforce a reciprocal weighted mixing according to the predicted category weights. Third, a reciprocal distribution alignment is applied to the mixed distributions to adjust the predicted distributions. Finally, we propose Implication Alignment Loss, which keeps consistency between the predictions of the same implications but from different versions. We empirically demonstrate the effectiveness of our proposed method in comparison with state-of-the-art benchmarks. Especially, our method achieves a 1. 18% error rate reduction over the latest state-of-the-art method MutexMatch on CIFAR-10 with 2 labels per class and exhibits robustness in the scenario of mismatched distribution.

NeurIPS Conference 2024 Conference Paper

SongCreator: Lyrics-based Universal Song Generation

  • Shun Lei
  • Yixuan Zhou
  • Boshi Tang
  • Max W. Lam
  • Feng Liu
  • Hangyu Liu
  • Jingcheng Wu
  • Shiyin Kang

Music is an integral part of human culture, embodying human intelligence and creativity, of which songs compose an essential part. While various aspects of song generation have been explored by previous works, such as singing voice, vocal composition and instrumental arrangement, etc. , generating songs with both vocals and accompaniment given lyrics remains a significant challenge, hindering the application of music generation models in the real world. In this light, we propose SongCreator, a song-generation system designed to tackle this challenge. The model features two novel designs: a meticulously designed dual-sequence language model (DSLM) to capture the information of vocals and accompaniment for song generation, and a series of attention mask strategies for DSLM, which allows our model to understand, generate and edit songs, making it suitable for various songrelated generation tasks by utilizing specific attention masks. Extensive experiments demonstrate the effectiveness of SongCreator by achieving state-of-the-art or competitive performances on all eight tasks. Notably, it surpasses previous works by a large margin in lyrics-to-song and lyrics-to-vocals. Additionally, it is able to independently control the acoustic conditions of the vocals and accompaniment in the generated song through different audio prompts, exhibiting its potential applicability. Our samples are available at https: //thuhcsi. github. io/SongCreator/.

NeurIPS Conference 2024 Conference Paper

Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?

  • Haoang Chi
  • He Li
  • Wenjing Yang
  • Feng Liu
  • Long Lan
  • Xiaoguang Ren
  • Tongliang Liu
  • Bo Han

Causal reasoning capability is critical in advancing large language models (LLMs) towards artificial general intelligence (AGI). While versatile LLMs appear to have demonstrated capabilities in understanding contextual causality and providing responses that obey the laws of causality, it remains unclear whether they perform genuine causal reasoning akin to humans. However, current evidence indicates the contrary. Specifically, LLMs are only capable of performing shallow (level-1) causal reasoning, primarily attributed to the causal knowledge embedded in their parameters, but they lack the capacity for genuine human-like (level-2) causal reasoning. To support this hypothesis, methodologically, we delve into the autoregression mechanism of transformer-based LLMs, revealing that it is not inherently causal. Empirically, we introduce a new causal Q&A benchmark named CausalProbe 2024, whose corpus is fresh and nearly unseen for the studied LLMs. Empirical results show a significant performance drop on CausalProbe 2024 compared to earlier benchmarks, indicating that LLMs primarily engage in level-1 causal reasoning. To bridge the gap towards level-2 causal reasoning, we draw inspiration from the fact that human reasoning is usually facilitated by general knowledge and intended goals. Inspired by this, we propose G$^2$-Reasoner, a LLM causal reasoning method that incorporates general knowledge and goal-oriented prompts into LLMs' causal reasoning processes. Experiments demonstrate that G$^2$-Reasoner significantly enhances LLMs' causal reasoning capability, particularly in fresh and fictitious contexts. This work sheds light on a new path for LLMs to advance towards genuine causal reasoning, going beyond level-1 and making strides towards level-2.

YNIMG Journal 2024 Journal Article

XDL-ESI: Electrophysiological Sources Imaging via explainable deep learning framework with validation on simultaneous EEG and iEEG

  • Meng Jiao
  • Xiaochen Xian
  • Boyu Wang
  • Yu Zhang
  • Shihao Yang
  • Spencer Chen
  • Hai Sun
  • Feng Liu

Electroencephalography (EEG) or Magnetoencephalography (MEG) source imaging aims to estimate the underlying activated brain sources to explain the observed EEG/MEG recordings. Solving the inverse problem of EEG/MEG Source Imaging (ESI) is challenging due to its ill-posed nature. To achieve a unique solution, it is essential to apply sophisticated regularization constraints to restrict the solution space. Traditionally, the design of regularization terms is based on assumptions about the spatiotemporal structure of the underlying source dynamics. In this paper, we propose a novel paradigm for ESI via an Explainable Deep Learning framework, termed as XDL-ESI, which connects the iterative optimization algorithm with deep learning architecture by unfolding the iterative updates with neural network modules. The proposed framework has the advantages of (1) establishing a data-driven approach to model the source solution structure instead of using hand-crafted regularization terms; (2) improving the robustness of source solutions by introducing a topological loss that leverages the geometric spatial information applying varying penalties on distinct localization errors; (3) improving the reconstruction efficiency and interpretability as it inherits the advantages from both the iterative optimization algorithms (interpretability) and deep learning approaches (function approximation). The proposed XDL-ESI framework provides an efficient, accurate, and interpretable paradigm to solve the ESI inverse problem with satisfactory performance in both simulated data and real clinical data. Specially, this approach is further validated using simultaneous EEG and intracranial EEG (iEEG).

YNIMG Journal 2023 Journal Article

Affine transformation edited and refined deep neural network for quantitative susceptibility mapping

  • Zhuang Xiong
  • Yang Gao
  • Feng Liu
  • Hongfu Sun

Deep neural networks have demonstrated great potential in solving dipole inversion for Quantitative Susceptibility Mapping (QSM). However, the performances of most existing deep learning methods drastically degrade with mismatched sequence parameters such as acquisition orientation and spatial resolution. We propose an end-to-end AFfine Transformation Edited and Refined (AFTER) deep neural network for QSM, which is robust against arbitrary acquisition orientation and spatial resolution up to 0.6 mm isotropic at the finest. The AFTER-QSM neural network starts with a forward affine transformation layer, followed by a Unet for dipole inversion, then an inverse affine transformation layer, followed by a Residual Dense Network (RDN) for QSM refinement. Simulation and in-vivo experiments demonstrated that the proposed AFTER-QSM network architecture had excellent generalizability. It can successfully reconstruct susceptibility maps from highly oblique and anisotropic scans, leading to the best image quality assessments in simulation tests and suppressed streaking artifacts and noise levels for in-vivo experiments compared with other methods. Furthermore, ablation studies showed that the RDN refinement network significantly reduced image blurring and susceptibility underestimation due to affine transformations. In addition, the AFTER-QSM network substantially shortened the reconstruction time from minutes using conventional methods to only a few seconds.

TMLR Journal 2023 Journal Article

Attacking Perceptual Similarity Metrics

  • Abhijay Ghildyal
  • Feng Liu

Perceptual similarity metrics have progressively become more correlated with human judgments on perceptual similarity; however, despite recent advances, the addition of an imperceptible distortion can still compromise these metrics. In our study, we systematically examine the robustness of these metrics to imperceptible adversarial perturbations. Following the two-alternative forced-choice experimental design with two distorted images and one reference image, we perturb the distorted image closer to the reference via an adversarial attack until the metric flips its judgment. We first show that all metrics in our study are susceptible to perturbations generated via common adversarial attacks such as FGSM, PGD, and the One-pixel attack. Next, we attack the widely adopted LPIPS metric using spatial-transformation-based adversarial perturbations (stAdv) in a white-box setting to craft adversarial examples that can effectively transfer to other similarity metrics in a black-box setting. We also combine the spatial attack stAdv with PGD ($\ell_\infty$-bounded) attack to increase transferability and use these adversarial examples to benchmark the robustness of both traditional and recently developed metrics. Our benchmark provides a good starting point for discussion and further research on the robustness of metrics to imperceptible adversarial perturbations.

NeurIPS Conference 2023 Conference Paper

Efficient Adversarial Contrastive Learning via Robustness-Aware Coreset Selection

  • Xilie Xu
  • Jingfeng Zhang
  • Feng Liu
  • Masashi Sugiyama
  • Mohan S. Kankanhalli

Adversarial contrastive learning (ACL) does not require expensive data annotations but outputs a robust representation that withstands adversarial attacks and also generalizes to a wide range of downstream tasks. However, ACL needs tremendous running time to generate the adversarial variants of all training data, which limits its scalability to large datasets. To speed up ACL, this paper proposes a robustness-aware coreset selection (RCS) method. RCS does not require label information and searches for an informative subset that minimizes a representational divergence, which is the distance of the representation between natural data and their virtual adversarial variants. The vanilla solution of RCS via traversing all possible subsets is computationally prohibitive. Therefore, we theoretically transform RCS into a surrogate problem of submodular maximization, of which the greedy search is an efficient solution with an optimality guarantee for the original problem. Empirically, our comprehensive results corroborate that RCS can speed up ACL by a large margin without significantly hurting the robustness transferability. Notably, to the best of our knowledge, we are the first to conduct ACL efficiently on the large-scale ImageNet-1K dataset to obtain an effective robust representation via RCS. Our source code is at https: //github. com/GodXuxilie/Efficient ACL via_RCS.

AAAI Conference 2023 Conference Paper

Electrophysiological Brain Source Imaging via Combinatorial Search with Provable Optimality

  • Guihong Wan
  • Meng Jiao
  • Xinglong Ju
  • Yu Zhang
  • Haim Schweitzer
  • Feng Liu

Electrophysiological Source Imaging (ESI) refers to reconstructing the underlying brain source activation from non-invasive Electroencephalography (EEG) and Magnetoencephalography (MEG) measurements on the scalp. Estimating the source locations and their extents is a fundamental tool in clinical and neuroscience applications. However, the estimation is challenging because of the ill-posedness and high coherence in the leadfield matrix as well as the noise in the EEG/MEG data. In this work, we proposed a combinatorial search framework to address the ESI problem with a provable optimality guarantee. Specifically, by exploiting the graph neighborhood information in the brain source space, we converted the ESI problem into a graph search problem and designed a combinatorial search algorithm under the framework of A* to solve it. The proposed algorithm is guaranteed to give an optimal solution to the ESI problem. Experimental results on both synthetic data and real epilepsy EEG data demonstrated that the proposed algorithm could faithfully reconstruct the source activation in the brain.

NeurIPS Conference 2023 Conference Paper

Enhancing Adversarial Contrastive Learning via Adversarial Invariant Regularization

  • Xilie Xu
  • Jingfeng Zhang
  • Feng Liu
  • Masashi Sugiyama
  • Mohan S. Kankanhalli

Adversarial contrastive learning (ACL) is a technique that enhances standard contrastive learning (SCL) by incorporating adversarial data to learn a robust representation that can withstand adversarial attacks and common corruptions without requiring costly annotations. To improve transferability, the existing work introduced the standard invariant regularization (SIR) to impose style-independence property to SCL, which can exempt the impact of nuisance style factors in the standard representation. However, it is unclear how the style-independence property benefits ACL-learned robust representations. In this paper, we leverage the technique of causal reasoning to interpret the ACL and propose adversarial invariant regularization (AIR) to enforce independence from style factors. We regulate the ACL using both SIR and AIR to output the robust representation. Theoretically, we show that AIR implicitly encourages the representational distance between different views of natural data and their adversarial variants to be independent of style factors. Empirically, our experimental results show that invariant regularization significantly improves the performance of state-of-the-art ACL methods in terms of both standard generalization and robustness on downstream tasks. To the best of our knowledge, we are the first to apply causal reasoning to interpret ACL and develop AIR for enhancing ACL-learned robust representations. Our source code is at https: //github. com/GodXuxilie/Enhancing ACL via_AIR.

TMLR Journal 2023 Journal Article

KRADA: Known-region-aware Domain Alignment for Open-set Domain Adaptation in Semantic Segmentation

  • Chenhong Zhou
  • Feng Liu
  • Chen Gong
  • Rongfei Zeng
  • Tongliang Liu
  • William Cheung
  • Bo Han

In semantic segmentation, we aim to train a pixel-level classifier to assign category labels to all pixels in an image, where labeled training images and unlabeled test images are from the same distribution and share the same label set. However, in an open world, the unlabeled test images probably contain unknown categories and have different distributions from the labeled images. Hence, in this paper, we consider a new, more realistic, and more challenging problem setting where the pixel-level classifier has to be trained with labeled images and unlabeled open-world images—we name it open world semantic segmentation (OSS). In OSS, the trained classifier is expected to identify unknown-class pixels and classify known-class pixels well. To solve OSS, we first investigate which distribution that unknown-class pixels obey. Then, motivated by the goodness-of-fit test, we use statistical measurements to show how a pixel fits the distribution of an unknown class and select highly-fitted pixels to form the unknown region in each test image. Eventually, we propose an end-to-end learning framework, known-region-aware domain alignment (KRADA), to distinguish unknown classes while aligning the distributions of known classes in labeled and unlabeled open-world images. The effectiveness of KRADA has been verified on two synthetic tasks and one COVID-19 segmentation task.

NeurIPS Conference 2023 Conference Paper

Learning to Augment Distributions for Out-of-distribution Detection

  • Qizhou Wang
  • Zhen Fang
  • Yonggang Zhang
  • Feng Liu
  • Yixuan Li
  • Bo Han

Open-world classification systems should discern out-of-distribution (OOD) data whose labels deviate from those of in-distribution (ID) cases, motivating recent studies in OOD detection. Advanced works, despite their promising progress, may still fail in the open world, owing to the lacking knowledge about unseen OOD data in advance. Although one can access auxiliary OOD data (distinct from unseen ones) for model training, it remains to analyze how such auxiliary data will work in the open world. To this end, we delve into such a problem from a learning theory perspective, finding that the distribution discrepancy between the auxiliary and the unseen real OOD data is the key to affect the open-world detection performance. Accordingly, we propose Distributional-Augmented OOD Learning (DAOL), alleviating the OOD distribution discrepancy by crafting an OOD distribution set that contains all distributions in a Wasserstein ball centered on the auxiliary OOD distribution. We justify that the predictor trained over the worst OOD data in the ball can shrink the OOD distribution discrepancy, thus improving the open-world detection performance given only the auxiliary OOD data. We conduct extensive evaluations across representative OOD detection setups, demonstrating the superiority of our DAOL over its advanced counterparts.

NeurIPS Conference 2023 Conference Paper

Out-of-distribution Detection Learning with Unreliable Out-of-distribution Sources

  • Haotian Zheng
  • Qizhou Wang
  • Zhen Fang
  • Xiaobo Xia
  • Feng Liu
  • Tongliang Liu
  • Bo Han

Out-of-distribution (OOD) detection discerns OOD data where the predictor cannot make valid predictions as in-distribution (ID) data, thereby increasing the reliability of open-world classification. However, it is typically hard to collect real out-of-distribution (OOD) data for training a predictor capable of discerning ID and OOD patterns. This obstacle gives rise to data generation-based learning methods, synthesizing OOD data via data generators for predictor training without requiring any real OOD data. Related methods typically pre-train a generator on ID data and adopt various selection procedures to find those data likely to be the OOD cases. However, generated data may still coincide with ID semantics, i. e. , mistaken OOD generation remains, confusing the predictor between ID and OOD data. To this end, we suggest that generated data (with mistaken OOD generation) can be used to devise an auxiliary OOD detection task to facilitate real OOD detection. Specifically, we can ensure that learning from such an auxiliary task is beneficial if the ID and the OOD parts have disjoint supports, with the help of a well-designed training procedure for the predictor. Accordingly, we propose a powerful data generation-based learning method named Auxiliary Task-based OOD Learning (ATOL) that can relieve the mistaken OOD generation. We conduct extensive experiments under various OOD detection setups, demonstrating the effectiveness of our method against its advanced counterparts.

NeurIPS Conference 2022 Conference Paper

Cluster and Aggregate: Face Recognition with Large Probe Set

  • Minchul Kim
  • Feng Liu
  • Anil K Jain
  • Xiaoming Liu

Feature fusion plays a crucial role in unconstrained face recognition where inputs (probes) comprise of a set of $N$ low quality images whose individual qualities vary. Advances in attention and recurrent modules have led to feature fusion that can model the relationship among the images in the input set. However, attention mechanisms cannot scale to large $N$ due to their quadratic complexity and recurrent modules suffer from input order sensitivity. We propose a two-stage feature fusion paradigm, Cluster and Aggregate, that can both scale to large $N$ and maintain the ability to perform sequential inference with order invariance. Specifically, Cluster stage is a linear assignment of $N$ inputs to $M$ global cluster centers, and Aggregation stage is a fusion over $M$ clustered features. The clustered features play an integral role when the inputs are sequential as they can serve as a summarization of past features. By leveraging the order-invariance of incremental averaging operation, we design an update rule that achieves batch-order invariance, which guarantees that the contributions of early image in the sequence do not diminish as time steps increase. Experiments on IJB-B and IJB-S benchmark datasets show the superiority of the proposed two-stage paradigm in unconstrained face recognition.

TMLR Journal 2022 Journal Article

Fingerprints of Super Resolution Networks

  • Jeremy Vonderfecht
  • Feng Liu

Several recent studies have demonstrated that deep-learning based image generation models, such as GANs, can be uniquely identified, and possibly even reverse-engineered, by the fingerprints they leave on their output images. We extend this research to single image super-resolution (SISR) networks. Compared to previously studied models, SISR networks are a uniquely challenging class of image generation model from which to extract and analyze fingerprints, as they can often generate images that closely match the corresponding ground truth and thus likely leave little flexibility to embed signatures. We take SISR models as examples to investigate if the findings from the previous work on fingerprints of GAN-based networks are valid for general image generation models. We show that SISR networks with a high upscaling factor or trained using adversarial loss leave highly distinctive fingerprints, and that under certain conditions, some SISR network hyperparameters can be reverse-engineered from these fingerprints.

YNICL Journal 2022 Journal Article

Frequency-dependent white-matter functional network changes associated with cognitive deficits in subcortical vascular cognitive impairment

  • Juanwei Ma
  • Feng Liu
  • Yang Wang
  • Lin Ma
  • Yali Niu
  • Jing Wang
  • Zhaoxiang Ye
  • Jing Zhang

Vascular cognitive impairment (VCI) refers to all forms of cognitive decline associated with cerebrovascular diseases, in which white matter (WM) is highly vulnerable. Although previous studies have shown that blood oxygen level-dependent (BOLD) signals inside WM can effectively reflect neural activities, whether WM BOLD signal alterations are present and their roles underlying cognitive impairment in VCI remain largely unknown. In this study, 36 subcortical VCI (SVCI) patients and 36 healthy controls were enrolled to evaluate WM dysfunction. Specifically, fourteen distinct WM networks were identified from resting-state functional MRI using K-means clustering analysis. Subsequently, between-network functional connectivity (FC) and within-network BOLD signal amplitude of WM networks were calculated in three frequency bands (band A: 0.01-0.15 Hz, band B: 0.08-0.15 Hz, and band C: 0.01-0.08 Hz). Patients with SVCI manifested decreased FC mainly in bilateral parietal WM regions, forceps major, superior and inferior longitudinal fasciculi. These connections extensively linked with distinct WM networks and with gray-matter networks such as frontoparietal control, dorsal and ventral attention networks, which exhibited frequency-specific alterations in SVCI. Additionally, extensive amplitude reductions were found in SVCI, showing frequency-dependent properties in parietal, anterior corona radiate, pre/post central, superior and inferior longitudinal fasciculus networks. Furthermore, these decreased FC and amplitudes showed significant positive correlations with cognitive performances in SVCI, and high diagnostic performances for SVCI especially combining all bands. Our study indicated that VCI-related cognitive deficits were characterized by frequency-dependent WM functional abnormalities, which offered novel applicable neuromarkers for VCI.

YNIMG Journal 2022 Journal Article

Instant tissue field and magnetic susceptibility mapping from MRI raw phase using Laplacian enhanced deep neural networks

  • Yang Gao
  • Zhuang Xiong
  • Amir Fazlollahi
  • Peter J Nestor
  • Viktor Vegh
  • Fatima Nasrallah
  • Craig Winter
  • G. Bruce Pike

Quantitative susceptibility mapping (QSM) is an MRI post-processing technique that produces spatially resolved magnetic susceptibility maps from phase data. However, the traditional QSM reconstruction pipeline involves multiple non-trivial steps, including phase unwrapping, background field removal, and dipole inversion. These intermediate steps not only increase the reconstruction time but accumulates errors. This study aims to overcome existing limitations by developing a Laplacian-of-Trigonometric-functions (LoT) enhanced deep neural network for near-instant quantitative field and susceptibility mapping (i.e., iQFM and iQSM) from raw MRI phase data. The proposed iQFM and iQSM methods were compared with established reconstruction pipelines on simulated and in vivo datasets. In addition, experiments on patients with intracranial hemorrhage and multiple sclerosis were also performed to test the generalization of the proposed neural networks. The proposed iQFM and iQSM methods in healthy subjects yielded comparable results to those involving the intermediate steps while dramatically improving reconstruction accuracies on intracranial hemorrhages with large susceptibilities. High susceptibility contrast between multiple sclerosis lesions and healthy tissue was also achieved using the proposed methods. Comparative studies indicated that the most significant contributor to iQFM and iQSM over conventional multi-step methods was the elimination of traditional Laplacian unwrapping. The reconstruction time on the order of minutes for traditional approaches was shortened to around 0.1 s using the trained iQFM and iQSM neural networks.

NeurIPS Conference 2022 Conference Paper

Is Out-of-Distribution Detection Learnable?

  • Zhen Fang
  • Yixuan Li
  • Jie Lu
  • Jiahua Dong
  • Bo Han
  • Feng Liu

Supervised learning aims to train a classifier under the assumption that training and test data are from the same distribution. To ease the above assumption, researchers have studied a more realistic setting: out-of-distribution (OOD) detection, where test data may come from classes that are unknown during training (i. e. , OOD data). Due to the unavailability and diversity of OOD data, good generalization ability is crucial for effective OOD detection algorithms. To study the generalization of OOD detection, in this paper, we investigate the probably approximately correct (PAC) learning theory of OOD detection, which is proposed by researchers as an open problem. First, we find a necessary condition for the learnability of OOD detection. Then, using this condition, we prove several impossibility theorems for the learnability of OOD detection under some scenarios. Although the impossibility theorems are frustrating, we find that some conditions of these impossibility theorems may not hold in some practical scenarios. Based on this observation, we next give several necessary and sufficient conditions to characterize the learnability of OOD detection in some practical scenarios. Lastly, we also offer theoretical supports for several representative OOD detection works based on our OOD theory.

NeurIPS Conference 2022 Conference Paper

Mix and Reason: Reasoning over Semantic Topology with Data Mixing for Domain Generalization

  • Chaoqi Chen
  • Luyao Tang
  • Feng Liu
  • Gangming Zhao
  • Yue Huang
  • Yizhou Yu

Domain generalization (DG) enables generalizing a learning machine from multiple seen source domains to an unseen target one. The general objective of DG methods is to learn semantic representations that are independent of domain labels, which is theoretically sound but empirically challenged due to the complex mixture of common and domain-specific factors. Although disentangling the representations into two disjoint parts has been gaining momentum in DG, the strong presumption over the data limits its efficacy in many real-world scenarios. In this paper, we propose Mix and Reason (MiRe), a new DG framework that learns semantic representations via enforcing the structural invariance of semantic topology. MiRe consists of two key components, namely, Category-aware Data Mixing (CDM) and Adaptive Semantic Topology Refinement (ASTR). CDM mixes two images from different domains in virtue of activation maps generated by two complementary classification losses, making the classifier focus on the representations of semantic objects. ASTR introduces relation graphs to represent semantic topology, which is progressively refined via the interactions between local feature aggregation and global cross-domain relational reasoning. Experiments on multiple DG benchmarks validate the effectiveness and robustness of the proposed MiRe.

JBHI Journal 2022 Journal Article

Understanding Dynamics of Pandemic Models to Support Predictions of COVID-19 Transmission: Parameter Sensitivity Analysis of SIR-Type Models

  • Chunfeng Ma
  • Xin Li
  • Zebin Zhao
  • Feng Liu
  • Kun Zhang
  • Adan Wu
  • Xiaowei Nie

Despite efforts made to model and predict COVID-19 transmission, large predictive uncertainty remains. Failure to understand the dynamics of the nonlinear pandemic prediction model is an important reason. To this end, local and multiple global sensitivity analysis approaches are synthetically applied to analyze the sensitivities of parameters and initial state variables and community size (N) in susceptible-infected-recovered (SIR) and its variant susceptible-exposed-infected-recovered (SEIR) models and basic reproduction number ( R0 ), aiming to provide prior information for parameter estimation and suggestions for COVID-19 prevention and control measures. We found that N influences both the maximum number of actively infected cases and the date on which the maximum number of actively infected cases is reached. The high effect of N on maximum actively infected cases and peak date suggests the necessity of isolating the infected cases in a small community. The protection rate and average quarantined time are most sensitive to the infected populations, with a summation of their first-order sensitivity indices greater than 0. 585, and their interactions are also substantial, being 0. 389 and 0. 334, respectively. The high sensitivities and interaction between the protection rate and average quarantined time suggest that protection and isolation measures should always be implemented in conjunction and started as early as possible. These findings provide insights into the predictability of the pandemic models by estimating influential parameters and suggest how to effectively prevent and control epidemic transmission.

NeurIPS Conference 2022 Conference Paper

Watermarking for Out-of-distribution Detection

  • Qizhou Wang
  • Feng Liu
  • Yonggang Zhang
  • Jing Zhang
  • Chen Gong
  • Tongliang Liu
  • Bo Han

Out-of-distribution (OOD) detection aims to identify OOD data based on representations extracted from well-trained deep models. However, existing methods largely ignore the reprogramming property of deep models and thus may not fully unleash their intrinsic strength: without modifying parameters of a well-trained deep model, we can reprogram this model for a new purpose via data-level manipulation (e. g. , adding a specific feature perturbation). This property motivates us to reprogram a classification model to excel at OOD detection (a new task), and thus we propose a general methodology named watermarking in this paper. Specifically, we learn a unified pattern that is superimposed onto features of original data, and the model's detection capability is largely boosted after watermarking. Extensive experiments verify the effectiveness of watermarking, demonstrating the significance of the reprogramming property of deep models in OOD detection.

YNIMG Journal 2021 Journal Article

Accelerating quantitative susceptibility and R2* mapping using incoherent undersampling and deep neural network reconstruction

  • Yang Gao
  • Martijn Cloos
  • Feng Liu
  • Stuart Crozier
  • G. Bruce Pike
  • Hongfu Sun

Quantitative susceptibility mapping (QSM) and R2* mapping are MRI post-processing methods that quantify tissue magnetic susceptibility and transverse relaxation rate distributions. However, QSM and R2* acquisitions are relatively slow, even with parallel imaging. Incoherent undersampling and compressed sensing reconstruction techniques have been used to accelerate traditional magnitude-based MRI acquisitions; however, most do not recover the full phase signal, as required by QSM, due to its non-convex nature. In this study, a learning-based Deep Complex Residual Network (DCRNet) is proposed to recover both the magnitude and phase images from incoherently undersampled data, enabling high acceleration of QSM and R2* acquisition. Magnitude, phase, R2*, and QSM results from DCRNet were compared with two iterative and one deep learning methods on retrospectively undersampled acquisitions from six healthy volunteers, one intracranial hemorrhage and one multiple sclerosis patients, as well as one prospectively undersampled healthy subject using a 7T scanner. Peak signal to noise ratio (PSNR), structural similarity (SSIM), root-mean-squared error (RMSE), and region-of-interest susceptibility and R2* measurements are reported for numerical comparisons. The proposed DCRNet method substantially reduced artifacts and blurring compared to the other methods and resulted in the highest PSNR, SSIM, and RMSE on the magnitude, R2*, local field, and susceptibility maps. Compared to two iterative and one deep learning methods, the DCRNet method demonstrated a 3.2% to 9.1% accuracy improvement in deep grey matter susceptibility when accelerated by a factor of four. The DCRNet also dramatically shortened the reconstruction time of single 2D brain images from 36-140 seconds using conventional approaches to only 15-70 milliseconds.

AAAI Conference 2021 Conference Paper

Adversarial Defence by Diversified Simultaneous Training of Deep Ensembles

  • Bo Huang
  • Zhiwei Ke
  • Yi Wang
  • Wei Wang
  • Linlin Shen
  • Feng Liu

Learning-based classifiers are susceptible to adversarial examples. Existing defence methods are mostly devised on individual classifiers. Recent studies showed that it is viable to increase adversarial robustness by promoting diversity over an ensemble of models. In this paper, we propose adversarial defence by encouraging ensemble diversity on learning high-level feature representations and gradient dispersion in simultaneous training of deep ensemble networks. We perform extensive evaluations under white-box and blackbox attacks including transferred examples and adaptive attacks. Our approach achieves a significant gain of up to 52% in adversarial robustness, compared with the baseline and the state-of-the-art method on image benchmarks with complex data scenes. The proposed approach complements the defence paradigm of adversarial training, and can further boost the performance. The source code is available at https: //github. com/ALIS-Lab/AAAI2021-PDD.

JBHI Journal 2021 Journal Article

Blood Pressure States Transition Inference Based on Multi-State Markov Model

  • Jingmei Yang
  • Feng Liu
  • Boyu Wang
  • Chaoyang Chen
  • Timothy Church
  • Lee Dukes
  • Jeffrey O. Smith

The investigation of risk factors associated with hypertension patients has been extensively studied in the past decades. However, the pattern of natural progressive trajectories to hypertension from nonhypertensive states was rarely explored. In this study, we are interested in discovering the underlying transition patterns between different blood pressure states, namely normal state, elevated state, and hypertensive state among the working population in the United States. A multi-state Markov model was built based on 88, 966 clinical records from 34, 719 participants we collected during the worksite preventive screening from 2012 to 2018. We first investigated the various risk factors, and we found that body mass index (BMI) is the most critical factor for developing new-onset hypertension. The transition probabilities, survival probabilities, and sojourn time of each state were derived given different levels of BMI, age groups, and gender categories. We found the underweight participants are more likely to remain in the current nonhypertensive states within 3 years, while extremely obese participants have a higher probability of developing hypertension. We discovered the distinct transition patterns among male and female participants. On average, the sojourn time in the normal state for normal-weight participants is 4. 33 years for females and 2. 18 years for their male counterparts. For the extremely obese participants, the average sojourn time in the normal state is 1. 38 years for females and 0. 71 years for males. In the end, a web-based graphical user interface (GUI) application was developed for clinicians to visualize the impact of behavioral interventions on delaying the progression of hypertension. Our analysis can provide a unique insight into hypertension research and proactive interventions.

YNIMG Journal 2021 Journal Article

Frequency drift in MR spectroscopy at 3T

  • Steve C.N. Hui
  • Mark Mikkelsen
  • Helge J. Zöllner
  • Vishwadeep Ahluwalia
  • Sarael Alcauter
  • Laima Baltusis
  • Deborah A. Barany
  • Laura R. Barlow

PURPOSE: field, especially when gradient intensive sequences are used. The aim of the study was to set a benchmark for typical drift encountered during MR spectroscopy (MRS) to assess the need for real-time field-frequency locking on MRI scanners by comparing field drift data from a large number of sites. METHOD: A standardized protocol was developed for 80 participating sites using 99 3T MR scanners from 3 major vendors. Phantom water signals were acquired before and after an EPI sequence. The protocol consisted of: minimal preparatory imaging; a short pre-fMRI PRESS; a ten-minute fMRI acquisition; and a long post-fMRI PRESS acquisition. Both pre- and post-fMRI PRESS were non-water suppressed. Real-time frequency stabilization/adjustment was switched off when appropriate. Sixty scanners repeated the protocol for a second dataset. In addition, a three-hour post-fMRI MRS acquisition was performed at one site to observe change of gradient temperature and drift rate. Spectral analysis was performed using MATLAB. Frequency drift in pre-fMRI PRESS data were compared with the first 5:20 minutes and the full 30:00 minutes of data after fMRI. Median (interquartile range) drifts were measured and showed in violin plot. Paired t-tests were performed to compare frequency drift pre- and post-fMRI. A simulated in vivo spectrum was generated using FID-A to visualize the effect of the observed frequency drifts. The simulated spectrum was convolved with the frequency trace for the most extreme cases. Impacts of frequency drifts on NAA and GABA were also simulated as a function of linear drift. Data from the repeated protocol were compared with the corresponding first dataset using Pearson's and intraclass correlation coefficients (ICC). RESULTS: Of the data collected from 99 scanners, 4 were excluded due to various reasons. Thus, data from 95 scanners were ultimately analyzed. For the first 5:20 min (64 transients), median (interquartile range) drift was 0.44 (1.29) Hz before fMRI and 0.83 (1.29) Hz after. This increased to 3.15 (4.02) Hz for the full 30 min (360 transients) run. Average drift rates were 0.29 Hz/min before fMRI and 0.43 Hz/min after. Paired t-tests indicated that drift increased after fMRI, as expected (p < 0.05). Simulated spectra convolved with the frequency drift showed that the intensity of the NAA singlet was reduced by up to 26%, 44 % and 18% for GE, Philips and Siemens scanners after fMRI, respectively. ICCs indicated good agreement between datasets acquired on separate days. The single site long acquisition showed drift rate was reduced to 0.03 Hz/min approximately three hours after fMRI. DISCUSSION: This study analyzed frequency drift data from 95 3T MRI scanners. Median levels of drift were relatively low (5-min average under 1 Hz), but the most extreme cases suffered from higher levels of drift. The extent of drift varied across scanners which both linear and nonlinear drifts were observed.

YNIMG Journal 2021 Journal Article

Genes associated with gray matter volume alterations in schizophrenia

  • Yuan Ji
  • Xue Zhang
  • Zirui Wang
  • Wen Qin
  • Huaigui Liu
  • Kaizhong Xue
  • Jie Tang
  • Qiang Xu

Although both schizophrenia and gray matter volume (GMV) show high heritability, however, genes accounting for GMV alterations in schizophrenia remain largely unknown. Based on risk genes identified in schizophrenia by the genome-wide association study of the Schizophrenia Working Group of the Psychiatric Genomics Consortium, we used transcription-neuroimaging association analysis to test that which of these genes are associated with GMV changes in schizophrenia. For each brain tissue sample, the expression profiles of 196 schizophrenia risk genes were extracted from six donated normal brains of the Allen Human Brain Atlas, and GMV differences between patients with schizophrenia and healthy controls were calculated based on five independent case-control structural MRI datasets (276 patients and 284 controls). Genes associated with GMV changes in schizophrenia were identified by performing cross-sample spatial correlations between expression levels of each gene and case-control GMV difference derived from the five MRI datasets integrated by harmonization and meta-analysis. We found that expression levels of 98 genes consistently showed significant cross-sample spatial correlations with GMV changes in schizophrenia. These genes were functionally enriched for chemical synaptic transmission, central nervous system development, and cell projection. Overall, this study provides a set of genes possibly associated with GMV changes in schizophrenia, which could be used as candidate genes to explore biological mechanisms underlying the structural impairments in schizophrenia.

AAAI Conference 2021 Conference Paper

How Does the Combined Risk Affect the Performance of Unsupervised Domain Adaptation Approaches?

  • Li Zhong
  • Zhen Fang
  • Feng Liu
  • Jie Lu
  • Bo Yuan
  • Guangquan Zhang

Unsupervised domain adaptation (UDA) aims to train a target classifier with labeled samples from the source domain and unlabeled samples from the target domain. Classical UDA learning bounds show that target risk is upper bounded by three terms: source risk, distribution discrepancy, and combined risk. Based on the assumption that the combined risk is a small fixed value, methods based on this bound train a target classifier by only minimizing estimators of the source risk and the distribution discrepancy. However, the combined risk may increase when minimizing both estimators, which makes the target risk uncontrollable. Hence the target classifier cannot achieve ideal performance if we fail to control the combined risk. To control the combined risk, the key challenge takes root in the unavailability of the labeled samples in the target domain. To address this key challenge, we propose a method named E-MixNet. E-MixNet employs enhanced mixup, a generic vicinal distribution, on the labeled source samples and pseudo-labeled target samples to calculate a proxy of the combined risk. Experiments show that the proxy can effectively curb the increase of the combined risk when minimizing the source risk and distribution discrepancy. Furthermore, we show that if the proxy of the combined risk is added into loss functions of four representative UDA methods, their performance is also improved.

NeurIPS Conference 2021 Conference Paper

Meta Two-Sample Testing: Learning Kernels for Testing with Limited Data

  • Feng Liu
  • Wenkai Xu
  • Jie Lu
  • Danica J. Sutherland

Modern kernel-based two-sample tests have shown great success in distinguishing complex, high-dimensional distributions by learning appropriate kernels (or, as a special case, classifiers). Previous work, however, has assumed that many samples are observed from both of the distributions being distinguished. In realistic scenarios with very limited numbers of data samples, it can be challenging to identify a kernel powerful enough to distinguish complex distributions. We address this issue by introducing the problem of meta two-sample testing (M2ST), which aims to exploit (abundant) auxiliary data on related tasks to find an algorithm that can quickly identify a powerful test on new target tasks. We propose two specific algorithms for this task: a generic scheme which improves over baselines, and a more tailored approach which performs even better. We provide both theoretical justification and empirical evidence that our proposed meta-testing schemes outperform learning kernel-based tests directly from scarce observations, and identify when such schemes will be successful.

EAAI Journal 2021 Journal Article

More intelligent and robust estimation of battery state-of-charge with an improved regularized extreme learning machine

  • Meng Jiao
  • Dongqing Wang
  • Yan Yang
  • Feng Liu

State-of-charge (SOC) is the key parameter for battery management, and the accurate estimation of SOC is pretty important for the safe and stable operation of lithium batteries. This paper investigates a regularized extreme learning machine trained with the spectral Fletcher–Reeves algorithm and tuned with the beetle antennae search algorithm (BAS-SFR-RELM) for intelligent and robust SOC estimation. In the experiment section, the urban dynamometer driving schedule (UDDS) profile and the Los Angeles 92 (LA92) profile are performed on a battery test platform for data collection. In the simulation section, the root mean squared error (RMSE) and the mean absolute error (MAE) are adopted to evaluate the performance of the model. Compared with the linear regression (LR), the back propagation (BP) network, the multi-layer perceptron (MLP), and the long short-term memory (LSTM) network, the BAS-SFR-RELM method can efficiently obtain the optimal regularization coefficient to effectively prevent overfitting with faster convergence speed. Increasing the number of hidden neurons in the BAS-SFR-RELM appropriately can improve the SOC estimation precision. Implementing the BAS-SFR-RELM with the noise-added data set gives high robustness for SOC estimation

NeurIPS Conference 2021 Conference Paper

Probabilistic Margins for Instance Reweighting in Adversarial Training

  • Qizhou Wang
  • Feng Liu
  • Bo Han
  • Tongliang Liu
  • Chen Gong
  • Gang Niu
  • Mingyuan Zhou
  • Masashi Sugiyama

Reweighting adversarial data during training has been recently shown to improve adversarial robustness, where data closer to the current decision boundaries are regarded as more critical and given larger weights. However, existing methods measuring the closeness are not very reliable: they are discrete and can take only a few values, and they are path-dependent, i. e. , they may change given the same start and end points with different attack paths. In this paper, we propose three types of probabilistic margin (PM), which are continuous and path-independent, for measuring the aforementioned closeness and reweighing adversarial data. Specifically, a PM is defined as the difference between two estimated class-posterior probabilities, e. g. , such a probability of the true label minus the probability of the most confusing label given some natural data. Though different PMs capture different geometric properties, all three PMs share a negative correlation with the vulnerability of data: data with larger/smaller PMs are safer/riskier and should have smaller/larger weights. Experiments demonstrated that PMs are reliable and PM-based reweighting methods outperformed state-of-the-art counterparts.

NeurIPS Conference 2021 Conference Paper

TOHAN: A One-step Approach towards Few-shot Hypothesis Adaptation

  • Haoang Chi
  • Feng Liu
  • Wenjing Yang
  • Long Lan
  • Tongliang Liu
  • Bo Han
  • William Cheung
  • James Kwok

In few-shot domain adaptation (FDA), classifiers for the target domain are trained with \emph{accessible} labeled data in the source domain (SD) and few labeled data in the target domain (TD). However, data usually contain private information in the current era, e. g. , data distributed on personal phones. Thus, the private data will be leaked if we directly access data in SD to train a target-domain classifier (required by FDA methods). In this paper, to prevent privacy leakage in SD, we consider a very challenging problem setting, where the classifier for the TD has to be trained using few labeled target data and a well-trained SD classifier, named few-shot hypothesis adaptation (FHA). In FHA, we cannot access data in SD, as a result, the private information in SD will be protected well. To this end, we propose a target-oriented hypothesis adaptation network (TOHAN) to solve the FHA problem, where we generate highly-compatible unlabeled data (i. e. , an intermediate domain) to help train a target-domain classifier. TOHAN maintains two deep networks simultaneously, in which one focuses on learning an intermediate domain and the other takes care of the intermediate-to-target distributional adaptation and the target-risk minimization. Experimental results show that TOHAN outperforms competitive baselines significantly.

NeurIPS Conference 2021 Conference Paper

Voxel-based 3D Detection and Reconstruction of Multiple Objects from a Single Image

  • Feng Liu
  • Xiaoming Liu

Inferring 3D locations and shapes of multiple objects from a single 2D image is a long-standing objective of computer vision. Most of the existing works either predict one of these 3D properties or focus on solving both for a single object. One fundamental challenge lies in how to learn an effective representation of the image that is well-suited for 3D detection and reconstruction. In this work, we propose to learn a regular grid of 3D voxel features from the input image which is aligned with 3D scene space via a 3D feature lifting operator. Based on the 3D voxel features, our novel CenterNet-3D detection head formulates the 3D detection as keypoint detection in the 3D space. Moreover, we devise an efficient coarse-to-fine reconstruction module, including coarse-level voxelization and a novel local PCA-SDF shape representation, which enables fine detail reconstruction and two orders of magnitude faster inference than prior methods. With complementary supervision from both 3D detection and reconstruction, one enables the 3D voxel features to be geometry and context preserving, benefiting both tasks. The effectiveness of our approach is demonstrated through 3D detection and reconstruction on single-object and multiple-object scenarios.

IJCAI Conference 2020 Conference Paper

Clarinet: A One-step Approach Towards Budget-friendly Unsupervised Domain Adaptation

  • Yiyang Zhang
  • Feng Liu
  • Zhen Fang
  • Bo Yuan
  • Guangquan Zhang
  • Jie Lu

In unsupervised domain adaptation (UDA), classifiers for the target domain are trained with massive true-label data from the source domain and unlabeled data from the target domain. However, it may be difficult to collect fully-true-label data in a source domain given limited budget. To mitigate this problem, we consider a novel problem setting where the classifier for the target domain has to be trained with complementary-label data from the source domain and unlabeled data from the target domain named budget-friendly UDA (BFUDA). The key benefit is that it is much less costly to collect complementary-label source data (required by BFUDA) than collecting the true-label source data (required by ordinary UDA). To this end, complementary label adversarial network (CLARINET) is proposed to solve the BFUDA problem. CLARINET maintains two deep networks simultaneously, where one focuses on classifying complementary-label source data and the other takes care of the source-to-target distributional adaptation. Experiments show that CLARINET significantly outperforms a series of competent baselines.

NeurIPS Conference 2020 Conference Paper

Learning Implicit Functions for Topology-Varying Dense 3D Shape Correspondence

  • Feng Liu
  • Xiaoming Liu

The goal of this paper is to learn dense 3D shape correspondence for topology-varying objects in an unsupervised manner. Conventional implicit functions estimate the occupancy of a 3D point given a shape latent code. Instead, our novel implicit function produces a part embedding vector for each 3D point, which is assumed to be similar to its densely corresponded point in another 3D shape of the same object category. Furthermore, we implement dense correspondence through an inverse function mapping from the part embedding to a corresponded 3D point. Both functions are jointly learned with several effective loss functions to realize our assumption, together with the encoder generating the shape latent code. During inference, if a user selects an arbitrary point on the source shape, our algorithm can automatically generate a confidence score indicating whether there is a correspondence on the target shape, as well as the corresponding semantic point if there is. Such a mechanism inherently benefits man-made objects with different part constitutions. The effectiveness of our approach is demonstrated through unsupervised 3D semantic correspondence and shape segmentation.

YNIMG Journal 2020 Journal Article

Neural mechanisms of AVPR1A RS3-RS1 haplotypes that impact verbal learning and memory

  • Yan Zhang
  • Dan Zhu
  • Peng Zhang
  • Wei Li
  • Wen Qin
  • Feng Liu
  • Jiayuan Xu
  • Qiang Xu

Converging evidence from both human and animal studies has highlighted the pervasive role of the neuropeptide arginine vasopressin (AVP), which is mediated by arginine vasopressin receptor 1A (AVPR1A), in both social and nonsocial learning and memory. However, the effect of genetic variants in AVPR1A on verbal learning and memory is unknown. The hippocampus is a heterogeneous structure that consists of several anatomically and functionally distinct subfields, and it is the principal target structure for the memory-enhancing effect of AVP. We tested the hypothesis that genetic variants in the RS3 and RS1 repeat polymorphisms may influence verbal learning and memory performance evaluated by the California Verbal Learning Test-II (CVLT-II) by modulating the gray matter volume (GMV) and resting-state functional connectivity (rsFC) of whole hippocampus and its subfields in a large cohort of young healthy subjects (n = 1001). Using a short/long classification scheme for the repeat length of RS3 and RS1, we found that the individuals carrying more short alleles of RS3-RS1 haplotypes had poorer learning and memory performance compared to that of those carrying more long alleles. We also revealed that individuals carrying more short alleles exhibited a significantly smaller GMV in the left cornu ammonis (CA)2/3 and weaker rsFC of the left CA2/3-bilateral thalamic (primarily in medial prefrontal subfields) compared to those carrying more long alleles. Furthermore, multiple mediation analysis confirmed that these two hippocampal imaging measures jointly and fully mediated the relationship between the genetic variants in AVPR1A RS3-RS1 haplotypes and the individual differences in verbal learning and memory performance. Our results suggest that genetic variants in AVPR1A RS3-RS1 haplotypes may affect verbal learning and memory performance in part by modulating the left hippocampal CA2/3 structure and its rsFC with the thalamus.

ICRA Conference 2020 Conference Paper

VALID: A Comprehensive Virtual Aerial Image Dataset

  • Lyujie Chen
  • Feng Liu
  • Yan Zhao
  • Wufan Wang
  • Xiaming Yuan
  • Jihong Zhu 0001

Aerial imagery plays an important role in land-use planning, population analysis, precision agriculture, and unmanned aerial vehicle tasks. However, existing aerial image datasets generally suffer from the problem of inaccurate labeling, single ground truth type, and few category numbers. In this work, we implement a simulator that can simultaneously acquire diverse visual ground truth data in the virtual environment. Based on that, we collect a comprehensive Virtual AeriaL Image Dataset named VALID, consisting of 6690 high-resolution images, all annotated with panoptic segmentation on 30 categories, object detection with oriented bounding box, and binocular depth maps, collected in 6 different virtual scenes and 5 various ambient conditions (sunny, dusk, night, snow and fog). To our knowledge, VALID is the first aerial image dataset that can provide panoptic level segmentation and complete dense depth maps. We analyze the characteristics of VALID and evaluate state-of-the-art methods for multiple tasks to provide reference baselines. The experiment results demonstrate that VALID is well presented and challenging. The dataset is available at https://sites.google.com/view/valid-dataset/.

YNIMG Journal 2019 Journal Article

Big GABA II: Water-referenced edited MR spectroscopy at 25 research sites

  • Mark Mikkelsen
  • Daniel L. Rimbault
  • Peter B. Barker
  • Pallab K. Bhattacharyya
  • Maiken K. Brix
  • Pieter F. Buur
  • Kim M. Cecil
  • Kimberly L. Chan

Accurate and reliable quantification of brain metabolites measured in vivo using 1H magnetic resonance spectroscopy (MRS) is a topic of continued interest. Aside from differences in the basic approach to quantification, the quantification of metabolite data acquired at different sites and on different platforms poses an additional methodological challenge. In this study, spectrally edited γ-aminobutyric acid (GABA) MRS data were analyzed and GABA levels were quantified relative to an internal tissue water reference. Data from 284 volunteers scanned across 25 research sites were collected using GABA+ (GABA + co-edited macromolecules (MM)) and MM-suppressed GABA editing. The unsuppressed water signal from the volume of interest was acquired for concentration referencing. Whole-brain T 1-weighted structural images were acquired and segmented to determine gray matter, white matter and cerebrospinal fluid voxel tissue fractions. Water-referenced GABA measurements were fully corrected for tissue-dependent signal relaxation and water visibility effects. The cohort-wide coefficient of variation was 17% for the GABA + data and 29% for the MM-suppressed GABA data. The mean within-site coefficient of variation was 10% for the GABA + data and 19% for the MM-suppressed GABA data. Vendor differences contributed 53% to the total variance in the GABA + data, while the remaining variance was attributed to site- (11%) and participant-level (36%) effects. For the MM-suppressed data, 54% of the variance was attributed to site differences, while the remaining 46% was attributed to participant differences. Results from an exploratory analysis suggested that the vendor differences were related to the unsuppressed water signal acquisition. Discounting the observed vendor-specific effects, water-referenced GABA measurements exhibit similar levels of variance to creatine-referenced GABA measurements. It is concluded that quantification using internal tissue water referencing is a viable and reliable method for the quantification of in vivo GABA levels.

IJCAI Conference 2019 Conference Paper

Cross-City Transfer Learning for Deep Spatio-Temporal Prediction

  • Leye Wang
  • Xu Geng
  • Xiaojuan Ma
  • Feng Liu
  • Qiang Yang

Spatio-temporal prediction is a key type of tasks in urban computing, e. g. , traffic flow and air quality. Adequate data is usually a prerequisite, especially when deep learning is adopted. However, the development levels of different cities are unbalanced, and still many cities suffer from data scarcity. To address the problem, we propose a novel cross-city transfer learning method for deep spatio-temporal prediction tasks, called RegionTrans. RegionTrans aims to effectively transfer knowledge from a data-rich source city to a data-scarce target city. More specifically, we first learn an inter-city region matching function to match each target city region to a similar source city region. A neural network is designed to effectively extract region-level representation for spatio-temporal prediction. Finally, an optimization algorithm is proposed to transfer learned features from the source city to the target city with the region matching function. Using citywide crowd flow prediction as a demonstration experiment, we verify the effectiveness of RegionTrans. Results show that RegionTrans can outperform the state-of-the-art fine-tuning deep spatio-temporal prediction models by reducing up to 10. 7% prediction error.

NeurIPS Conference 2019 Conference Paper

Meta Learning with Relational Information for Short Sequences

  • Yujia Xie
  • Haoming Jiang
  • Feng Liu
  • Tuo Zhao
  • Hongyuan Zha

This paper proposes a new meta-learning method -- named HARMLESS (HAwkes Relational Meta Learning method for Short Sequences) for learning heterogeneous point process models from a collection of short event sequence data along with a relational network. Specifically, we propose a hierarchical Bayesian mixture Hawkes process model, which naturally incorporates the relational information among sequences into point process modeling. Compared with existing methods, our model can capture the underlying mixed-community patterns of the relational network, which simultaneously encourages knowledge sharing among sequences and facilitates adaptively learning for each individual sequence. We further propose an efficient stochastic variational meta-EM algorithm, which can scale to large problems. Numerical experiments on both synthetic and real data show that HARMLESS outperforms existing methods in terms of predicting the future events.

AAAI Conference 2017 Conference Paper

A Sparse Dictionary Learning Framework to Discover Discriminative Source Activations in EEG Brain Mapping

  • Feng Liu
  • Shouyi Wang
  • Jay Rosenberger
  • Jianzhong Su
  • Hanli Liu

Electroencephalography (EEG) source analysis is one of the most important noninvasive human brain imaging tools that provides millisecond temporal accuracy. However, discovering essential activated brain sources associated with different brain status is still a challenging problem. In this study, we propose for the first time that the ill-posed EEG inverse problem can be formulated and solved as a sparse over-complete dictionary learning problem. In particular, a novel supervised sparse dictionary learning framework was developed for EEG source reconstruction. A revised version of discriminative K-SVD (DK-SVD) algorithm is exploited to solve the formulated supervised dictionary learning problem. As the proposed learning framework incorporated the EEG label information of different brain status, it is capable of learning a sparse representation that reveal the most discriminative brain activity sources among different brain states. Compared to the state-of-the-art EEG source analysis methods, proposed sparse dictionary learning framework achieved significant superior performance in both computing speed and accuracy for the challenging EEG source reconstruction problem through extensive numerical experiments. More importantly, the experimental results also validated that the proposed sparse learning framework is effective to discover the discriminative task-related brain activation sources, which shows the potential to advance the high resolution EEG source analysis for real-time non-invasive brain imaging research.

AAAI Conference 2017 Short Paper

A Supervised Sparse Learning Framework to Solve EEG Inverse Problem for Discriminative Activations Pattern

  • Feng Liu

Electroencephalography (EEG) is one of the most important noninvasive neuroimaging tools that provides excellent temporal accuracy. As the EEG electrode sensors measure electrical potentials on the scalp instead of direct measuring activities of brain voxels deep inside the head, many approaches are proposed to infer the activated brain regions due to its significance in neuroscience research and clinical application. However, since mostly part of the brain activity is composed of the spontaneous neural activities or non-task related activations, task related activation patterns will be corrupted in strong background signal/noises. In our research, we proposed a sparse learning framework for solving EEG inverse problem which aims to explicitly extract the discriminative sources for different cognitive tasks by fusing the label information into the inverse model. The proposed framework is capable of estimation the discriminative brain sources under given different brain states where traditional inverse methods failed. We introduced two models, one is formulated as supervised sparse dictionary learning and the other one is the graph regularized discriminative source estimation model to promote the consistency within same class. Preliminary experimental results also validated that the proposed sparse learning framework is effective to discover the discriminative task-related brain activation sources, which shows the potential to advance the high resolution EEG source analysis for real-time non-invasive brain imaging research.

YNIMG Journal 2017 Journal Article

Big GABA: Edited MR spectroscopy at 24 research sites

  • Mark Mikkelsen
  • Peter B. Barker
  • Pallab K. Bhattacharyya
  • Maiken K. Brix
  • Pieter F. Buur
  • Kim M. Cecil
  • Kimberly L. Chan
  • David Y.-T. Chen

Magnetic resonance spectroscopy (MRS) is the only biomedical imaging method that can noninvasively detect endogenous signals from the neurotransmitter γ-aminobutyric acid (GABA) in the human brain. Its increasing popularity has been aided by improvements in scanner hardware and acquisition methodology, as well as by broader access to pulse sequences that can selectively detect GABA, in particular J-difference spectral editing sequences. Nevertheless, implementations of GABA-edited MRS remain diverse across research sites, making comparisons between studies challenging. This large-scale multi-vendor, multi-site study seeks to better understand the factors that impact measurement outcomes of GABA-edited MRS. An international consortium of 24 research sites was formed. Data from 272 healthy adults were acquired on scanners from the three major MRI vendors and analyzed using the Gannet processing pipeline. MRS data were acquired in the medial parietal lobe with standard GABA+ and macromolecule- (MM-) suppressed GABA editing. The coefficient of variation across the entire cohort was 12% for GABA+ measurements and 28% for MM-suppressed GABA measurements. A multilevel analysis revealed that most of the variance (72%) in the GABA+ data was accounted for by differences between participants within-site, while site-level differences accounted for comparatively more variance (20%) than vendor-level differences (8%). For MM-suppressed GABA data, the variance was distributed equally between site- (50%) and participant-level (50%) differences. The findings show that GABA+ measurements exhibit strong agreement when implemented with a standard protocol. There is, however, increased variability for MM-suppressed GABA measurements that is attributed in part to differences in site-to-site data acquisition. This study's protocol establishes a framework for future methodological standardization of GABA-edited MRS, while the results provide valuable benchmarks for the MRS community.

YNIMG Journal 2014 Journal Article

Inter-modality relationship constrained multi-modality multi-task feature selection for Alzheimer's Disease and mild cognitive impairment identification

  • Feng Liu
  • Chong-Yaw Wee
  • Huafu Chen
  • Dinggang Shen

Previous studies have demonstrated that the use of integrated information from multi-modalities could significantly improve diagnosis of Alzheimer's Disease (AD). However, feature selection, which is one of the most important steps in classification, is typically performed separately for each modality, which ignores the potentially strong inter-modality relationship within each subject. Recent emergence of multi-task learning approach makes the joint feature selection from different modalities possible. However, joint feature selection may unfortunately overlook different yet complementary information conveyed by different modalities. We propose a novel multi-task feature selection method to preserve the complementary inter-modality information. Specifically, we treat feature selection from each modality as a separate task and further impose a constraint for preserving the inter-modality relationship, besides separately enforcing the sparseness of the selected features from each modality. After feature selection, a multi-kernel support vector machine (SVM) is further used to integrate the selected features from each modality for classification. Our method is evaluated using the baseline PET and MRI images of subjects obtained from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database. Our method achieves a good performance, with an accuracy of 94. 37% and an area under the ROC curve (AUC) of 0. 9724 for AD identification, and also an accuracy of 78. 80% and an AUC of 0. 8284 for mild cognitive impairment (MCI) identification. Moreover, the proposed method achieves an accuracy of 67. 83% and an AUC of 0. 6957 for separating between MCI converters and MCI non-converters (to AD). These performances demonstrate the superiority of the proposed method over the state-of-the-art classification methods.

IJCAI Conference 2009 Conference Paper

  • Feng Liu
  • Yuzhen Niu
  • Michael Gleicher

In this paper, we present a method that uses web photos for measuring frame interestingness of a travel video. Web photo collections, such as those on Flickr, tend to contain interesting images because their images are more carefully taken, composed, and selected. Because these photos have already been chosen as subjectively interesting, they serve as evidence that similar images are also interesting. Our idea is to leverage these web photos to measure the interestingness of video frames. Specifically, we measure the interestingness of each video frame according to its similarity to web photos. The similarity is defined based on the scene content and composition. We characterize the scene content using scale invariant local features, specifically SIFT keypoints. We characterize composition by feature distribution. Accordingly, we measure the similarity between a web photo and a video frame based on the co-occurrence of the SIFT features, and the similarity between their spatial distribution. Interestingness of a video frame is measured by considering how many photos it is similar to, and how similar it is to them. Our experiments on measuring frame interestingness of videos from YouTube using photos from Flickr show the initial success of our method.

YNIMG Journal 2008 Journal Article

Study of the development of fetal baboon brain using magnetic resonance imaging at 3 Tesla

  • Feng Liu
  • Marianne Garland
  • Yunsuo Duan
  • Raymond I. Stark
  • Dongrong Xu
  • Zhengchao Dong
  • Ravi Bansal
  • Bradley S. Peterson

Direct observational data on the development of the brains of human and nonhuman primates is on remarkably scant, and most of our understanding of primate brain development is extrapolated from findings in rodent models. Magnetic resonance imaging (MRI) is a promising tool for the noninvasive, longitudinal study of the developing primate brain. We devised a protocol to scan pregnant baboons serially at 3 T for up to 3 h per session. Seven baboons were scanned 1–6 times, beginning as early as 56 days post-conceptional age, and as late as 185 days (term ∼185 days). Successful scanning of the fetal baboon required careful animal preparation and anesthesia, in addition to optimization of the scanning protocol. We successfully acquired maps of relaxation times (T 1 and T 2) and high-resolution anatomical images of the brains of fetal baboons at multiple time points during the course of gestation. These images demonstrated the convergence of gray and white matter contrast near term, and furthermore demonstrated that the loss of contrast at that age is a consequence of the continuous change in relaxation times during fetal brain development. These data furthermore demonstrate that maps of relaxation times have clear advantages over the relaxation time weighted images for the tracking of the changes in brain structure during fetal development. This protocol for in utero MRI of fetal baboon brains will help to advance the use of nonhuman primate models to study fetal brain development longitudinally.

v2026.09.13