Arrow Research search

Author name cluster

Rushi Lan

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

9 papers
1 author row

Possible papers

9

AAAI Conference 2026 Conference Paper

Integral-based Knockoffs Inference for Partially Linear Models

  • Hao Wang
  • Biqin Song
  • Rushi Lan
  • Hong Chen

Partial linear models (PLM) have attracted much attention for regression estimation and variable selection due to their feasibility on utilizing linear and nonlinear approximations jointly. However, theoretical understanding of how they control the false discovery rate (FDR) during variable selection remains limited. To address this issue, we formulate a new integral-based knockoffs (IKO) inference scheme for controlled variable selection in PLM, where integral-based knockoff statistics are used to measure the variable importance and B-splines (or random Fourier features) are employed for approximating nonlinear components. In theory, FDR control is guaranteed for both linear and nonlinear parts, and the statistical analysis for its power is established. Empirical evaluations validate the effectiveness of our proposed approach.

AAAI Conference 2025 Conference Paper

CA-MLIF: Cross-Attention and Multimodal Low-Rank Interaction Fusion Framework for Tumor Prognostic Prediction

  • Yajun An
  • Jiale Chen
  • Huan Lin
  • Zhenbing Liu
  • Siyang Feng
  • Hualong Zhang
  • Rushi Lan
  • Zaiyi Liu

Cancer is a leading cause of death worldwide due to its aggressive nature and complex variability. Accurate prognosis is therefore challenging but essential for guiding personalized treatment and follow-up. Previous research often relied on single data sources, missing the opportunity to combine various types of patient information for more comprehensive survival predictions. To address these challenges, we propose a two-stage fusion method named Cross-Attention and Multimodal Low-Rank Interaction Fusion Framework (CA-MLIF). In the first stage, we propose a CA mechanism for real-time feature updates and cross-modal mutual learning to capture rich semantic information. In the second stage, we design a novel multimodal low-rank interaction fusion method for survival prediction. Specifically, we present modal attention mechanism (MAM) for feature filtration, low-rank multimodal fusion (LMF) for model complexity reduction, and optimal weight concatenation (OWC) for maximizing feature integration. Extensive experiments on two public datasets TCGA-GBMLGG and TCGA-KIRC, as well as a multi-center in-house lung adenocarcinoma (LUAD) dataset validate the effectiveness of CA-MLIF, which demonstrate that our method outperforms existing approaches in survival prediction under both pathology-gene fusion and CT-pathology fusion scenarios.

AAAI Conference 2025 Conference Paper

Multi-View Collaborative Learning Network for Speech Deepfake Detection

  • Kuiyuan Zhang
  • Zhongyun Hua
  • Rushi Lan
  • Yifang Guo
  • Yushu Zhang
  • Guoai Xu

As deep learning techniques advance rapidly, deepfake speech synthesized through text-to-speech or voice conversion networks is becoming increasingly realistic, posing significant challenges for detection and raising potential threats to social security. This growing realism has prompted extensive research in speech deepfake detection. However, current detection methods primarily focus on extracting features from either the raw waveform or the spectrogram, often overlooking the valuable correspondences between these two modalities that could enhance the detection of previously unseen types of deepfakes. In this work, we propose a multi-view collaborative learning network for speech deepfake detection, which jointly learns robust speech representations from both raw waveforms and spectrograms. Specifically, we first design a Dual-Branch Contrastive Learning (DBCL) framework for learning different view features. DBCL consists of two branches that learn representations from the raw waveform or the spectrogram and utilizes contrastive learning to enhance inter- and inner-view correlations. Additionally, we introduce a Waveform-Spectrogram Fusion Module (WSFM) to exchange multi-view information for collaborative learning. In the feature learning process, WSFM converts features between views and merges them adaptively using waveform-spectrogram cross-attention. The final detection is conducted based on the concatenation of the waveform and spectrogram features. We conduct extensive experiments on four benchmark deepfake speech detection datasets, and the experimental results demonstrate that our method can achieve better detection performance than current state-of-the-art detection methods.

AAAI Conference 2025 Conference Paper

Phoneme-Level Feature Discrepancies: A Key to Detecting Sophisticated Speech Deepfakes

  • Kuiyuan Zhang
  • Zhongyun Hua
  • Rushi Lan
  • Yushu Zhang
  • Yifang Guo

Recent advancements in text-to-speech and speech conversion technologies have enabled the creation of highly convincing synthetic speech. While these innovations offer numerous practical benefits, they also cause significant security challenges when maliciously misused. Therefore, there is an urgent need to detect these synthetic speech signals. Phoneme features provide a powerful speech representation for deepfake detection. However, previous phoneme-based detection approaches typically focused on specific phonemes, overlooking temporal inconsistencies across the entire phoneme sequence. In this paper, we develop a new mechanism for detecting speech deepfakes by identifying the inconsistencies of phoneme-level speech features. We design an adaptive phoneme pooling technique that extracts sample-specific phoneme-level features from frame-level speech data. By applying this technique to features extracted by pre-trained audio models on previously unseen deepfake datasets, we demonstrate that deepfake samples often exhibit phoneme-level inconsistencies when compared to genuine speech. To further enhance detection accuracy, we propose a deepfake detector that uses a graph attention network to model the temporal dependencies of phoneme-level features. Additionally, we introduce a random phoneme substitution augmentation technique to increase feature diversity during training. Extensive experiments on four benchmark datasets demonstrate the superior performance of our method over existing state-of-the-art detection methods.

JBHI Journal 2025 Journal Article

Semi-Supervised Gland Segmentation via Feature-Enhanced Contrastive Learning and Dual-Consistency Strategy

  • Jiejiang Yu
  • Bingbing Li
  • Xipeng Pan
  • Zhenwei Shi
  • Huadeng Wang
  • Rushi Lan
  • Xiaonan Luo

In the field of gland segmentation in histopathology, deep-learning methods have made significant progress. However, most existing methods not only require a large amount of high-quality annotated data but also tend to confuse the internal of the gland with the background. To address this challenge, we propose a new semi-supervised method named DCCL-Seg for gland segmentation, which follows the teacher-student framework. Our approach can be divided into follows steps. First, we design a contrastive learning module to improve the ability of the student model's feature extractor to distinguish between gland and background features. Then, we introduce a Signed Distance Field (SDF) prediction task and employ dual-consistency strategy (across tasks and models) to better reinforce the learning of gland internal. Next, we proposed a pseudo label filtering and reweighting mechanism, which filters and reweights the pseudo labels generated by the teacher model based on confidence. However, even after reweighting, the pseudo labels may still be influenced by unreliable pixels. Finally, we further designed an assistant predictor to learn the reweighted pseudo labels, which do not interfere with the student model's predictor and ensure the reliability of the student model's predictions. Experimental results on the publicly available GlaS and CRAG datasets demonstrate that our method outperforms other semi-supervised medical image segmentation methods.

AAAI Conference 2025 Conference Paper

Weakly Supervised Gland Segmentation with Class Semantic Consistency and Purified Labels Filtration

  • Siyang Feng
  • Huadeng Wang
  • Chu Han
  • Zhenbing Liu
  • Hualong Zhang
  • Rushi Lan
  • Xipeng Pan

Image-level weakly supervised semantic segmentation (WSSS) reduces the dependence on high-quality data annotation, which plays a crucial role in computational pathology. Benefit from the ability to localize the objects with only binary labels, Class Activation Map (CAM) is a widely used method to initial pseudo masks. However, due to the low contrast among different tissues in histopathological images, most existing CAM-based methods perform poorly in gland segmentation. We retrospect this process and find that class consistency and semantic consistency can guide the network to effectively distinguish confusing pixels and generate fine-grained pseudo masks. Specifically, for class consistency, we propose Consistency Correlation Attention (CCA) to encourage the network to focus on the contribution of class features to semantic dependencies. For semantic consistency, we propose Multi-scale Pyramid Fusion Pooling (MPFP) to aggregate coarse-to-fine global semantic information from CAMs at multiple spatial resolutions, thus identifying class localization. Additionally, we introduce a Purified Labels Filtration (PLF) strategy during the segmentation phase to mitigate the noisy supervision signal and improve the segmentation quality of the model. Extensive experiments show that the our method achieves new state-of-the-art results on three publicly available gland datasets. Furthermore, our method demonstrates impressive domain adaptation capability, achieving satisfactory results with only a small portion of samples when faced with unseen domain data.

EAAI Journal 2025 Journal Article

Weakly supervised histopathology tissue semantic segmentation with multi-scale voting and online noise suppression

  • Xipeng Pan
  • Hualong Zhang
  • Huahu Deng
  • Huadeng Wang
  • Lingqiao Li
  • Zhenbing Liu
  • Lin Wang
  • Yajun An

The development of an Artificial Intelligence (AI) assisted tissue segmentation method of digital pathology images is critical for cancer diagnosis and prognosis. Excellent performance has been achieved with the current fully supervised segmentation approach, which relies on a huge number of annotated data. However, drawing dense pixel-level annotations on the giga-pixel whole slide image (WSI) is extremely time-consuming and labor-intensive. To this end, we propose a tissue segmentation method using only patch-level classification labels to reduce such annotation burden and significantly improve the quality of the pseudo-masks. We introduce a framework with two phases of classification and segmentation. In the classification phase, we propose a multi-scale voting method on the Class Activation Map (CAM) based model to obtain more stable pseudo masks. In the segmentation phase, an Online Noise Suppression Strategy (ONSS) is proposed to encourage the model to focus on more reliable signals in the pseudo mask rather than noisy signals. Extensive experiments on two weakly supervised pathology image tissue segmentation datasets Lung Adenocarcinoma (LUAD-HistoSeg) and Breast Cancer Semantic Segmentation (BCSS-WSSS) demonstrate our model outperforms state-of-the-art weakly-supervised semantic segmentation (WSSS) methods using patch-level labels. Furthermore, our method exhibits superior generalization ability compared to other models, and demonstrates promising adaptation performance on unseen domains with only small amounts of data.

AIIM Journal 2025 Journal Article

Weakly supervised nuclei segmentation based on pseudo label correction and uncertainty denoising

  • Xipeng Pan
  • Shilong Song
  • Zhenbing Liu
  • Huadeng Wang
  • Lingqiao Li
  • Haoxiang Lu
  • Rushi Lan
  • Xiaonan Luo

Nuclei segmentation plays a vital role in computer-aided histopathology image analysis. Numerous fully supervised learning approaches exhibit amazing performance relying on pathological image with precisely annotations. Whereas, it is difficult and time-consuming in accurate manual labeling on pathological images. Hence, this paper presents a two-stage weakly supervised model including coarse and fine phases, which can achieve nuclei segmentation on whole slide images using only point annotations. In the coarse segmentation step, Voronoi diagram and K-means cluster results are generated based on the point annotations to supervise the training network. In order to cope with the different imaging conditions, an image adaptive clustering pseudo label algorithm is proposed to adapt the color distribution of different images. A Multi-scale Feature Fusion (MFF) module is designed in the decoder to better fusion the feature outputs. Additionally, to reduce the interference of erroneous cluster label, an Exponential Moving Average for cluster label Correction (EMAC) strategy is proposed. After the first step, an uncertainty estimation pseudo label denoising strategy is introduced to denoise Voronoi diagram and adaptive cluster label. In the fine segmentation step, the optimized labels are used for training to obtain the final predicted probability map. Extensive experiments are performed on MoNuSeg and TNBC public benchmarks, which demonstrate our proposed method is superior to other existing nuclei segmentation methods based on point labels. Codes are available at: https: //github. com/SSL-droid/WNS-PLCUD.

JBHI Journal 2017 Journal Article

Medical Image Retrieval via Histogram of Compressed Scattering Coefficients

  • Rushi Lan
  • Yicong Zhou

The features used in many current medical image retrieval systems are usually low-level hand-crafted features. This limitation may adversely affect the retrieval performance. To address this problem, this paper proposes a simple yet discriminative feature, called histogram of compressed scattering coefficients (HCSC), for medical image retrieval. In the proposed work, the scattering transform, a particular variation of deep convolutional networks, is first performed to yield more abstract representations of a medical image. A projection operation is then conducted to compress the obtained scattering coefficients for efficient processing. Finally, a bag-of-words (BoW) histogram is derived from the compressed scattering coefficients as the features of the medical image. The proposed HCSC takes the advantages of both scattering transform and BoW model. Experiments on three benchmark medical computer tomography image databases demonstrate that HCSC outperforms several state-of-the-art features.

v2026.09.13