Arrow Research search

Author name cluster

Youyong Kong

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

13 papers
2 author rows

Possible papers

13

AIIM Journal 2026 Journal Article

SpineCLUE: Automatic vertebrae identification using contrastive learning and uncertainty estimation

  • Sheng Zhang
  • Hongxuan Li
  • Minheng Chen
  • Mingying Li
  • Miao Liu
  • Junxian Wu
  • Cheng Xue
  • Youyong Kong

Vertebrae identification in arbitrary fields-of-view plays a crucial role in diagnosing spine disease. Most spine CT contain only local regions, such as the neck, chest, and abdomen. Existing spine-level methods, which rely on a priori on the specific number of target vertebrae, are less able to cope with this challenge. In this paper, we propose a three-stage vertebra-level method to address the challenges in 3D CT vertebrae identification with arbitrary fields-of-view. In order to integrate contextual prior information during identification, rather than identifying independently at the vertebrae-level, we perform the vertebrae localization, segmentation and identification tasks sequentially, thus making effective use of anatomical prior information about the vertebrae throughout the process. Specifically, to improve the stability of localization and prevent failures caused by abnormal vertebral positions in 3D space, we introduce a dual-factor density clustering algorithm to acquire localization information for individual vertebrae, thereby facilitating the subsequent segmentation and recognition processes. In addition, to tackle the issue of inter-class similarity and intra-class variability, we pretrain our identification network by using a supervised contrastive learning method. To further optimize the identification results, we estimated the uncertainty of the classification network and utilized the message fusion module to combine the uncertainty scores, while aggregating global information about the spine. Our method achieves state-of-the-art results on the VerSe20 challenge benchmark.

AAAI Conference 2025 Conference Paper

Exploring Rationale Learning for Continual Graph Learning

  • Lei Song
  • Jiaxing Li
  • Qinghua Si
  • Shihan Guan
  • Youyong Kong

Catastrophic forgetting poses a significant challenge for graph neural networks in continuously updating their knowledge base with data streams. To address this issue, much of the research has focused on node-level continual learning using parameter regularization or rehearsal-based strategies, while little attention given to graph-level tasks. Furthermore, current paradigms for continual graph learning may inadvertently capture spurious correlations for specific tasks through shortcuts, thereby exacerbating the forgetting of previous knowledge when new tasks are introduced. To tackle these challenges, we propose a novel paradigm, Rationale Learning GNN (RL-GNN), for graph-level continual graph learning. Specifically, we harness the invariant learning principle to incorporate environmental interventions into both the current and historical distributions, aiming to uncover rationales by minimizing empirical risk across all environments. The rationale serves as the sole factor guiding the learning process. Therefore, continual graph learning is redefined as capturing these invariant rationales within task sequences, alleviating catastrophic forgetting caused by spurious features. Extensive experiments on real-world datasets with varying task lengths demonstrate the effectiveness of our RL-GNN in continuous knowledge assimilation and reduction of catastrophic forgetting.

AAAI Conference 2025 Conference Paper

HePa: Heterogeneous Graph Prompting for All-Level Classification Tasks

  • Jia Jinghong
  • Lei Song
  • Jiaxing Li
  • Youyong Kong

Heterogeneous graphs, which are common in real-world downstream tasks, have recently sparked a wave of research interest. The performance of end-to-end heterogeneous graph neural networks (HGNNs) greatly relies on supervised training for specific tasks. To reduce the labeling cost, the "pretrain-finetune" paradigm has been widely adopted, but it leads to a knowledge gap between the pre-trained model and downstream tasks. In an effort to address this gap, the "pretrain-prompt" paradigm has emerged as a promising approach. This involves fine-tuning randomly initialized learnable vectors in downstream tasks. However, this approach may result in an insufficient representation of downstream task features. Existing techniques for heterogeneous graph prompting restructure the heterogeneous graph to align with the homogeneous graph prompting scheme. This can potentially introduce the same limitations as homogeneous graph prompt learning. In this paper, we propose HePa, short for Heterogeneous Graph Prompting for all-level classification tasks. It not only includes a unified prompt template-graph adapted for heterogeneous graphs but also introduces a novel pre-prompt token optimized during the pre-training phase to convey task information downstream. With these designs, HePa can complete all levels of classification tasks toward few-shot scenarios while activating in-context learning. Finally, we conducted a comprehensive experimental analysis of HePa on three benchmark datasets.

IJCAI Conference 2025 Conference Paper

Suit the Node Pair to the Case: A Multi-Scale Node Pair Grouping Strategy for Graph-MLP Distillation

  • Rui Dong
  • Jiaxing Li
  • Weihuang Zheng
  • Youyong Kong

Graph Neural Network (GNN) is powerful in solving various graph-related tasks, while its message passing mechanism may lead to latency during inference time. Multi-Layer-Perceptron (MLP) can achieve fast inference speed but with limited performance. One solution to fill this gap is through Knowledge Distillation. However, current distillation methods follow a ''node-to-node'' paradigm, while considering the complex relationships between different node pairs, direct distillation fails to capture these multiple-granularity features in GNN. Furthermore, current methods which focuses on the alignment of logits in the final layer ignores further learning within layers inside student MLP. Therefore, in this paper, we introduce a multi-scale knowledge distillation method (MSN-GDM) aiming to capture multiple knowledge from GNN to MLP. We firstly propose a multi-scale node-pair grouping strategy to assign node pairs to different-scale groups according to node pair similarity metrics. The similarity metrics consider both node features and topological structures of the given node pair. Then based on the preprocessed node-set groups, we design a multi-scale distillation method that can capture comprehensive knowledge in the corresponding node-set groups. The hierarchical weighted sum of each layer is applied as the final output. Extensive experiments on eight real-world datasets demonstrate the effectiveness of our proposed method.

ICML Conference 2025 Conference Paper

Topology-Aware Dynamic Reweighting for Distribution Shifts on Graph

  • Weihuang Zheng
  • Jiashuo Liu
  • Jiaxing Li
  • Jiayun Wu
  • Peng Cui 0001
  • Youyong Kong

Graph Neural Networks (GNNs) are widely used for node classification tasks but often fail to generalize when training and test nodes come from different distributions, limiting their practicality. To address this challenge, recent approaches have adopted invariant learning and sample reweighting techniques from the out-of-distribution (OOD) generalization field. However, invariant learning-based methods face difficulties when applied to graph data, as they rely on the impractical assumption of obtaining real environment labels and strict invariance, which may not hold in real-world graph structures. Moreover, current sample reweighting methods tend to overlook topological information, potentially leading to suboptimal results. In this work, we introduce the Topology-Aware Dynamic Reweighting (TAR) framework to address distribution shifts by leveraging the inherent graph structure. TAR dynamically adjusts sample weights through gradient flow on the graph edges during training. Instead of relying on strict invariance assumptions, we theoretically prove that our method is able to provide distributional robustness, thereby enhancing the out-of-distribution generalization performance on graph data. Our framework’s superiority is demonstrated through standard testing on extensive node classification OOD datasets, exhibiting marked improvements over existing methods.

NeurIPS Conference 2024 Conference Paper

Leveraging Tumor Heterogeneity: Heterogeneous Graph Representation Learning for Cancer Survival Prediction in Whole Slide Images

  • Junxian Wu
  • Xinyi Ke
  • Xiaoming Jiang
  • Huanwen Wu
  • Youyong Kong
  • Lizhi Shao

Survival prediction is a significant challenge in cancer management. Tumor micro-environment is a highly sophisticated ecosystem consisting of cancer cells, immune cells, endothelial cells, fibroblasts, nerves and extracellular matrix. The intratumor heterogeneity and the interaction across multiple tissue types profoundly impacts the prognosis. However, current methods often neglect the fact that the contribution to prognosis differs with tissue types. In this paper, we propose ProtoSurv, a novel heterogeneous graph model for WSI survival prediction. The learning process of ProtoSurv is not only driven by data but also incorporates pathological domain knowledge, including the awareness of tissue heterogeneity, the emphasis on prior knowledge of prognostic-related tissues, and the depiction of spatial interaction across multiple tissues. We validate ProtoSurv across five different cancer types from TCGA (i. e. , BRCA, LGG, LUAD, COAD and PAAD), and demonstrate the superiority of our method over the state-of-the-art methods.

AAAI Conference 2024 Conference Paper

Multiscale Low-Frequency Memory Network for Improved Feature Extraction in Convolutional Neural Networks

  • Fuzhi Wu
  • Jiasong Wu
  • Youyong Kong
  • Chunfeng Yang
  • Guanyu Yang
  • Huazhong Shu
  • Guy Carrault
  • Lotfi Senhadji

Deep learning and Convolutional Neural Networks (CNNs) have driven major transformations in diverse research areas. However, their limitations in handling low-frequency in-formation present obstacles in certain tasks like interpreting global structures or managing smooth transition images. Despite the promising performance of transformer struc-tures in numerous tasks, their intricate optimization com-plexities highlight the persistent need for refined CNN en-hancements using limited resources. Responding to these complexities, we introduce a novel framework, the Mul-tiscale Low-Frequency Memory (MLFM) Network, with the goal to harness the full potential of CNNs while keep-ing their complexity unchanged. The MLFM efficiently preserves low-frequency information, enhancing perfor-mance in targeted computer vision tasks. Central to our MLFM is the Low-Frequency Memory Unit (LFMU), which stores various low-frequency data and forms a parallel channel to the core network. A key advantage of MLFM is its seamless compatibility with various prevalent networks, requiring no alterations to their original core structure. Testing on ImageNet demonstrated substantial accuracy improvements in multiple 2D CNNs, including ResNet, MobileNet, EfficientNet, and ConvNeXt. Furthermore, we showcase MLFM's versatility beyond traditional image classification by successfully integrating it into image-to-image translation tasks, specifically in semantic segmenta-tion networks like FCN and U-Net. In conclusion, our work signifies a pivotal stride in the journey of optimizing the ef-ficacy and efficiency of CNNs with limited resources. This research builds upon the existing CNN foundations and paves the way for future advancements in computer vision. Our codes are available at https://github.com/AlphaWuSeu/MLFM.

JBHI Journal 2024 Journal Article

STANet: Spatio-Temporal Adaptive Network and Clinical Prior Embedding Learning for 3D+T CMR Segmentation

  • Xiaoming Qi
  • Yuting He
  • Yaolei Qi
  • Youyong Kong
  • Guanyu Yang
  • Shuo Li

The segmentation of cardiac structure in magnetic resonance images (CMR) is paramount in diagnosing and managing cardiovascular illnesses, given its 3D+Time (3D+T) sequence. The existing deep learning methods are constrained in their ability to 3D+T CMR segmentation, due to: (1) Limited motion perception. The complexity of heart beating renders the motion perception in 3D+T CMR, including the long-range and cross-slice motions. The existing methods' local perception and slice-fixed perception directly limit the performance of 3D+T CMR perception. (2) Lack of labels. Due to the expensive labeling cost of the 3D+T CMR sequence, the labels of 3D+T CMR only contain the end-diastolic and end-systolic frames. The incomplete labeling scheme causes inefficient supervision. Hence, we propose a novel spatio-temporal adaptation network with clinical prior embedding learning (STANet) to ensure efficient spatio-temporal perception and optimization on 3D+T CMR segmentation. (1) A spatio-temporal adaptive convolution (STAC) treats the 3D+T CMR sequence as a whole for perception. The long-distance motion correlation is embedded into the structural perception by learnable weight regularization to balance long-range motion perception. The structural similarity is measured by cross-attention to adaptively correlate the cross-slice motion. (2) A clinical prior embedding learning strategy (CPE) is proposed to optimize the partially labeled 3D+T CMR segmentation dynamically by embedding clinical priors into optimization. STANet achieves outstanding performance with Dice of 0. 917 and 0. 94 on two public datasets (ACDC and STACOM), which indicates STANet has the potential to be incorporated into computer-aided diagnosis tools for clinical application.

YNICL Journal 2023 Journal Article

Predicting treatment response in adolescents and young adults with major depressive episodes from fMRI using graph isomorphism network

  • Jia Duan
  • Yueying Li
  • Xiaotong Zhang
  • Shuai Dong
  • Pengfei Zhao
  • Jie Liu
  • Junjie Zheng
  • Rongxin Zhu

BACKGROUND: Major depressive episode (MDE) is the main clinical feature of mood disorders (major depressive disorder and bipolar disorder) in adolescents and young adults and accounts for most of the disease course. However, 30%-40% of MDE patients not responding to clinical first-line interventions. It is crucial to predict treatment response in the early stages and identify biomarkers associated with treatment response. Graph Isomorphism Network (GIN), a deep learning method, is promising for predicting treatment response for individual MDE patients with more powerful representation ability to capture the features of brain functional connectivity. METHODS: In this study, GIN was used to predict individual treatment response in 198 adolescents and young adults with MDE. The most discriminating regions were also identified for the treatment response prediction. RESULTS: Using GIN approach, the baseline functional connectivity could predict 79.8% responders and 67.4% non-responders to treatment (accuracy 74.24%). Furthermore, the most discriminating brain regions were mainly involved in paralimbic and subcortical areas. CONCLUSIONS: GIN has shown potential in predicting treatment response for individual patients, which may enable personalized treatment decisions. Furthermore, targeted interventions focused on modulating the activity and connectivity within paralimbic and subcortical regions could potentially improve treatment outcomes and enable personalized interventions for adolescents and young adults with MDE.

NeurIPS Conference 2023 Conference Paper

RH-BrainFS: Regional Heterogeneous Multimodal Brain Networks Fusion Strategy

  • Hongting Ye
  • Yalu Zheng
  • Yueying Li
  • Ke Zhang
  • Youyong Kong
  • Yonggui Yuan

Multimodal fusion has become an important research technique in neuroscience that completes downstream tasks by extracting complementary information from multiple modalities. Existing multimodal research on brain networks mainly focuses on two modalities, structural connectivity (SC) and functional connectivity (FC). Recently, extensive literature has shown that the relationship between SC and FC is complex and not a simple one-to-one mapping. The coupling of structure and function at the regional level is heterogeneous. However, all previous studies have neglected the modal regional heterogeneity between SC and FC and fused their representations via "simple patterns", which are inefficient ways of multimodal fusion and affect the overall performance of the model. In this paper, to alleviate the issue of regional heterogeneity of multimodal brain networks, we propose a novel Regional Heterogeneous multimodal Brain networks Fusion Strategy (RH-BrainFS). Briefly, we introduce a brain subgraph networks module to extract regional characteristics of brain networks, and further use a new transformer-based fusion bottleneck module to alleviate the issue of regional heterogeneity between SC and FC. To the best of our knowledge, this is the first paper to explicitly state the issue of structural-functional modal regional heterogeneity and to propose asolution. Extensive experiments demonstrate that the proposed method outperforms several state-of-the-art methods in a variety of neuroscience tasks.

JBHI Journal 2022 Journal Article

Few-Shot Learning for Deformable Medical Image Registration With Perception-Correspondence Decoupling and Reverse Teaching

  • Yuting He
  • Tiantian Li
  • Rongjun Ge
  • Jian Yang
  • Youyong Kong
  • Jian Zhu
  • Huazhong Shu
  • Guanyu Yang

Deformable medical image registration estimates corresponding deformation to align the regions of interest (ROIs) of two images to a same spatial coordinate system. However, recent unsupervised registration models only have correspondence ability without perception, making misalignment on blurred anatomies and distortion on task-unconcerned backgrounds. Label-constrained (LC) registration models embed the perception ability via labels, but the lack of texture constraints in labels and the expensive labeling costs causes distortion internal ROIs and overfitted perception. We propose the first few-shot deformable medical image registration framework, Perception-Correspondence Registration (PC-Reg), which embeds perception ability to registration models only with few labels, thus greatly improving registration accuracy and reducing distortion. 1) We propose the Perception-Correspondence Decoupling which decouples the perception and correspondence actions of registration to two CNNs. Therefore, independent optimizations and feature representations are available avoiding interference of the correspondence due to the lack of texture constraints. 2) For few-shot learning, we propose Reverse Teaching which aligns labeled and unlabeled images to each other to provide supervision information to the structure and style knowledge in unlabeled images, thus generating additional training data. Therefore, these data will reversely teach our perception CNN more style and structure knowledge, improving its generalization ability. Our experiments on three datasets with only five labels demonstrate that our PC-Reg has competitive registration accuracy and effective distortion-reducing ability. Compared with LC-VoxelMorph( $\lambda =1$ ), we achieve the 12. 5%, 6. 3% and 1. 0% Reg-DSC improvements on three datasets, revealing our framework with great potential in clinical application.

IJCAI Conference 2022 Conference Paper

Hierarchical Diffusion Scattering Graph Neural Network

  • Ke Zhang
  • Xinyan Pu
  • Jiaxing Li
  • Jiasong Wu
  • Huazhong Shu
  • Youyong Kong

Graph neural network (GNN) is popular now to solve the tasks in non-Euclidean space and most of them learn deep embeddings by aggregating the neighboring nodes. However, these methods are prone to some problems such as over-smoothing because of the single-scale perspective field and the nature of low-pass filter. To address these limitations, we introduce diffusion scattering network (DSN) to exploit high-order patterns. With observing the complementary relationship between multi-layer GNN and DSN, we propose Hierarchical Diffusion Scattering Graph Neural Network (HDS-GNN) to efficiently bridge DSN and GNN layer by layer to supplement GNN with multi-scale information and band-pass signals. Our model extracts node-level scattering representations by intercepting the low-pass filtering, and adaptively tunes the different scales to regularize multi-scale information. Then we apply hierarchical representation enhancement to improve GNN with the scattering features. We benchmark our model on nine real-world networks on the transductive semi-supervised node classification task. The experimental results demonstrate the effectiveness of our method.

JBHI Journal 2022 Journal Article

Landmark Localization for Cephalometric Analysis Using Multiscale Image Patch-Based Graph Convolutional Networks

  • Gang Lu
  • Yuanxiu Zhang
  • Youyong Kong
  • Chen Zhang
  • Jean-Louis Coatrieux
  • Huazhong Shu

Accurate and robust cephalometric image analysis plays an essential role in orthodontic diagnosis, treatment assessment and surgical planning. This paper proposes a novel landmark localization method for cephalometric analysis using multiscale image patch-based graph convolutional networks. In detail, image patches with the same size are hierarchically sampled from the Gaussian pyramid to well preserve multiscale context information. We combine local appearance and shape information into spatialized features with an attention module to enrich node representations in graph. The spatial relationships of landmarks are built with the incorporation of three-layer graph convolutional networks, and multiple landmarks are simultaneously updated and moved toward the targets in a cascaded coarse-to-fine process. Quantitative results obtained on publicly available cephalometric X-ray images have exhibited superior performance compared with other state-of-the-art methods in terms of mean radial error and successful detection rate within various precision ranges. Our approach performs significantly better especially in the clinically accepted range of 2 mm and this makes it suitable in cephalometric analysis and orthognathic surgery.

v2026.09.13