Arrow Research search

Author name cluster

Tong Tong

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

10 papers
1 author row

Possible papers

10

EAAI Journal 2025 Journal Article

A universal parameter-efficient fine-tuning approach for stereo image super-resolution

  • Yuanbo Zhou
  • Yuyang Xue
  • Xinlin Zhang
  • Wei Deng
  • Tao Wang
  • Tao Tan
  • Qinquan Gao
  • Tong Tong

Despite advances in the use of the strategy of pre-training then fine-tuning in low-level vision tasks, the increasing size of models presents significant challenges for this paradigm, particularly in terms of training time and memory consumption. In addition, unsatisfactory results may occur when pre-trained single-image models are directly applied to a multi-image domain. In this paper, we propose an efficient method for transferring a pre-trained single-image super-resolution transformer network to the domain of stereo image super-resolution (SteISR) using a parameter-efficient fine-tuning approach. Specifically, the concept of stereo adapters and spatial adapters are introduced, which are incorporated into the pre-trained single-image super-resolution transformer network. Subsequently, only the inserted adapters are trained on stereo datasets. Compared with the classical full fine-tuning paradigm, our method can effectively reduce training time and memory consumption by 57% and 15%, respectively. Moreover, this method allows us to train only 4. 8% of the original model parameters, achieving state-of-the-art performance on four commonly used SteISR benchmarks. This technology is expected to improve stereo image resolution in various fields such as medical imaging and autonomous driving, thereby indirectly enhancing the accuracy of depth estimation and object recognition tasks.

JBHI Journal 2025 Journal Article

UniMRISegNet: Universal 3D Network for Various Organs and Cancers Segmentation on Multi-Sequence MRI

  • Zhuoneng Zhang
  • Luyi Han
  • Tianyu Zhang
  • Zehui Lin
  • Qinquan Gao
  • Tong Tong
  • Yue Sun
  • Tao Tan

Three-dimensional organ and cancer segmentation based on multi-sequence MRI is crucial for assisting clinical diagnosis. However, current automated segmentation methods often focus on specific sequences, specific organs, and specific cancers, i. e. , lack of generality. To address this issue, we propose a universal segmentation network for multi-sequence MRI (UniMRISegNet) that can segment multiple organs and cancers. UniMRISegNet features a shared encoder-decoder architecture equipped with contextual prompt generation (CPG) and prompt-conditioned dynamic convolution (PCDC) modules. The CPG module encodes sequence-specific, position-specific, and organ/cancer-specific text prompts as prior information to inform UniMRISegNet about the specific task to be executed. The PCDC module can adaptively generate model weights based on the assigned prompts, enhancing the segmentation capabilities of the UniMRISegNet for specific tasks. To mitigate discrepancies between different sequences of the same organ and capture similarities between related sequences, we design a novel loss function called Semantic-Aware Cosine Similarity Loss (SACSL), which integrates the cosine similarity of text embeddings to reconcile discrepancies and similarities between MRI sequences of the same organ. We created a large-scale annotated multi-sequence, multi-organ, and multi-cancer segmentation workflow (MSOCS), and demonstrated that our UniMRISegNet outperforms other universal networks and single-task networks on MSOCS. Furthermore, the universal weights from MSOCS can be transferred to never-before-seen downstream tasks, achieving superior performance compared to training from scratch.

JBHI Journal 2024 Journal Article

Weakly Supervised Classification for Nasopharyngeal Carcinoma With Transformer in Whole Slide Images

  • Ziwei Hu
  • Jianchao Wang
  • Qinquan Gao
  • Zhida Wu
  • Hanchuan Xu
  • Zhechen Guo
  • Jiawei Quan
  • Lihua Zhong

Pathological examination of nasopharyngeal carcinoma (NPC) is an indispensable factor for diagnosis, guiding clinical treatment and judging prognosis. Traditional and fully supervised NPC diagnosis algorithms require manual delineation of regions of interest on the gigapixel of whole slide images (WSIs), which however is laborious and often biased. In this paper, we propose a weakly supervised framework based on Tokens-to-Token Vision Transformer (WS-T2T-ViT) for accurate NPC classification with only a slide-level label. The label of tile images is inherited from their slide-level label. Specifically, WS-T2T-ViT is composed of the multi-resolution pyramid, T2T-ViT and multi-scale attention module. The multi-resolution pyramid is designed for imitating the coarse-to-fine process of manual pathological analysis to learn features from different magnification levels. The T2T module captures the local and global features to overcome the lack of global information. The multi-scale attention module improves classification performance by weighting the contributions of different granularity levels. Extensive experiments are performed on the 802-patient NPC and CAMELYON16 dataset. WS-T2T-ViT achieves an area under the receiver operating characteristic curve (AUC) of 0. 989 for NPC classification on the NPC dataset. The experiment results of CAMELYON16 dataset demonstrate the robustness and generalizability of WS-T2T-ViT in WSI-level classification.

JBHI Journal 2023 Journal Article

CUSS-Net: A Cascaded Unsupervised-Based Strategy and Supervised Network for Biomedical Image Diagnosis and Segmentation

  • Xiaogen Zhou
  • Zhiqiang Li
  • Yuyang Xue
  • Shun Chen
  • Meijuan Zheng
  • Cong Chen
  • Yue Yu
  • Xingqing Nie

Biomedical image segmentation and classification are critical components in a computer-aided diagnosis system. However, various deep convolutional neural networks are trained by a single task, ignoring the potential contribution of mutually performing multiple tasks. In this paper, we propose a cascaded unsupervised-based strategy to boost the supervised CNN framework for automated white blood cell (WBC) and skin lesion segmentation and classification, called CUSS-Net. Our proposed CUSS-Net consists of an unsupervised-based strategy (US) module, an enhanced segmentation network named E-SegNet, and a mask-guided classification network called MG-ClsNet. On the one hand, the proposed US module produces coarse masks that provide a prior localization map for the proposed E-SegNet to enhance it in locating and segmenting a target object accurately. On the other hand, the enhanced coarse masks predicted by the proposed E-SegNet are then fed into the proposed MG-ClsNet for accurate classification. Moreover, a novel cascaded dense inception module is presented to capture more high-level information. Meanwhile, we adopt a hybrid loss by combining a dice loss and a cross-entropy loss to alleviate the imbalance training problem. We evaluate our proposed CUSS-Net on three public medical image datasets. Experiments show that our proposed CUSS-Net outperforms representative state-of-the-art approaches.

JBHI Journal 2023 Journal Article

Dual-Input Transformer: An End-to-End Model for Preoperative Assessment of Pathological Complete Response to Neoadjuvant Chemotherapy in Breast Cancer Ultrasonography

  • Tong Tong
  • Dongyang Li
  • Jionghui Gu
  • Guo Chen
  • Guotao Bai
  • Xin Yang
  • Kun Wang
  • Tianan Jiang

Neoadjuvant chemotherapy (NAC) is the primary method to reduce the burden of tumor and metastasis; in the treatment of breast cancer, it may provide additional opportunities for breast-conserving surgery. Preoperative assessment of pathological complete response (PCR) to NAC is important for developing individualized treatment approaches and predicting patient prognosis. Compared to magnetic resonance imaging (MRI) and mammography, ultrasonography (US) has the advantages of simplicity, flexibility, and real-time imaging. Moreover, it does not require radiation and can provide multi-time acquisition of the tumor during NAC treatment. Recently, deep learning radiomics models based on multi-time-point US images for the prediction of NAC effectiveness have been proposed. To further improve the prediction performance, we carefully designed four supporting modules for our proposed dual-input transformer (DiT): isolated tokens-to-token patch embedding module, shared position embedding, time embedding, and weighted average pooling feature representation modules. The design of each module considers the characteristics of the US images at multiple time points. We validated our model on our retrospective US dataset composed of 484 cases from two centers whose consistency is not sufficiently high. Patients were allocated to training (n = 297), validation (n = 99), and external test (n = 88) sets. The results show that our model can achieve better performance than the Siamese CNN and the standard tokens-to-token vision transformer without using multi-time-point images. The ablation study also proved the effectiveness of each module designed for DiT.

YNICL Journal 2017 Journal Article

Five-class differential diagnostics of neurodegenerative diseases using random undersampling boosting

  • Tong Tong
  • Christian Ledig
  • Ricardo Guerrero
  • Andreas Schuh
  • Juha Koikkalainen
  • Antti Tolonen
  • Hanneke Rhodius
  • Frederik Barkhof

Differentiating between different types of neurodegenerative diseases is not only crucial in clinical practice when treatment decisions have to be made, but also has a significant potential for the enrichment of clinical trials. The purpose of this study is to develop a classification framework for distinguishing the four most common neurodegenerative diseases, including Alzheimer's disease, frontotemporal lobe degeneration, Dementia with Lewy bodies and vascular dementia, as well as patients with subjective memory complaints. Different biomarkers including features from images (volume features, region-wise grading features) and non-imaging features (CSF measures) were extracted for each subject. In clinical practice, the prevalence of different dementia types is imbalanced, posing challenges for learning an effective classification model. Therefore, we propose the use of the RUSBoost algorithm in order to train classifiers and to handle the class imbalance training problem. Furthermore, a multi-class feature selection method based on sparsity is integrated into the proposed framework to improve the classification performance. It also provides a way for investigating the importance of different features and regions. Using a dataset of 500 subjects, the proposed framework achieved a high accuracy of 75.2% with a balanced accuracy of 69.3% for the five-class classification using ten-fold cross validation, which is significantly better than the results using support vector machine or random forest, demonstrating the feasibility of the proposed framework to support clinical decision making.

YNIMG Journal 2017 Journal Article

Functional density and edge maps: Characterizing functional architecture in individuals and improving cross-subject registration

  • Tong Tong
  • Iman Aganj
  • Tian Ge
  • Jonathan R. Polimeni
  • Bruce Fischl

Population-level inferences and individual-level analyses are two important aspects in functional magnetic resonance imaging (fMRI) studies. Extracting reliable and informative features from fMRI data that capture biologically meaningful inter-subject variation is critical for aligning and comparing functional networks across subjects, and connecting the properties of functional brain organization with variations in behavior, cognition and genetics. In this study, we derive two new measures, which we term functional density map and edge map, and demonstrate their usefulness in characterizing the function of individual brains. Specifically, using data from the Human Connectome Project (HCP), we show that (1) both functional maps capture intrinsic properties of the functional connectivity pattern in individuals while exhibiting large variation across subjects; (2) functional maps derived from either resting-state or task-evoked fMRI can be used to accurately identify subjects from a population; and (3) cross-subject alignment using these functional maps considerably reduces functional variation and improves functional correspondence across subjects over state-of-the-art multimodal registration algorithms. Our results suggest that the proposed functional density and edge maps are promising features in characterizing the functional architecture in individuals and provide an alternative way to explore the functional variation across subjects.

YNICL Journal 2016 Journal Article

Differential diagnosis of neurodegenerative diseases using structural MRI data

  • Juha Koikkalainen
  • Hanneke Rhodius-Meester
  • Antti Tolonen
  • Frederik Barkhof
  • Betty Tijms
  • Afina W. Lemstra
  • Tong Tong
  • Ricardo Guerrero

Different neurodegenerative diseases can cause memory disorders and other cognitive impairments. The early detection and the stratification of patients according to the underlying disease are essential for an efficient approach to this healthcare challenge. This emphasizes the importance of differential diagnostics. Most studies compare patients and controls, or Alzheimer's disease with one other type of dementia. Such a bilateral comparison does not resemble clinical practice, where a clinician is faced with a number of different possible types of dementia. Here we studied which features in structural magnetic resonance imaging (MRI) scans could best distinguish four types of dementia, Alzheimer's disease, frontotemporal dementia, vascular dementia, and dementia with Lewy bodies, and control subjects. We extracted an extensive set of features quantifying volumetric and morphometric characteristics from T1 images, and vascular characteristics from FLAIR images. Classification was performed using a multi-class classifier based on Disease State Index methodology. The classifier provided continuous probability indices for each disease to support clinical decision making. A dataset of 504 individuals was used for evaluation. The cross-validated classification accuracy was 70.6% and balanced accuracy was 69.1% for the five disease groups using only automatically determined MRI features. Vascular dementia patients could be detected with high sensitivity (96%) using features from FLAIR images. Controls (sensitivity 82%) and Alzheimer's disease patients (sensitivity 74%) could be accurately classified using T1-based features, whereas the most difficult group was the dementia with Lewy bodies (sensitivity 32%). These results were notable better than the classification accuracies obtained with visual MRI ratings (accuracy 44.6%, balanced accuracy 51.6%). Different quantification methods provided complementary information, and consequently, the best results were obtained by utilizing several quantification methods. The results prove that automatic quantification methods and computerized decision support methods are feasible for clinical practice and provide comprehensive information that may help clinicians in the diagnosis making.

YNIMG Journal 2015 Journal Article

Standardized evaluation of algorithms for computer-aided diagnosis of dementia based on structural MRI: The CADDementia challenge

  • Esther E. Bron
  • Marion Smits
  • Wiesje M. van der Flier
  • Hugo Vrenken
  • Frederik Barkhof
  • Philip Scheltens
  • Janne M. Papma
  • Rebecca M.E. Steketee

Algorithms for computer-aided diagnosis of dementia based on structural MRI have demonstrated high performance in the literature, but are difficult to compare as different data sets and methodology were used for evaluation. In addition, it is unclear how the algorithms would perform on previously unseen data, and thus, how they would perform in clinical practice when there is no real opportunity to adapt the algorithm to the data at hand. To address these comparability, generalizability and clinical applicability issues, we organized a grand challenge that aimed to objectively compare algorithms based on a clinically representative multi-center data set. Using clinical practice as the starting point, the goal was to reproduce the clinical diagnosis. Therefore, we evaluated algorithms for multi-class classification of three diagnostic groups: patients with probable Alzheimer's disease, patients with mild cognitive impairment and healthy controls. The diagnosis based on clinical criteria was used as reference standard, as it was the best available reference despite its known limitations. For evaluation, a previously unseen test set was used consisting of 354 T1-weighted MRI scans with the diagnoses blinded. Fifteen research teams participated with a total of 29 algorithms. The algorithms were trained on a small training set (n=30) and optionally on data from other sources (e. g. , the Alzheimer's Disease Neuroimaging Initiative, the Australian Imaging Biomarkers and Lifestyle flagship study of aging). The best performing algorithm yielded an accuracy of 63. 0% and an area under the receiver-operating-characteristic curve (AUC) of 78. 8%. In general, the best performances were achieved using feature extraction based on voxel-based morphometry or a combination of features that included volume, cortical thickness, shape and intensity. The challenge is open for new submissions via the web-based framework: http: //caddementia. grand-challenge. org.

YNIMG Journal 2013 Journal Article

Segmentation of MR images via discriminative dictionary learning and sparse coding: Application to hippocampus labeling

  • Tong Tong
  • Robin Wolz
  • Pierrick Coupé
  • Joseph V. Hajnal
  • Daniel Rueckert

We propose a novel method for the automatic segmentation of brain MRI images by using discriminative dictionary learning and sparse coding techniques. In the proposed method, dictionaries and classifiers are learned simultaneously from a set of brain atlases, which can then be used for the reconstruction and segmentation of an unseen target image. The proposed segmentation strategy is based on image reconstruction, which is in contrast to most existing atlas-based labeling approaches that rely on comparing image similarities between atlases and target images. In addition, we propose a Fixed Discriminative Dictionary Learning for Segmentation (F-DDLS) strategy, which can learn dictionaries offline and perform segmentations online, enabling a significant speed-up in the segmentation stage. The proposed method has been evaluated for the hippocampus segmentation of 80 healthy ICBM subjects and 202 ADNI images. The robustness of the proposed method, especially of our F-DDLS strategy, was validated by training and testing on different subject groups in the ADNI database. The influence of different parameters was studied and the performance of the proposed method was also compared with that of the nonlocal patch-based approach. The proposed method achieved a median Dice coefficient of 0. 879 on 202 ADNI images and 0. 890 on 80 ICBM subjects, which is competitive compared with state-of-the-art methods.

v2026.09.13