Arrow Research search

Author name cluster

Weidong Cai

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

26 papers
1 author row

Possible papers

26

YNIMG Journal 2026 Journal Article

DeepMultiConnectome: Deep multi-task prediction of structural connectomes directly from diffusion MRI tractography

  • Marcus J. Vroemen
  • Yuqian Chen
  • Yui Lo
  • Tengfei Xue
  • Weidong Cai
  • Fan Zhang
  • Josien P.W. Pluim
  • Lauren J. O'Donnell

Diffusion MRI (dMRI) tractography enables in vivo mapping of brain structural connections, but traditional connectome generation is time-consuming and requires gray matter parcellation, posing challenges for large-scale studies. We introduce DeepMultiConnectome, a deep-learning model that predicts structural connectomes directly from tractography, bypassing the need for gray matter parcellation while supporting multiple parcellation schemes. Using a point-cloud-based neural network with multi-task learning, the model classifies streamlines according to their connected regions across two parcellation schemes, sharing a learned representation. By classifying individual streamlines, our method's output serves as a flexible prerequisite for constructing a wide range of differently weighted connectomes. We train and validate DeepMultiConnectome on tractography from the Human Connectome Project Young Adult dataset (N = 1000), labeled with an 84 and 164 region gray matter parcellation scheme. DeepMultiConnectome predicts multiple structural connectomes from a 3-million-streamline tractogram in ∼40 seconds. DeepMultiConnectome is evaluated by comparing predicted connectomes with traditional connectomes generated using the conventional method of labeling streamlines using a gray matter parcellation. The predicted connectomes show high agreement with traditionally generated connectomes across two parcellation schemes and multiple weighting strategies, and largely preserve network properties. Pearson correlations were r = 0.992 and 0.986 for streamline-count-weighted connectomes, r = 0.995 and 0.992 for SIFT2-weighted connectomes, and r = 0.775 and 0.727 for mean-FA-weighted connectomes. Test-retest analysis and downstream predictions of age and cognitive function demonstrate performance and reproducibility comparable to traditionally generated connectomes. Overall, DeepMultiConnectome provides a fast and scalable model for generating subject-specific connectomes across multiple parcellation and weighting schemes.

JBHI Journal 2026 Journal Article

Edge Extension for Missing Anatomical Features: A Mask-Guided Spatial Diffusion Framework for Ultrasound Scoliosis Image Outpainting

  • Chen Zhang
  • Wei Guo
  • De Yang
  • Weidong Cai
  • Yongping Zheng
  • Sai Ho Ling

Accurate scoliosis diagnosis relies on precise spinal curvature measurement, traditionally using radiographic Cobb's angle. Ultrasound imaging offers a radiation-free alternative via ultrasound curve angle (UCA) estimation, but its clinical utility is limited by incomplete anatomical information due to the restricted field of view (FOV) during scanning. This hinders key tasks like segmentation and landmark detection, restricting ultrasound's broader adoption in scoliosis assessment. To address this challenge, we propose an edge-aware outpainting diffusion framework that restores missing spinal anatomy by integrating mask-guided spatial diffusion. Specifically, the model is trained to predict noise between randomly selected target windows and anchor regions using spatially encoded masks. During inference, a dedicated edge-preservation mechanism guides the generation of anatomical structures. In addition to mitigating hallucinations in diffusion-based generation and ensuring perceptual consistency between generated and retained regions, we incorporate a total variation loss to enforce structural smoothness and coherence across the entire output. This approach effectively constrains reconstruction within masked regions, improving the recovery of anatomical features—particularly in cases of complex S-shaped spinal deformities commonly seen in scoliosis ultrasound imaging, where the field of view is inherently limited. Extensive experiments demonstrate our approach achieves the lowest Fréchet Inception Distance (180. 97) and highest Inception Score (1. 87 $\pm$ 0. 12), while improving thoracic and lumbar UCA estimation accuracy by 47. 8% and 24. 6%, respectively. Thoracic structure detection also increases by 8. 1% compared to Swin-Unet. Low KL divergence and Wasserstein distance confirm strong distributional alignment between generated and real anatomy. Overall, our framework enables anatomically consistent outpainting under limited FOV, enhancing ultrasound's reliability for clinical scoliosis assessment.

AAAI Conference 2026 Conference Paper

Gotta Hear Them All: Towards Sound Source Aware Audio Generation

  • Wei Guo
  • Heng Wang
  • Jianbo Ma
  • Weidong Cai

Audio synthesis has broad applications in multimedia. Recent advancements have made it possible to generate relevant audios from inputs describing an audio scene, such as images or texts. However, the immersiveness and expressiveness of the generation are limited. One possible problem is that existing methods solely rely on the global scene and overlook details of local sounding objects (i.e., sound sources). To address this issue, we propose a Sound Source-Aware Audio (SS2A) generator. SS2A is able to locally perceive multimodal sound sources from a scene with visual detection and cross-modality translation. It then contrastively learns a Cross-Modal Sound Source (CMSS) Manifold to semantically disambiguate each source. Finally, we attentively mix their CMSS semantics into a rich audio representation, from which a pretrained audio generator outputs the sound. To model the CMSS manifold, we curate a novel single-sound-source visual-audio dataset VGGS3 from VGGSound. We also design a Sound Source Matching Score to clearly measure localized audio relevance. With the effectiveness of explicit sound source modeling, SS2A achieves state-of-the-art performance in extensive image-to-audio tasks. We also qualitatively demonstrate SS2A's ability to achieve intuitive synthesis control by compositing vision, text, and audio conditions. Furthermore, we show that our sound source modeling can achieve competitive video-to-audio performance with a straightforward temporal aggregation mechanism.

AAAI Conference 2026 Conference Paper

HiFusion: Hierarchical Intra-Spot Alignment and Regional Context Fusion for Spatial Gene Expression Prediction from Histopathology

  • Ziqiao Weng
  • Yaoyu Fang
  • Jiahe Qian
  • Xinkun Wang
  • Lee A D Cooper
  • Weidong Cai
  • Bo Zhou

Spatial transcriptomics (ST) bridges gene expression and tissue morphology but faces clinical adoption barriers due to technical complexity and prohibitive costs. While computational methods predict gene expression from H&E-stained whole-slide images (WSIs), existing approaches often fail to capture the intricate biological heterogeneity within spots and are susceptible to morphological noise when integrating contextual information from surrounding tissue. To overcome these limitations, we propose HiFusion, a novel deep learning framework that integrates two complementary components. First, we introduce the Hierarchical Intra-Spot Modeling module that extracts fine-grained morphological representations through multi-resolution sub-patch decomposition, guided by a feature alignment loss to ensure semantic consistency across scales. Concurrently, we present the Context-aware Cross-scale Fusion module, which employs cross-attention to selectively incorporate biologically relevant regional context, thereby enhancing representational capacity. This architecture enables comprehensive modeling of both cellular-level features and tissue microenvironmental cues, which are essential for accurate gene expression prediction. Extensive experiments on two benchmark ST datasets demonstrate that HiFusion achieves state-of-the-art performance across both 2D slide-wise cross-validation and more challenging 3D sample-specific scenarios. These results underscore HiFusion’s potential as a robust, accurate, and scalable solution for ST inference from routine histopathology.

YNICL Journal 2026 Journal Article

Study of sex differences in the whole brain white matter using diffusion MRI tractography and suprathreshold fiber cluster statistics

  • Fan Zhang
  • Jarrett Rushmore
  • Yijie Li
  • Suheyla Cetin-Karayumak
  • Yang Song
  • Weidong Cai
  • Carl-Fredrik Westin
  • James J. Levitt

Sex-specific characteristics demonstrate a substantial influence on the brain white matter, suggesting distinct structural connectivity patterns between females and males. Diffusion MRI tractography is an important tool in assessing white matter connectivity and brain tissue microstructure across different populations. Whole brain tractography analysis for group statistical comparison is a challenging task due to the large number of white matter connections. This work studies whole-brain white matter connectivity differences between females and males using dMRI tractography. We study a large cohort of 707 subjects from the Human Connectome Project Young Adult dataset. By applying a well-established fiber clustering pipeline and a suprathreshold fiber cluster statistical method, we analyze tracts in the cerebral cortex and understudied pathways like those connecting to the cerebellum. We identify several tracts with significant sex differences in terms of their fractional anisotropy and/or mean diffusivity. These include deep tracts like the arcuate fasciculus, corticospinal tract, and corpus callosum, superficial tracts in the frontal lobe, and cerebellar tracts. Finally, a canonical correlation analysis (CCA) that identifies covariance patterns between white matter and behavior measures reveals that these white matter differences are associated with a range of neurobehavioral measures, with the strongest and most consistent associations observed for motor function, suggesting motor circuits as a potential key focus for future research.

AAAI Conference 2026 Conference Paper

UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning

  • Tiancheng Gu
  • Kaicheng Yang
  • Kaichen Zhang
  • Xiang An
  • Ziyong Feng
  • Yueyi Zhang
  • Weidong Cai
  • Jiankang Deng

Universal multimodal embedding models are essential in various tasks. Existing approaches typically use in-batch mining to identify hard negatives by measuring the similarity of query-candidate pairs. However, these methods often struggle to capture subtle semantic differences among candidates and lack diversity in negative samples. Moreover, the embeddings exhibit limited discriminative ability in distinguishing false and hard negatives. In this paper, we leverage the advanced understanding capabilities of MLLMs to enhance representation learning, and present a novel Universal Multimodal Embedding(UniME-V2) model. Our approach first constructs a potential hard negative set through global retrieval. We then introduce the MLLM-as-a-Judge mechanism, which utilizes MLLMs to assess the semantic alignment of query-candidate pairs and generate soft semantic matching scores. These scores serve as a foundation for hard negative mining, mitigating the impact of false negatives and enabling the identification of diverse, high-quality hard negatives. Furthermore, the semantic matching scores are used as soft labels to mitigate the rigid one-to-one mapping constraint. By aligning the similarity matrix with the soft semantic matching score matrix, the model learns semantic distinctions among candidates, significantly enhancing its discriminative capacity. To further improve performance, we propose UniME-V2, a reranking model trained on our mined hard negatives through a joint pairwise and listwise optimization approach. We conduct comprehensive experiments on the MMEB benchmark and multiple retrieval tasks, demonstrating that our method achieves state-of-the-art performance across all tasks.

AAAI Conference 2025 Conference Paper

CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination

  • Kaicheng Yang
  • Tiancheng Gu
  • Xiang An
  • Haiqiang Jiang
  • Xiangzi Dai
  • Ziyong Feng
  • Weidong Cai
  • Jiankang Deng

Contrastive Language-Image Pre-training (CLIP) has achieved excellent performance over a wide range of tasks. However, the effectiveness of CLIP heavily relies on a substantial corpus of pre-training data, resulting in notable consumption of computational resources. Although knowledge distillation has been widely applied in single modality models, how to efficiently expand knowledge distillation to vision-language foundation models with extensive data remains relatively unexplored. In this paper, we introduce CLIP-CID, a novel distillation mechanism that effectively transfers knowledge from a large vision-language foundation model to a smaller model. We initially propose a simple but efficient image semantic balance method to reduce transfer learning bias and improve distillation efficiency. This method filters out 43.7% of image-text pairs from the LAION400M while maintaining superior performance. After that, we leverage cluster-instance discrimination to facilitate knowledge transfer from the teacher model to the student model, thereby empowering the student model to acquire a holistic semantic comprehension of the pre-training data. Experimental results demonstrate that CLIP-CID achieves state-of-the-art performance on various downstream tasks including linear probe and zero-shot classification.

AIIM Journal 2024 Journal Article

Improving multiple sclerosis lesion segmentation across clinical sites: A federated learning approach with noise-resilient training

  • Lei Bai
  • Dongang Wang
  • Hengrui Wang
  • Michael Barnett
  • Mariano Cabezas
  • Weidong Cai
  • Fernando Calamante
  • Kain Kyle

Accurately measuring the evolution of Multiple Sclerosis (MS) with magnetic resonance imaging (MRI) critically informs understanding of disease progression and helps to direct therapeutic strategy. Deep learning models have shown promise for automatically segmenting MS lesions, but the scarcity of accurately annotated data hinders progress in this area. Obtaining sufficient data from a single clinical site is challenging and does not address the heterogeneous need for model robustness. Conversely, the collection of data from multiple sites introduces data privacy concerns and potential label noise due to varying annotation standards. To address this dilemma, we explore the use of the federated learning framework while considering label noise. Our approach enables collaboration among multiple clinical sites without compromising data privacy under a federated learning paradigm that incorporates a noise-robust training strategy based on label correction. Specifically, we introduce a Decoupled Hard Label Correction (DHLC) strategy that considers the imbalanced distribution and fuzzy boundaries of MS lesions, enabling the correction of false annotations based on prediction confidence. We also introduce a Centrally Enhanced Label Correction (CELC) strategy, which leverages the aggregated central model as a correction teacher for all sites, enhancing the reliability of the correction process. Extensive experiments conducted on two multi-site datasets demonstrate the effectiveness and robustness of our proposed methods, indicating their potential for clinical applications in multi-site collaborations to train better deep learning models with lower cost in data collection and annotation.

AAAI Conference 2024 Conference Paper

PaintHuman: Towards High-Fidelity Text-to-3D Human Texturing via Denoised Score Distillation

  • Jianhui Yu
  • Hao Zhu
  • Liming Jiang
  • Chen Change Loy
  • Weidong Cai
  • Wayne Wu

Recent advances in zero-shot text-to-3D human generation, which employ the human model prior (e.g., SMPL) or Score Distillation Sampling (SDS) with pre-trained text-to-image diffusion models, have been groundbreaking. However, SDS may provide inaccurate gradient directions under the weak diffusion guidance, as it tends to produce over-smoothed results and generate body textures that are inconsistent with the detailed mesh geometry. Therefore, directly leveraging existing strategies for high-fidelity text-to-3D human texturing is challenging. In this work, we propose a model called PaintHuman to addresses the challenges from two perspectives. We first propose a novel score function, Denoised Score Distillation (DSD), which directly modifies the SDS by introducing negative gradient components to iteratively correct the gradient direction and generate high-quality textures. In addition, we use the depth map as a geometric guide to ensure that the texture is semantically aligned to human mesh surfaces. To guarantee the quality of rendered results, we employ geometry-aware networks to predict surface materials and render realistic human textures. Extensive experiments, benchmarked against state-of-the-art (SoTA) methods, validate the efficacy of our approach.Project page: https://painthuman.github.io/.

AAAI Conference 2024 Conference Paper

V2A-Mapper: A Lightweight Solution for Vision-to-Audio Generation by Connecting Foundation Models

  • Heng Wang
  • Jianbo Ma
  • Santiago Pascual
  • Richard Cartwright
  • Weidong Cai

Building artificial intelligence (AI) systems on top of a set of foundation models (FMs) is becoming a new paradigm in AI research. Their representative and generative abilities learnt from vast amounts of data can be easily adapted and transferred to a wide range of downstream tasks without extra training from scratch. However, leveraging FMs in cross-modal generation remains under-researched when audio modality is involved. On the other hand, automatically generating semantically-relevant sound from visual input is an important problem in cross-modal generation studies. To solve this vision-to-audio (V2A) generation problem, existing methods tend to design and build complex systems from scratch using modestly sized datasets. In this paper, we propose a lightweight solution to this problem by leveraging foundation models, specifically CLIP, CLAP, and AudioLDM. We first investigate the domain gap between the latent space of the visual CLIP and the auditory CLAP models. Then we propose a simple yet effective mapper mechanism (V2A-Mapper) to bridge the domain gap by translating the visual input between CLIP and CLAP spaces. Conditioned on the translated CLAP embedding, pretrained audio generative FM AudioLDM is adopted to produce high-fidelity and visually-aligned sound. Compared to previous approaches, our method only requires a quick training of the V2A-Mapper. We further analyze and conduct extensive experiments on the choice of the V2A-Mapper and show that a generative mapper is better at fidelity and variability (FD) while a regression mapper is slightly better at relevance (CS). Both objective and subjective evaluation on two V2A datasets demonstrate the superiority of our proposed method compared to current state-of-the-art approaches - trained with 86% fewer parameters but achieving 53% and 19% improvement in FD and CS, respectively. Supplementary materials such as audio samples are provided at our demo website: https://v2a-mapper.github.io/.

YNIMG Journal 2023 Journal Article

Deep fiber clustering: Anatomically informed fiber clustering with self-supervised deep learning for fast and effective tractography parcellation

  • Yuqian Chen
  • Chaoyi Zhang
  • Tengfei Xue
  • Yang Song
  • Nikos Makris
  • Yogesh Rathi
  • Weidong Cai
  • Fan Zhang

White matter fiber clustering is an important strategy for white matter parcellation, which enables quantitative analysis of brain connections in health and disease. In combination with expert neuroanatomical labeling, data-driven white matter fiber clustering is a powerful tool for creating atlases that can model white matter anatomy across individuals. While widely used fiber clustering approaches have shown good performance using classical unsupervised machine learning techniques, recent advances in deep learning reveal a promising direction toward fast and effective fiber clustering. In this work, we propose a novel deep learning framework for white matter fiber clustering, Deep Fiber Clustering (DFC), which solves the unsupervised clustering problem as a self-supervised learning task with a domain-specific pretext task to predict pairwise fiber distances. This process learns a high-dimensional embedding feature representation for each fiber, regardless of the order of fiber points reconstructed during tractography. We design a novel network architecture that represents input fibers as point clouds and allows the incorporation of additional sources of input information from gray matter parcellation. Thus, DFC makes use of combined information about white matter fiber geometry and gray matter anatomy to improve the anatomical coherence of fiber clusters. In addition, DFC conducts outlier removal naturally by rejecting fibers with low cluster assignment probability. We evaluate DFC on three independently acquired cohorts, including data from 220 individuals across genders, ages (young and elderly adults), and different health conditions (healthy control and multiple neuropsychiatric disorders). We compare DFC to several state-of-the-art white matter fiber clustering algorithms. Experimental results demonstrate superior performance of DFC in terms of cluster compactness, generalization ability, anatomical coherence, and computational efficiency.

AAAI Conference 2023 Conference Paper

PaRot: Patch-Wise Rotation-Invariant Network via Feature Disentanglement and Pose Restoration

  • Dingxin Zhang
  • Jianhui Yu
  • Chaoyi Zhang
  • Weidong Cai

Recent interest in point cloud analysis has led rapid progress in designing deep learning methods for 3D models. However, state-of-the-art models are not robust to rotations, which remains an unknown prior to real applications and harms the model performance. In this work, we introduce a novel Patch-wise Rotation-invariant network (PaRot), which achieves rotation invariance via feature disentanglement and produces consistent predictions for samples with arbitrary rotations. Specifically, we design a siamese training module which disentangles rotation invariance and equivariance from patches defined over different scales, e.g., the local geometry and global shape, via a pair of rotations. However, our disentangled invariant feature loses the intrinsic pose information of each patch. To solve this problem, we propose a rotation-invariant geometric relation to restore the relative pose with equivariant information for patches defined over different scales. Utilising the pose information, we propose a hierarchical module which implements intra-scale and inter-scale feature aggregation for 3D shape learning. Moreover, we introduce a pose-aware feature propagation process with the rotation-invariant relative pose information embedded. Experiments show that our disentanglement module extracts high-quality rotation-robust features and the proposed lightweight model achieves competitive results in rotated 3D object classification and part segmentation tasks.

AAAI Conference 2023 Conference Paper

Rethinking Rotation Invariance with Point Cloud Registration

  • Jianhui Yu
  • Chaoyi Zhang
  • Weidong Cai

Recent investigations on rotation invariance for 3D point clouds have been devoted to devising rotation-invariant feature descriptors or learning canonical spaces where objects are semantically aligned. Examinations of learning frameworks for invariance have seldom been looked into. In this work, we review rotation invariance (RI) in terms of point cloud registration (PCR) and propose an effective framework for rotation invariance learning via three sequential stages, namely rotation-invariant shape encoding, aligned feature integration, and deep feature registration. We first encode shape descriptors constructed with respect to reference frames defined over different scales, e.g., local patches and global topology, to generate rotation-invariant latent shape codes. Within the integration stage, we propose an Aligned Integration Transformer (AIT) to produce a discriminative feature representation by integrating point-wise self- and cross-relations established within the shape codes. Meanwhile, we adopt rigid transformations between reference frames to align the shape codes for feature consistency across different scales. Finally, the deep integrated feature is registered to both rotation-invariant shape codes to maximize their feature similarities, such that rotation invariance of the integrated feature is preserved and shared semantic information is implicitly extracted from shape codes. Experimental results on 3D shape classification, part segmentation, and retrieval tasks prove the feasibility of our framework. Our project page is released at: https://rotation3d.github.io/.

YNIMG Journal 2022 Journal Article

Insights from an autism imaging biomarker challenge: Promises and threats to biomarker discovery

  • Nicolas Traut
  • Katja Heuer
  • Guillaume Lemaître
  • Anita Beggiato
  • David Germanaud
  • Monique Elmaleh
  • Alban Bethegnies
  • Laurent Bonnasse-Gahot

MRI has been extensively used to identify anatomical and functional differences in Autism Spectrum Disorder (ASD). Yet, many of these findings have proven difficult to replicate because studies rely on small cohorts and are built on many complex, undisclosed, analytic choices. We conducted an international challenge to predict ASD diagnosis from MRI data, where we provided preprocessed anatomical and functional MRI data from > 2,000 individuals. Evaluation of the predictions was rigorously blinded. 146 challengers submitted prediction algorithms, which were evaluated at the end of the challenge using unseen data and an additional acquisition site. On the best algorithms, we studied the importance of MRI modalities, brain regions, and sample size. We found evidence that MRI could predict ASD diagnosis: the 10 best algorithms reliably predicted diagnosis with AUC∼0.80 - far superior to what can be currently obtained using genotyping data in cohorts 20-times larger. We observed that functional MRI was more important for prediction than anatomical MRI, and that increasing sample size steadily increased prediction accuracy, providing an efficient strategy to improve biomarkers. We also observed that despite a strong incentive to generalise to unseen data, model development on a given dataset faces the risk of overfitting: performing well in cross-validation on the data at hand, but not generalising. Finally, we were able to predict ASD diagnosis on an external sample added after the end of the challenge (EU-AIMS), although with a lower prediction accuracy (AUC=0.72). This indicates that despite being based on a large multisite cohort, our challenge still produced biomarkers fragile in the face of dataset shifts.

YNIMG Journal 2022 Journal Article

Methylphenidate remediates aberrant brain network dynamics in children with attention‐deficit/hyperactivity disorder: A randomized controlled trial

  • Yoshifumi Mizuno
  • Weidong Cai
  • Kaustubh Supekar
  • Kai Makita
  • Shinichiro Takiguchi
  • Akemi Tomoda
  • Vinod Menon

Methylphenidate is a widely used first-line treatment for attention deficit/hyperactivity disorder (ADHD), but the underlying circuit mechanisms are poorly understood. Here we investigate whether a single dose of osmotic release oral system methylphenidate can remediate attention deficits and aberrancies in functional circuit dynamics in cognitive control networks, which have been implicated in ADHD. In a randomized placebo-controlled double-blind crossover design, 27 children with ADHD were scanned twice with resting-state functional MRI and sustained attention was examined using a continuous performance task under methylphenidate and placebo conditions; 49 matched typically-developing (TD) children were scanned once for comparison. Dynamic time-varying cross-network interactions between the salience (SN), frontoparietal (FPN), and default mode (DMN) networks were examined in children with ADHD under both administration conditions and compared with TD children. Methylphenidate improved sustained attention on a continuous performance task in children with ADHD, when compared to the placebo condition. Children with ADHD under placebo showed aberrancies in dynamic time-varying cross-network interactions between the SN, FPN and DMN, which were remediated by methylphenidate. Multivariate classification analysis confirmed that methylphenidate remediates aberrant dynamic brain network interactions. Furthermore, dynamic time-varying network interactions under placebo conditions predicted individual differences in methylphenidate-induced improvements in sustained attention in children with ADHD. These findings suggest that a single dose of methylphenidate can remediate deficits in sustained attention and aberrant brain circuit dynamics in cognitive control circuits in children with ADHD. Findings identify a novel brain circuit mechanism underlying a first-line pharmacological treatment for ADHD, and may inform clinically useful biomarkers for evaluating treatment outcomes.

JBHI Journal 2022 Journal Article

Multiple Sclerosis Lesion Analysis in Brain Magnetic Resonance Images: Techniques and Clinical Applications

  • Yang Ma
  • Chaoyi Zhang
  • Mariano Cabezas
  • Yang Song
  • Zihao Tang
  • Dongnan Liu
  • Weidong Cai
  • Michael Barnett

Multiple sclerosis (MS) is a chronic inflammatory and degenerative disease of the central nervous system, characterized by the appearance of focal lesions in the white and gray matter that topographically correlate with an individual patient’s neurological symptoms and signs. Magnetic resonance imaging (MRI) provides detailed in-vivo structural information, permitting the quantification and categorization of MS lesions that critically inform disease management. Traditionally, MS lesions have been manually annotated on 2D MRI slices, a process that is inefficient and prone to inter-/intra-observer errors. Recently, automated statistical imaging analysis techniques have been proposed to detect and segment MS lesions based on MRI voxel intensity. However, their effectiveness is limited by the heterogeneity of both MRI data acquisition techniques and the appearance of MS lesions. By learning complex lesion representations directly from images, deep learning techniques have achieved remarkable breakthroughs in the MS lesion segmentation task. Here, we provide a comprehensive review of state-of-the-art automatic statistical and deep-learning MS segmentation methods and discuss current and future clinical applications. Further, we review technical strategies, such as domain adaptation, to enhance MS lesion segmentation in real-world clinical settings.

IJCAI Conference 2022 Conference Paper

Spatiality-guided Transformer for 3D Dense Captioning on Point Clouds

  • Heng Wang
  • Chaoyi Zhang
  • Jianhui Yu
  • Weidong Cai

Dense captioning in 3D point clouds is an emerging vision-and-language task involving object-level 3D scene understanding. Apart from coarse semantic class prediction and bounding box regression as in traditional 3D object detection, 3D dense captioning aims at producing a further and finer instance-level label of natural language description on visual appearance and spatial relations for each scene object of interest. To detect and describe objects in a scene, following the spirit of neural machine translation, we propose a transformer-based encoder-decoder architecture, namely SpaCap3D, to transform objects into descriptions, where we especially investigate the relative spatiality of objects in 3D scenes and design a spatiality-guided encoder via a token-to-token spatial relation learning objective and an object-centric decoder for precise and spatiality-enhanced object caption generation. Evaluated on two benchmark datasets, ScanRefer and ReferIt3D, our proposed SpaCap3D outperforms the baseline method Scan2Cap by 4. 94% and 9. 61% in CIDEr@0. 5IoU, respectively. Our project page with source code and supplementary files is available at https: //SpaCap3D. github. io/.

AAAI Conference 2020 Conference Paper

Shape-Oriented Convolution Neural Network for Point Cloud Analysis

  • Chaoyi Zhang
  • Yang Song
  • Lina Yao
  • Weidong Cai

Point cloud is a principal data structure adopted for 3D geometric information encoding. Unlike other conventional visual data, such as images and videos, these irregular points describe the complex shape features of 3D objects, which makes shape feature learning an essential component of point cloud analysis. To this end, a shape-oriented message passing scheme dubbed ShapeConv is proposed to focus on the representation learning of the underlying shape formed by each local neighboring point. Despite this intra-shape relationship learning, ShapeConv is also designed to incorporate the contextual effects from the inter-shape relationship through capturing the long-ranged dependencies between local underlying shapes. This shape-oriented operator is stacked into our hierarchical learning architecture, namely Shape-Oriented Convolutional Neural Network (SOCNN), developed for point cloud analysis. Extensive experiments have been performed to evaluate its significance in the tasks of point cloud classification and part segmentation.

IJCAI Conference 2019 Conference Paper

Nuclei Segmentation via a Deep Panoptic Model with Semantic Feature Fusion

  • Dongnan Liu
  • Donghao Zhang
  • Yang Song
  • Chaoyi Zhang
  • Fan Zhang
  • Lauren O'Donnell
  • Weidong Cai

Automated detection and segmentation of individual nuclei in histopathology images is important for cancer diagnosis and prognosis. Due to the high variability of nuclei appearances and numerous overlapping objects, this task still remains challenging. Deep learning based semantic and instance segmentation models have been proposed to address the challenges, but these methods tend to concentrate on either the global or local features and hence still suffer from information loss. In this work, we propose a panoptic segmentation model which incorporates an auxiliary semantic segmentation branch with the instance branch to integrate global and local features. Furthermore, we design a feature map fusion mechanism in the instance branch and a new mask generator to prevent information loss. Experimental results on three different histopathology datasets demonstrate that our method outperforms the state-of-the-art nuclei segmentation methods and popular semantic and instance segmentation models by a large margin.

YNIMG Journal 2018 Journal Article

Suprathreshold fiber cluster statistics: Leveraging white matter geometry to enhance tractography statistical analysis

  • Fan Zhang
  • Weining Wu
  • Lipeng Ning
  • Gloria McAnulty
  • Deborah Waber
  • Borjan Gagoski
  • Kiera Sarill
  • Hesham M. Hamoda

This work presents a suprathreshold fiber cluster (STFC) method that leverages the whole brain fiber geometry to enhance statistical group difference analyses. The proposed method consists of 1) a well-established study-specific data-driven tractography parcellation to obtain white matter tract parcels and 2) a newly proposed nonparametric, permutation-test-based STFC method to identify significant differences between study populations. The basic idea of our method is that a white matter parcel's neighborhood (nearby parcels with similar white matter anatomy) can support the parcel's statistical significance when correcting for multiple comparisons. We propose an adaptive parcel neighborhood strategy to allow suprathreshold fiber cluster formation that is robust to anatomically varying inter-parcel distances. The method is demonstrated by application to a multi-shell diffusion MRI dataset from 59 individuals, including 30 attention deficit hyperactivity disorder patients and 29 healthy controls. Evaluations are conducted using both synthetic and in-vivo data. The results indicate that the STFC method gives greater sensitivity in finding group differences in white matter tract parcels compared to several traditional multiple comparison correction methods.

YNIMG Journal 2018 Journal Article

Whole brain white matter connectivity analysis using machine learning: An application to autism

  • Fan Zhang
  • Peter Savadjiev
  • Weidong Cai
  • Yang Song
  • Yogesh Rathi
  • Birkan Tunç
  • Drew Parker
  • Tina Kapur

In this paper, we propose an automated white matter connectivity analysis method for machine learning classification and characterization of white matter abnormality via identification of discriminative fiber tracts. The proposed method uses diffusion MRI tractography and a data-driven approach to find fiber clusters corresponding to subdivisions of the white matter anatomy. Features extracted from each fiber cluster describe its diffusion properties and are used for machine learning. The method is demonstrated by application to a pediatric neuroimaging dataset from 149 individuals, including 70 children with autism spectrum disorder (ASD) and 79 typically developing controls (TDC). A classification accuracy of 78. 33% is achieved in this cross-validation study. We investigate the discriminative diffusion features based on a two-tensor fiber tracking model. We observe that the mean fractional anisotropy from the second tensor (associated with crossing fibers) is most affected in ASD. We also find that local along-tract (central cores and endpoint regions) differences between ASD and TDC are helpful in differentiating the two groups. These altered diffusion properties in ASD are associated with multiple robustly discriminative fiber clusters, which belong to several major white matter tracts including the corpus callosum, arcuate fasciculus, uncinate fasciculus and aslant tract; and the white matter structures related to the cerebellum, brain stem, and ventral diencephalon. These discriminative fiber clusters, a small part of the whole brain tractography, represent the white matter connections that could be most affected in ASD. Our results indicate the potential of a machine learning pipeline based on white matter fiber clustering.

YNIMG Journal 2017 Journal Article

Bayesian switching factor analysis for estimating time-varying functional connectivity in fMRI

  • Jalil Taghia
  • Srikanth Ryali
  • Tianwen Chen
  • Kaustubh Supekar
  • Weidong Cai
  • Vinod Menon

There is growing interest in understanding the dynamical properties of functional interactions between distributed brain regions. However, robust estimation of temporal dynamics from functional magnetic resonance imaging (fMRI) data remains challenging due to limitations in extant multivariate methods for modeling time-varying functional interactions between multiple brain areas. Here, we develop a Bayesian generative model for fMRI time-series within the framework of hidden Markov models (HMMs). The model is a dynamic variant of the static factor analysis model (Ghahramani and Beal, 2000). We refer to this model as Bayesian switching factor analysis (BSFA) as it integrates factor analysis into a generative HMM in a unified Bayesian framework. In BSFA, brain dynamic functional networks are represented by latent states which are learnt from the data. Crucially, BSFA is a generative model which estimates the temporal evolution of brain states and transition probabilities between states as a function of time. An attractive feature of BSFA is the automatic determination of the number of latent states via Bayesian model selection arising from penalization of excessively complex models. Key features of BSFA are validated using extensive simulations on carefully designed synthetic data. We further validate BSFA using fingerprint analysis of multisession resting-state fMRI data from the Human Connectome Project (HCP). Our results show that modeling temporal dependencies in the generative model of BSFA results in improved fingerprinting of individual participants. Finally, we apply BSFA to elucidate the dynamic functional organization of the salience, central-executive, and default mode networks—three core neurocognitive systems with central role in cognitive and affective information processing (Menon, 2011). Across two HCP sessions, we demonstrate a high level of dynamic interactions between these networks and determine that the salience network has the highest temporal flexibility among the three networks. Our proposed methods provide a novel and powerful generative model for investigating dynamic brain connectivity.

NeurIPS Conference 2017 Conference Paper

Regularized Modal Regression with Applications in Cognitive Impairment Prediction

  • Xiaoqian Wang
  • Hong Chen
  • Weidong Cai
  • Dinggang Shen
  • Heng Huang

Linear regression models have been successfully used to function estimation and model selection in high-dimensional data analysis. However, most existing methods are built on least squares with the mean square error (MSE) criterion, which are sensitive to outliers and their performance may be degraded for heavy-tailed noise. In this paper, we go beyond this criterion by investigating the regularized modal regression from a statistical learning viewpoint. A new regularized modal regression model is proposed for estimation and variable selection, which is robust to outliers, heavy-tailed noise, and skewed noise. On the theoretical side, we establish the approximation estimate for learning the conditional mode function, the sparsity analysis for variable selection, and the robustness characterization. On the application side, we applied our model to successfully improve the cognitive impairment prediction using the Alzheimer’s Disease Neuroimaging Initiative (ADNI) cohort data.

AAAI Conference 2017 Conference Paper

Video Recovery via Learning Variation and Consistency of Images

  • Zhouyuan Huo
  • Shangqian Gao
  • Weidong Cai
  • Heng Huang

Matrix completion algorithms have been popularly used to recover images with missing entries, and they are proved to be very effective. Recent works utilized tensor completion models in video recovery assuming that all video frames are homogeneous and correlated. However, real videos are made up of different episodes or scenes, i. e. heterogeneous. Therefore, a video recovery model which utilizes both video spatiotemporal consistency and variation is necessary. To solve this problem, we propose a new video recovery method Sectional Trace Norm with Variation and Consistency Constraints (STN-VCC). In our model, capped 1-norm regularization is utilized to learn the spatial-temporal consistency and variation between consecutive frames in video clips. Meanwhile, we introduce a new low-rank model to capture the low-rank structure in video frames with a better approximation of rank minimization than traditional trace norm. An efficient optimization algorithm is proposed, and we also provide a proof of convergence in the paper. We evaluate the proposed method via several video recovery tasks and experiment results show that our new method consistently outperforms other related approaches.

NeurIPS Conference 2016 Conference Paper

Error Analysis of Generalized Nyström Kernel Regression

  • Hong Chen
  • Haifeng Xia
  • Heng Huang
  • Weidong Cai

Nystr\"{o}m method has been used successfully to improve the computational efficiency of kernel ridge regression (KRR). Recently, theoretical analysis of Nystr\"{o}m KRR, including generalization bound and convergence rate, has been established based on reproducing kernel Hilbert space (RKHS) associated with the symmetric positive semi-definite kernel. However, in real world applications, RKHS is not always optimal and kernel function is not necessary to be symmetric or positive semi-definite. In this paper, we consider the generalized Nystr\"{o}m kernel regression (GNKR) with $\ell_2$ coefficient regularization, where the kernel just requires the continuity and boundedness. Error analysis is provided to characterize its generalization performance and the column norm sampling is introduced to construct the refined hypothesis space. In particular, the fast learning rate with polynomial decay is reached for the GNKR. Experimental analysis demonstrates the satisfactory performance of GNKR with the column norm sampling.

YNIMG Journal 2012 Journal Article

Roles for the pre-supplementary motor area and the right inferior frontal gyrus in stopping action: Electrophysiological responses and functional and structural connectivity

  • Nicole C. Swann
  • Weidong Cai
  • Christopher R. Conner
  • Thomas A. Pieters
  • Michael P. Claffey
  • Jobi S. George
  • Adam R. Aron
  • Nitin Tandon

Both the pre-supplementary motor area (preSMA) and the right inferior frontal gyrus (rIFG) are important for stopping action outright. These regions are also engaged when preparing to stop. We aimed to elucidate the roles of these regions by harnessing the high spatio-temporal resolution of electrocorticography (ECoG), and by using a task that engages both preparing to stop and stopping outright. First, we validated the task using fMRI in 16 healthy control participants to confirm that both the preSMA and the rIFG were active. Next, we studied a rare patient with intracranial grid coverage of both these regions, using macrostimulation, diffusion tractography, cortico-cortical evoked potentials (CCEPs) and task-based ECoG. Macrostimulation of the preSMA induced behavioral motor arrest. Diffusion tractography revealed a structural connection between the preSMA and rIFG. CCEP analysis showed that stimulation of the preSMA evoked strong local field potentials within 30ms in rIFG. During the task, when preparing to stop, there was increased high gamma amplitude (~70–250Hz) in both regions, with preSMA preceding rIFG by ~750ms. For outright stopping there was also a high gamma amplitude increase in both regions, again with preSMA preceding rIFG. Further, at the time of stopping, there was an increase in beta band activity (~16Hz) in both regions, with significantly stronger inter-regional coherence for successful vs. unsuccessful stop trials. The results complement earlier reports of a structural/functional action control network between the preSMA and rIFG. They go further by revealing between-region timing differences in the high gamma band when preparing to stop and stopping outright. They also reveal strong between-region coherence in the beta band when stopping is successful. Implications for theories of action control are discussed.

v2026.09.13