Arrow Research search

Author name cluster

Geng Chen

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

19 papers
2 author rows

Possible papers

19

AAAI Conference 2026 Conference Paper

AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning

  • Jirong Zha
  • Yuxuan Fan
  • Tianyu Zhang
  • Geng Chen
  • Yingfeng Chen
  • Chen Gao
  • Xinlei Chen

Multimodal Large Language Models (MLLMs) have shown promise in single-agent vision tasks, yet benchmarks for evaluating multi-agent collaborative perception remain scarce. This gap is critical, as multi-drone systems provide enhanced coverage, robustness, and collaboration compared to single-sensor setups. Existing multi-image benchmarks mainly target basic perception tasks using high-quality single-agent images, thus failing to evaluate MLLMs in more complex, egocentric collaborative scenarios, especially under real-world degraded perception conditions. To address these challenges, we introduce AirCopBench, the first comprehensive benchmark designed to evaluate MLLMs in embodied aerial collaborative perception under challenging perceptual conditions. AirCopBench includes 14.6k+ questions derived from both simulator and real-world data, spanning four key task dimensions: Scene Understanding, Object Understanding, Perception Assessment, and Collaborative Decision, across 14 task types. We construct the benchmark using data from challenging degraded-perception scenarios with annotated collaborative events, generating large-scale questions through model-, rule-, and human-based methods under rigorous quality control. Evaluations on 40 MLLMs show significant performance gaps in collaborative perception tasks, with the best model trailing humans by 24.38% on average and exhibiting inconsistent results across tasks. Fine-tuning experiments further confirm the feasibility of sim-to-real transfer in aerial collaborative perception.

EAAI Journal 2026 Journal Article

Laplacian-guided contextual instance learning for whole slide image classification

  • Jian Chen
  • Ziyuan Chen
  • Geng Chen
  • Mengyu Liu
  • Sohaib Asif
  • He Zhang
  • Jun Jin

Classification plays an important role in the diagnosis and prognosis of cancers such as endometrial and breast cancer. Achieving satisfactory performance in classifying cancer molecular subtypes from whole slide images presents a substantial challenge. This difficulty arises from diverse and complex inter-instance relationships and feature homogeneity among different molecular subtypes. To address these issues, this paper presents a novel Laplacian-guided contextual instance learning (LapCIL) framework, which focuses on learning inter-instance relationships to effectively identify molecular subtypes. The LapCIL framework consists of a dynamic contiguous masking strategy, a contextual instance learning block, and a Laplacian channel classification head. In the LapCIL framework, a dynamic contiguous masking strategy is proposed to generate more inter-instance relationships from finite data. Considering the diversity and complexity of inter-instance relationships, a contextual instance learning block is introduced, which leverages a contextual self-attention mechanism to capture the relationships between different instances. To enhance the distinguishing capability between different instances even further, especially in scenarios where feature homogeneity renders it challenging to differentiate morphologically similar cell types, the LapCIL framework incorporates a Laplacian channel classification head. The Laplacian channel classification head focuses on structured local features and dynamically attends to discriminative channel groups. Extensive experiments are conducted on the CAncer MEtastases in LYmphnOdes challeNge (CAMELYON16) breast cancer dataset, the BReAst Carcinoma Subtyping dataset, and a clinical endometrial cancer dataset to evaluate the proposed LapCIL framework. Our framework achieves significant advantages over state-of-the-art methods, both on the clinical dataset and the CAMELYON16 breast cancer dataset.

JBHI Journal 2026 Journal Article

MuST: Multi-Scale Transformer Incorporating Hierarchical Attention and TCN for EEG Decoding

  • Kui Zhao
  • Enze Shi
  • Di Zhu
  • Sigang Yu
  • Geng Chen
  • Shijie Zhao
  • Dingwen Zhang
  • Shu Zhang

Electroencephalography (EEG) signals exhibit significant and inherent time scales differences across individuals and tasks. Despite notable successes in decoding EEG signals in single-tasks (e. g. , detection of epilepsy), where the time scales are relatively consistent, substantial differences in temporal characteristics among various tasks pose a significant challenge. To address these limitations, we propose the MuST, which stands for Mu lti- S cale T ransformer, aiming to dynamically learn characteristics of EEG signals on different time scales. Building on the conventional Convolutional Neural Network (CNN)-Transformer model, the MuST introduces two innovations: (1) A hierarchical Transformer structure to dynamically capture global dependencies and long-range information from EEG signals at different scales. (2) A novel temporal convolutional network (TCN) module to replace the original feed forward network (FFN) module in the Transformer, effectively capturing local temporal patterns and short-term dependencies from EEG signals. To validate the performance of the MuST, we conducted experiments on five public EEG datasets with extreme time-scale differences. The experimental results on these datasets demonstrate that we have achieved an average classification accuracy of 91. 69% under identical parameter settings. This surpasses the baseline EEGNet by 5. 65%, highlighting its superior capability in handling multi-scale EEG signals for diverse tasks. More critically, MuST demonstrates a successful unified modeling of EEG temporal heterogeneity through mixed dataset training (epilepsy detection and sleep staging classification). This breakthrough validates our multi-scale architecture's capability to dynamically reconcile divergent neurophysiological timescales within a single model. Our code can be found at https://github.com/wisercc/MuST.

AIIM Journal 2026 Journal Article

Precise estimation of tissue microstructure with hybrid graph transformer

  • Haotian Jiang
  • Geng Chen
  • Jiquan Ma
  • Hui Cui
  • Shu Zhang
  • Yong Xia
  • Pew-Thian Yap

The accurate estimation of tissue microstructure requires a sufficient amount of Diffusion MRI (DMRI) data, however, the clinical acquisition of this is challenging. Deep learning therefore improves the inference of tissue microstructure by highly undersampled DMRI. However, existing methods typically suffer from the lack of consideration of joint information in the spatial domain (x-space) and the diffusion wavevector domain (q-space). Here, we propose a hybrid graph transformer (HGT) for combined q-space learning and x-space guidance for precise estimation of tissue microstructure. The HGT consists of a q-space learning module, which explicitly considers the geometrical data structure in q-space based on a graph convolutional network, and an x-space guidance module, which learns long-range spatial dependencies based on residual dense transformer blocks. The x-space guidance module provides anatomical context to regularize the estimation of microstructure from undersampled q -space data. Extensive experiments on data from the human connectome project and high-quality diffusion-weighted imaging of Parkinson’s disease indicate that HGT performs better than cutting-edge methods.

IJCAI Conference 2025 Conference Paper

A³-Net: Calibration-Free Multi-View 3D Hand Reconstruction for Enhanced Musical Instrument Learning

  • Geng Chen
  • Xufeng Jian
  • Yuchen Chen
  • Pengfei Ren
  • Jingyu Wang
  • Haifeng Sun
  • Qi Qi
  • Jing Wang

Precise 3D hand posture is essential for learning musical instruments. Reconstructing highly precise 3D hand gestures enables learners to correct and master proper techniques through 3D simulation and Extended Reality. However, exsiting methods typically rely on precisely calibrated multi-camera systems, which are not easily deployable in everyday environments. In this paper, we focus on calibration-free multi-view 3D hand reconstruction in unconstrained scenarios. Establishing correspondences between multi-view images is particularly challenging without camera extrinsics. To address this, we propose A^3-Net, a multi-level alignment framework that utilizes 3D structural representations with hierarchical geometric and explicit semantic information as alignment proxies, facilitating multi-view feature interaction in both 3D geometric space and 2D visual space. Specifically, we first perfrom global geometric alignment to map multi-view features into a canonical space. Subsequently, we aggregate information into predefined sparse and dense proxies to further integrate cross-view semantics through mutual interaction. Finnaly, we perfrom 2D alignment to align projected 2D visual features with 2D observations. Our method achieves state-of-the-art results in the multi-view 3D hand reconstruction task, demonstrating the effectiveness of our proposed framework.

AIIM Journal 2025 Journal Article

Mixture-attention Siamese transformer for video polyp segmentation

  • Geng Chen
  • Junqing Yang
  • Xiaozhou Pu
  • Ge-Peng Ji
  • Huan Xiong
  • Yongsheng Pan
  • Hengfei Cui
  • Yong Xia

Accurate segmentation of polyps from colonoscopy videos is of great significance to polyp treatment and early prevention of colorectal cancer. However, it is challenging due to the difficulties associated with modeling long-range spatio-temporal relationships within a colonoscopy video. In this paper, we address this challenging task with a novel Mixture-Attention Siamese Transformer (MAST), which explicitly models the long-range spatio-temporal relationships with a mixture-attention mechanism for accurate polyp segmentation. Specifically, we first construct a Siamese transformer architecture to jointly encode paired video frames for their feature representations. We then design a mixture-attention module to exploit the intra-frame and inter-frame correlations, enhancing the features with rich spatio-temporal relationships. Finally, the enhanced features are fed to two parallel decoders for predicting the segmentation maps. Extensive experiments on the large-scale SUN-SEG benchmark demonstrate the superior performance of MAST in comparison with the cutting-edge competitors. Our code is publicly available at https: //github. com/Junqing-Yang/MAST.

AAAI Conference 2025 Conference Paper

mmFAS: Multimodal Face Anti-Spoofing Using Multi-Level Alignment and Switch-Attention Fusion

  • Geng Chen
  • Wuyuan Xie
  • Di Lin
  • Ye Liu
  • Miaohui Wang

The increasing number of presentation attacks on reliable face matching has raised concerns and garnered attention towards face anti-spoofing (FAS). However, existing methods for FAS modeling commonly fuse multiple visual modalities (e.g., RGB, Depth, and Infrared) in a straightforward manner, disregarding latent feature gaps that can hinder representation learning. To address this challenge, we propose a novel multimodal FAS framework (mmFAS) that focuses on explicit alignment and fusion of latent features across different modalities. Specifically, we develop a multimodal alignment module to alleviate the latent feature gap by using instance-level contrastive learning and class-level matching simultaneously. Further, we explore a new switch-attention based fusion module to automatically aggregate complementary information and control model complexity. To evaluate the anti-spoofing performance more effectively, we adopt a challenging yet meaningful cross-database protocol involving four benchmark multimodal FAS datasets to simulate realworld scenarios. Extensive experimental results demonstrate the effectiveness of mmFAS in improving the accuracy of FAS systems, outperforming 10 representative methods.

JBHI Journal 2025 Journal Article

Super-resolution Reconstruction of Fetal Brain MRI with Multi-view Interpolation Weight Learning

  • Shijie Huang
  • DengQiang Jia
  • Kai Zhang
  • Lingnan Kong
  • Fangmei Zhu
  • Zhongxiang Ding
  • Geng Chen
  • Dinggang Shen

Super-resolution reconstruction (SRR) of isotropic fetal brain MR images is critical for prenatal ex aminations but is hindered by fetal motion and misalignment of thick-slice scans. To address these challenges comprehensively, we introduce an innovative deep learning model, namely 3D-WISE, a 3D Weighted Interpolation for Super-resolution Estimation of fetal brain MRI. The model generates high-quality isotropic fetal brain MR images by learning the interpolation weights to correct misalignments between slices and volumes. These misalignments are estimated by extracting deep features from multiple motion corrupted stacks. Specifically, 3D-WISE incorporates two key components: (1) a weight learning module for multi view interpolation and (2) a feature extraction module guided by multi-type attention mechanisms. The weight learning module first maps motion-corrupted thick-slice stacks into latent feature spaces. The resulting features are then fed to an implicit decoding block to estimate interpolation weights of the surrounding points for a given coordinate. We further enhance our approach by incorporating convolutional block attention and atlas-induced cross-attention mechanisms. Extensive experiments on two benchmark datasets show that our 3D-WISE achieves remarkably improved performance compared to the widely adoptedregistration-reconstruction framework. We also ex tend the experiments on anatomical structure reconstruction and achieve promising results, highlighting the significant potential of our 3D-WISE for fetal brain MR images SRR in clinical settings. Our code is available at https://github.com/sj-huang/3D-WISE.

NeurIPS Conference 2025 Conference Paper

Unified 2D-3D Discrete Priors for Noise-Robust and Calibration-Free Multiview 3D Human Pose Estimation

  • Geng Chen
  • Pengfei Ren
  • Xufeng Jian
  • Haifeng Sun
  • Menghao Zhang
  • Qi Qi
  • Zirui Zhuang
  • Jing Wang

Multi-view 3D human pose estimation (HPE) leverages complementary information across views to improve accuracy and robustness. Traditional methods rely on camera calibration to establish geometric correspondences, which is sensitive to calibration accuracy and lacks flexibility in dynamic settings. Calibration-free approaches address these limitations by learning adaptive view interactions, typically leveraging expressive and flexible continuous representations. However, as the multiview interaction relationship is learned entirely from data without constraint, they are vulnerable to noisy input, which can propagate, amplify and accumulate errors across all views, severely corrupting the final estimated pose. To mitigate this, we propose a novel framework that integrates a noise-resilient discrete prior into the continuous representation-based model. Specifically, we introduce the \textit{UniCodebook}, a unified, compact, robust, and discrete representation complementary to continuous features, allowing the model to benefit from robustness to noise while preserving regression capability. Furthermore, we further propose an attribute-preserving and complementarity-enhancing Discrete-Continuous Spatial Attention (DCSA) mechanism to facilitate interaction between discrete priors and continuous pose features. Extensive experiments on three representative datasets demonstrate that our approach outperforms both calibration-required and calibration-free methods, achieving state-of-the-art performance.

JBHI Journal 2024 Journal Article

Exploratory Training for Universal Lesion Detection: Enhancing Lesion Mining Quality Through Temporal Verification

  • Xiaoyu Bai
  • Geng Chen
  • Benteng Ma
  • Changyang Li
  • Jingfeng Zhang
  • Yong Xia

Universal lesion detection (ULD) has great value in clinical practice as it can detect various lesions across multiple organs. Deep learning-based detectors have great potential but require high-quality annotated training data. In practice, due to cost, expertise requirements, and the diverse nature of lesions, incomplete annotations are encountered. Directly training ULD detectors under this condition can yield suboptimal results. Leading pseudo-label methods rely on a dynamic lesion-mining mechanism operating at the mini-batch level to address this issue. However, the quality of mined lesions is inconsistent across different iterations, potentially limiting performance enhancement. Inspired by the observation that deep models learn concepts with increasing complexity, we propose an exploratory-training-based ULD (ET-ULD) method to assess the reliability of mined lesions over time. Our approach uses a teacher-student detection model where the teacher mines suspicious lesions, which are then combined with incomplete annotations to train the student. On top of that, we design a bounding-box bank to record the mining timestamps. Each image is trained in several rounds, allowing us to get a sequence of timestamps for the mined lesions. If a mined lesion consistently appears, it is likely to be a true lesion, otherwise, it may just be a noise. This serves as a crucial criterion for selecting reliable mined lesions for retraining. Experimental results show that ET-ULD surpass existing state-of-the-art methods on two distinct lesion image datasets. Notably, on the DeepLesion dataset, ET-ULD achieved a 5. 4% improvement in Average Precision (AP) over the previous methods, demonstrating its superior performance.

JBHI Journal 2024 Journal Article

Image Recovery Matters: A Recovery-Extraction Framework for Robust Fetal Brain Extraction From MR Images

  • Jian Chen
  • Ranlin Lu
  • Shilin Ye
  • Mengting Guang
  • Tewodros Megabiaw Tassew
  • Bin Jing
  • Guofu Zhang
  • Geng Chen

The extraction of the fetal brain from magnetic resonance (MR) images is a challenging task. In particular, fetal MR images suffer from different kinds of artifacts introduced during the image acquisition. Among those artifacts, intensity inhomogeneity is a common one affecting brain extraction. In this work, we propose a deep learning-based recovery-extraction framework for fetal brain extraction, which is particularly effective in handling fetal MR images with intensity inhomogeneity. Our framework involves two stages. First, the artifact-corrupted images are recovered with the proposed generative adversarial learning-based image recovery network with a novel region-of-darkness discriminator that enforces the network focusing on artifacts of the images. Second, we propose a brain extraction network for more effective fetal brain segmentation by strengthening the association between lower- and higher-level features as well as suppressing task-irrelevant features. Thanks to the proposed recovery-extraction strategy, our framework is able to accurately segment fetal brains from artifact-corrupted MR images. The experiments show that our framework achieves promising performance in both quantitative and qualitative evaluations, and outperforms state-of-the-art methods in both image recovery and fetal brain extraction.

ICLR Conference 2024 Conference Paper

Learning Planning Abstractions from Language

  • Weiyu Liu
  • Geng Chen
  • Joy Hsu
  • Jiayuan Mao
  • Jiajun Wu 0001

This paper presents a framework for learning state and action abstractions in sequential decision-making domains. Our framework, planning abstraction from language (PARL), utilizes language-annotated demonstrations to automatically discover a symbolic and abstract action space and induce a latent state abstraction based on it. PARL consists of three stages: 1) recovering object-level and action concepts, 2) learning state abstractions, abstract action feasibility, and transition models, and 3) applying low-level policies for abstract actions. During inference, given the task description, PARL first makes abstract action plans using the latent transition and feasibility functions, then refines the high-level plan using low-level policies. PARL generalizes across scenarios involving novel object instances and environments, unseen concept compositions, and tasks that require longer planning horizons than settings it is trained on.

NeurIPS Conference 2024 Conference Paper

Zipper: Addressing Degeneracy in Algorithm-Agnostic Inference

  • Geng Chen
  • Yinxu Jia
  • Guanghui Wang
  • Changliang Zou

The widespread use of black box prediction methods has sparked an increasing interest in algorithm/model-agnostic approaches for quantifying goodness-of-fit, with direct ties to specification testing, model selection and variable importance assessment. A commonly used framework involves defining a predictiveness criterion, applying a cross-fitting procedure to estimate the predictiveness, and utilizing the difference in estimated predictiveness between two models as the test statistic. However, even after standardization, the test statistic typically fails to converge to a non-degenerate distribution under the null hypothesis of equal goodness, leading to what is known as the degeneracy issue. To addresses this degeneracy issue, we present a simple yet effective device, Zipper. It draws inspiration from the strategy of additional splitting of testing data, but encourages an overlap between two testing data splits in predictiveness evaluation. Zipper binds together the two overlapping splits using a slider parameter that controls the proportion of overlap. Our proposed test statistic follows an asymptotically normal distribution under the null hypothesis for any fixed slider value, guaranteeing valid size control while enhancing power by effective data reuse. Finite-sample experiments demonstrate that our procedure, with a simple choice of the slider, works well across a wide range of settings.

IJCAI Conference 2023 Conference Paper

Dichotomous Image Segmentation with Frequency Priors

  • Yan Zhou
  • Bo Dong
  • Yuanfeng Wu
  • Wentao Zhu
  • Geng Chen
  • Yanning Zhang

Dichotomous image segmentation (DIS) has a wide range of real-world applications and gained increasing research attention in recent years. In this paper, we propose to tackle DIS with informative frequency priors. Our model, called FP-DIS, stems from the fact that prior knowledge in the frequency domain can provide valuable cues to identify fine-grained object boundaries. Specifically, we propose a frequency prior generator to jointly utilize a fixed filter and learnable filters to extract informative frequency priors. Before embedding the frequency priors into the network, we first harmonize the multi-scale side-out features to reduce their heterogeneity. This is achieved by our feature harmonization module, which is based on a gating mechanism to harmonize the grouped features. Finally, we propose a frequency prior embedding module to embed the frequency priors into multi-scale features through an adaptive modulation strategy. Extensive experiments on the benchmark dataset, DIS5K, demonstrate that our FP-DIS outperforms state-of-the-art methods by a large margin in terms of key evaluation metrics.

IJCAI Conference 2021 Conference Paper

Context-aware Cross-level Fusion Network for Camouflaged Object Detection

  • Yujia Sun
  • Geng Chen
  • Tao Zhou
  • Yi Zhang
  • Nian Liu

Camouflaged object detection (COD) is a challenging task due to the low boundary contrast between the object and its surroundings. In addition, the appearance of camouflaged objects varies significantly, e. g. , object size and shape, aggravating the difficulties of accurate COD. In this paper, we propose a novel Context-aware Crosslevel Fusion Network (C2F-Net) to address the challenging COD task. Specifically, we propose an Attention-induced Cross-level Fusion Module (ACFM) to integrate the multi-level features with informative attention coefficients. The fused features are then fed to the proposed Dual-branch Global Context Module (DGCM), which yields multi-scale feature representations for exploiting rich global context information. In C2F-Net, the two modules are conducted on high-level features using a cascaded manner. Extensive experiments on three widely used benchmark datasets demonstrate that our C2F-Net is an effective COD model and outperforms state-of-the-art models remarkably. Our code is publicly available at: https: //github. com/thograce/C2FNet.

AAAI Conference 2021 Conference Paper

Dual-Octave Convolution for Accelerated Parallel MR Image Reconstruction

  • Chun-Mei Feng
  • Zhanyuan Yang
  • Geng Chen
  • Yong Xu
  • Ling Shao

Magnetic resonance (MR) image acquisition is an inherently prolonged process, whose acceleration by obtaining multiple undersampled images simultaneously through parallel imaging has always been the subject of research. In this paper, we propose the Dual-Octave Convolution (Dual-OctConv), which is capable of learning multi-scale spatial-frequency features from both real and imaginary components, for fast parallel MR image reconstruction. By reformulating the complex operations using octave convolutions, our model shows a strong ability to capture richer representations of MR images, while at the same time greatly reducing the spatial redundancy. More specifically, the input feature maps and convolutional kernels are first split into two components (i. e. , real and imaginary), which are then divided into four groups according to their spatial frequencies. Then, our Dual-OctConv conducts intra-group information updating and inter-group information exchange to aggregate the contextual information across different groups. Our framework provides two appealing benefits: (i) it encourages interactions between real and imaginary components at various spatial frequencies to achieve richer representational capacity, and (ii) it enlarges the receptive field by learning multiple spatial-frequency features of both the real and imaginary components. We evaluate the performance of the proposed model on the acceleration of multi-coil MR image reconstruction. Extensive experiments are conducted on an in vivo knee dataset under different undersampling patterns and acceleration factors. The experimental results demonstrate the superiority of our model in accelerated parallel MR image reconstruction. Our code is available at: github. com/chunmeifeng/Dual-OctConv.

JBHI Journal 2019 Journal Article

Automatic Retinal Layer Segmentation of OCT Images With Central Serous Retinopathy

  • Dehui Xiang
  • Geng Chen
  • Fei Shi
  • Weifang Zhu
  • Qinghuai Liu
  • Songtao Yuan
  • Xinjian Chen

In this paper, an automatic method is reported for simultaneously segmenting layers and fluid in 3-D OCT retinal images of subjects suffering from central serous retinopathy. To enhance contrast between adjacent layers, multiscale bright and dark layer detection filters are proposed. Due to appearance of serous fluid or pigment epithelial detachment caused fluid, contrast between adjacent layers is often reduced, and also large morphological changes are caused. In addition, 24 features are designed for random forest classifiers. Then, 8 coarse surfaces are obtained based on the trained random forest classifiers. Finally, a hypergraph is constructed based on the smoothed image and the layer structure detection responses. A modified live wire algorithm is proposed to accurately detect surfaces between retinal layers, even though OCT images with fluids are of low contrast and layers are largely deformed. The proposed method was evaluated on 48 spectral domain OCT images with central serous retinopathy. The experimental results showed that the proposed method outperformed the state-of-art methods with regard to layers and fluid segmentation.

YNIMG Journal 2019 Journal Article

The UNC/UMN Baby Connectome Project (BCP): An overview of the study design and protocol development

  • Brittany R. Howell
  • Martin A. Styner
  • Wei Gao
  • Pew-Thian Yap
  • Li Wang
  • Kristine Baluyot
  • Essa Yacoub
  • Geng Chen

The human brain undergoes extensive and dynamic growth during the first years of life. The UNC/UMN Baby Connectome Project (BCP), one of the Lifespan Connectome Projects funded by NIH, is an ongoing study jointly conducted by investigators at the University of North Carolina at Chapel Hill and the University of Minnesota. The primary objective of the BCP is to characterize brain and behavioral development in typically developing infants across the first 5 years of life. The ultimate goals are to chart emerging patterns of structural and functional connectivity during this period, map brain-behavior associations, and establish a foundation from which to further explore trajectories of health and disease. To accomplish these goals, we are combining state of the art MRI acquisition and analysis techniques, including high-resolution structural MRI (T1-and T2-weighted images), diffusion imaging (dMRI), and resting state functional connectivity MRI (rfMRI). While the overall design of the BCP largely is built on the protocol developed by the Lifespan Human Connectome Project (HCP), given the unique age range of the BCP cohort, additional optimization of imaging parameters and consideration of an age appropriate battery of behavioral assessments were needed. Here we provide the overall study protocol, including approaches for subject recruitment, strategies for imaging typically developing children 0–5 years of age without sedation, imaging protocol and optimization, a description of the battery of behavioral assessments, and QA/QC procedures. Combining HCP inspired neuroimaging data with well-established behavioral assessments during this time period will yield an invaluable resource for the scientific community.

YNIMG Journal 2017 Journal Article

Joint prediction of longitudinal development of cortical surfaces and white matter fibers from neonatal MRI

  • Islem Rekik
  • Gang Li
  • Pew-Thian Yap
  • Geng Chen
  • Weili Lin
  • Dinggang Shen

The human brain can be modeled as multiple interrelated shapes (or a multishape), each for characterizing one aspect of the brain, such as the cortex and white matter pathways. Predicting the developing multishape is a very challenging task due to the contrasting nature of the developmental trajectories of the constituent shapes: smooth for the cortical surface and non-smooth for white matter tracts due to changes such as bifurcation. We recently addressed this problem and proposed an approach for predicting the multishape developmental spatiotemporal trajectories of infant brains based only on neonatal MRI data using a set of geometric, dynamic, and fiber-to-surface connectivity features. In this paper, we propose two key innovations to further improve the prediction of multishape evolution. First, for a more accurate cortical surface prediction, instead of simply relying on one neonatal atlas to guide the prediction of the multishape, we propose to use multiple neonatal atlases to build a spatially heterogeneous atlas using the multidirectional varifold representation. This individualizes the atlas by locally maximizing its similarity to the testing baseline cortical shape for each cortical region, thereby better representing the baseline testing cortical surface, which founds the multishape prediction process. Second, for temporally consistent fiber prediction, we propose to reliably estimate spatiotemporal connectivity features using low-rank tensor completion, thereby capturing the variability and richness of the temporal development of fibers. Experimental results confirm that the proposed variants significantly improve the prediction performance of our original multishape prediction framework for both cortical surfaces and fiber tracts shape at 3, 6, and 9 months of age. Our pioneering model will pave the way for learning how to predict the evolution of anatomical shapes with abnormal changes. Ultimately, devising accurate shape evolution prediction models that can help quantify and predict the severity of a brain disorder as it progresses will be of great aid in individualized treatment planning.

v2026.09.13