Arrow Research search

Author name cluster

Xuesong Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

10 papers
2 author rows

Possible papers

10

AAAI Conference 2025 Conference Paper

CLIP-MSM: A Multi-Semantic Mapping Brain Representation for Human High-Level Visual Cortex

  • Guoyuan Yang
  • Mufan Xue
  • Ziming Mao
  • Haofang Zheng
  • Jia Xu
  • Dabin Sheng
  • Ruotian Sun
  • Ruoqi Yang

Prior work employing deep neural networks (DNNs) with explainable techniques has identified human visual cortical selective representation to specific categories. However, constructing high-performing encoding models that accurately capture brain responses to coexisting multi-semantics remains elusive. Here, we used CLIP models combined with CLIP Dissection to establish a multi-semantic mapping framework (CLIP-MSM) for hypothesis-free analysis in human high-level visual cortex. First, we utilize CLIP models to construct voxel-wise encoding models for predicting visual cortical responses to natural scene images. Then, we apply CLIP Dissection and normalize the semantic mapping score to achieve the mapping of single brain voxels to multiple semantics. Our findings indicate that CLIP Dissection applied to DNNs modeling the human high-level visual cortex demonstrates better interpretability accuracy compared to Network Dissection. In addition, to demonstrate how our method enables fine-grained discovery in hypothesis-free analysis, we quantify the accuracy between CLIP-MSM’s reconstructed brain activation in response to categories of faces, bodies, places, words and food, and the ground truth of brain activation. We demonstrate that CLIP-MSM provides more accurate predictions of visual responses compared to CLIP Dissection. Our results have been validated using two large natural image datasets: the Natural Scenes Dataset (NSD) and the Natural Object Dataset (NOD).

IROS Conference 2025 Conference Paper

DynamicGSG: Dynamic 3D Gaussian Scene Graphs for Environment Adaptation

  • Luzhou Ge
  • Xiangyu Zhu
  • Zhuo Yang
  • Xuesong Li

In real-world scenarios, environment changes caused by human or agent activities make it extremely challenging for robots to perform various long-term tasks. Recent works typically struggle to effectively understand and adapt to dynamic environments due to the inability to update their environment representations in memory in response to environment changes and lack of fine-grained reconstruction of the environments. To address these challenges, we propose DynamicGSG, a dynamic, high-fidelity, open-vocabulary scene graph construction system leveraging Gaussian Splatting. DynamicGSG builds hierarchical scene graphs using advanced vision language models to represent the spatial hierarchy and semantic relationships between objects in the environments, utilizes a joint feature loss to supervise Gaussian instance grouping while optimizing the Gaussian maps, and locally updates the Gaussian scene graphs according to real environment changes for long-term environment adaptation. Experiments and ablation studies demonstrate the performance and efficacy of our proposed method in terms of semantic segmentation, language-guided object retrieval, and reconstruction quality. In addition, we validate the dynamic updating capabilities of our system within real-world laboratory settings. The source code and supplementary materials will be available at: https://github.com/GeLuzhou/Dynamic-Gsg.

EAAI Journal 2025 Journal Article

Neural dynamic fluid reconstruction technique for four-dimensional imaging of combustion flame based on deep learning

  • Fuhao Zhang
  • Zhiyin Ma
  • Can Gao
  • Gang Xun
  • Qingchun Lei
  • Xuesong Li

Three-dimensional optical diagnostic techniques based on the principles of tomographic imaging enable the acquisition of rich three-dimensional information in experimental flow fields through reconstruction calculations. However, for tasks involving the reconstruction of three-dimensional flow fields at high temporal resolutions, existing methods incur high computational costs, low reconstruction efficiency, and struggle to achieve high spatiotemporal resolution measurements. This paper proposes the Neural Dynamic Fluid Reconstruction Technique (NDFRT) based on deep learning. NDFRT incorporates the time dimension into the reconstruction scope to achieve the four-dimensional reconstruction of dynamic flow fields using neural networks. NDFRT has the following technical advantages: (1) ultra-high spatiotemporal reconstruction resolution; (2) good computational efficiency, with the reconstruction parameter scale only half that of traditional voxel-based methods; (3) the ability to perform three-dimensional frame prediction of dynamic fluids. We validated the proposed method using numerical simulation and experimental jet flame reconstruction and compared it with the traditional algebraic reconstruction technique (ART). Experimental results demonstrate that NDFRT outperforms traditional ART methods in terms of computational efficiency, reconstruction resolution, and reconstruction accuracy.

IROS Conference 2025 Conference Paper

Vibration-Based Energy Metric for Restoring Needle Alignment in Autonomous Robotic Ultrasound

  • Zhongyu Chen
  • Chenyang Li 0004
  • Xuesong Li
  • Dianye Huang
  • Zhongliang Jiang
  • Stefanie Speidel
  • Xiangyu Chu
  • Kwok Wai Samuel Au

Precise needle alignment is essential for percutaneous needle insertion in robotic ultrasound-guided procedures. However, inherent challenges such as speckle noise, needle-like artifacts, and low image resolution complicate robust needle detection, which is essential for alignment in ultrasound images. These issues become particularly problematic when visibility is reduced or lost, diminishing the effectiveness of visual-based needle alignment methods. In this paper, we propose a method to restore effectively when the ultrasound imaging plane and the needle insertion plane are misaligned. Unlike many existing approaches that rely heavily on needle visibility in ultrasound images, our method uses a more robust feature by periodically vibrating the needle using a mechanical system. Specifically, we propose a new vibration-based energy metric that remains effective even when the needle is fully out of plane. Using this metric, we develop an elegant control strategy to reposition the ultrasound probe in response to misalignments between the imaging plane and the needle insertion plane in both translation and rotation. Experiments conducted on ex-vivo porcine tissue samples using a dual-arm robotic ultrasound-guided needle insertion system demonstrate the effectiveness of the proposed approach. The experimental results show the translational error of 0. 41±0. 27 mm and the rotational error of 0. 51±0. 19 degrees.

AAAI Conference 2024 Conference Paper

A Convolutional Neural Network Interpretable Framework for Human Ventral Visual Pathway Representation

  • Mufan Xue
  • Xinyu Wu
  • Jinlong Li
  • Xuesong Li
  • Guoyuan Yang

Recently, convolutional neural networks (CNNs) have become the best quantitative encoding models for capturing neural activity and hierarchical structure in the ventral visual pathway. However, the weak interpretability of these black-box models hinders their ability to reveal visual representational encoding mechanisms. Here, we propose a convolutional neural network interpretable framework (CNN-IF) aimed at providing a transparent interpretable encoding model for the ventral visual pathway. First, we adapt the feature-weighted receptive field framework to train two high-performing ventral visual pathway encoding models using large-scale functional Magnetic Resonance Imaging (fMRI) in both goal-driven and data-driven approaches. We find that network layer-wise predictions align with the functional hierarchy of the ventral visual pathway. Then, we correspond feature units to voxel units in the brain and successfully quantify the alignment between voxel responses and visual concepts. Finally, we conduct Network Dissection along the ventral visual pathway including the fusiform face area (FFA), and discover variations related to the visual concept of `person'. Our results demonstrate the CNN-IF provides a new perspective for understanding encoding mechanisms in the human ventral visual pathway, and the combination of ante-hoc interpretable structure and post-hoc interpretable approaches can achieve fine-grained voxel-wise correspondence between model and brain. The source code is available at: https://github.com/BIT-YangLab/CNN-IF.

JBHI Journal 2024 Journal Article

Adaptive Knowledge Distillation for High-Quality Unsupervised MRI Reconstruction With Model-Driven Priors

  • Zhengliang Wu
  • Xuesong Li

Magnetic Resonance Imaging (MRI) reconstruction has made significant progress with the introduction of Deep Learning (DL) technology combined with Compressed Sensing (CS). However, most existing methods require large fully sampled training datasets to supervise the training process, which may be unavailable in many applications. Current unsupervised models also show limitations in performance or speed and may face unaligned distributions during testing. This paper proposes an unsupervised method to train competitive reconstruction models that can generate high-quality samples in an end-to-end style. Firstly teacher models are trained by filling the re-undersampled images and compared with the undersampled images in a self-supervised manner. The teacher models are then distilled to train another cascade model that can leverage the entire undersampled k-space during its training and testing. Additionally, we propose an adaptive distillation method to re-weight the samples based on the variance of teachers, which represents the confidence of the reconstruction results, to improve the quality of distillation. Experimental results on multiple datasets demonstrate that our method significantly accelerates the inference process while preserving or even improving the performance compared to the teacher model. In our tests, the distilled models show 5%–10% improvements in PSNR and SSIM compared with no distillation and are 10 times faster than the teacher.

IROS Conference 2023 Conference Paper

Thoracic Cartilage Ultrasound-CT Registration Using Dense Skeleton Graph

  • Zhongliang Jiang
  • Chenyang Li 0004
  • Xuesong Li
  • Nassir Navab

Autonomous ultrasound (US) imaging has gained increased interest recently, and it has been seen as a potential solution to overcome the limitations of free-hand US exami-nations, such as inter-operator variations. However, it is still challenging to accurately map planned paths from a generic atlas to individual patients, particularly for thoracic applications with high acoustic-impedance bone structures below the skin. To address this challenge, a dense graph-based non-rigid registration is proposed to transfer planned paths from the atlas to the current setup by explicitly considering subcutaneous bone surface. To this end, the sternum and cartilage branches are segmented using a template matching to assist coarse alignment of US and CT point clouds. Afterward, a directed graph is generated based on the CT template. Then, the self-organizing map using geographical distance is successively performed twice to extract the optimal graph representations for CT and US point clouds, individually. To evaluate the proposed approach, five cartilage point clouds from distinct patients are employed. The results demonstrate that the proposed graph-based registration can effectively map trajectories from CT to the current setup to do US examination through limited intercostal space. The non-rigid registration results in terms of Hausdorff distance (Mean±SD) is $9. 48 \pm 0. 27$ mm and the path transferring error in terms of Euclidean distance is $2. 21\pm 1. 11\ mm$. The code 1 1 https://github.com/marslicy/Cartilage-graph-based-US-CT-Registration and video 2 2 Video: https://www.youtube.com/watch?v=QJz2fkwgbP8 can be publicly accessed.

YNIMG Journal 2021 Journal Article

The divided brain: Functional brain asymmetry underlying self-construal

  • Gen Shi
  • Xuesong Li
  • Yifan Zhu
  • Ruihong Shang
  • Yang Sun
  • Hua Guo
  • Jie Sui

Self-construal (orientations of independence and interdependence) is a fundamental concept that guides human behaviour, and it is linked to a large number of brain regions. However, understanding the connectivity of these regions and the critical principles underlying these self-functions are lacking. Because brain activity linked to self-related processes are intrinsic, the resting-state method has received substantial attention. Here, we focused on resting-state functional connectivity matrices based on brain asymmetry as indexed by the differential partition of the connectivity located in mirrored positions of the two hemispheres, hemispheric specialization measured using the intra-hemispheric (left or right) connectivity, brain communication via inter-hemispheric interactions, and global connectivity as the sum of the two intra-hemispheric connectivity. Combining machine learning techniques with hypothesis-driven network mapping approaches, we demonstrated that orientations of independence and interdependence were best predicted by the asymmetric matrix compared to brain communication, hemispheric specialization, and global connectivity matrices. The network results revealed that there were distinct asymmetric connections between the default mode network, the salience network and the executive control network which characterise independence and interdependence. These analyses shed light on the importance of brain asymmetry in understanding how complex self-functions are optimally represented in the brain networks.

YNIMG Journal 2019 Journal Article

Downward cross-modal plasticity in single-sided deafness

  • Yufei Qiao
  • Xuesong Li
  • Hang Shen
  • Xue Zhang
  • Yang Sun
  • Wenyang Hao
  • Bingya Guo
  • Daofeng Ni

The auditory cortex has been shown to participate in visual processing in individuals with complete auditory deprivation. However, it remains unclear whether partial hearing deprivation like single-sided deafness (SSD) leads to similar cross-modal plasticity. To investigate this, we enrolled individuals with long-term SSD, into functional MRI scans under resting-state and a visuo-spatial working memory task. Contrary to previous findings in bilateral deafness, our study revealed decreased activation in the auditory cortex in both left (LSSD) and right (RSSD) single-sided deafness compared to normal hearing controls, with statistical significance in RSSD. The degree of involvement was correlated with residual hearing ability in RSSD. These observations suggest that SSD can lead to a downward cross-modal plasticity: the more hearing ability lost, the fewer brain resources in the auditory cortex can be applied to visual tasks. In addition, the fronto-parietal cortex was observed to be less activated during the visual task in RSSD while the resting-state fMRI revealed increased functional connectivity between the fronto-parietal cortex and the auditory cortex, suggesting fronto-parietal resources may be recruited less by vision but more by hearing. The LSSD showed a similar alteration trend with RSSD, but without statistical significance. Together these findings may indicate that when hearing is partially deprived in SSD, there may be redistribution for brain resources between hearing and vision, and vision tends to allocate less resources. Our findings in this pilot study of unilateral auditory-deprived individuals enrich the understanding of cross-modal plasticity in the brain.

YNIMG Journal 2018 Journal Article

Dual-TRACER: High resolution fMRI with constrained evolution reconstruction

  • Xuesong Li
  • Xiaodong Ma
  • Lyu Li
  • Zhe Zhang
  • Xue Zhang
  • Yan Tong
  • Lihong Wang
  • Sen Song

fMRI with high spatial resolution is beneficial for studies in psychology and neuroscience, but is limited by various factors such as prolonged imaging time, low signal to noise ratio and scarcity of advanced facilities. Compressed Sensing (CS) based methods for accelerating fMRI data acquisition are promising. Other advanced algorithms like k-t FOCUSS or PICCS have been developed to improve performance. This study aims to investigate a new method, Dual-TRACER, based on Temporal Resolution Acceleration with Constrained Evolution Reconstruction (TRACER), for accelerating fMRI acquisitions using golden angle variable density spiral. Both numerical simulations and in vivo experiments at 3T were conducted to evaluate and characterize this method. Results show that Dual-TRACER can provide functional images with a high spatial resolution (1×1mm2) under an acceleration factor of 20 while maintaining hemodynamic signals well. Compared with other investigated methods, dual-TRACER provides a better signal recovery, higher fMRI sensitivity and more reliable activation detection.

v2026.09.13