Arrow Research search

Author name cluster

Tao Jin

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

29 papers
2 author rows

Possible papers

29

AAAI Conference 2026 Conference Paper

Scene-Aware Spatiotemporal Generalization: Towards Robust Temporal Action Detection Across Domains

  • Fangming Feng
  • Sihang Cai
  • Zequn Xie
  • Yangyang Wu
  • Tao Jin

Temporal Action Detection (TAD) aims to identify specific actions in long, untrimmed videos by determining their start, end times and categories, yet existing models suffer from performance degradation under out-of-distribution scenarios due to unrealistic i.i.d. assumptions. While domain generalization (DG) offers a promising solution, image-based DG methods fail to address the unique spatiotemporal challenges in video-based TAD, including the spatiotemporal complexities and significant variations in action instance scales and densities across domains. To bridge this gap, we propose the first DG framework tailored for TAD. We propose Scene-Aware Video Segmentation, which segments videos based on semantic similarity, addressing cross-domain action instance density and scale discrepancies. Additionally, we present Temporal-Aware Normalization Perturbation to generate diverse video features while preserving temporal integrity. We establish the first DG-TAD benchmark, evaluating 11 state-of-the-art DG methods across four datasets. The experiments demonstrate that our framework consistently outperforms existing approaches, achieving superior generalization on unseen domains. The proposed modules are architecture-agnostic, offering plug-and-play compatibility for broader video understanding tasks.

EAAI Journal 2026 Journal Article

Virtual evaluation method of bridge load-bearing capacity based on dynamic load test and intelligent algorithm

  • Pengzhen Lu
  • Yuchao Liu
  • Tao Jin
  • Ying Wu
  • Xianglong Zheng
  • Tong Guo

Accurate and timely assessment of bridge bearing capacity is crucial for ensuring structural safety, maintaining traffic flow, and extending service life. Traditional static load testing suffers from limitations like traffic disruption, long duration, and high costs. To address this, this paper proposes an innovative bridge load-bearing capacity assessment framework integrating dynamic load testing with intelligent algorithms. Specifically, key structural parameters are identified through global sensitivity analysis. Dynamic test data are combined with a Bayesian Ridge Regression model to iteratively update the finite element model via back-calculated input parameters. Based on the validation coefficient approach, this method indirectly predicts theoretical static responses by leveraging relationships between static-dynamic characteristics to obtain equivalent static test results, enabling rapid and intelligent evaluation of bridge load performance. Additionally, dynamic assessment and prediction of bearing capacity are achieved using the adjusted FEM and Feature Mode Decomposition (FMD). Through optimization of input variables and training samples, the model demonstrates high accuracy and robust generalization. Case studies validate the method's effectiveness, showing low cost, minimal traffic disruption, and high safety standards, making it particularly suitable for rapid assessment of medium and small span bridges. This approach provides new insights for bridge operation and maintenance, reducing socio-economic costs while enhancing understanding of bridge performance.

EAAI Journal 2025 Journal Article

A multi-temporal granularity feature driven convolutional ensemble model for electricity theft detection

  • Mingfa Yang
  • Qinyu Huang
  • Yulong Liu
  • Xidong Zheng
  • Tao Jin
  • Mohamed A. Mohamed

Electricity theft causes substantial economic losses and safety hazards. While the widespread adoption of advanced metering infrastructure has significantly reduced electricity theft, perpetrators continue to find ways to exploit the system, employing increasingly covert and intricate methods. To address the ongoing challenge, this paper proposes an attention mechanism optimized multi-temporal granularity feature driven convolutional ensemble model for enhanced accuracy and robustness in electricity theft detection (ETD). For comprehensive feature extraction across diverse temporal scales, the proposed framework integrates two specialized feature extraction modules. The first module, a squeeze-and-excitation network-optimized temporal convolutional network, selectively focuses on informative temporal features within the electricity consumption data. The second module, a dual-dimensional attention enhanced deep residual network composed of residual blocks embedded with the convolutional block attention module, facilitates the model's concurrent learning of informative spatial and temporal features. Then, the features from each module are fused and classified through a fully connected layer. To validate the effectiveness of the proposed ETD method, this paper conducted simulation experiments using the publicly available dataset from the State Grid Corporation of China. The experimental results show that the model optimized with the attention mechanism significantly improves the performance of ETD. Compared to other ETD models, the proposed model performs excellently in various indicators under different training set ratios and sample imbalance scenarios, demonstrating good generalization and robustness. Additionally, the model was deployed on a Raspberry Pi edge computing device to further verify its feasibility in practical engineering applications.

AAAI Conference 2025 Conference Paper

A Wander Through the Multimodal Landscape: Efficient Transfer Learning via Low-rank Sequence Multimodal Adapter

  • Zirun Guo
  • Xize Cheng
  • Yangyang Wu
  • Tao Jin

Efficient transfer learning methods such as adapter-based methods have shown great success in unimodal models and vision-language models. However, existing methods have two main challenges in fine-tuning multimodal models. Firstly, they are designed for vision-language tasks and fail to extend to situations where there are more than two modalities. Secondly, they exhibit limited exploitation of interactions between modalities and lack efficiency. To address these issues, in this paper, we propose the loW-rank sequence multimodal adapter (Wander). We first use the outer product to fuse the information from different modalities in an element-wise way effectively. For efficiency, we use CP decomposition to factorize tensors into rank-one components and achieve substantial parameter reduction. Furthermore, we implement a token-level low-rank decomposition to extract more fine-grained features and sequence relationships between modalities. With these designs, Wander enables token-level interactions between sequences of different modalities in a parameter-efficient way. We conduct extensive experiments on datasets with different numbers of modalities, where Wander outperforms state-of-the-art efficient transfer learning methods consistently. The results fully demonstrate the effectiveness, efficiency and universality of Wander.

NeurIPS Conference 2025 Conference Paper

AHa-Bench: Benchmarking Audio Hallucinations in Large Audio-Language Models

  • Xize Cheng
  • Dongjie Fu
  • Chenyuhao Wen
  • Shannon Yu
  • Zehan Wang
  • Shengpeng Ji
  • Siddhant Arora
  • Tao Jin

Hallucinations present a significant challenge in the development and evaluation of large language models (LLMs), directly affecting their reliability and accuracy. While notable advancements have been made in research on textual and visual hallucinations, there is still a lack of a comprehensive benchmark for evaluating auditory hallucinations in large audio language models (LALMs). To fill this gap, we introduce AHa-Bench, a systematic and comprehensive benchmark for audio hallucinations. Audio data, in particular, uniquely combines the multi-attribute complexity of visual data with the semantic richness of textual data, leading to auditory hallucinations that share characteristics with both visual and textual hallucinations. Based on the source of these hallucinations, AHa-Bench categorizes them into semantic hallucinations, acoustic hallucinations, and semantic-acoustic confusion hallucinations. In addition, we systematically evaluate seven open-source local perception language models (LALMs), demonstrating the challenges these models face in audio understanding, especially when it comes to jointly understanding semantic and acoustic information. Through the development of a comprehensive evaluation framework, AHa-Bench aims to enhance the robustness and stability of LALMs, fostering more reliable and nuanced audio understanding in LALMs. The benchmark dataset is available at \url{https: //huggingface. co/datasets/ahabench/AHa-Bench}.

AAAI Conference 2025 Conference Paper

Bridging the Gap for Test-Time Multimodal Sentiment Analysis

  • Zirun Guo
  • Tao Jin
  • Wenlong Xu
  • Wang Lin
  • Yangyang Wu

Multimodal sentiment analysis (MSA) is an emerging research topic that aims to understand and recognize human sentiment or emotions through multiple modalities. However, in real-world dynamic scenarios, the distribution of target data is always changing and different from the source data used to train the model, which leads to performance degradation. Common adaptation methods usually need source data, which could pose privacy issues or storage overheads. Therefore, test-time adaptation (TTA) methods are introduced to improve the performance of the model at inference time. Existing TTA methods are always based on probabilistic models and unimodal learning, and thus can not be applied to MSA which is often considered as a multimodal regression task. In this paper, we propose two strategies: Contrastive Adaptation and Stable Pseudo-label generation (CASP) for test-time adaptation for multimodal sentiment analysis. The two strategies deal with the distribution shifts for MSA by enforcing consistency and minimizing empirical risk, respectively. Extensive experiments show that CASP brings significant and consistent improvements to the performance of the model across various distribution shift settings and with different backbones, demonstrating its effectiveness and versatility.

AAAI Conference 2025 Conference Paper

Speech Watermarking with Discrete Intermediate Representations

  • Shengpeng Ji
  • Ziyue Jiang
  • Jialong Zuo
  • Minghui Fang
  • Yifu Chen
  • Tao Jin
  • Zhou Zhao

Speech watermarking techniques can proactively mitigate the potential harmful consequences of instant voice cloning techniques. These techniques involve the insertion of signals into speech that are imperceptible to humans but can be detected by algorithms. Previous approaches typically embed watermark messages into continuous space. However, intuitively, embedding watermark information into robust discrete latent space can significantly improve the robustness of watermarking systems. In this paper, we propose DiscreteWM, a novel speech watermarking framework that injects watermarks into the discrete intermediate representations of speech. Specifically, we map speech into discrete latent space with a vector-quantized autoencoder and inject watermarks by changing the modular arithmetic relation of discrete IDs. To ensure the imperceptibility of watermarks, we also propose a manipulator model to select the candidate tokens for watermark embedding. Experimental results demonstrate that our framework achieves state-of-the-art performance in robustness and imperceptibility, simultaneously. Moreover, our flexible frame-wise approach can serve as an efficient solution for both voice cloning detection and information hiding. Additionally, DiscreteWM can encode 1 to 150 bits of watermark information within a 1-second speech clip, indicating its encoding capacity.

NeurIPS Conference 2024 Conference Paper

$E^3$: Exploring Embodied Emotion Through A Large-Scale Egocentric Video Dataset

  • Wang Lin
  • Yueying Feng
  • Wenkang Han
  • Tao Jin
  • Zhou Zhao
  • Fei Wu
  • Chang Yao
  • Jingyuan Chen

Understanding human emotions is fundamental to enhancing human-computer interaction, especially for embodied agents that mimic human behavior. Traditional emotion analysis often takes a third-person perspective, limiting the ability of agents to interact naturally and empathetically. To address this gap, this paper presents $E^3$ for Exploring Embodied Emotion, the first massive first-person view video dataset. $E^3$ contains more than $50$ hours of video, capturing $8$ different emotion types in diverse scenarios and languages. The dataset features videos recorded by individuals in their daily lives, capturing a wide range of real-world emotions conveyed through visual, acoustic, and textual modalities. By leveraging this dataset, we define $4$ core benchmark tasks - emotion recognition, emotion classification, emotion localization, and emotion reasoning - supported by more than $80$k manually crafted annotations, providing a comprehensive resource for training and evaluating emotion analysis models. We further present Emotion-LlaMa, which complements visual modality with acoustic modality to enhance the understanding of emotion in first-person videos. The results of comparison experiments with a large number of baselines demonstrate the superiority of Emotion-LlaMa and set a new benchmark for embodied emotion analysis. We expect that $E^3$ can promote advances in multimodal understanding, robotics, and augmented reality, and provide a solid foundation for the development of more empathetic and context-aware embodied agents.

NeurIPS Conference 2024 Conference Paper

Action Imitation in Common Action Space for Customized Action Image Synthesis

  • Wang Lin
  • Jingyuan Chen
  • Jiaxin Shi
  • Zirun Guo
  • Yichen Zhu
  • Zehan Wang
  • Tao Jin
  • Zhou Zhao

We propose a novel method, \textbf{TwinAct}, to tackle the challenge of decoupling actions and actors in order to customize the text-guided diffusion models (TGDMs) for few-shot action image generation. TwinAct addresses the limitations of existing methods that struggle to decouple actions from other semantics (e. g. , the actor's appearance) due to the lack of an effective inductive bias with few exemplar images. Our approach introduces a common action space, which is a textual embedding space focused solely on actions, enabling precise customization without actor-related details. Specifically, TwinAct involves three key steps: 1) Building common action space based on a set of representative action phrases; 2) Imitating the customized action within the action space; and 3) Generating highly adaptable customized action images in diverse contexts with action similarity loss. To comprehensively evaluate TwinAct, we construct a novel benchmark, which provides sample images with various forms of actions. Extensive experiments demonstrate TwinAct's superiority in generating accurate, context-independent customized actions while maintaining the identity consistency of different subjects, including animals, humans, and even customized actors.

NeurIPS Conference 2024 Conference Paper

Classifier-guided Gradient Modulation for Enhanced Multimodal Learning

  • Zirun Guo
  • Tao Jin
  • Jingyuan Chen
  • Zhou Zhao

Multimodal learning has developed very fast in recent years. However, during the multimodal training process, the model tends to rely on only one modality based on which it could learn faster, thus leading to inadequate use of other modalities. Existing methods to balance the training process always have some limitations on the loss functions, optimizers and the number of modalities and only consider modulating the magnitude of the gradients while ignoring the directions of the gradients. To solve these problems, in this paper, we present a novel method to balance multimodal learning with C lassifier- G uided G radient M odulation (CGGM), considering both the magnitude and directions of the gradients. We conduct extensive experiments on four multimodal datasets: UPMC-Food 101, CMU-MOSI, IEMOCAP and BraTS 2021, covering classification, regression and segmentation tasks. The results show that CGGM outperforms all the baselines and other state-of-the-art methods consistently, demonstrating its effectiveness and versatility. Our code is available at https: //github. com/zrguo/CGGM.

NeurIPS Conference 2024 Conference Paper

Extending Multi-modal Contrastive Representations

  • Ziang Zhang
  • Zehan Wang
  • Luping Liu
  • Rongjie Huang
  • Xize Cheng
  • Zhenhui Ye
  • Wang Lin
  • Huadai Liu

Multi-modal contrastive representation (MCR) of more than three modalities is critical in multi-modal learning. Although recent methods showcase impressive achievements, the high dependence on large-scale, high-quality paired data and the expensive training costs limit their further development. Inspired by recent C-MCR, this paper proposes $\textbf{Ex}$tending $\textbf{M}$ultimodal $\textbf{C}$ontrastive $\textbf{R}$epresentation (Ex-MCR), a training-efficient and paired-data-free method to build unified contrastive representation for many modalities. Since C-MCR is designed to learn a new latent space for the two non-overlapping modalities and projects them onto this space, a significant amount of information from their original spaces is lost in the projection process. To address this issue, Ex-MCR proposes to extend one modality's space into the other's, rather than mapping both modalities onto a completely new space. This method effectively preserves semantic alignment in the original space. Experimentally, we extend pre-trained audio-text and 3D-image representations to the existing vision-text space. Without using paired data, Ex-MCR achieves comparable performance to advanced methods on a series of audio-image-text and 3D-image-text tasks and achieves superior performance when used in parallel with data-driven methods. Moreover, semantic alignment also emerges between the extended modalities (e. g. , audio and 3D).

ICRA Conference 2023 Conference Paper

High Resolution Point Clouds from mmWave Radar

  • Akarsh Prabhakara
  • Tao Jin
  • Arnav Das 0001
  • Gantavya Bhatt
  • Lilly Kumari
  • Elahe Soltanaghai
  • Jeff A. Bilmes
  • Swarun Kumar

This paper explores a machine learning approach on data from a single-chip mmWave radar for generating high resolution point clouds – a key sensing primitive for robotic applications such as mapping, odometry and localization. Unlike lidar and vision-based systems, mmWave radar can operate in harsh environments and see through occlusions like smoke, fog, and dust. Unfortunately, current mmWave processing techniques offer poor spatial resolution compared to lidar point clouds. This paper presents RadarHD, an end-to-end neural network that constructs lidar-like point clouds from low resolution radar input. Enhancing radar images is challenging due to the presence of specular and spurious reflections. Radar data also doesn't map well to traditional image processing techniques due to the signal's sinc-like spreading pattern. We overcome these challenges by training RadarHD on a large volume of raw I/Q radar data paired with lidar point clouds across diverse indoor settings. Our experiments show the ability to generate rich point clouds even in scenes unobserved during training and in the presence of heavy smoke occlusion. Further, RadarHD's point clouds are high-quality enough to work with existing lidar odometry and mapping workflows.

NeurIPS Conference 2022 Conference Paper

Active Ranking without Strong Stochastic Transitivity

  • Hao Lou
  • Tao Jin
  • Yue Wu
  • Pan Xu
  • Quanquan Gu
  • Farzad Farnoud

Ranking from noisy comparisons is of great practical interest in machine learning. In this paper, we consider the problem of recovering the exact full ranking for a list of items under ranking models that do *not* assume the Strong Stochastic Transitivity property. We propose a $$\delta$$-correct algorithm, Probe-Rank, that actively learns the ranking of the items from noisy pairwise comparisons. We prove a sample complexity upper bound for Probe-Rank, which only depends on the preference probabilities between items that are adjacent in the true ranking. This improves upon existing sample complexity results that depend on the preference probabilities for all pairs of items. Probe-Rank thus outperforms existing methods over a large collection of instances that do not satisfy Strong Stochastic Transitivity. Thorough numerical experiments in various settings are conducted, demonstrating that Probe-Rank is significantly more sample-efficient than the state-of-the-art active ranking method.

ICRA Conference 2022 Conference Paper

Collision Avoidance for Multiple Quadrotors Using Elastic Safety Clearance Based Model Predictive Control

  • Tao Jin
  • Xinghu Wang
  • Haibo Ji
  • Jian Di
  • Han Yan

When multiple quadrotors fly in a cluttered environment, collision-free flight must be assured. In this paper, we propose a novel elastic safety clearance based model predictive control (ESC-MPC) for multiple maneuverable quadrotors to avoid collisions in the presence of disturbance. This is accomplished through leveraging tube based model predictive control to maintain the quadrotor in a tube of trajectories. Exponential control barrier function (ECBF) is integrated to realize the elastic safety clearance mechanism which offers a dynamic safety margin in maneuverable flight. We validate the superiority of our approach with laboratory experiments.

ICRA Conference 2022 Conference Paper

LADC: Learning-Based Anti-Disturbance Control for Washing Drone

  • Jian Di
  • Shaofeng Chen
  • Han Yan
  • Xinghu Wang
  • Hepeng Zhang
  • Haibo Ji
  • Tao Jin

Disturbance mainly caused by recoil force in-evitably makes washing drone seriously deviate from the desired position, thereby reducing the cleaning efficiency. It is neces-sary to develop an effective anti-disturbance control method. Although some progresses have been made, the position error thereof is still large, rendering existing methods inapplicable in washing drone. In this paper, we propose a learning-based anti-disturbance control (LADC) method to significantly reduce the position error by combining robust nonlinear control and partial differential equation network (PDENet). Taking data noise into account, we use differential spectral normalization in the training of the PDENet. A distinguishing feature of our method is to directly learn PDENet parameters from flight logs without installing extra sensors. Experimental results indicate that the proposed method outperforms classical PD method and extended state observer (ESO) based control method with 70 % and 50 % reduced position error, respectively, and can be further applied in variable scenarios. Video: https:// youtu.be/gNfLFAXalkI

NeurIPS Conference 2021 Conference Paper

Generalizable Multi-linear Attention Network

  • Tao Jin
  • Zhou Zhao

The majority of existing multimodal sequential learning methods focus on how to obtain effective representations and ignore the importance of multimodal fusion. Bilinear attention network (BAN) is a commonly used fusion method, which leverages tensor operations to associate the features of different modalities. However, BAN has a poor compatibility for more modalities, since the computational complexity of the attention map increases exponentially with the number of modalities. Based on this concern, we propose a new method called generalizable multi-linear attention network (MAN), which can associate as many modalities as possible in linear complexity with hierarchical approximation decomposition (HAD). Besides, considering the fact that softmax attention kernels cannot be decomposed as linear operation directly, we adopt the addition random features (ARF) mechanism to approximate the non-linear softmax functions with enough theoretical analysis. We conduct extensive experiments on four datasets of three tasks (multimodal sentiment analysis, multimodal speaker traits recognition, and video retrieval), the experimental results show that MAN could achieve competitive results compared with the state-of-the-art methods, showcasing the effectiveness of the approximation decomposition and addition random features mechanism.

AAAI Conference 2020 Conference Paper

Rank Aggregation via Heterogeneous Thurstone Preference Models

  • Tao Jin
  • Pan Xu
  • Quanquan Gu
  • Farzad Farnoud

We propose the Heterogeneous Thurstone Model (HTM) for aggregating ranked data, which can take the accuracy levels of different users into account. By allowing different noise distributions, the proposed HTM model maintains the generality of Thurstone’s original framework, and as such, also extends the Bradley-Terry-Luce (BTL) model for pairwise comparisons to heterogeneous populations of users. Under this framework, we also propose a rank aggregation algorithm based on alternating gradient descent to estimate the underlying item scores and accuracy levels of different users simultaneously from noisy pairwise comparisons. We theoretically prove that the proposed algorithm converges linearly up to a statistical error which matches that of the state-of-the-art method for the single-user BTL model. We evaluate the proposed HTM model and algorithm on both synthetic and real data, demonstrating that it outperforms existing methods.

IJCAI Conference 2020 Conference Paper

SBAT: Video Captioning with Sparse Boundary-Aware Transformer

  • Tao Jin
  • Siyu Huang
  • Ming Chen
  • Yingming Li
  • Zhongfei Zhang

In this paper, we focus on the problem of applying the transformer structure to video captioning effectively. The vanilla transformer is proposed for uni-modal language generation task such as machine translation. However, video captioning is a multimodal learning problem, and the video features have much redundancy between different time steps. Based on these concerns, we propose a novel method called sparse boundary-aware transformer (SBAT) to reduce the redundancy in video representation. SBAT employs boundary-aware pooling operation for scores from multihead attention and selects diverse features from different scenarios. Also, SBAT includes a local correlation scheme to compensate for the local information loss brought by sparse operation. Based on SBAT, we further propose an aligned cross-modal encoding scheme to boost the multimodal interaction. Experimental results on two benchmark datasets show that SBAT outperforms the state-of-the-art methods under most of the metrics.

YNIMG Journal 2017 Journal Article

Enhancing sensitivity of pH-weighted MRI with combination of amide and guanidyl CEST

  • Tao Jin
  • Ping Wang
  • T. Kevin Hitchens
  • Seong-Gi Kim

Amide-proton-transfer weighted (APTw) MRI has emerged as a non-invasive pH-weighted imaging technique for studies of several diseases such as ischemic stroke. However, its pH-sensitivity is relatively low, limiting its capability to detect small pH changes. In this work, computer simulations, protamine phantom experiments, and in vivo gas challenge and experimental stroke in rats showed that, with judicious selection of the saturation pulse power, the amide-CEST at 3. 6ppm and guanidyl-CEST signals at 2. 0ppm changed in opposite directions with decreased pH. Thus, the difference between amide-CEST and guanidyl-CEST can enhance the pH measurement sensitivity, and is dubbed as pHenh. Acidification induced a negative contrast in APTw, but a positive contrast in pHenh. In vivo experiments showed that pHenh can detect hypercapnia-induced acidosis with about 3-times higher sensitivity than APTw. Also, pHenh slightly reduced gray and white matter contrast compared to APTw. In stroke animals, the CEST contrast between the ipsilateral ischemic core and contralateral normal tissue was −1. 85 ± 0. 42% for APTw and 3. 04 ± 0. 61% (n = 5) for pHenh, and the contrast to noise was 2. 9 times higher for pHenh than APTw. Our results suggest that pHenh can be a useful tool for non-invasive pH-weighted imaging.

YNIMG Journal 2016 Journal Article

Glucose metabolism-weighted imaging with chemical exchange-sensitive MRI of 2-deoxyglucose (2DG) in brain: Sensitivity and biological sources

  • Tao Jin
  • Hunter Mehrens
  • Ping Wang
  • Seong-Gi Kim

Recent proof-of-principle studies have demonstrated the feasibility of measuring the uptake and metabolism of non-labeled 2-deoxy-D-glucose (2DG) by a chemical exchange-sensitive spin-lock (CESL) MRI approach. In order to gain better understanding of this new approach, we performed dynamic in vivo CESL MRI on healthy rat brains with an intravenous injection of 2DG under various conditions at 9. 4T. For three 2DG doses of 0. 25, 0. 5 and 1g/kg, we found that 2DG-CESL signals increased linearly with injection dose at the initial (<20min) but not the later period (>40min) suggesting time-dependent differential weightings of 2DG transport and metabolism. Remaining 2DG-CESL studies were performed with 0. 25g/kg 2DG. Since a higher isoflurane level reduces glucose metabolism and increases blood flow, 2DG-CESL was measured under 0. 5%, 1. 5% and 2. 2% isoflurane. The 2DG-CESL signal was reduced at higher isoflurane levels correlating well with the 2DG phosphorylation in the intracellular space. To detect regional heterogeneities of glucose metabolism, 2DG-CESL with 0. 33×0. 33×1. 50mm3 resolution was obtained, which indeed showed a higher response in the cortex compared to the corpus callosum. Lastly, unlike CESL MRI with the injection of non-transportable mannitol, the 2DG-CESL response decreased with an increased spin-lock pulse power confirming that 2DG-CESL is dominated by chemical exchange processes in the extravascular space. Taken together, our results showed that 2DG-CESL MRI signals mainly indicate glucose transport and metabolism and may be a useful biomarker for metabolic studies of normal and diseased brains.

YNIMG Journal 2013 Journal Article

Characterization of non-hemodynamic functional signal measured by spin-lock fMRI

  • Tao Jin
  • Seong-Gi Kim

Current functional MRI techniques measure hemodynamic changes induced by neural activity. Alternative measurement of signals originated from tissue is desirable and may be achieved using T1ρ, the spin-lattice relaxation time in the rotating-frame, which is measured by spin-lock MRI. Functional T1ρ changes in the brain can have contributions from vascular dilation, tissue acidosis, and potentially other contributions. When the blood contributions were suppressed with a contrast agent at 9. 4 T, a small tissue-originated T1ρ change was consistently observed at the middle cortical layers of cat visual cortex during visual stimulation, which had different dynamic characteristics compared to hemodynamic fMRI such as a faster response and no post-stimulus undershoot. Functional tissue T1ρ is highly dependent on the magnetic field strength and experimental parameters such as the power of the spin-locking pulse. With a 500Hz spin-locking pulse, the tissue T1ρ without the blood contribution increased during visual stimulation, but decreased during acidosis-inducing hypercapnia and global ischemia, indicating different signal origins. Phantom studies suggest that it may have contribution from concentration decrease in metabolites. Even though the sensitivity is much weaker than BOLD and its exact interpretation needs further investigation, our results show that non-hemodynamic functional signal can be consistently observed by spin-lock fMRI.

YNIMG Journal 2012 Journal Article

Magnetic resonance imaging of the Amine–Proton EXchange (APEX) dependent contrast

  • Tao Jin
  • Ping Wang
  • Xiaopeng Zong
  • Seong-Gi Kim

Chemical exchange between water and labile protons from amino-acids, proteins and other molecules can be exploited to provide tissue contrast with magnetic resonance imaging (MRI) techniques. Using an off-resonance Spin-Locking (SL) scheme for signal preparation is advantageous because the image contrast can be tuned to specific exchange rates by adjusting SL pulse parameters. While the amide–proton transfer (APT) contrast is obtained optimally with steady-state preparation, using a low power and long irradiation pulse, image contrast from the faster amine–water proton exchange (APEX) is optimized in the transient state with a higher power and a shorter SL pulse. Our phantom experiments show that the APEX contrast is sensitive to protein and amino acid concentration, as well as pH. In vivo 9. 4-T SL MRI data of rat brains with irradiation parameters optimized to slow exchange rates have a sharp peak at 3. 5ppm and also broad peak at −2 to −5ppm, inducing negative contrast in APT-weighted images, while the APEX image has large positive signal resulting from a weighted summation of many different amine-groups. Brain ischemia induced by cardiac arrest decreases pure APT signal from ~1. 7% to ~0%, and increases the APEX signal from ~8% to ~16%. In the middle cerebral artery occlusion (MCAO) model, the APEX signal shows different spatial and temporal patterns with large inter-animal variations compared to APT and water diffusion maps. Because of the similarity between the chemical exchange saturation transfer (CEST) and SL techniques, APEX contrast can also be obtained by a CEST approach using similar irradiation parameters. APEX may provide useful information for many diseases involving a change in levels of proteins, peptides, amino-acids, or pH, and may serve as a sensitive neuroimaging biomarker.

YNIMG Journal 2010 Journal Article

Change of the cerebrospinal fluid volume during brain activation investigated by T1ρ-weighted fMRI

  • Tao Jin
  • Seong-Gi Kim

A voxel in MRI often contains tissue as well as cerebrospinal fluid (CSF). During functional stimulation, volume fractions of these different water compartments may change. To directly image the CSF volume fraction and measure its functional change, we utilized a rotating-frame longitudinal relaxation time (T 1ρ)-weighted MRI technique. At 9. 4T with a spin-locking frequency of ∼500Hz, T 1ρ of tissue water and CSF are about 48 and 450ms, respectively. Therefore, the parenchyma signal becomes negligible when a long spin-locking time (e. g. , 200ms) is applied, leaving only the CSF signal. Baseline CSF volume fraction (V csf) and its change induced by visual stimulation were mapped in isoflurane-anesthetized cats (n =6). In both T 1ρ-weighted fMRI with spin locking times of 200 and 300ms, negative changes with similar magnitudes were observed, indicating that a decrease in V csf is a dominant contributor. In the region with voxels containing the visual cortex and CSF compartments, an average baseline V csf was 24. 6±2%, an average CSF volume fraction change (ΔV csf/V csf) was −2. 45±0. 6%, and an absolute change in CSF volume fraction (ΔV csf) was −0. 6±0. 15%. A negative correlation was observed between pixel-wise baseline V csf and ΔV csf/V csf, which can be explained by similar ΔV csf among voxels. Our results suggest that the functional reduction of CSF volume fraction could contribute to fMRI signals, especially when the tissue signal is significantly reduced as compared to the CSF with certain experimental techniques or parameters.

YNIMG Journal 2008 Journal Article

Cortical layer-dependent dynamic blood oxygenation, cerebral blood flow and cerebral blood volume responses during visual stimulation

  • Tao Jin
  • Seong-Gi Kim

The spatiotemporal characteristics of cerebral blood volume (CBV) and flow (CBF) responses are important for understanding neurovascular coupling mechanisms and blood oxygenation level-dependent (BOLD) signals. For this, cortical layer-dependent BOLD, CBV and CBF responses were measured at the cat visual cortex using fMRI. Major findings are: (i) the time-dependent fMRI cortical profile is dependent on imaging modality. Overall, the peak across the cortex occurs at the cortical surface for BOLD, but at the middle cortical layer for CBV and CBF. Compared to an initial stimulation period (4–10 s), the spatial specificity of CBV to the middle cortical layer increases significantly at a later time, while the specificity of BOLD and CBF slightly changes. (ii) The CBV response at the upper cortical area containing large pial vessels has a faster onset time and time to peak than the BOLD response at the same area, and a faster time to peak than CBV at the middle cortical area with microvessels. This suggests that the dilation of microvessels at the middle cortical area follows arterial volume increase at the surface of the cortex. (iii) For all three modalities, the post-stimulus undershoot was observed with the 60-s stimulation paradigm, indicating that the post-stimulus BOLD undershoot cannot be explained by the delayed venous CBV recovery theory under our experimental conditions. (iv) The relationship between CBV and CBF responses is both spatially and temporally dependent. Thus, a single power-law scaling constant (gamma value) may not be applicable for high-resolution study.

YNIMG Journal 2008 Journal Article

Functional changes of apparent diffusion coefficient during visual stimulation investigated by diffusion-weighted gradient-echo fMRI

  • Tao Jin
  • Seong-Gi Kim

The signal source of apparent diffusion coefficient (ADC) changes induced by neural activity is not fully understood. To examine this issue, ADC-fMRI in response to a visual stimulus was obtained in isoflurane-anesthetized cats at 9. 4 T. A gradient-echo technique was used for minimizing the coupling between diffusion and background field gradients, which was experimentally confirmed. In the small b-value domain (b =5 and 200 s/mm2), a functional ADC increase was detected at the middle of the visual cortex and at the cortical surface, which was caused mainly by an increase in cerebral blood volume (CBV) and inflow. With higher b-values (b =200 and 1000−1200 s/mm2), a functional ADC decrease was observed in the parenchyma and also at the cortical surface. Within the parenchyma, the ADC decrease responded faster than the BOLD signal, but was not well localized to the middle of visual cortex and almost disappeared when the intravascular signal was removed with a susceptibility contrast agent, suggesting that the decrease in ADC without contrast agent was mostly of vascular origin. At the cortical surface, an average ADC decrease of 0. 5% remained after injection of the contrast agent, which may have arisen from a functional reduction of the partial volume of cerebrospinal fluid. Overall, a functional ADC change of tissue origin could not be detected under our experimental conditions.

YNIMG Journal 2008 Journal Article

Improved cortical-layer specificity of vascular space occupancy fMRI with slab inversion relative to spin-echo BOLD at 9.4 T

  • Tao Jin
  • Seong-Gi Kim

Cerebral blood volume (CBV)-weighted endogenous functional contrast can be obtained by the vascular space occupancy (VASO) technique. VASO relies on nonselective inversion for nulling blood signals, but the implementation of VASO at magnetic fields higher than 3 T is difficult due to converging T 1 values of tissue and blood water and a stronger counteracting blood oxygen level-dependent (BOLD) effect. To improve functional CBV sensitivity, we proposed to use VASO with slab-selective inversion (SI-VASO). Computer simulations showed that the SI-VASO approach significantly increases functional sensitivity compared to the original VASO in a stronger magnetic field with a shorter repetition time. To examine layer-dependent specificity, SI-VASO and spin-echo BOLD (SE-BOLD) functional magnetic resonance imaging (fMRI) experiments were performed on isoflurane-anesthetized cats during visual stimulation at 9. 4 T. Unlike simultaneously acquired SE-BOLD signal, the SI-VASO signal peaked at 0. 9 mm from the surface of the cortex and was localized to the middle cortical layer. The full-width at half maximal response across the cortex was narrower for SI-VASO than for SE-BOLD (1. 7 mm vs. 2. 5 mm, respectively), suggesting that SI-VASO is better localized to neuronally active sites than SE-BOLD fMRI. The magnitude of the SI-VASO change in the middle cortex was −1. 45% with our experimental parameters, corresponding to a relative CBV change of ∼8% when baseline CBV was assumed to be 5% in a two-compartment model. The signal response profile across the cortex, the calculated CBV change, and the time course of SI-VASO fMRI were similar to those of previously obtained CBV-weighted fMRI with contrast agent in the same animal model, suggesting that SI-VASO measures predominately functional CBV responses.

YNIMG Journal 2007 Journal Article

Improved spatial localization of post-stimulus BOLD undershoot relative to positive BOLD

  • Fuqiang Zhao
  • Tao Jin
  • Ping Wang
  • Seong-Gi Kim

The negative blood oxygenation level-dependent (BOLD) signal following the cessation of stimulation (post-stimulus BOLD undershoot) is observed in functional magnetic resonance imaging (fMRI) studies. However, its spatial characteristics are unknown. To investigate this, gradient-echo BOLD fMRI in response to visual stimulus was obtained in isoflurane-anesthetized cats at 9. 4 T. Since the middle cortical layer (layer 4) is known to have the highest metabolic and cerebral blood volume (CBV) responses, images were obtained to view the cortical cross-section. Robust post-stimulus BOLD undershoot was observed in all studies, and lasted longer than 30 s after the cessation of 40–60 s stimulation. The magnitude of post-stimulus BOLD undershoot was linearly dependent on echo time with little intercept when extrapolating to TE=0, indicating that the T 2* change is the major cause of the BOLD undershoot. The post-stimulus BOLD undershoot was observed within the cortex and near the surface of the cortex, while the prolonged CBV elevation was observed only at the middle of the cortex. Within the cortex, the largest post-stimulus undershoot was detected at the middle of the cortex, similar to the CBV increase during the stimulation period. Our findings demonstrate that, even though there is significant contribution from pial vessel signals, the post-stimulus undershoot BOLD signal is useful to improve the spatial localization of fMRI to active cortical sites.

v2026.09.13