Arrow Research search

Author name cluster

Ping Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

44 papers
2 author rows

Possible papers

44

AAAI Conference 2026 Conference Paper

Breaking Measurement Barriers: From Compressed Sensing to Deep Reconstruction

  • Gang Qu
  • Ping Wang
  • Siming Zheng
  • Xin Yuan

Deep learning methods have achieved remarkable success in image compressed sensing (CS) task, namely reconstructing a high-fidelity image from its compressed measurement. However, existing methods are deficient in incoherent compressed measurement at sensing phase and implicit measurement representations at reconstruction phase, limiting the overall performance. In this work, we answer two questions: (i) how to improve the measurement incoherence for decreasing the ill-posedness; (ii) how to learn informative representations from measurements. To this end, we propose a novel asymmetric Kronecker CS (AKCS) model and theoretically present its better incoherence than previous Kronecker CS with minimal increase of complexity. Moreover, apart from the explicit measurement representations in gradient descent projection in unfolding networks, we further propose a measurement-aware cross attention (MACA) mechanism to learn implicit measurement representations. We integrate AKCS and MACA into a widely-used unfolding architecture to get a measurement-enhanced unfolding network (MEUNet). Extensive experiments demonstrate that the proposed MEUNet achieves state-of-the-art (SOTA) performance in reconstruction accuracy with high efficiency.

AAAI Conference 2026 Conference Paper

End-to-End Contrastive Language-Speech Pretraining Model for Long-Form Spoken Question Answering

  • Jiliang Hu
  • Zuchao Li
  • Baoyuan Qi
  • Guoming Liu
  • Ping Wang

Significant progress has been made in spoken question answering (SQA) in recent years. However, many existing methods, including large audio language models, struggle with processing long audio. Follow the success of retrieval augmented generation, a speech-related retriever shows promising in help preprocessing long-form speech. But the performance of existing speech-related retrievers is lacking. To address this challenge, we propose CLSR, an end-to-end contrastive language-speech retriever that efficiently extracts question-relevant segments from long audio recordings for downstream SQA task. Unlike conventional speech-text contrastive models, CLSR incorporates an intermediate step that converts acoustic features into text-like representations prior to alignment, thereby more effectively bridging the gap between modalities. Experimental results across four cross-modal retrieval datasets demonstrate that CLSR surpasses both end-to-end speech related retrievers and pipeline approaches combining speech recognition with text retrieval, providing a robust foundation for advancing practical long-form SQA applications.

AAAI Conference 2026 Conference Paper

FinMathBench: A Formula-Driven Benchmark for Evaluating LLMs’ Math Reasoning Capabilities in Finance

  • Yi He
  • Ping Wang
  • Shiqiang Xiong
  • Chao Chen
  • Haixiang Hu

Many existing financial math reasoning benchmarks suffer from data contamination and high manual construction costs. To address this, we propose a novel formula-driven approach to dynamically construct math reasoning benchmarks in finance. Our two-stage approach: (1) generates single-formula questions by LLMs using a "Mask-for-Solve" paradigm for ground truth answers, and (2) synthesizes multi-formula questions through hierarchical tree-based DAGs. Our approach ensures novelty (via LLMs' creativity) and controllability of difficulty (via DAG structure). Based on a self-constructed financial formula bank, we utilize the proposed method to build FinMathBench, the first formula-driven and fully LLM-generated benchmark aimed at assessing LLMs' math reasoning abilities in finance, containing 946 questions across 4 complexity levels. Evaluation results on 40 LLMs demonstrate significant accuracy drops in multi-formula questions, e.g., 72.9% (1-Formula) to 14.0% (4-Formula) for GPT-4o under Chain-of-Thought prompting. Three critical flaws of LLMs are also observed: poor direct calculation performance, bias toward frequently solved variables in formulas, and erroneous "correction" of valid but extreme financial values. These findings highlight gaps in current LLMs' domain-specific reasoning and underscore FinMathBench's value for advancing robust financial LLMs.

AAAI Conference 2026 Conference Paper

GlitchMiner: Mining Glitch Tokens in Large Language Models via Gradient-based Discrete Optimization

  • Zihui Wu
  • Haichang Gao
  • Ping Wang
  • Shudong Zhang
  • Zhaoxiang Liu
  • Shiguo Lian

Glitch tokens—inputs that trigger unpredictable or anomalous behavior in Large Language Models (LLMs)—pose significant challenges to model reliability and safety. Existing detection methods primarily rely on heuristic embedding patterns or statistical anomalies within internal representations, limiting their generalizability across different model architectures and potentially missing anomalies that deviate from observed patterns. We introduce GlitchMiner, an behavior-driven framework designed to identify glitch tokens by maximizing predictive entropy. Leveraging a gradient-guided local search strategy, GlitchMiner efficiently explores the discrete token space without relying on model-specific heuristics or large-batch sampling. Extensive experiments across ten LLMs from five major model families demonstrate that GlitchMiner consistently outperforms existing approaches in detection accuracy and query efficiency, providing a generalizable and scalable solution for effective glitch token discovery.

AAAI Conference 2026 Conference Paper

High-Speed FHD Full-Color Video Computer-Generated Holography

  • Haomiao Zhang
  • Miao Cao
  • Xuan Yu
  • Hui Luo
  • Yanling Piao
  • Mengjie Qin
  • Zhangyuan Li
  • Ping Wang

Computer-generated holography (CGH) is a promising technology for next-generation displays. However, generating high-speed, high-quality holographic video requires both high frame rate display and efficient computation, but is constrained by two key limitations: (i) Learning-based models often produce over-smoothed phases with narrow angular spectra, causing severe color crosstalk in high frame rate full-color displays such as depth-division multiplexing and thus resulting in a trade-off between frame rate and color fidelity. (ii) Existing frame-by-frame optimization methods typically optimize frames independently, neglecting spatial-temporal correlations between consecutive frames and leading to computationally inefficient solutions. To overcome these challenges, in this paper, we propose a novel high-speed full-color video CGH generation scheme. First, we introduce Spectrum-Guided Depth Division Multiplexing (SGDDM), which optimizes phase distributions via frequency modulation, enabling high-fidelity full-color display at high frame rates. Second, we present HoloMamba, a lightweight asymmetric Mamba-Unet architecture that explicitly models spatial-temporal correlations across video sequences to enhance reconstruction quality and computational efficiency. Extensive simulated and real-world experiments demonstrate that SGDDM achieves high-fidelity full-color display without compromise in frame rate, while HoloMamba generates FHD (1080p) full-color holographic video at over 260 FPS, more than 2.6 times faster than the prior state-of-the-art Divide-Conquer-and-Merge Strategy.

AAAI Conference 2026 Conference Paper

MHB: Medical Hallucination Benchmark for Large Language Models in Complex Clinical Tasks

  • Jianrong Lu
  • Junwei Liu
  • Xingyun Zheng
  • Minghui Yang
  • Jian Wang
  • Ping Wang
  • Yechao Zhang

The integration of Large Language Models (LLMs) into clinical applications presents transformative potential but is undermined by the critical risk of hallucination, the generation of plausible but factually incorrect information. Such failures pose a direct threat to patient safety and the integrity of clinical decision-making. To address this challenge, we introduce MHB, a novel and comprehensive benchmark framework designed to evaluate LLM reliability in two complex, high-stakes clinical contexts: multi-turn medical dialogues and clinical case report analysis. The core of our contribution is a systematic methodology for generating adversarial test cases by injecting ``hallucination traps" into realistic medical data, guided by a fine-grained taxonomy of clinical errors. MHB, comprising 4,695 samples and 20,288 evaluation rubrics, underwent a rigorous, two-stage validation by a panel of 60 licensed physicians from top-tier hospitals, ensuring high clinical realism and consistency. This comprehensive assessment of leading LLMs revealed significant, clinically relevant shortcomings across the board. Even the best-performing model, Claude-4-Sonnet, exhibited a hallucination rate of 29.1%, with some open-source models exceeding 57.0%. All models struggled with specific traps, like fabricated medical data or non-existent guidelines, highlighting prevalent systemic weaknesses.

AAAI Conference 2026 Conference Paper

PulseMind: A Multi-Modal Medical Model for Real-World Clinical Diagnosis

  • Jiao Xu
  • Junwei Liu
  • Jiangwei Lao
  • Qi Zhu
  • Yunpeng Zhao
  • Congyun Jin
  • Shinan Liu
  • Zhihong Lu

Recent advances in medical multi-modal models focus on specialized image analysis like dermatology, pathology, or radiology. However, they do not fully capture the complexity of real-world clinical diagnostics, which involve heterogeneous inputs and require ongoing contextual understanding during patient-physician interactions. To bridge this gap, we introduce PulseMind, a new family of multi-modal diagnostic models that integrates a systematically curated dataset, a comprehensive evaluation benchmark, and a tailored training framework. Specifically, we first construct a diagnostic dataset, MediScope, which comprises 98,000 real-world multi-turn consultations and 601,500 medical images, spanning over 10 major clinical departments and more than 200 sub-specialties. Then, to better reflect the requirements of real-world clinical diagnosis, we develop the PulseMind Benchmark, a multi-turn diagnostic consultation benchmark with a four-dimensional evaluation protocol comprising proactiveness, accuracy, usefulness, and language quality. Finally, we design a training framework tailored for multi-modal clinical diagnostics, centered around a core component named Comparison-based Reinforcement Policy Optimization (CRPO). Compared to absolute score rewards, CRPO uses relative preference signals from multi-dimensional comparisons to provide stable and human-aligned training guidance. Extensive experiments demonstrate that PulseMind achieves competitive performance on both the diagnostic consultation benchmark and public medical benchmarks.

EAAI Journal 2025 Journal Article

A bidirectional bi-objective graph search model for sustainable urban railway alignment optimization

  • Tianlong Zhang
  • Yan Gao
  • Shuangting Xu
  • Ting Deng
  • Qing He
  • Paul Schonfeld
  • Yang Zou
  • Dong Liang

Designing railway alignments in building-dense urban areas is a challenging task, requiring consideration of both costs and impacts on existing buildings and the environment. Achieving a viable solution necessitates the application of computer-aided techniques for three-dimensional (3D) global path searches while simultaneously optimizing multiple objectives. To tackle this challenge, this study proposes a bidirectional bi-objective graph search model. This model efficiently searches the 3D space to generate high-quality railway alignment solutions that simultaneously consider both comprehensive costs (including railway construction, ecological, and affected building costs) and carbon emissions (covering emissions from railways and buildings). It provides valuable reference solutions for designers, enhancing the design efficiency. The model includes two main innovations: (1) the ability to quickly search the entire 3D space using a graph-based strategy, generating multiple alignment solutions that meet design constraints in a single optimization process, and (2) the ability to accurately and efficiently account for the impact of railway alignments on existing buildings during optimization. Testing the model on a real-world urban case demonstrates its capability to generate multiple alternative railway alignments within minutes. The Pareto balanced solution achieves an 18. 91 % reduction in comprehensive costs and a 13. 46 % decrease in carbon emissions compared to manual design. The estimation error of affected building areas is approximately 2 %–4 % along the approximately 40 km alignment. Overall, the significance of this study lies in exploring the application of efficient graph search algorithms in the multi-objective optimization design of railway alignments in urban areas, advancing ongoing research in this field.

NeurIPS Conference 2025 Conference Paper

Efficient RAW Image Deblurring with Adaptive Frequency Modulation

  • Wenlong Jiao
  • Binglong Li
  • Wei Shang
  • Ping Wang
  • Dongwei Ren

Image deblurring plays a crucial role in enhancing visual clarity across various applications. Although most deep learning approaches primarily focus on sRGB images, which inherently lose critical information during the image signal processing pipeline, RAW images, being unprocessed and linear, possess superior restoration potential but remain underexplored. Deblurring RAW images presents unique challenges, particularly in handling frequency-dependent blur while maintaining computational efficiency. To address these issues, we propose Frequency Enhanced Network (FrENet), a framework specifically designed for RAW-to-RAW deblurring that operates directly in the frequency domain. We introduce a novel Adaptive Frequency Positional Modulation module, which dynamically adjusts frequency components according to their spectral positions, thereby enabling precise control over the deblurring process. Additionally, frequency domain skip connections are adopted to further preserve high-frequency details. Experimental results demonstrate that FrENet surpasses state-of-the-art deblurring methods in RAW image deblurring, achieving significantly better restoration quality while maintaining high efficiency in terms of reduced MACs. Furthermore, FrENet's adaptability enables it to be extended to sRGB images, where it delivers comparable or superior performance compared to methods specifically designed for sRGB data. The source code will be publicly available.

ICLR Conference 2025 Conference Paper

Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

  • Yang Tian
  • Sizhe Yang
  • Jia Zeng
  • Ping Wang
  • Dahua Lin
  • Hao Dong 0003
  • Jiangmiao Pang

Current efforts to learn scalable policies in robotic manipulation primarily fall into two categories: one focuses on "action," which involves behavior cloning from extensive collections of robotic data, while the other emphasizes "vision," enhancing model generalization by pre-training representations or generative models, also referred to as world models, using large-scale visual datasets. This paper presents an end-to-end paradigm that predicts actions using inverse dynamics models conditioned on the robot's forecasted visual states, named Predictive Inverse Dynamics Models (PIDM). By closing the loop between vision and action, the end-to-end PIDM can be a better scalable action learner. In practice, we use Transformers to process both visual states and actions, naming the model Seer. It is initially pre-trained on large-scale robotic datasets, such as DROID, and can be adapted to real-world scenarios with a little fine-tuning data. Thanks to large-scale, end-to-end training and the continuous synergy between vision and action at each execution step, Seer significantly outperforms state-of-the-art methods across both simulation and real-world experiments. It achieves improvements of 13% on the LIBERO-LONG benchmark, 22% on CALVIN ABC-D, and 43% in real-world tasks. Notably, it demonstrates superior generalization for novel objects, lighting conditions, and environments under high-intensity disturbances. Code and models will be publicly available.

EAAI Journal 2025 Journal Article

Pseudo-label attention-based multiple instance learning for whole slide image classification

  • Jing He
  • Ping Wang
  • Jingwen Cai
  • Dan Tang
  • Shaowen Yao
  • Renyang Liu

Automating disease classification in whole slide images (WSIs) is crucial for improving clinical diagnostic efficiency. However, existing multiple instance learning (MIL) approaches for this task often struggle with challenges such as insufficient focus on positive regions and data imbalance between positive and negative regions. These issues can lead to suboptimal performance in practical applications. To address these problems, in this paper, we propose a novel embedding-based MIL technique called pseudo-label attention-based multiple instance learning (PAMIL). PAMIL aggregates each instance’s features regarding their contributions to improving downstream classification performance. The key insight of PAMIL involves training the model in a supervised manner by introducing pseudo-labels to emphasize positive regions. Additionally, we propose a fine-tuning strategy to effectively refine the dataset, eliminating the interference of false-positive data and alleviating data imbalance. The effectiveness of PAMIL was demonstrated through comparisons with six state-of-the-art MIL techniques across two large-scale, real-world datasets. Empirical results show that the proposed method outperforms other methods, achieving up to a 2. 15% improvement in accuracy and a 1. 61% increase in area under the curve (AUC) on the Cancer Genome Atlas Non-Small Cell Lung Cancer (TCGA-NSCLC) dataset, highlighting the superiority of our method in practical applications, such as helping clinicians diagnose quickly.

AAAI Conference 2025 Conference Paper

SongSong: A Time Phonograph for Chinese SongCi Music from Thousand of Years Away

  • Jiliang Hu
  • Jiajia Li
  • Ziyi Pan
  • Chong Chen
  • Zuchao Li
  • Ping Wang
  • Lefei Zhang

Recently, there have been significant advancements in music generation. However, existing models primarily focus on creating modern pop songs, making it challenging to produce ancient music with distinct rhythms and styles, such as ancient Chinese SongCi. In this paper, we introduce SongSong, the first music generation model capable of restoring Chinese SongCi to our knowledge. Our model first predicts the melody from the input SongCi, then separately generates the singing voice and accompaniment based on that melody, and finally combines all elements to create the final piece of music. Additionally, to address the lack of ancient music datasets, we create OpenSongSong, a comprehensive dataset of ancient Chinese SongCi music, featuring 29.9 hours of compositions by various renowned SongCi music masters. To assess SongSong's proficiency in performing SongCi, we randomly select 85 SongCi sentences that were not part of the training set for evaluation against SongSong and music generation platforms such as Suno and SkyMusic. The subjective and objective outcomes indicate that our proposed model achieves leading performance in generating high-quality SongCi music.

NeurIPS Conference 2025 Conference Paper

Spectral Compressive Imaging via Chromaticity-Intensity Decomposition

  • Xiaodong Wang
  • Zijun He
  • Ping Wang
  • Lishun Wang
  • Yanan Hu
  • Xin Yuan

In coded aperture snapshot spectral imaging (CASSI), the captured measurement entangles spatial and spectral information, posing a severely ill-posed inverse problem for hyperspectral images (HSIs) reconstruction. Moreover, the captured radiance inherently depends on scene illumination, making it difficult to recover the intrinsic spectral reflectance that remains invariant to lighting conditions. To address these challenges, we propose a chromaticity-intensity decomposition framework, which disentangles an HSI into a spatially smooth intensity map and a spectrally variant chromaticity cube. The chromaticity encodes lighting-invariant reflectance, enriched with high-frequency spatial details and local spectral sparsity. Building on this decomposition, we develop CIDNet—a Chromaticity-Intensity Decomposition unfolding network within a dual-camera CASSI system. CIDNet integrates a hybrid spatial-spectral Transformer tailored to reconstruct fine-grained and sparse spectral chromaticity and a degradation-aware, spatially-adaptive noise estimation module that captures anisotropic noise across iterative stages. Extensive experiments on both synthetic and real-world CASSI datasets demonstrate that our method achieves superior performance in both spectral and chromaticity fidelity. Code is released at: \url{https: //github. com/xiaodongwo/CIDNet}.

JBHI Journal 2025 Journal Article

Swallow-PPG: Photoplethysmography Templates for Comprehensive Temporal Analysis of Swallowing Anatomical Actions

  • Ying Zhang
  • Junjie Li
  • Ping Wang
  • Huaiyu Zhu
  • Bo Wang
  • Wei Luo
  • Yun Pan

In clinical practice, Videofluoroscopic Swallowing Study (VFSS) is commonly used to monitor the activity of anatomical structures during swallowing. However, it is limited by ionizing radiation exposure, adverse effects of barium contrast agents, and the high cost of specialized equipment. In this study, we propose a framework for analyzing swallowing behaviors in photoplethysmography (PPG) waveforms, which includes generalizing the manifestation of swallowing in PPG (i. e. , swallowing templates generation) and conducting comprehensive temporal analysis of swallowing anatomical actions (TASAA). For swallowing templates generation, we cluster and average the samples to obtain waveforms of templates, followed by conducting shape-based mapping and averaging on 28 time indicators to derive template unified time indicators (TUTIs). For comprehensive TASAA, we leverage templates waveforms and TUTIs to estimate time indicators based on the mapping relationship between samples and their respective templates. We evaluate the proposed framework on 357 swallowing PPG samples from 41 elderly subjects. The average relative error across all time indicators is 0. 123, and 6 indicators notably excel with errors below 0. 1. The proposed template-based swallowing analysis framework is expected to become a low-cost and non-ionizing alternative to VFSS for comprehensive TASAA.

NeurIPS Conference 2025 Conference Paper

UnCLe: Towards Scalable Dynamic Causal Discovery in Non-linear Temporal Systems

  • Tingzhu Bi
  • Yicheng Pan
  • Xinrui Jiang
  • Huize Sun
  • Meng Ma
  • Ping Wang

Uncovering cause-effect relationships from observational time series is fundamental to understanding complex systems. While many methods infer static causal graphs, real-world systems often exhibit dynamic causality —where relationships evolve over time. Accurately capturing these temporal dynamics requires time-resolved causal graphs. We propose UnCLe, a novel deep learning method for scalable dynamic causal discovery. UnCLe employs a pair of Uncoupler and Recoupler networks to disentangle input time series into semantic representations and learns inter-variable dependencies via auto-regressive Dependency Matrices. It estimates dynamic causal influences by analyzing datapoint-wise prediction errors induced by temporal perturbations. Extensive experiments demonstrate that UnCLe not only outperforms state-of-the-art baselines on static causal discovery benchmarks but, more importantly, exhibits a unique capability to accurately capture and represent evolving temporal causality in both synthetic and real-world dynamic systems (e. g. , human motion). UnCLe offers a promising approach for revealing the underlying, time-varying mechanisms of complex phenomena.

AAAI Conference 2024 Conference Paper

A Novel Energy Based Model Mechanism for Multi-Modal Aspect-Based Sentiment Analysis

  • Tianshuo Peng
  • Zuchao Li
  • Ping Wang
  • Lefei Zhang
  • Hai Zhao

Multi-modal aspect-based sentiment analysis (MABSA) has recently attracted increasing attention. The span-based extraction methods, such as FSUIE, demonstrate strong performance in sentiment analysis due to their joint modeling of input sequences and target labels. However, previous methods still have certain limitations: (i) They ignore the difference in the focus of visual information between different analysis targets (aspect or sentiment). (ii) Combining features from uni-modal encoders directly may not be sufficient to eliminate the modal gap and can cause difficulties in capturing the image-text pairwise relevance. (iii) Existing span-based methods for MABSA ignore the pairwise relevance of target span boundaries. To tackle these limitations, we propose a novel framework called DQPSA. Specifically, our model contains a Prompt as Dual Query (PDQ) module that uses the prompt as both a visual query and a language query to extract prompt-aware visual information and strengthen the pairwise relevance between visual information and the analysis target. Additionally, we introduce an Energy-based Pairwise Expert (EPE) module that models the boundaries pairing of the analysis target from the perspective of an Energy-based Model. This expert predicts aspect or sentiment span based on pairwise stability. Experiments on three widely used benchmarks demonstrate that DQPSA outperforms previous approaches and achieves a new state-of-the-art performance. The code will be released at https://github.com/pengts/DQPSA.

AAAI Conference 2024 Conference Paper

Multi-Modal Latent Space Learning for Chain-of-Thought Reasoning in Language Models

  • Liqi He
  • Zuchao Li
  • Xiantao Cai
  • Ping Wang

Chain-of-thought (CoT) reasoning has exhibited impressive performance in language models for solving complex tasks and answering questions. However, many real-world questions require multi-modal information, such as text and images. Previous research on multi-modal CoT has primarily focused on extracting fixed image features from off-the-shelf vision models and then fusing them with text using attention mechanisms. This approach has limitations because these vision models were not designed for complex reasoning tasks and do not align well with language thoughts. To overcome this limitation, we introduce a novel approach for multi-modal CoT reasoning that utilizes latent space learning via diffusion processes to generate effective image features that align with language thoughts. Our method fuses image features and text representations at a deep level and improves the complex reasoning ability of multi-modal CoT. We demonstrate the efficacy of our proposed method on multi-modal ScienceQA and machine translation benchmarks, achieving state-of-the-art performance on ScienceQA. Overall, our approach offers a more robust and effective solution for multi-modal reasoning in language models, enhancing their ability to tackle complex real-world problems.

AAAI Conference 2024 Conference Paper

N-gram Unsupervised Compoundation and Feature Injection for Better Symbolic Music Understanding

  • Jinhao Tian
  • Zuchao Li
  • Jiajia Li
  • Ping Wang

The first step to apply deep learning techniques for symbolic music understanding is to transform musical pieces (mainly in MIDI format) into sequences of predefined tokens like note pitch, note velocity, and chords. Subsequently, the sequences are fed into a neural sequence model to accomplish specific tasks. Music sequences exhibit strong correlations between adjacent elements, making them prime candidates for N-gram techniques from Natural Language Processing (NLP). Consider classical piano music: specific melodies might recur throughout a piece, with subtle variations each time. In this paper, we propose a novel method, NG-Midiformer, for understanding symbolic music sequences that leverages the N-gram approach. Our method involves first processing music pieces into word-like sequences with our proposed unsupervised compoundation, followed by using our N-gram Transformer encoder, which can effectively incorporate N-gram information to enhance the primary encoder part for better understanding of music sequences. The pre-training process on large-scale music datasets enables the model to thoroughly learn the N-gram information contained within music sequences, and subsequently apply this information for making inferences during the fine-tuning stage. Experiment on various datasets demonstrate the effectiveness of our method and achieved state-of-the-art performance on a series of music understanding downstream tasks. The code and model weights will be released at https://github.com/CinqueOrigin/NG-Midiformer.

ICRA Conference 2024 Conference Paper

RoboKeyGen: Robot Pose and Joint Angles Estimation via Diffusion-based 3D Keypoint Generation

  • Yang Tian
  • Jiyao Zhang
  • Guowei Huang 0002
  • Bin Wang
  • Ping Wang
  • Jiangmiao Pang
  • Hao Dong 0003

Estimating robot pose and joint angles is significant in advanced robotics, enabling applications like robot collaboration and online hand-eye calibration. However, the introduction of unknown joint angles makes prediction more complex than simple robot pose estimation, due to its higher dimensionality. Previous methods either regress 3D keypoints directly or utilise a render&compare strategy. These approaches often falter in terms of performance or efficiency and grapple with the cross-camera gap problem. This paper presents a novel framework that bifurcates the high-dimensional prediction task into two manageable subtasks: 2D keypoints detection and lifting 2D keypoints to 3D. This separation promises enhanced performance without sacrificing the efficiency innate to keypoint-based techniques. A vital component of our method is the lifting of 2D keypoints to 3D keypoints. Common deterministic regression methods may falter when faced with uncertainties from 2D detection errors or self-occlusions. Leveraging the robust modeling potential of diffusion models, we reframe this issue as a conditional 3D keypoints generation task. To bolster cross-camera adaptability, we introduce the Normalised Camera Coordinate Space (NCCS), ensuring alignment of estimated 2D keypoints across varying camera intrinsics. Experimental results demonstrate that the proposed method outperforms the state-of-the-art render&compare method and achieves higher inference speed. Furthermore, the tests accentuate our method’s robust cross-camera generalisation capabilities. We intend to release both the dataset and code in https://nimolty.github.io/Robokeygen/.

EAAI Journal 2023 Journal Article

Application of effective gravitational search algorithm with constraint priority and expert experience in optimal allocation problems of distribution network

  • Jie Qian
  • Ping Wang
  • Chenggen Pu
  • Xiaoli Peng
  • Gonggui Chen

Optimal allocation problem of distribution network (OAPDN), which attracts much attention of electric enterprise, contributes to the flexible and environmental-friendly power supply by rationally introducing distributed generations (DGs) and shunt capacitors (SCs). To smoothly solve OAPDN problems, several effective measures such as advantageous schemes guidance (ASG) mechanism are presented and integrated into the proposed modified gravitational search algorithm with expert experience (MGSA-EE). Multiple OAPDN experiments essentially indicate that the MGSA-EE with higher efficiency and stronger exploration capability reduces the power loss on 33, 69 and 119 node networks by 94. 15%, 98. 10% and 84. 59%, which is superior to most published technologies. Furthermore, multi-objective OAPDN problem which simultaneously considers two or more goals is also studied. Compared with the single-objective one, it is more in line with the diverse demands of actual electricity market, but the difficulty is greatly increased as well. On this basis, this paper extends MGSA-EE to an innovative multi-objective MGSA-EE (MMGSA-EE) algorithm by the suggested non-inferior sorting strategy with constraints-prior. Most typically in multi-objective OAPDN experiments on large scale networks, MMGSA-EE achieves high-quality scheme that concurrently reduces power loss and voltage deviation of 69-node network by 94. 07% and 98. 69%. Meanwhile, several quantitative indicators also prove that the proposed MMGSA-EE has competitive advantages over the original algorithm in terms of Pareto fronts, node voltage profiles and execution time. In general, MGSA-EE and MMGSA-EE algorithms provide efficient tools for exploring superior DG/SC configuration schemes, and are of great significance to fill research gaps in multi-objective optimizations of distribution network.

NeurIPS Conference 2023 Conference Paper

Towards Semi-Structured Automatic ICD Coding via Tree-based Contrastive Learning

  • Chang Lu
  • Chandan Reddy
  • Ping Wang
  • Yue Ning

Automatic coding of International Classification of Diseases (ICD) is a multi-label text categorization task that involves extracting disease or procedure codes from clinical notes. Despite the application of state-of-the-art natural language processing (NLP) techniques, there are still challenges including limited availability of data due to privacy constraints and the high variability of clinical notes caused by different writing habits of medical professionals and various pathological features of patients. In this work, we investigate the semi-structured nature of clinical notes and propose an automatic algorithm to segment them into sections. To address the variability issues in existing ICD coding models with limited data, we introduce a contrastive pre-training approach on sections using a soft multi-label similarity metric based on tree edit distance. Additionally, we design a masked section training strategy to enable ICD coding models to locate sections related to ICD codes. Extensive experimental results demonstrate that our proposed training strategies effectively enhance the performance of existing ICD coding methods.

AAAI Conference 2022 Conference Paper

Inferring Prototypes for Multi-Label Few-Shot Image Classification with Word Vector Guided Attention

  • Kun Yan
  • Chenbin Zhang
  • Jun Hou
  • Ping Wang
  • Zied Bouraoui
  • Shoaib Jameel
  • Steven Schockaert

Multi-label few-shot image classification (ML-FSIC) is the task of assigning descriptive labels to previously unseen images, based on a small number of training examples. A key feature of the multi-label setting is that images often have multiple labels, which typically refer to different regions of the image. When estimating prototypes, in a metric-based setting, it is thus important to determine which regions are relevant for which labels, but the limited amount of training data makes this highly challenging. As a solution, in this paper, we propose to use word embeddings as a form of prior knowledge about the meaning of the labels. In particular, visual prototypes are obtained by aggregating the local feature maps of the support images, using an attention mechanism that relies on the label embeddings. As an important advantage, our model can infer prototypes for unseen labels without the need for fine-tuning any model parameters, which demonstrates its strong generalization abilities. Experiments on COCO and PASCAL VOC furthermore show that our model substantially improves the current state-of-the-art.

AAAI Conference 2021 Conference Paper

A Simple and Effective Self-Supervised Contrastive Learning Framework for Aspect Detection

  • Tian Shi
  • Liuqing Li
  • Ping Wang
  • Chandan K. Reddy

Unsupervised aspect detection (UAD) aims at automatically extracting interpretable aspects and identifying aspect-specific segments (such as sentences) from online reviews. However, recent deep learning based topic models, specifically aspectbased autoencoder, suffer from several problems such as extracting noisy aspects and poorly mapping aspects discovered by models to the aspects of interest. To tackle these challenges, in this paper, we first propose a self-supervised contrastive learning framework and an attention-based model equipped with a novel smooth self-attention (SSA) module for the UAD task in order to learn better representations for aspects and review segments. Secondly, we introduce a high-resolution selective mapping (HRSMap) method to efficiently assign aspects discovered by the model to the aspects of interest. We also propose using a knowledge distillation technique to further improve the aspect detection performance. Our methods outperform several recent unsupervised and weakly supervised approaches on publicly available benchmark user review datasets. Aspect interpretation results show that extracted aspects are meaningful, have a good coverage, and can be easily mapped to aspects of interest. Ablation studies and attention weight visualization also demonstrate effectiveness of SSA and the knowledge distillation method.

NeurIPS Conference 2021 Conference Paper

Self-Paced Contrastive Learning for Semi-supervised Medical Image Segmentation with Meta-labels

  • Jizong Peng
  • Ping Wang
  • Christian Desrosiers
  • Marco Pedersoli

The contrastive pre-training of a recognition model on a large dataset of unlabeled data often boosts the model’s performance on downstream tasks like image classification. However, in domains such as medical imaging, collecting unlabeled data can be challenging and expensive. In this work, we consider the task of medical image segmentation and adapt contrastive learning with meta-label annotations to scenarios where no additional unlabeled data is available. Meta-labels, such as the location of a 2D slice in a 3D MRI scan, often come for free during the acquisition process. We use these meta-labels to pre-train the image encoder, as well as in a semi-supervised learning step that leverages a reduced set of annotated data. A self-paced learning strategy exploiting the weak annotations is proposed to furtherhelp the learning process and discriminate useful labels from noise. Results on five medical image segmentation datasets show that our approach: i) highly boosts the performance of a model trained on a few scans, ii) outperforms previous contrastive and semi-supervised approaches, and iii) reaches close to the performance of a model trained on the full data.

AAAI Conference 2021 Conference Paper

Towards Universal Physical Attacks on Single Object Tracking

  • Li Ding
  • Yongwei Wang
  • Kaiwen Yuan
  • Minyang Jiang
  • Ping Wang
  • Hua Huang
  • Z. Jane Wang

Recent studies show that small perturbations in video frames could misguide single object trackers. However, such attacks have been mainly designed for digital-domain videos (i. e. , perturbation on full images), which makes them practically infeasible to evaluate the adversarial vulnerability of trackers in real-world scenarios. Here we made the first step towards physically feasible adversarial attacks against visual tracking in real scenes with a universal patch to camouflage single object trackers. Fundamentally different from physical object detection, the essence of single object tracking lies in the feature matching between the search image and templates, and we therefore specially design the maximum textural discrepancy (MTD), a resolution-invariant and target location-independent feature de-matching loss. The MTD distills global textural information of the template and search images at hierarchical feature scales prior to performing feature attacks. Moreover, we evaluate two shape attacks, the regression dilation and shrinking, to generate stronger and more controllable attacks. Further, we employ a set of transformations to simulate diverse visual tracking scenes in the wild. Experimental results show the effectiveness of the physically feasible attacks on SiamMask and SiamRPN++ visual trackers both in digital and physical scenes.

AAAI Conference 2020 Conference Paper

Distance-IoU Loss: Faster and Better Learning for Bounding Box Regression

  • Zhaohui Zheng
  • Ping Wang
  • Wei Liu
  • Jinze Li
  • Rongguang Ye
  • Dongwei Ren

Bounding box regression is the crucial step in object detection. In existing methods, while ℓn-norm loss is widely adopted for bounding box regression, it is not tailored to the evaluation metric, i.e., Intersection over Union (IoU). Recently, IoU loss and generalized IoU (GIoU) loss have been proposed to benefit the IoU metric, but still suffer from the problems of slow convergence and inaccurate regression. In this paper, we propose a Distance-IoU (DIoU) loss by incorporating the normalized distance between the predicted box and the target box, which converges much faster in training than IoU and GIoU losses. Furthermore, this paper summarizes three geometric factors in bounding box regression, i.e., overlap area, central point distance and aspect ratio, based on which a Complete IoU (CIoU) loss is proposed, thereby leading to faster convergence and better performance. By incorporating DIoU and CIoU losses into state-of-the-art object detection algorithms, e.g., YOLO v3, SSD and Faster R-CNN, we achieve notable performance gains in terms of not only IoU metric but also GIoU metric. Moreover, DIoU can be easily adopted into non-maximum suppression (NMS) to act as the criterion, further boosting performance improvement. The source code and trained models are available at https://github.com/Zzh-tju/DIoU.

AAAI Conference 2020 Conference Paper

K-BERT: Enabling Language Representation with Knowledge Graph

  • Weijie Liu
  • Peng Zhou
  • Zhe Zhao
  • Zhiruo Wang
  • Qi Ju
  • Haotang Deng
  • Ping Wang

Pre-trained language representation models, such as BERT, capture a general language representation from large-scale corpora, but lack domain-specific knowledge. When reading a domain text, experts make inferences with relevant knowledge. For machines to achieve this capability, we propose a knowledge-enabled language representation model (K-BERT) with knowledge graphs (KGs), in which triples are injected into the sentences as domain knowledge. However, too much knowledge incorporation may divert the sentence from its correct meaning, which is called knowledge noise (KN) issue. To overcome KN, K-BERT introduces softposition and visible matrix to limit the impact of knowledge. K-BERT can easily inject domain knowledge into the models by being equipped with a KG without pre-training by itself because it is capable of loading model parameters from the pre-trained BERT. Our investigation reveals promising results in twelve NLP tasks. Especially in domain-specific tasks (including finance, law, and medicine), K-BERT significantly outperforms BERT, which demonstrates that K-BERT is an excellent choice for solving the knowledge-driven problems that require experts.

NeurIPS Conference 2019 Conference Paper

Gate Decorator: Global Filter Pruning Method for Accelerating Deep Convolutional Neural Networks

  • Zhonghui You
  • Kun Yan
  • Jinmian Ye
  • Meng Ma
  • Ping Wang

Filter pruning is one of the most effective ways to accelerate and compress convolutional neural networks (CNNs). In this work, we propose a global filter pruning algorithm called Gate Decorator, which transforms a vanilla CNN module by multiplying its output by the channel-wise scaling factors (i. e. gate). When the scaling factor is set to zero, it is equivalent to removing the corresponding filter. We use Taylor expansion to estimate the change in the loss function caused by setting the scaling factor to zero and use the estimation for the global filter importance ranking. Then we prune the network by removing those unimportant filters. After pruning, we merge all the scaling factors into its original module, so no special operations or structures are introduced. Moreover, we propose an iterative pruning framework called Tick-Tock to improve pruning accuracy. The extensive experiments demonstrate the effectiveness of our approaches. For example, we achieve the state-of-the-art pruning ratio on ResNet-56 by reducing 70% FLOPs without noticeable loss in accuracy. For ResNet-50 on ImageNet, our pruned model with 40% FLOPs reduction outperforms the baseline model by 0. 31% in top-1 accuracy. Various datasets are used, including CIFAR-10, CIFAR-100, CUB-200, ImageNet ILSVRC-12 and PASCAL VOC 2011.

IJCAI Conference 2019 Conference Paper

Solving the Satisfiability Problem of Modal Logic S5 Guided by Graph Coloring

  • Pei Huang
  • Minghao Liu
  • Ping Wang
  • Wenhui Zhang
  • Feifei Ma
  • Jian Zhang

Modal logic S5 has found various applications in artificial intelligence. With the advances in modern SAT solvers, SAT-based approach has shown great potential in solving the satisfiability problem of S5. The scale of the SAT encoding for S5 is strongly influenced by the upper bound on the number of possible worlds. In this paper, we present a novel SAT-based approach for S5 satisfiability problem. We show a normal form for S5 formulas. Based on this normal form, a conflict graph can be derived whose chromatic number provides an upper bound of the possible worlds and a lot of unnecessary search spaces can be eliminated in this process. A heuristic graph coloring algorithm is adopted to balance the efficiency and optimality. The number of possible worlds can be significantly reduced for many practical instances. Extensive experiments demonstrate that our approach outperforms state-of-the-art S5-SAT solvers.

YNIMG Journal 2017 Journal Article

Enhancing sensitivity of pH-weighted MRI with combination of amide and guanidyl CEST

  • Tao Jin
  • Ping Wang
  • T. Kevin Hitchens
  • Seong-Gi Kim

Amide-proton-transfer weighted (APTw) MRI has emerged as a non-invasive pH-weighted imaging technique for studies of several diseases such as ischemic stroke. However, its pH-sensitivity is relatively low, limiting its capability to detect small pH changes. In this work, computer simulations, protamine phantom experiments, and in vivo gas challenge and experimental stroke in rats showed that, with judicious selection of the saturation pulse power, the amide-CEST at 3. 6ppm and guanidyl-CEST signals at 2. 0ppm changed in opposite directions with decreased pH. Thus, the difference between amide-CEST and guanidyl-CEST can enhance the pH measurement sensitivity, and is dubbed as pHenh. Acidification induced a negative contrast in APTw, but a positive contrast in pHenh. In vivo experiments showed that pHenh can detect hypercapnia-induced acidosis with about 3-times higher sensitivity than APTw. Also, pHenh slightly reduced gray and white matter contrast compared to APTw. In stroke animals, the CEST contrast between the ipsilateral ischemic core and contralateral normal tissue was −1. 85 ± 0. 42% for APTw and 3. 04 ± 0. 61% (n = 5) for pHenh, and the contrast to noise was 2. 9 times higher for pHenh than APTw. Our results suggest that pHenh can be a useful tool for non-invasive pH-weighted imaging.

IJCAI Conference 2017 Conference Paper

Hierarchical Feature Selection with Recursive Regularization

  • Hong Zhao
  • Pengfei Zhu
  • Ping Wang
  • Qinghua Hu

In the big data era, the sizes of datasets have increased dramatically in terms of the number of samples, features, and classes. In particular, there exists usually a hierarchical structure among the classes. This kind of task is called hierarchical classification. Various algorithms have been developed to select informative features for flat classification. However, these algorithms ignore the semantic hyponymy in the directory of hierarchical classes, and select a uniform subset of the features for all classes. In this paper, we propose a new technique for hierarchical feature selection based on recursive regularization. This algorithm takes the hierarchical information of the class structure into account. As opposed to flat feature selection, we select different feature subsets for each node in a hierarchical tree structure using the parent-children relationships and the sibling relationships for hierarchical regularization. By imposing $\ell_{2, 1}$-norm regularization to different parts of the hierarchical classes, we can learn a sparse matrix for the feature ranking of each node. Extensive experiments on public datasets demonstrate the effectiveness of the proposed algorithm.

TCS Journal 2017 Journal Article

The entire chromatic number of graphs embedded on the torus with large maximum degree

  • Xiaoxue Hu
  • Ping Wang
  • Yiqiao Wang
  • Weifan Wang

An embedded graph G = ( V, E, F ) on the torus is entirely k-colorable if V ∪ E ∪ F can be colored with k colors such that any two adjacent or incident elements receive different colors. In this paper, we prove that every embedded graph G on the torus with maximum degree Δ ≥ 10 is entirely ( Δ + 2 ) -colorable.

YNIMG Journal 2016 Journal Article

Glucose metabolism-weighted imaging with chemical exchange-sensitive MRI of 2-deoxyglucose (2DG) in brain: Sensitivity and biological sources

  • Tao Jin
  • Hunter Mehrens
  • Ping Wang
  • Seong-Gi Kim

Recent proof-of-principle studies have demonstrated the feasibility of measuring the uptake and metabolism of non-labeled 2-deoxy-D-glucose (2DG) by a chemical exchange-sensitive spin-lock (CESL) MRI approach. In order to gain better understanding of this new approach, we performed dynamic in vivo CESL MRI on healthy rat brains with an intravenous injection of 2DG under various conditions at 9. 4T. For three 2DG doses of 0. 25, 0. 5 and 1g/kg, we found that 2DG-CESL signals increased linearly with injection dose at the initial (<20min) but not the later period (>40min) suggesting time-dependent differential weightings of 2DG transport and metabolism. Remaining 2DG-CESL studies were performed with 0. 25g/kg 2DG. Since a higher isoflurane level reduces glucose metabolism and increases blood flow, 2DG-CESL was measured under 0. 5%, 1. 5% and 2. 2% isoflurane. The 2DG-CESL signal was reduced at higher isoflurane levels correlating well with the 2DG phosphorylation in the intracellular space. To detect regional heterogeneities of glucose metabolism, 2DG-CESL with 0. 33×0. 33×1. 50mm3 resolution was obtained, which indeed showed a higher response in the cortex compared to the corpus callosum. Lastly, unlike CESL MRI with the injection of non-transportable mannitol, the 2DG-CESL response decreased with an increased spin-lock pulse power confirming that 2DG-CESL is dominated by chemical exchange processes in the extravascular space. Taken together, our results showed that 2DG-CESL MRI signals mainly indicate glucose transport and metabolism and may be a useful biomarker for metabolic studies of normal and diseased brains.

YNIMG Journal 2012 Journal Article

Magnetic resonance imaging of the Amine–Proton EXchange (APEX) dependent contrast

  • Tao Jin
  • Ping Wang
  • Xiaopeng Zong
  • Seong-Gi Kim

Chemical exchange between water and labile protons from amino-acids, proteins and other molecules can be exploited to provide tissue contrast with magnetic resonance imaging (MRI) techniques. Using an off-resonance Spin-Locking (SL) scheme for signal preparation is advantageous because the image contrast can be tuned to specific exchange rates by adjusting SL pulse parameters. While the amide–proton transfer (APT) contrast is obtained optimally with steady-state preparation, using a low power and long irradiation pulse, image contrast from the faster amine–water proton exchange (APEX) is optimized in the transient state with a higher power and a shorter SL pulse. Our phantom experiments show that the APEX contrast is sensitive to protein and amino acid concentration, as well as pH. In vivo 9. 4-T SL MRI data of rat brains with irradiation parameters optimized to slow exchange rates have a sharp peak at 3. 5ppm and also broad peak at −2 to −5ppm, inducing negative contrast in APT-weighted images, while the APEX image has large positive signal resulting from a weighted summation of many different amine-groups. Brain ischemia induced by cardiac arrest decreases pure APT signal from ~1. 7% to ~0%, and increases the APEX signal from ~8% to ~16%. In the middle cerebral artery occlusion (MCAO) model, the APEX signal shows different spatial and temporal patterns with large inter-animal variations compared to APT and water diffusion maps. Because of the similarity between the chemical exchange saturation transfer (CEST) and SL techniques, APEX contrast can also be obtained by a CEST approach using similar irradiation parameters. APEX may provide useful information for many diseases involving a change in levels of proteins, peptides, amino-acids, or pH, and may serve as a sensitive neuroimaging biomarker.

IROS Conference 2011 Conference Paper

A subject-based motion generation model with adjustable walking pattern for a gait robotic trainer: NaTUre-gaits

  • Ping Wang
  • Kin Huat Low
  • Alison H. McGregor

A gait trainer, NaTUre-gaits (natural and tunable rehabilitation gait system), has been developed to provide assistance for the gait rehabilitation. The gait rehabilitation system can provide mobility for an overground locomotion in forward walking and turning by a mobile platform. The exoskeleton modules are mounted on the mobile platform and attached to the lower limbs and pelvis in parallel. The synchronized motion generation for the exoskeleton modules is provided by virtue of the inverse kinematic model analysis. The pelvic trajectory is predefined and ten points are specified within one gait cycle to obtain the foot trajectory from the designated step length and height. The trajectory is obtained via curve fitting based on these specified points. On the other hand, in order to keep the desired foot trajectory, the pelvic motion is compensated to accommodate hip and knee joint angles. The proposed generation method is tested on the gait system and it has demonstrated the effectiveness and smoothness of the system.

YNIMG Journal 2011 Journal Article

Sensitivity and specificity of high-resolution balanced steady-state free precession fMRI at high field of 9.4T

  • Sung-Hong Park
  • Tae Kim
  • Ping Wang
  • Seong-Gi Kim

Balanced steady-state free precession (bSSFP) is an attractive fMRI method at high fields due to minimal spatial distortion. To examine sensitivity and specificity of bSSFP fMRI at ultrahigh magnetic field of 9. 4T, we performed high-resolution pass-band high flip-angle (16°) bSSFP fMRI with four phase cycling (PC) angles at two repetition times (TR) of 10ms and 20ms and conventional gradient-recalled-echo (GRE) fMRI with TR of 20ms on rat brain during forepaw stimulation. The sensitivity of bSSFP fMRI with TR of 20ms was higher than that of GRE fMRI regardless of PC angle. Because of magnetic field inhomogeneity, fMRI foci were changed with PC angle in bSSFP fMRI, which was more prominent when TR was shorter. Within a middle cortical layer region where magnetic field inhomogeneity was relatively small, the homogeneity of bSSFP fMRI signals was higher at shorter TR. Acquisition of baseline transition-band bSSFP images helped to identify pass- and transition-band regions and to understand corresponding bSSFP fMRI signals. Fourier analysis of the multiple PC bSSFP datasets provided echoes of multiple pathways separately, and the main echo component showed lower sensitivity and better homogeneity than the free induction decay component. In summary, pass-band bSSFP techniques would have advantages over GRE-based fMRI in terms of sensitivity, and may be a good choice for fMRI at ultrahigh fields.

TCS Journal 2010 Journal Article

The surviving rate of an infected network

  • Weifan Wang
  • Stephen Finbow
  • Ping Wang

Let G be a connected network. Let k ≥ 1 be an integer. Suppose that a vertex v of G becomes infected. A program is then installed on k -nodes not yet infected. Afterwards, the virus spreads to all its unprotected neighbors in each time interval. The virus and the network administrator take turns until the virus can no longer spread further. Let s n k ( v ) denote the maximum number of vertices in G the network administrator can save when a virus infects v. The k -surviving rate ρ k ( G ) of G is defined to be the average value ∑ v ∈ V ( G ) s n k ( v ) / n 2. In particular, we write ρ ( G ) = ρ 1 ( G ). In this paper, we first use a probabilistic method to show that almost all networks have k -surviving rate arbitrarily close to 0. Then, we prove the following results: (1) ρ ( G ) ≥ 2 35 for a planar network G of girth at least 9; (2) ρ 2 ( G ) ≥ 1 16 for a series–parallel network G; and (3) ρ 2 d − 1 ( G ) ≥ 2 5 d for a d -degenerate network G.

YNIMG Journal 2008 Journal Article

Trial-by-trial relationship between neural activity, oxygen consumption, and blood flow responses

  • Kazuto Masamoto
  • Alberto Vazquez
  • Ping Wang
  • Seong-Gi Kim

Trial-by-trial variability in local field potential (LFP), tissue partial pressure of oxygen (PO2), cerebral blood flow (CBF), and deoxyhemoglobin-weighted optical imaging of intrinsic signals (OIS) were tested in the rat somatosensory cortex while fixed electrical forepaw stimulation (1. 0-ms pulses with amplitude of 1. 2 mA at a frequency of 6 Hz) was repeatedly applied. The changes in the cerebral metabolic rate of oxygen (CMRO2) were also evaluated using a hypotension condition established by our group based on the administration of a vasodilator. Under normal conditions, CBF, PO2, and OIS showed positive signal changes (48%, 32%, and 0. 42%, respectively) following stimulation. Over multiple trials, the CBF responses were well correlated with the integral of the LFP amplitudes (∑LFP) (R mean =0. 78), whereas a lower correlation was found between PO2 and ∑LFP (R mean =0. 60) and between OIS and ∑LFP (R mean =0. 54). Under the hypotension condition the LFP responses were preserved, but the CBF responses were suppressed and the PO2 and OIS changes were negative (−12% and −0. 28%, respectively). In this condition, the trial-by-trial variations in PO2 and OIS were well correlated with the variability in ∑LFPs (R mean =−0. 77 and −0. 76, respectively), indicating a single trial coupling between CMRO2 changes and ∑LFP. These findings show that CBF and CMRO2 signals are more directly correlated with neural activity compared to blood oxygen-sensitive methods such as OIS and BOLD fMRI.

YNIMG Journal 2007 Journal Article

Improved spatial localization of post-stimulus BOLD undershoot relative to positive BOLD

  • Fuqiang Zhao
  • Tao Jin
  • Ping Wang
  • Seong-Gi Kim

The negative blood oxygenation level-dependent (BOLD) signal following the cessation of stimulation (post-stimulus BOLD undershoot) is observed in functional magnetic resonance imaging (fMRI) studies. However, its spatial characteristics are unknown. To investigate this, gradient-echo BOLD fMRI in response to visual stimulus was obtained in isoflurane-anesthetized cats at 9. 4 T. Since the middle cortical layer (layer 4) is known to have the highest metabolic and cerebral blood volume (CBV) responses, images were obtained to view the cortical cross-section. Robust post-stimulus BOLD undershoot was observed in all studies, and lasted longer than 30 s after the cessation of 40–60 s stimulation. The magnitude of post-stimulus BOLD undershoot was linearly dependent on echo time with little intercept when extrapolating to TE=0, indicating that the T 2* change is the major cause of the BOLD undershoot. The post-stimulus BOLD undershoot was observed within the cortex and near the surface of the cortex, while the prolonged CBV elevation was observed only at the middle of the cortex. Within the cortex, the largest post-stimulus undershoot was detected at the middle of the cortex, similar to the CBV increase during the stimulation period. Our findings demonstrate that, even though there is significant contribution from pial vessel signals, the post-stimulus undershoot BOLD signal is useful to improve the spatial localization of fMRI to active cortical sites.

YNIMG Journal 2006 Journal Article

Cortical layer-dependent BOLD and CBV responses measured by spin-echo and gradient-echo fMRI: Insights into hemodynamic regulation

  • Fuqiang Zhao
  • Ping Wang
  • Kristy Hendrich
  • Kamil Ugurbil
  • Seong-Gi Kim

Spatial specificity of functional magnetic resonance imaging (fMRI) signals to sub-millimeter functional architecture remains controversial. To investigate this issue, high-resolution fMRI in response to visual stimulus was obtained in isoflurane-anesthetized cats at 9. 4 T using conventional gradient-echo (GE) and spin-echo (SE) techniques; blood oxygenation-level dependent (BOLD) and cerebral blood volume (CBV)-weighted data were acquired without and with injection of 10 mg Fe/kg monocrystalline iron oxide nanoparticles (MION), respectively. Studies after MION injection at two SE times show that the T 2′ contribution to SE fMRI is minimal. GE and SE BOLD changes were spread across the cortical layers. GE and SE CBV-weighted fMRI responses peaked at the middle cortical layer, which has the highest experimentally-determined microvascular volume; full-width at half-maximum was <1. 0 mm. Parenchymal sensitivity of GE CBV-weighted fMRI was ∼3 times higher than that of SE CBV-weighted fMRI and ∼1. 5 times higher than that of BOLD fMRI. It is well known that GE CBV-weighted fMRI detects a volume change in vessels of all sizes, while SE CBV-weighted fMRI is heavily weighted toward microvascular changes. Peak CBV change of 10% at the middle of the cortex in GE measurements was 1. 8 times higher than that in SE measurements, indicating that CBV changes occur predominantly for vasculature connecting the intracortical vessels and capillaries. Our data supports the notion of laminar-dependent CBV regulation at a sub-millimeter scale.

YNIMG Journal 2006 Journal Article

Spatial specificity of the enhanced dip inherently induced by prolonged oxygen consumption in cat visual cortex: Implication for columnar resolution functional MRI

  • Mitsuhiro Fukuda
  • Ping Wang
  • Chan-Hong Moon
  • Manabu Tanifuji
  • Seong-Gi Kim

Since changes in oxygen consumption induced by active neurons are specific to cortical columns, the small and transient “dip” of deoxyhemoglobin signal, which indicates an increase in oxygen consumption, has been of great interest. In this study, we succeeded in enhancing and sustaining the dip in the deoxyhemoglobin-weighted 620-nm intrinsic optical imaging signals from a 10-s orientation-selective stimulation in cat visual cortex by reducing arterial blood pressure with sodium nitroprusside (a vasodilator) to mitigate the contribution of stimulus-induced blood supply. During this condition, intact spiking activity and a significant reduction of stimulus-induced blood volume changes (570-nm intrinsic signals) were confirmed. The deoxyhemoglobin signal from the prolonged dip was highly localized to iso-orientation domains only during the initial ∼2 s; the signal specificity weakened over time although the domains were still resolvable after 2 s. The most plausible explanation for this time-dependent spatial specificity is that deoxyhemoglobin induced by oxygen consumption drains from active sites, where spiking activity occurs, to spatially non-specific downstream vessels over time. Our results suggest that the draining effect of pial and intracortical veins in dHb-based imaging techniques, such as blood oxygenation-level dependent (BOLD) functional MRI, is intrinsically unavoidable and reduces its spatial specificity of dHb signal regardless of whether the stimulus-induced blood supply is spatially specific.

YNIMG Journal 2005 Journal Article

Spatial specificity of cerebral blood volume-weighted fMRI responses at columnar resolution

  • Fuqiang Zhao
  • Ping Wang
  • Kristy Hendrich
  • Seong-Gi Kim

The spatial specificity of functional magnetic resonance imaging (fMRI) signals to columnar architecture remains uncertain. At columnar resolution, the specificity of intrinsic cerebral blood volume (CBV) response to orientation-selective columns in isoflurane-anesthetized cats was determined for CBV-weighted fMRI signals after injection of iron oxide at a dose of 10 mg Fe/kg. CBV-weighted fMRI data were acquired at 9. 4 T with an in-plane resolution of 156 × 156 μm2 in area 18 during visual stimulation at two orthogonal orientations. A 1-mm-thick imaging slice was selected tangential to the cortical surface. Regions with large CBV changes in response to two orthogonal orientation gratings were highly complementary. Maps of iso-orientation domains in response to these gratings were highly reproducible, suggesting that CBV-weighted fMRI has high sensitivity and specificity. The average distance between iso-orientation domains was 1. 37 ± 0. 28 mm (n = 10 orientations) in an anterior–posterior direction. CBV-weighted fMRI signal change in the iso-orientation domains induced by preferred orientation was 1. 69 ± 0. 24 (n = 10) times larger than that induced by orthogonal orientation. Our data demonstrate that CBV regulates at a submillimeter columnar scale and CBV-weighted fMRI has sufficient specificity to map columnar organization in animals.

v2026.09.13