Arrow Research search

Author name cluster

Yang Zheng

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

19 papers
2 author rows

Possible papers

19

EAAI Journal 2026 Journal Article

Physics-guided efficient automotive radar object detection framework via multi-task joint optimization

  • Siqi Pang
  • Kaitai Guo
  • Yang Zheng
  • Jimin Liang

Millimeter-wave radar is essential for robust perception in autonomous systems, as its performance remains independent of visual conditions. However, conventional multi-stage radar object detection algorithms often face challenges in real-time deployment due to computational inefficiency and limited task adaptability. To solve this engineering application challenge, we propose an efficient and physics-guided framework that integrates domain knowledge with data-driven learning to address these efficiency gaps. Unlike conventional methods, our proposed method operates directly on the compact complex-valued range-Doppler radar spectrum. To effectively extract information from this representation, our framework incorporates three jointly optimized modules. First, a complex-valued virtual channel decoupling module utilizes waveform structure priors to decouple virtual channels and eliminate Doppler aliasing. Second, a learnable complex-valued Angle of Arrival (AoA) estimation module, initialized from the Fast Fourier Transform (FFT) basis, replaces static angle-FFT to provide task-specific angular refinement at significantly lower cost. Finally, an Inter-Dimensional-aware attention mechanism exploits the anisotropic nature of radar data to enhance feature expressiveness. This end-to-end, multi-task joint optimization mechanism allows the entire system to learn task-aligned representations, boosting both robustness and efficiency. Extensive experiments demonstrate that our method significantly reduces overhead (up to 48. 9% fewer parameters, 52. 8% reduction in floating-point operations, and 7. 6-fold faster inference) while achieving an 18. 6% average improvement in detection accuracy over advanced multi-stage approaches. These results highlight a scalable and high-precision solution for real-time, intelligent radar perception on resource-constrained platforms.

YNIMG Journal 2026 Journal Article

The impact of chronic psychosocial stress on corticomuscular responses to thermal pain stimulation

  • Lili Li
  • Hui Ma
  • Yang Zheng
  • Zhongliang Yu

Psychosocial stress refers to the subjective experience and appraisal of real or potentially threatening psychosocial conditions. Pain is a complex physiological and psychological phenomenon, referring to unpleasant sensory and emotional experience associated with actual or potential tissue damage. To investigate the impact of chronic psychosocial stress on neural circuit responses to thermal pain, this study analyzed corticomuscular activities by electroencephalograph and electromyography. The results demonstrate that chronic psychosocial stress can change the cortical synchronization at α and γ frequencies and enhance motor unit recruitment under heat pain at both initial and progressive stages. Moreover, under progressive stimulation, pain-related neural feedback and localization functions may be disrupted, presenting impaired corticomuscular coherence. Therefore, chronic psychosocial stress can alter the integrated processing across both central and peripheral pathways, indicating a disintegration of the cortical-motor integration function of heat pain at initial and progressive stages.

YNIMG Journal 2025 Journal Article

Corticomuscular coherence existed at the single motor unit level

  • Yang Zheng
  • Bofang Zheng
  • Wei Qiang
  • Yu Peng
  • Guanghua Xu
  • Gang Wang
  • Lili Li
  • Henry Shin

The monosynaptic cortico-motoneuronal connections suggest the possibility of individual motor units (MUs) receiving independent commands from motor cortex. However, previous studies that used corticomuscular coherence (CMC) between electroencephalogram (EEG) signals and electromyogram (EMG) signals have not directly explored the corticospinal functionality at the single motoneuron level. The objective of this study is to find out whether synchronous activities exist between the motor cortex and individual MUs. Corticomuscular coherence was calculated between the EEG signals and the MU firing event trains which were extracted using the EMG decomposition technique. The results showed that some but not all MUs indeed had significant coherent activities with the contralateral motor cortex, which we named the cortico-motoneuronal coherence (CMnC). In contrast to the CMC only occurring in β and γ bands, CMnC occurred across the four common EEG frequency bands (θ, α, β and γ). Further, we identified individual MUs that showed significant interactions with the motor cortex. These coherent MUs (CohMU) could still be found even when the EMG signals were not coupled with the cortical activities. Compared with conventional CMC, our preliminary results indicated that the CMnC could potentially help to investigate the complex coupling between cortical and muscular activities due to its ability to separate different correlated components. This study proves that corticomuscular coherence exists at a single MU level, which provides a new perspective for the research on corticomuscular coupling. Further study on the CMnC could help deepen our understanding of the neural control of movement.

AAAI Conference 2025 Conference Paper

CustomContrast: A Multilevel Contrastive Perspective for Subject-Driven Text-to-Image Customization

  • Nan Chen
  • Mengqi Huang
  • Zhuowei Chen
  • Yang Zheng
  • Lei Zhang
  • Zhendong Mao

Subject-driven text-to-image (T2I) customization has drawn significant interest in academia and industry. This task enables pre-trained models to generate novel images based on unique subjects. Existing studies adopt a self-reconstructive perspective, focusing on capturing all details of a single image, which will misconstrue the specific image's irrelevant attributes (e.g., view, pose, and background) as the subject intrinsic attributes. This misconstruction leads to both overfitting or underfitting of irrelevant and intrinsic attributes of the subject, i.e., these attributes are over-represented or under-represented simultaneously, causing a trade-off between similarity and controllability. In this study, we argue an ideal subject representation can be achieved by a cross-differential perspective, i.e., decoupling subject intrinsic attributes from irrelevant attributes via contrastive learning, which allows the model to focus more on intrinsic attributes through intra-consistency (features of the same subject are spatially closer) and inter-distinctiveness (features of different subjects have distinguished differences). Specifically, we propose CustomContrast, a novel framework, which includes a Multilevel Contrastive Learning (MCL) paradigm and a Multimodal Feature Injection (MFI) Encoder. The MCL paradigm is used to extract intrinsic features of subjects from high-level semantics to low-level appearance through crossmodal semantic contrastive learning and multiscale appearance contrastive learning. To facilitate contrastive learning, we introduce the MFI encoder to capture cross-modal representations. Extensive experiments show the effectiveness of CustomContrast in subject similarity and text controllability.

JBHI Journal 2025 Journal Article

Ensembled-SAMs for Enhanced Small Coronary Artery Segmentation in CCTA Images

  • Fei Chen
  • Junyao Ge
  • Yang Zheng
  • Kaitai Guo
  • Feng Cao
  • Jimin Liang

Accurate coronary artery segmentation is crucial for quantitative analysis of coronary arteries in noninvasive coronary computed tomography angiography (CCTA) images. However, current segmentation algorithms often have unsatisfactory recall due to the small size and complex morphology of coronary arteries, particularly in the distal segments. To address this issue, we introduce a new fully automated method named Ensembled-SAMs, which harnesses the strengths of the Segment Anything Model (SAM) and the no-new-U-Net (nnU-Net). First, noisy bounding box prompts are automatically generated by a vesselness algorithm that highlights the tubular structures in the CCTA images. These noisy prompts are then used to fine-tune the SAM and its two variants separately. The SAM variants introduce a classification head in their mask decoder to alleviate the false positives. In addition, an nnU-Net segmentation network is trained from scratch. Finally, the outputs of the SAMs and the nnU-Net are strategically aggregated to obtain the final segmentation result. Experiments on both a self-built dataset and the public Automated Segmentation of Coronary Arteries (ASOCA) challenge dataset demonstrate that the proposed Ensembled-SAMs outperforms the state-of-the-arts, achieving precise segmentation of coronary arteries, with particular enhancement in delineating small coronary artery segments.

ICML Conference 2025 Conference Paper

Integrating Intermediate Layer Optimization and Projected Gradient Descent for Solving Inverse Problems with Diffusion Models

  • Yang Zheng
  • Wen Li
  • Zhaoqiang Liu

Inverse problems (IPs) involve reconstructing signals from noisy observations. Recently, diffusion models (DMs) have emerged as a powerful framework for solving IPs, achieving remarkable reconstruction performance. However, existing DM-based methods frequently encounter issues such as heavy computational demands and suboptimal convergence. In this work, building upon the idea of the recent work DMPlug, we propose two novel methods, DMILO and DMILO-PGD, to address these challenges. Our first method, DMILO, employs intermediate layer optimization (ILO) to alleviate the memory burden inherent in DMPlug. Additionally, by introducing sparse deviations, we expand the range of DMs, enabling the exploration of underlying signals that may lie outside the range of the diffusion model. We further propose DMILO-PGD, which integrates ILO with projected gradient descent (PGD), thereby reducing the risk of suboptimal convergence. We provide an intuitive theoretical analysis of our approaches under appropriate conditions and validate their superiority through extensive experiments on diverse image datasets, encompassing both linear and nonlinear IPs. Our results demonstrate significant performance gains over state-of-the-art methods, highlighting the effectiveness of DMILO and DMILO-PGD in addressing common challenges in DM-based IP solvers.

NeurIPS Conference 2025 Conference Paper

Pro3D-Editor: A Progressive-Views Perspective for Consistent and Precise 3D Editing

  • Yang Zheng
  • Mengqi Huang
  • Nan Chen
  • Zhendong Mao

Text-guided 3D editing aims to precisely edit semantically relevant local 3D regions, which has significant potential for various practical applications ranging from 3D games to film production. Existing methods typically follow a view-indiscriminate paradigm: editing 2D views indiscriminately and projecting them back into 3D space. However, they overlook the different cross-view interdependencies, resulting in inconsistent multi-view editing. In this study, we argue that ideal consistent 3D editing can be achieved through a progressive-views paradigm, which propagates editing semantics from the editing-salient view to other editing-sparse views. Specifically, we propose Pro3D-Editor, a novel framework, which mainly includes Primary-view Sampler, Key-view Render, and Full-view Refiner. Primary-view Sampler dynamically samples and edits the most editing-salient view as the primary view. Key-view Render accurately propagates editing semantics from the primary view to other key views through its Mixture-of-View-Experts Low-Rank Adaption (MoVE-LoRA). Full-view Refiner edits and refines the 3D object based on the edited multi-views. Extensive experiments demonstrate that our method outperforms existing methods in editing accuracy and spatial consistency.

NeurIPS Conference 2025 Conference Paper

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing

  • Mingfei Chen
  • Zijun Cui
  • Xiulong Liu
  • Jinlin Xiang
  • Yang Zheng
  • Jingyuan Li
  • Eli Shlizerman

3D spatial reasoning in dynamic, audio-visual environments is a cornerstone of human cognition yet remains largely unexplored by existing Audio-Visual Large Language Models (AV-LLMs) and benchmarks, which predominantly focus on static or 2D scenes. We introduce SAVVY-Bench, the first benchmark for 3D spatial reasoning in dynamic scenes with synchronized spatial audio. SAVVY-Bench is comprised of thousands of carefully curated question–answer pairs probing both directional and distance relationships involving static and moving objects, and requires fine-grained temporal grounding, consistent 3D localization, and multi-modal annotation. To tackle this challenge, we propose SAVVY, a novel training-free reasoning pipeline that consists of two stages: (i) Egocentric Spatial Tracks Estimation, which leverages AV-LLMs as well as other audio-visual methods to track the trajectories of key objects related to the query using both visual and spatial audio cues, and (ii) Dynamic Global Map Construction, which aggregates multi-modal queried object trajectories and converts them into a unified global dynamic map. Using the constructed map, a final QA answer is obtained through a coordinate transformation that aligns the global map with the queried viewpoint. Empirical evaluation demonstrates that SAVVY substantially enhances performance of state-of-the-art AV-LLMs, setting a new standard and stage for approaching dynamic 3D spatial reasoning in AV-LLMs.

JBHI Journal 2024 Journal Article

A Feature Fusion Model Based on Temporal Convolutional Network for Automatic Sleep Staging Using Single-Channel EEG

  • Jiameng Bao
  • Guangming Wang
  • Tianyu Wang
  • Ning Wu
  • Shimin Hu
  • Won Hee Lee
  • Sio-Long Lo
  • Xiangguo Yan

Sleep staging is a crucial task in sleep monitoring and diagnosis, but clinical sleep staging is both time-consuming and subjective. In this study, we proposed a novel deep learning algorithm named feature fusion temporal convolutional network (FFTCN) for automatic sleep staging using single-channel EEG data. This algorithm employed a one-dimensional convolutional neural network (1D-CNN) to extract temporal features from raw EEG, and a two-dimensional CNN (2D-CNN) to extract time-frequency features from spectrograms generated through continuous wavelet transform (CWT) at the epoch level. These features were subsequently fused and further fed into a temporal convolutional network (TCN) to classify sleep stages at the sequence level. Moreover, a two-step training strategy was used to enhance the model's performance on an imbalanced dataset. Our proposed method exhibits superior performance in the 5-class classification task for healthy subjects, as evaluated on the SHHS-1, Sleep-EDF-153, and ISRUC-S1 datasets. This work provided a straightforward and promising method for improving the accuracy of automatic sleep staging using only single-channel EEG, and the proposed method exhibited great potential for future applications in professional sleep monitoring, which could effectively alleviate the workload of sleep technicians.

JBHI Journal 2024 Journal Article

Classification of Three Anesthesia Stages Based on Near-Infrared Spectroscopy Signals

  • Zhian Liu
  • Lichengxi Si
  • Shaoxian Shi
  • Jing Li
  • Jing Zhu
  • Won Hee Lee
  • Sio-Long Lo
  • Xiangguo Yan

Proper monitoring of anesthesia stages can guarantee the safe performance of clinical surgeries. In this study, different anesthesia stages were classified using near-infrared spectroscopy (NIRS) signals with machine learning. The cerebral hemodynamic variables of right proximal oxyhemoglobin (HbO 2 ) in maintenance (MNT), emergence (EM) and the consciousness (CON) stage were collected and then the differences between the three stages were compared by phase-amplitude coupling (PAC). Then combined with time-domain including linear (mean, standard deviation, max, min and range), nonlinear (sample entropy) and power in frequency-domain signal features, feature selection was performed and finally classification was performed by support vector machine (SVM) classifier. The results show that the PAC of the NIRS signal was gradually enhanced with the deepening of anesthesia level. A good three-classification accuracy of 69. 27% was obtained, which exceeded the result of classification of any single category feature. These results indicate the feasibility of NIRS signals in performing three or even more anesthesia stage classifications, providing insight into the development of new anesthesia monitoring modalities.

NeurIPS Conference 2024 Conference Paper

Inexact Augmented Lagrangian Methods for Conic Optimization: Quadratic Growth and Linear Convergence

  • Feng-Yi Liao
  • Lijun Ding
  • Yang Zheng

Augmented Lagrangian Methods (ALMs) are widely employed in solving constrained optimizations, and some efficient solvers are developed based on this framework. Under the quadratic growth assumption, it is known that the dual iterates and the Karush–Kuhn–Tucker (KKT) residuals of ALMs applied to conic programs converge linearly. In contrast, the convergence rate of the primal iterates has remained elusive. In this paper, we resolve this challenge by establishing new $\textit{quadratic growth}$ and $\textit{error bound}$ properties for primal and dual conic programs under the standard strict complementarity condition. Our main results reveal that both primal and dual iterates of the ALMs converge linearly contingent solely upon the assumption of strict complementarity and a bounded solution set. This finding provides a positive answer to an open question regarding the asymptotically linear convergence of the primal iterates of ALMs applied to conic optimization.

NeurIPS Conference 2023 Conference Paper

Inferring Hybrid Neural Fluid Fields from Videos

  • Hong-Xing Yu
  • Yang Zheng
  • Yuan Gao
  • Yitong Deng
  • Bo Zhu
  • Jiajun Wu

We study recovering fluid density and velocity from sparse multiview videos. Existing neural dynamic reconstruction methods predominantly rely on optical flows; therefore, they cannot accurately estimate the density and uncover the underlying velocity due to the inherent visual ambiguities of fluid velocity, as fluids are often shapeless and lack stable visual features. The challenge is further pronounced by the turbulent nature of fluid flows, which calls for properly designed fluid velocity representations. To address these challenges, we propose hybrid neural fluid fields (HyFluid), a neural approach to jointly infer fluid density and velocity fields. Specifically, to deal with visual ambiguities of fluid velocity, we introduce a set of physics-based losses that enforce inferring a physically plausible velocity field, which is divergence-free and drives the transport of density. To deal with the turbulent nature of fluid velocity, we design a hybrid neural velocity representation that includes a base neural velocity field that captures most irrotational energy and a vortex particle-based velocity that models residual turbulent velocity. We show that our method enables recovering vortical flow details. Our approach opens up possibilities for various learning and reconstruction applications centered around 3D incompressible flow, including fluid re-simulation and editing, future prediction, and neural dynamic scene composition. Project website: https: //kovenyu. com/HyFluid/

AAAI Conference 2023 Conference Paper

Iteratively Enhanced Semidefinite Relaxations for Efficient Neural Network Verification

  • Jianglin Lan
  • Yang Zheng
  • Alessio Lomuscio

We propose an enhanced semidefinite program (SDP) relaxation to enable the tight and efficient verification of neural networks (NNs). The tightness improvement is achieved by introducing a nonlinear constraint to existing SDP relaxations previously proposed for NN verification. The efficiency of the proposal stems from the iterative nature of the proposed algorithm in that it solves the resulting non-convex SDP by recursively solving auxiliary convex layer-based SDP problems. We show formally that the solution generated by our algorithm is tighter than state-of-the-art SDP-based solutions for the problem. We also show that the solution sequence converges to the optimal solution of the non-convex enhanced SDP relaxation. The experimental results on standard benchmarks in the area show that our algorithm achieves the state-of-the-art performance whilst maintaining an acceptable computational cost.

AAAI Conference 2023 Conference Paper

PatchNAS: Repairing DNNs in Deployment with Patched Network Architecture Search

  • Yuchu Fang
  • Wenzhong Li
  • Yao Zeng
  • Yang Zheng
  • Zheng Hu
  • Sanglu Lu

Despite being widely deployed in safety-critical applications such as autonomous driving and health care, deep neural networks (DNNs) still suffer from non-negligible reliability issues. Numerous works had reported that DNNs were vulnerable to either natural environmental noises or man-made adversarial noises. How to repair DNNs in deployment with noisy samples is a crucial topic for the robustness of neural networks. While many network repairing methods based on data argumentation and weight adjustment have been proposed, they require retraining and redeploying the whole model, which causes high overhead and is infeasible for varying faulty cases on different deployment environments. In this paper, we propose a novel network repairing framework called PatchNAS from the architecture perspective, where we freeze the pretrained DNNs and introduce a small patch network to deal with failure samples at runtime. PatchNAS introduces a novel network instrumentation method to determine the faulty stage of the network structure given the collected failure samples. A small patch network structure is searched unsupervisedly using neural architecture search (NAS) technique with data samples from deployment environment. The patch network repairs the DNNs by correcting the output feature maps of the faulty stage, which helps to maintain network performance on normal samples and enhance robustness in noisy environments. Extensive experiments based on several DNNs across 15 types of natural noises show that the proposed PatchNAS outperforms the state-of-the-arts with significant performance improvement as well as much lower deployment overhead.

AAAI Conference 2022 Conference Paper

Tight Neural Network Verification via Semidefinite Relaxations and Linear Reformulations

  • Jianglin Lan
  • Yang Zheng
  • Alessio Lomuscio

We present a novel semidefinite programming (SDP) relaxation that enables tight and efficient verification of neural networks. The tightness is achieved by combining SDP relaxations with valid linear cuts, constructed by using the reformulation-linearisation technique (RLT). The computational efficiency results from a layerwise SDP formulation and an iterative algorithm for incrementally adding RLTgenerated linear cuts to the verification formulation. The layer RLT-SDP relaxation here presented is shown to produce the tightest SDP relaxation for ReLU neural networks available in the literature. We report experimental results based on MNIST neural networks showing that the method outperforms the state-of-the-art methods while maintaining acceptable computational overheads. For networks of approximately 10k nodes (1k, respectively), the proposed method achieved an improvement in the ratio of certified robustness cases from 0% to 82% (from 35% to 70%, respectively).

IJCAI Conference 2021 Conference Paper

Efficient Neural Network Verification via Layer-based Semidefinite Relaxations and Linear Cuts

  • Ben Batten
  • Panagiotis Kouvaros
  • Alessio Lomuscio
  • Yang Zheng

We introduce an efficient and tight layer-based semidefinite relaxation for verifying local robustness of neural networks. The improved tightness is the result of the combination between semidefinite relaxations and linear cuts. We obtain a computationally efficient method by decomposing the semidefinite formulation into layerwise constraints. By leveraging on chordal graph decompositions, we show that the formulation here presented is provably tighter than current approaches. Experiments on a set of benchmark networks show that the approach here proposed enables the verification of more instances compared to other relaxation methods. The results also demonstrate that the SDP relaxation here proposed is one order of magnitude faster than previous SDP methods.

IROS Conference 2021 Conference Paper

RRT-Based Path Planning for Follow-the-Leader Motion of Hyper-Redundant Manipulators

  • Hanghang Wei
  • Yang Zheng
  • Guoying Gu

Hyper-redundant manipulators with slender body and high dexterity are widely applied for operations in confined spaces. Among the motion planning methods for these operations, the follow-the-leader motion controller is generally developed to avoid the obstacles, while the path trajectories are usually given. In this paper, we present an autonomous motion planner with a specialized rapidly exploring random tree (Sp-RRT) approach for follow-the-leader motion of hyper-redundant manipulators. Starting from the target pose in the workspace, the exploring tree can expand to multiple entrances while guaranteeing the final pose of the manipulator’s end-effector. Meanwhile, the dexterity of hyper-redundant manipulators (even with different segments) can be utilized sufficiently with customized expanding parameters. Simulation results compared with existing methods are conducted to demonstrate the aforementioned characteristics and effectiveness. For further validation, we experimentally verify the development with our custom-built hyper-redundant manipulator to realize the generated path with follow-the-leader motion.

EAAI Journal 2017 Journal Article

Design of integrated synergetic controller for the excitation and governing system of hydraulic generator unit

  • Wenlong Zhu
  • Yang Zheng
  • Jisheng Dai
  • Jianzhong Zhou

Synergetic control theory is introduced into hydraulic generator excitation system (HGES) and hydraulic generator governing system (HGRS) in this paper. Synergetic excitation controller (SEC), synergetic governing controller (SGC) of HGU have been designed. In order to enhance the terminal voltage control and mechanical power tracking performances simultaneously, the integrated synergetic controller (ISC) is also proposed. ISC implements synergetic control of terminal voltage, rotor speed, mechanical input power and guide vane opening. Namely, the ISC is considering both of the excitation system and governing system of hydraulic generator unit (HGU), which can provide control function instead of SEC and SGC. In addition, the control rules of the aforementioned three controllers are deduced from the nonlinear mathematical analytic model of hydroelectric generator unit. At the end of this paper, comparative case studies between the proposed SGC, SEC, ISC and classic PID controller are presented. The results show that the proposed ISC improves the nonlinear HGU system performance with a more accurate precision and shorter settling time in different operating conditions.

IROS Conference 2014 Conference Paper

Pedalvatar: An IMU-based real-time body motion capture system using foot rooted kinematic model

  • Yang Zheng
  • Ka-Chun Chan
  • Charlie C. L. Wang

In this paper, we present a low-cost IMU-based system, Pedalvatar, which can capture the full-body motion of users in real-time. Unlike the prior approaches using the hip-joint as the root of forward kinematic model, a foot-rooted kinematic model is developed in this work. A state change mechanism has also been investigated to allow dynamically switching the root of kinematic trees between the left and the right foot. Benefitted from this, full-body motions can be well captured in our system as long as there is at least one static foot in the movement. The ‘floating’ artifact of hip-joint rooted methods has been eliminated in our approach, and more complicated motions such as climbing stairs can be successfully captured in real-time. Comparing to those vision-based systems, this IMU-based system provides more flexibility on capturing outdoor motions that are important for many robotic applications.

v2026.09.13