Arrow Research search

Author name cluster

Yang Zhao

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

59 papers
2 author rows

Possible papers

59

EAAI Journal 2026 Journal Article

A lightweight pruning framework with minimal retraining using Taylor expansion and multi-knowledge preservation strategy

  • Suyun Lian
  • Yang Zhao
  • Jiajian Cai
  • Muxin Liao
  • Stefan Poslad
  • Jihong Pei

Artificial intelligence(AI) has become increasingly important in resource-constrained environments, where efficient model design directly impacts real-world applications. This paper implements a structured pruning framework as an AI technique, aiming to enable the deployment of deep convolutional neural networks(CNNs) in practical tasks. CNNs remain foundational in applications demanding spatial precision, such as medical imaging and autonomous driving. However, their substantial computational demands, for example the inference of visual geometry group network(VGG) of 16 layers requiring 15. 5 billion floating-point operations(FLOPs), conflict with the stringent resource restrictions of edge devices. To overcome the limitations of existing pruning methodologies, we propose a structured pruning method that integrates gradient-aware importance assessment with hierarchical error reconstruction. The framework comprises two core innovations: (1) the Hybrid-Order Importance Criterion, which quantifies channel redundancy through fused first and second-order Taylor expansion, and (2) the Multi-Knowledge Synergistic Reconstruction, which replaces traditional “prune-and-retrain” paradigms with hierarchical multi-scale feature reconstruction to preserve representational capacity. Extensive experiments across diverse architectures and various scale datasets demonstrate that our method achieves state-of-the-art(SOTA) model compression while maintaining competitive accuracy, outperforming existing structured pruning algorithms. Furthermore, the pruned models attain a four times acceleration in inference speeds, validating their practical efficacy.

IS Journal 2026 Journal Article

A Survey on Continuous Unlearning in Generative AI: Approaches and Tradeoffs

  • Yang Zhao
  • Hongyang Du
  • Yijing Lin
  • Keyi Xiang
  • Dusit Niyato
  • H. Vincent Poor

Generative artificial intelligence (GenAI) models have innovated content creation but raise concerns about privacy, security, and regulatory compliance such as General Data Protection Regulation. In response, unlearning techniques have emerged to selectively remove data while preserving the utility of the model. This article reviews unlearning methods in centralized and decentralized settings. These strategies mitigate risks, such as data leakage, membership inference, and bias amplification. By integrating unlearning with continuous or lifelong learning paradigms, GenAI models can adapt dynamically while honoring the “right to be forgotten. ” In existing unlearning methods, we explore key tradeoffs involving computational overhead, accuracy retention, generative quality, and thorough data deletion. Our review covers technical and ethical considerations and future directions, highlighting a balanced path toward responsible GenAI systems.

AAAI Conference 2026 Conference Paper

Deep Clustering Based on Sparse Kolmogorov-Arnold Network and Spectral Constraint

  • Zixuan Bi
  • Yang Zhao
  • Ganchao Liu

At present, spectral clustering is an important branch of unsupervised learning, and its application in deep learning has been widely concerned. However, for high-dimensional sparse datasets, the complexity of network scale leads to parameter explosion, and static Gaussian kernel often has wrong preset data structure. To overcome these challenges, we propose a novel deep clustering model, Deep Clustering Based on Sparse Kolmogorov-Arnold Network (KAN) and Spectral Constraint. It contains a deep sparse clustering framework, in which sparse KAN and the orthogonal layer are designed to enhance the sparsity of the activation function matrix, reduce the number of parameters and improve the stability of model convergence. Additionally, we add an adaptive optimized affinity matrix based on spectral constraint, which overcomes the limitations of static Gaussian kernels, and improves the performance and stability of spectral constraint. Experimental results on both synthetic and real datasets demonstrate that our model outperforms existing methods in clustering performance, computational efficiency, and stability.

AAAI Conference 2026 Conference Paper

Event-Guided Scene Text Image Super-Resolution

  • Zihan Qi
  • Zeyu Xiao
  • Haoyi Zhao
  • Yang Zhao
  • Feng Xue
  • Wei Jia

Scene text image super-resolution aims to enhance text legibility by recovering high-resolution text images from low-resolution inputs. However, maintaining fine details such as text strokes, edges, and textual accuracy remains challenging, particularly in low-light environments and high-speed motion scenarios, where degradation is more severe. Event cameras, with their high temporal resolution and ability to capture intensity changes, offer a promising solution for restoring lost fine details and mitigating degradation in these challenging conditions. In this paper, we propose EvTSR, the first framework that integrates Event data for scene Text image Super-Resolution. The core of EvTSR is the dual-stream frequency boost (DSFB) mechanism, which separates image features into high- and low-frequency components. High-frequency details like edges and strokes are enhanced using event data via the event-guided high-frequency (EGH) mechanism, while low-frequency components, responsible for global structure, are refined using the Text-Guided Low-frequency (TGL) mechanism with a pre-trained text recognizer, ensuring textual coherence. To further improve cross-modal integration, we introduce the cross-modal fusion (CMF) mechanism, which effectively aligns event and image features, enabling robust information fusion. Extensive experiments demonstrate that EvTSR achieves superior performance over existing methods.

TMLR Journal 2026 Journal Article

Incorporating New Knowledge into Federated Learning: Advances, Insights, and Future Directions

  • Lixu Wang
  • Sun Yinggang
  • Yang Zhao
  • Jiaqi Wu
  • Jiahua Dong
  • Ating Yin
  • Qinbin Li
  • Qingqing Ye

Federated Learning (FL) is a distributed learning approach that allows participants to collaboratively train machine learning models without sharing the raw data. It is rapidly developing in an era where privacy protection is increasingly valued. It is this rapid development trend, along with the continuous emergence of new demands for FL in the real world, that prompts us to focus on a very important problem: How to Incorporate New Knowledge into Federated Learning? The primary challenge here is to effectively and timely incorporate various new knowledge into existing FL systems and evolve these systems to reduce costs, upgrade functionalities, and facilitate sustainable development. In the meantime, established FL systems should preserve existing functionalities during the incorporation of new knowledge. In this paper, we systematically define the main sources of new knowledge in FL, including new features, tasks, models, and algorithms. For each source, we thoroughly analyze and discuss the technical approaches for incorporating new knowledge into existing FL systems and examine the impact of the form and timing of new knowledge arrival on the incorporation process. Unlike prior surveys that primarily catalogue FL techniques under a fixed system specification, we adopt a lifecycle evolution perspective and synthesize methods that enable time-varying integration of new features, tasks, models, and aggregation algorithms while preserving existing functionality. Furthermore, we comprehensively discuss the potential future directions for FL, incorporating new knowledge and considering a variety of factors, including scenario setups, security and privacy threats, and incentives.

AAAI Conference 2026 Conference Paper

LSAP-PV: High-Fidelity Palm Vein Image Synthesis via Layered Spectral Absorption Projection-Guided Diffusion Model

  • Sheng Shang
  • Chenglong Zhao
  • Ruixin Zhang
  • Jianlong Jin
  • Jingyun Zhang
  • Jun Wang
  • Yang Zhao
  • Shouhong Ding

Palm vein recognition has emerged as a promising biometric technology, yet its development remains constrained by the scarcity of large-scale publicly available datasets. Several methods of palm vein image generation have been proposed to address this issue. These methods usually focus on the anatomical realism of palm vein patterns, but overlook the biophysical correlation between identities and vein patterns, particularly in simulating identity-specific vein contrast. To tackle this limitation, we propose a novel biophysics-driven synthesis method. Our method constructs a 3D palm vascular tree via established modeling method. Then, a projection model is proposed to map the 3D tree into 2D space to derive palm vein patterns. The projection model is based on skin spectral absorption and simulates the natural attenuation of light passing through the skin using a layer integration method. For different identities, we sample different skin parameters, resulting in varying degrees of attenuation. This method effectively simulates the variation in vein contrast across different identities. Furthermore, we introduce a conditional diffusion model that uses the projected patterns as identity conditions to generate palm vein images. To the best of our knowledge, this is the first palm vein generation method based on the diffusion model. Experimental results demonstrate that our method not only outperforms existing methods, but also enables a recognition model trained on our synthetic data to achieve superior performance compared to a model trained on real-world data at a scale of 2,000 IDs under an open-set protocol with a TAR@FAR=1:1 of 1e-4.

JBHI Journal 2025 Journal Article

A Novel Dynamic Latent Variables-Based Framework for Enhancing Freezing of Gait Detection in Parkinson's Disease Patients

  • Xuan Wang
  • Lisha Yu
  • S. Joe Qin
  • Yang Zhao

Freezing of Gait (FOG) is one of the most severe symptoms of Parkinson's disease (PD), which often lead to life-threatening falls. Wearable sensor-based technologies coupled with data driven methods have advanced the detection of FOG in a timely fashion. However, most existing monitoring methods overlook the dynamics of processes when extracting effective information from high-dimensional sensor data. To tackle these problems, we develop a novel framework for FOG detection by integrating Dynamic Latent Variable (DLV)-based dimensionality reduction strategies and personalized monitoring. First, a multi-channel sliding window mechanism is adopted to extract the multiple potentially effective feature sequences. Second, an interpretable DLV-based method incorporating time-lagged terms is designed for the subspace representation of complex high-dimensional sequences. Third, the extracted DLVs are integrated with threshold-based methods or the Statistical Process Control (SPC) method for anomaly detection. We identified distinct variations in gait patterns among individuals, underscoring the importance of personalized approaches. The proposed framework demonstrates its effectiveness in FOG detection via validating on real world dataset, achieving a sensitivity of $\mathbf {0. 845} \pm \mathbf {0. 254}$ and a specificity of $\mathbf {0. 842} \pm \mathbf {0. 211}$.

NeurIPS Conference 2025 Conference Paper

Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation

  • Shanchuan Lin
  • Ceyuan Yang
  • Hao He
  • Jianwen Jiang
  • Yuxi Ren
  • Xin Xia
  • Yang Zhao
  • Xuefeng Xiao

Existing large-scale video generation models are computationally intensive, preventing adoption in real-time and interactive applications. In this work, we propose autoregressive adversarial post-training (AAPT) to turn a pre-trained latent video diffusion model into a real-time, interactive, streaming video generator. Our model autoregressively generates a latent frame at a time using a single neural function evaluation (1NFE). The model can stream the result to the user in real time and receive interactive responses as control to generate the next latent frame. Unlike existing approaches, our method explores adversarial training as an effective paradigm for autoregressive generation. This allows us to design a more efficient architecture for one-step generation and to train the model in a student-forcing way to mitigate error accumulation. The adversarial approach also enables us to train the model for long-duration generation fully utilizing the KV cache. As a result, our 8B model achieves real-time, 24fps, nonstop, streaming video generation at 736x416 resolution on a single H100, or 1280x720 on 8xH100 up to a minute long (1440 frames).

AAAI Conference 2025 Conference Paper

E4: Energy-Efficient DNN Inference for Edge Video Analytics via Early Exiting and DVFS

  • Ziyang Zhang
  • Yang Zhao
  • Ming-Ching Chang
  • Changyao Lin
  • Jie Liu

Deep neural network (DNN) models are increasingly popular in edge video analytic applications. However, the computeintensive nature of DNN models pose challenges for energyefficient inference on resource-constrained edge devices. Most existing solutions focus on optimizing DNN inference latency and accuracy, often overlooking energy efficiency. They also fail to account for the varying complexity of video frames, leading to sub-optimal performance in edge video analytics. In this paper, we propose an EnergyEfficient Early-Exit (E4) framework that enhances DNN inference efficiency for edge video analytics by integrating a novel early-exit mechanism with dynamic voltage and frequency scaling (DVFS) governors. It employs an attentionbased cascade module to analyze video frame diversity and automatically determine optimal DNN exit points. Additionally, E4 features a just-in-time (JIT) profiler that uses coordinate descent search to co-optimize CPU and GPU clock frequencies for each layer before the DNN exit points. Extensive evaluations demonstrate that E4 outperforms current state-of-the-art methods, achieving up to 2.8× speedup and 26% average energy saving while maintaining high accuracy.

NeurIPS Conference 2025 Conference Paper

How to Auto-optimize Prompts for Domain Tasks? Adaptive Prompting and Reasoning through Evolutionary Domain Knowledge Adaptation

  • Yang Zhao
  • Pu Wang
  • Hao Frank Yang

Designing optimal prompts and reasoning processes for large language models (LLMs) on domain-specific tasks is both necessary and challenging in real-world applications. Determining how to integrate domain knowledge, enhance reasoning efficiency, and even provide domain experts with refined knowledge integration hints are particularly crucial yet unresolved tasks. In this research, we propose Evolutionary Graph Optimization for Prompting (EGO-Prompt), an automated framework to designing better prompts, efficient reasoning processes and providing enhanced causal-informed process. EGO-Prompt begins with a general prompt and fault-tolerant initial Semantic Causal Graph (SCG) descriptions, constructed by human experts, which is then automatically refined and optimized to guide LLM reasoning. Recognizing that expert-defined SCGs may be partial or imperfect and that their optimal integration varies across LLMs, EGO-Prompt integrates a novel causal-guided textual gradient process in two steps: first, generating nearly deterministic reasoning guidance from the SCG for each instance, and second, adapting the LLM to effectively utilize the guidance alongside the original input. The iterative optimization algorithm further refines both the SCG and the reasoning mechanism using textual gradients with ground-truth. We tested the framework on real-world public health, transportation and human behavior tasks. EGO-Prompt achieves 7. 32\%–12. 61\% higher F1 than cutting-edge methods, and allows small models to reach the performence of larger models at under 20\% of the original cost. It also outputs a refined, domain-specific SCG that improves interpretability.

EAAI Journal 2025 Journal Article

Medical artificial intelligence for early detection of lung cancer: A survey

  • Guohui Cai
  • Ying Cai
  • Zeyu Zhang
  • Yuanzhouhan Cao
  • Lin Wu
  • Daji Ergu
  • Zhibin Liao
  • Yang Zhao

Lung cancer remains one of the leading causes of morbidity and mortality worldwide, making early diagnosis critical for improving therapeutic outcomes and patient prognosis. Computer-aided diagnosis systems, which analyze computed tomography images, have proven effective in detecting and classifying pulmonary nodules, significantly enhancing the detection rate of early-stage lung cancer. Although traditional machine learning algorithms have been valuable, they exhibit limitations in handling complex sample data. The recent emergence of deep learning has revolutionized medical image analysis, driving substantial advancements in this field. This review focuses on recent progress in deep learning for pulmonary nodule detection, segmentation, and classification. Traditional machine learning methods, such as support vector machines and k-nearest neighbors, have shown limitations, paving the way for advanced approaches like Convolutional Neural Networks, Recurrent Neural Networks, and Generative Adversarial Networks. The integration of ensemble models and novel techniques is also discussed, emphasizing the latest developments in lung cancer diagnosis. Deep learning algorithms, combined with various analytical techniques, have markedly improved the accuracy and efficiency of pulmonary nodule analysis, surpassing traditional methods, particularly in nodule classification. Although challenges remain, continuous technological advancements are expected to further strengthen the role of deep learning in medical diagnostics, especially for early lung cancer detection and diagnosis. A comprehensive list of lung cancer detection models reviewed in this work is available at https: //github. com/CaiGuoHui123/Awesome-Lung-Cancer-Detection.

NeurIPS Conference 2025 Conference Paper

Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet Architectures

  • Shuqing Luo
  • Ye Han
  • Pingzhi Li
  • Jiayin Qin
  • Jie Peng
  • Yang Zhao
  • Yu Cao
  • Tianlong Chen

Mixture-of-Experts (MoE) architecture offers enhanced efficiency for Large Language Models (LLMs) with modularized computation, yet its inherent sparsity poses significant hardware deployment challenges, including memory locality issues, communication overhead, and inefficient computing resource utilization. Inspired by the modular organization of the human brain, we propose $\texttt{Mozart}$, a novel algorithm-hardware co-design framework tailored for efficient training of MoE-based LLMs on 3. 5D wafer-scale chiplet architectures. On the algorithm side, $\texttt{Mozart}$ exploits the inherent modularity of chiplets and introduces: ($1$) an expert allocation strategy that enables efficient on-package all-to-all communication, and ($2$) a fine-grained scheduling mechanism that improves communication-computation overlap through streaming tokens and experts. On the architecture side, $\texttt{Mozart}$ adaptively co-locates heterogeneous modules on specialized chiplets with a 2. 5D NoP-Tree topology and hierarchical memory structure. Evaluation across three popular MoE models demonstrates significant efficiency gains, enabling more effective parallelization and resource utilization for large-scale modularized MoE-LLMs.

IJCAI Conference 2025 Conference Paper

MSMAR-RL: Multi-Step Masked-Attention Recovery Reinforcement Learning for Safe Maneuver Decision in High-Speed Pursuit-Evasion Game

  • Yang Zhao
  • Wenzhe Zhao
  • Xuelong Li

Ensuring the safety of high-speed agent in dynamic adversarial environments, such as pursuit-evasion games with target-purchase and obstacle-avoidance, is a significant challenge. Existing reinforcement learning methods often fail to balance safety and reward under strict safety constraints and diverse environmental conditions. To address these limitations, this paper proposes a novel zero-constraint-violation recovery RL framework tailored for high-speed uav pursuit-evasion combat games. The framework includes three key innovations. (1) An extendable multi-step reach-avoid theory: we provide a zero-constraint-violation safety guarantee for multi-strategy reinforcement learning and enabling early danger detection in high speed game. (2) A masked-attention recovery strategy: we introduce a padding-mask attention architecture to handle spatiotemporal variations in dynamic obstacles with varying threat levels. (3) Experimental validation: we validate the framework in obstacle-rich pursuit-evasion scenarios, demonstrating its superiority through comparison with other algorithm and ablation studies. Our approach also shows potential for extension to other rapid-motion tasks and more complex hazardous scenarios. Details and code could be found at https: //msmar-rl. github. io.

JBHI Journal 2025 Journal Article

MSTG-Transformer: Multivariate Spatial-Temporal Gated Transformer Model for 3D Skeleton Data-based Fall Risk Prediction

  • Junjie Cao
  • Xuan Wang
  • Keyi Huang
  • Lisha Yu
  • Xiaomao Fan
  • Yang Zhao

As the aging population continues to grow, falls among older adults have become a significant public health concern worldwide. Data-driven approaches for effective fall risk prediction, which integrate standard functional tests with 3D skeleton data from depth sensors, are gaining increasing attention. However, the complex physiological and functional interactions among skeletal keypoints during ambulation pose challenges for multidimensional feature extraction in most predictive models. In this study, we developed a novel approach based on preprocessed 3D skeleton data, named Multivariate SpatialTemporal Gated Transformer (MSTG-Transformer). This approach consists of three main stages. First, gait cycle sequences are constructed to sophisticatedly depict the movement patterns of subjects, amplifying the distinctions between groups. Then, spatial and topological features are extracted via convolutional modules, and a dual-stream encoder block is employed to encode the features of 3D skeleton data across both time steps and time channels. Finally, a voting scheme is used to determine fall risk by integrating the classification results of individual gait cycle segments. Validation experiments on a real-world dataset demonstrate that our proposed approach outperforms classical methods, achieving a superior prediction accuracy of 0. 9510 ± 0. 0240. Additionally, our study highlights the crucial role of potential interactions between skeletal keypoints in accurately predicting fall risk

AAAI Conference 2025 Conference Paper

Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing

  • Pengcheng Zhao
  • Jinxing Zhou
  • Yang Zhao
  • Dan Guo
  • Yanxiang Chen

The Audio-Visual Video Parsing task aims to recognize and temporally localize all events occurring in either the audio or visual stream, or both. Capturing accurate event semantics for each audio/visual segment is vital. Prior works directly utilize the extracted holistic audio and visual features for intra- and cross-modal temporal interactions. However, each segment may contain multiple events, resulting in semantically mixed holistic features that can lead to semantic interference during intra- or cross-modal interactions: the event semantics of one segment may incorporate semantics of unrelated events from other segments. To address this issue, our method begins with a Class-Aware Feature Decoupling (CAFD) module, which explicitly decouples the semantically mixed features into distinct class-wise features, including multiple event-specific features and a dedicated background feature. The decoupled class-wise features enable our model to selectively aggregate useful semantics for each segment from clearly matched classes contained in other segments, preventing semantic interference from irrelevant classes. Specifically, we further design a Fine-Grained Semantic Enhancement module for encoding intra- and cross-modal relations. It comprises a Segment-wise Event Co-occurrence Modeling (SECM) block and a Local-Global Semantic Fusion (LGSF) block. The SECM exploits inter-class dependencies of concurrent events within the same timestamp with the aid of a novel event co-occurrence loss. The LGSF further enhances the event semantics of each segment by incorporating relevant semantics from more informative global video features. Extensive experiments validate the effectiveness of the proposed modules and loss functions, resulting in a new state-of-the-art parsing performance.

NeurIPS Conference 2025 Conference Paper

Personalized Decision Modeling: Utility Optimization or Textualized-Symbolic Reasoning

  • Yibo Zhao
  • Yang Zhao
  • Hongru Du
  • Hao Frank Yang

Decision-making models for individuals, particularly in high-stakes scenarios like vaccine uptake, often diverge from population optimal predictions. This gap arises from the uniqueness of the individual decision-making process, shaped by numerical attributes (e. g. , cost, time) and linguistic influences (e. g. , personal preferences and constraints). Developing upon Utility Theory and leveraging the textual-reasoning capabilities of Large Language Models (LLMs), this paper proposes an Adaptive Textual-symbolic Human-centric Reasoning framework (ATHENA) to address the optimal information integration. ATHENA uniquely integrates two stages: First, it discovers robust, group-level symbolic utility functions via LLM-augmented symbolic discovery; Second, it implements individual-level semantic adaptation, creating personalized semantic templates guided by the optimal utility to model personalized choices. Validated on real-world travel mode and vaccine choice tasks, ATHENA consistently outperforms utility-based, machine learning, and other LLM-based models, lifting F1 score by at least 6. 5\% over the strongest cutting-edge models. Further, ablation studies confirm that both stages of ATHENA are critical and complementary, as removing either clearly degrades overall predictive performance. By organically integrating symbolic utility modeling and semantic adaptation, ATHENA provides a new scheme for modeling human-centric decisions. The project page can be found at https: //yibozh. github. io/Athena.

AAAI Conference 2025 Conference Paper

PVTree: Realistic and Controllable Palm Vein Generation for Recognition Tasks

  • Sheng Shang
  • Chenglong Zhao
  • Ruixin Zhang
  • Jianlong Jin
  • Jingyun Zhang
  • Rizen Guo
  • Shouhong Ding
  • Yunsheng Wu

Palm vein recognition is an emerging biometric technology that offers enhanced security and privacy. However, acquiring sufficient palm vein data for training deep learning-based recognition models is challenging due to the high costs of data collection and privacy protection constraints. This has led to a growing interest in generating pseudo-palm vein data using generative models. Existing methods, however, often produce unrealistic palm vein patterns or struggle with controlling identity and style attributes. To address these issues, we propose a novel palm vein generation framework named PVTree. First, the palm vein identity is defined by a complex and authentic 3D palm vascular tree, created using an improved Constrained Constructive Optimization (CCO) algorithm. Second, palm vein patterns of the same identity are generated by projecting the same 3D vascular tree into 2D images from different views and converting them into realistic images using a generative model. As a result, PVTree satisfies the need for both identity consistency and intra-class diversity. Extensive experiments conducted on several publicly available datasets demonstrate that our proposed palm vein generation method surpasses existing methods and achieves a higher TAR@FAR=1e-4 under the 1:1 Open-set protocol. To the best of our knowledge, this is the first time that the performance of a recognition model trained on synthetic palm vein data exceeds that of the recognition model trained on real data, which indicates that palm vein image generation research has a promising future.

NeurIPS Conference 2025 Conference Paper

QiMeng-CodeV-R1: Reasoning-Enhanced Verilog Generation

  • Yaoyu Zhu
  • Di Huang
  • Hanqi Lyu
  • Xiaoyun Zhang
  • Chongxiao Li
  • Wenxuan Shi
  • Yutong Wu
  • Jianan Mu

Large language models (LLMs) trained via reinforcement learning with verifiable reward (RLVR) have achieved breakthroughs on tasks with explicit, automatable verification, such as software programming and mathematical problems. Extending RLVR to electronic design automation (EDA), especially automatically generating hardware description languages (HDLs) like Verilog from natural-language (NL) specifications, however, poses three key challenges: the lack of automated and accurate verification environments, the scarcity of high-quality NL-code pairs, and the prohibitive computation cost of RLVR. To this end, we introduce CodeV-R1, an RLVR framework for training Verilog generation LLMs. First, we develop a rule-based testbench generator that performs robust equivalence checking against golden references. Second, we propose a round-trip data synthesis method that pairs open-source Verilog snippets with LLM-generated NL descriptions, verifies code–NL–code consistency via the generated testbench, and filters out inequivalent examples to yield a high-quality dataset. Third, we employ a two-stage "distill-then-RL" training pipeline: distillation for the cold start of reasoning abilities, followed by adaptive DAPO, our novel RLVR algorithm that can reduce training cost by adaptively adjusting sampling rate. The resulting model, CodeV-R1-7B, achieves 68. 6 \% and 72. 9 \% pass@1 on VerilogEval v2 and RTLLM v1. 1, respectively, surpassing prior state-of-the-art by 12$\sim$20 \%, while even exceeding the performance of 671B DeepSeek-R1 on RTLLM. We have released our model, training code, and dataset to facilitate research in EDA and LLM communities.

ICML Conference 2025 Conference Paper

RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models

  • Quan Wei 0001
  • Chung-Yiu Yau
  • Hoi-To Wai
  • Yang Zhao
  • Dongyeop Kang
  • Youngsuk Park
  • Mingyi Hong 0001

Supervised fine-tuning is a standard method for adapting pre-trained large language models (LLMs) to downstream tasks. Quantization has been recently studied as a post-training technique for efficient LLM deployment. To obtain quantized fine-tuned LLMs, conventional pipelines would first fine-tune the pre-trained models, followed by post-training quantization. This often yields suboptimal performance as it fails to leverage the synergy between fine-tuning and quantization. To effectively realize low-bit quantization of weights, activations and KV caches in LLMs, we propose an algorithm named Rotated Straight-Through-Estimator (RoSTE), which combines quantization-aware supervised fine-tuning (QA-SFT) with an adaptive rotation strategy that identifies an effective rotation configuration to reduce activation outliers. We provide theoretical insights on RoSTE by analyzing its prediction error when applied to an overparameterized least square quantized training problem. Our findings reveal that the prediction error is directly proportional to the quantization error of the converged weights, which can be effectively managed through an optimized rotation configuration. Experiments on Pythia, Qwen and Llama models of different sizes demonstrate the effectiveness of RoSTE. Compared to existing post-SFT quantization baselines, our method consistently achieves superior performances across various tasks and different LLM architectures. Our code is available at https: //github. com/OptimAI-Lab/RoSTE.

JBHI Journal 2025 Journal Article

Sensor-Based Digital Biomarkers for Early Identification of Cognitive Frailty: A Systematic Review

  • Ruiyuan Wu
  • Keyi Huang
  • Kwok-Leung Tsui
  • Jianbang Xiang
  • Yang Zhao

Cognitive Frailty (CF), characterized by the co-occurrence of Physical Frailty (PF) and Mild Cognitive Impairment (MCI), is increasingly recognized as a critical predictor of adverse health outcomes in aging populations. Despite its clinical significance, early identification of CF remains challenging due to heterogeneous diagnostic criteria and reliance on resource-intensive assessments. This systematic review synthesizes current evidence on sensor-based digital biomarkers for CF detection across eight databases following PRISMA 2020 guidelines. A combined qualitative and quantitative synthesis is conducted based on 20 eligible studies. Key findings reveal that Inertial Measurement Units (IMUs) are the most frequently utilized sensors (n = 6 studies), primarily capturing gait and motor function parameters under various functional tasks. Among digital biomarkers, gait features, particularly Stride Length under single-task conditions and Velocity under dual-task or single-task conditions, show significant associations with CF. Body composition metrics, particularly Appendicular Skeletal Muscle Mass Index (ASM), and physical activity parameters, such as Moderate-to-Vigorous-intensity Physical Activity (MVPA), show consistent negative associations with CF, while cardiovascular indices like Cardio-Ankle Vascular Index (CAVI) reveal positive correlations. Eighteen studies develop a total of 23 CF-related models, of which only four focus on predictive classification. These four studies report ten distinct models, with AUCs ranging from 0. 696 to 0. 911 and accuracy from 0. 480 to 0. 863. A quantitative synthesis of six models from three studies yields pooled estimates of sensitivity at 0. 81 (95% CI: 0. 55-0. 94) and specificity at 0. 80 (95% CI: 0. 49-0. 95), suggesting moderate-to-good discriminative performance despite notable heterogeneity in diagnostic definitions, sample sizes, sensor types, and modeling approaches. These results systematically map the evolving landscape of digital biomarkers for CF, emphasizing the need for standardized sensor protocols and rigorous validation to advance clinical translation.

NeurIPS Conference 2025 Conference Paper

Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study

  • Guanlin (Frank) Wu
  • Boyan Su
  • Yang Zhao
  • Pu Wang
  • Yichen Lin
  • Hao Frank Yang

How to integrate and verify spatial intelligence in foundation models remains an open challenge. Current practice often proxies Visual-Spatial Intelligence (VSI) with purely textual prompts and VQA-style scoring, which obscures geometry, invites linguistic shortcuts, and weakens attribution to genuinely spatial skills. We introduce Spatial Intelligence Grid (SIG): a structured, grid-based schema that explicitly encodes object layouts, inter-object relations, and physically grounded priors. As a complementary channel to text, SIG provides a faithful, compositional representation of scene structure for foundation-model reasoning. Building on SIG, we derive SIG-informed evaluation metrics that quantify a model’s intrinsic VSI, which separates spatial capability from language priors. In few-shot in-context learning with state-of-the-art multimodal LLMs (e. g. GPT- and Gemini-family models), SIG yields consistently larger, more stable, and more comprehensive gains across all VSI metrics compared to VQA-only representations, indicating its promise as a data-labeling and training schema for learning VSI. We also release SIGBench, a benchmark of 1. 4K driving frames annotated with ground-truth SIG labels and human gaze traces, supporting both grid-based machine VSI tasks and attention-driven, human-like VSI tasks in autonomous-driving scenarios.

YNIMG Journal 2025 Journal Article

Transcranial brain atlas based on photon measurement density function in a triple-parameter standard channel space

  • Lijiang Wei
  • Yang Zhao
  • Farui Liu
  • Yuanyuan Chen
  • Yilong Xu
  • Zheng Li
  • Chaozhe Zhu

Functional near-infrared spectroscopy (fNIRS) is a widely-used transcranial brain imaging technique in neuroscience research. Nevertheless, the lack of anatomical information from recordings poses challenges for designing appropriate optode montages and for localizing fNIRS signals to underlying anatomical regions. The photon measurement density function (PMDF) is often employed to address these issues, as it accurately measures the sensitivity of an fNIRS channel to perturbations of absorption coefficients at any brain location. However, existing PMDF-based localization methods have two limitations: (1) limited channel space, and (2) estimation based on a single standard head model, which usually differ anatomically from individuals. To overcome these limitations, this study proposes a continuous standard channel space for fNIRS and constructs a PMDF-based transcranial brain atlas (PMDF-TBA) by calculating PMDFs using MRI data from 48 adults. The PMDF-TBA contains group-averaged sensitivities of channels to gray matter and brain regions as defined in 3 atlases: Brodmann, AAL2, and LPBA40. We evaluated the prediction ability of PMDF-TBA for sensitivity of unseen individuals. The results show that it outperformed PMDFs based on single standard head models, making PMDF-TBA a more generalizable fNIRS spatial localization tool. Therefore, in the absence of individual sMRI data, PMDF-TBA can optimize optode montage design, enhance channel sensitivity in target brain regions, and assist in source localization for fNIRS data, thereby facilitating the application of fNIRS in neuroscience research.

NeurIPS Conference 2025 Conference Paper

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection

  • Yang Zhao
  • Kai Xiong
  • Xiao Ding
  • Li Du
  • Yangou Ouyang
  • Zhouhao Sun
  • Jiannan Guan
  • Wenbin Zhang

A primary impediment to scaling reinforcement learning (RL) for large language model (LLM) training is the substantial computational cost, predominantly arising from the necessity of multi-sampling for policy optimization and evaluation. This underscores the critical yet challenging nature of efficient training data selection. Drawing inspiration from the Zone of Proximal Development (ZPD) theory, which posits that learners acquire knowledge more effectively from tasks of intermediate difficulty, we hypothesize that LLMs exhibit optimal learning from data they have not yet mastered but demonstrate the potential to comprehend. Conventional methodologies for assessing data difficulty or informativeness typically rely on computationally intensive multi-sampling or iterative procedures. To address this limitation, we introduce UFO-RL (**U**ncertainty-**F**ocused **O**ptimization for **R**einforcement **L**earning), a novel framework that employs a computationally efficient single-pass uncertainty estimation technique to identify informative training instances. This method, requiring only a single forward pass and obviating the need for iterative next-token computation, achieves a significant acceleration (up to 185$\times$) in data evaluation compared to multi-sampling approaches. UFO-RL leverages this efficient metric to select data within the model's estimated ZPD for training. Extensive experimentation across diverse LLMs and mathematical benchmarks demonstrates that training with a mere 10\% of the data, carefully selected by UFO-RL, yields performance comparable to or even surpassing that of full-data training. Furthermore, this targeted data selection results in up to a 16$\times$ reduction in overall training time, concurrently enhancing training stability and improving generalization capabilities. Thus, UFO-RL presents a practical and highly efficient strategy for scaling RL fine-tuning of LLMs by focusing learning efforts on the most informative and valuable data, thereby mitigating the computational bottlenecks associated with traditional RL training.

NeurIPS Conference 2025 Conference Paper

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

  • Jiaming Han
  • Hao Chen
  • Yang Zhao
  • Hanyu Wang
  • Qi Zhao
  • Ziyan Yang
  • Hao He
  • Xiangyu Yue

This paper presents a multimodal framework that attempts to unify visual understanding and generation within a shared discrete semantic representation. At its core is the Text-Aligned Tokenizer (TA-Tok), which converts images into discrete tokens using a text-aligned codebook projected from a large language model's (LLM) vocabulary. By integrating vision and text into a unified space with an expanded vocabulary, our multimodal LLM, Tar, enables cross-modal input and output through a shared interface, without the need for modality-specific designs. Additionally, we propose scale-adaptive encoding and decoding to balance efficiency and visual detail, along with a generative de-tokenizer to produce high-fidelity visual outputs. To address diverse decoding needs, we utilize two complementary de-tokenizers: a fast autoregressive model and a diffusion-based model. To enhance modality fusion, we investigate advanced pre-training tasks, demonstrating improvements in both visual understanding and generation. Experiments across benchmarks show that Tar matches or surpasses existing multimodal LLM methods, achieving faster convergence and greater training efficiency. All code, models, and data will be made publicly available.

IROS Conference 2024 Conference Paper

An Optimization-Based Planner with B-spline Parameterized Continuous-Time Reference Signals

  • Chuyuan Tao
  • Sheng Cheng 0001
  • Fanxin Wang
  • Yang Zhao
  • Naira Hovakimyan

For the cascaded planning and control modules implemented for robot navigation, the frequency gap between the planner and controller has received limited attention. In this study, we introduce a novel B-spline parameterized optimization-based planner (BSPOP) designed to address the frequency gap challenge with limited onboard computational power in robots. The proposed planner generates continuous-time control inputs for low-level controllers running at arbitrary frequencies to track. Furthermore, when considering the convex control action sets, BSPOP uses the convex hull property to automatically constrain the continuous-time control inputs within the convex set. Consequently, compared with the discrete-time optimization-based planners, BSPOP reduces the number of decision variables and inequality constraints, which improves computational efficiency as a byproduct. Simulation results demonstrate that our approach can achieve a comparable planning performance to the high-frequency baseline optimization-based planners while demanding less computational power. Both simulation and experiment results show that the proposed method performs better in planning compared with baseline planners in the same frequency.

NeurIPS Conference 2024 Conference Paper

Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

  • Haifeng Huang
  • Yilun Chen
  • Zehan Wang
  • Rongjie Huang
  • Runsen Xu
  • Tai WANG
  • Luping Liu
  • Xize Cheng

Recent advancements in 3D Large Language Models (LLMs) have demonstrated promising capabilities for 3D scene understanding. However, previous methods exhibit deficiencies in general referencing and grounding capabilities for intricate scene comprehension. In this paper, we introduce the use of object identifiers and object-centric representations to interact with scenes at the object level. Specifically, we decompose the input 3D scene into a set of object proposals, each assigned a unique identifier token, which enables efficient object referencing and grounding during user-assistant interactions. Given the scarcity of scene-language data, we model the scene embeddings as a sequence of explicit object-level embeddings, derived from semantic-rich 2D or 3D representations. By employing object identifiers, we transform diverse 3D scene-language tasks into a unified question-answering format, facilitating joint training without the need for additional task-specific heads. With minimal fine-tuning on all downstream tasks, our model significantly outperforms existing methods on benchmarks including ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, and SQA3D.

NeurIPS Conference 2024 Conference Paper

Extending Multi-modal Contrastive Representations

  • Ziang Zhang
  • Zehan Wang
  • Luping Liu
  • Rongjie Huang
  • Xize Cheng
  • Zhenhui Ye
  • Wang Lin
  • Huadai Liu

Multi-modal contrastive representation (MCR) of more than three modalities is critical in multi-modal learning. Although recent methods showcase impressive achievements, the high dependence on large-scale, high-quality paired data and the expensive training costs limit their further development. Inspired by recent C-MCR, this paper proposes $\textbf{Ex}$tending $\textbf{M}$ultimodal $\textbf{C}$ontrastive $\textbf{R}$epresentation (Ex-MCR), a training-efficient and paired-data-free method to build unified contrastive representation for many modalities. Since C-MCR is designed to learn a new latent space for the two non-overlapping modalities and projects them onto this space, a significant amount of information from their original spaces is lost in the projection process. To address this issue, Ex-MCR proposes to extend one modality's space into the other's, rather than mapping both modalities onto a completely new space. This method effectively preserves semantic alignment in the original space. Experimentally, we extend pre-trained audio-text and 3D-image representations to the existing vision-text space. Without using paired data, Ex-MCR achieves comparable performance to advanced methods on a series of audio-image-text and 3D-image-text tasks and achieves superior performance when used in parallel with data-driven methods. Moreover, semantic alignment also emerges between the extended modalities (e. g. , audio and 3D).

NeurIPS Conference 2024 Conference Paper

Image Understanding Makes for A Good Tokenizer for Image Generation

  • Luting Wang
  • Yang Zhao
  • Zijian Zhang
  • Jiashi Feng
  • Si Liu
  • Bingyi Kang

Modern image generation (IG) models have been shown to capture rich semantics valuable for image understanding (IU) tasks. However, the potential of IU models to improve IG performance remains uncharted. We address this issue using a token-based IG framework, which relies on effective tokenizers to project images into token sequences. Currently, **pixel reconstruction** (e. g. , VQGAN) dominates the training objective for image tokenizers. In contrast, our approach adopts the **feature reconstruction** objective, where tokenizers are trained by distilling knowledge from pretrained IU encoders. Comprehensive comparisons indicate that tokenizers with strong IU capabilities achieve superior IG performance across a variety of metrics, datasets, tasks, and proposal networks. Notably, VQ-KD CLIP achieves $4. 10$ FID on ImageNet-1k (IN-1k). Visualization suggests that the superiority of VQ-KD can be partly attributed to the rich semantics within the VQ-KD codebook. We further introduce a straightforward pipeline to directly transform IU encoders into tokenizers, demonstrating exceptional effectiveness for IG tasks. These discoveries may energize further exploration into image tokenizer research and inspire the community to reassess the relationship between IU and IG. The code is released at https: //github. com/magic-research/vector_quantization.

AAAI Conference 2024 Conference Paper

PCE-Palm: Palm Crease Energy Based Two-Stage Realistic Pseudo-Palmprint Generation

  • Jianlong Jin
  • Lei Shen
  • Ruixin Zhang
  • Chenglong Zhao
  • Ge Jin
  • Jingyun Zhang
  • Shouhong Ding
  • Yang Zhao

The lack of large-scale data seriously hinders the development of palmprint recognition. Recent approaches address this issue by generating large-scale realistic pseudo palmprints from Bézier curves. However, the significant difference between Bézier curves and real palmprints limits their effectiveness. In this paper, we divide the Bézier-Real difference into creases and texture differences, thus reducing the generation difficulty. We introduce a new palm crease energy (PCE) domain as a bridge from Bézier curves to real palmprints and propose a two-stage generation model. The first stage generates PCE images (realistic creases) from Bézier curves, and the second stage outputs realistic palmprints (realistic texture) with PCE images as input. In addition, we also design a lightweight plug-and-play line feature enhancement block to facilitate domain transfer and improve recognition performance. Extensive experimental results demonstrate that the proposed method surpasses state-of-the-art methods. Under extremely few data settings like 40 IDs (only 2.5% of the total training set), our model achieves a 29% improvement over RPG-Palm and outperforms ArcFace with 100% training set by more than 6% in terms of TAR@FAR=1e-6.

JBHI Journal 2024 Journal Article

Sensor-Based Multifaceted Feature Extraction and Ensemble Elastic Net Approach for Assessing Fall Risk in Community-Dwelling Older Adults

  • Xuan Wang
  • Lisha Yu
  • Hailiang Wang
  • Kwok Leung Tsui
  • Yang Zhao

Accurate identification of community-dwelling older adults at high fall risk can facilitate timely intervention and significantly reduce fall incidents. Analyzing gait and balance capabilities via feature extraction and modeling through sensor-based motion data has emerged as a viable approach for fall risk assessment. However, the existing approaches for extracting key features related to fall risk lack inclusiveness, with limited consideration of the non-linear characteristics of sensor signals, such as signal complexity, self-similarity, and local stability. In this study, we developed a multifaceted feature extraction scheme employing diverse feature types, including demographic, descriptive statistical, non-linear, spatiotemporal and spectral features, derived from three-axis accelerometers and gyroscope data. This study is the first attempt to investigate non-linear features related to fall risk in multi-task scenarios from a dynamic system perspective. Based on the extracted multifaceted features, we propose an ensemble elastic net (E-E-N) approach for handling imbalanced data and offering high model interpretability. The E-E-N utilizes bootstrap sampling to construct base classifiers and employs a weighting mechanism to aggregate the base classifiers. We conducted a set of validation experiments using real-world data for comprehensive comparative analysis. The results demonstrate that the E-E-N approach exhibits superior predictive performance on fall risk classification. Our proposed approach offers a cost-effective tool for accurately assessing fall risk and alleviating the burden of continuous health monitoring in the long term.

AAAI Conference 2024 Conference Paper

Stereo Vision Conversion from Planar Videos Based on Temporal Multiplane Images

  • Shanding Diao
  • Yuan Chen
  • Yang Zhao
  • Wei Jia
  • Zhao Zhang
  • Ronggang Wang

With the rapid development of 3D movie and light-field displays, there is a growing demand for stereo videos. However, generating high-quality stereo videos from planar videos remains a challenging task. Traditional depth-image-based rendering techniques struggle to effectively handle the problem of occlusion exposure, which occurs when the occluded contents become visible in other views. Recently, the single-view multiplane images (MPI) representation has shown promising performance for planar video stereoscopy. However, the MPI still lacks real details that are occluded in the current frame, resulting in blurry artifacts in occlusion exposure regions. In fact, planar videos can leverage complementary information from adjacent frames to predict a more complete scene representation for the current frame. Therefore, this paper extends the MPI from still frames to the temporal domain, introducing the temporal MPI (TMPI). By extracting complementary information from adjacent frames based on optical flow guidance, obscured regions in the current frame can be effectively repaired. Additionally, a new module called masked optical flow warping (MOFW) is introduced to improve the propagation of pixels along optical flow trajectories. Experimental results demonstrate that the proposed method can generate high-quality stereoscopic or light-field videos from a single view and reproduce better occluded details than other state-of-the-art (SOTA) methods. https://github.com/Dio3ding/TMPI

NeurIPS Conference 2023 Conference Paper

Connecting Multi-modal Contrastive Representations

  • Zehan Wang
  • Yang Zhao
  • Xize 成
  • Haifeng Huang
  • Jiageng Liu
  • Aoxiong Yin
  • Li Tang
  • Linjun Li

Multi-modal Contrastive Representation (MCR) learning aims to encode different modalities into a semantically aligned shared space. This paradigm shows remarkable generalization ability on numerous downstream tasks across various modalities. However, the reliance on massive high-quality data pairs limits its further development on more modalities. This paper proposes a novel training-efficient method for learning MCR without paired data called Connecting Multi-modal Contrastive Representations (C-MCR). Specifically, given two existing MCRs pre-trained on $(\mathcal{A}$, $\mathcal{B})$ and $(\mathcal{B}$, $\mathcal{C})$ modality pairs, we project them to a new space and use the data from the overlapping modality $\mathcal{B}$ to aligning the two MCRs in the new space. Meanwhile, since the modality pairs $(\mathcal{A}$, $\mathcal{B})$ and $(\mathcal{B}$, $\mathcal{C})$ are already aligned within each MCR, the connection learned by overlapping modality can also be transferred to non-overlapping modality pair $(\mathcal{A}$, $\mathcal{C})$. To unleash the potential of C-MCR, we further introduce a semantic-enhanced inter- and intra-MCR connection method. We first enhance the semantic consistency and completion of embeddings across different modalities for more robust alignment. Then we utilize the inter-MCR alignment to establish the connection, and employ the intra-MCR alignment to better maintain the connection for inputs from non-overlapping modalities. To demonstrate the effectiveness of C-MCR, we take the field of audio-visual and 3D-language learning as examples. Specifically, we connect CLIP and CLAP via texts to derive audio-visual representations, and integrate CLIP and ULIP via images for 3D-language representations. Remarkably, without using any paired data, C-MCR for audio-visual achieves state-of-the-art performance on audio-image retrieval, audio-visual source localization, and counterfactual audio-image recognition tasks. Furthermore, C-MCR for 3D-language also attains advanced zero-shot 3D point cloud classification accuracy on ModelNet40. Our project page is available at \url{https: //c-mcr. github. io/C-MCR/}

AAAI Conference 2023 Conference Paper

CoopInit: Initializing Generative Adversarial Networks via Cooperative Learning

  • Yang Zhao
  • Jianwen Xie
  • Ping Li

Numerous research efforts have been made to stabilize the training of the Generative Adversarial Networks (GANs), such as through regularization and architecture design. However, we identify the instability can also arise from the fragile balance at the early stage of adversarial learning. This paper proposes the CoopInit, a simple yet effective cooperative learning-based initialization strategy that can quickly learn a good starting point for GANs, with a very small computation overhead during training. The proposed algorithm consists of two learning stages: (i) Cooperative initialization stage: The discriminator of GAN is treated as an energy-based model (EBM) and is optimized via maximum likelihood estimation (MLE), with the help of the GAN's generator to provide synthetic data to approximate the learning gradients. The EBM also guides the MLE learning of the generator via MCMC teaching; (ii) Adversarial finalization stage: After a few iterations of initialization, the algorithm seamlessly transits to the regular mini-max adversarial training until convergence. The motivation is that the MLE-based initialization stage drives the model towards mode coverage, which is helpful in alleviating the issue of mode dropping during the adversarial learning stage. We demonstrate the effectiveness of the proposed approach on image generation and one-sided unpaired image-to-image translation tasks through extensive experiments.

NeurIPS Conference 2023 Conference Paper

De novo Drug Design using Reinforcement Learning with Multiple GPT Agents

  • Xiuyuan Hu
  • Guoqing Liu
  • Yang Zhao
  • Hao Zhang

De novo drug design is a pivotal issue in pharmacology and a new area of focus in AI for science research. A central challenge in this field is to generate molecules with specific properties while also producing a wide range of diverse candidates. Although advanced technologies such as transformer models and reinforcement learning have been applied in drug design, their potential has not been fully realized. Therefore, we propose MolRL-MGPT, a reinforcement learning algorithm with multiple GPT agents for drug molecular generation. To promote molecular diversity, we encourage the agents to collaborate in searching for desirable molecules in diverse directions. Our algorithm has shown promising results on the GuacaMol benchmark and exhibits efficacy in designing inhibitors against SARS-CoV-2 protein targets. The codes are available at: https: //github. com/HXYfighter/MolRL-MGPT.

EAAI Journal 2023 Journal Article

Lightning risk assessment of offshore wind farms by semi-supervised learning

  • Qibin Zhou
  • Jingjie Ye
  • Guohua Yang
  • Ruanming Huang
  • Yang Zhao
  • Yudan Gu
  • Xiaoyan Bian

The wind turbine has rapidly developed worldwide with increasing height and scale, resulting in the increased risk of lightning strikes. When wind turbines were stroke by the lightning, they will be damaged, causing economic loss and outage. Lightning risk assessment can guide the improvement of lightning protection and the design of the wind farms to efficiently prevent lightning damages. The traditional lightning risk assessment methods rely on subjective features to some extent. The existing lightning risk assessment methods based on machine learning demand abundant labeled data. It is extremely difficult to label and acquire the data. This paper proposed a lightning risk assessment method based on semi-supervised learning to address the challenges of labeling negative samples and limited labeled data. The semi-supervised K-means algorithm is proposed to divide all data into three parts. The Laplacian support vector machine (LapSVM) with hyperparameters optimized by the particle swarm optimization (PSO) is used to assess the lightning risk. The proposed method has better performance than the standard SVM and neural network (NN). Moreover, previous researches did not consider lightning protection ability. This paper introduces the receptor number into lightning risk assessment. The assessment results suggest that there is higher lightning risk in areas with a great number of wind turbines, high lightning density, and strong lightning strength. The ocean area is more likely to have low lightning risk. The results are valuable for lightning protection optimization of existing wind farms and can give guidance for the plan of new wind farms.

JBHI Journal 2023 Journal Article

Semi-Supervised Medical Image Segmentation With Voxel Stability and Reliability Constraints

  • Yang Zhao
  • Ke Lu
  • Jian Xue
  • Shuhua Wang
  • Jian Lu

Semi-supervised learning is becoming an effective solution in medical image segmentation because annotations are costly and tedious to acquire. Methods based on the teacher-student model use consistency regularization and uncertainty estimation and have shown good potential in dealing with limited annotated data. Nevertheless, the existing teacher-student model is seriously limited by the exponential moving average algorithm, which leads to the optimization trap. Moreover, the classic uncertainty estimation method calculates the global uncertainty for images but does not consider local region-level uncertainty, which is unsuitable for medical images with blurry regions. In this article, the Voxel Stability and Reliability Constraint (VSRC) model is proposed to address these issues. Specifically, the Voxel Stability Constraint (VSC) strategy is introduced to optimize parameters and exchange effective knowledge between two independent initialized models, which can break through the performance bottleneck and avoid model collapse. Moreover, a new uncertainty estimation strategy, the Voxel Reliability Constraint (VRC), is proposed for use in our semi-supervised model to consider the uncertainty at the local region level. We further extend our model to auxiliary tasks and propose a task-level consistency regularization with uncertainty estimation. Extensive experiments on two 3D medical image datasets demonstrate that our method outperforms other state-of-the-art semi-supervised medical image segmentation methods under limited supervision.

JBHI Journal 2022 Journal Article

Dual-Branch Network With Dual-Sampling Modulated Dice Loss for Hard Exudate Segmentation in Color Fundus Images

  • Qing Liu
  • Haotian Liu
  • Yang Zhao
  • Yixiong Liang

Automated segmentation of hard exudates in colour fundus images is a challenge task due to issues of extreme class imbalance and enormous size variation. This paper aims to tackle these issues and proposes a dual-branch network with dual-sampling modulated Dice loss. It consists of two branches: large hard exudate biased segmentation branch and small hard exudate biased segmentation branch. Both of them are responsible for their own duties separately. Furthermore, we propose a dual-sampling modulated Dice loss for the training such that our proposed dual-branch network is able to segment hard exudates in different sizes. In detail, for the first branch, we use a uniform sampler to sample pixels from predicted segmentation mask for Dice loss calculation, which leads to this branch naturally be biased in favour of large hard exudates as Dice loss generates larger cost on misidentification of large hard exudates than small hard exudates. For the second branch, we use a re-balanced sampler to oversample hard exudate pixels and undersample background pixels for loss calculation. In this way, cost on misidentification of small hard exudates is enlarged, which enforces the parameters in the second branch fit small hard exudates well. Considering that large hard exudates are much easier to be correctly identified than small hard exudates, we propose an easy-to-difficult learning strategy by adaptively modulating the losses of two branches. We evaluate our proposed method on two public datasets and the results demonstrate that ours achieves state-of-the-art performance.

NeurIPS Conference 2022 Conference Paper

Towards Effective Multi-Modal Interchanges in Zero-Resource Sounding Object Localization

  • Yang Zhao
  • Chen Zhang
  • Haifeng Huang
  • Haoyuan Li
  • Zhou Zhao

Aiming to locate the object that emits a specified sound in complex scenes, the task of sounding object localization bridges two perception-oriented modalities of vision and acoustics, and brings enormous research value to the comprehensive perceptual understanding of machine intelligence. Although there are massive training data collected in this field, few of them contain accurate bounding box annotations, hindering the learning process and further application of proposed models. In order to address this problem, we try to explore an effective multi-modal knowledge transfer strategy to obtain precise knowledge from other similar tasks and transfer it through well-aligned multi-modal data to deal with this task in a zero-resource manner. Concretely, we design and propose a novel \textit{Two-stream Universal Referring localization Network} (TURN), which is composed of a localization stream and an alignment stream to carry out different functions. The former is utilized to extract the knowledge related to referring object localization from the image grounding task, while the latter is devised to learn a universal semantic space shared between texts and audios. Moreover, we further develop an adaptive sampling strategy to automatically identify the overlap between different data domains, thus boosting the performance and stability of our model. The extensive experiments on various publicly-available benchmarks demonstrate that TURN can achieve competitive performance compared with the state-of-the-art approaches without using any data in this field, which verifies the feasibility of our proposed mechanisms and strategies.

ICLR Conference 2021 Conference Paper

Learning Energy-Based Generative Models via Coarse-to-Fine Expanding and Sampling

  • Yang Zhao
  • Jianwen Xie
  • Ping Li 0001

Energy-based models (EBMs) parameterized by neural networks can be trained by the Markov chain Monte Carlo (MCMC) sampling-based maximum likelihood estimation. Despite the recent significant success of EBMs in image generation, the current approaches to train EBMs are unstable and have difficulty synthesizing diverse and high-fidelity images. In this paper, we propose to train EBMs via a multistage coarse-to-fine expanding and sampling strategy, which starts with learning a coarse-level EBM from images at low resolution and then gradually transits to learn a finer-level EBM from images at higher resolution by expanding the energy function as the learning progresses. The proposed framework is computationally efficient with smooth learning and sampling. It achieves the best performance on image generation amongst all EBMs and is the first successful EBM to synthesize high-fidelity images at $512\times512$ resolution. It can also be useful for image restoration and out-of-distribution detection. Lastly, the proposed framework is further generalized to the one-sided unsupervised image-to-image translation and beats baseline methods in terms of model size and training budget. We also present a gradient-based generative saliency method to interpret the translation dynamics.

AAAI Conference 2021 Conference Paper

Synchronous Interactive Decoding for Multilingual Neural Machine Translation

  • Hao He
  • Qian Wang
  • Zhipeng Yu
  • Yang Zhao
  • Jiajun Zhang
  • Chengqing Zong

To simultaneously translate a source language into multiple different target languages is one of the most common scenarios of multilingual translation. However, existing methods cannot make full use of translation model information during decoding, such as intra-lingual and inter-lingual future information, and therefore may suffer from some issues like the unbalanced outputs. In this paper, we present a new approach for synchronous interactive multilingual neural machine translation (SimNMT), which predicts each target language output simultaneously and interactively using historical and future information of all target languages. Specifically, we first propose a synchronous cross-interactive decoder in which generation of each target output does not only depend on its generated sequences, but also relies on its future information, as well as history and future contexts of other target languages. Then, we present a new interactive multilingual beam search algorithm that enables synchronous interactive decoding of all target languages in a single model. We take two target languages as an example to illustrate and evaluate the proposed SimNMT model on IWSLT datasets. The experimental results demonstrate that our method achieves significant improvements over several advanced NMT and M- NMT models.

IROS Conference 2020 Conference Paper

A Bottom-up Framework for Construction of Structured Semantic 3D Scene Graph

  • Bangguo Yu
  • Chongyu Chen
  • Fengyu Zhou
  • Fang Wan
  • Wenmi Zhuang
  • Yang Zhao

For high-level human-robot interaction tasks, 3D scene understanding is important and non-trivial for autonomous robots. However, parsing and utilizing effective environment information of the 3D scene is not trivial due to the complexity of the 3D environment and the limited ability for reasoning about our visual world. Although there have been great efforts on semantic detection and scene analysis, the existing solutions for parsing and representation of the 3D scene still fail to preserve accurate semantic information and equip sufficient applicability. This study proposes a bottomup construction framework for structured 3D scene graph generation, which efficiently describes the objects, relations and attributes of the 3D indoor environment with structured representation. In the proposed method, we adopt visual perception to capture the semantic information and inference from scene priors to calculate the optimal parse graph. Afterwards, an improved probabilistic grammar model is used to represent the scene priors. Experiment results demonstrate that the proposed framework significantly outperforms existing methods in terms of accuracy, and a demonstration is provided to verify the applicability in applying to high-level human-robot interaction tasks. The supplementary video can be accessed at the following link: https://youtu.be/vEWNxnSwmKI.

ICLR Conference 2020 Conference Paper

Bayesian Meta Sampling for Fast Uncertainty Adaptation

  • Zhenyi Wang 0001
  • Yang Zhao
  • Ping Yu
  • Ruiyi Zhang 0002
  • Changyou Chen

Meta learning has been making impressive progress for fast model adaptation. However, limited work has been done on learning fast uncertainty adaption for Bayesian modeling. In this paper, we propose to achieve the goal by placing meta learning on the space of probability measures, inducing the concept of meta sampling for fast uncertainty adaption. Specifically, we propose a Bayesian meta sampling framework consisting of two main components: a meta sampler and a sample adapter. The meta sampler is constructed by adopting a neural-inverse-autoregressive-flow (NIAF) structure, a variant of the recently proposed neural autoregressive flows, to efficiently generate meta samples to be adapted. The sample adapter moves meta samples to task-specific samples, based on a newly proposed and general Bayesian sampling technique, called optimal-transport Bayesian sampling. The combination of the two components allows a simple learning procedure for the meta sampler to be developed, which can be efficiently optimized via standard back-propagation. Extensive experimental results demonstrate the efficiency and effectiveness of the proposed framework, obtaining better sample quality and faster uncertainty adaption compared to related methods.

EAAI Journal 2020 Journal Article

Fault diagnosis using novel AdaBoost based discriminant locality preserving projection with resamples

  • Yan-Lin He
  • Yang Zhao
  • Xiao Hu
  • Xiao-Na Yan
  • Qun-Xiong Zhu
  • Yuan Xu

Fault diagnosis plays a pivotal role in ensuring the safety of process industries. However, due to the diversity of process faults and the high coupling of fault data, it becomes very difficult to achieve high accuracy in the fault diagnosis of complex industrial processes. To address this concern, in this article, a novel AdaBoost-based discriminant locality preserving projection (DLPP) with resamples (A-DLPPR) model is proposed. The proposed A-DLPPR model has two features: to address the problem of matrix decomposition in DLPP, the bootstrap method is utilized to generate groups of resample data, and to obtain high classification accuracy, the AdaBoost-based classification technique is adopted. Finally, an effective fault diagnosis model using the proposed A-DLPPR model can be established. To validate the effectiveness of the proposed A-DLPPR model, the Tennessee Eastman process (TEP) is selected, and case studies using different kinds of TEP faults are conducted. The simulation results indicate that the proposed A-DLPPR model can achieve higher fault diagnosis accuracy than some other models, which verifies that in the field of complex industrial processes, the proposed A-DLPPR method can be used as an effective model for fault diagnosis.

ICML Conference 2020 Conference Paper

Feature Quantization Improves GAN Training

  • Yang Zhao
  • Chunyuan Li
  • Ping Yu
  • Jianfeng Gao 0001
  • Changyou Chen

The instability in GANs’ training has been a long-standing problem despite remarkable research efforts. We identify that instability issues stem from difficulties of performing feature matching with mini-batch statistics, due to a fragile balance between the fixed target distribution and the progressively generated distribution. In this work, we propose feature quantizatoin (FQ) for the discriminator, to embed both true and fake data samples into a shared discrete space. The quantized values of FQ are constructed as an evolving dictionary, which is consistent with feature statistics of the recent distribution history. Hence, FQ implicitly enables robust feature matching in a compact space. Our method can be easily plugged into existing GAN models, with little computational overhead in training. Extensive experimental results show that the proposed FQ-GAN can improve the FID scores of baseline methods by a large margin on a variety of tasks, including three representative GAN models on 10 benchmarks, achieving new state-of-the-art performance.

NeurIPS Conference 2020 Conference Paper

FracTrain: Fractionally Squeezing Bit Savings Both Temporally and Spatially for Efficient DNN Training

  • Yonggan Fu
  • Haoran You
  • Yang Zhao
  • Yue Wang
  • Chaojian Li
  • Kailash Gopalakrishnan
  • Zhangyang Wang
  • Yingyan Lin

Recent breakthroughs in deep neural networks (DNNs) have fueled a tremendous demand for intelligent edge devices featuring on-site learning, while the practical realization of such systems remains a challenge due to the limited resources available at the edge and the required massive training costs for state-of-the-art (SOTA) DNNs. As reducing precision is one of the most effective knobs for boosting training time/energy efficiency, there has been a growing interest in low-precision DNN training. In this paper, we explore from an orthogonal direction: how to fractionally squeeze out more training cost savings from the most redundant bit level, progressively along the training trajectory and dynamically per input. Specifically, we propose FracTrain that integrates (i) progressive fractional quantization which gradually increases the precision of activations, weights, and gradients that will not reach the precision of SOTA static quantized DNN training until the final training stage, and (ii) dynamic fractional quantization which assigns precisions to both the activations and gradients of each layer in an input-adaptive manner, for only "fractionally" updating layer parameters. Extensive simulations and ablation studies (six models, four datasets, and three training settings including standard, adaptation, and fine-tuning) validate the effectiveness of FracTrain in reducing computational cost and hardware-quantified energy/latency of DNN training while achieving a comparable or better (-0. 12%~+1. 87%) accuracy. For example, when training ResNet-74 on CIFAR-10, FracTrain achieves 77. 6% and 53. 5% computational cost and training latency savings, respectively, compared with the best SOTA baseline, while achieving a comparable (-0. 07%) accuracy. Our codes are available at: https: //github. com/RICE-EIC/FracTrain.

IJCAI Conference 2020 Conference Paper

Knowledge Graphs Enhanced Neural Machine Translation

  • Yang Zhao
  • Jiajun Zhang
  • Yu Zhou
  • Chengqing Zong

Knowledge graphs (KGs) store much structured information on various entities, many of which are not covered by the parallel sentence pairs of neural machine translation (NMT). To improve the translation quality of these entities, in this paper we propose a novel KGs enhanced NMT method. Specifically, we first induce the new translation results of these entities by transforming the source and target KGs into a unified semantic space. We then generate adequate pseudo parallel sentence pairs that contain these induced entity pairs. Finally, NMT model is jointly trained by the original and pseudo sentence pairs. The extensive experiments on Chinese-to-English and Englishto-Japanese translation tasks demonstrate that our method significantly outperforms the strong baseline models in translation quality, especially in handling the induced entities.

AAAI Conference 2020 Conference Paper

Learning Diverse Stochastic Human-Action Generators by Learning Smooth Latent Transitions

  • Zhenyi Wang
  • Ping Yu
  • Yang Zhao
  • Ruiyi Zhang
  • Yufan Zhou
  • Junsong Yuan
  • Changyou Chen

Human-motion generation is a long-standing challenging task due to the requirement of accurately modeling complex and diverse dynamic patterns. Most existing methods adopt sequence models such as RNN to directly model transitions in the original action space. Due to high dimensionality and potential noise, such modeling of action transitions is particularly challenging. In this paper, we focus on skeleton-based action generation and propose to model smooth and diverse transitions on a latent space of action sequences with much lower dimensionality. Conditioned on a latent sequence, actions are generated by a frame-wise decoder shared by all latent action-poses. Specifically, an implicit RNN is defined to model smooth latent sequences, whose randomness (diversity) is controlled by noise from the input. Different from standard action-prediction methods, our model can generate action sequences from pure noise without any conditional action poses. Remarkably, it can also generate unseen actions from mixed classes during training. Our model is learned with a bi-directional generative-adversarial-net framework, which can not only generate diverse action sequences of a particular class or mix classes, but also learns to classify action sequences within the same model. Experimental results show the superiority of our method in both diverse action-sequence generation and classification, relative to existing methods.

IJCAI Conference 2020 Conference Paper

Learning From Multi-Dimensional Partial Labels

  • Haobo Wang
  • Weiwei Liu
  • Yang Zhao
  • Tianlei Hu
  • Ke Chen
  • Gang Chen

Multi-dimensional classification has attracted huge attention from the community. Though most studies consider fully annotated data, in real practice obtaining fully labeled data in MDC tasks is usually intractable. In this paper, we propose a novel learning paradigm: MultiDimensional Partial Label Learning (MDPL) where the ground-truth labels of each instance are concealed in multiple candidate label sets. We first introduce the partial hamming loss for MDPL that incurs a large loss if the predicted labels are not in candidate label sets, and provide an empirical risk minimization (ERM) framework. Theoretically, we rigorously prove the conditions for ERM learnability of MDPL in both independent and dependent cases. Furthermore, we present two MDPL algorithms under our proposed ERM framework. Comprehensive experiments on both synthetic and real-world datasets validate the effectiveness of our proposals.

AAAI Conference 2020 Conference Paper

Patchy Image Structure Classification Using Multi-Orientation Region Transform

  • Xiaohan Yu
  • Yang Zhao
  • Yongsheng Gao
  • Shengwu Xiong
  • Xiaohui Yuan

Exterior contour and interior structure are both vital features for classifying objects. However, most of the existing methods consider exterior contour feature and internal structure feature separately, and thus fail to function when classifying patchy image structures that have similar contours and flexible structures. To address above limitations, this paper proposes a novel Multi-Orientation Region Transform (MORT), which can effectively characterize both contour and structure features simultaneously, for patchy image structure classification. MORT is performed over multiple orientation regions at multiple scales to effectively integrate patchy features, and thus enables a better description of the shape in a coarse-to-fine manner. Moreover, the proposed MORT can be extended to combine with the deep convolutional neural network techniques, for further enhancement of classification accuracy. Very encouraging experimental results on the challenging ultra-fine-grained cultivar recognition task, insect wing recognition task, and large variation butterfly recognition task are obtained, which demonstrate the effectiveness and superiority of the proposed MORT over the state-of-theart methods in classifying patchy image structures. Our code and three patchy image structure datasets are available at: https: //github. com/XiaohanYu-GU/MReT2019.

AIJ Journal 2020 Journal Article

Synchronous bidirectional inference for neural sequence generation

  • Jiajun Zhang
  • Long Zhou
  • Yang Zhao
  • Chengqing Zong

In sequence to sequence generation tasks (e. g. machine translation and abstractive summarization), inference is generally performed in a left-to-right manner to produce the result token by token. The neural approaches, such as LSTM and self-attention networks, are now able to make full use of all the predicted history hypotheses from left side during inference, but cannot meanwhile access any future (right side) information and usually generate unbalanced outputs (e. g. left parts are much more accurate than right ones in Chinese-English translation). In this work, we propose a synchronous bidirectional inference model to generate outputs using both left-to-right and right-to-left decoding simultaneously and interactively. First, we introduce a novel beam search algorithm that facilitates synchronous bidirectional decoding. Then, we present the core approach which enables left-to-right and right-to-left decoding to interact with each other, so as to utilize both the history and future predictions simultaneously during inference. We apply the proposed model to both LSTM and self-attention networks. Furthermore, we propose a novel fine-tuning based parameter optimization algorithm in addition to the simple two-pass strategy. The extensive experiments on machine translation and abstractive summarization demonstrate that our synchronous bidirectional inference model can achieve remarkable improvements over the strong baselines.

YNIMG Journal 2020 Journal Article

Targeting brain functions from the scalp: Transcranial brain atlas based on large-scale fMRI data synthesis

  • Yihan Jiang
  • Zheng Li
  • Yang Zhao
  • Xiang Xiao
  • Wei Zhang
  • Peipei Sun
  • Yihong Yang
  • Chaozhe Zhu

Transcranial brain mapping techniques, such as functional near-infrared spectroscopy (fNIRS) and transcranial magnetic stimulation (TMS), have been playing an increasingly important role in studies of human brain functions. Given a brain function of interest, fNIRS probes and TMS coils should be properly placed on the scalp to ensure that the function is effectively measured or modulated. However, since brain activity is inside the skull and invisible to the researcher during placement, this blind targeting may cause the device to partially or completely miss the functional target, resulting in inconsistent experimental results and divergent clinical outcomes, especially when participants’ structural MRI data are not available. To address this issue, we propose here a framework for targeting a designated function directly from the scalp. First, a functional brain atlas for the targeted brain function is constructed via a meta-analysis of large-scale functional magnetic resonance imaging datasets. Second, the functional brain atlas is presented on the scalp surface by using a transcranial mapping previously established from an structural MRI dataset (n ​= ​114), resulting in a novel functional transcranial brain atlas (fTBA). Finally, a low-cost, portable scalp-navigation system is used to localize the transcranial device on the individual’s scalp with the guidance of the fTBA. To demonstrate the feasibility of the targeting framework, both fNIRS and TMS mapping experiments were conducted. The results show that fTBA-guided fNIRS positioning can detect functional activity with high sensitivity and specificity for working memory and motor systems; Moreover, compared with traditional TMS targeting approaches (e. g. the International 10–20 System and the conventional 5-cm rule), the fTBA suggested motor stimulation site is closesr to both the motor hotspot and the center of gravity of motor evoked potentials (MEP-COG). In summary, the proposed method unblinds the transcranial function targeting process using prior information, providing an effective and straightforward approach to transcranial brain mapping studies, especially those without participants’ structural MRI data.

ICML Conference 2020 Conference Paper

Variance Reduction in Stochastic Particle-Optimization Sampling

  • Jianyi Zhang
  • Yang Zhao
  • Changyou Chen

Stochastic particle-optimization sampling (SPOS) is a recently-developed scalable Bayesian sampling framework unifying stochastic gradient MCMC (SG-MCMC) and Stein variational gradient descent (SVGD) algorithms based on Wasserstein gradient flows. With a rigorous non-asymptotic convergence theory developed, SPOS can avoid the particle-collapsing pitfall of SVGD. However, the variance-reduction effect in SPOS has not been clear. In this paper, we address this gap by presenting several variance-reduction techniques for SPOS. Specifically, we propose three variants of variance-reduced SPOS, called SAGA particle-optimization sampling (SAGA-POS), SVRG particle-optimization sampling (SVRG-POS) and a variant of SVRG-POS which avoids full gradient computations, denoted as SVRG-POS$^+$. Importantly, we provide non-asymptotic convergence guarantees for these algorithms in terms of the 2-Wasserstein metric and analyze their complexities. The results show our algorithms yield better convergence rates than existing variance-reduced variants of stochastic Langevin dynamics, though more space is required to store the particles in training. Our theory aligns well with experimental results on both synthetic and real datasets.

AAAI Conference 2019 Conference Paper

Addressing the Under-Translation Problem from the Entropy Perspective

  • Yang Zhao
  • Jiajun Zhang
  • Chengqing Zong
  • Zhongjun He
  • Hua Wu

Neural Machine Translation (NMT) has drawn much attention due to its promising translation performance in recent years. However, the under-translation problem still remains a big challenge. In this paper, we focus on the under-translation problem and attempt to find out what kinds of source words are more likely to be ignored. Through analysis, we observe that a source word with a large translation entropy is more inclined to be dropped. To address this problem, we propose a coarse-to-fine framework. In coarse-grained phase, we introduce a simple strategy to reduce the entropy of highentropy words through constructing the pseudo target sentences. In fine-grained phase, we propose three methods, including pre-training method, multitask method and two-pass method, to encourage the neural model to correctly translate these high-entropy words. Experimental results on various translation tasks show that our method can significantly improve the translation quality and substantially reduce the under-translation cases of high-entropy words.

IJCAI Conference 2019 Conference Paper

Discriminative and Correlative Partial Multi-Label Learning

  • Haobo Wang
  • Weiwei Liu
  • Yang Zhao
  • Chen Zhang
  • Tianlei Hu
  • Gang Chen

In partial label learning (PML), each instance is associated with a candidate label set that contains multiple relevant labels and other false positive labels. The most challenging issue for the PML is that the training procedure is prone to be affected by the labeling noise. We observe that state-of-the-art PML methods are either powerless to disambiguate the correct labels from the candidate labels or incapable of extracting the label correlations sufficiently. To fill this gap, a two-stage DiscRiminative and correlAtive partial Multi-label leArning (DRAMA) algorithm is presented in this work. In the first stage, a confidence value is learned for each label by utilizing the feature manifold, which indicates how likely a label is correct. In the second stage, a gradient boosting model is induced to fit the label confidences. Specifically, to explore the label correlations, we augment the feature space by the previously elicited labels on each boosting round. Extensive experiments on various real-world datasets clearly validate the superiority of our proposed method.

NeurIPS Conference 2019 Conference Paper

E2-Train: Training State-of-the-art CNNs with Over 80% Energy Savings

  • Yue Wang
  • Ziyu Jiang
  • Xiaohan Chen
  • Pengfei Xu
  • Yang Zhao
  • Yingyan Lin
  • Zhangyang Wang

Convolutional neural networks (CNNs) have been increasingly deployed to edge devices. Hence, many efforts have been made towards efficient CNN inference on resource-constrained platforms. This paper attempts to explore an orthogonal direction: how to conduct more energy-efficient training of CNNs, so as to enable on-device training? We strive to reduce the energy cost during training, by dropping unnecessary computations, from three complementary levels: stochastic mini-batch dropping on the data level; selective layer update on the model level; and sign prediction for low-cost, low-precision back-propagation, on the algorithm level. Extensive simulations and ablation studies, with real energy measurements from an FPGA board, confirm the superiority of our proposed strategies and demonstrate remarkable energy savings for training. For example, when training ResNet-74 on CIFAR-10, we achieve aggressive energy savings of >90% and >60%, while incurring a top-1 accuracy loss of only about 2% and 1. 2%, respectively. When training ResNet-110 on CIFAR-100, an over 84% training energy saving is achieved without degrading inference accuracy.

AAAI Conference 2019 Conference Paper

Self-Adversarially Learned Bayesian Sampling

  • Yang Zhao
  • Jianyi Zhang
  • Changyou Chen

Scalable Bayesian sampling is playing an important role in modern machine learning, especially in the fast-developed unsupervised-(deep)-learning models. While tremendous progresses have been achieved via scalable Bayesian sampling such as stochastic gradient MCMC (SG-MCMC) and Stein variational gradient descent (SVGD), the generated samples are typically highly correlated. Moreover, their sample-generation processes are often criticized to be inefficient. In this paper, we propose a novel self-adversarial learning framework that automatically learns a conditional generator to mimic the behavior of a Markov kernel (transition kernel). High-quality samples can be efficiently generated by direct forward passes though a learned generator. Most importantly, the learning process adopts a self-learning paradigm, requiring no information on existing Markov kernels, e. g. , knowledge of how to draw samples from them. Specifically, our framework learns to use current samples, either from the generator or pre-provided training data, to update the generator such that the generated samples progressively approach a target distribution, thus it is called self-learning. Experiments on both synthetic and real datasets verify advantages of our framework, outperforming related methods in terms of both sampling efficiency and sample quality.

AAAI Conference 2018 Conference Paper

An Euclidean Distance Based on Tensor Product Graph Diffusion Related Attribute Value Embedding for Nominal Data Clustering

  • Lei Gu
  • Ningning Zhou
  • Yang Zhao

Not like numerical data clustering, nominal data clustering is a very difficult problem because there exists no natural relative ordering between nominal attribute values. This paper mainly aims to make the Euclidean distance measure appropriate to nominal data clustering, and the core idea is the attribute value embedding, namely, transforming each nominal attribute value into a numerical vector. This embedding method consists of four steps. In the first step, the weights, which can quantify the amount of information in attribute values, is calculated for each value in each nominal attribute based on each object and its k nearest neighbors. In the second step, an intra-attribute value similarity matrix is created for each nominal attribute by using the attribute value’s weights. In the third step, for each nominal attribute, we find another attribute with the maximal dependence on it, and build an inter-attribute value similarity matrix on the basis of the attribute value’s weights related to these two attributes. In the last step, a diffusion matrix of each nominal attribute is constructed by the tensor product graph diffusion process, and this step can cause the acquired value embedding to contain simultaneously the intra- and inter-attribute value similarities information. To evaluate the effectiveness of our proposed method, experiments are done on 10 data sets. Experimental results demonstrate that our method not only enables the Euclidean distance to be used for nominal data clustering, but also can acquire the better clustering performance than several existing state-of-the-art approaches.

IJCAI Conference 2018 Conference Paper

Phrase Table as Recommendation Memory for Neural Machine Translation

  • Yang Zhao
  • Yining Wang
  • Jiajun Zhang
  • Chengqing Zong

Neural Machine Translation (NMT) has drawn much attention due to its promising translation performance recently. However, several studies indicate that NMT often generates fluent but unfaithful translations. In this paper, we propose a method to alleviate this problem by using a phrase table as recommendation memory. The main idea is to add bonus to words worthy of recommendation, so that NMT can make correct predictions. Specifically, we first derive a prefix tree to accommodate all the candidate target phrases by searching the phrase translation table according to the source sentence. Then, we construct a recommendation word set by matching between candidate target phrases and previously translated target words by NMT. After that, we determine the specific bonus value for each recommendable word by using the attention vector and phrase translation probability. Finally, we integrate this bonus value into NMT to improve the translation results. The extensive experiments demonstrate that the proposed methods obtain remarkable improvements over the strong attention based NMT.

v2026.09.13