Arrow Research search

Author name cluster

Lei Zhao

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

48 papers
2 author rows

Possible papers

48

EAAI Journal 2026 Journal Article

Dual-channel machine learning proxy for pseudo-two-dimensional model with enhanced extrapolation correction

  • Yue Cui
  • Yaxuan Wang
  • Shilong Guo
  • Liang Deng
  • Junfu Li
  • Lei Zhao
  • Zhenbo Wang

Physics-based electrochemical models are essential for analyzing and predicting the performance of lithium metal batteries (LMBs), yet their high computational cost restricts their use in real-time applications. To overcome this limitation, a dual-channel agent model (DCAM) is proposed, which decouples the mapping between electrochemical parameters and discharge duration and voltage profile, serving as an efficient surrogate for the pseudo-two-dimensional (P2D) model. Unlike physics-constrained or purely data-driven approaches, DCAM does not rely on explicit physical equations but rather bypasses regions with poor generalization, enabling fast and accurate extrapolation from partial discharge data. Furthermore, a hybrid modeling framework is developed by embedding a multilayer perceptron (MLP) into the Butler–Volmer (B–V) kinetics, replacing the iterative Newton process and thereby improving computational efficiency. Experimental and simulation results demonstrate that the proposed framework accurately reproduces the P2D voltage behavior and achieves high-fidelity extrapolation and robustness under limited-data conditions.

AAAI Conference 2026 Conference Paper

Forgetting by Pruning: Data Deletion in Join Cardinality Estimation

  • Chaowei He
  • Yuanjun Liu
  • Qingzhi Ma
  • Shenyuan Ren
  • Xizhao Luo
  • Lei Zhao
  • An Liu

Machine unlearning in learned cardinality estimation (CE) systems presents unique challenges due to the complex distributional dependencies in multi-table relational data. Specifically, data deletion, a core component of machine unlearning, faces three critical challenges in learned CE models: attribute-level sensitivity, inter-table propagation and domain disappearance leading to severe overestimation in multi-way joins. We propose Cardinality Estimation Pruning (CEP), the first unlearning framework specifically designed for multi-table learned CE systems. CEP introduces Distribution Sensitivity Pruning, which constructs semi-join deletion results and computes sensitivity scores to guide parameter pruning, and Domain Pruning, which removes support for value domains entirely eliminated by deletion. We evaluate CEP on state-of-the-art architectures NeuroCard and FACE across IMDb and TPC-H datasets. Results demonstrate CEP consistently achieves the lowest Q-error in multi-table scenarios, particularly under high deletion ratios, often outperforming full retraining. Furthermore, CEP significantly reduces convergence iterations, incurring negligible computational overhead of 0.3%-2.5% of fine-tuning time.

AAAI Conference 2026 Conference Paper

Inpaint-Anywhere: Zero-Shot Multi-Identity Inpainting with Efficient Diffusion Transformer

  • Junsheng Luan
  • Lei Zhao
  • Wei Xing

Subject-driven generation, which aims to synthesize visual content for a given identity V* with specific attributes, has garnered increasing attention in recent years. While existing methods demonstrate impressive identity consistency for both single and multiple identities, they often lack user-specified spatial control. Recent approaches, such as OminiControl-2 and EasyControl, enable inpainting conditioned on a single identity but fall short in multi-identity scenarios. In this paper, we introduce BoundID, a dataset synthesis pipeline for generating multi-identity images with bounding box annotations, and introduce Inpaint-Anywhere, a diffusion transformer framework for multi-identity inpainting. Given multiple identity references and corresponding masks, our method simultaneously generates all desired identities at precise locations while achieving both high identity and prompt fidelity. Extensive experiments show that Inpaint-Anywhere achieves state-of-the-art performance in multi-identity inpainting.

AAAI Conference 2026 Conference Paper

MPA: Multimodal Prototype Augmentation for Few-Shot Learning

  • Liwen Wu
  • Wei Wang
  • Lei Zhao
  • Zhan Gao
  • Qika Lin
  • Shaowen Yao
  • Zuozhu Liu
  • Bin Pu

Recently, Few-shot Learning (FSL) has become a popular task that aims to recognize new classes from only a few labeled examples and has been widely applied in fields such as natural science, remote sensing, and medical images. However, most existing methods focus only on the visual modality and compute prototypes directly from raw support images, which lack comprehensive and rich multimodal information. To address these limitations, we propose a novel Multimodal Prototype Augmentation FSL framework called MPA, including LLM-based Multi-Variant Semantic Enhancement (LMSE), Hierarchical Multi-View Augmentation (HMA), and an Adaptive Uncertain Class Absorber (AUCA). LMSE leverages large language models to generate diverse paraphrased category descriptions, enriching the support set with additional semantic cues. HMA exploits both natural and multi-view augmentations to enhance feature diversity (e.g., changes in viewing distance, camera angles, and lighting conditions). AUCA models uncertainty by introducing uncertain classes via interpolation and Gaussian sampling, effectively absorbing uncertain samples. Extensive experiments on four single-domain and six cross-domain FSL benchmarks demonstrate that MPA achieves superior performance compared to existing state-of-the-art methods across most settings. Notably, MPA surpasses the second-best method by 12.29% and 24.56% in the single-domain and cross-domain setting, respectively, in the 5-way 1-shot setting.

AAAI Conference 2026 Conference Paper

Organ-Aware Routing Mixture-of-Retrieval Augmented Generation for Fetal Ultrasound Reporting

  • Bin Pu
  • Siyu Wang
  • Rongbin Li
  • Xinpeng Ding
  • Lei Zhao
  • Chaoqi Chen
  • Shengli Li
  • Kenli Li

Fetal ultrasound screening is a uniquely complex diagnostic task involving the simultaneous assessment of multiple fetal organs—each with its own anatomical and clinical context—within a single examination. Automating report generation for such cases poses a significant challenge: unlike existing methods that focus on single-organ radiology tasks (e.g., chest X-rays), fetal ultrasound requires reasoning over a structured, multiple-to-multiple setting, i.e., multi-organ images corresponding to a multi-section report. In this paper, we introduce FetusR, the first large-scale dataset for multi-organ fetal ultrasound reporting, containing 15,594 real-world cases with rich organ-wise annotations. To address the intrinsic image-report alignment, we propose Organ-Aware Routing Mixture-of-Retrieval Augmented Generation (ORM-RAG) inspired by the Mixture-of-Experts paradigm. Our method decomposes the complex alignment problem into multiple one-to-one sub-retrieval tasks. Specifically, ORM-RAG integrates (1) an organ-aware mixture-of-retrieval module that partitions the retrieval space into organ-specific corpora for independent retrieval, and (2) a dynamic routing mechanism that selectively aggregates high-confidence organ-specific reports while filtering uncertain ones. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art baselines across both textual similarity and clinical accuracy metrics. Our work opens a new direction for long-form, structured report generation in real-world, multi-organ medical imaging scenarios.

EAAI Journal 2026 Journal Article

Research on transparent working face based on dynamic interpretation technology of three-dimensional seismic data

  • Wei Li
  • Lei Zhao
  • Zaibin Liu
  • Wenming Liu
  • Bo Li
  • Junsheng Yan
  • Huahui Wang

As coal mining extends to greater depths, accurately detecting coal seam floor undulations, identifying coal thickness variations, and recognizing complex geological features such as collapse columns has become increasingly essential. These challenges raise higher demands for safety and efficiency in mining operations. This study proposes a dynamic interpretation method for transparent mining faces based on three-dimensional (3D) seismic data to enhance the accuracy of detecting coal seam geological structures. The method comprehensively applies target processing and dynamic interpretation to conduct an accurate analysis of the coal seam floor elevation and average velocity field. The data model is dynamically updated by integrating surface drilling and underground roadway data, which significantly enhances model accuracy and reliability. The study shows that the improved data correction method significantly enhances model accuracy and reliability. The accuracy of channel wave detection in structural prediction reaches 75%, while maintaining the maximum absolute error for floor profile prediction within 1. 33 m. The random forest model, a machine learning approach, was improved by combining gray correlation and Particle Swarm Optimization (PSO) algorithms, further revealing the complex relationship between coal seam floor elevation and Two-Way Travel Time (TWTT). The method proposed not only enhances the precision and efficiency of transparent face detection using Artificial Intelligence (AI) techniques but also offers reliable geological support for safe coal mining operations. As geological data are continuously and dynamically updated, the method enables real-time optimization of mining decisions and reduces the risk of geological hazards.

AAAI Conference 2026 Conference Paper

Topology-Inspired Backward-Free Framework for Test-Time Adaptation in Medical Detection

  • Bin Pu
  • Xingguo Lv
  • Jiewen Yang
  • Kai Xu
  • Lei Zhao
  • Zuozhu Liu
  • Kenli Li

Recently, Test-Time Adaptation (TTA) has gained increasing attention in medical imaging due to its ability to improve model generalization under domain shifts without retraining. In particular, directly applying a well-trained model across various medical centers faces significant performance degradation caused by variations in equipment, operators, imaging conditions, and scanning skill levels of sonographers. Existing TTA methods either rely on parameter adaptation that increases computational cost or apply simple prediction fusion that ignores anatomical structure knowledge. To address these limitations, we propose a novel backward-free Topology-aware TTA framework named T^3 that integrates Structural Perception Modeling (SPM) and Box Regression Adaptation (BRA). SPM is implemented through an organ space heatmap generated via Gaussian kernel superposition. This heatmap encodes anatomical topology without requiring additional training or source data. BRA further improves localization and classification by fusing detection outputs based on the contribution of detected results to anatomically meaningful peak points from the heatmaps. Extensive experiments were conducted across six cross-domain scenarios, and the results demonstrate that our method achieves state-of-the-art cross-domain detection performance while maintaining high efficiency, offering a practical and robust solution for real-world medical diagnostic applications.

AAAI Conference 2026 Conference Paper

Wavelet Enhanced Adaptive Frequency Filter for Sequential Recommendation

  • Huayang Xu
  • Huanhuan Yuan
  • Guanfeng Liu
  • Junhua Fang
  • Lei Zhao
  • Pengpeng Zhao

Sequential recommendation has garnered significant attention for its ability to capture dynamic preferences by mining users’ historical interaction data. Given that users’ complex and intertwined periodic preferences are difficult to disentangle in the time domain, recent research is exploring frequency domain analysis to identify these hidden patterns. However, current frequency-domain-based methods suffer from two key limitations: (i) They primarily employ static filters with fixed characteristics, overlooking the personalized nature of behavioral patterns; (ii) While the global discrete Fourier transform excels at modeling long-range dependencies, it can blur non-stationary signals and short-term fluctuations. To overcome these limitations, we propose a novel method called Wavelet Enhanced Adaptive Frequency Filter for Sequential Recommendation (WEARec). Specifically, it consists of two vital modules: dynamic frequency-domain filtering and wavelet feature enhancement. The former is used to dynamically adjust filtering operations based on behavioral sequences to extract personalized global information, and the latter integrates wavelet transform to reconstruct sequences, enhancing blurred non-stationary signals and short-term fluctuations. Finally, these two modules work synergistically to achieve comprehensive performance and efficiency optimization in long sequential recommendation scenarios. Extensive experiments on four widely-used benchmark datasets demonstrate the superiority of WEARec.

EAAI Journal 2025 Journal Article

Adaptive detection method for driver fatigue using facial multisource dynamic behavior fusion

  • Guoxin Zhang
  • Fei Yang
  • Xin Fang
  • Lili Wang
  • Lei Zhao
  • Chaoning Yu

Driving while fatigued is a leading cause of traffic accidents. This study proposed an adaptive detection model to recognize driver fatigue based on the dynamic facial behavior information of drivers. First, drivers’ facial fatigue features were extracted to establish a general feature space, including pupil movement, eye state, and fatigue expression parameters. A differentiated feature space was then built based on individual drivers, taking into account the homogeneity, regularity, and individual variances in drivers' facial behavior at various states. A complete adaptive fatigue feature space was built by integrating the general feature space and differentiated feature space. Finally, a driver adaptive fatigue discrimination model was constructed to classify the general and adaptive fatigue feature space to detect driver fatigue states adaptively. A driver fatigue detection dataset from real scenarios had been established to validate the performance of the proposed model. Experimental results demonstrated that the proposed method significantly improved the detection accuracy of driver fatigue. In terms of artificial intelligence, this study contributes a novel adaptive feature space construction method based on multimodal dynamic feature fusion for facial fatigue recognition; in engineering application, it develops an adaptive driver fatigue detection system grounded in multimodal dynamic behaviors, which provides real-time alerts upon detecting driver fatigue and ensures driving safety.

AAAI Conference 2025 Conference Paper

Anatomical Knowledge Mining and Matching for Semi-supervised Medical Multi-structure Detection

  • Bin Pu
  • Liwen Wang
  • Jiewen Yang
  • Xingbo Dong
  • Benteng Ma
  • Zhuangzhuang Chen
  • Lei Zhao
  • Shengli Li

In medical image analysis, detecting multiple structures is crucial for evaluations and diagnosis but is often limited by the lack of high-quality annotations. Semi-supervised object detection emerges as a potent methodology to enhance model performance and generalization by leveraging a vast pool of unlabeled data alongside a minimal set of labeled data. A striking observation is that both unlabelled and labeled medical images contain a priori anatomical knowledge from human screening. In this work, we introduce a novel semi-supervised approach named Semi-akmm for mining and matching anatomical knowledge in ultrasound images. We develop an Adaptive Prior Knowledge Transfer (APKT) module to mine and explore the distribution and knowledge of potential proposal boxes by proposal proportion constraint. Furthermore, within a teacher-student learning framework, we put forward an Anatomical Structure Matching (ASM) module to facilitate co-learning consistent topological prior knowledge between the student and teacher models. To our knowledge, this marks the inception of an efficient semi-supervised medical multi-structure detection model. Our experiments across five publicly available ultrasound datasets demonstrate that Semi-akmm sets a new benchmark in performance with solid results that outperform existing methods.

AAAI Conference 2025 Conference Paper

Cascaded Diffusion Models for Virtual Try-On: Improving Control and Resolution

  • Guangyuan Li
  • Yongkang Wang
  • Junsheng Luan
  • Lei Zhao
  • Wei Xing
  • Huaizhong Lin
  • Binkai Ou

Previous virtual try-on methods have employed ControlNet architecture in exemplar-based inpainting diffusion models to guide the generation of try-on images, preserving the garment's features and enhancing the realism of the generated images. While these methods have maintained the identity of the garment and improved the naturalness of the generated images, they still face the following limitations: (1) For garments with complex features, such as intricate text, patterns, and uncommon styles, they struggle to retain these detailed features in the generated try-on images. (2) They are limited to generating try-on images at a maximum resolution of 1K, which may not meet the demands of real-world scenarios, where higher resolutions might be required. To address the aforementioned issues, in this paper, we propose a Cascaded Diffusion Model for virtual try-on to enhance both image controllability and resolution. We call it CDM-VTON. Specifically, we design two diffusion models: the Multi-Conditioned Diffusion Model (MC-DM) and the Super-Resolution Diffusion Model (SR-DM). The former generates low-resolution try-on images while preserving the garment's complex features, and the latter enhances the resolution of these images. Additionally, we incorporate a multi-control integration module in the MC-DM, which injects multiple control conditions into a frozen denoising U-Net to ensure that the generated try-on images retain complex garment features. Our experimental results demonstrate that our method outperforms previous approaches in preserving garment details and generating authentic virtual try-on images, both qualitatively and quantitatively.

TMLR Journal 2025 Journal Article

Federated Learning with Efficient Local Adaptation for Realized Volatility Prediction

  • Lei Zhao
  • Lin Cai
  • Wu-Sheng Lu

Financial markets present unique challenges for Federated Learning (FL) due to fragmented datasets, dynamic participation, and the critical need for precise and reliable predictions. Isolated local datasets often fail to capture the full spectrum of market dynamics, blocking accurate realized volatility predictions. Unlike traditional FL methods that focus on improving convergence during the training process, we propose Federated Learning with Adaptive Robustness and Efficiency for Local Adaptation (FLARE-LA), a novel framework designed to optimize predictive performance after the global training phase. FLARE-LA leverages Taylor-based local linearization and probabilistic optimization to efficiently adapt global models to local data distributions, enabling fast responsiveness to new market conditions. This adaptability ensures trained local models align with real-world scenarios, making FLARE-LA particularly suited to dynamic financial applications. Extensive experimental evaluations demonstrate FLARE-LA's superior performance, showcasing its ability to significantly enhance post-FL outcomes compared to state-of-the-art FL algorithms. The results underscore FLARE-LA's unique capability to drive advancements in financial forecasting and other high-stakes, rapidly evolving domains.

IJCAI Conference 2025 Conference Paper

GPL4SRec: Graph Multi-Level Aware Prompt Learning for Streaming Recommendation

  • Hao Cang
  • Huanhuan Yuan
  • Jiaqing Fan
  • Lei Zhao
  • Guanfeng Liu
  • Pengpeng Zhao

Streaming Recommendation (SRec) aims to capture evolving user preferences in the streaming scenarios. Recently, Graph Prompt Learning (GPL) methods have demonstrated their effectiveness and adaptability within SRec. However, existing graph prompt solutions rarely consider the evolution of multi-hop cascading relationships between users and items, which are crucial for modeling the shifts in user preferences. To address this problem, we propose a novel Graph Multi-Level Aware Prompt Learning for Streaming Recommendation, named GPL4SRec. Specifically, a graph encoder is first pre-trained on extensive historical data to capture user long-term preferences. Then, we design three types of prompts, namely node-aware, structure-aware, and layer-aware prompts, which are used to guide the pre-trained encoder to better capture user short-term preferences. This is accomplished by accounting for both the incremental changes in users and items, as well as the cascading evolution in multi-hop relationships. Furthermore, we provide a theoretical analysis showing that our prompt templates are critical to achieving superior performance. Finally, experimental results also prove that our model significantly outperforms the state-of-the-art approaches in SRec.

EAAI Journal 2025 Journal Article

Personalized text-to-image generation with Large Language and Vision Assistant enhanced training

  • Junsheng Luan
  • Zhanjie Zhang
  • Wei Xing
  • Lei Zhao

Personalized image generation aims to synthesize images of a specific identity. The identity, denoted as V ∗, refers to an entity with distinctive visual attributes, such as a dog-shaped backpack. However, existing methods like DreamBooth and Custom Diffusion often struggle to generate images of V ∗ that accurately match the input prompts. In this work, we analyze two key issues underlying this limitation: (1) the overbinding problem, where the prompt tokens used to represent V ∗ unintentionally bind to irrelevant visual details from the reference image during training; and (2) the low language prior problem, where insufficient use of pre-trained language prior limits the model’s ability to faithfully generate all the prompt words. To overcome these challenges, we propose LLaVA-Booth, a novel personalization method for diverse, identity-preserving image generation, based on Large Language and Vision Assistant (LLaVA) enhanced training. Our method alleviates the overbinding problem by disentangling background information and solves the low language prior problem by enriching the language context. Additionally, we introduce two auxiliary objectives: (1) an identity (ID) binding loss to strengthen the identity binding and (2) a prior preservation loss to prevent language drift and encourage generation diversity. Experiments demonstrate that LLaVA-Booth effectively mitigates overbinding and enhances language priors to improve prompt fidelity, then generates diverse, high-quality, and identity-preserving images of V ∗.

TAAS Journal 2025 Journal Article

Privacy-Preserving Group-by-Aggregation Queries for Data Federation under V2X environment

  • Zicheng Cao
  • Guanfeng Liu
  • Qingzhi Ma
  • Wei Chen
  • Lei Zhao
  • An Liu

Vehicle-to-everything (V2X) technology enables vehicles to communicate with each other, infrastructure, and the cloud, facilitating intelligent traffic management and vehicle interconnection. However, the data generated by vehicles raises concerns regarding personal privacy and corporate interests. With the rapid development of V2X technology, data security issues are becoming increasingly prominent. Data federation, as an emerging data-sharing model, utilizes secure multi-party computation techniques to enable collaboration among data owners without disclosing raw data, offering a new approach to addressing privacy and security concerns in the data exchange process of V2X. This paper proposes a group-by-aggregation query algorithm for data federation, aiming to protect personal privacy data while facilitating effective data sharing and analysis. The algorithm reverses the traditional group-by-aggregation queries process by not transmitting grouping results but rather passing encrypted aggregated attribute values to relevant data owners. By leveraging encryption algorithms with additive homomorphic or order-preserving properties to encrypt the aggregated attribute values, the algorithm ensures the correctness of mathematical operations performed under encryption, such as addition and comparison operations. Finally, the effectiveness and practicality of the algorithm are validated through experimental evaluations.

NeurIPS Conference 2025 Conference Paper

StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs

  • Qijun Luo
  • Mengqi Li
  • Lei Zhao
  • Xiao Li

Training language models on long sequence data is a demanding requirement for enhancing the model's capability on complex tasks, e. g. , long-chain reasoning. However, as the sequence length scales up, the memory cost for storing activation values becomes huge during the Backpropagation (BP) process, even with the application of gradient checkpointing technique. To tackle this challenge, we propose a *memory-efficient* and *exact* BP method called **StreamBP**, which performs a linear decomposition of the chain rule along the sequence dimension in a layer-wise manner, significantly reducing the memory cost of activation values and logits. The proposed method is applicable to common objectives such as SFT, GRPO, and DPO. From an implementation perspective, StreamBP achieves less computational FLOPs and faster BP speed by leveraging the causal structure of the language model. Compared to gradient checkpointing, StreamBP scales up the maximum sequence length of BP by $2. 8-5. 5 \times$ larger, while using comparable or even less BP time. Note that StreamBP's sequence length scaling ability can be directly transferred to batch size scaling for accelerating training. We further develop a communication-efficient distributed StreamBP to effectively support multi-GPU training and broaden its applicability. Our code can be easily integrated into the training pipeline of any transformer models and is available at https: //github. com/Ledzy/StreamBP.

JBHI Journal 2025 Journal Article

TKR-FSOD: Fetal Anatomical Structure Few-Shot Detection Utilizing Topological Knowledge Reasoning

  • Xi Li
  • Ying Tan
  • Bocheng Liang
  • Bin Pu
  • Jiewen Yang
  • Lei Zhao
  • Yanqing Kong
  • Lixian Yang

Fetal multi-anatomical structure detection in ultrasound (US) images can clearly present the relationship and influence between anatomical structures, providing more comprehensive information about fetal organ structures and assisting sonographers in making more accurate diagnoses, widely used in structure evaluation. Recently, deep learning methods have shown superior performance in detecting various anatomical structures in ultrasound images, but still have the potential for performance improvement in categories where it is difficult to obtain samples, such as rare diseases. Few-shot learning has attracted a lot of attention in medical image analysis due to its ability to solve the problem of data scarcity. However, existing few-shot learning research in medical image analysis focuses on classification and segmentation, and the research on object detection has been neglected. In this paper, we propose a novel fetal anatomical structure few-shot detection method in ultrasound images, TKR-FSOD, which learns topological knowledge through a Topological Knowledge Reasoning Module to help the model reason about and detect anatomical structures. Furthermore, we propose a Discriminate Ability Enhanced Feature Learning Module that extracts abundant discriminative features to enhance the model's discriminative ability. Experimental results demonstrate that our method outperforms the state-of-the-art baseline methods, exceeding the second-best method with a maximum margin of 4. 8% on 5-shot of split 1 under four-chamber cardiac view.

EAAI Journal 2025 Journal Article

VectorSketcher: Learning to create a vector-based free-hand sketch

  • Zhanjie Zhang
  • Quanwei Zhang
  • Junsheng Luan
  • Mengyuan Yang
  • Yun Wang
  • Lei Zhao

Sketch synthesis refers to converting a given content image into a sketch. Existing sketch synthesis methods are generally divided into pixel-based and vector-based sketch synthesis. Pixel-based sketch synthesis methods always introduce obvious artifacts and disharmonious patterns. Besides, they cannot support generating sketches with different levels of abstraction. The vector-based sketch synthesis methods have limitations in describing the structure and semantics of the content image. To tackle these problems, we propose a novel framework called VectorSketcher, which can create vectorized sketches that accurately describe the structure and semantics of the content images without introducing obvious artifacts and disharmonious patterns. Specifically, we proposed a Multi-scale Feature-based Stroke Initialization (MFSI) to speed up the optimization and essential visual details of the given image. We introduce a Controllable Score Distillation Sampling loss (CSDS) to further learn the content image’s detail. Extensive quantitative and qualitative experiments show that VectorSketcher can generate more accurate vector-based sketches than existing state-of-the-art (SOTA) sketch synthesis methods.

AAAI Conference 2024 Conference Paper

ArtBank: Artistic Style Transfer with Pre-trained Diffusion Model and Implicit Style Prompt Bank

  • Zhanjie Zhang
  • Quanwei Zhang
  • Wei Xing
  • Guangyuan Li
  • Lei Zhao
  • Jiakai Sun
  • Zehua Lan
  • Junsheng Luan

Artistic style transfer aims to repaint the content image with the learned artistic style. Existing artistic style transfer methods can be divided into two categories: small model-based approaches and pre-trained large-scale model-based approaches. Small model-based approaches can preserve the content strucuture, but fail to produce highly realistic stylized images and introduce artifacts and disharmonious patterns; Pre-trained large-scale model-based approaches can generate highly realistic stylized images but struggle with preserving the content structure. To address the above issues, we propose ArtBank, a novel artistic style transfer framework, to generate highly realistic stylized images while preserving the content structure of the content images. Specifically, to sufficiently dig out the knowledge embedded in pre-trained large-scale models, an Implicit Style Prompt Bank (ISPB), a set of trainable parameter matrices, is designed to learn and store knowledge from the collection of artworks and behave as a visual prompt to guide pre-trained large-scale models to generate highly realistic stylized images while preserving content structure. Besides, to accelerate training the above ISPB, we propose a novel Spatial-Statistical-based self-Attention Module (SSAM). The qualitative and quantitative experiments demonstrate the superiority of our proposed method over state-of-the-art artistic style transfer methods. Code is available at https://github.com/Jamie-Cheung/ArtBank.

AAAI Conference 2024 Conference Paper

Attack Deterministic Conditional Image Generative Models for Diverse and Controllable Generation

  • Tianyi Chu
  • Wei Xing
  • Jiafu Chen
  • Zhizhong Wang
  • Jiakai Sun
  • Lei Zhao
  • Haibo Chen
  • Huaizhong Lin

Existing generative adversarial network (GAN) based conditional image generative models typically produce fixed output for the same conditional input, which is unreasonable for highly subjective tasks, such as large-mask image inpainting or style transfer. On the other hand, GAN-based diverse image generative methods require retraining/fine-tuning the network or designing complex noise injection functions, which is computationally expensive, task-specific, or struggle to generate high-quality results. Given that many deterministic conditional image generative models have been able to produce high-quality yet fixed results, we raise an intriguing question: is it possible for pre-trained deterministic conditional image generative models to generate diverse results without changing network structures or parameters? To answer this question, we re-examine the conditional image generation tasks from the perspective of adversarial attack and propose a simple and efficient plug-in projected gradient descent (PGD) like method for diverse and controllable image generation. The key idea is attacking the pre-trained deterministic generative models by adding a micro perturbation to the input condition. In this way, diverse results can be generated without any adjustment of network structures or fine-tuning of the pre-trained models. In addition, we can also control the diverse results to be generated by specifying the attack direction according to a reference text or image. Our work opens the door to applying adversarial attack to low-level vision tasks, and experiments on various conditional image generation tasks demonstrate the effectiveness and superiority of the proposed method.

NeurIPS Conference 2024 Conference Paper

CogVLM: Visual Expert for Pretrained Language Models

  • Weihan Wang
  • Qingsong Lv
  • Wenmeng Yu
  • Wenyi Hong
  • Ji Qi
  • Yan Wang
  • Junhui Ji
  • Zhuoyi Yang

We introduce CogVLM, a powerful open-source visual language foundation model. Different from the popular \emph{shallow alignment} method which maps image features into the input space of language model, CogVLM bridges the gap between the frozen pretrained language model and image encoder by a trainable visual expert module in the attention and FFN layers. As a result, CogVLM enables a deep fusion of vision language features without sacrificing any performance on NLP tasks. CogVLM-17B achieves state-of-the-art performance on 17 classic cross-modal benchmarks, including 1) image captioning datasets: NoCaps, Flicker30k, 2) VQA datasets: OKVQA, TextVQA, OCRVQA, ScienceQA, 3) LVLM benchmarks: MM-Vet, MMBench, SEED-Bench, LLaVABench, POPE, MMMU, MathVista, 4) visual grounding datasets: RefCOCO, RefCOCO+, RefCOCOg, Visual7W. Codes and checkpoints are available at Github.

JBHI Journal 2024 Journal Article

FARN: Fetal Anatomy Reasoning Network for Detection With Global Context Semantic and Local Topology Relationship

  • Lei Zhao
  • Guanghua Tan
  • Qianghui Wu
  • Bin Pu
  • Hongliang Ren
  • Shengli Li
  • Kenli Li

Accurate recognition of fetal anatomical structure is a pivotal task in ultrasound (US) image analysis. Sonographers naturally apply anatomical knowledge and clinical expertise to recognizing key anatomical structures in complex US images. However, mainstream object detection approaches usually treat each structure recognition separately, overlooking anatomical correlations between different structures in fetal US planes. In this work, we propose a Fetal Anatomy Reasoning Network (FARN) that incorporates two kinds of relationship forms: a global context semantic block summarized with visual similarity and a local topology relationship block depicting structural pair constraints. Specifically, by designing the Adaptive Relation Graph Reasoning (ARGR) module, anatomical structures are treated as nodes, with two kinds of relationships between nodes modeled as edges. The flexibility of the model is enhanced by constructing the adaptive relationship graph in a data-driven way, enabling adaptation to various data samples without the need for predefined additional constraints. The feature representation is further refined by aggregating the outputs of the ARGR module. Comprehensive experimental results demonstrate that FARN achieves promising performance in detecting 37 anatomical structures across key US planes in tertiary obstetric screening. FARN effectively utilizes key relationships to improve detection performance, demonstrates robustness to small-scale, similar, and indistinct structures, and avoids some detection errors that deviate from anatomical norms. Overall, our study serves as a resource for developing efficient and concise approaches to model inter-anatomy relationships.

NeurIPS Conference 2024 Conference Paper

How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression

  • Xingwu Chen
  • Lei Zhao
  • Difan Zou

Despite the remarkable success of transformer-based models in various real-world tasks, their underlying mechanisms remain poorly understood. Recent studies have suggested that transformers can implement gradient descent as an in-context learner for linear regression problems and have developed various theoretical analyses accordingly. However, these works mostly focus on the expressive power of transformers by designing specific parameter constructions, lacking a comprehensive understanding of their inherent working mechanisms post-training. In this study, we consider a sparse linear regression problem and investigate how a trained multi-head transformer performs in-context learning. We experimentally discover that the utilization of multi-heads exhibits different patterns across layers: multiple heads are utilized and essential in the first layer, while usually only a single head is sufficient for subsequent layers. We provide a theoretical explanation for this observation: the first layer preprocesses the context data, and the following layers execute simple optimization steps based on the preprocessed context. Moreover, we demonstrate that such a preprocess-then-optimize algorithm can significantly outperform naive gradient descent and ridge regression algorithms. Further experimental results support our explanations. Our findings offer insights into the benefits of multi-head attention and contribute to understanding the more intricate mechanisms hidden within trained transformers.

ICML Conference 2024 Conference Paper

Is Inverse Reinforcement Learning Harder than Standard Reinforcement Learning? A Theoretical Perspective

  • Lei Zhao
  • Mengdi Wang 0001
  • Yu Bai 0017

Inverse Reinforcement Learning (IRL)—the problem of learning reward functions from demonstrations of an expert policy —plays a critical role in developing intelligent systems. While widely used in applications, theoretical understandings of IRL present unique challenges and remain less developed compared with standard RL. For example, it remains open how to do IRL efficiently in standard offline settings with pre-collected data, where states are obtained from a behavior policy (which could be the expert policy itself), and actions are sampled from the expert policy. This paper provides the first line of results for efficient IRL in vanilla offline and online settings using polynomial samples and runtime. Our algorithms and analyses seamlessly adapt the pessimism principle commonly used in offline RL, and achieve IRL guarantees in stronger metrics than considered in existing work. We provide lower bounds showing that our sample complexities are nearly optimal. As an application, we also show that the learned rewards can transfer to another target MDP with suitable guarantees when the target MDP satisfies certain similarity assumptions with the original (source) MDP.

AAAI Conference 2024 Conference Paper

PNeSM: Arbitrary 3D Scene Stylization via Prompt-Based Neural Style Mapping

  • Jiafu Chen
  • Wei Xing
  • Jiakai Sun
  • Tianyi Chu
  • Yiling Huang
  • Boyan Ji
  • Lei Zhao
  • Huaizhong Lin

3D scene stylization refers to transform the appearance of a 3D scene to match a given style image, ensuring that images rendered from different viewpoints exhibit the same style as the given style image, while maintaining the 3D consistency of the stylized scene. Several existing methods have obtained impressive results in stylizing 3D scenes. However, the mod- els proposed by these methods need to be re-trained when applied to a new scene. In other words, their models are cou- pled with a specific scene and cannot adapt to arbitrary other scenes. To address this issue, we propose a novel 3D scene stylization framework to transfer an arbitrary style to an ar- bitrary scene, without any style-related or scene-related re- training. Concretely, we first map the appearance of the 3D scene into a 2D style pattern space, which realizes complete disentanglement of the geometry and appearance of the 3D scene and makes our model be generalized to arbitrary 3D scenes. Then we stylize the appearance of the 3D scene in the 2D style pattern space via a prompt-based 2D stylization al- gorithm. Experimental results demonstrate that our proposed framework is superior to SOTA methods in both visual qual- ity and generalization.

IJCAI Conference 2024 Conference Paper

Towards Highly Realistic Artistic Style Transfer via Stable Diffusion with Step-aware and Layer-aware Prompt

  • Zhanjie Zhang
  • Quanwei Zhang
  • Huaizhong Lin
  • Wei Xing
  • Juncheng Mo
  • Shuaicheng Huang
  • Jinheng Xie
  • Guangyuan Li

Artistic style transfer aims to transfer the learned artistic style onto an arbitrary content image, generating artistic stylized images. Existing generative adversarial network-based methods fail to generate highly realistic stylized images and always introduce obvious artifacts and disharmonious patterns. Recently, large-scale pre-trained diffusion models opened up a new way for generating highly realistic artistic stylized images. However, diffusion model-based methods generally fail to preserve the content structure of input content images well, introducing some undesired content structure and style patterns. To address the above problems, we propose a novel pre-trained diffusion-based artistic style transfer method, called LSAST, which can generate highly realistic artistic stylized images while preserving the content structure of input content images well, without bringing obvious artifacts and disharmonious style patterns. Specifically, we introduce a Step-aware and Layer-aware Prompt Space, a set of learnable prompts, which can learn the style information from the collection of artworks and dynamically adjusts the input images' content structure and style pattern. To train our prompt space, we propose a novel inversion method, called Step-ware and Layer-aware Prompt Inversion, which allows the prompt space to learn the style information of the artworks collection. In addition, we inject a pre-trained conditional branch of ControlNet into our LSAST, which further improved our framework's ability to maintain content structure. Extensive experiments demonstrate that our proposed method can generate more highly realistic artistic stylized images than the state-of-the-art artistic style transfer methods. Code is available at https: //github. com/Jamie-Cheung/LSAST.

JBHI Journal 2024 Journal Article

TransFSM: Fetal Anatomy Segmentation and Biometric Measurement in Ultrasound Images Using a Hybrid Transformer

  • Lei Zhao
  • Guanghua Tan
  • Bin Pu
  • Qianghui Wu
  • Hongliang Ren
  • Kenli Li

Biometric parameter measurements are powerful tools for evaluating a fetus's gestational age, growth pattern, and abnormalities in a 2D ultrasound. However, it is still challenging to measure fetal biometric parameters automatically due to the indiscriminate confusing factors, limited foreground-background contrast, variety of fetal anatomy shapes at different gestational ages, and blurry anatomical boundaries in ultrasound images. The performance of a standard CNN architecture is limited for these tasks due to the restricted receptive field. We propose a novel hybrid Transformer framework, TransFSM, to address fetal multi-anatomy segmentation and biometric measurement tasks. Unlike the vanilla Transformer based on a single-scale input, TransFSM has a deformable self-attention mechanism, so it can effectively process multi-scale information to segment fetal anatomy with irregular shapes and different sizes. We devised a boundary-aware decoder (BAD) to capture more intrinsic local details using boundary-wise prior knowledge, which compensates for the defects of the Transformer in extracting local features. In addition, a Transformer auxiliary segment head is designed to improve mask prediction by learning the semantic correspondence of the same pixel categories and feature discriminability among different pixel categories. Extensive experiments were conducted on clinical cases and benchmark datasets for anatomy segmentation and biometric measurement tasks. The experiment results indicate that our method achieves state-of-the-art performance in seven evaluation metrics compared with CNN-based, Transformer-based, and hybrid approaches. By knowledge distillation, the proposed TransFSM can create a more compact and efficient model with high deploying potential in resource-constrained scenarios. Our study serves as a unified framework for biometric estimation across multiple anatomical regions to monitor fetal growth in clinical practice.

EAAI Journal 2023 Journal Article

Automatic waste detection with few annotated samples: Improving waste management efficiency

  • Wei Zhou
  • Lei Zhao
  • Hongpu Huang
  • Yuzhi Chen
  • Sixuan Xu
  • Chen Wang

Automatic waste detection in natural environments exhibits a great potential to improve the efficiency and reduce the labor cost of waste management. Recent deep learning-based waste detectors rely heavily on substantial annotated samples for training, but annotating sufficient samples for various categories of waste is labor-intensive and time-consuming. To address this issue, this paper simulates the visual system of human beings and develops a few-shot waste detection framework. To enable the proposed framework more suitable for waste detection, a waste proposal module using a comprehensive feature fusion manner is designed to allow the features of support images to fully interact with those of query images, guiding the framework to generate more potential region proposals containing waste. Also, a waste classification module using soft attention mechanism and foreground mask is designed to alleviate the issue of spatial misalignment and achieve the fine-grained classification towards waste-related proposals. The proposed framework is a general detection framework which can flexibly detect various categories of waste with few labeled samples (i. e. , less than 30 instances per category). Experimental results show that the proposed framework achieves a mean average precision of 31. 16% over 12 waste categories when only few samples (i. e. , 30 instances per category) are provided, surpassing a state-of-the-art few-shot detector named AFDNet by 1. 68%. This data scale-insensitive nature allows humans to reduce the effort and time required for laborious waste image collection and annotation, significantly increasing the flexibility of automatic waste detection and boosting the efficiency of waste management.

AAAI Conference 2023 Conference Paper

Generative Image Inpainting with Segmentation Confusion Adversarial Training and Contrastive Learning

  • Zhiwen Zuo
  • Lei Zhao
  • Ailin Li
  • Zhizhong Wang
  • Zhanjie Zhang
  • Jiafu Chen
  • Wei Xing
  • Dongming Lu

This paper presents a new adversarial training framework for image inpainting with segmentation confusion adversarial training (SCAT) and contrastive learning. SCAT plays an adversarial game between an inpainting generator and a segmentation network, which provides pixel-level local training signals and can adapt to images with free-form holes. By combining SCAT with standard global adversarial training, the new adversarial training framework exhibits the following three advantages simultaneously: (1) the global consistency of the repaired image, (2) the local fine texture details of the repaired image, and (3) the flexibility of handling images with free-form holes. Moreover, we propose the textural and semantic contrastive learning losses to stabilize and improve our inpainting model's training by exploiting the feature representation space of the discriminator, in which the inpainting images are pulled closer to the ground truth images but pushed farther from the corrupted images. The proposed contrastive losses better guide the repaired images to move from the corrupted image data points to the real image data points in the feature representation space, resulting in more realistic completed images. We conduct extensive experiments on two benchmark datasets, demonstrating our model's effectiveness and superiority both qualitatively and quantitatively.

ICRA Conference 2023 Conference Paper

Implementation and Optimization of Grasping Learning with Dual-modal Soft Gripper

  • Lei Zhao
  • Haoyue Liu
  • Feihan Li
  • Xingyu Ding
  • Yuhao Sun
  • Fuchun Sun 0001
  • Jianhua Shan
  • Qi Ye

Robust and efficient grasping of different objects is still an open problem due to the difficulty of integrating multidisciplinary knowledge such as gripper ontology design, perception, control, and learning. In recent years, learning-based methods have achieved excellent results in grasping various novel objects. However, current methods are usually limited to a single grasping mode or rely on different end effectors to grasp objects of different shapes. For human beings, our hands are capable of grasping various objects with changes in grasping methods and form of hands. In light of this, developing a gripper with similar performance could possibly improve the robot's gripping ability. In this paper, we design a dual-modal soft gripper (DSG) and propose a deep reinforcement learning (DRL) framework to implement the operations. Both of our grasping modes, namely enveloping and pinching, are achieved through the tendon drive system and the deformation of the spring steel plate, which enables the gripper to switch between the two grasping modes in real time. We also combined the cutting-edge achievements of deep learning and reinforcement learning to design an autonomous grasping algorithm based on Q-learning and a deep Q network. Moreover, to fully utilize the visual input from the sensor, we added semantic embeddings of target objects to facilitate the learning, which is especially useful in deciding the grasping method for objects previously unseen. We also evaluate our DRL framework in different scenarios, offering a detailed comparison of each grasping mode and the mixed method (with or without semantic information). Our design has proved efficient in reducing the number of failing grasping actions and improving the success rate when facing novel and tricky objects.

AAAI Conference 2023 Conference Paper

MicroAST: Towards Super-fast Ultra-Resolution Arbitrary Style Transfer

  • Zhizhong Wang
  • Lei Zhao
  • Zhiwen Zuo
  • Ailin Li
  • Haibo Chen
  • Wei Xing
  • Dongming Lu

Arbitrary style transfer (AST) transfers arbitrary artistic styles onto content images. Despite the recent rapid progress, existing AST methods are either incapable or too slow to run at ultra-resolutions (e.g., 4K) with limited resources, which heavily hinders their further applications. In this paper, we tackle this dilemma by learning a straightforward and lightweight model, dubbed MicroAST. The key insight is to completely abandon the use of cumbersome pre-trained Deep Convolutional Neural Networks (e.g., VGG) at inference. Instead, we design two micro encoders (content and style encoders) and one micro decoder for style transfer. The content encoder aims at extracting the main structure of the content image. The style encoder, coupled with a modulator, encodes the style image into learnable dual-modulation signals that modulate both intermediate features and convolutional filters of the decoder, thus injecting more sophisticated and flexible style signals to guide the stylizations. In addition, to boost the ability of the style encoder to extract more distinct and representative style signals, we also introduce a new style signal contrastive loss in our model. Compared to the state of the art, our MicroAST not only produces visually superior results but also is 5-73 times smaller and 6-18 times faster, for the first time enabling super-fast (about 0.5 seconds) AST at 4K ultra-resolutions.

IJCAI Conference 2023 Conference Paper

Sequential Recommendation with Probabilistic Logical Reasoning

  • Huanhuan Yuan
  • Pengpeng Zhao
  • Xuefeng Xian
  • Guanfeng Liu
  • Yanchi Liu
  • Victor S. Sheng
  • Lei Zhao

Deep learning and symbolic learning are two frequently employed methods in Sequential Recommendation (SR). Recent neural-symbolic SR models demonstrate their potential to enable SR to be equipped with concurrent perception and cognition capacities. However, neural-symbolic SR remains a challenging problem due to open issues like representing users and items in logical reasoning. In this paper, we combine the Deep Neural Network (DNN) SR models with logical reasoning and propose a general framework named Sequential Recommendation with Probabilistic Logical Reasoning (short for SR-PLR). This framework allows SR-PLR to benefit from both similarity matching and logical reasoning by disentangling feature embedding and logic embedding in the DNN and probabilistic logic network. To better capture the uncertainty and evolution of user tastes, SR-PLR embeds users and items with a probabilistic method and conducts probabilistic logical reasoning on users' interaction patterns. Then the feature and logic representations learned from the DNN and logic network are concatenated to make the prediction. Finally, experiments on various sequential recommendation models demonstrate the effectiveness of the SR-PLR. Our code is available at https: //github. com/Huanhuaneryuan/SR-PLR.

IJCAI Conference 2023 Conference Paper

TeSTNeRF: Text-Driven 3D Style Transfer via Cross-Modal Learning

  • Jiafu Chen
  • Boyan Ji
  • Zhanjie Zhang
  • Tianyi Chu
  • Zhiwen Zuo
  • Lei Zhao
  • Wei Xing
  • Dongming Lu

Text-driven 3D style transfer aims at stylizing a scene according to the text and generating arbitrary novel views with consistency. Simply combining image/video style transfer methods and novel view synthesis methods results in flickering when changing viewpoints, while existing 3D style transfer methods learn styles from images instead of texts. To address this problem, we for the first time design an efficient text-driven model for 3D style transfer, named TeSTNeRF, to stylize the scene using texts via cross-modal learning: we leverage an advanced text encoder to embed the texts in order to control 3D style transfer and align the input text and output stylized images in latent space. Furthermore, to obtain better visual results, we introduce style supervision, learning feature statistics from style images and utilizing 2D stylization results to rectify abrupt color spill. Extensive experiments demonstrate that TeSTNeRF significantly outperforms existing methods and provides a new way to guide 3D style transfer.

IJCAI Conference 2023 Conference Paper

VGOS: Voxel Grid Optimization for View Synthesis from Sparse Inputs

  • Jiakai Sun
  • Zhanjie Zhang
  • Jiafu Chen
  • Guangyuan Li
  • Boyan Ji
  • Lei Zhao
  • Wei Xing

Neural Radiance Fields (NeRF) has shown great success in novel view synthesis due to its state-of-the-art quality and flexibility. However, NeRF requires dense input views (tens to hundreds) and a long training time (hours to days) for a single scene to generate high-fidelity images. Although using the voxel grids to represent the radiance field can significantly accelerate the optimization process, we observe that for sparse inputs, the voxel grids are more prone to overfitting to the training views and will have holes and floaters, which leads to artifacts. In this paper, we propose VGOS, an approach for fast (3-5 minutes) radiance field reconstruction from sparse inputs (3-10 views) to address these issues. To improve the performance of voxel-based radiance field in sparse input scenarios, we propose two methods: (a) We introduce an incremental voxel training strategy, which prevents overfitting by suppressing the optimization of peripheral voxels in the early stage of reconstruction. (b) We use several regularization techniques to smooth the voxels, which avoids degenerate solutions. Experiments demonstrate that VGOS achieves state-of-the-art performance for sparse inputs with super-fast convergence. Code will be available at https: //github. com/SJoJoK/VGOS.

JBHI Journal 2022 Journal Article

A Cascaded Multi-Task Generative Framework for Detecting Aortic Dissection on 3-D Non-Contrast-Enhanced Computed Tomography

  • Xiangyu Xiong
  • Yan Ding
  • Chuanqi Sun
  • Zhuoneng Zhang
  • Xiuhong Guan
  • Tianjing Zhang
  • Hao Chen
  • Hongyan Liu

Contrast-enhanced computed tomography (CE-CT) is the gold standard for diagnosing aortic dissection (AD). However, contrast agents can cause allergic reactions or renal failure in some patients. Moreover, AD diagnosis by radiologists using non-contrast-enhanced CT (NCE-CT) images has poor sensitivity. To address this issue, we propose a novel cascaded multi-task generative framework for AD detection using NCE-CT volumes. The framework includes a 3D nnU-Net and a 3D multi-task generative architecture (3D MTGA). Specifically, the 3D nnU-Net was employed to segment aortas from NCE-CT volumes. The 3D MTGA was then employed to simultaneously synthesize CE-CT volumes, segment true & false lumen, and classify the patient as AD or non-AD. A theoretical formulation demonstrated that the 3D MTGA could increase the Jensen–Shannon Divergence (JSD) between AD and non-AD for each NCE-CT volume, thus indirectly improving the AD detection performance. Experiments also showed that the proposed framework could achieve an average accuracy of 0. 831, a sensitivity of 0. 938, and an F1-score of 0. 847 in comparison with seven state-of-the-art classification models used by three radiologists with junior, intermediate, and senior experiences, respectively. The experimental results indicate that the proposed framework obtains superior performance to state-of-the-art models in AD detection. Thus, it has great potential to reduce the misdiagnosis of AD using NCE-CT in clinical practice. The source codes and supplementary materials for our framework are available at https://github.com/yXiangXiong/CMTGF.

YNICL Journal 2022 Journal Article

Altered frequency-specific/universal amplitude characteristics of spontaneous brain oscillations in patients with bipolar disorder

  • Zhi-Fang Zhang
  • Qi-Jing Bo
  • Feng Li
  • Lei Zhao
  • Peng Gao
  • Yun Wang
  • Rui Liu
  • Xiong-Ying Chen

The human brain is a dynamic system with intrinsic oscillations in spontaneous neural activity. Whether the dynamic characteristics of these spontaneous oscillations are differentially altered across different frequency bands in patients with bipolar disorder (BD) remains unclear. This study recruited 65 patients with BD and 85 healthy controls (HCs). The entire frequency range of resting-state fMRI data was decomposed into four frequency intervals. Two-way repeated-measures ANCOVA was employed to detect frequency-specific/universal alterations in the dynamic oscillation amplitude in BD. The patients were then divided into two subgroups according to their mood states to explore whether these alterations were independent of their mood states. Finally, other window sizes, step sizes, and window types were tested to replicate all analyses. Frequency-specific abnormality of the dynamic oscillation amplitude was detected within the posterior medial parietal cortex (centered at the precuneus extending to the posterior cingulate cortex). This specific profile indicates decreased amplitudes in the lower frequency bands (slow-5/4) and no amplitude changes in the higher frequency bands (slow-3/2) compared with HCs. Frequency-universal abnormalities of the dynamic oscillation amplitude were also detectable, indicating increased amplitudes in the thalamus and left cerebellum anterior lobe but decreased amplitudes in the medial superior frontal gyrus. These alterations were independent of the patients' mood states and replicable across multiple analytic and parametric settings. In short, frequency-specific/universal amplitude characteristics of spontaneous oscillations were observed in patients with BD. These abnormal characteristics have important implications for specific functional changes in BD from multiple frequency and dynamic perspectives.

NeurIPS Conference 2022 Conference Paper

An Investigation into Whitening Loss for Self-supervised Learning

  • Xi Weng
  • Lei Huang
  • Lei Zhao
  • Rao Anwer
  • Salman H. Khan
  • Fahad Shahbaz Khan

A desirable objective in self-supervised learning (SSL) is to avoid feature collapse. Whitening loss guarantees collapse avoidance by minimizing the distance between embeddings of positive pairs under the conditioning that the embeddings from different views are whitened. In this paper, we propose a framework with an informative indicator to analyze whitening loss, which provides a clue to demystify several interesting phenomena as well as a pivoting point connecting to other SSL methods. We reveal that batch whitening (BW) based methods do not impose whitening constraints on the embedding, but they only require the embedding to be full-rank. This full-rank constraint is also sufficient to avoid dimensional collapse. Based on our analysis, we propose channel whitening with random group partition (CW-RGP), which exploits the advantages of BW-based methods in preventing collapse and avoids their disadvantages requiring large batch size. Experimental results on ImageNet classification and COCO object detection reveal that the proposed CW-RGP possesses a promising potential for learning good representations. The code is available at https: //github. com/winci-ai/CW-RGP.

IJCAI Conference 2022 Conference Paper

DivSwapper: Towards Diversified Patch-based Arbitrary Style Transfer

  • Zhizhong Wang
  • Lei Zhao
  • Haibo Chen
  • Zhiwen Zuo
  • Ailin Li
  • Wei Xing
  • Dongming Lu

Gram-based and patch-based approaches are two important research lines of style transfer. Recent diversified Gram-based methods have been able to produce multiple and diverse stylized outputs for the same content and style images. However, as another widespread research interest, the diversity of patch-based methods remains challenging due to the stereotyped style swapping process based on nearest patch matching. To resolve this dilemma, in this paper, we dive into the crux of existing patch-based methods and propose a universal and efficient module, termed DivSwapper, for diversified patch-based arbitrary style transfer. The key insight is to use an essential intuition that neural patches with higher activation values could contribute more to diversity. Our DivSwapper is plug-and-play and can be easily integrated into existing patch-based and Gram-based methods to generate diverse results for arbitrary styles. We conduct theoretical analyses and extensive experiments to demonstrate the effectiveness of our method, and compared with state-of-the-art algorithms, it shows superiority in diversity, quality, and efficiency.

YNICL Journal 2022 Journal Article

Network impact score is an independent predictor of post-stroke cognitive impairment: A multicenter cohort study in 2341 patients with acute ischemic stroke

  • J. Matthijs Biesbroek
  • Nick A. Weaver
  • Hugo P. Aben
  • Hugo J. Kuijf
  • Jill Abrigo
  • Hee-Joon Bae
  • Mélanie Barbay
  • Jonathan G. Best

BACKGROUND: Post-stroke cognitive impairment (PSCI) is a common consequence of stroke. Accurate prediction of PSCI risk is challenging. The recently developed network impact score, which integrates information on infarct location and size with brain network topology, may improve PSCI risk prediction. AIMS: To determine if the network impact score is an independent predictor of PSCI, and of cognitive recovery or decline. METHODS: We pooled data from patients with acute ischemic stroke from 12 cohorts through the Meta VCI Map consortium. PSCI was defined as impairment in ≥ 1 cognitive domain on neuropsychological examination, or abnormal Montreal Cognitive Assessment. Cognitive recovery was defined as conversion from PSCI 24 months) and cognitive recovery or decline using logistic regression. Models were adjusted for age, sex, education, prior stroke, infarct volume, and study site. RESULTS: We included 2341 patients with 4657 cognitive assessments. PSCI was present in 398/844 patients (47%) 24 months. Cognitive recovery occurred in 64/181 (35%) patients and cognitive decline in 26/287 (9%). The network impact score predicted PSCI in the univariable (OR 1.50, 95%CI 1.34-1.68) and multivariable (OR 1.27, 95%CI 1.10-1.46) GEE model, with similar ORs in the logistic regression models for specified post-stroke intervals. The network impact score was not associated with cognitive recovery or decline. CONCLUSIONS: The network impact score is an independent predictor of PSCI. As such, the network impact score may contribute to a more precise and individualized cognitive prognostication in patients with ischemic stroke. Future studies should address if multimodal prediction models, combining the network impact score with demographics, clinical characteristics and other advanced brain imaging biomarkers, will provide accurate individualized prediction of PSCI. A tool for calculating the network impact score is freely available at https://metavcimap.org/features/software-tools/lsm-viewer/.

IJCAI Conference 2022 Conference Paper

Style Fader Generative Adversarial Networks for Style Degree Controllable Artistic Style Transfer

  • Zhiwen Zuo
  • Lei Zhao
  • Shuobin Lian
  • Haibo Chen
  • Zhizhong Wang
  • Ailin Li
  • Wei Xing
  • Dongming Lu

Artistic style transfer is the task of synthesizing content images with learned artistic styles. Recent studies have shown the potential of Generative Adversarial Networks (GANs) for producing artistically rich stylizations. Despite the promising results, they usually fail to control the generated images' style degree, which is inflexible and limits their applicability for practical use. To address the issue, in this paper, we propose a novel method that for the first time allows adjusting the style degree for existing GAN-based artistic style transfer frameworks in real time after training. Our method introduces two novel modules into existing GAN-based artistic style transfer frameworks: a Style Scaling Injection (SSI) module and a Style Degree Interpretation (SDI) module. The SSI module accepts the value of Style Degree Factor (SDF) as the input and outputs parameters that scale the feature activations in existing models, offering control signals to alter the style degrees of the stylizations. And the SDI module interprets the output probabilities of a multi-scale content-style binary classifier as the style degrees, providing a mechanism to parameterize the style degree of the stylizations. Moreover, we show that after training our method can enable existing GAN-based frameworks to produce over-stylizations. The proposed method can facilitate many existing GAN-based artistic style transfer frameworks with marginal extra training overheads and modifications. Extensive qualitative evaluations on two typical GAN-based style transfer models demonstrate the effectiveness of the proposed method for gaining style degree control for them.

AAAI Conference 2022 Conference Paper

Texture Reformer: Towards Fast and Universal Interactive Texture Transfer

  • Zhizhong Wang
  • Lei Zhao
  • Haibo Chen
  • Ailin Li
  • Zhiwen Zuo
  • Wei Xing
  • Dongming Lu

In this paper, we present the texture reformer, a fast and universal neural-based framework for interactive texture transfer with user-specified guidance. The challenges lie in three aspects: 1) the diversity of tasks, 2) the simplicity of guidance maps, and 3) the execution efficiency. To address these challenges, our key idea is to use a novel feed-forward multiview and multi-stage synthesis procedure consisting of I) a global view structure alignment stage, II) a local view texture refinement stage, and III) a holistic effect enhancement stage to synthesize high-quality results with coherent structures and fine texture details in a coarse-to-fine fashion. In addition, we also introduce a novel learning-free view-specific texture reformation (VSTR) operation with a new semantic map guidance strategy to achieve more accurate semanticguided and structure-preserved texture transfer. The experimental results on a variety of application scenarios demonstrate the effectiveness and superiority of our framework. And compared with the state-of-the-art interactive texture transfer algorithms, it not only achieves higher quality results but, more remarkably, also is 2-5 orders of magnitude faster.

NeurIPS Conference 2021 Conference Paper

Artistic Style Transfer with Internal-external Learning and Contrastive Learning

  • Haibo Chen
  • Lei Zhao
  • Zhizhong Wang
  • Huiming Zhang
  • Zhiwen Zuo
  • Ailin Li
  • Wei Xing
  • Dongming Lu

Although existing artistic style transfer methods have achieved significant improvement with deep neural networks, they still suffer from artifacts such as disharmonious colors and repetitive patterns. Motivated by this, we propose an internal-external style transfer method with two contrastive losses. Specifically, we utilize internal statistics of a single style image to determine the colors and texture patterns of the stylized image, and in the meantime, we leverage the external information of the large-scale style dataset to learn the human-aware style information, which makes the color distributions and texture patterns in the stylized image more reasonable and harmonious. In addition, we argue that existing style transfer methods only consider the content-to-stylization and style-to-stylization relations, neglecting the stylization-to-stylization relations. To address this issue, we introduce two contrastive losses, which pull the multiple stylization embeddings closer to each other when they share the same content or style, but push far away otherwise. We conduct extensive experiments, showing that our proposed method can not only produce visually more harmonious and satisfying artistic images, but also promote the stability and consistency of rendered video clips.

AAAI Conference 2021 Conference Paper

Context-Guided Adaptive Network for Efficient Human Pose Estimation

  • Lei Zhao
  • Jun Wen
  • Pengfei Wang
  • Nenggan Zheng

Although recent work has achieved great progress in human pose estimation (HPE), most methods show limitations in either inference speed or accuracy. In this paper, we propose a fast and accurate end-to-end HPE method, which is specifically designed to overcome the commonly encountered jitter box, defective box and ambiguous box problems of boxbased methods, e. g. Mask R-CNN. Concretely, 1) we propose the ROIGuider to aggregate box instance features from all feature levels under the guidance of global context instance information. Further, 2) the proposed Center Line Branch is equipped with a Dichotomy Extended Area algorithm to adaptively expand each instance box area, and Ambiguity Alleviation strategy to eliminate duplicated keypoints. Finally, 3) to achieve efficient multi-scale feature fusion and real-time inference, we design a novel Trapezoidal Network (TNet) backbone. Experimenting on the COCO dataset, our method achieves 68. 1 AP at 25. 4 fps, and outperforms Mask- RCNN by 8. 9 AP at a similar speed. The competitive performance on the HPE and person instance segmentation tasks over the state-of-the-art models show the promise of the proposed method. The source code will be made available at https: //github. com/zlcnup/CGANet.

TIST Journal 2021 Journal Article

TAML: A Traffic-aware Multi-task Learning Model for Estimating Travel Time

  • Jiajie Xu
  • Saijun Xu
  • Rui Zhou
  • Chengfei Liu
  • An Liu
  • Lei Zhao

Travel time estimation has been recognized as an important research topic that can find broad applications. Existing approaches aim to explore mobility patterns via trajectory embedding for travel time estimation. Though state-of-the-art methods utilize estimated traffic condition (by explicit features such as average traffic speed) for auxiliary supervision of travel time estimation, they fail to model their mutual influence and result in inaccuracy accordingly. To this end, in this article, we propose an improved traffic-aware model, called TAML, which adopts a multi-task learning network to integrate a travel time estimator and a traffic estimator in a shared space and improves the accuracy of estimation by enhanced representation of traffic condition, such that more meaningful implicit features are fully captured. In TAML, multi-task learning is further applied for travel time estimation in multi-granularities (including road segment, sub-path, and entire path). The multiple loss functions are combined by considering the homoscedastic uncertainty of each task. Extensive experiments on two real trajectory datasets demonstrate the effectiveness of our proposed methods.

ECAI Conference 2020 Conference Paper

An Efficient Agreement Mechanism in CapsNets by Pairwise Product

  • Lei Zhao
  • Xiaohui Wang
  • Lei Huang 0015

Capsule networks (CapsNets) are capable of modeling visual hierarchical relationships, which is achieved by the “routing-by-agreement” mechanism. This paper proposes a pairwise agreement mechanism to build capsules, inspired by the feature interactions of factorization machines (FMs). The proposed method has a much lower computation complexity. We further proposed a new CapsNet architecture that combines the strengths of residual networks in representing low-level visual features and CapsNets in modeling the relationships of parts to wholes. We conduct comprehensive experiments to compare the routing algorithms, including dynamic routing, EM routing, and our proposed FM agreement, based on both architectures of original CapsNet and our proposed one, and the results show that our method achieves both excellent performance and efficiency under a variety of situations.

YNICL Journal 2019 Journal Article

Structural covariance in subcortical stroke patients measured by automated MRI-based volumetry

  • Caihong Wang
  • Lei Zhao
  • Yishan Luo
  • Jingchun Liu
  • Peifang Miao
  • Sen Wei
  • Lin Shi
  • Jingliang Cheng

A network-level investigation of the volumetric changes of subcortical stroke patients is still lacking. Here, we explored the alterations of structural covariance caused by subcortical stroke with automated brain volumetry. T1-weighed brain MRI scans were obtained from 63 normal controls (NC), 46 stroke patients with infarct in left internal capsule (CI_L), 33 stroke patients with infarct in right internal capsule (CI_R). We performed automatic anatomical segmentation of the T1-weighted brain images with AccuBrain. Volumetric structural covariance analyses were first performed within the basal ganglia structures that were both identified by voxel-based morphometry with AAL atlas and AccuBrain. Subsequently, we additionally included the infratentorial regions that were particularly quantified by AccuBrain for the structural covariance analyses and investigated the alterations of anatomical connections within these subcortical regions in CI_L and CI_R compared with NC. The association between the regional brain volumetry and motor function was also evaluated in stroke groups. There were significant and extensive volumetric differences in stroke patients. These significant regions were generally symmetric for CI_L and CI_R group depending on the side of stroke, involving both regions close to lesions and remote regions. The structural covariance analyses revealed the synergy volume alteration in subcortical regions both in CI_L and CI_R group. In addition, the alterations of volumetric structural covariance were more extensive in CI_L group than CI_R group. Moreover, we found that the subcortical regions with atrophy contributed to the deficits of motor function in CI_R group but not CI_L group, indicating a lesion-side effect of brain volumetric changes after stroke. These findings indicated that the chronic subcortical stroke patients have extensive disordered anatomical connections involving the whole-brain level network, and the connections patterns depend on the lesion-side.

YNIMG Journal 1999 Journal Article

Real-Time Adaptive Functional MRI

  • Seung-Schik Yoo
  • Charles R.G. Guttmann
  • Lei Zhao
  • Lawrence P. Panych

Adaptively limiting image acquisition to areas of interest will allow more efficient data acquisition time for in-depth characterization of areas of brain activation. We designed and implemented an adaptive image acquisition scheme that uses a multiresolution-based strategy to zoom into the regions of cortical activity. Real-time pulse prescription and data processing capabilities were combined with spatially selective radiofrequency encoding. The method was successfully demonstrated in volunteers performing simple sensorimotor paradigms for simultaneous activation of primary motor and cerebellar areas. We believe that real-time adaptation of spatial and temporal sampling to task-related changes will increase the efficiency and flexibility of functional mapping experiments. Contrast-to-noise analysis in selected regions-of-interest was performed to quantitatively assess the multiresolution adaptive approach.

v2026.09.13