Arrow Research search

Author name cluster

Guoli Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

12 papers
1 author row

Possible papers

12

AAAI Conference 2026 Conference Paper

ClearAIR: A Human-Visual-Perception-Inspired All-in-One Image Restoration

  • Xu Zhang
  • Huan ZHang
  • Guoli Wang
  • Qian Zhang
  • Lefei Zhang

Recently, All-in-One image restoration (AiOIR) has advanced significantly, offering promising solutions for complex real-world degradations. However, most existing approaches heavily rely on degradation-specific representation learning, which can lead to oversmoothing and artifacts in the restored images. To address this limitation, we propose ClearAIR, a novel AiOIR framework inspired by human visual perception and designed with a hierarchical restoration strategy in a coarse-to-fine manner. First, leveraging the global priority characteristic of early human visual perception, we employ an image quality assessment model to evaluate the overall image structure and degradation level. Next, we introduce a Semantic Guidance Unit to provide coarse semantic region guidance and a Task Identifier to predict local degradation types, enabling a more informed characterization of local degradation patterns. Finally, aiming at the challenge of local detail restoration, we propose an Internal Clue Reuse Mechanism that deeply mines the internal information of the image in a self-supervised manner to enhance the model’s capacity for fine-detail recovery. Experimental results demonstrate that ClearAIR achieves superior restoration performance across diverse synthetic and real-world datasets.

AAAI Conference 2026 Conference Paper

EASE: Practical and Efficient Safety Alignment for Small Language Models

  • Haonan Shi
  • Guoli Wang
  • Tu Ouyang
  • An Wang

Small language models (SLMs) are increasingly deployed on edge devices, making their safety alignment crucial yet challenging. Current shallow alignment methods that rely on direct refusal of malicious queries fail to provide robust protection, particularly against adversarial jailbreaks. While deliberative safety reasoning alignment offers deeper alignment for defending against sophisticated attacks, effectively implanting such reasoning capability in SLMs with limited capabilities remains an open challenge. Moreover, safety reasoning incurs significant computational overhead as models apply reasoning to nearly all queries, making it impractical for resource-constrained edge deployment scenarios that demand rapid responses. We propose EASE, a novel framework that enables practical and Efficient safety Alignment for Small languagE models. Our approach first identifies the optimal safety reasoning teacher that can effectively distill safety reasoning capabilities to SLMs. We then align models to selectively activate safety reasoning for dangerous adversarial jailbreak queries while providing direct responses to straightforward malicious queries and general helpful tasks. This selective mechanism enables small models to maintain robust safety guarantees against sophisticated attacks while preserving computational efficiency for benign interactions. Experimental results demonstrate that EASE reduces jailbreak attack success rates by up to 17% compared to shallow alignment methods while reducing inference overhead by up to 90% compared to deliberative safety reasoning alignment, making it practical for SLMs real-world edge deployments.

JBHI Journal 2025 Journal Article

MedKAFormer: When Kolmogorov–Arnold Theorem Meets Vision Transformer for Medical Image Representation

  • Guoli Wang
  • Qikui Zhu
  • Chaoda Song
  • Benzheng Wei
  • Shuo Li

Vision Transformers (ViTs) suffer from high parameter complexity because they rely on Multi-layer Perceptrons (MLPs) for nonlinear representation. This issue is particularly challenging in medical image analysis, where labeled data is limited, leading to inadequate feature representation. Existing methods have attempted to optimize either the patch embedding stage or the non-embedding stage of ViTs. Still, they have struggled to balance effective modeling, parameter complexity, and data availability. Recently, the Kolmogorov–Arnold Network (KAN) was introduced as an alternative to MLPs, offering a potential solution to the large parameter issue in ViTs. However, KAN cannot be directly integrated into ViT due to challenges such as handling 2D structured data and dimensionality catastrophe. To solve this problem, we propose MedKAFormer, the first ViT model to incorporate the Kolmogorov–Arnold (KA) theorem for medical image representation. It includes a Dynamic Kolmogorov–Arnold Convolution (DKAC) layer for flexible nonlinear modeling in the patch embedding stage. Additionally, it introduces a Nonlinear Sparse Token Mixer (NSTM) and a Nonlinear Dynamic Filter (NDF) in the non-embedding stage. These components provide comprehensive nonlinear representation while reducing model overfitting. MedKAFormer reduces parameter complexity by 85. 61% compared to ViT-Base and achieves competitive results on 14 medical datasets across various imaging modalities and structures.

AAAI Conference 2024 Conference Paper

Enhancing RAW-to-sRGB with Decoupled Style Structure in Fourier Domain

  • Xuanhua He
  • Tao Hu
  • Guoli Wang
  • Zejin Wang
  • Run Wang
  • Qian Zhang
  • Keyu Yan
  • Ziyi Chen

RAW to sRGB mapping, which aims to convert RAW images from smartphones into RGB form equivalent to that of Digital Single-Lens Reflex (DSLR) cameras, has become an important area of research. However, current methods often ignore the difference between cell phone RAW images and DSLR camera RGB images, a difference that goes beyond the color matrix and extends to spatial structure due to resolution variations. Recent methods directly rebuild color mapping and spatial structure via shared deep representation, limiting optimal performance. Inspired by Image Signal Processing (ISP) pipeline, which distinguishes image restoration and enhancement, we present a novel Neural ISP framework, named FourierISP. This approach breaks the image down into style and structure within the frequency domain, allowing for independent optimization. FourierISP is comprised of three subnetworks: Phase Enhance Subnet for structural refinement, Amplitude Refine Subnet for color learning, and Color Adaptation Subnet for blending them in a smooth manner. This approach sharpens both color and structure, and extensive evaluations across varied datasets confirm that our approach realizes state-of-the-art results. Code will be available at https://github.com/alexhe101/FourierISP.

IJCAI Conference 2023 Conference Paper

Long-term Wind Power Forecasting with Hierarchical Spatial-Temporal Transformer

  • Yang Zhang
  • Lingbo Liu
  • Xinyu Xiong
  • Guanbin Li
  • Guoli Wang
  • Liang Lin

Wind power is attracting increasing attention around the world due to its renewable, pollution-free, and other advantages. However, safely and stably integrating the high permeability intermittent power energy into electric power systems remains challenging. Accurate wind power forecasting (WPF) can effectively reduce power fluctuations in power system operations. Existing methods are mainly designed for short-term predictions and lack effective spatial-temporal feature augmentation. In this work, we propose a novel end-to-end wind power forecasting model named Hierarchical Spatial-Temporal Transformer Network (HSTTN) to address the long-term WPF problems. Specifically, we construct an hourglass-shaped encoder-decoder framework with skip-connections to jointly model representations aggregated in hierarchical temporal scales, which benefits long-term forecasting. Based on this framework, we capture the inter-scale long-range temporal dependencies and global spatial correlations with two parallel Transformer skeletons and strengthen the intra-scale connections with downsampling and upsampling operations. Moreover, the complementary information from spatial and temporal features is fused and propagated in each other via Contextual Fusion Blocks (CFBs) to promote the prediction further. Extensive experimental results on two large-scale real-world datasets demonstrate the superior performance of our HSTTN over existing solutions.

AAAI Conference 2022 Conference Paper

AdaptivePose: Human Parts as Adaptive Points

  • Yabo Xiao
  • Xiao Juan Wang
  • Dongdong Yu
  • Guoli Wang
  • Qian Zhang
  • Mingshu HE

Multi-person pose estimation methods generally follow topdown and bottom-up paradigms, both of which can be considered as two-stage approaches thus leading to the high computation cost and low efficiency. Towards a compact and efficient pipeline for multi-person pose estimation task, in this paper, we propose to represent the human parts as points and present a novel body representation, which leverages an adaptive point set including the human center and seven humanpart related points to represent the human instance in a more fine-grained manner. The novel representation is more capable of capturing the various pose deformation and adaptively factorizes the long-range center-to-joint displacement thus delivers a single-stage differentiable network to more precisely regress multi-person pose, termed as AdaptivePose. For inference, our proposed network eliminates the grouping as well as refinements and only needs a single-step disentangling process to form multi-person pose. Without any bells and whistles, we achieve the best speed-accuracy trade-offs of 67. 4% AP / 29. 4 fps with DLA-34 and 71. 3% AP / 9. 1 fps with HRNet-W48 on COCO test-dev dataset.

AAAI Conference 2022 Conference Paper

Learning Quality-Aware Representation for Multi-Person Pose Regression

  • Yabo Xiao
  • Dongdong Yu
  • Xiao Juan Wang
  • Lei Jin
  • Guoli Wang
  • Qian Zhang

Off-the-shelf single-stage multi-person pose regression methods generally leverage the instance score (i. e. , confidence of the instance localization) to indicate the pose quality for selecting the pose candidates. We consider that there are two gaps involved in existing paradigm: 1) The instance score is not well interrelated with the pose regression quality. 2) The instance feature representation, which is used for predicting the instance score, does not explicitly encode the structural pose information to predict the reasonable score that represents pose regression quality. To address the aforementioned issues, we propose to learn the pose regression quality-aware representation. Concretely, for the first gap, instead of using the previous instance confidence label (e. g. , discrete {1, 0} or Gaussian representation) to denote the position and confidence for person instance, we firstly introduce the Consistent Instance Representation (CIR) that unifies the pose regression quality score of instance and the confidence of background into a pixel-wise score map to calibrates the inconsistency between instance score and pose regression quality. To fill the second gap, we further present the Query Encoding Module (QEM) including the Keypoint Query Encoding (KQE) to encode the positional and semantic information for each keypoint and the Pose Query Encoding (PQE) which explicitly encodes the predicted structural pose information to better fit the Consistent Instance Representation (CIR). By using the proposed components, we significantly alleviate the above gaps. Our method outperforms previous single-stage regression-based even bottom-up methods and achieves the state-of-the-art result of 71. 7 AP on MS COCO test-dev set.

EAAI Journal 2020 Journal Article

Sensor-based activity recognition of solitary elderly via stigmergy and two-layer framework

  • Zimin Xu
  • Guoli Wang
  • Xuemei Guo

With the acceleration of aging process of population structure, the single resident lifestyle is increasing on account of the high cost of care services and the privacy invasion concern. It is essential to monitor the activities of solitary elderly to find the emergency and lifestyle deviation, as independent life cannot be maintained due to physical or mental problems. The unobtrusive systems are the most preferred choice for the real-life long-term monitoring, while the camera and wearable devices based systems are not suitable due to the privacy and uncomfortableness, respectively. We propose a novel sensor-based activity recognition model based on the two-layer multi-granularity framework and the emergent paradigm with marker-based stigmergy. The stigmergy based marking subsystem builds features by aggregating the context-aware information and generating the two-dimensional activity pheromone trail. The two-layer framework consists of coarse-grained and fine-grained classification subsystems. The coarse-grained subsystem identifies whether the input completed activity segmented by the traditional method is easily-confused, and utilizes our generalized segmentation method to increase the inter-cluster distance. The fine-grained subsystem employs machine learning or deep learning classifiers to realize the activity recognition task. The proposed model is a data-driven model based on the information self-organization. It does not need sophisticated domain knowledge, and can fully mine the hidden feature structure containing semantically related information and spatio-temporal characteristics. The experimental results demonstrate the effectiveness of the proposed method.

IJCAI Conference 2019 Conference Paper

Neurons Merging Layer: Towards Progressive Redundancy Reduction for Deep Supervised Hashing

  • Chaoyou Fu
  • Liangchen Song
  • Xiang Wu
  • Guoli Wang
  • Ran He

Deep supervised hashing has become an active topic in information retrieval. It generates hashing bits by the output neurons of a deep hashing network. During binary discretization, there often exists much redundancy between hashing bits that degenerates retrieval performance in terms of both storage and accuracy. This paper proposes a simple yet effective Neurons Merging Layer (NMLayer) for deep supervised hashing. A graph is constructed to represent the redundancy relationship between hashing bits that is used to guide the learning of a hashing network. Specifically, it is dynamically learned by a novel mechanism defined in our active and frozen phases. According to the learned relationship, the NMLayer merges the redundant neurons together to balance the importance of each output neuron. Moreover, multiple NMLayers are progressively trained for a deep hashing network to learn a more compact hashing code from a long redundant code. Extensive experiments on four datasets demonstrate that our proposed method outperforms state-of-the-art hashing methods.

AAAI Conference 2019 Conference Paper

Self-Ensembling Attention Networks: Addressing Domain Shift for Semantic Segmentation

  • Yonghao Xu
  • Bo Du
  • Lefei Zhang
  • Qian Zhang
  • Guoli Wang
  • Liangpei Zhang

Recent years have witnessed the great success of deep learning models in semantic segmentation. Nevertheless, these models may not generalize well to unseen image domains due to the phenomenon of domain shift. Since pixel-level annotations are laborious to collect, developing algorithms which can adapt labeled data from source domain to target domain is of great significance. To this end, we propose self-ensembling attention networks to reduce the domain gap between different datasets. To the best of our knowledge, the proposed method is the first attempt to introduce selfensembling model to domain adaptation for semantic segmentation, which provides a different view on how to learn domain-invariant features. Besides, since different regions in the image usually correspond to different levels of domain gap, we introduce the attention mechanism into the proposed framework to generate attention-aware features, which are further utilized to guide the calculation of consistency loss in the target domain. Experiments on two benchmark datasets demonstrate that the proposed framework can yield competitive performance compared with the state of the art methods.

EAAI Journal 2018 Journal Article

Online activity recognition and daily habit modeling for solitary elderly through indoor position-based stigmergy

  • Zhichao Tan
  • Liwen Xu
  • Wei Zhong
  • Xuemei Guo
  • Guoli Wang

This paper concerns the issue of monitoring elderly behavior in the context of ambient assisted living (AAL). Under the framework of online daily habit modeling (ODHM), we employ the emergent representation for activities of daily living (ADLs) with position-based stigmergy, and then combine it with convolution neural networks (CNNs) to accomplish the tasks of recognizing ADLs. In addition, we propose a new paradigm of activity summarization with the robustness to break interruptions. Radio tomographic imaging (RTI) is promoted as a simple yet flexible way of facilitating the required position-based stigmergy. Such position-based AAL systems can benefit the advantages of having no need any sophisticated domain models in analyzing and understanding ADLs while no burden training is involved in ODHM. Moreover, the emergent based data aggregation and deep learning of CNN together allow the recognition of ADLs at a fine-grained level, which contributes to the performance improvement of ODHM. Experimental results demonstrate the effectiveness of the proposed approach.

JBHI Journal 2015 Journal Article

Stroke Parameters Identification Algorithm in Handwriting Movements Analysis by Synthesis

  • Min Liu
  • Xuemei Guo
  • Guoli Wang

This paper presents a new approach to identify the stroke parameters in handwriting movement data understanding. A two-step analysis by synthesis paradigm is employed to facilitate the coarse-to-fine parameter identification for all strokes. One is the stroke data extraction, the other is the coarse-to-fine stroke parameter identification. The new consideration of using this two-step paradigm is that the nonnegative primitive factorization technique is incorporated to decouple the overlapped strokes from the measurement data. In comparison to the existing paradigms of using the heuristic stroke data decoupling techniques, our paradigm presented here contributes to alleviating the difficulty of local optimum traps with the well-shaped initializations in the global optimization for jointly identifying stroke parameters. Moreover, our paradigm excludes the iteration between two steps, which contributes to the enhancement of computational efficiency. Experimental results are reported to validate the proposed approach.

v2026.09.13