Arrow Research search

Author name cluster

Zeyu Xiao

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

14 papers
2 author rows

Possible papers

14

AAAI Conference 2026 Conference Paper

Event-Guided Scene Text Image Super-Resolution

  • Zihan Qi
  • Zeyu Xiao
  • Haoyi Zhao
  • Yang Zhao
  • Feng Xue
  • Wei Jia

Scene text image super-resolution aims to enhance text legibility by recovering high-resolution text images from low-resolution inputs. However, maintaining fine details such as text strokes, edges, and textual accuracy remains challenging, particularly in low-light environments and high-speed motion scenarios, where degradation is more severe. Event cameras, with their high temporal resolution and ability to capture intensity changes, offer a promising solution for restoring lost fine details and mitigating degradation in these challenging conditions. In this paper, we propose EvTSR, the first framework that integrates Event data for scene Text image Super-Resolution. The core of EvTSR is the dual-stream frequency boost (DSFB) mechanism, which separates image features into high- and low-frequency components. High-frequency details like edges and strokes are enhanced using event data via the event-guided high-frequency (EGH) mechanism, while low-frequency components, responsible for global structure, are refined using the Text-Guided Low-frequency (TGL) mechanism with a pre-trained text recognizer, ensuring textual coherence. To further improve cross-modal integration, we introduce the cross-modal fusion (CMF) mechanism, which effectively aligns event and image features, enabling robust information fusion. Extensive experiments demonstrate that EvTSR achieves superior performance over existing methods.

AAAI Conference 2026 Conference Paper

Exploiting Blurry Representations for Event-guided Video Super-Resolution

  • Zeyu Xiao
  • Xinchao Wang

Blurry video super-resolution (BVSR) remains fundamentally ill-posed due to the simultaneous loss of high-frequency spatial details and reliable motion cues in blurry low-resolution frames. While cascade-based and joint BVSR methods struggle under severe blur, existing event-guided VSR approaches largely assume clean inputs and are ineffective against complex motion degradation. These methods fail to model blurry representations or leverage event signals for blur-aware motion cues, leading to sub-optimal performance. We propose BluR-EVSR, a unified framework that implicitly models Blurry Representations and leverages Event cameras to jointly address both blur and resolution degradation for VSR. The framework begins with a self-supervised degradation learning strategy guided by event streams and neighboring frames, enabling adaptive blur representation without requiring explicit supervision. A dynamic routing mechanism encodes spatially varying degradations, while a motion-saliency degradation-aware attention module injects motion saliency priors to facilitate efficient RGB-event fusion. Integrated into a bidirectional recurrent framework, BluR-EVSR enables temporally consistent and detail-preserving restoration with low computational cost. Extensive experiments across multiple benchmarks show that our method significantly outperforms prior BVSR and event-based approaches.

AAAI Conference 2026 Conference Paper

FreLay: Frequency-aware Energy Function for Training-free Layout-to-Image Generation

  • Bonan Li
  • Yinhan Hu
  • Songhua Liu
  • Zeyu Xiao
  • Xinchao Wang

Layout-to-Image generation has significantly advanced content creation by enabling the rendering of visual text under predefined spatial layouts. Current approaches achieve training-free layout guidance by constructing attention-based energy functions to derive correction gradients. In this paper, we demonstrate that vanilla energy functions suffer from two limitations, resulting in imprecise layout control and visually unrealistic artifacts. First, the normalizing factor of the Boltzmann distribution defined by the energy functions is non-negligible when calculating correction gradients, yet current energy functions cannot compute this factor exactly. Furthermore, while attention varies over time during the denoising process, existing approaches employ a fixed formulation. To address these challenges, we introduce FreLay, a novel training-free approach equipped with a frequency-aware energy function. Our method first reformulates the energy function to handle the normalization factor, enabling accurate computation of correction gradients. Simultaneously, leveraging the prior knowledge that low-frequency information deteriorates slower during noise addition, we design a time-specific energy function for each timestep from a frequency-domain perspective. Experimental results demonstrate that FreLay consistently outperforms existing state-of-the-art training-free methods by a large margin both qualitatively and quantitatively across multiple datasets.

AAAI Conference 2026 Conference Paper

KineST: A Kinematics-guided Spatiotemporal State Space Model for Human Motion Tracking from Sparse Signals

  • Shuting Zhao
  • Zeyu Xiao
  • Xinrong Chen

Full-body motion tracking plays an essential role in AR/VR applications, bridging physical and virtual interactions. However, it is challenging to reconstruct realistic and diverse full-body poses based on sparse signals obtained by head-mounted displays, which are the main devices in AR/VR scenarios. Existing methods for pose reconstruction often incur high computational costs or rely on separately modeling spatial and temporal dependencies, making it difficult to balance accuracy, temporal coherence, and efficiency. To address this problem, we propose KineST, a novel kinematics-guided state space model, which effectively extracts spatiotemporal dependencies while integrating local and global pose perception. The innovation comes from two core ideas. Firstly, in order to better capture intricate joint relationships, the scanning strategy within the State Space Duality framework is reformulated into kinematics-guided bidirectional scanning, which embeds kinematic priors. Secondly, a mixed spatiotemporal representation learning approach is employed to tightly couple spatial and temporal contexts, balancing accuracy and smoothness. Additionally, a geometric angular velocity loss is introduced to impose physically meaningful constraints on rotational variations for further improving motion stability. Extensive experiments demonstrate that KineST has superior performance in both accuracy and temporal consistency within a lightweight framework.

AAAI Conference 2026 Conference Paper

Seeing the Unseen: Zooming in the Dark with Event Cameras

  • Dachun Kai
  • Zeyu Xiao
  • Huyue Zhu
  • Jiaxiao Wang
  • Yueyi Zhang
  • Xiaoyan Sun

This paper addresses low-light video super-resolution (LVSR), aiming to restore high-resolution videos from low-light, low-resolution (LR) inputs. Existing LVSR methods often struggle to recover fine details due to limited contrast and insufficient high-frequency information. To overcome these challenges, we present RetinexEVSR, the first event-driven LVSR framework that leverages high-contrast event signals and Retinex-inspired priors to enhance video quality under low-light scenarios. Unlike previous approaches that directly fuse degraded signals, RetinexEVSR introduces a novel bidirectional cross-modal fusion strategy to extract and integrate meaningful cues from noisy event data and degraded RGB frames. Specifically, an illumination-guided event enhancement module is designed to progressively refine event features using illumination maps derived from the Retinex model, thereby suppressing low-light artifacts while preserving high-contrast details. Furthermore, we propose an event-guided reflectance enhancement module that utilizes the enhanced event features to dynamically recover reflectance details via a multi-scale fusion mechanism. Experimental results show that our RetinexEVSR achieves state-of-the-art performance on three datasets. Notably, on the SDSD benchmark, our method can get up to 2.95 dB gain while reducing runtime by 65% compared to prior event-based methods.

AAAI Conference 2026 Conference Paper

UniMGS: Unifying Mesh and 3D Gaussian Splatting with Single-Pass Rasterization and Proxy-Based Deformation

  • Zeyu Xiao
  • Mingyang Sun
  • Yimin Cong
  • Lintao Wang
  • Dongliang Kou
  • Zhenyi Wu
  • Dingkang Yang
  • Peng Zhai

Joint rendering and deformation of mesh and 3D Gaussian Splatting (3DGS) have significant value as both representations offer complementary advantages for graphics applications. However, due to differences in representation and rendering pipelines, existing studies render meshes and 3DGS separately, making it difficult to accurately handle occlusions and transparency. Moreover, the deformed 3DGS still suffers from visual artifacts due to the sensitivity to the topology quality of the proxy mesh. These issues pose serious obstacles to the joint use of 3DGS and meshes, making it difficult to adapt 3DGS to conventional mesh-oriented graphics pipelines. We propose UniMGS, the first unified framework for rasterizing mesh and 3DGS in a single-pass anti-aliased manner, with a novel binding strategy for 3DGS deformation based on proxy mesh. Our key insight is to blend the colors of both triangle and Gaussian fragments by anti-aliased α-blending in a single pass, achieving visually coherent results with precise handling of occlusion and transparency. To improve the visual appearance of the deformed 3DGS, our Gaussian-centric binding strategy employs a proxy mesh and spatially associates Gaussians with the mesh faces, significantly reducing rendering artifacts. With these two components, UniMGS enables the visualization and manipulation of 3D objects represented by mesh or 3DGS within a unified framework, opening up new possibilities in embodied AI, virtual reality, and gaming. We will release our source code to facilitate future research.

NeurIPS Conference 2025 Conference Paper

Asymmetric Dual-Lens Video Deblurring

  • Zeyu Xiao
  • Xinchao Wang

Modern smartphones often feature asymmetric dual-lens systems, capturing wide-angle and ultra-wide views with complementary perspectives and details. Motion and shake can blur the wide lens, while the ultra-wide lens, despite lower resolution, retains sharper details. This natural complementarity offers valuable cues for video deblurring. However, existing methods focus mainly on single-camera inputs or symmetric stereo pairs, neglecting the cross-lens redundancy in mobile dual-camera systems. In this paper, we propose a practical video deblurring method, AsLeD-Net, which recurrently aligns and propagates temporal reference features from ultra-wide views fused with features extracted from wide-angle blurry frames. AsLeD-Net consists of two key modules: the adaptive local matching (ALM) module, which refines blurry features using $K$-nearest neighbor reference features, and the difference compensation (DC) module, which ensures spatial consistency and reduces misalignment. Additionally, AsLeD-Net uses the reference-guided motion compensation (RMC) module for temporal alignment, further improving frame-to-frame consistency in the deblurring process. We validate the effectiveness of AsLeD-Net through extensive experiments, benchmarking it against potential solutions for asymmetric lens deblurring.

YNIMG Journal 2025 Journal Article

Brain development during the lifespan of cynomolgus monkeys

  • Zhiqiang Tan
  • Binbin Nie
  • Huanhua Wu
  • Bang Li
  • Jingjie Shang
  • Tianhao Zhang
  • Zeyu Xiao
  • Chenchen Dong

F]FDG PET-MRI data from 228 healthy cynomolgus monkeys spanning the age range of 0.5-29.5 years to construct an age-specific multimodal image brain template toolset tailored to cynomolgus monkeys. Their brain volume and glucose metabolism were quantitatively analyzed by utilizing an individualized spatial segmentation algorithm. Our findings encapsulated the growth and development trends, sex differences, and asymmetrical variations in brain volume and glucose metabolism in cynomolgus monkeys, and analyzed the correlation between the brain volume and glucose metabolism. This endeavor enhances our capacity to leverage the cynomolgus monkey model in neuroscience research by providing a valuable resource for researchers. The age-specific brain template toolset and associated data offer a robust foundation for future investigations, facilitating a nuanced understanding of brain development in this primate species and, consequently, informing and advancing neuroscience research employing cynomolgus monkeys.

AAAI Conference 2025 Conference Paper

Event-Enhanced Blurry Video Super-Resolution

  • Dachun Kai
  • Yueyi Zhang
  • Jin Wang
  • Zeyu Xiao
  • Zhiwei Xiong
  • Xiaoyan Sun

In this paper, we tackle the task of blurry video super-resolution (BVSR), aiming to generate high-resolution (HR) videos from low-resolution (LR) and blurry inputs. Current BVSR methods often fail to restore sharp details at high resolutions, resulting in noticeable artifacts and jitter due to insufficient motion information for deconvolution and the lack of high-frequency details in LR frames. To address these challenges, we introduce event signals into BVSR and propose a novel event-enhanced network, Ev-DeblurVSR. To effectively fuse information from frames and events for feature deblurring, we introduce a reciprocal feature deblurring module that leverages motion information from intra-frame events to deblur frame features while reciprocally using global scene context from the frames to enhance event features. Furthermore, to enhance temporal consistency, we propose a hybrid deformable alignment module that fully exploits the complementary motion information from inter-frame events and optical flow to improve motion estimation in the deformable alignment process. Extensive evaluations demonstrate that Ev-DeblurVSR establishes a new state-of-the-art performance on both synthetic and real-world datasets. Notably, on real data, our method is 2.59 dB more accurate and 7.28× faster than the recent best BVSR baseline FMA-Net.

EAAI Journal 2025 Journal Article

Large-scale analytic hierarchy process method based on fuzzy-rough-advantage relation

  • Bin Yu
  • Zeyu Xiao
  • Yinglong Dai
  • Zeshui Xu

In the big data era, the complexity of decision-making problems on a large scale has risen, particularly when confronted with large-scale alternatives (objects). Traditional multi-attribute decision-making methods, however, are becoming less effective in handling these problems. To address this, the present study proposes a Large-scale Analytic Hierarchy Process (LAHP) approach predicated on fuzzy-rough-advantage relations. Initially, an unsupervised learning technique, known as K-means++, is applied to cluster the alternatives (objects). Subsequently, the fuzzy-rough-advantage relationship is utilized to form the judgment matrix for these clusters. Following this, the Analytic Hierarchy Process (AHP) principles are employed to determine the weights of each cluster based on various evaluation criteria. These weights are then utilized to compute the overall scores of each cluster and identify the optimal group. Finally, the applicability of the LAHP method is confirmed through case analysis, with the ranking outcomes being subsequently analyzed and discussed.

ECAI Conference 2025 Conference Paper

MDCF-Net: Modality Decomposition and Compensation Fusion Network for Infrared-Visible Object Detection

  • Jiangtao Fan
  • Zeyu Xiao
  • Anish Jindal
  • Amir Atapour-Abarghouei

Infrared-visible object detection aims to leverage the complementary information between infrared and visible modalities to improve detection performance in challenging environments. However, existing infrared-visible object detection methods face several limitations: (1) difficulty in effectively extracting and decomposing modality-common and modality-specific features; (2) interference from modality-irrelevant or redundant information; and (3) insufficient fusion of cross-modal complementary cues. To address these issues, we propose a novel Modality Decomposition and Compensation Fusion Network (MDCF-Net). Specifically, MDCF-Net first decomposes the common and unique features across different modalities. It then performs selective enhancement and interaction between these features via cross-modality compensation. Finally, a dynamic fusion strategy based on spatial and channel attention is applied to adaptively integrate the enhanced features. Extensive experiments on two public datasets, LLVIP and FLIR, demonstrate that our proposed method achieves superior detection performance and exhibits robust generalisation across various challenging conditions. The Code is available at https: //github. com/fanjiangtao666/MDCF-Net/tree/main

AAAI Conference 2025 Conference Paper

Occlusion-Embedded Hybrid Transformer for Light Field Super-Resolution

  • Zeyu Xiao
  • Zhuoyuan Li
  • Wei Jia

Transformer-based networks have set new benchmarks in light field super-resolution (SR), but adapting them to capture both global and local spatial-angular correlations efficiently remains challenging. Moreover, many methods fail to account for geometric details like occlusions, leading to performance drops. To tackle these issues, we introduce OHT. This hybrid network leverages occlusion maps through an occlusion-embedded mix layer. It combines the strengths of convolutional networks and Transformers via spatial-angular separable convolution (SASep-Conv) and angular self-attention (ASA). SASep-Conv offers a lightweight alternative to 3D convolution for capturing spatial-angular correlations, while the ASA mechanism applies 3D self-attention across the angular dimension. These designs allow OHT to capture global angular correlations effectively. Extensive experiments on multiple datasets demonstrate OHT's superior performance.

AAAI Conference 2022 Conference Paper

Efficient Model-Driven Network for Shadow Removal

  • Yurui Zhu
  • Zeyu Xiao
  • Yanchi Fang
  • Xueyang Fu
  • Zhiwei Xiong
  • Zheng-Jun Zha

Deep Convolutional Neural Networks (CNNs) based methods have achieved significant breakthroughs in the task of single image shadow removal. However, the performance of these methods remains limited for several reasons. First, the existing shadow illumination model ignores the spatially variant property of the shadow images, hindering their further performance. Second, most deep CNNs based methods directly estimate the shadow free results from the input shadow images like a black box, thus losing the desired interpretability. To address these issues, we first propose a new shadow illumination model for the shadow removal task. This new shadow illumination model ensures the identity mapping among unshaded regions, and adaptively performs fine grained spatial mapping between shadow regions and their references. Then, based on the shadow illumination model, we reformulate the shadow removal task as a variational optimization problem. To effectively solve the variational problem, we design an iterative algorithm and unfold it into a deep network, naturally increasing the interpretability of the deep model. Experiments show that our method could achieve SOTA performance with less than half parameters, one-fifth of floating-point of operations (FLOPs), and over seventeen times faster than SOTA method (DHAN).

NeurIPS Conference 2021 Conference Paper

Unfolding Taylor's Approximations for Image Restoration

  • Man Zhou
  • Xueyang Fu
  • Zeyu Xiao
  • Gang Yang
  • Aiping Liu
  • Zhiwei Xiong

Deep learning provides a new avenue for image restoration, which demands a delicate balance between fine-grained details and high-level contextualized information during recovering the latent clear image. In practice, however, existing methods empirically construct encapsulated end-to-end mapping networks without deepening into the rationality, and neglect the intrinsic prior knowledge of restoration task. To solve the above problems, inspired by Taylor’s Approximations, we unfold Taylor’s Formula to construct a novel framework for image restoration. We find the main part and the derivative part of Taylor’s Approximations take the same effect as the two competing goals of high-level contextualized information and spatial details of image restoration respectively. Specifically, our framework consists of two steps, which are correspondingly responsible for the mapping and derivative functions. The former first learns the high-level contextualized information and the later combines it with the degraded input to progressively recover local high-order spatial details. Our proposed framework is orthogonal to existing methods and thus can be easily integrated with them for further improvement, and extensive experiments demonstrate the effectiveness and scalability of our proposed framework.

v2026.09.13