Arrow Research search

Author name cluster

Wei Jia

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

11 papers
1 author row

Possible papers

11

EAAI Journal 2026 Journal Article

Advanced technology-driven few-shot relation extraction: Challenges, opportunities, and future outlook

  • Daiyi Li
  • Yaoyao Liang
  • Shenyi Qian
  • Yifan Sun
  • Huaiguang Wu
  • Wei Jia
  • Wenjie Han
  • Yilong Fu

Relation Extraction (RE), as one of the core tasks in Natural Language Processing (NLP), plays a significant role in information structuring, knowledge discovery, and intelligent system construction. However, when labeled data are scarce, traditional RE methods not only struggle to model effectively but also suffer from a notable decline in the recognition performance of low-frequency relations. Therefore, developing efficient and stable Few-Shot RE (FSRE) methods in the context of the scarcity and high-cost of labeled data has become an important research hotspot. To provide new research ideas for current researchers, this review systematically summarizes the fundamental theories and methodological frameworks in the field of FSRE. Firstly, existing FSRE methods are classified into three categories based on optimization strategies and knowledge utilization methods: parameter optimization-based methods, metric learning-based methods, and large language models (LLMs)-based methods. Secondly, a comprehensive review and summary of these three types of FSRE approaches are presented, analyzing their advantages and limitations from both theoretical and experimental perspectives. Finally, we comprehensively explore the challenges and future directions of the development of this research technology, offering a theoretical foundation and practical reference to guide subsequent research, and promote the advancement of artificial intelligence (AI).

AAAI Conference 2026 Conference Paper

Bidirectional Counterfactual Distillation for Review-Based Recommendation

  • Sheng Sang
  • Shujie Li
  • Shuaiyang Li
  • Kang Liu
  • Teng Li
  • Wei Jia
  • Dan Guo
  • Feng Xue

Review-based recommendation methods typically integrate multiple behaviors, including interactions, reviews, and ratings, to model user preferences. To effectively extract preference signals from diverse behaviors, some studies train multiple student models to capture distinct behavioral patterns, and leverage online distillation to facilitate collaborative learning among them. However, we argue that these techniques suffer from bias contamination from rating distributions and feature homogenization during cross-behavior knowledge transfer: (1) Rating distribution bias, arising from non-uniform historical ratings, propagates across behaviors through distillation, contaminating the true preference representations of other behaviors. (2) Static distillation strategies often lead to homogenized behavioral features, hindering the learning of behavior-specific preferences. To address these issues, we propose a novel Bidirectional Counterfactual Distillation (BiCoD) framework for review-based recommendation. In BiCoD, we first design an adversarial counterfactual distillation module to suppress the impact of non-uniform rating distributions on distillation, thereby preventing it from contaminating the user's true preference representations across behaviors. Subsequently, we introduce a stage-aware bidirectional distillation strategy to enhance the distinctiveness of behavioral features, facilitating the effective learning of behavior-specific preferences. Extensive experiments on five real-world datasets validate the effectiveness and superiority of the proposed framework.

AAAI Conference 2026 Conference Paper

Event-Guided Scene Text Image Super-Resolution

  • Zihan Qi
  • Zeyu Xiao
  • Haoyi Zhao
  • Yang Zhao
  • Feng Xue
  • Wei Jia

Scene text image super-resolution aims to enhance text legibility by recovering high-resolution text images from low-resolution inputs. However, maintaining fine details such as text strokes, edges, and textual accuracy remains challenging, particularly in low-light environments and high-speed motion scenarios, where degradation is more severe. Event cameras, with their high temporal resolution and ability to capture intensity changes, offer a promising solution for restoring lost fine details and mitigating degradation in these challenging conditions. In this paper, we propose EvTSR, the first framework that integrates Event data for scene Text image Super-Resolution. The core of EvTSR is the dual-stream frequency boost (DSFB) mechanism, which separates image features into high- and low-frequency components. High-frequency details like edges and strokes are enhanced using event data via the event-guided high-frequency (EGH) mechanism, while low-frequency components, responsible for global structure, are refined using the Text-Guided Low-frequency (TGL) mechanism with a pre-trained text recognizer, ensuring textual coherence. To further improve cross-modal integration, we introduce the cross-modal fusion (CMF) mechanism, which effectively aligns event and image features, enabling robust information fusion. Extensive experiments demonstrate that EvTSR achieves superior performance over existing methods.

AAAI Conference 2026 Conference Paper

LinProVSR: Linguistics-Knowledge Guided Progressive Disambiguation Network for Visual Speech Recognition

  • Feng Xue
  • Baochao Zhu
  • Wei Jia
  • Shujie Li
  • Yu Li
  • Jinrui Zhang
  • Shengeng Tang
  • Dan Guo

Visual Speech Recognition (VSR), commonly known as lipreading, enables the recognition of spoken text by analyzing lip visual features. Due to the subtlety of lip movements, its recognition is much harder than other motion recognition tasks. Existing VSR models face the challenge of viseme ambiguity when processing phonemes with similar pronunciations—multiple phonemes share similar viseme features, leading to a notable drop in lipreading accuracy. To address this issue, this study proposes a Linguistics-Knowledge Guided Progressive Disambiguation Network for Visual Speech Recognition(LinProVSR) framework. First, an ambiguous sample set is constructed based on linguistic knowledge to provide supervisory signals for the model's training. Then, a Progressive Contrastive Disambiguation Network (PCDN) is designed, which progressively enhances the model's ability to capture the subtle viseme differences corresponding to similar phonemes through viseme-phoneme contrastive disambiguation in the encoding stage and text contrastive disambiguation in the decoding stage. Furthermore, we pioneer the Ambiguous Word Error Rate (AWER) metric specifically for evaluating recognition of phonetically ambiguous text, and verify the effectiveness of the proposed method on multiple public datasets, achieving a significant breakthrough especially in distinguishing visually similar phonemes.

AAAI Conference 2026 Conference Paper

LSAP-PV: High-Fidelity Palm Vein Image Synthesis via Layered Spectral Absorption Projection-Guided Diffusion Model

  • Sheng Shang
  • Chenglong Zhao
  • Ruixin Zhang
  • Jianlong Jin
  • Jingyun Zhang
  • Jun Wang
  • Yang Zhao
  • Shouhong Ding

Palm vein recognition has emerged as a promising biometric technology, yet its development remains constrained by the scarcity of large-scale publicly available datasets. Several methods of palm vein image generation have been proposed to address this issue. These methods usually focus on the anatomical realism of palm vein patterns, but overlook the biophysical correlation between identities and vein patterns, particularly in simulating identity-specific vein contrast. To tackle this limitation, we propose a novel biophysics-driven synthesis method. Our method constructs a 3D palm vascular tree via established modeling method. Then, a projection model is proposed to map the 3D tree into 2D space to derive palm vein patterns. The projection model is based on skin spectral absorption and simulates the natural attenuation of light passing through the skin using a layer integration method. For different identities, we sample different skin parameters, resulting in varying degrees of attenuation. This method effectively simulates the variation in vein contrast across different identities. Furthermore, we introduce a conditional diffusion model that uses the projected patterns as identity conditions to generate palm vein images. To the best of our knowledge, this is the first palm vein generation method based on the diffusion model. Experimental results demonstrate that our method not only outperforms existing methods, but also enables a recognition model trained on our synthetic data to achieve superior performance compared to a model trained on real-world data at a scale of 2,000 IDs under an open-set protocol with a TAR@FAR=1:1 of 1e-4.

EAAI Journal 2025 Journal Article

Cross-domain facial expression recognition: Bi-Directional Fusion of Active and Stable Information

  • Yanan Zhu
  • Jiaqiu Ai
  • Weibao Xue
  • Mingyang Wu
  • Sen Yang
  • Wei Jia
  • Min Hu

Facial expression recognition (FER) algorithms often encounter obstacles in cross-domain scenarios, attributed to variations in collection conditions such as lighting, weather, age, gender, and skin color of subjects. Unlike existing approaches that primarily focus on extracting globally invariant features and aligning domain distributions, we propose a novel framework that fundamentally shifts the approach to cross-domain FER. Our proposed algorithm, termed Bi-Directional Fusion of Active and Stable Information (FER-DAS), uniquely combines three innovative components: the Active Assessment Strategy (AAS), Cross-Domain Dynamic Class Threshold (CD-DCT), and Weighted Cross-Domain Alignment (WCDA). The AAS component selectively identifies and enhances active samples in the target domain, providing precise annotations for improved model robustness. Samples with the highest uncertainty are deemed active, indicating low prediction confidence and high informational value for model training. These are then filtered using a predefined threshold to ensure only the most informative samples are included in training iterations. In contrast to conventional static threshold techniques, our dynamic class threshold strategy (CD-DCT) adaptively filters stable samples across domains, thereby ensuring that only the most reliable information is utilized in training. The WCDA strategy further refines this process by dynamically assessing and weighting the contribution of target domain samples to class centers, effectively mitigating domain distribution discrepancies. Extensive experiments on multiple benchmark datasets confirm that FER-DAS sets a new standard in cross-domain FER, consistently outperforming existing state-of-the-art methods.

AAAI Conference 2025 Conference Paper

Occlusion-Embedded Hybrid Transformer for Light Field Super-Resolution

  • Zeyu Xiao
  • Zhuoyuan Li
  • Wei Jia

Transformer-based networks have set new benchmarks in light field super-resolution (SR), but adapting them to capture both global and local spatial-angular correlations efficiently remains challenging. Moreover, many methods fail to account for geometric details like occlusions, leading to performance drops. To tackle these issues, we introduce OHT. This hybrid network leverages occlusion maps through an occlusion-embedded mix layer. It combines the strengths of convolutional networks and Transformers via spatial-angular separable convolution (SASep-Conv) and angular self-attention (ASA). SASep-Conv offers a lightweight alternative to 3D convolution for capturing spatial-angular correlations, while the ASA mechanism applies 3D self-attention across the angular dimension. These designs allow OHT to capture global angular correlations effectively. Extensive experiments on multiple datasets demonstrate OHT's superior performance.

AAAI Conference 2025 Conference Paper

PVTree: Realistic and Controllable Palm Vein Generation for Recognition Tasks

  • Sheng Shang
  • Chenglong Zhao
  • Ruixin Zhang
  • Jianlong Jin
  • Jingyun Zhang
  • Rizen Guo
  • Shouhong Ding
  • Yunsheng Wu

Palm vein recognition is an emerging biometric technology that offers enhanced security and privacy. However, acquiring sufficient palm vein data for training deep learning-based recognition models is challenging due to the high costs of data collection and privacy protection constraints. This has led to a growing interest in generating pseudo-palm vein data using generative models. Existing methods, however, often produce unrealistic palm vein patterns or struggle with controlling identity and style attributes. To address these issues, we propose a novel palm vein generation framework named PVTree. First, the palm vein identity is defined by a complex and authentic 3D palm vascular tree, created using an improved Constrained Constructive Optimization (CCO) algorithm. Second, palm vein patterns of the same identity are generated by projecting the same 3D vascular tree into 2D images from different views and converting them into realistic images using a generative model. As a result, PVTree satisfies the need for both identity consistency and intra-class diversity. Extensive experiments conducted on several publicly available datasets demonstrate that our proposed palm vein generation method surpasses existing methods and achieves a higher TAR@FAR=1e-4 under the 1:1 Open-set protocol. To the best of our knowledge, this is the first time that the performance of a recognition model trained on synthetic palm vein data exceeds that of the recognition model trained on real data, which indicates that palm vein image generation research has a promising future.

AAAI Conference 2024 Conference Paper

PCE-Palm: Palm Crease Energy Based Two-Stage Realistic Pseudo-Palmprint Generation

  • Jianlong Jin
  • Lei Shen
  • Ruixin Zhang
  • Chenglong Zhao
  • Ge Jin
  • Jingyun Zhang
  • Shouhong Ding
  • Yang Zhao

The lack of large-scale data seriously hinders the development of palmprint recognition. Recent approaches address this issue by generating large-scale realistic pseudo palmprints from Bézier curves. However, the significant difference between Bézier curves and real palmprints limits their effectiveness. In this paper, we divide the Bézier-Real difference into creases and texture differences, thus reducing the generation difficulty. We introduce a new palm crease energy (PCE) domain as a bridge from Bézier curves to real palmprints and propose a two-stage generation model. The first stage generates PCE images (realistic creases) from Bézier curves, and the second stage outputs realistic palmprints (realistic texture) with PCE images as input. In addition, we also design a lightweight plug-and-play line feature enhancement block to facilitate domain transfer and improve recognition performance. Extensive experimental results demonstrate that the proposed method surpasses state-of-the-art methods. Under extremely few data settings like 40 IDs (only 2.5% of the total training set), our model achieves a 29% improvement over RPG-Palm and outperforms ArcFace with 100% training set by more than 6% in terms of TAR@FAR=1e-6.

AAAI Conference 2024 Conference Paper

Stereo Vision Conversion from Planar Videos Based on Temporal Multiplane Images

  • Shanding Diao
  • Yuan Chen
  • Yang Zhao
  • Wei Jia
  • Zhao Zhang
  • Ronggang Wang

With the rapid development of 3D movie and light-field displays, there is a growing demand for stereo videos. However, generating high-quality stereo videos from planar videos remains a challenging task. Traditional depth-image-based rendering techniques struggle to effectively handle the problem of occlusion exposure, which occurs when the occluded contents become visible in other views. Recently, the single-view multiplane images (MPI) representation has shown promising performance for planar video stereoscopy. However, the MPI still lacks real details that are occluded in the current frame, resulting in blurry artifacts in occlusion exposure regions. In fact, planar videos can leverage complementary information from adjacent frames to predict a more complete scene representation for the current frame. Therefore, this paper extends the MPI from still frames to the temporal domain, introducing the temporal MPI (TMPI). By extracting complementary information from adjacent frames based on optical flow guidance, obscured regions in the current frame can be effectively repaired. Additionally, a new module called masked optical flow warping (MOFW) is introduced to improve the propagation of pixels along optical flow trajectories. Experimental results demonstrate that the proposed method can generate high-quality stereoscopic or light-field videos from a single view and reproduce better occluded details than other state-of-the-art (SOTA) methods. https://github.com/Dio3ding/TMPI

AAAI Conference 2023 Conference Paper

Universal Information Extraction as Unified Semantic Matching

  • Jie Lou
  • Yaojie Lu
  • Dai Dai
  • Wei Jia
  • Hongyu Lin
  • Xianpei Han
  • Le Sun
  • Hua Wu

The challenge of information extraction (IE) lies in the diversity of label schemas and the heterogeneity of structures. Traditional methods require task-specific model design and rely heavily on expensive supervision, making them difficult to generalize to new schemas. In this paper, we decouple IE into two basic abilities, structuring and conceptualizing, which are shared by different tasks and schemas. Based on this paradigm, we propose to universally model various IE tasks with Unified Semantic Matching (USM) framework, which introduces three unified token linking operations to model the abilities of structuring and conceptualizing. In this way, USM can jointly encode schema and input text, uniformly extract substructures in parallel, and controllably decode target structures on demand. Empirical evaluation on 4 IE tasks shows that the proposed method achieves state-of-the-art performance under the supervised experiments and shows strong generalization ability in zero/few-shot transfer settings.

v2026.09.13