Arrow Research search

Author name cluster

Zixuan Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

8 papers
2 author rows

Possible papers

8

EAAI Journal 2026 Journal Article

A dual-path lightweight detector with hybrid attention for real-time object detection

  • Boyang Yu
  • Zixuan Li
  • Yue Cao
  • Xu Zhang
  • Wansu Lim
  • William Liu

Unmanned aerial vehicle (UAV)-based object detection presents significant challenges, including pronounced variations in object scale and limited computational resources. To address these issues, this paper proposes the Dual-path and Bimodal-attention-enhanced Network (DBYNet), a real-time detection framework optimized for UAV applications. DBYNet adopts a deployment-oriented design that integrates a dual-path backbone for spatial–semantic feature decoupling, a hybrid attention mechanism for enhanced contextual modeling, and lightweight optimization strategies to improve inference efficiency. Specifically, a shallow lightweight branch preserves fine-grained spatial details, while a deep branch with deformable convolutions captures high-level semantic features, and the proposed hybrid attention combines Overlapping Cross Attention (OCA) and Channel-spatial Bimodal Attention (CAB) to strengthen feature interaction. In addition, Quantization-Aware Training and temperature-aware distillation are employed to reduce model complexity without compromising accuracy. Extensive experiments on the VisDrone2019 dataset demonstrate that DBYNet achieves a favorable accuracy–efficiency trade-off, particularly improving robustness for small, densely distributed, and low-visibility targets in challenging UAV scenarios.

AAAI Conference 2026 Conference Paper

Emotion and Intention Guided Multi-Modal Learning for Sticker Response Selection

  • Yuxuan Hu
  • Jian Chen
  • Yuhao Wang
  • Zixuan Li
  • Jing Xiong
  • Pengyue Jia
  • Wei Wang
  • Chengming Li

Stickers are widely used in online communication to convey emotions and implicit intentions. The Sticker Response Selection (SRS) task aims to select the most contextually appropriate sticker based on the dialogue. However, existing methods typically rely on semantic matching and model emotional and intentional cues separately, which can lead to mismatches when emotions and intentions are misaligned. To address this issue, we propose Emotion and Intention Guided Multi-Modal Learning (EIGML). This framework is the first to jointly model emotion and intention, effectively reducing the bias caused by isolated modeling and significantly improving selection accuracy. Specifically, we introduce Dual-Level Contrastive Framework to perform both intra-modality and inter-modality alignment, ensuring consistent representation of emotional and intentional features within and across modalities. In addition, we design an Intention-Emotion Guided Multi-Modal Fusion module that integrates emotional and intentional information progressively through three components: Emotion-Guided Intention Knowledge Selection, Intention-Emotion Guided Attention Fusion, and Similarity-Adjusted Matching Mechanism. This design injects rich, effective information into the model and enables a deeper understanding of the dialogue, ultimately enhancing sticker selection performance. Experimental results on two public datasets show that EIGML outperforms state-of-the-art baselines, achieving higher accuracy and a better understanding of emotional and intentional features.

AAAI Conference 2026 Conference Paper

Explicit Modeling of Causal Factors and Confounders for Image Classification

  • Wei Wu
  • Lei Meng
  • Zhuang Qi
  • Zixuan Li
  • Yachong Zhang
  • Xiaoshuo Yan
  • Xiangxu Meng

Causal inference has emerged as a promising approach for identifying decisive semantic factors and eliminating spurious correlations in visual representation learning. However, most existing methods rely on latent, data-driven confounder modeling, normally attributing the source of bias to background information while neglecting object-level semantic confusions that commonly occur in complex scenes. This limits their effectiveness in disentangling causal factors from confounding semantics. To address this challenge, we propose an explicit modeling approach for both causal factors and confounders, termed Explicit Modeling Causal Model (EMCM). The proposed framework consists of three key components. The Features Stability Estimation module explicitly models the relationship between visual semantics and class labels by leveraging clustering patterns to perform class-aware separation of causal and confounding factors. It produces class-specific causal factors and confounding factors linked to ambiguous categories. Subsequently, the Discriminative Features Enhancing module integrates causal factors into fused patch features via front-door intervention for stable semantics. In parallel, the Explicit Confounder Modeling and Debiasing Module learns confounders under clear label guidance and derives debiased context features by TDE modeling. This framework leverages two complementary causal perspectives to construct a unified semantic representation that facilitates improved generalization. Extensive experiments on two datasets demonstrate that EMCM effectively disentangles causal and confounding factors in complex scenarios, consistently outperforming state-of-the-art causal debiasing methods and text-guided methods in all metrics.

IJCAI Conference 2025 Conference Paper

Empowering Vision Transformers with Multi-Scale Causal Intervention for Long-Tailed Image Classification

  • Xiaoshuo Yan
  • Zhaochuan Li
  • Lei Meng
  • Zhuang Qi
  • Wei Wu
  • Zixuan Li
  • Xiangxu Meng

Causal inference has emerged as a promising approach to mitigate long-tail classification by handling the biases introduced by class imbalance. However, along with the change of advanced backbone models from Convolutional Neural Networks (CNNs) to Visual Transformers (ViT), existing causal models may not achieve an expected performance gain. This paper investigates the influence of existing causal models on CNNs and ViT variants, highlighting that ViT's global feature representation makes it hard for causal methods to model associations between fine-grained features and predictions, which leads to difficulties in classifying tail classes with similar visual appearance. To address these issues, this paper proposes TSCNet, a two-stage causal modeling method to discover fine-grained causal associations through multi-scale causal interventions. Specifically, in the hierarchical causal representation learning stage (HCRL), it decouples the background and objects, applying backdoor interventions at both the patch and feature level to prevent model from using class-irrelevant areas to infer labels which enhances fine-grained causal representation. In the counterfactual logits' bias calibration stage (CLBC), it refines the optimization of model's decision boundary by adaptive constructing counterfactual balanced data distribution to remove the spurious associations in the logits caused by data distribution. Extensive experiments conducted on various long-tail benchmarks demonstrate that the proposed TSCNet can eliminate multiple biases introduced by data imbalance, which outperforms existing methods.

EAAI Journal 2025 Journal Article

Reliable federated learning based on delayed gradient aggregation for intelligent connected vehicles

  • Zhigang Yang
  • Cheng Cheng
  • Zixuan Li
  • Ruyan Wang
  • Xuhua Zhang

As an organic combination of the Internet of Vehicles and intelligent vehicles, Intelligent Connected Vehicles (ICVs) have very high research and application value. Traditional data application methods require the local aggregation of sensitive user data, which poses a threat to user data privacy. Federated learning (FL) is a promising machine learning method that leverages distributed, personalized datasets to enhance performance while preserving user privacy. However, in mobile environments, unreliable client data can degrade the global model, reducing accuracy. Additionally, the mobility of ICVs can destabilize the training process, prolonging model updates and diminishing aggregation accuracy. To address these challenges, this paper proposes a dynamic asynchronous aggregation method that improves both reliability and training efficiency in FL for mobile networks. Therefore, it becomes crucial to find reliable aggregation of mobile device participation in FL tasks. To this end, we propose a reliable FL scheme, which only selects reliable mobile devices to participate in model aggregation to improve the generalization ability of the model. In addition, we design a dynamic asynchronous aggregation method based on reputation scores without affecting the model. Reduce model training time without compromising performance. Through experimental analysis, it is proved that this method can improve the reliability and effectiveness of FL tasks in mobile networks.

IJCAI Conference 2025 Conference Paper

Semantic-Space-Intervened Diffusive Alignment for Visual Classification

  • Zixuan Li
  • Lei Meng
  • Guoqing Chao
  • Wei Wu
  • Yimeng Yang
  • Xiaoshuo Yan
  • Zhuang Qi
  • Xiangxu Meng

Cross-modal alignment is an effective approach to improving visual classification. Existing studies typically enforce a one-step mapping that uses deep neural networks to project the visual features to mimic the distribution of textual features. However, they typically face difficulties in finding such a projection due to the two modalities in both the distribution of class-wise samples and the range of their feature values. To address this issue, this paper proposes a novel Semantic-Space-Intervened Diffusive Alignment method, termed SeDA, models a semantic space as a bridge in the visual-to-textual projection, considering both types of features share the same class-level information in classification. More importantly, a bi-stage diffusion framework is developed to enable the progressive alignment between the two modalities. Specifically, SeDA first employs a Diffusion-Controlled Semantic Learner to model the semantic feature space of visual features by constraining the interactive features of the diffusion model and the category centers of visual features. In the later stage of SeDA, the Diffusion-Controlled Semantic Translator focuses on learning the distribution of textual features from the semantic space. Meanwhile, the Progressive Feature Interaction Network introduces stepwise feature interactions at each alignment step, progressively integrating textual information into mapped features. Experimental results show that SeDA achieves stronger cross-modal feature alignment, leading to superior performance over existing methods across multiple scenarios.

ICLR Conference 2024 Conference Paper

DQ-LoRe: Dual Queries with Low Rank Approximation Re-ranking for In-Context Learning

  • Jing Xiong
  • Zixuan Li
  • Chuanyang Zheng
  • Zhijiang Guo
  • Yichun Yin
  • Enze Xie
  • Zhicheng Yang
  • Qingxing Cao

Recent advances in natural language processing, primarily propelled by Large Language Models (LLMs), have showcased their remarkable capabilities grounded in in-context learning. A promising avenue for guiding LLMs in intricate reasoning tasks involves the utilization of intermediate reasoning steps within the Chain-of-Thought (CoT) paradigm. Nevertheless, the central challenge lies in the effective selection of exemplars for facilitating in-context learning. In this study, we introduce a framework that leverages Dual Queries and Low-rank approximation Re-ranking (DQ-LoRe) to automatically select exemplars for in-context learning. Dual Queries first query LLM to obtain LLM-generated knowledge such as CoT, then query the retriever to obtain the final exemplars via both question and the knowledge. Moreover, for the second query, LoRe employs dimensionality reduction techniques to refine exemplar selection, ensuring close alignment with the input question's knowledge. Through extensive experiments, we demonstrate that DQ-LoRe significantly outperforms prior state-of-the-art methods in the automatic selection of exemplars for GPT-4, enhancing performance from 92.5\% to 94.2\%. Our comprehensive analysis further reveals that DQ-LoRe consistently outperforms retrieval-based approaches in terms of both performance and adaptability, especially in scenarios characterized by distribution shifts. DQ-LoRe pushes the boundaries of in-context learning and opens up new avenues for addressing complex reasoning challenges.

AAAI Conference 2023 Conference Paper

Rich Event Modeling for Script Event Prediction

  • Long Bai
  • Saiping Guan
  • Zixuan Li
  • Jiafeng Guo
  • Xiaolong Jin
  • Xueqi Cheng

Script is a kind of structured knowledge extracted from texts, which contains a sequence of events. Based on such knowledge, script event prediction aims to predict the subsequent event. To do so, two aspects should be considered for events, namely, event description (i.e., what the events should contain) and event encoding (i.e., how they should be encoded). Most existing methods describe an event by a verb together with a few core arguments (i.e., subject, object, and indirect object), which are not precise enough. In addition, existing event encoders are limited to a fixed number of arguments, which are not flexible enough to deal with extra information. Thus, in this paper, we propose the Rich Event Prediction (REP) framework for script event prediction. Fundamentally, it is based on the proposed rich event description, which enriches the existing ones with three kinds of important information, namely, the senses of verbs, extra semantic roles, and types of participants. REP contains an event extractor to extract such information from texts. Based on the extracted rich information, a predictor then selects the most probable subsequent event. The core component of the predictor is a transformer-based event encoder that integrates the above information flexibly. Experimental results on the widely used Gigaword Corpus show the effectiveness of the proposed framework.

v2026.09.13