Arrow Research search

Author name cluster

Xin Zhao

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

53 papers
2 author rows

Possible papers

53

AAAI Conference 2026 Conference Paper

Analyzing and Mitigating Object Hallucination: A Training Bias Perspective

  • Yifan Li
  • Kun Zhou
  • Xin Zhao
  • Lei Fang
  • Jirong Wen

As scaling up training data has significantly improved the general multimodal capabilities of Large Vision-Language Models (LVLMs), they still suffer from the hallucination issue, generating text that is inconsistent with the visual input. This phenomenon motivates us to systematically investigate the role of training data in hallucination. We introduce a new benchmark, POPEv2, which consists of counterfactual images collected from the training data of LVLMs with certain objects masked. Through comprehensive evaluation on POPEv2, we find that current LVLMs suffer from training bias: they fail to fully leverage their training data and hallucinate more frequently on images seen during training. Specifically, they perform poorly on counterfactual images, often incorrectly answering “Yes” to questions about masked objects. To understand this issue, we conduct probing experiments on the models’ internal components, revealing that this training bias is primarily located in the language modeling (LM) head, which fails to correctly translate accurate visual representations into textual outputs. Based on these findings, we propose Obliviate, an efficient and lightweight unlearning method designed to mitigate object hallucination via training bias unlearning. Obliviate identifies the discrepancy between ground-truth labels and model outputs on the training data as a proxy for bias and adopts a parameter- and data-efficient fine-tuning strategy that only updates the LM head. Extensive experiments demonstrate the effectiveness of our approach. While only reusing the training data and updating approximately 2% of the parameters, Obliviate significantly reduces hallucination across both discriminative and generative tasks. Furthermore, it demonstrates strong scalability with respect to both model size (2B to 72B) and training data volume, and exhibits promising generalization to hallucination types beyond object-level hallucination.

AAAI Conference 2026 Conference Paper

ARGH-Mark: Anchor-Synchronized Watermarking with Hamming Correction for Robust and Quality-Preserving LLM Attribution

  • He Li
  • Xiaojun Chen
  • Jingcheng He
  • Zhendong Zhao
  • Shuguang Yuan
  • Xin Zhao
  • Yunfei Yang

The proliferation of large language models has intensified demands for reliable content attribution, yet existing watermarking techniques face a fundamental trilemma: they cannot simultaneously optimize for robustness against attacks, minimal text quality degradation, and detection efficiency. To resolve this challenge, we propose ARGH-Mark, a novel watermarking framework that integrates three synergistic innovations: (1) Anchor-synchronized phase recovery for maintaining detection integrity under insertion/deletion attacks, (2) RG-balanced vocabulary modulation that dynamically partitions lexicons via contextual hashing to preserve generation quality, and (3) Hamming-based error correction enabling single-bit error rectification through algebraic coding. Comprehensive evaluations across question answering (ELI5), summarization (CNN/DailyMail), and text generation (C4) demonstrate state-of-the-art performance: the proposed ARGH-Mark framework achieves near-perfect match rate and bit accuracy across diverse configurations, while preserving the quality of the generated text. It significantly reduces detection latency, enabling real-time extraction, and maintains high robustness against token tampering attacks through integrated Hamming error correction, ensuring reliable attribution in adversarial settings. ARGH-Mark achieves a new Pareto frontier in the watermarking design space and advances trustworthy deployment of generative AI in alignment-critical applications.

AAAI Conference 2026 Conference Paper

BEE-RAG: Balanced Entropy Engineering for Retrieval-Augmented Generation

  • Yuhao Wang
  • Ruiyang Ren
  • Yucheng Wang
  • Jing Liu
  • Xin Zhao
  • Hua Wu
  • Haifeng Wang

With the rapid advancement of large language models (LLMs), retrieval-augmented generation (RAG) has emerged as a critical approach to supplement the inherent knowledge limitations of LLMs. However, due to the typically large volume of retrieved information, RAG tends to operate with long context lengths. From the perspective of entropy engineering, we identify unconstrained entropy growth and attention dilution due to long retrieval context as significant factors affecting RAG performance. In this paper, we propose the balanced entropy-engineered RAG (BEE-RAG) framework, which improves the adaptability of RAG systems to varying context lengths through the principle of entropy invariance. By leveraging balanced context entropy to reformulate attention dynamics, BEE-RAG separates attention sensitivity from context length, ensuring a stable entropy level. Building upon this, we introduce a zero-shot inference strategy for multi-importance estimation and a parameter-efficient adaptive fine-tuning mechanism to obtain the optimal balancing factor for different settings. Extensive experiments across multiple RAG tasks demonstrate the effectiveness of BEE-RAG.

AAAI Conference 2026 Conference Paper

DeepTracer: Tracing Stolen Model via Deep Coupled Watermarks

  • Yunfei Yang
  • Xiaojun Chen
  • Yuexin Xuan
  • Zhendong Zhao
  • Xin Zhao
  • He Li

Model watermarking techniques can embed watermark information into the protected model for ownership declaration by constructing specific input-output pairs. However, existing watermarks are easily removed when facing model stealing attacks, and make it difficult for model owners to effectively verify the copyright of stolen models. In this paper, we analyze the root cause of the failure of current watermarking methods under model stealing scenarios and then explore potential solutions. Specifically, we introduce a robust watermarking framework, DeepTracer, which leverages a novel watermark samples construction method and a same-class coupling loss constraint. DeepTracer can incur a high-coupling model between watermark task and primary task that makes adversaries inevitably learn the hidden watermark task when stealing the primary task functionality. Furthermore, we propose an effective watermark samples filtering mechanism that elaborately select watermark key samples used in model ownership verification to enhance the reliability of watermarks. Extensive experiments across multiple datasets and models demonstrate that our method surpasses existing approaches in defending against various model stealing attacks, as well as watermark attacks, and achieves new state-of-the-art effectiveness and robustness.

AAAI Conference 2026 Conference Paper

Hide and Seek with LLMs: An Adversarial Game for Sneaky Error Generation and Self-Improving Diagnosis

  • Rui Zou
  • Mengqi Wei
  • Yutao Zhu
  • Jirong Wen
  • Xin Zhao
  • Jing Chen

Large Language Models (LLMs) excel in reasoning and generation across domains, but still struggle with identifying and diagnosing complex errors. This stems mainly from training objectives that prioritize correct answers, limiting exposure to and learning from errors. While recent studies have begun to address this by introducing error signals, most rely on shallow, static errors, restricting improvement in deep diagnostic ability. To overcome this, we propose Hide and Seek Game (HSG), a dynamic adversarial framework for error generation and diagnosis, and evaluate it on mathematical problem-solving. HSG involves two adversarial roles: Sneaky, which hides by generating subtle, deceptive reasoning errors, and Diagnosis, which seeks to accurately detect them. Through adversarial co-evolution, both error stealth and diagnostic precision are enhanced. Experiments on three mathematical reasoning datasets demonstrate that HSG significantly boosts error diagnosis, achieving 16.8%-31.4% higher accuracy than baselines like GPT-4o. We also release a challenging dataset of deceptive errors and diagnostic annotations as a benchmark for future research.

EAAI Journal 2026 Journal Article

Multi-vehicle cooperative decision-making for ramp merging in mixed traffic: A pre-trial behavior-informed reinforcement learning approach

  • Qingsong Wei
  • Xianyi Xie
  • Lisheng Jin
  • Xin Zhao
  • Yewei Shi
  • Guofeng Luo
  • Yaping Liao

In mixed-traffic environments, the non-cooperative behavior of human-driven vehicles (HDVs) undermines the ability of autonomous vehicles (AVs) to make safe and timely decisions. Existing methods enhance individual AV responses but often fall short in ensuring robust cooperation among multiple agents in the presence of human-induced uncertainty. To address this, this paper proposes a Multi-Agent Proximal Policy Optimization framework with Pre-Trial Behavior Adjustment (MAPPO-PTBA), contributing to AI by enhancing AV coordination and decision-making in dynamic mixed-traffic environments. From an engineering perspective, the framework is applied to a highway ramp merging scenario, significantly improving AV performance in safety and efficiency. Based on the MAPPO algorithm, the framework integrates a Pre-Trial Behavior Adjustment (PTBA) model and a Multi-Dimensional Driving Behavior Assessment (MDBA) module. The PTBA model simulates exploratory trajectories and utilizes future states of surrounding vehicles to proactively filter unsafe actions, identifying potential conflicts under varying traffic conditions. The MDBA module extracts operational features from naturalistic driving data to dynamically assess AV behavior, enabling an adaptive action-filtering mechanism that discards unsafe actions while retaining feasible ones in interactions with uncooperative HDVs. Simulation results show MAPPO-PTBA achieves success rates of 100%, 97%, and 94% under low, medium, and high traffic densities, respectively, and improves average vehicle speed by up to 8. 26% compared to baselines. The framework significantly enhances AV merging safety and efficiency through coordinated decision-making, predictive exploration, and adaptive safety assessments.

EAAI Journal 2026 Journal Article

Physics-informed deep learning for predictive risk perception in tunnel construction

  • Penghui Lin
  • Jilang Wang
  • Limao Zhang
  • Robert L.K. Tiong
  • Cheng Meng
  • Xin Zhao

Ensuring a safe and stable excavation process is essential in tunnel construction, particularly when using Earth Pressure Balance (EPB) Tunnel Boring Machines (TBMs). This study aims to enhance the prediction of soil chamber pressure (SCP), a key parameter for maintaining ground stability, by developing a physics-informed deep learning model. The proposed physics-informed deep neural Network (PDNN) embeds an ordinary differential equation (ODE) representing the pressure balance mechanism into the model's loss function, ensuring physical consistency. The PDNN model is evaluated against other advanced deep learning models using real-world TBM data. Results show that the PDNN achieves a high predictive accuracy, with the coefficient of determination ( R 2 ) values of 0. 97 and 0. 96 on training and testing data, respectively, while demonstrating strong generalization under small datasets and multi-step forecasting conditions. By incorporating physics-based constraints, the model improves both interpretability and reliability, as further validated through SHapley Additive explanation (SHAP) analysis. This work represents a novel and effective application of physics-informed machine learning in tunnel construction, bridging the gap between data-driven modeling and engineering domain knowledge to support proactive safety risk management.

AAAI Conference 2026 Conference Paper

Reasoning with Exploration: An Entropy Perspective

  • Daixuan Cheng
  • Shaohan Huang
  • Xuekai Zhu
  • Bo Dai
  • Xin Zhao
  • Zhenliang Zhang
  • Furu Wei

Balancing exploration and exploitation is a central goal in reinforcement learning (RL). Despite recent advances in enhancing language model (LM) reasoning, most methods lean toward exploitation, and increasingly encounter performance plateaus. In this work, we revisit entropy -- a signal of exploration in RL -- and examine its relationship to exploratory reasoning in LMs. Through empirical analysis, we uncover positive correlations between high-entropy regions and three types of exploratory reasoning actions: (1) pivotal tokens that determine or connect logical steps, (2) reflective actions such as self-verification and correction, and (3) rare behaviors under-explored by the base LMs. Motivated by this, we introduce a minimal modification to standard RL with only one line of code: augmenting the advantage function with an entropy-based term. Unlike traditional maximum-entropy methods which encourage exploration by promoting uncertainty, we encourage exploration by promoting deeper and longer reasoning chains. Notably, our method achieves significant gains on the Pass@K metric -- an upper-bound estimator of LM reasoning capabilities -- even when evaluated with extremely large K values, pushing the boundaries of LM reasoning.

AAAI Conference 2026 Conference Paper

Towards Effective Code-Integrated Reasoning

  • Fei Bai
  • Yingqian Min
  • Beichen Zhang
  • Zhipeng Chen
  • Xin Zhao
  • Lei Fang
  • Zheng Liu
  • Zhongyuan Wang

In this paper, we investigate code-integrated reasoning (CIR), where models generate code when necessary and integrate feedback by executing it through a code interpreter. To acquire this capability, models must learn when and how to use external code tools effectively, which is supported by tool-augmented reinforcement learning (RL). Despite its benefits, tool-augmented RL can still suffer from potential instability in the learning dynamics. In light of this challenge, we present a systematic approach ETIR (Effective TIR) to improving the training effectiveness and stability of tool-augmented RL for code-integrated reasoning. Specifically, we develop enhanced training strategies that balance exploration and stability, progressively building tool-use capabilities while improving reasoning performance. Through extensive experiments on five mainstream mathematical reasoning benchmarks, our model demonstrates significant performance improvements over multiple competitive baselines. Furthermore, we conduct an in-depth analysis of the mechanism of code-integrated reasoning, revealing several key insights, such as the extension of model’s capability boundaries and the simultaneous improvement of reasoning efficiency through code integration. These findings underscore the potential of code-integrated reasoning as a scalable paradigm for advancing robust and efficient language model reasoning.

AAAI Conference 2026 Conference Paper

Value-Aligned Prompt Moderation via Zero-Shot Agentic Rewriting for Safe Image Generation

  • Xin Zhao
  • Xiaojun Chen
  • Bingshan Liu
  • Zeyao Liu
  • Zhendong Zhao
  • Xiaoyan Gu

Generative vision-language models like Stable Diffusion demonstrate remarkable capabilities in creative media synthesis, but they also pose substantial risks of producing unsafe, offensive, or culturally inappropriate content when prompted adversarially. Current defenses struggle to align outputs with human values without sacrificing generation quality or incurring high costs. To address these challenges, we introduce VALOR (Value-Aligned LLM-Overseen Rewriter), a modular, zero-shot agentic framework for safer and more helpful text-to-image generation. VALOR integrates layered prompt analysis with human-aligned value reasoning: a multi-level NSFW detector filters lexical and semantic risks; a cultural value alignment module identifies violations of social norms, legality, and representational ethics; and an intention disambiguator detects subtle or indirect unsafe implications. When unsafe content is detected, prompts are selectively rewritten by a large language model under dynamic, role-specific instructions designed to preserve user intent while enforcing alignment. If the generated image still fails a safety check, VALOR optionally performs a stylistic regeneration to steer the output toward a safer visual domain without altering core semantics. Experiments across adversarial, ambiguous, and value-sensitive prompts show that VALOR significantly reduces unsafe outputs by up to 100.00% while preserving prompt usefulness and creativity. These results highlight VALOR as a scalable and effective approach for deploying safe, aligned, and helpful image generation systems in open-world settings.

IROS Conference 2025 Conference Paper

A Bio-inspired Spherical Soft Magnetic Millirobot for Gastrointestinal Applications

  • Yulin Li
  • Zhaorui Hong
  • Yuhao Zhao
  • Shuohao Zhang
  • Xin Zhao
  • Liu Yang

Gastroscopy and colonoscopy have become the fundamental tools for gastrointestinal (GI) tract diagnosis and treatment. Conventional tethered devices usually lead to the use of anesthetic agents and patient discomfort. Capsule endoscopy is becoming an ideal alternative, however, the smooth capsule shape may hinder its active locomotion and retention in the GI tract. Here, we propose a spherical soft magnetic millirobot (S 2 M 2 robot) that integrates a virus-like spherical body with protrusions and an octopus-inspired sucker design for enhanced physical capabilities. The protruding suckers (radius: 2. 2-4. 8 mm) on its surface enable efficient locomotion performance (maximum angular velocity: 8 r/s, maximum speed: 180 mm/s) with strong adhesive ability (maximum force: 3. 5 N). The ex vivo experiments in a swine stomach demonstrate the robot’s motion on the slippery surface of the gastric mucosa and the effectiveness of the pressure-based drug delivery system. The in vitro and ex vivo results highlight the superior mobility and controllability, showcasing its potential as a carrier robot for the next-generation capsule endoscopy.

EAAI Journal 2025 Journal Article

A small object detection algorithm for mine environment

  • Dong Liu
  • Xin Zhao
  • Weiqiang Fan

The detection of protective equipment carried by underground mine operators is a crucial measure for preventing safety accidents and safeguarding personal life and property. However, current challenges include low object detection accuracy and difficulty detecting small objects, we propose a small object detection algorithm based on the improved You Only Look Once Version 8 (YOLOv8) for the mine environment. To minimize the semantic gap between features at different levels and enhance the feature fusion effect, the Asymptotic Feature Pyramid Network-Four (AFPN-F) has been designed to replace the Neck component of YOLOv8, enabling the detection model to better adapt to semantic information across varying levels. To enhance the model's sensitivity to small objects in the mine environment, a superficial feature output layer has been added to the model. This addition helps to prevent the loss of small-sized objects, which may contain limited feature information, during successive convolution operations. To address the significant differences in the scales of various objects in the mine, the More Focused Intersection over Union Loss (Focaler-IoU) is introduced as a loss function. This modification is intended to improve the handling of different types of regression samples, enhance training accuracy, and ensure that the model is more focused on small objects in the mine environment. The experimental results show that the proposed model outperforms other mainstream models. Compared to the baseline model YOLOv8, achieving an improvement of 3. 7 percent in mean Average Precision (mAP), the number of parameters has been reduced by 30 percent, resulting in a model size of only 5 Megabytes. This study provides an effective solution for detecting small objects in underground mines.

EAAI Journal 2025 Journal Article

Adaptive self-evolving extreme learning machine-based terminal sliding mode control with application in retinal vein injection

  • Bo Hu
  • Shiyu Xu
  • Lu Liu
  • Rongxin Liu
  • Mingzhu Sun
  • Xin Zhao

Retinal vein occlusion (RVO) is a serious condition that can lead to blindness. Injecting drugs into the retinal vein is a promising procedure for treating RVO. Due to the fragility of the retinal tissue, maintaining a precise drug flow rate (DFR) with a fast response is critical. Considering the unknown disturbance from piston dynamic and the drug-vein interaction, an adaptive self-evolving neural terminal sliding mode (ASNTSM) controller is proposed for DFR tracking. The integral terminal sliding surface is adopted to track the desired DFR in finite-time. The extreme learning machine (ELM) is utilized to estimate overall disturbances, and the adaptive switching gain is employed to compensate for the estimation error without requiring prior bounds. To achieve a compact ELM structure, a self-evolving mechanism is designed to implement the growth or pruning strategy of the hidden neurons. Theoretical analysis has proven that the ASNTSM controller can guarantee finite-time stability. Comparative experiments are conducted using a silicon phantom with simulated blood flow disturbances. The experimental results illustrate that the ASNTSM controller not only achieves lower transient time and average steady-state error, but also exhibits lower fluctuation and chattering effect. The self-evolving mechanism enhances the practicability of neural network in artificial intelligence-based medical engineering. Therefore, the ASNTSM controller is suitable for retinal vein injection tasks to improve surgical efficiency.

AAAI Conference 2025 Conference Paper

Alleviating Shifted Distribution in Human Preference Alignment through Meta-Learning

  • Shihan Dou
  • Yan Liu
  • Enyu Zhou
  • Songyang Gao
  • Tianlong Li
  • Limao Xiong
  • Xin Zhao
  • Haoxiang Jia

The capability of the reward model (RM) is crucial for the success of Reinforcement Learning from Human Feedback (RLHF) in aligning with human preferences. However, as training progresses, the output space distribution of the policy model shifts. The RM, initially trained on responses sampled from the output distribution of the early policy model, gradually loses its ability to distinguish between responses from the newly shifted distribution. This issue is further compounded when the RM, trained on a specific data distribution, struggles to generalize to examples outside of that distribution. These two issues can be united as a challenge posed by the shifted distribution of the environment. To surmount this challenge, we introduce MetaRM, a novel method leveraging meta-learning to adapt the RM to the shifted environment distribution. MetaRM optimizes the RM in an alternating way, by preserving both the preferences of the original preference pairs, as well as maximizing discrimination power over new examples of the shifted distribution. Extensive experiments demonstrate that MetaRM can iteratively enhance the performance of human preference alignment by improving the RM's capacity to identify subtle differences in samples of shifted distributions.

NeurIPS Conference 2025 Conference Paper

Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations

  • Zican Dong
  • Han Peng
  • Peiyu Liu
  • Xin Zhao
  • Dong Wu
  • Feng Xiao
  • Zhifeng Wang

Mixture-of-Experts (MoE) models achieve a favorable trade-off between performance and inference efficiency by activating only a subset of experts. However, the memory overhead of storing all experts remains a major limitation, especially in large-scale MoE models such as DeepSeek-R1 (671B). In this study, we investigate domain specialization and expert redundancy in large-scale MoE models and uncover a consistent behavior we term~\emph{few-shot expert localization}, with only a few in-domain demonstrations, the model consistently activates a sparse and stable subset of experts on tasks within the same domain. Building on this observation, we propose a simple yet effective pruning framework, \textbf{EASY-EP}, that leverages a few domain-specific demonstrations to identify and retain only the most relevant experts. EASY-EP comprises two key components: \textbf{output-aware expert importance assessment} and \textbf{expert-level token contribution estimation}. The former evaluates the importance of each expert for the current token by considering the gating scores and L2 norm of the outputs of activated experts, while the latter assesses the contribution of tokens based on representation similarities before and after routed experts. Experiments on DeepSeek-R1 and DeepSeek-V3-0324 show that our method can achieve comparable performances and $2. 99\times$ throughput under the same memory budget as the full model, with only half the experts. Our code is available at https: //github. com/RUCAIBox/EASYEP.

NeurIPS Conference 2025 Conference Paper

ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming Contests

  • Shiyi Xu
  • Hu Yiwen
  • Yingqian Min
  • Zhipeng Chen
  • Xin Zhao
  • Ji-Rong Wen

With the significant progress of large reasoning models in complex coding and reasoning tasks, existing benchmarks, like LiveCodeBench and CodeElo, are insufficient to evaluate the coding capabilities of large language models (LLMs) in real competition environments. Moreover, current evaluation metrics such as Pass@K fail to capture the reflective abilities of reasoning models. To address these challenges, we propose ICPC-Eval, a top-level competitive coding benchmark designed to probing the frontiers of LLM reasoning. ICPC-Eval includes 118 carefully curated problems from 11 recent ICPC contests held in various regions of the world, offering three key contributions: 1) A challenging realistic ICPC competition scenario, featuring a problem type and difficulty distribution consistent with actual contests. 2) A robust test case generation method and a corresponding local evaluation toolkit, enabling efficient and accurate local evaluation. 3) An effective test-time scaling evaluation metric, Refine@K, which allows iterative repair of solutions based on execution feedback. The results underscore the significant challenge in evaluating complex reasoning abilities: top-tier reasoning models like DeepSeek-R1 often rely on multi-turn code feedback to fully unlock their in-context reasoning potential when compared to non-reasoning counterparts. Furthermore, despite recent advancements in code generation, these models still lag behind top-performing human teams. We release the benchmark at: https: //github. com/RUCAIBox/ICPC-Eval

NeurIPS Conference 2025 Conference Paper

Incentivizing Dual Process Thinking for Efficient Large Language Model Reasoning

  • Xiaoxue Cheng
  • Junyi Li
  • Zhenduo Zhang
  • Xinyu Tang
  • Xin Zhao
  • Xinyu Kong
  • Zhiqiang Zhang

Large reasoning models (LRMs) have demonstrated strong performance on complex reasoning tasks, but often suffer from overthinking, generating redundant content regardless of task difficulty. Inspired by the dual process theory in cognitive science, we propose Adaptive Cognition Policy Optimization (ACPO), a reinforcement learning framework that enables LRMs to achieve efficient reasoning through adaptive cognitive allocation and dynamic system switch. ACPO incorporates two key components: (1) introducing system-aware reasoning tokens to explicitly represent the thinking modes thereby making the model's cognitive process transparent, and (2) integrating online difficulty estimation and token length budget to guide adaptive system switch and reasoning during reinforcement learning. To this end, we propose a two-stage training strategy. The first stage begins with supervised fine-tuning to cold start the model, enabling it to generate reasoning paths with explicit thinking modes. In the second stage, we apply ACPO to further enhance adaptive system switch for difficulty-aware reasoning. Experimental results demonstrate that ACPO effectively reduces redundant reasoning while adaptively adjusting cognitive allocation based on task complexity, achieving efficient hybrid reasoning.

NeurIPS Conference 2025 Conference Paper

Irrational Complex Rotations Empower Low-bit Optimizers

  • Zhen Tian
  • Xin Zhao
  • Ji-Rong Wen

In this paper, we propose a novel optimizer state compression algorithm, namely \textbf{$\pi$-Quant}, which leverages the properties of irrational numbers (\eg $\pi$) for memory-efficient training. The core idea is based on our mathematical findings, which show that a pair of parameters can be represented by a single rotation angle using the complex rotation scheme. Building on this insight, we map the parameters into a complex space and perform quantization using the corresponding rotation angles. To efficiently integrate it into optimization process, we develop an efficient system of geometric equations that computes the precise rotation angles with linear complexity. We evaluate $\pi$-Quant on a wide range of tasks. Our experiments show that it can reduce the bit-width of parameters to 3. 32-bit, achieving a 41. 8\% decrease in GPU memory usage, all while maintaining full accuracy. \textcolor{blue}{We have submitted the code in supplementary materials}.

IJCAI Conference 2025 Conference Paper

L2M2: A Hierarchical Framework Integrating Large Language Model and Multi-agent Reinforcement Learning

  • Minghong Geng
  • Shubham Pateria
  • Budhitama Subagdja
  • Lin Li
  • Xin Zhao
  • Ah-Hwee Tan

Multi-agent reinforcement learning (MARL) has demonstrated remarkable success in collaborative tasks, yet faces significant challenges in scaling to complex scenarios requiring sustained planning and coordination across long horizons. While hierarchical approaches help decompose these tasks, they typically rely on hand-crafted subtasks and domain-specific knowledge, limiting their generalizability. We present L2M2, a novel hierarchical framework that leverages large language models (LLMs) for high-level strategic planning and MARL for low-level execution. L2M2 enables zero-shot planning that supports both end-to-end training and direct integration with pre-trained MARL models. Experiments in the VMAS environment demonstrate that L2M2's LLM-guided MARL achieves superior performance while requiring less than 20% of the training samples compared to baseline methods. In the MOSMAC environment, L2M2 demonstrates strong performance with pre-defined subgoals and maintains substantial effectiveness without subgoals - scenarios where baseline methods consistently fail. Analysis through kernel density estimation reveals L2M2's ability to automatically generate appropriate navigation plans, demonstrating its potential for addressing complex multi-agent coordination tasks.

IROS Conference 2025 Conference Paper

Modeling and Simulation of Single-micropipette Cell Rotation for Imitation Learning

  • Zefu Wang
  • Yuchen Hua
  • Huiying Gong
  • Yujie Zhang
  • Zhanli Yang
  • Yaowei Liu
  • Xin Zhao
  • Mingzhu Sun

Cell rotation plays a crucial role in micromanipulation. Among manual cell rotation techniques, single-micropipette cell rotation is widely adopted due to its high efficiency and flexibility. However, there is currently no method capable of achieving automated single-micropipette cell rotation. In this study, we developed the first three-dimensional (3D) simulation system for single-micropipette cell rotation. Based on this simulation system, we successfully achieved single-micropipette cell rotation imitation learning (IL) for the first time. Specifically, we first analyze the forces acting on cells in the fluid, establishing a dynamic model that describes the cell's behavior in response to the flow velocity at the holding micropipette's orifice, the relative position of the micropipette, and time. We then developed the cell rotation simulation environment by discretizing the model and designing the simulation's cell and holding micropipette models based on real-world conditions. Finally, we designed a network architecture for IL using this model, achieving single-micropipette cell rotation in simulation. The results demonstrate that the simulation system exhibits a relative error range of 5. 34% to 12. 21% compared to real-world experiments, indicating a high degree of accuracy. Additionally, the single-micropipette cell rotation task achieved a success rate of 69% with an average completion time of 17. 13 seconds, closely matching the expert data's average time of 17. 69 seconds, confirming the feasibility of the simulation system.

NeurIPS Conference 2025 Conference Paper

Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts Transformers

  • Xin Zhao
  • Xiaojun Chen
  • Bingshan Liu
  • Haoyu Gao
  • Zhendong Zhao
  • Yilong Chen

Large language models (LLMs) with Mixture-of-Experts (MoE) architectures achieve impressive performance and efficiency by dynamically routing inputs to specialized subnetworks, known as experts. However, this sparse routing mechanism inherently exhibits task preferences due to expert specialization, introducing a new and underexplored vulnerability to backdoor attacks. In this work, we investigate the feasibility and effectiveness of injecting backdoors into MoE-based LLMs by exploiting their inherent expert routing preferences. We thus propose \textbf{BadSwitch}, a novel backdoor framework that integrates task-coupled dynamic trigger optimization with a sensitivity-guided Top-S expert tracing mechanism. Our approach jointly optimizes trigger embeddings during pretraining while identifying S most sensitive experts, subsequently constraining the Top-K gating mechanism to these targeted experts. Unlike traditional backdoor attacks that rely on superficial data poisoning or model editing, BadSwitch primarily embeds malicious triggers into expert routing paths with strong task affinity, enabling precise and stealthy model manipulation. Through comprehensive evaluations across three prominent MoE architectures (Switch Transformer, QwenMoE, and DeepSeekMoE), we demonstrate that BadSwitch can efficiently hijack pre-trained models with up to 100\% success rate (ASR) while maintaining the highest clean accuracy (ACC) among all baselines. Furthermore, BadSwitch exhibits strong resilience against both text-level and model-level defense mechanisms, achieving 94. 07\% ASR and 87. 18\% ACC on the AGNews dataset. Our analysis of expert activation patterns reveals fundamental insights into MoE vulnerabilities. We anticipate this work will expose security risks in MoE systems and contribute to advancing AI safety.

EAAI Journal 2024 Journal Article

A least squares–support vector machine for learning solution to multi-physical transient-state field coupled problems

  • Xiaoming Han
  • Xin Zhao
  • Yecheng Wu
  • Zhengwei Qu
  • Guofeng Li

The least squares–support vector machine (LS-SVM) method has achieved remarkable success in solving electromagnetic equations. However, the boundaries of the entire computational domain for solving multi-physical transient-state field coupled problems are varied. The shape functions used in mesh-based methods (such as the finite element method and the finite volume method) are constructed on meshes, so it is difficult to obtain an accurate solution using mesh-based methods. To overcome this disadvantage of mesh-based methods, the LS-SVM method is presented in this paper for solving multi-physical transient-state field coupled problems. First, the time step of the transient field is iterated by the Crank–Nicolson (C-N) method. Following that, the Karush–Kuhn–Tucker (KKT) optimality conditions are used, and the quadratic programming problem is transformed into the solution of a system of equations. Finally, an immune algorithm is used to determine the shape parameters, and the accuracy of the solution is improved. The efficiency of the LS-SVM method was demonstrated by solving a two-dimensional transient-state electrothermal coupled problem and a two-dimensional transient-state electromagnetic–fluid coupled problem. The method was compared with the finite element method (or finite volume method), and the same order of calculation accuracy was obtained by the LS-SVM method. Compared to the physics-informed neural network, a more accurate solution was obtained and shorter computation times were required by the LS-SVM method.

AAAI Conference 2024 Conference Paper

DDAE: Towards Deep Dynamic Vision BERT Pretraining

  • Honghao Chen
  • Xiangwen Kong
  • Xiangyu Zhang
  • Xin Zhao
  • Kaiqi Huang

Recently, masked image modeling (MIM) has demonstrated promising prospects in self-supervised representation learning. However, existing MIM frameworks recover all masked patches equivalently, ignoring that the reconstruction difficulty of different patches can vary sharply due to their diverse distance from visible patches. In this paper, we propose a novel deep dynamic supervision to enable MIM methods to dynamically reconstruct patches with different degrees of difficulty at different pretraining phases and depths of the model. Our deep dynamic supervision helps to provide more locality inductive bias for ViTs especially in deep layers, which inherently makes up for the absence of local prior for self-attention mechanism. Built upon the deep dynamic supervision, we propose Deep Dynamic AutoEncoder (DDAE), a simple yet effective MIM framework that utilizes dynamic mechanisms for pixel regression and feature self-distillation simultaneously. Extensive experiments across a variety of vision tasks including ImageNet classification, semantic segmentation on ADE20K and object detection on COCO demonstrate the effectiveness of our approach.

AAMAS Conference 2024 Conference Paper

JDRec: Practical Actor-Critic Framework for Online Combinatorial Recommender System

  • Xin Zhao
  • Jiaxin Li
  • Zhiwei Fang
  • Yuchen Guo
  • Jinyuan Zhao
  • Jie He
  • Wenlong Chen
  • Changping Peng

In the realm of online recommendation systems, the Combinatorial Recommender (CR) system stands out for its unique approach. It presents users with a list of items on a result page, where user behavior is simultaneously influenced by contextual information and the items listed. Formulated as a combinatorial optimization problem, the objective of the CR system is to maximize the recommendation reward across the entire list of items. Despite the significant potential of CR systems, developing a practical and efficient model remains substantial challenges. These challenges stem from the dynamic nature of online environments and the pressing need for personalized recommendations. To tackle these challenges, we decompose the overarching problem into two sub-problems: list generation and list evaluation. We propose novel and pragmatic model architectures for each sub-problem aiming to concurrently enhance both effectiveness and efficiency. To further adapt the CR system to online scenarios, we integrate a bootstrap algorithm into an actor-critic reinforcement framework. This innovative approach called JD Recommender System (JDRec) is designed to continuously refine the recommendation mode through sustained user interaction, ensuring the system’s adaptability and relevance. The proposed JDRec framework, tested through rigorous offline and online experiments, has shown promising results. It has been successfully deployed in online JD recommendation systems, yielding a notable improvement in click-through rate by 2. 6% and augmenting the total value of the platform by 5. 03%. Besides, we release the large scale dataset used in our work to facilitate further research. This work is licensed under a Creative Commons Attribution International 4. 0 License. *Equal contribution. Proc. of the 23rd International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2024), N. Alechina, V. Dignum, M. Dastani, J. S. Sichman (eds.), May 6 – 10, 2024, Auckland, New Zealand. © 2024 International Foundation for Autonomous Agents and Multiagent Systems (www. ifaamas. org).

YNIMG Journal 2024 Journal Article

Neural correlates of working memory training: An fMRI meta-analysis

  • Yao Zhang
  • Junjun Fu
  • Xin Zhao

Working memory (WM) can be improved by cognitive training. Numerous studies examined neural mechanisms underlying WM training, although with differing conclusions. Therefore, we conducted a meta-analysis to examine the neural substrates underlying WM training in healthy adults. Findings from global analyses showed substantial neural changes in the frontoparietal and subcortical regions. Results from training dosage analyses of WM training showed that shorter WM training could produce neural changes in the frontoparietal regions, whereas longer WM training could produce changes in the subcortical regions (striatum, anterior cingulate cortex, and insula). WM training-induced neural changes were also moderated by the type of training task, with updating tasks inducing neural changes in more regions than maintenance tasks. Overall, these results indicate that the neural changes associated with WM training occur in the frontoparietal network and dopamine-related brain areas, extending previous meta-analyses on WM training and advancing our understanding of the neural underpinnings of WM training effects.

NeurIPS Conference 2023 Conference Paper

A Multi-modal Global Instance Tracking Benchmark (MGIT): Better Locating Target in Complex Spatio-temporal and Causal Relationship

  • Shiyu Hu
  • Dailing Zhang
  • wu meiqi
  • Xiaokun Feng
  • Xuchen Li
  • Xin Zhao
  • Kaiqi Huang

Tracking an arbitrary moving target in a video sequence is the foundation for high-level tasks like video understanding. Although existing visual-based trackers have demonstrated good tracking capabilities in short video sequences, they always perform poorly in complex environments, as represented by the recently proposed global instance tracking task, which consists of longer videos with more complicated narrative content. Recently, several works have introduced natural language into object tracking, desiring to address the limitations of relying only on a single visual modality. However, these selected videos are still short sequences with uncomplicated spatio-temporal and causal relationships, and the provided semantic descriptions are too simple to characterize video content. To address these issues, we (1) first propose a new multi-modal global instance tracking benchmark named MGIT. It consists of 150 long video sequences with a total of 2. 03 million frames, aiming to fully represent the complex spatio-temporal and causal relationships coupled in longer narrative content. (2) Each video sequence is annotated with three semantic grains (i. e. , action, activity, and story) to model the progressive process of human cognition. We expect this multi-granular annotation strategy can provide a favorable environment for multi-modal object tracking research and long video understanding. (3) Besides, we execute comparative experiments on existing multi-modal object tracking benchmarks, which not only explore the impact of different annotation methods, but also validate that our annotation method is a feasible solution for coupling human understanding into semantic labels. (4) Additionally, we conduct detailed experimental analyses on MGIT, and hope the explored performance bottlenecks of existing algorithms can support further research in multi-modal object tracking. The proposed benchmark, experimental results, and toolkit will be released gradually on http: //videocube. aitestunion. com/.

IROS Conference 2023 Conference Paper

Disentangled Discriminator for Unsupervised Domain Adaptation on Object Detection

  • Yangguang Zhu
  • Ping Guo
  • Haoran Wei
  • Xin Zhao
  • Xiangbin Wu

Object detection plays an important role in computer vision tasks such as autonomous driving, robotics, etc. Typically, a detection model is firstly trained on collected data and then deployed in real world. However, the discrepancy exists between training (source) and testing (target) data, which degrades the detection model's performance in the real world. To mitigate the negative effects, Unsupervised Domain Adaptation (UDA) methods learn the features of a shared domain via a discriminator. However, existing discriminators consider only the in-distribution adversarial learning, which ignore the out-of-distribution data of individual domains. In this paper, we propose a disentangled discriminator to consider the in-distribution and outliers separately. It aligns the source and target data with split branches under a gated strategy. We combine the disentangled discriminator with a Teacher-Student (T-S) framework that trains the student using labeled source data and unlabeled target data under a self-training mechanism. Specifically, the teacher network, that is updated with the parameters of student network via the exponential moving average, predicts pseudo labels for unlabeled data. The quality of pseudo labels can be improved after alleviating the domain discrepancy thanks to the disentangled discriminator. Extensive experiments on benchmarks demonstrate the superiority of the proposed method. Specifically, we achieve 53. 9% mAP on Foggy Cityscapes, which is 7. 2% higher than the Oracle.

NeurIPS Conference 2023 Conference Paper

Evaluating and Improving Tool-Augmented Computation-Intensive Math Reasoning

  • Beichen Zhang
  • Kun Zhou
  • Xilin Wei
  • Xin Zhao
  • Jing Sha
  • Shijin Wang
  • Ji-Rong Wen

Chain-of-thought prompting (CoT) and tool augmentation have been validated in recent work as effective practices for improving large language models (LLMs) to perform step-by-step reasoning on complex math-related tasks. However, most existing math reasoning datasets may not be able to fully evaluate and analyze the ability of LLMs in manipulating tools and performing reasoning, as they often only require very few invocations of tools or miss annotations for evaluating intermediate reasoning steps, thus supporting only outcome evaluation. To address the issue, we construct CARP, a new Chinese dataset consisting of 4, 886 computation-intensive algebra problems with formulated annotations on intermediate steps, facilitating the evaluation of the intermediate reasoning process. In CARP, we test four LLMs with CoT prompting, and find that they are all prone to make mistakes at the early steps of the solution, leading to incorrect answers. Based on this finding, we propose a new approach that can facilitate the deliberation on reasoning steps with tool interfaces, namely DELI. In DELI, we first initialize a step-by-step solution based on retrieved exemplars, then iterate two deliberation procedures that check and refine the intermediate steps of the generated solution, from both tool manipulation and natural language reasoning perspectives, until solutions converge or the maximum iteration is achieved. Experimental results on CARP and six other datasets show that the proposed DELI mostly outperforms competitive baselines, and can further boost the performance of existing CoT methods. Our data and code are available at https: //github. com/RUCAIBox/CARP.

NeurIPS Conference 2023 Conference Paper

REASONER: An Explainable Recommendation Dataset with Comprehensive Labeling Ground Truths

  • Xu Chen
  • Jingsen Zhang
  • Lei Wang
  • Quanyu Dai
  • Zhenhua Dong
  • Ruiming Tang
  • Rui Zhang
  • Li Chen

Explainable recommendation has attracted much attention from the industry and academic communities. It has shown great potential to improve the recommendation persuasiveness, informativeness and user satisfaction. In the past few years, while a lot of promising explainable recommender models have been proposed, the datasets used to evaluate them still suffer from several limitations, for example, the explanation ground truths are not labeled by the real users, the explanations are mostly single-modal and around only one aspect. To bridge these gaps, in this paper, we build a new explainable recommendation dataset, which, to our knowledge, is the first contribution that provides a large amount of real user labeled multi-modal and multi-aspect explaination ground truths. In specific, we firstly develop a video recommendation platform, where a series of questions around the recommendation explainability are carefully designed. Then, we recruit about 3000 high-quality labelers with different backgrounds to use the system, and collect their behaviors and feedback to our questions. In this paper, we detail the construction process of our dataset and also provide extensive analysis on its characteristics. In addition, we develop a library, where ten well-known explainable recommender models are implemented in a unified framework. Based on this library, we build several benchmarks for different explainable recommendation tasks. At last, we present many new opportunities brought by our dataset, which are expected to promote the field of explainable recommendation. Our dataset, library and the related documents have been released at https: //reasoner2023. github. io/.

IJCAI Conference 2023 Conference Paper

STS-GAN: Can We Synthesize Solid Texture with High Fidelity from Arbitrary 2D Exemplar?

  • Xin Zhao
  • Jifeng Guo
  • Lin Wang
  • Fanqi Li
  • Jiahao Li
  • Junteng Zheng
  • Bo Yang

Solid texture synthesis (STS), an effective way to extend a 2D exemplar to a 3D solid volume, exhibits advantages in computational photography. However, existing methods generally fail to accurately learn arbitrary textures, which may result in the failure to synthesize solid textures with high fidelity. In this paper, we propose a novel generative adversarial nets-based framework (STS-GAN) to extend the given 2D exemplar to arbitrary 3D solid textures. In STS-GAN, multi-scale 2D texture discriminators evaluate the similarity between the given 2D exemplar and slices from the generated 3D texture, promoting the 3D texture generator synthesizing realistic solid textures. Finally, experiments demonstrate that the proposed method can generate high-fidelity solid textures with similar visual characteristics to the 2D exemplar.

AAAI Conference 2023 Conference Paper

Unsupervised Domain Adaptation for Medical Image Segmentation by Selective Entropy Constraints and Adaptive Semantic Alignment

  • Wei Feng
  • Lie Ju
  • Lin Wang
  • Kaimin Song
  • Xin Zhao
  • Zongyuan Ge

Generalizing a deep learning model to new domains is crucial for computer-aided medical diagnosis systems. Most existing unsupervised domain adaptation methods have made significant progress in reducing the domain distribution gap through adversarial training. However, these methods may still produce overconfident but erroneous results on unseen target images. This paper proposes a new unsupervised domain adaptation framework for cross-modality medical image segmentation. Specifically, We first introduce two data augmentation approaches to generate two sets of semantics-preserving augmented images. Based on the model's predictive consistency on these two sets of augmented images, we identify reliable and unreliable pixels. We then perform a selective entropy constraint: we minimize the entropy of reliable pixels to increase their confidence while maximizing the entropy of unreliable pixels to reduce their confidence. Based on the identified reliable and unreliable pixels, we further propose an adaptive semantic alignment module which performs class-level distribution adaptation by minimizing the distance between same class prototypes between domains, where unreliable pixels are removed to derive more accurate prototypes. We have conducted extensive experiments on the cross-modality cardiac structure segmentation task. The experimental results show that the proposed method significantly outperforms the state-of-the-art comparison algorithms. Our code and data are available at https://github.com/fengweie/SE_ASA.

NeurIPS Conference 2022 Conference Paper

InsPro: Propagating Instance Query and Proposal for Online Video Instance Segmentation

  • Fei He
  • Haoyang Zhang
  • Naiyu Gao
  • Jian Jia
  • Yanhu Shan
  • Xin Zhao
  • Kaiqi Huang

Video instance segmentation (VIS) aims at segmenting and tracking objects in videos. Prior methods typically generate frame-level or clip-level object instances first and then associate them by either additional tracking heads or complex instance matching algorithms. This explicit instance association approach increases system complexity and fails to fully exploit temporal cues in videos. In this paper, we design a simple, fast and yet effective query-based framework for online VIS. Relying on an instance query and proposal propagation mechanism with several specially developed components, this framework can perform accurate instance association implicitly. Specifically, we generate frame-level object instances based on a set of instance query-proposal pairs propagated from previous frames. This instance query-proposal pair is learned to bind with one specific object across frames through conscientiously developed strategies. When using such a pair to predict an object instance on the current frame, not only the generated instance is automatically associated with its precursors on previous frames, but the model gets a good prior for predicting the same object. In this way, we naturally achieve implicit instance association in parallel with segmentation and elegantly take advantage of temporal clues in videos. To show the effectiveness of our method InsPro, we evaluate it on two popular VIS benchmarks, i. e. , YouTube-VIS 2019 and YouTube-VIS 2021. Without bells-and-whistles, our InsPro with ResNet-50 backbone achieves 43. 2 AP and 37. 6 AP on these two benchmarks respectively, outperforming all other online VIS methods.

NeurIPS Conference 2022 Conference Paper

Prompt Certified Machine Unlearning with Randomized Gradient Smoothing and Quantization

  • Zijie Zhang
  • Yang Zhou
  • Xin Zhao
  • Tianshi Che
  • Lingjuan Lyu

The right to be forgotten calls for efficient machine unlearning techniques that make trained machine learning models forget a cohort of data. The combination of training and unlearning operations in traditional machine unlearning methods often leads to the expensive computational cost on large-scale data. This paper presents a prompt certified machine unlearning algorithm, PCMU, which executes one-time operation of simultaneous training and unlearning in advance for a series of machine unlearning requests, without the knowledge of the removed/forgotten data. First, we establish a connection between randomized smoothing for certified robustness on classification and randomized smoothing for certified machine unlearning on gradient quantization. Second, we propose a prompt certified machine unlearning model based on randomized data smoothing and gradient quantization. We theoretically derive the certified radius R regarding the data change before and after data removals and the certified budget of data removals about R. Last but not least, we present another practical framework of randomized gradient smoothing and quantization, due to the dilemma of producing high confidence certificates in the first framework. We theoretically demonstrate the certified radius R' regarding the gradient change, the correlation between two types of certified radii, and the certified budget of data removals about R'.

AAAI Conference 2022 Conference Paper

QueryProp: Object Query Propagation for High-Performance Video Object Detection

  • Fei He
  • Naiyu Gao
  • Jian Jia
  • Xin Zhao
  • Kaiqi Huang

Video object detection has been an important yet challenging topic in computer vision. Traditional methods mainly focus on designing the image-level or box-level feature propagation strategies to exploit temporal information. This paper argues that with a more effective and efficient feature propagation framework, video object detectors can gain improvement in terms of both accuracy and speed. For this purpose, this paper studies object-level feature propagation, and proposes an object query propagation (QueryProp) framework for high-performance video object detection. The proposed QueryProp contains two propagation strategies: 1) query propagation is performed from sparse key frames to dense non-key frames to reduce the redundant computation on nonkey frames; 2) query propagation is performed from previous key frames to the current key frame to improve feature representation by temporal context modeling. To further facilitate query propagation, an adaptive propagation gate is designed to achieve flexible key frame selection. We conduct extensive experiments on the ImageNet VID dataset. QueryProp achieves comparable accuracy with state-of-the-art methods and strikes a decent accuracy/speed trade-off.

ICML Conference 2021 Conference Paper

Expressive 1-Lipschitz Neural Networks for Robust Multiple Graph Learning against Adversarial Attacks

  • Xin Zhao
  • Zeru Zhang
  • Zijie Zhang 0001
  • Lingfei Wu 0001
  • Jiayin Jin
  • Yang Zhou 0001
  • Ruoming Jin
  • Dejing Dou

Recent findings have shown multiple graph learning models, such as graph classification and graph matching, are highly vulnerable to adversarial attacks, i. e. small input perturbations in graph structures and node attributes can cause the model failures. Existing defense techniques often defend specific attacks on particular multiple graph learning tasks. This paper proposes an attack-agnostic graph-adaptive 1-Lipschitz neural network, ERNN, for improving the robustness of deep multiple graph learning while achieving remarkable expressive power. A K_l-Lipschitz Weibull activation function is designed to enforce the gradient norm as K_l at layer l. The nearest matrix orthogonalization and polar decomposition techniques are utilized to constraint the weight norm as 1/K_l and make the norm-constrained weight close to the original weight. The theoretical analysis is conducted to derive lower and upper bounds of feasible K_l under the 1-Lipschitz constraint. The combination of norm-constrained weight and activation function leads to the 1-Lipschitz neural network for expressive and robust multiple graph learning.

ICML Conference 2021 Conference Paper

Integrated Defense for Resilient Graph Matching

  • Jiaxiang Ren 0001
  • Zijie Zhang 0001
  • Jiayin Jin
  • Xin Zhao
  • Sixing Wu
  • Yang Zhou 0001
  • Yelong Shen
  • Tianshi Che

A recent study has shown that graph matching models are vulnerable to adversarial manipulation of their input which is intended to cause a mismatching. Nevertheless, there is still a lack of a comprehensive solution for further enhancing the robustness of graph matching against adversarial attacks. In this paper, we identify and study two types of unique topology attacks in graph matching: inter-graph dispersion and intra-graph assembly attacks. We propose an integrated defense model, IDRGM, for resilient graph matching with two novel defense techniques to defend against the above two attacks simultaneously. A detection technique of inscribed simplexes in the hyperspheres consisting of multiple matched nodes is proposed to tackle inter-graph dispersion attacks, in which the distances among the matched nodes in multiple graphs are maximized to form regular simplexes. A node separation method based on phase-type distribution and maximum likelihood estimation is developed to estimate the distribution of perturbed graphs and separate the nodes within the same graphs over a wide space, for defending intra-graph assembly attacks, such that the interference from the similar neighbors of the perturbed nodes is significantly reduced. We evaluate the robustness of our IDRGM model on real datasets against state-of-the-art algorithms.

JBHI Journal 2021 Journal Article

Synergic Adversarial Label Learning for Grading Retinal Diseases via Knowledge Distillation and Multi-Task Learning

  • Lie Ju
  • Xin Wang
  • Xin Zhao
  • Huimin Lu
  • Dwarikanath Mahapatra
  • Paul Bonnington
  • Zongyuan Ge

The need for comprehensive and automated screening methods for retinal image classification has long been recognized. Well-qualified doctors annotated images are very expensive and only a limited amount of data is available for various retinal diseases such as diabetic retinopathy (DR) and age-related macular degeneration (AMD). Some studies show that some retinal diseases such as DR and AMD share some common features like haemorrhages and exudation but most classification algorithms only train those disease models independently when the only single label for one image is available. Inspired by multi-task learning where additional monitoring signals from various sources is beneficial to train a robust model. We propose a method called synergic adversarial label learning (SALL) which leverages relevant retinal disease labels in both semantic and feature space as additional signals and train the model in a collaborative manner using knowledge distillation. Our experiments on DR and AMD fundus image classification task demonstrate that the proposed method can significantly improve the accuracy of the model for grading diseases by 5. 91% and 3. 69% respectively. In addition, we conduct additional experiments to show the effectiveness of SALL from the aspects of reliability and interpretability in the context of medical imaging application.

NeurIPS Conference 2021 Conference Paper

Validating the Lottery Ticket Hypothesis with Inertial Manifold Theory

  • Zeru Zhang
  • Jiayin Jin
  • Zijie Zhang
  • Yang Zhou
  • Xin Zhao
  • Jiaxiang Ren
  • Ji Liu
  • Lingfei Wu

Despite achieving remarkable efficiency, traditional network pruning techniques often follow manually-crafted heuristics to generate pruned sparse networks. Such heuristic pruning strategies are hard to guarantee that the pruned networks achieve test accuracy comparable to the original dense ones. Recent works have empirically identified and verified the Lottery Ticket Hypothesis (LTH): a randomly-initialized dense neural network contains an extremely sparse subnetwork, which can be trained to achieve similar accuracy to the former. Due to the lack of theoretical evidence, they often need to run multiple rounds of expensive training and pruning over the original large networks to discover the sparse subnetworks with low accuracy loss. By leveraging dynamical systems theory and inertial manifold theory, this work theoretically verifies the validity of the LTH. We explore the possibility of theoretically lossless pruning as well as one-time pruning, compared with existing neural network pruning and LTH techniques. We reformulate the neural network optimization problem as a gradient dynamical system and reduce this high-dimensional system onto inertial manifolds to obtain a low-dimensional system regarding pruned subnetworks. We demonstrate the precondition and existence of pruned subnetworks and prune the original networks in terms of the gap in their spectrum that make the subnetworks have the smallest dimensions.

AAAI Conference 2020 Conference Paper

GlobalTrack: A Simple and Strong Baseline for Long-Term Tracking

  • Lianghua Huang
  • Xin Zhao
  • Kaiqi Huang

A key capability of a long-term tracker is to search for targets in very large areas (typically the entire image) to handle possible target absences or tracking failures. However, currently there is a lack of such a strong baseline for global instance search. In this work, we aim to bridge this gap. Specifically, we propose GlobalTrack, a pure global instance search based tracker that makes no assumption on the temporal consistency of the target's positions and scales. GlobalTrack is developed based on two-stage object detectors, and it is able to perform full-image and multi-scale search of arbitrary instances with only a single query as the guide. We further propose a cross-query loss to improve the robustness of our approach against distractors. With no online learning, no punishment on position or scale changes, no scale smoothing and no trajectory refinement, our pure global instance search based tracker achieves comparable, sometimes much better performance on four large-scale tracking benchmarks (i.e., 52.1% AUC on LaSOT, 63.8% success rate on TLP, 60.3% MaxGM on OxUvA and 75.4% normalized precision on TrackingNet), compared to state-of-the-art approaches that typically require complex post-processing. More importantly, our tracker runs without cumulative errors, i.e., any type of temporary tracking failures will not affect its performance on future frames, making it ideal for long-term tracking. We hope this work will be a strong baseline for long-term tracking and will stimulate future works in this area.

AAAI Conference 2020 Conference Paper

TANet: Robust 3D Object Detection from Point Clouds with Triple Attention

  • Zhe Liu
  • Xin Zhao
  • Tengteng Huang
  • Ruolan Hu
  • Yu Zhou
  • Xiang Bai

In this paper, we focus on exploring the robustness of the 3D object detection in point clouds, which has been rarely discussed in existing approaches. We observe two crucial phenomena: 1) the detection accuracy of the hard objects, e. g. , Pedestrians, is unsatisfactory, 2) when adding additional noise points, the performance of existing approaches decreases rapidly. To alleviate these problems, a novel TANet is introduced in this paper, which mainly contains a Triple Attention (TA) module, and a Coarse-to-Fine Regression (CFR) module. By considering the channel-wise, point-wise and voxel-wise attention jointly, the TA module enhances the crucial information of the target while suppresses the unstable cloud points. Besides, the novel stacked TA further exploits the multi-level feature attention. In addition, the CFR module boosts the accuracy of localization without excessive computation cost. Experimental results on the validation set of KITTI dataset demonstrate that, in the challenging noisy cases, i. e. , adding additional random noisy points around each object, the presented approach goes far beyond state-of-theart approaches. Furthermore, for the 3D object detection task of the KITTI benchmark, our approach ranks the first place on Pedestrian class, by using the point clouds as the only input. The running speed is around 29 frames per second.

AAAI Conference 2020 Conference Paper

Temporal Context Enhanced Feature Aggregation for Video Object Detection

  • Fei He
  • Naiyu Gao
  • Qiaozhe Li
  • Senyao Du
  • Xin Zhao
  • Kaiqi Huang

Video object detection is a challenging task because of the presence of appearance deterioration in certain video frames. One typical solution is to aggregate neighboring features to enhance per-frame appearance features. However, such a method ignores the temporal relations between the aggregated frames, which is critical for improving video recognition accuracy. To handle the appearance deterioration problem, this paper proposes a temporal context enhanced network (TCENet) to exploit temporal context information by temporal aggregation for video object detection. To handle the displacement of the objects in videos, a novel DeformAlign module is proposed to align the spatial features from frame to frame. Instead of adopting a fixed-length window fusion strategy, a temporal stride predictor is proposed to adaptively select video frames for aggregation, which facilitates exploiting variable temporal information and requiring fewer video frames for aggregation to achieve better results. Our TCENet achieves state-of-the-art performance on the ImageNet VID dataset and has a faster runtime. Without bellsand-whistles, our TCENet achieves 80. 3% mAP by only aggregating 3 frames.

AAAI Conference 2019 Conference Paper

3D Object Detection Using Scale Invariant and Feature Reweighting Networks

  • Xin Zhao
  • Zhe Liu
  • Ruolan Hu
  • Kaiqi Huang

3D object detection plays an important role in a large number of real-world applications. It requires us to estimate the localizations and the orientations of 3D objects in real scenes. In this paper, we present a new network architecture which focuses on utilizing the front view images and frustum point clouds to generate 3D detection results. On the one hand, a PointSIFT module is utilized to improve the performance of 3D segmentation. It can capture the information from different orientations in space and the robustness to different scale shapes. On the other hand, our network obtains the useful features and suppresses the features with less information by a SENet module. This module reweights channel features and estimates the 3D bounding boxes more effectively. Our method is evaluated on both KITTI dataset for outdoor scenes and SUN-RGBD dataset for indoor scenes. The experimental results illustrate that our method achieves better performance than the state-of-the-art methods especially when point clouds are highly sparse.

IJCAI Conference 2019 Conference Paper

Pedestrian Attribute Recognition by Joint Visual-semantic Reasoning and Knowledge Distillation

  • Qiaozhe Li
  • Xin Zhao
  • Ran He
  • Kaiqi Huang

Pedestrian attribute recognition in surveillance is a challenging task in computer vision due to significant pose variation, viewpoint change and poor image quality. To achieve effective recognition, this paper presents a graph-based global reasoning framework to jointly model potential visual-semantic relations of attributes and distill auxiliary human parsing knowledge to guide the relational learning. The reasoning framework models attribute groups on a graph and learns a projection function to adaptively assign local visual features to the nodes of the graph. After feature projection, graph convolution is utilized to perform global reasoning between the attribute groups to model their mutual dependencies. Then, the learned node features are projected back to visual space to facilitate knowledge transfer. An additional regularization term is proposed by distilling human parsing knowledge from a pre-trained teacher model to enhance feature representations. The proposed framework is verified on three large scale pedestrian attribute datasets including PETA, RAP, and PA-100k. Experiments show that our method achieves state-of-the-art results.

AAAI Conference 2019 Conference Paper

Recurrent Attention Model for Pedestrian Attribute Recognition

  • Xin Zhao
  • Liufang Sang
  • Guiguang Ding
  • Jungong Han
  • Na Di
  • Chenggang Yan

Pedestrian attribute recognition is to predict attribute labels of pedestrian from surveillance images, which is a very challenging task for computer vision due to poor imaging quality and small training dataset. It is observed that many semantic pedestrian attributes to be recognised tend to show spatial locality and semantic correlations by which they can be grouped while previous works mostly ignore this phenomenon. Inspired by Recurrent Neural Network (RNN)’s super capability of learning context correlations and Attention Model’s capability of highlighting the region of interest on feature map, this paper proposes end-to-end Recurrent Convolutional (RC) and Recurrent Attention (RA) models, which are complementary to each other. RC model mines the correlations among different attribute groups with convolutional LSTM unit, while RA model takes advantage of the intra-group spatial locality and inter-group attention correlation to improve the performance of pedestrian attribute recognition. Our RA method combines the Recurrent Learning and Attention Model to highlight the spatial position on feature map and mine the attention correlations among different attribute groups to obtain more precise attention. Extensive empirical evidence shows that our recurrent model frameworks achieve state-of-the-art results, based on pedestrian attribute datasets, i. e. standard PETA and RAP datasets.

AAAI Conference 2019 Conference Paper

Visual-Semantic Graph Reasoning for Pedestrian Attribute Recognition

  • Qiaozhe Li
  • Xin Zhao
  • Ran He
  • Kaiqi Huang

Pedestrian attribute recognition in surveillance is a challenging task due to poor image quality, significant appearance variations and diverse spatial distribution of different attributes. This paper treats pedestrian attribute recognition as a sequential attribute prediction problem and proposes a novel visual-semantic graph reasoning framework to address this problem. Our framework contains a spatial graph and a directed semantic graph. By performing reasoning using the Graph Convolutional Network (GCN), one graph captures spatial relations between regions and the other learns potential semantic relations between attributes. An end-to-end architecture is presented to perform mutual embedding between these two graphs to guide the relational learning for each other. We verify the proposed framework on three large scale pedestrian attribute datasets including PETA, RAP, and PA- 100k. Experiments show superiority of the proposed method over state-of-the-art methods and effectiveness of our joint GCN structures for sequential attribute prediction.

IJCAI Conference 2018 Conference Paper

Densely Cascaded Shadow Detection Network via Deeply Supervised Parallel Fusion

  • Yupei Wang
  • Xin Zhao
  • Yin Li
  • Xuecai Hu
  • Kaiqi Huang

Shadow detection is an important and challenging problem in computer vision. Recently, single image shadow detection had achieved major progress with the development of deep convolutional networks. However, existing methods are still vulnerable to background clutters, and often fail to capture the global context of an input image. These global contextual and semantic cues are essential for accurately localizing the shadow regions. Moreover, rich spatial details are required to segment shadow regions with precise shape. To this end, this paper presents a novel model characterized by a deeply supervised parallel fusion (DSPF) network and a densely cascaded learning scheme. The DSPF network achieves a comprehensive fusion of global semantic cues and local spatial details by multiple stacked parallel fusion branches, which are learned in a deeply supervised manner. Moreover, the densely cascaded learning scheme is employed to refine the spatial details. Our method is evaluated on two widely used shadow detection benchmarks. Experimental results show that our method outperforms state-of-the-arts by a large margin.

IJCAI Conference 2018 Conference Paper

Grouping Attribute Recognition for Pedestrian with Joint Recurrent Learning

  • Xin Zhao
  • Liufang Sang
  • Guiguang Ding
  • Yuchen Guo
  • Xiaoming Jin

Pedestrian attributes recognition is to predict attribute labels of pedestrian from surveillance images, which is a very challenging task for computer vision due to poor imaging quality and small training dataset. It is observed that semantic pedestrian attributes to be recognised tend to show semantic or visual spatial correlation. Attributes can be grouped by the correlation while previous works mostly ignore this phenomenon. Inspired by Recurrent Neural Network (RNN)'s super capability of learning context correlations, this paper proposes an end-to-end Grouping Recurrent Learning (GRL) model that takes advantage of the intra-group mutual exclusion and inter-group correlation to improve the performance of pedestrian attribute recognition. Our GRL method starts with the detection of precise body region via Body Region Proposal followed by feature extraction from detected regions. These features, along with the semantic groups, are fed into RNN for recurrent grouping attribute recognition, where intra group correlations can be learned. Extensive empirical evidence shows that our GRL model achieves state-of-the-art results, based on pedestrian attribute datasets, i. e. standard PETA and RAP datasets.

AAAI Conference 2018 Conference Paper

On Trivial Solution and High Correlation Problems in Deep Supervised Hashing

  • Yuchen Guo
  • Xin Zhao
  • Guiguang Ding
  • Jungong Han

Deep supervised hashing (DSH), which combines binary learning and convolutional neural network, has attracted considerable research interests and achieved promising performance for highly efficient image retrieval. In this paper, we show that the widely used loss functions, pair-wise loss and triplet loss, suffer from the trivial solution problem and usually lead to highly correlated bits in practice, limiting the performance of DSH. One important reason is that it is difficult to incorporate proper constraints into the loss functions under the mini-batch based optimization algorithm. To tackle these problems, we propose to adopt ensemble learning strategy for deep model training. We found out that this simple strategy is capable of effectively decorrelating different bits, making the hashcodes more informative. Moreover, it is very easy to parallelize the training and support incremental model learning, which are very useful for real-world applications but usually ignored by existing DSH approaches. Experiments on benchmarks demonstrate the proposed ensemble based DSH can improve the performance of DSH approaches significant.

IJCAI Conference 2017 Conference Paper

TUCH: Turning Cross-view Hashing into Single-view Hashing via Generative Adversarial Nets

  • Xin Zhao
  • Guiguang Ding
  • Yuchen Guo
  • Jungong Han
  • Yue Gao

Cross-view retrieval, which focuses on searching images as response to text queries or vice versa, has received increasing attention recently. Cross-view hashing is to efficiently solve the cross-view retrieval problem with binary hash codes. Most existing works on cross-view hashing exploit multi-view embedding method to tackle this problem, which inevitably causes the information loss in both image and text domains. Inspired by the Generative Adversarial Nets (GANs), this paper presents a new model that is able to Turn Cross-view Hashing into single-view hashing (TUCH), thus enabling the information of image to be preserved as much as possible. TUCH is a novel deep architecture that integrates a language model network T for text feature extraction, a generator network G to generate fake images from text feature and a hashing network H for learning hashing functions to generate compact binary codes. Our architecture effectively unifies joint generative adversarial learning and cross-view hashing. Extensive empirical evidence shows that our TUCH approach achieves state-of-the-art results, especially on text to image retrieval, based on image-sentences datasets, i. e. standard IAPRTC-12 and large-scale Microsoft COCO.

IJCAI Conference 2016 Conference Paper

Semi-Supervised Multimodal Deep Learning for RGB-D Object Recognition

  • Yanhua Cheng
  • Xin Zhao
  • Rui Cai
  • Zhiwei Li
  • Kaiqi Huang
  • Yong Rui

This paper studies the problem of RGB-D object recognition. Inspired by the great success of deep convolutional neural networks (DCNN) in AI, researchers have tried to apply it to improve the performance of RGB-D object recognition. However, DCNN always requires a large-scale annotated dataset to supervise its training. Manually labeling such a large RGB-D dataset is expensive and time consuming, which prevents DCNN from quickly promoting this research area. To address this problem, we propose a semi-supervised multimodal deep learning framework to train DCNN effectively based on very limited labeled data and massive unlabeled data. The core of our framework is a novel diversity preserving co-training algorithm, which can successfully guide DCNN to learn from the unlabeled RGB-D data by making full use of the complementary cues of the RGB and depth data in object representation. Experiments on the benchmark RGB-D dataset demonstrate that, with only 5% labeled training data, our approach achieves competitive performance for object recognition compared with those state-of-the-art results reported by fully-supervised methods.

YNIMG Journal 2016 Journal Article

Use of a steady-state baseline to address evoked vs. oscillation models of visual evoked potential origin

  • Minpeng Xu
  • Yihong Jia
  • Hongzhi Qi
  • Yong Hu
  • Feng He
  • Xin Zhao
  • Peng Zhou
  • Lixin Zhang

There has been a long debate about the neural mechanism of event-related potentials (ERPs). Previously, no evidence or method was apparent to validate the two competing models, the evoked model and the oscillation model. One argument is whether the pre-stimulus brain oscillation could influence the following ERP. This study carried out an innovative visual oddball task experiment to investigate the dynamic process of visual evoked potentials. A period of stable oscillations of specified dominant frequencies and initial phases, i. e. the steady-state baseline, would be induced before responses to transient stimuli of different contrasts, which could overcome the artifact problem caused by the ‘sorting’ method. The result first revealed a ‘three-period-transition’ for the generation of visual evoked potentials by an objective decomposition. The ERP almost retained the preceding oscillation during the first period, provided an unstable negative potential in the second period, and generated the N1 component in the third period. The cross term analysis showed that the evoked model couldn't be the whole explanation for the ERP generation. Furthermore, the component analysis revealed that the N1 latency was sensitive to the initial phase under the low stimulus contrast (supporting the oscillation model) but not under the high stimulus contrast (supporting the evoked model). It demonstrated that the external stimulus contrast is a significant factor deciding the explicit model for ERPs. Our method and preliminary results may help reconcile the previous, seemly contradictory findings on the ERP mechanism.

AAAI Conference 2015 Conference Paper

Mining User Intents in Twitter: A Semi-Supervised Approach to Inferring Intent Categories for Tweets

  • Jinpeng Wang
  • Gao Cong
  • Xin Zhao
  • Xiaoming Li

In this paper, we propose to study the problem of identifying and classifying tweets into intent categories. For example, a tweet “I wanna buy a new car” indicates the user’s intent for buying a car. Identifying such intent tweets will have great commercial value among others. In particular, it is important that we can distinguish different types of intent tweets. We propose to classify intent tweets into six categories, namely Food & Drink, Travel, Career & Education, Goods & Services, Event & Activities and Trifle. We propose a semi-supervised learning approach to categorizing intent tweets into the six categories. We construct a test collection by using a bootstrap method. Our experimental results show that our approach is effective in inferring intent categories for tweets.

IROS Conference 2006 Conference Paper

Voxel-Based Modeling and Rendering for Virtual MEMS Fabrication Process

  • Guangyi Sun
  • Xin Zhao
  • Guizhang Lu

This paper puts forward a novel approach based on voxel-based modeling and volume rendering to Virtual MEMS Fabrication Process simulation. Voxel-based modeling with 3D mathematical morphology and expert system can produce realistic representation of complex MEMS device and is really stable, robust, and accurate compared to the traditional solid modeling. Additionally, volume rendering technique enables the user to fully visualize the internal structure of 3D geometry. A voxel-based modeling and visualization prototyping for Virtual MEMS Fabrication Process, called ZProcess, has been successfully developed. The capabilities of ZProcess are validated for various fabrication processes and device design.

v2026.09.13