Arrow Research search

Author name cluster

Xiaohan Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

18 papers
2 author rows

Possible papers

18

AAAI Conference 2026 Conference Paper

Modality-Balanced Collaborative Distillation for Multi-Modal Domain Generalization

  • Xiaohan Wang
  • Zhangtao Cheng
  • Ting Zhong
  • Leiting Chen
  • Fan Zhou

Weight Averaging (WA) has emerged as a powerful technique for enhancing generalization by promoting convergence to a flat loss landscape, which correlates with stronger out-of-distribution performance. However, applying WA directly to multi-modal domain generalization (MMDG) is challenging: differences in optimization speed across modalities lead WA to overfit to faster-converging ones in early stages, suppressing the contribution of slower yet complementary modalities, thereby hindering effective modality fusion and skewing the loss surface toward sharper, less generalizable minima. To address this issue, we propose MBCD, a unified collaborative distillation framework that retains WA's flatness-inducing advantages while overcoming its shortcomings in multi-modal contexts. MBCD begins with adaptive modality dropout in the student model to curb early-stage bias toward dominant modalities. A gradient consistency constraint then aligns learning signals between uni-modal branches and the fused representation, encouraging coordinated and smoother optimization. Finally, a WA-based teacher conducts cross-modal distillation by transferring fused knowledge to each uni-modal branch, which strengthens cross-modal interactions and steer convergence toward flatter solutions. Extensive experiments on MMDG benchmarks show that MBCD consistently outperforms existing methods, achieving superior accuracy and robustness across diverse unseen domains.

JBHI Journal 2025 Journal Article

Design of a Multi-Parameter Fusion Sensor and System for Respiratory Monitoring of Mechanically Ventilated Patients in the ICU

  • Shuai Ren
  • Xiaohan Wang
  • Maolin Cai
  • Yan Shi
  • Tao Wang
  • Zujin Luo

In order to achieve precise respiratory therapy for mechanically ventilated patients, real-time monitoring of the state parameters of inhaled and exhaled gases is required. These parameters are primarily measured by ventilators, with limitations such as insufficient monitoring parameters, circuit leaks, and constraints imposed by distance and obstacles. This paper designs a low-power wireless sensor for multi-parameter monitoring near the patient, which can be used continuously for approximately 60 days. Based on this sensor, an intelligent respiratory monitoring system with a distributed architecture is proposed to achieve intelligent patient-ventilator asynchrony (PVA) perception. Experimental results show that the system can stably and accurately collect and transmit data, with measurement errors for pressure, flow, temperature, humidity, and CO $_{2}$ concentration being $\pm$ 1. 3%, $\pm$ 2. 1%, $\pm$ 0. 6 $^\circ$, $\pm$ 1% RH, $\pm$ 0. 3 mmHg respectively. The proposed sensor and system have the potential to enhance the efficiency and intelligence of medical care significantly.

JBHI Journal 2025 Journal Article

Improving Patient-Ventilator Synchrony During Pressure Support Ventilation Based on Reinforcement Learning Algorithm

  • Liming Hao
  • Xiaohan Wang
  • Shuai Ren
  • Yan Shi
  • Maolin Cai
  • Tao Wang
  • Zujin Luo

Mechanical ventilation is an effective treatment for critically ill patients and those with pulmonary diseases. However, patient-ventilator asynchrony (PVA) remains a significant challenge, potentially leading to high mortality. Improving patient-ventilator synchrony poses a complex decision-making problem in clinical practice. Traditional methods rely heavily on clinicians' experience, often resulting in inefficiencies, delayed ventilator adjustments, and resource shortages. This paper proposes a novel approach using a deep reinforcement learning (RL) algorithm based on deep Q-learning (DQN) to enhance patient-ventilator synchrony during pressure support ventilation. The action space and reward function are established from clinical experience, and a pneumatic model of the mechanical ventilation system is constructed to simulate various patient conditions and types of PVAs. Clinical data are used to evaluate the RL algorithm qualitatively and quantitatively. The RL-optimized ventilation strategy reduces the proportion of breaths containing PVAs from 37. 52% to 7. 08%, demonstrating its effectiveness in assisting clinical decision-making, improving synchrony, and enabling intelligent ventilator control, bedside monitoring, and automatic weaning.

ICLR Conference 2025 Conference Paper

Innovative Thinking, Infinite Humor: Humor Research of Large Language Models through Structured Thought Leaps

  • Han Wang
  • Yilin Zhao
  • Dian Li
  • Xiaohan Wang
  • Sinbadliu
  • Xuguang Lan
  • Hui Wang

Humor is previously regarded as a gift exclusive to humans for the following reasons. Humor is a culturally nuanced aspect of human language, presenting challenges for its understanding and generation. Humor generation necessitates a multi-hop reasoning process, with each hop founded on proper rationales. Although many studies, such as those related to GPT-o1, focus on logical reasoning with reflection and correction, they still fall short in humor generation. Due to the sparsity of the knowledge graph in creative thinking, it is arduous to achieve multi-hop reasoning. Consequently, in this paper, we propose a more robust framework for addressing the humor reasoning task, named LoL. LoL aims to inject external information to mitigate the sparsity of the knowledge graph, thereby enabling multi-hop reasoning. In the first stage of LoL, we put forward an automatic instruction-evolution method to incorporate the deeper and broader thinking processes underlying humor. Judgment-oriented instructions are devised to enhance the model's judgment capability, dynamically supplementing and updating the sparse knowledge graph. Subsequently, through reinforcement learning, the reasoning logic for each online-generated response is extracted using GPT-4o. In this process, external knowledge is re-introduced to aid the model in logical reasoning and the learning of human preferences. Finally, experimental results indicate that the combination of these two processes can enhance both the model's judgment ability and its generative capacity. These findings deepen our comprehension of the creative capabilities of large language models (LLMs) and offer approaches to boost LLMs' creative abilities for cross-domain innovative applications.

ICRA Conference 2025 Conference Paper

LAFNET: Lightweight Aerial Fire Detection Model for Onboard Edge Computing

  • Haozhou Zhai
  • Weiming Yan
  • Xiaohan Wang
  • Tuhao Zhao
  • Tianjiang Hu

Fire poses significant threats to life and property, necessitating efficient inspection and accurate identification. Although aerial computer vision algorithms hold great promise, the deployment and computational limitations of onboard platforms prevent existing algorithms from meeting high standards of accuracy and real-time performance. To address these challenges, we propose an lightweight aerial fire detection model, LAFNET. This model incorporates the EffiDarknetLight backbone, optimized for both lightweight design and ease of deployment, integrates specially designed LightGhost(LG) block components within the LightGhost-Path Aggregation Network(LG-PAN) neck, resulting in a model Params of only 1. 3 M. Experimental results demonstrate that our method attains a good trade-off between lightweight design and detection accuracy. Compared to the smallest standard YOLO series' model YOLOv5n, LAFNET improves MAP by $\mathbf{2. 1 \%}$, while reducing Params and FLOPs by $\mathbf{2 7. 8 \%}$ and $\mathbf{2 9. 3 \%}$, the inference speed on Nvidia Orin Nano edge computing side improves 24. 8 %. These experiments indicate that LAFNET offers a highly efficient solution for aerial fire detection, combining speed and accuracy.

UAI Conference 2025 Conference Paper

Targeted Learning for Variable Importance

  • Xiaohan Wang
  • Yunzhe Zhou
  • Giles Hooker

Variable importance is one of the most widely used measures for interpreting machine learning with significant interest from both statistics and machine learning communities. Recently, increasing attention has been directed toward uncertainty quantification in these metrics. Current approaches largely rely on one-step procedures, which, while asymptotically efficient, can present higher sensitivity and instability in finite sample settings. To address these limitations, we propose a novel method by employing the targeted learning (TL) framework, designed to enhance robustness in inference for variable importance metrics. Our approach is particularly suited for conditional permutation variable importance. We show that it (i) retains the asymptotic efficiency of traditional methods, (ii) maintains comparable computational complexity, and (iii) delivers improved accuracy, especially in finite sample contexts. We further support these findings with numerical experiments that illustrate the practical advantages of our method and validate the theoretical results.

ICLR Conference 2025 Conference Paper

Video Action Differencing

  • James Burgess
  • Xiaohan Wang
  • Yuhui Zhang
  • Anita Rau
  • Alejandro Lozano
  • Lisa Dunlap
  • Trevor Darrell
  • Serena Yeung

How do two individuals differ when performing the same action? In this work, we introduce Video Action Differencing (VidDiff), the novel task of identifying subtle differences between videos of the same action, which has numerous applications, such as coaching and skill learning. To enable development on this new task, we first create VidDiffBench, a benchmark dataset containing 549 video pairs, with human annotations of 4,469 fine-grained action differences and 2,075 timestamps indicating where these differences occur. Our experiments demonstrate that VidDiffBench poses a significant challenge for state-of-the-art large multimodal models (LMMs), such as GPT-4o and Qwen2-VL. By analyzing the failure cases of LMMs on VidDiffBench, we highlight two key challenges for this task: localizing relevant sub-actions over two videos and fine-grained frame comparison. To overcome these, we propose the VidDiff method, an agentic workflow that breaks the task into three stages: action difference proposal, keyframe localization, and frame differencing, each stage utilizing specialized foundation models. To encourage future research in this new task, we release the benchmark and code.

ICLR Conference 2025 Conference Paper

Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

  • Orr Zohar
  • Xiaohan Wang
  • Yonatan Bitton
  • Idan Szpektor
  • Serena Yeung

The performance and reasoning capabilities of Large Multi-modal Models (LMMs) is dependent on the size and quality of their training datasets. However, collecting datasets that support chain-of-thought instruction tuning is highly challenging. Existing video instruction tuning datasets are often derived by prompting large language models with video captions to generate question-answer pairs, which makes them predominantly descriptive rather than reasoning-focused. Meanwhile, many labeled video datasets with diverse labels and supervision exist -- however, we find that their integration into LMMs is non-trivial. Herein, we present $\underline{\text{Video}}$ $\underline{\text{S}}\text{elf}$-$\underline{\text{T}}\text{raining}$ $\text{with}$ $\underline{\text{a}}\text{ugmented}$ $\underline{\text{R}}\text{easoning}$ (Video-STaR), the first self-training approach for video instruction tuning. Video-STaR allows the utilization of *any* labeled video dataset for video instruction tuning. In Video-STaR, an LMM cycles between instruction generation and finetuning, which we show (I) improves general video understanding and (II) adapts LMMs to novel downstream tasks with existing supervision. During instruction generation, an LMM is prompted to propose an answer. The answers are then filtered only to those that contain the original video labels, and the LMM is then re-trained on the generated dataset. By training exclusively on generated answers containing the correct video labels, Video-STaR leverages these existing labels as weak supervision for video instruction tuning. Our results demonstrate that Video-STaR-augmented LMMs achieve notable improvements in (I) general Video QA, where TempCompass performance improved by 6.1%, *and* (II) downstream tasks, with a 9.9% increase in Kinetics700-QA accuracy and a 4.0% improvement in action quality assessment on FineDiving, while also exhibiting better interpretability.

IJCAI Conference 2024 Conference Paper

Continual Multimodal Knowledge Graph Construction

  • Xiang Chen
  • Jingtian Zhang
  • Xiaohan Wang
  • Ningyu Zhang
  • Tongtong Wu
  • Yuxiang Wang
  • Yongheng Wang
  • Huajun Chen

Current Multimodal Knowledge Graph Construction (MKGC) models struggle with the real-world dynamism of continuously emerging entities and relations, often succumbing to catastrophic forgetting—loss of previously acquired knowledge. This study introduces benchmarks aimed at fostering the development of the continual MKGC domain. We further introduce the MSPT framework, designed to surmount the shortcomings of existing MKGC approaches during multimedia data processing. MSPT harmonizes the retention of learned knowledge (stability) and the integration of new data (plasticity), outperforming current continual learning and multimodal methods. Our results confirm MSPT's superior performance in evolving knowledge environments, showcasing its capacity to navigate the balance between stability and plasticity.

AAAI Conference 2024 Conference Paper

Cross-Sentence Gloss Consistency for Continuous Sign Language Recognition

  • Qi Rao
  • Ke Sun
  • Xiaohan Wang
  • Qi Wang
  • Bang Zhang

Continuous sign language recognition (CSLR) aims to recognize gloss sequences from continuous sign videos. Recent works enhance the gloss representation consistency by mining correlations between visual and contextual modules within individual sentences. However, there still remain much richer correlations among glosses across different sentences. In this paper, we present a simple yet effective Cross-Sentence Gloss Consistency (CSGC), which enforces glosses belonging to a same category to be more consistent in representation than those belonging to different categories, across all training sentences. Specifically, in CSGC, a prototype is maintained for each gloss category and benefits the gloss discrimination in a contrastive way. Thanks to the well-distinguished gloss prototype, an auxiliary similarity classifier is devised to enhance the recognition clues, thus yielding more accurate results. Extensive experiments conducted on three CSLR datasets show that our proposed CSGC significantly boosts the performance of CSLR, surpassing existing state-of-the-art works by large margins (i.e., 1.6% on PHOENIX14, 2.4% on PHOENIX14-T, and 5.7% on CSL-Daily).

AAAI Conference 2024 Conference Paper

DGL: Dynamic Global-Local Prompt Tuning for Text-Video Retrieval

  • Xiangpeng Yang
  • Linchao Zhu
  • Xiaohan Wang
  • Yi Yang

Text-video retrieval is a critical multi-modal task to find the most relevant video for a text query. Although pretrained models like CLIP have demonstrated impressive potential in this area, the rising cost of fully finetuning these models due to increasing model size continues to pose a problem. To address this challenge, prompt tuning has emerged as an alternative. However, existing works still face two problems when adapting pretrained image-text models to downstream video-text tasks: (1) The visual encoder could only encode frame-level features and failed to extract global-level general video information. (2) Equipping the visual and text encoder with separated prompts failed to mitigate the visual-text modality gap. To this end, we propose DGL, a cross-modal Dynamic prompt tuning method with Global-Local video attention. In contrast to previous prompt tuning methods, we employ the shared latent space to generate local-level text and frame prompts that encourage inter-modal interaction. Furthermore, we propose modeling video in a global-local attention mechanism to capture global video information from the perspective of prompt tuning. Extensive experiments reveal that when only 0.67% parameters are tuned, our cross-modal prompt tuning strategy DGL outperforms or is comparable to fully finetuning methods on MSR-VTT, VATEX, LSMDC, and ActivityNet datasets. Code will be available at https://github.com/knightyxp/DGL.

AAAI Conference 2024 Conference Paper

Interpretable3D: An Ad-Hoc Interpretable Classifier for 3D Point Clouds

  • Tuo Feng
  • Ruijie Quan
  • Xiaohan Wang
  • Wenguan Wang
  • Yi Yang

3D decision-critical tasks urgently require research on explanations to ensure system reliability and transparency. Extensive explanatory research has been conducted on 2D images, but there is a lack in the 3D field. Furthermore, the existing explanations for 3D models are post-hoc and can be misleading, as they separate explanations from the original model. To address these issues, we propose an ad-hoc interpretable classifier for 3D point clouds (i.e., Interpretable3D). As an intuitive case-based classifier, Interpretable3D can provide reliable ad-hoc explanations without any embarrassing nuances. It allows users to understand how queries are embedded within past observations in prototype sets. Interpretable3D has two iterative training steps: 1) updating one prototype with the mean of the embeddings within the same sub-class in Prototype Estimation, and 2) penalizing or rewarding the estimated prototypes in Prototype Optimization. The mean of embeddings has a clear statistical meaning, i.e., class sub-centers. Moreover, we update prototypes with their most similar observations in the last few epochs. Finally, Interpretable3D classifies new samples according to prototypes. We evaluate the performance of Interpretable3D on four popular point cloud models: DGCNN, PointNet2, PointMLP, and PointNeXt. Our Interpretable3D demonstrates comparable or superior performance compared to softmax-based black-box models in the tasks of 3D shape classification and part segmentation. Our code is released at: github.com/FengZicai/Interpretable3D.

ICLR Conference 2024 Conference Paper

Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models

  • Shuai Zhao 0006
  • Xiaohan Wang
  • Linchao Zhu
  • Yi Yang 0001

One fascinating aspect of pre-trained vision-language models (VLMs) learning under language supervision is their impressive zero-shot generalization capability. However, this ability is hindered by distribution shifts between the training and testing data. Previous test time adaptation (TTA) methods for VLMs in zero-shot classification rely on minimizing the entropy of model outputs, tending to be stuck in incorrect model predictions. In this work, we propose TTA with feedback to rectify the model output and prevent the model from becoming blindly confident. Specifically, a CLIP model is adopted as the reward model during TTA and provides feedback for the VLM. Given a single test sample, the VLM is forced to maximize the CLIP reward between the input and sampled results from the VLM output distribution. The proposed \textit{reinforcement learning with CLIP feedback~(RLCF)} framework is highly flexible and universal. Beyond the classification task, with task-specific sampling strategies and a proper reward baseline choice, RLCF can be easily extended to not only discrimination tasks like retrieval but also generalization tasks like image captioning, improving the zero-shot generalization capacity of VLMs. According to the characteristics of these VL tasks, we build different fully TTA pipelines with RLCF to improve the zero-shot generalization ability of various VLMs. Extensive experiments along with promising empirical results demonstrate the effectiveness of RLCF. The code is available at https://github.com/mzhaoshuai/RLCF.

NeurIPS Conference 2024 Conference Paper

Why are Visually-Grounded Language Models Bad at Image Classification?

  • Yuhui Zhang
  • Alyssa Unell
  • Xiaohan Wang
  • Dhruba Ghosh
  • Yuchang Su
  • Ludwig Schmidt
  • Serena Yeung-Levy

Image classification is one of the most fundamental capabilities of machine vision intelligence. In this work, we revisit the image classification task using visually-grounded language models (VLMs) such as GPT-4V and LLaVA. We find that existing proprietary and public VLMs, despite often using CLIP as a vision encoder and having many more parameters, significantly underperform CLIP on standard image classification benchmarks like ImageNet. To understand the reason, we explore several hypotheses concerning the inference algorithms, training objectives, and data processing in VLMs. Our analysis reveals that the primary cause is data-related: critical information for image classification is encoded in the VLM's latent space but can only be effectively decoded with enough training data. Specifically, there is a strong correlation between the frequency of class exposure during VLM training and instruction-tuning and the VLM's performance in those classes; when trained with sufficient data, VLMs can match the accuracy of state-of-the-art classification models. Based on these findings, we enhance a VLM by integrating classification-focused datasets into its training, and demonstrate that the enhanced classification performance of the VLM transfers to its general capabilities, resulting in an improvement of 11. 8% on the newly collected ImageWikiQA dataset.

NeurIPS Conference 2023 Conference Paper

CaMP: Causal Multi-policy Planning for Interactive Navigation in Multi-room Scenes

  • Xiaohan Wang
  • Yuehu Liu
  • Xinhang Song
  • Beibei Wang
  • Shuqiang Jiang

Visual navigation has been widely studied under the assumption that there may be several clear routes to reach the goal. However, in more practical scenarios such as a house with several messy rooms, there may not. Interactive Navigation (InterNav) considers agents navigating to their goals more effectively with object interactions, posing new challenges of learning interaction dynamics and extra action space. Previous works learn single vision-to-action policy with the guidance of designed representations. However, the causality between actions and outcomes is prone to be confounded when the attributes of obstacles are diverse and hard to measure. Learning policy for long-term action planning in complex scenes also leads to extensive inefficient exploration. In this paper, we introduce a causal diagram of InterNav clarifying the confounding bias caused by obstacles. To address the problem, we propose a multi-policy model that enables the exploration of counterfactual interactions as well as reduces unnecessary exploration. We develop a large-scale dataset containing 600k task episodes in 12k multi-room scenes based on the ProcTHOR simulator and showcase the effectiveness of our method with the evaluations on our dataset.

IJCAI Conference 2023 Conference Paper

Open Anomalous Trajectory Recognition via Probabilistic Metric Learning

  • Qiang Gao
  • Xiaohan Wang
  • Chaoran Liu
  • Goce Trajcevski
  • Li Huang
  • Fan Zhou

Typically, trajectories considered anomalous are the ones deviating from usual (e. g. , traffic-dictated) driving patterns. However, this closed-set context fails to recognize the unknown anomalous trajectories, resulting in an insufficient self-motivated learning paradigm. In this study, we investigate the novel Anomalous Trajectory Recognition problem in an Open-world scenario (ATRO) and introduce a novel probabilistic Metric learning model, namely ATROM, to address it. Specifically, ATROM can detect the presence of unknown anomalous behavior in addition to identifying known behavior. It has a Mutual Interaction Distillation that uses contrastive metric learning to explore the interactive semantics regarding the diverse behavioral intents and a Probabilistic Trajectory Embedding that forces the trajectories with distinct behaviors to follow different Gaussian priors. More importantly, ATROM offers a probabilistic metric rule to discriminate between known and unknown behavioral patterns by taking advantage of the approximation of multiple priors. Experimental results on two large-scale trajectory datasets demonstrate the superiority of ATROM in addressing both known and unknown anomalous patterns.

ICRA Conference 2022 Conference Paper

Multi-robot Cooperative Pursuit via Potential Field-Enhanced Reinforcement Learning

  • Zheng Zhang
  • Xiaohan Wang
  • Qingrui Zhang
  • Tianjiang Hu

It is of great challenge, though promising, to coordinate collective robots for hunting an evader in a decentralized manner purely in light of local observations. In this paper, this challenge is addressed by a novel hybrid cooperative pursuit algorithm that combines reinforcement learning with the artificial potential field method. In the proposed algorithm, decentralized deep reinforcement learning is employed to learn cooperative pursuit policies that are adaptive to dynamic environments. The artificial potential field method is integrated into the learning process as predefined rules to improve the data efficiency and generalization ability. It is shown by numerical simulations that the proposed hybrid design outperforms the pursuit policies either learned from vanilla reinforcement learning or designed by the potential field method. Furthermore, experiments are conducted by transferring the learned pursuit policies into real-world mobile robots. Experimental results demonstrate the feasibility and potential of the proposed algorithm in learning multiple cooperative pursuit strategies.

AAAI Conference 2020 Conference Paper

Symbiotic Attention with Privileged Information for Egocentric Action Recognition

  • Xiaohan Wang
  • Yu Wu
  • Linchao Zhu
  • Yi Yang

Egocentric video recognition is a natural testbed for diverse interaction reasoning. Due to the large action vocabulary in egocentric video datasets, recent studies usually utilize a twobranch structure for action recognition, i. e. , one branch for verb classification and the other branch for noun classification. However, correlation study between the verb and the noun branches have been largely ignored. Besides, the two branches fail to exploit local features due to the absence of position-aware attention mechanism. In this paper, we propose a novel Symbiotic Attention framework leveraging Privileged information (SAP) for egocentric video recognition. Finer position-aware object detection features can facilitate the understanding of actor’s interaction with the object. We introduce these features in action recognition and regard them as privileged information. Our framework enables mutual communication among the verb branch, the noun branch, and the privileged information. This communication process not only injects local details into global features, but also exploits implicit guidance about the spatio-temporal position of an on-going action. We introduce a novel symbiotic attention (SA) to enable effective communication. It first normalizes the detection guided features on one branch to underline the action-relevant information from the other branch. SA adaptively enhances the interactions among the three sources. To further catalyze this communication, spatial relations are uncovered for the selection of most action-relevant information. It identifies the most valuable and discriminative feature for classification. We validate the effectiveness of our SAP quantitatively and qualitatively. Notably, it achieves the state-ofthe-art on two large-scale egocentric video datasets.

v2026.09.13