Arrow Research search

Author name cluster

Haochen Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

8 papers
2 author rows

Possible papers

8

AAAI Conference 2026 Conference Paper

CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models

  • Jingyao Li
  • Jingyun Wang
  • Molin Tan
  • Haochen Wang
  • Cilin Yan
  • Likun Shi
  • Jiayin Cai
  • Xiaolong Jiang

Cross-Video Reasoning (CVR) presents a significant challenge in video understanding, which requires simultaneous understanding of multiple videos to aggregate and compare information across groups of videos. Most existing video understanding benchmarks focus on single-video analysis, failing to assess the ability of multimodal large language models (MLLMs) to simultaneously reason over various videos. Recent benchmarks evaluate MLLMs' capabilities on multi-view videos that capture different perspectives of the same scene. However, their limited tasks hinder a thorough assessment of MLLMs in diverse real-world CVR scenarios. To this end, we introduce CrossVid, the first benchmark designed to comprehensively evaluate MLLMs' spatial-temporal reasoning ability in cross-video contexts. Firstly, CrossVid encompasses a wide spectrum of hierarchical tasks, comprising four high-level dimensions and ten specific tasks, thereby closely reflecting the complex and varied nature of real-world video understanding. Secondly, CrossVid provides 5,331 videos, along with 9,015 challenging question-answering pairs, spanning single-choice, multiple-choice, and open-ended question formats. Through extensive experiments on various open-source and closed-source MLLMs, we observe that Gemini-2.5-Pro performs best on CrossVid, achieving an average accuracy of 50.4%. Notably, our in-depth case study demonstrates that most current MLLMs struggle with CVR tasks, primarily due to their inability to integrate or compare evidence distributed across multiple videos for reasoning. These insights highlight the potential of CrossVid to guide future advancements in enhancing MLLMs’ CVR capabilities.

JBHI Journal 2025 Journal Article

Acupuncture State Detection at Zusanli (ST-36) Based on Scalp EEG and Transformer

  • Wenhao Rao
  • Meiyan Xu
  • Haochen Wang
  • Weicheng Hua
  • Jiayang Guo
  • Yongheng Zhang
  • Haibin Zhu
  • Ziqiu Zhou

In clinical acupuncture practice, needle twirling (NT) and needle retention (NR) are strategically combined to achieve different therapeutic effects, highlighting the importance of distinguishing between different acupuncture states. Scalp EEG has been proven significantly relevant to brain activity and acupuncture stimulation. In this work, we designed an acupuncture paradigm to collect scalp EEG to study the differences in EEG changes during different acupuncture states. Since deep learning (DL) has been increasingly used in EEG analysis, we propose the Acupuncture Transformer Detector (ATD), a model based on Convolutional Neural Networks (CNN) and Transformer technology. ATD encapsulates the local and global features of EEG under the acupuncture states of Zusanli acupoint (ST-36) in an end-to-end classification framework. The experiment results from 28 healthy participants show that the proposed model can efficiently classify the EEG in different states, with an accuracy of $85. 47\pm 0. 73\%$. In this study, time-frequency analysis revealed that power changes were mainly confined to the delta frequency band under different acupuncture states. Brain topography revealed that ST-36 was activated primarily on the left frontal and parieto-occipital areas. This method provides new ideas for automatic recognition of acupuncture status from the perspective of DL, offering new solutions for standardizing acupuncture procedures.

ICRA Conference 2025 Conference Paper

Integrating Learning-Based Manipulation and Physics-Based Locomotion for Whole-Body Badminton Robot Control

  • Haochen Wang
  • Zhiwei Shi
  • Chengxi Zhu
  • Yafei Qiao
  • Cheng Zhang 0014
  • Fan Yang 0059
  • Pengjie Ren
  • Lan Lu

Learning-based methods, such as imitation learning (IL) and reinforcement learning (RL), can produce excel control policies over challenging agile robot tasks, such as sports robot. However, no existing work has harmonized learning-based policy with model-based methods to reduce training complexity and ensure the safety and stability for agile badminton robot control. In this paper, we introduce Hamlet, a novel hybrid control system for agile badminton robots. Specifically, we propose a model-based strategy for chassis locomotion which provides a base for arm policy. We introduce a physics-informed “IL+RL” training framework for learning-based arm policy. In this train framework, a modelbased strategy with privileged information is used to guide arm policy training during both IL and RL phases. In addition, we train the critic model during IL phase to alleviate the performance drop issue when transitioning from IL to RL. We present results on our self-engineered badminton robot, achieving 94. 5% success rate against the serving machine and $\mathbf{9 0. 7 \%}$ success rate against human players. Our system can be easily generalized to other agile mobile manipulation tasks e. g. , agile catching, table tennis. A video demonstrating our system can be viewed at https://youtu.be/8-ixKAD18Mk.

NeurIPS Conference 2025 Conference Paper

MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs

  • Tianhao Peng
  • Haochen Wang
  • Yuanxing Zhang
  • Noah Wang
  • Zili Wang
  • Ge Zhang
  • Jian Yang
  • Shihao Li

The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e. g. , sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark for evaluating M ulti- V ideo U nderstanding for MLLMs. Specifically, our MVU-Eval mainly assesses eight core competencies through 1, 824 meticulously curated question-answer pairs spanning 4, 959 videos from diverse domains, addressing both fundamental perception tasks and high-order reasoning tasks. These capabilities are rigorously aligned with real-world applications such as multi-sensor synthesis in autonomous systems and cross-angle sports analytics. Through extensive evaluation of state-of-the-art open-source and closed-source models, we reveal significant performance discrepancies and limitations in current MLLMs' ability to perform understanding across multiple videos. The benchmark will be made publicly available to foster future research.

ICLR Conference 2025 Conference Paper

Reconstructive Visual Instruction Tuning

  • Haochen Wang
  • Anlin Zheng
  • Yucheng Zhao
  • Tiancai Wang
  • Zheng Ge
  • Xiangyu Zhang 0005
  • Zhaoxiang Zhang 0001

This paper introduces reconstructive visual instruction tuning (ROSS), a family of Large Multimodal Models (LMMs) that exploit vision-centric supervision signals. In contrast to conventional visual instruction tuning approaches that exclusively supervise text outputs, ROSS prompts LMMs to supervise visual outputs via reconstructing input images. By doing so, it capitalizes on the inherent richness and detail present within input images themselves, which are often lost in pure text supervision. However, producing meaningful feedback from natural images is challenging due to the heavy spatial redundancy of visual signals. To address this issue, ROSS employs a denoising objective to reconstruct latent representations of input images, avoiding directly regressing exact raw RGB values. This intrinsic activation design inherently encourages LMMs to maintain image detail, thereby enhancing their fine-grained comprehension capabilities and reducing hallucinations. Empirically, ROSS consistently brings significant improvements across different visual encoders and language models. In comparison with extrinsic assistance state-of-the-art alternatives that aggregate multiple visual experts, ROSS delivers competitive performance with a single SigLIP visual encoder, demonstrating the efficacy of our vision-centric supervision tailored for visual outputs. The code will be made publicly available upon acceptance.

NeurIPS Conference 2024 Conference Paper

OpenSatMap: A Fine-grained High-resolution Satellite Dataset for Large-scale Map Construction

  • Hongbo Zhao
  • Lue Fan
  • Yuntao Chen
  • Haochen Wang
  • Yuran Yang
  • Xiaojuan Jin
  • Yixin Zhang
  • Gaofeng Meng

In this paper, we propose OpenSatMap, a fine-grained, high-resolution satellite dataset for large-scale map construction. Map construction is one of the foundations of the transportation industry, such as navigation and autonomous driving. Extracting road structures from satellite images is an efficient way to construct large-scale maps. However, existing satellite datasets provide only coarse semantic-level labels with a relatively low resolution (up to level 19), impeding the advancement of this field. In contrast, the proposed OpenSatMap (1) has fine-grained instance-level annotations; (2) consists of high-resolution images (level 20); (3) is currently the largest one of its kind; (4) collects data with high diversity. Moreover, OpenSatMap covers and aligns with the popular nuScenes dataset and Argoverse 2 dataset to potentially advance autonomous driving technologies. By publishing and maintaining the dataset, we provide a high-quality benchmark for satellite-based map construction and downstream tasks like autonomous driving.

NeurIPS Conference 2023 Conference Paper

DropPos: Pre-Training Vision Transformers by Reconstructing Dropped Positions

  • Haochen Wang
  • Junsong Fan
  • Yuxi Wang
  • Kaiyou Song
  • Tong Wang
  • ZHAO-XIANG ZHANG

As it is empirically observed that Vision Transformers (ViTs) are quite insensitive to the order of input tokens, the need for an appropriate self-supervised pretext task that enhances the location awareness of ViTs is becoming evident. To address this, we present DropPos, a novel pretext task designed to reconstruct Dropped Positions. The formulation of DropPos is simple: we first drop a large random subset of positional embeddings and then the model classifies the actual position for each non-overlapping patch among all possible positions solely based on their visual appearance. To avoid trivial solutions, we increase the difficulty of this task by keeping only a subset of patches visible. Additionally, considering there may be different patches with similar visual appearances, we propose position smoothing and attentive reconstruction strategies to relax this classification problem, since it is not necessary to reconstruct their exact positions in these cases. Empirical evaluations of DropPos show strong capabilities. DropPos outperforms supervised pre-training and achieves competitive results compared with state-of-the-art self-supervised alternatives on a wide range of downstream benchmarks. This suggests that explicitly encouraging spatial reasoning abilities, as DropPos does, indeed contributes to the improved location awareness of ViTs. The code is publicly available at https: //github. com/Haochen-Wang409/DropPos.

NeurIPS Conference 2022 Conference Paper

Learning from Future: A Novel Self-Training Framework for Semantic Segmentation

  • Ye Du
  • Yujun Shen
  • Haochen Wang
  • Jingjing Fei
  • Wei Li
  • Liwei Wu
  • Rui Zhao
  • Zehua Fu

Self-training has shown great potential in semi-supervised learning. Its core idea is to use the model learned on labeled data to generate pseudo-labels for unlabeled samples, and in turn teach itself. To obtain valid supervision, active attempts typically employ a momentum teacher for pseudo-label prediction yet observe the confirmation bias issue, where the incorrect predictions may provide wrong supervision signals and get accumulated in the training process. The primary cause of such a drawback is that the prevailing self-training framework acts as guiding the current state with previous knowledge because the teacher is updated with the past student only. To alleviate this problem, we propose a novel self-training strategy, which allows the model to learn from the future. Concretely, at each training step, we first virtually optimize the student (i. e. , caching the gradients without applying them to the model weights), then update the teacher with the virtual future student, and finally ask the teacher to produce pseudo-labels for the current student as the guidance. In this way, we manage to improve the quality of pseudo-labels and thus boost the performance. We also develop two variants of our future-self-training (FST) framework through peeping at the future both deeply (FST-D) and widely (FST-W). Taking the tasks of unsupervised domain adaptive semantic segmentation and semi-supervised semantic segmentation as the instances, we experimentally demonstrate the effectiveness and superiority of our approach under a wide range of settings. Code is available at https: //github. com/usr922/FST.

v2026.09.13