Arrow Research search

Author name cluster

Xin Wei

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

11 papers
2 author rows

Possible papers

11

JBHI Journal 2026 Journal Article

Electrooculography-Based Detection of Refractive Vision Problems

  • Xin Wei
  • Huakun Liu
  • Yutaro Hirao
  • Monica Perusquia-Hernandez
  • Katsutoshi Masai
  • Hideaki Uchiyama
  • Kiyoshi Kiyokawa

Early detection of visual impairments remains a persistent challenge, especially due to the subtle and often unnoticed nature of early-stage symptoms. Recent works have attempted to transition clinical tests to home-based services or develop innovative diagnostic methods, but most approaches remain self-initiated and discrete. In this study, we focused on refractive disorders and explored the feasibility of using electrooculography (EOG) to detect changes in refractive power passively. Thirty-nine participants used optometry trial lenses to simulate different refractive conditions. Participants performed a series of visual tasks while their EOG signals were recorded. We trained classification models to predict simulated refractive power levels relative to baseline visual condition across multiple evaluation settings, including within-subject, temporal generalization, and across-subject scenarios. The findings reveal that refractive power classification models achieve a mean accuracy of $0. 950 \pm 0. 034$ in within-subject, within-condition scenarios. Within-subject models tested on data from a different time point showed highly variable performance. While some participants achieved promising results, overall accuracy remained low, with a mean of $0. 159 \pm 0. 285$. We employed three strategies to evaluate the across-subject models. Naive models performed poorly ( $0. 161 \pm 0. 063$ ) and linear normalization provided limited improvement ( $0. 175 \pm 0. 062$ ). However, the fine-tuning strategy substantially improved the model's performance ( $0. 785 \pm 0. 123$ ). EOG signals contain useful information for refractive power classification, particularly in personalized contexts. However, generalizing across time and individuals remains challenging. Overall, this work offers valuable insights for advancing EOG-based systems aimed at passive, real-time monitoring of visual conditions.

AAAI Conference 2026 Conference Paper

TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding

  • Canhui Tang
  • Zifan Han
  • Hongbo Sun
  • Sanping Zhou
  • Xuchong Zhang
  • Xin Wei
  • Ye Yuan
  • Huayu Zhang

Multimodal Large Language Models (MLLMs) have demonstrated significant progress in vision-language tasks, yet they still face challenges when processing long-duration video inputs. The limitation arises from MLLMs' context limit and training costs, necessitating sparse frame sampling before feeding videos into MLLMs. However, building a trainable sampling method remains challenging due to the unsupervised and non-differentiable nature of sparse frame sampling in Video-MLLMs. To address these problems, we propose Temporal Sampling Policy Optimization (**TSPO**), advancing MLLMs' long-form video-language understanding via reinforcement learning. Specifically, we first propose a trainable event-aware temporal agent, which captures event-query correlation for performing probabilistic keyframe selection. Then, we propose the TSPO reinforcement learning paradigm, which models keyframe selection and language generation as a joint decision-making process, enabling end-to-end group relative optimization for the temporal sampling policy. Furthermore, we propose a dual-style long video training data construction pipeline, balancing comprehensive temporal understanding and key segment localization. Finally, we incorporate rule-based answering accuracy and temporal locating reward mechanisms to optimize the temporal sampling policy. Comprehensive experiments show that our TSPO achieves state-of-the-art performance across multiple long video understanding benchmarks, and shows transferable ability across different cutting-edge Video-MLLMs.

ECAI Conference 2025 Conference Paper

Constructing WiFi-Video-Fused Multi-Modal Synthetic Datasets for Crowd Counting

  • Bing Jia
  • Xin Wei
  • Shaowen Sun
  • Lifei Hao
  • Baoqi Huang

Recent progress in crowd counting has underscored its potential across diverse real-world applications. Nevertheless, the majority of existing approaches remain constrained by reliance on unimodal data, thereby limiting both robustness and generalizability. To address these challenges, we present a novel virtual simulation framework for generating synchronized multi-modal synthetic datasets. The proposed framework supports automated data generation with high-fidelity ground-truth annotations and provides programmatic control over environmental conditions and crowd dynamics. Through comprehensive experiments employing pre-training and fine-tuning strategies, we demonstrate that the synthetic datasets produced by our framework substantially improve model performance and generalization in crowd counting tasks. The work contributes a scalable and reproducible solution to the problem of data scarcity in multi-modal crowd counting research.

JBHI Journal 2025 Journal Article

Label-Aware Dual Graph Neural Networks for Multi-Label Fundus Image Classification

  • Yanbei Liu
  • Xinwen Peng
  • Xin Wei
  • Lei Geng
  • Fang Zhang
  • Zhitao Xiao
  • Jerry Chun-Wei Lin

Fundus disease is a complex and universal disease involving a variety of pathologies. Its early diagnosis using fundus images can effectively prevent further diseases and provide targeted treatment plans for patients. Recent deep learning models for classification of this disease are gradually emerging as a critical research field, which is attracting widespread attention. However, in practice, most of the existing methods only focus on local visual cues of a single image, and ignore the underlying explicit interaction similarity between subjects and correlation information among pathologies in fundus diseases. In this paper, we propose a novel label-aware dual graph neural networks for multi-label fundus image classification that consists of population-based graph representation learning and pathology-based graph representation learning modules. Specifically, we first construct a population-based graph by integrating image features and non-image information to learn patient's representations by incorporating associations between subjects. Then, we represent pathologies as a sparse graph where its nodes are associated with pathology-based feature vectors and the edges correspond to probability of the co-occurrence of labels to generate a set of classifier scores by the propagation of multi-layer graph information. Finally, our model can adaptively recalibrate multi-label outputs. Detailed experiments and analysis of our results show the effectiveness of our method compared with state-of-the-art multi-label fundus image classification methods.

AAAI Conference 2023 Conference Paper

Feature Distribution Fitting with Direction-Driven Weighting for Few-Shot Images Classification

  • Xin Wei
  • Wei Du
  • Huan Wan
  • Weidong Min

Few-shot learning has received increasing attention and witnessed significant advances in recent years. However, most of the few-shot learning methods focus on the optimization of training process, and the learning of metric and sample generating networks. They ignore the importance of learning the ground-truth feature distributions of few-shot classes. This paper proposes a direction-driven weighting method to make the feature distributions of few-shot classes precisely fit the ground-truth distributions. The learned feature distributions can generate an unlimited number of training samples for the few-shot classes to avoid overfitting. Specifically, the proposed method consists of two optimization strategies. The direction-driven strategy is for capturing more complete direction information that can describe the feature distributions. The similarity-weighting strategy is proposed to estimate the impact of different classes in the fitting procedure and assign corresponding weights. Our method outperforms the current state-of-the-art performance by an average of 3% for 1-shot on standard few-shot learning benchmarks like miniImageNet, CIFAR-FS, and CUB. The excellent performance and compelling visualization show that our method can more accurately estimate the ground-truth distributions.

JBHI Journal 2023 Journal Article

SA-RPN: A Spacial Aware Region Proposal Network for Acne Detection

  • Jianwei Zhang
  • Lei Zhang
  • Junyou Wang
  • Xin Wei
  • Jiaqi Li
  • Xian Jiang
  • Dan Du

Automated detection of skin lesions offers excellent potential for interpretative diagnosis and precise treatment of acne vulgar. However, the blurry boundary and small size of lesions make it challenging to detect acne lesions with traditional object detection methods. To better understand the acne detection task, we construct a new benchmark dataset named AcneSCU, consisting of 276 facial images with 31777 instance-level annotations from clinical dermatology. To the best of our knowledge, AcneSCU is the first acne dataset with high-resolution imageries, precise annotations, and fine-grained lesion categories, which enables the comprehensive study of acne detection. More importantly, we propose a novel method called Spatial Aware Region Proposal Network (SA-RPN) to improve the proposal quality of two-stage detection methods. Specifically, the representation learning for the classification and localization task is disentangled with a double head component to promote the proposals for hard samples. Then, Normalized Wasserstein Distance of each proposal is predicted to improve the correlation between the classification scores and the proposals' intersection-over-unions (IoUs). SA-RPN can serve as a plug-and-play module to enhance standard two-stage detectors. Extensive experiments are conducted on both AcneSCU and the public dataset ACNE04, and the results show that the proposed method can consistently outperform state-of-the-art methods.

AAAI Conference 2022 System Paper

A Synthetic Prediction Market for Estimating Confidence in Published Work

  • Sarah Rajtmajer
  • Christopher Griffin
  • Jian Wu
  • Robert Fraleigh
  • Laxmaan Balaji
  • Anna Squicciarini
  • Anthony Kwasnica
  • David Pennock

Explainably estimating confidence in published scholarly work offers opportunity for faster and more robust scientific progress. We develop a synthetic prediction market to assess the credibility of published claims in the social and behavioral sciences literature. We demonstrate our system and detail our findings using a collection of known replication projects. We suggest that this work lays the foundation for a research agenda that creatively uses AI for peer review.

ICRA Conference 2022 Conference Paper

GCLO: Ground Constrained LiDAR Odometry with Low-drifts for GPS-denied Indoor Environments

  • Xin Wei
  • Jixin Lv
  • Jie Sun
  • Erbao Dong
  • Shiliang Pu

LiDAR is widely adopted in Simultaneous Localization And Mapping (SLAM) and High Definition (HD) map production. The accuracy of LiDAR Odometry (LO) is of great importance, especially in GPS-denied environments. However, we found typical LO results are prone to drift upwards along the vertical direction in underground parking lots, leading to poor mapping results. This paper proposes a Ground Constrained LO method named GCLO, which exploits planar grounds in these specific environments to compress the vertical pose drifts. GCLO is divided into three parts. First, a sensor-centric sliding map is maintained, and the point-to-plane ICP method is implemented to perform the scan-to-map registration. Then, at each key-frame, the sliding map is recorded as a local map. Ground points nearby are segmented and modeled as a planar landmark in the form of Closest Point (CP) parameterization. Finally, planar ground landmarks observed at different key-frames are associated. The ground landmark observation constraints are fused into the pose graph optimization framework to improve the LO performance. Experimental results in HIK and KITTI datasets demonstrate GCLO's superior performances in terms of accuracy in indoor multi-floor parking lots and flat outdoor sites. The limitation of GCLO in adaptability for other environments is also discussed.

JBHI Journal 2022 Journal Article

GREN: Graph-Regularized Embedding Network for Weakly-Supervised Disease Localization in X-Ray Images

  • Baolian Qi
  • Gangming Zhao
  • Xin Wei
  • Changde Du
  • Chengwei Pan
  • Yizhou Yu
  • Jinpeng Li

Locating diseases in chest X-ray images with few careful annotations saves large human effort. Recent works approached this task with innovative weakly-supervised algorithms such as multi-instance learning (MIL) and class activation maps (CAM), however, these methods often yield inaccurate or incomplete regions. One of the reasons is the neglection of the pathological implications hidden in the relationship across anatomical regions within each image and the relationship across images. In this paper, we argue that the cross-region and cross-image relationship, as contextual and compensating information, is vital to obtain more consistent and integral regions. To model the relationship, we propose the Graph Regularized Embedding Network (GREN), which leverages the intra-image and inter-image information to locate diseases on chest X-ray images. GREN uses a pre-trained U-Net to segment the lung lobes, and then models the intra-image relationship between the lung lobes using an intra-image graph to compare different regions. Meanwhile, the relationship between in-batch images is modeled by an inter-image graph to compare multiple images. This process mimics the training and decision-making process of a radiologist: comparing multiple regions and images for diagnosis. In order for the deep embedding layers of the neural network to retain structural information (important in the localization task), we use the Hash coding and Hamming distance to compute the graphs, which are used as regularizers to facilitate training. By means of this, our approach achieves the state-of-the-art result on NIH chest X-ray dataset for weakly-supervised disease localization. Our codes are accessible online.

NeurIPS Conference 2022 Conference Paper

Learning Generalizable Part-based Feature Representation for 3D Point Clouds

  • Xin Wei
  • Xiang Gu
  • Jian Sun

Deep networks on 3D point clouds have achieved remarkable success in 3D classification, while they are vulnerable to geometry variations caused by inconsistent data acquisition procedures. This results in a challenging 3D domain generalization (3DDG) problem, that is to generalize a model trained on source domain to an unseen target domain. Based on the observation that local geometric structures are more generalizable than the whole shape, we propose to reduce the geometry shift by a generalizable part-based feature representation and design a novel part-based domain generalization network (PDG) for 3D point cloud classification. Specifically, we build a part-template feature space shared by source and target domains. Shapes from distinct domains are first organized to part-level features and then represented by part-template features. The transformed part-level features, dubbed aligned part-based representations, are then aggregated by a part-based feature aggregation module. To improve the robustness of the part-based representations, we further propose a contrastive learning framework upon part-based shape representation. Experiments and ablation studies on 3DDA and 3DDG benchmarks justify the efficacy of the proposed approach for domain generalization, compared with the previous state-of-the-art methods. Our code will be available on http: //github. com/weixmath/PDG.

IJCAI Conference 2019 Conference Paper

Explore Truthful Incentives for Tasks with Heterogenous Levels of Difficulty in the Sharing Economy

  • Pengzhan Zhou
  • Xin Wei
  • Cong Wang
  • Yuanyuan Yang

Incentives are explored in the sharing economy to inspire users for better resource allocation. Previous works build a budget-feasible incentive mechanism to learn users' cost distribution. However, they only consider a special case that all tasks are considered as the same. The general problem asks for finding a solution when the cost for different tasks varies. In this paper, we investigate this general problem by considering a system with k levels of difficulty. We present two incentivizing strategies for offline and online implementation, and formally derive the ratio of utility between them in different scenarios. We propose a regret-minimizing mechanism to decide incentives by dynamically adjusting budget assignment and learning from users' cost distributions. Our experiment demonstrates utility improvement about 7 times and time saving of 54% to meet a utility objective compared to the previous works.

v2026.09.13