Arrow Research search

Author name cluster

Xiaobo Zhang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

9 papers
2 author rows

Possible papers

9

AAAI Conference 2026 Conference Paper

Diffusion Reconstruction-based Data Likelihood Estimation for Core-Set Selection

  • Mingyang Chen
  • Jiawei Du
  • Bo Huang
  • Yi Wang
  • Xiaobo Zhang
  • Wei Wang

Existing core-set selection methods predominantly rely on heuristic scoring signals such as training dynamics or model uncertainty, lacking explicit modeling of data likelihood. This omission may hinder the constructed subset from capturing subtle yet critical distributional structures that underpin effective model training. In this work, we propose a novel, theoretically grounded approach that leverages diffusion models to estimate data likelihood via reconstruction deviation induced by partial reverse denoising. Specifically, we establish a formal connection between reconstruction error and data likelihood, grounded in the Evidence Lower Bound (ELBO) of Markovian diffusion processes, thereby enabling a principled, distribution-aware scoring criterion for data selection. Complementarily, we introduce an efficient information-theoretic method to identify the optimal reconstruction timestep, ensuring that the deviation provides a reliable signal indicative of underlying data likelihood. Extensive experiments on ImageNet demonstrate that reconstruction deviation offers an effective scoring criterion, consistently outperforming existing baselines across selection ratios, and closely matching full-data training using only 50% of the data. Further analysis shows that the likelihood-informed nature of our score reveals informative insights in data selection, shedding light on the interplay between data distributional characteristics and model learning preferences.

AIIM Journal 2026 Journal Article

Machine learning-based methods for predicting postpartum depression: A review

  • Xiaobo Zhang
  • Jinyu Bao
  • Jianzhong Ye
  • Xiaoping Qiu
  • Yi Pan

Postpartum depression (PPD) is a widespread mental illness after delivery, which has a substantial impact on the health of both mothers and infants. Machine learning (ML) has developed rapidly and plays a vital function in disease prediction. This article summarizes and reviews ML techniques used to predict PPD, aiming to investigate their potential for predicting the risk of PPD. We performed a bibliographic search on China National Knowledge Infrastructure (CNKI), China Science and Technology Journal Database (CQVIP), Web of Science and Google Scholar looking for studies aimed at the prediction of PPD using ML techniques. Of the 103 articles collected, 25 fulfilled the inclusion criteria. Supervised learning was the primary ML technique applied and the most prevalent ML models were gradient boosting, random forest, and support vector machine. Notably, the PPD prediction model based on gradient boosting has the best effect and the vast majority of studies have ended up in an area under the curve that exceeds 0. 7. All studies indicate that it is feasible to use ML techniques to predict PPD. We focused on the research of ML techniques used for PPD prediction, and did not delve into the medical knowledge related to PPD prediction. ML has great potential in the field of PPD prediction. Nevertheless, further research is needed to fully realize this prospect, including standardizing data collection, improving the robustness of feature selection, and encouraging interdisciplinary collaboration. This will help improve the stability and accuracy of the model and provide more personalized medical services for patients.

JBHI Journal 2025 Journal Article

A Medical Multimodal Large Language Model for Pediatric Pneumonia

  • Weiwei Tian
  • Xinyu Huang
  • Tianhao Cheng
  • Wen He
  • Jinwu Fang
  • Rui Feng
  • Daoying Geng
  • Xiaobo Zhang

Pediatric pneumonia is the leading cause of death among children under five years worldwide, imposing a substantial burden on affected families. Currently, there are three significant hurdles in diagnosing and treating pediatric pneumonia. Firstly, pediatric pneumonia shares similar symptoms with other respiratory diseases, making rapid and accurate differential diagnosis challenging. Secondly, primary hospitals often lack sufficient medical resources and experienced doctors. Lastly, providing personalized diagnostic reports and treatment recommendations is labor-intensive and time-consuming. To tackle these challenges, we proposed a Med ical M ultimodal L arge L anguage M odel for P ediatric P neumonia (P2Med-MLLM). It was capable of handling diverse clinical tasks—such as generating free-text medical records and radiology reports—within a unified framework. Specifically, P2Med-MLLM was trained on a large-scale dataset, including real clinical information from 163, 999 outpatient and 8, 684 inpatient cases. It can process both plain text data (e. g. , outpatient and inpatient records) and interleaved image-text pairs (e. g. , 2D chest X-ray images, 3D chest Computed Tomography images, and corresponding radiology reports). We designed a three-stage training strategy to enable P2Med-MLLM to comprehend medical knowledge and follow instructions for various clinical decision-support tasks. To rigorously evaluate P2Med-MLLM's performance, we conducted automatic scoring by the large language model and manual scoring by the specialist on the test set of 642 samples, meticulously verified by pediatric pulmonology specialists. The results demonstrated the reliability of automated scoring and the superiority of P2Med-MLLM. This work plays a crucial role in assisting doctors with prompt diagnosis and treatment planning, reducing severe symptom mortality rates, and optimizing the allocation of medical resources.

NeurIPS Conference 2025 Conference Paper

AOR: Anatomical Ontology-Guided Reasoning for Medical Large Multimodal Model in Chest X-Ray Interpretation

  • Qingqiu Li
  • Zihang Cui
  • Seongsu Bae
  • Jilan Xu
  • Runtian Yuan
  • Yuejie Zhang
  • Rui Feng
  • Quanli Shen

Chest X-rays (CXRs) are the most frequently performed imaging examinations in clinical settings. Recent advancements in Medical Large Multimodal Models (MLMMs) have enabled automated CXR interpretation, improving diagnostic accuracy and efficiency. However, despite their strong visual understanding, current MLMMs still face two major challenges: (1) insufficient region-level understanding and interaction, and (2) limited accuracy and interpretability due to single-step prediction. In this paper, we address these challenges by empowering MLMMs with anatomy-centric reasoning capabilities to enhance their interactivity and explainability. Specifically, we propose an Anatomical Ontology-Guided Reasoning (AOR) framework that accommodates both textual and optional visual prompts, centered on region-level information to enable multimodal multi-step reasoning. We also develop AOR-Instruction, a large instruction dataset for MLMMs training, under the guidance of expert physicians. Our experiments demonstrate AOR's superior performance in both Visual Question Answering (VQA) and report generation tasks. Code and data are available at: https: //github. com/Liqq1/AOR.

ICLR Conference 2025 Conference Paper

Influence-Guided Diffusion for Dataset Distillation

  • Mingyang Chen
  • Jiawei Du
  • Bo Huang 0017
  • Yi Wang 0017
  • Xiaobo Zhang
  • Wei Wang 0011

Dataset distillation aims to streamline the training process by creating a compact yet effective dataset for a much larger original dataset. However, existing methods often struggle with distilling large, high-resolution datasets due to prohibitive resource costs and limited performance, primarily stemming from sample-wise optimizations in the pixel space. Motivated by the remarkable capabilities of diffusion generative models in learning target dataset distributions and controllably sampling high-quality data tailored to user needs, we propose framing dataset distillation as a controlled diffusion generation task aimed at generating data specifically tailored for effective training purposes. By establishing a correlation between the overarching objective of dataset distillation and the trajectory influence function, we introduce the Influence-Guided Diffusion (IGD) sampling framework to generate training-effective data without the need to retrain diffusion models. An efficient guided function is designed by leveraging the trajectory influence function as an indicator to steer diffusions to produce data with influence promotion and diversity enhancement. Extensive experiments show that the training performance of distilled datasets generated by diffusions can be significantly improved by integrating with our IGD method and achieving state-of-the-art performance in distilling ImageNet datasets. Particularly, an exceptional result is achieved on the ImageNet-1K, reaching 60.3\% at IPC=50. Our code is available at https://github.com/mchen725/DD_IGD.

AAAI Conference 2024 Conference Paper

Let All Be Whitened: Multi-Teacher Distillation for Efficient Visual Retrieval

  • Zhe Ma
  • Jianfeng Dong
  • Shouling Ji
  • Zhenguang Liu
  • Xuhong Zhang
  • Zonghui Wang
  • Sifeng He
  • Feng Qian

Visual retrieval aims to search for the most relevant visual items, e.g., images and videos, from a candidate gallery with a given query item. Accuracy and efficiency are two competing objectives in retrieval tasks. Instead of crafting a new method pursuing further improvement on accuracy, in this paper we propose a multi-teacher distillation framework Whiten-MTD, which is able to transfer knowledge from off-the-shelf pre-trained retrieval models to a lightweight student model for efficient visual retrieval. Furthermore, we discover that the similarities obtained by different retrieval models are diversified and incommensurable, which makes it challenging to jointly distill knowledge from multiple models. Therefore, we propose to whiten the output of teacher models before fusion, which enables effective multi-teacher distillation for retrieval models. Whiten-MTD is conceptually simple and practically effective. Extensive experiments on two landmark image retrieval datasets and one video retrieval dataset demonstrate the effectiveness of our proposed method, and its good balance of retrieval performance and efficiency. Our source code is released at https://github.com/Maryeon/whiten_mtd.

AAAI Conference 2024 Conference Paper

Towards Evidential and Class Separable Open Set Object Detection

  • Ruofan Wang
  • Rui-Wei Zhao
  • Xiaobo Zhang
  • Rui Feng

Detecting in open-world scenarios poses a formidable challenge for models intended for real-world deployment. The advanced closed set object detectors achieve impressive performance under the closed set setting, but often produce overconfident misprediction on unknown objects due to the lack of supervision. In this paper, we propose a novel Evidential Object Detector (EOD) to formulate the Open Set Object Detection (OSOD) problem from the perspective of Evidential Deep Learning (EDL) theory, which quantifies classification uncertainty by placing the Dirichlet Prior over the categorical distribution parameters. The task-specific customized evidential framework, equipped with meticulously designed model architecture and loss function, effectively bridges the gap between EDL theory and detection tasks. Moreover, we utilize contrastive learning as an implicit means of evidential regularization and to encourage the class separation in the latent space. Alongside, we innovatively model the background uncertainty to further improve the unknown discovery ability. Extensive experiments on benchmark datasets demonstrate the outperformance of the proposed method over existing ones.

AAAI Conference 2023 Conference Paper

TransVCL: Attention-Enhanced Video Copy Localization Network with Flexible Supervision

  • Sifeng He
  • Yue He
  • Minlong Lu
  • Chen Jiang
  • Xudong Yang
  • Feng Qian
  • Xiaobo Zhang
  • Lei Yang

Video copy localization aims to precisely localize all the copied segments within a pair of untrimmed videos in video retrieval applications. Previous methods typically start from frame-to-frame similarity matrix generated by cosine similarity between frame-level features of the input video pair, and then detect and refine the boundaries of copied segments on similarity matrix under temporal constraints. In this paper, we propose TransVCL: an attention-enhanced video copy localization network, which is optimized directly from initial frame-level features and trained end-to-end with three main components: a customized Transformer for feature enhancement, a correlation and softmax layer for similarity matrix generation, and a temporal alignment module for copied segments localization. In contrast to previous methods demanding the handcrafted similarity matrix, TransVCL incorporates long-range temporal information between feature sequence pair using self- and cross- attention layers. With the joint design and optimization of three components, the similarity matrix can be learned to present more discriminative copied patterns, leading to significant improvements over previous methods on segment-level labeled datasets (VCSL and VCDB). Besides the state-of-the-art performance in fully supervised setting, the attention architecture facilitates TransVCL to further exploit unlabeled or simply video-level labeled data. Additional experiments of supplementing video-level labeled datasets including SVD and FIVR reveal the high flexibility of TransVCL from full supervision to semi-supervision (with or without video-level annotation). Code is publicly available at https://github.com/transvcl/TransVCL.

TIST Journal 2022 Journal Article

Jointly Optimizing Expressional and Residual Models for 3D Facial Expression Removal

  • Qian Zheng
  • Yueming Wang
  • Zhenfang Hu
  • Xiaobo Zhang
  • Zhaohui Wu
  • Gang Pan

This article proposes a facial expression removal method to recover a 3D neutral face from a single 3D expressional or non-neutral face. We treat a 3D non-neutral face as the sum of its neutral one and the residual. This can be satisfied if the correspondence between 3D vertices of expressional faces and those of neutral faces is established. We propose a non-rigid deformation method to establish the correspondence between 3D faces. Then, according to algebra inequality, the minimization of a neutral face model can be replaced by the minimization of its upper bound, i.e., the errors of an expressional face model and a residual model. Thus, we co-optimize the representation errors of the latter two models and build the relationship between the representation coefficients of the two models. Given an expressional face as the input, its corresponding neutral face can be inferred by the associative representation parameters in these two models. In the testing stage, we use an iterative joint fitting scheme to obtain a more accurate recovery. Extensive experiments are conducted to evaluate our method. The results show that our method obtains considerably better performance than existing methods in terms of average root mean square errors and recognition rates, and also better visual effects.

v2026.09.13