Arrow Research search

Author name cluster

Xiao Han

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

38 papers
2 author rows

Possible papers

38

JBHI Journal 2026 Journal Article

Advanced Camera-Based Scoliosis Screening via Deep Learning Detection and Fusion of Trunk, Limb, and Skeleton Features

  • Ziyan Wang
  • Yi Zhou
  • Ninghui Xu
  • Yuqin Zhou
  • Heran Zhao
  • Zhiyong Chang
  • Zhigang Hu
  • Xiao Han

Scoliosis significantly impacts quality of life, highlighting the need for effective early scoliosis screening (SS) and intervention. However, current SS methods often involve physical contact, undressing, or radiation exposure. This study introduces an innovative, non-invasive SS approach utilizing a monocular RGB camera that eliminates the need for undressing, sensor attachment, and radiation exposure. We introduce a novel approach that employs Parameterized Human 3D Reconstruction (PH3DR) to reconstruct 3D human models, thereby effectively eliminating clothing obstructions, seamlessly integrated with an ISANet segmentation network, which has been enhanced by Multi-Scale Fusion Attention (MSFA) module we proposed for facilitating the segmentation of distinct human trunk and limb features (HTLF), capturing body surface asymmetries related to scoliosis. Additionally, we propose a Swin Transformer-enhanced CMU-Pose to extract human skeleton features (HSF), identifying skeletal asymmetries crucial for SS. Finally, we develop a fusion model that integrates the HTLF and HSF, combining surface morphology and skeletal features to improve the precision of SS. The experiments demonstrated that PH3DR and MSFA significantly improved the segmentation and extraction of HTLF, whereas ST-based CMU-Pose substantially enhanced the extraction of HSF. Our final model achieved a comparable F1 (0. 895 $\pm$ 0. 014) to the best-performing baseline model, with only 0. 79% of the parameters and 1. 64% of the FLOPs, achieving 36 FPS–significantly higher than the best-performing baseline model (10 FPS). Moreover, our model outperformed two spine surgeons, one less experienced and the other moderately experienced. With its patient-friendly, privacy-preserving, and easily deployable solution, this approach is particularly well-suited for early SS and routine monitoring.

TMLR Journal 2026 Journal Article

Algorithmic Recourse in Abnormal Multivariate Time Series

  • Xiao Han
  • Lu Zhang
  • Yongkai Wu
  • Shuhan Yuan

Algorithmic recourse provides actionable recommendations to alter unfavorable predictions of machine learning models, enhancing transparency through counterfactual explanations. While significant progress has been made in algorithmic recourse for static data, such as tabular and image data, limited research explores recourse for multivariate time series, particularly for reversing abnormal time series. This paper introduces Recourse in time series Anomaly Detection (RecAD), a framework for addressing anomalies in multivariate time series using backtracking counterfactual reasoning. By modeling the causes of anomalies as external interventions on exogenous variables, RecAD predicts recourse actions to restore normal status as counterfactual explanations, where the recourse function, responsible for generating actions based on observed data, is trained using an end-to-end approach. Experiments on synthetic and real-world datasets demonstrate its effectiveness.

EAAI Journal 2026 Journal Article

Multi-task time series forecasting with adaptive graph neural networks based on feature uncertainty

  • Xiao Han
  • Zhisong Pan
  • Yongjie Huang

Multi-task time series forecasting aims to enhance prediction accuracy by leveraging shared knowledge among related tasks, finding widespread applications in critical domains such as financial risk analysis and medical monitoring. However, existing methods often overlook the impact of feature uncertainty on knowledge reliability and fail to dynamically model cross-timestep task relationships. This leads to challenges like negative transfer and the inability of static sharing mechanisms to adapt to temporal dynamics. To address these issues, this paper proposes DPG-Net, a Dynamic Probabilistic Graph Network that utilizes a Bayesian framework to model task features as Gaussian random variables, thereby quantifying their uncertainty. This uncertainty guides a gated attention mechanism to dynamically construct a cross-timestep probabilistic graph, enabling adaptive and reliable knowledge sharing among tasks. Experimental validation on multiple clinical risk prediction datasets demonstrates that DPG-Net achieves superior average performance in terms of AUROC on the MIMIC-III Infection, PhysioNet, MIMIC-III Heart Failure, and MIMIC-III Respiratory Failure datasets compared to state-of-the-art models, with improvements of 11. 48%, 8. 30%, 10. 48%, and 7. 42%, respectively, highlighting its capability to improve prediction accuracy and mitigate negative transfer. Ablation studies further confirm the effectiveness of the probabilistic modeling and dynamic knowledge-sharing mechanisms.

EAAI Journal 2026 Journal Article

Power load forecasting based on time-frequency domain feature fusion

  • Wenhua Jiao
  • Xiao Han
  • Ce Yu
  • Xiaowei Xu
  • Qing Zhang
  • Xiang Zhang
  • Bin Wang
  • Lijuan Li

Power Load Forecasting (PLF) provides decision support for grid planning and efficiency improvement and plays a pivotal role in reducing energy consumption. Current PLF methodologies face significant challenges in effectively coordinating the complex interdependencies among multiple factors affecting power system load. To address this fundamental limitation, we propose the dual-stream architecture named Time -Frequency Interaction Network (TF-interactionNet) that synergistically integrates the time-frequency domain, enabling it to consider both autocorrelation and interaction of time series. Our approach introduces three specialized modules: 1) The Time-Mamba (T-Mamba) module employs advanced state space models to capture long-range temporal dependencies and global trend patterns in power load sequences; 2) The Frequency-Temporal Convolutional Network (F-TCN) module utilizes frequency-optimized temporal convolutional network to extract localized frequency characteristics and identify subtle load fluctuation patterns. 3) The proposed Feature Attention Fusion Network (FAFN) innovatively fuses complementary features in the time-frequency domain, retaining the key effective features while avoiding the information redundancy that occurs during fusion. Extensive experimental evaluations show that TF-interactionNet achieves state-of-the-art performance across multiple evaluation metrics. On three datasets, compared to the latest forecasting model, the proposed model reduced the mean squared error of 24-h load forecasting by 0. 1%, 0. 1%, and 1%, respectively, while the mean absolute percentage error decreased by 0. 198%, 0. 066%, and 0. 106%, respectively. It still maintains high effectiveness in predicting power load changes over longer time periods. This model improves forecasting accuracy across multiple prediction horizons, offering new insights for advancing PLF and holding sustained research value.

EAAI Journal 2025 Journal Article

A novel short-term prediction method for distributed photovoltaic power generation considering extreme weather

  • Xin Guan
  • Xiao Han
  • Jun Wang
  • Tao Wang

Distributed photovoltaic power plants are often impacted by various factors such as weather conditions and geographical locations, making it challenging to fully capture the spatial correlation characteristics among multiple photovoltaic plants. Furthermore, the failure to consider meteorological factors that influence photovoltaic output results in larger prediction errors during extreme weather events. To reduce prediction errors, this paper proposes a short-term photovoltaic forecasting method that considers meteorological factors, explores spatial correlations among photovoltaic plants, and captures temporal characteristics. Firstly, a Graph Attention Network is established to obtain spatial correlations between different plants while a Convolutional Neural Network is employed to extract feature information of meteorological factors. Then, the feature information from these two sources is integrated and input into a Long Short-Term Memory network, which is enhanced based on Spiking Neural P Systems to extract temporal characteristics of photovoltaic output and complete the prediction task. Finally, real-world power station datasets are utilized for validation and comparison with several typical photovoltaic prediction models. The results clearly show that the application of artificial intelligence in this proposed method can effectively improve the accuracy of distributed photovoltaic power forecasting, demonstrating the great potential of AI in the field of photovoltaic power prediction.

AAAI Conference 2025 Conference Paper

Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion Models

  • Fei Shen
  • Hu Ye
  • Sibo Liu
  • Jun Zhang
  • Cong Wang
  • Xiao Han
  • Yang Wei

Recent research showcases the considerable potential of conditional diffusion models for generating consistent stories. However, current methods, which primarily generate stories in a caption-dependent manner, often overlook the importance of contextual consistency and the relevance of frames during sequential generation. To address this, we propose a novel Rich-contextual Conditional Diffusion Models (RCDMs), a two-stage approach designed to enhance story generation's semantic consistency and temporal consistency. Specifically, in the first stage, the frame-prior transformer diffusion model is presented to predict the frame semantic embedding of the unknown clip by aligning the semantic correlations between the captions and frames of the known clip. The second stage establishes a robust model with rich contextual conditions, including reference images of the known clip, the predicted frame semantic embedding of the unknown clip, and text embeddings of all captions. By jointly injecting these rich contextual conditions at the image and feature levels, RCDMs can generate semantic and temporal consistency stories. Moreover, RCDMs can generate consistent stories with a single forward inference compared to autoregressive models. Our qualitative and quantitative results demonstrate that our proposed RCDMs outperform in challenging scenarios.

NeurIPS Conference 2025 Conference Paper

Data Efficient Adaptation in Large Language Models via Continuous Low-Rank Fine-Tuning

  • Xiao Han
  • ZIMO ZHAO
  • Wanyu Wang
  • Maolin Wang
  • Zitao Liu
  • Yi Chang
  • Xiangyu Zhao

Recent advancements in Large Language Models (LLMs) have emphasized the critical role of fine-tuning (FT) techniques in adapting LLMs to specific tasks, especially when retraining from scratch is computationally infeasible. Fine-tuning enables LLMs to leverage task- or domain-specific data, producing models that more effectively meet the requirements of targeted applications. However, conventional FT approaches often suffer from catastrophic forgetting and suboptimal data efficiency, limiting their real-world applicability. To address these challenges, this paper proposes DEAL, a novel framework that integrates Low-Rank Adaptation (LoRA) with a continuous fine-tuning strategy. By incorporating knowledge retention and adaptive parameter update modules, the framework mitigates the limitations of existing FT methods while maintaining efficiency. Experiments on 15 diverse datasets show that DEAL consistently outperforms baseline methods, yielding substantial gains in task accuracy and resource efficiency. These findings demonstrate the potential of our approach to advance continual adaptation in LLMs by enhancing task performance while improving resource efficiency. The source code is publicly available at https: //github. com/Applied-Machine-Learning-Lab/DEAL.

ICML Conference 2025 Conference Paper

Eigen Analysis of Conjugate Kernel and Neural Tangent Kernel

  • Xiangchao Li
  • Xiao Han
  • Qing Yang

In this paper, we investigate deep feedforward neural networks with random weights. The input data matrix $\boldsymbol{X}$ is drawn from a Gaussian mixture model. We demonstrate that certain eigenvalues of the conjugate kernel and neural tangent kernel may lie outside the support of their limiting spectral measures in the high-dimensional regime. The existence and asymptotic positions of such isolated eigenvalues are rigorously analyzed. Furthermore, we provide a precise characterization of the entrywise limit of the projection matrix onto the eigenspace associated with these isolated eigenvalues. Our findings reveal that the eigenspace captures inherent group features present in $\boldsymbol{X}$. This study offers a quantitative analysis of how group features from the input data evolve through hidden layers in randomly weighted neural networks.

AAAI Conference 2025 Conference Paper

GARLIC: GPT-Augmented Reinforcement Learning with Intelligent Control for Vehicle Dispatching

  • Xiao Han
  • Zijian Zhang
  • Xiangyu Zhao
  • Yuanshao Zhu
  • Guojiang Shen
  • Xiangjie Kong
  • Xuetao Wei
  • Liqiang Nie

As urban residents demand higher travel quality, vehicle dispatch has become a critical component of online ride-hailing services. However, current vehicle dispatch systems struggle to navigate the complexities of urban traffic dynamics, including unpredictable traffic conditions, diverse driver behaviors, and fluctuating supply and demand patterns. These challenges have resulted in travel difficulties for passengers in certain areas, while many drivers in other areas are unable to secure orders, leading to a decline in the overall quality of urban transportation services. To address these issues, this paper introduces GARLIC: a framework of GPT-Augmented Reinforcement Learning with Intelligent Control for vehicle dispatching. GARLIC utilizes multiview graphs to capture hierarchical traffic states, and learns a dynamic reward function that accounts for individual driving behaviors. The framework further integrates a GPT model trained with a custom loss function to enable high-precision predictions and optimize dispatching policies in real-world scenarios. Experiments conducted on two real-world datasets demonstrate that GARLIC effectively aligns with driver behaviors while reducing the empty load rate of vehicles.

IJCAI Conference 2025 Conference Paper

Let’s Group: A Plug-and-Play SubGraph Learning Method for Memory-Efficient Spatio-Temporal Graph Modeling

  • Wenchao Weng
  • Hanyu Jiang
  • Mei Wu
  • Xiao Han
  • Haidong Gao
  • Guojiang Shen
  • Xiangjie Kong

Spatio-temporal graph modeling is widely applied to spatio-temporal data, analyzing the relationships between data to achieve accurate predictions. However, despite the excellent predictive performance of increasingly complex models, their intricate architectures result in significant memory overhead and computational complexity when handling spatio-temporal data, which limits their practical applications. To address these challenges, we propose a plug-and-play SubGraph Learning (SGL) method to reduce the memory overhead without compromising performance. Specifically, we introduce a SubGraph Partition Module (SGPM), which leverages a set of learnable memory vectors to select node groups with similar features from the graph, effectively partitioning the graph into smaller subgraphs. Noting that partitioning the graph may lead to feature redundancy, as overlapping information across subgraphs can occur. To overcome this, we design a SubGraph Feature Aggregation Module (SGFAM), which mitigates redundancy by averaging node features from different subgraphs. Experiments on four traffic network datasets of various scales demonstrate that SGL can significantly reduce memory overhead, achieving up to a 56. 4\% reduction in average GPU memory overhead, while maintaining robust prediction performance. The source code is available at https: //github. com/wengwenchao123/SubGraph-Learning.

TMLR Journal 2025 Journal Article

MarDini: Masked Auto-regressive Diffusion for Video Generation at Scale

  • Haozhe Liu
  • Shikun Liu
  • Zijian Zhou
  • Mengmeng Xu
  • Yanping Xie
  • Xiao Han
  • Juan Camilo Perez
  • Ding Liu

We introduce MarDini, a new family of video diffusion models that integrate the advantages of masked auto-regression (MAR) into a unified diffusion model (DM) framework. Here, MAR handles temporal planning, while DM focuses on spatial generation in an asymmetric network design: i) a MAR-based planning model containing most of the parameters generates planning signals for each masked frame using low-resolution input; ii) a lightweight generation model uses these signals to produce high-resolution frames via diffusion de-noising. MarDini’s MAR enables video generation conditioned on any number of masked frames at any frame positions: a single model can handle video interpolation (e.g., masking middle frames), image-to-video generation (e.g., masking from the second frame onward), and video expansion (e.g., masking half the frames). The efficient design allocates most of the computational resources to the low-resolution planning model, making computationally expensive but important spatio-temporal attention feasible at scale. MarDini sets a new state-of-the-art for video interpolation; meanwhile, within few inference steps, it efficiently generates videos on par with those of much more expensive advanced image-to-video models.

ICLR Conference 2025 Conference Paper

Root Cause Analysis of Anomalies in Multivariate Time Series through Granger Causal Discovery

  • Xiao Han
  • Saima Absar
  • Lu Zhang
  • Shuhan Yuan

Identifying the root causes of anomalies in multivariate time series is challenging due to the complex dependencies among the series. In this paper, we propose a comprehensive approach called AERCA that inherently integrates Granger causal discovery with root cause analysis. By defining anomalies as interventions on the exogenous variables of time series, AERCA not only learns the Granger causality among time series but also explicitly models the distributions of exogenous variables under normal conditions. AERCA then identifies the root causes of anomalies by highlighting exogenous variables that significantly deviate from their normal states. Experiments on multiple synthetic and real-world datasets demonstrate that AERCA can accurately capture the causal relationships among time series and effectively identify the root causes of anomalies.

NeurIPS Conference 2024 Conference Paper

G3: An Effective and Adaptive Framework for Worldwide Geolocalization Using Large Multi-Modality Models

  • Pengyue Jia
  • Yiding Liu
  • Xiaopeng Li
  • Yuhao Wang
  • Yantong Du
  • Xiao Han
  • Xuetao Wei
  • Shuaiqiang Wang

Worldwide geolocalization aims to locate the precise location at the coordinate level of photos taken anywhere on the Earth. It is very challenging due to 1) the difficulty of capturing subtle location-aware visual semantics, and 2) the heterogeneous geographical distribution of image data. As a result, existing studies have clear limitations when scaled to a worldwide context. They may easily confuse distant images with similar visual contents, or cannot adapt to various locations worldwide with different amounts of relevant data. To resolve these limitations, we propose G3, a novel framework based on Retrieval-Augmented Generation (RAG). In particular, G3 consists of three steps, i. e. , G eo-alignment, G eo-diversification, and G eo-verification to optimize both retrieval and generation phases of worldwide geolocalization. During Geo-alignment, our solution jointly learns expressive multi-modal representations for images, GPS and textual descriptions, which allows us to capture location-aware semantics for retrieving nearby images for a given query. During Geo-diversification, we leverage a prompt ensembling method that is robust to inconsistent retrieval performance for different image queries. Finally, we combine both retrieved and generated GPS candidates in Geo-verification for location prediction. Experiments on two well-established datasets IM2GPS3k and YFCC4k verify the superiority of G3 compared to other state-of-the-art methods. Our code is available online https: //github. com/Applied-Machine-Learning-Lab/G3 for reproduction.

JMLR Journal 2024 Journal Article

Individual-centered Partial Information in Social Networks

  • Xiao Han
  • Y. X. Rachel Wang
  • Qing Yang
  • Xin Tong

In statistical network analysis, we often assume either the full network is available or multiple subgraphs can be sampled to estimate various global properties of the network. However, in a real social network, people frequently make decisions based on their local view of the network alone. Here, we consider a partial information framework that characterizes the local network centered at a given individual by path length $L$ and gives rise to a partial adjacency matrix. Under $L=2$, we focus on the problem of (global) community detection using the popular stochastic block model (SBM) and its degree-corrected variant (DCSBM). We derive theoretical properties of the eigenvalues and eigenvectors from the signal term of the partial adjacency matrix and propose new spectral-based community detection algorithms that achieve consistency under appropriate conditions. Our analysis also allows us to propose a new centrality measure that assesses the importance of an individual's partial information in determining global community structure. Using simulated and real networks, we demonstrate the performance of our algorithms and compare our centrality measure with other popular alternatives to show it captures unique nodal information. Our results illustrate that the partial information framework enables us to compare the viewpoints of different individuals regarding the global structure. [abs] [ pdf ][ bib ] &copy JMLR 2024. ( edit, beta )

IJCAI Conference 2024 Conference Paper

KDDC: Knowledge-Driven Disentangled Causal Metric Learning for Pre-Travel Out-of-Town Recommendation

  • Yinghui Liu
  • Guojiang Shen
  • Chengyong Cui
  • Zhenzhen Zhao
  • Xiao Han
  • Jiaxin Du
  • Xiangyu Zhao
  • Xiangjie Kong

Pre-travel recommendation is developed to provide a variety of out-of-town Point-of-Interests (POIs) for users planning to travel away from their hometowns but have not yet decided on their destination. Existing out-of-town recommender systems work on constructing users' latent preferences and inferring travel intentions from their check-in sequences. However, there are still two challenges that hamper the performance of these approaches: i) Users' interactive data (including hometown and out-of-town check-ins) tend to be rare, and while candidate POIs that come from different regions contain various semantic information; ii) The causes for user check-in include not only interest but also conformity, which are easily entangled and overlooked. To fill these gaps, we propose a Knowledge-Driven Disentangled Causal metric learning framework (KDDC) that mitigates interaction data sparsity by enhancing POI semantic representation and considers the distributions of two causes (i. e. , conformity and interest) for pre-travel recommendation. Specifically, we pretrain a constructed POI attribute knowledge graph through a segmented interaction method and POI semantic information is aggregated via relational heterogeneity. In addition, we devise a disentangled causal metric learning to model and infer userrelated representations. Extensive experiments on two real-world nationwide datasets display the consistent superiority of our KDDC over state-of-theart baselines.

AAAI Conference 2024 Conference Paper

VIGC: Visual Instruction Generation and Correction

  • Bin Wang
  • Fan Wu
  • Xiao Han
  • Jiahui Peng
  • Huaping Zhong
  • Pan Zhang
  • Xiaoyi Dong
  • Weijia Li

The integration of visual encoders and large language models (LLMs) has driven recent progress in multimodal large language models (MLLMs). However, the scarcity of high-quality instruction-tuning data for vision-language tasks remains a challenge. The current leading paradigm, such as LLaVA, relies on language-only GPT-4 to generate data, which requires pre-annotated image captions and detection bounding boxes, suffering from understanding image details. A practical solution to this problem would be to utilize the available multimodal large language models to generate instruction data for vision-language tasks. However, it's worth noting that the currently accessible MLLMs are not as powerful as their LLM counterparts, as they tend to produce inadequate responses and generate false information. As a solution for addressing the current issue, this paper proposes the Visual Instruction Generation and Correction (VIGC) framework that enables multimodal large language models to generate instruction-tuning data and progressively enhance its quality on-the-fly. Specifically, Visual Instruction Generation (VIG) guides the vision-language model to generate diverse instruction-tuning data. To ensure generation quality, Visual Instruction Correction (VIC) adopts an iterative update mechanism to correct any inaccuracies in data produced by VIG, effectively reducing the risk of hallucination. Leveraging the diverse, high-quality data generated by VIGC, we finetune mainstream models and validate data quality based on various evaluations. Experimental results demonstrate that VIGC not only compensates for the shortcomings of language-only data generation methods, but also effectively enhances the benchmark performance. The models, datasets, and code are available at https://opendatalab.github.io/VIGC

EAAI Journal 2023 Journal Article

A Fermatean fuzzy Fine–Kinney for occupational risk evaluation using extensible MARCOS with prospect theory

  • Weizhong Wang
  • Xiao Han
  • Weiping Ding
  • Qun Wu
  • Xiaoqing Chen
  • Muhammet Deveci

The extant Fine–Kinney frameworks are insufficient to tackle the risk evaluation problem with Fermatean fuzzy information, in which the prioritization degrees and psychological characteristics of decision-makers are considered. Hence, this study develops a hybrid Fine–Kinney-based occupational risk evaluation framework with an extended Fermatean fuzzy MARCOS method (measurement of alternatives and ranking to Compromise solution). Such a MARCOS method improves conventional MARCOS by integrating Fermatean fuzzy prioritized weighted average operator and prospect theory. This improved method has the capability to handle the occupational risk analysis problem with Fermatean fuzzy data in the risk ranking procedure considering the prioritization degrees and bounded rational behavior of decision-makers. In addition, the Fermatean fuzzy numbers-based risk rating scales are established to transform the linguistic risk scores from decision-makers, it allows for handling uncertain risk rating information from decision-makers more effectively. Further, the improved MARCOS method is incorporated into the occupational risk ranking procedure, as it considers the decision-maker’s prioritization relationships among decision-makers and their reference point effect in occupational risk priority calculation. After that, an occupational risk analysis case for construction operations is selected to test the applicability and validity of the proposed framework. The result indicates that the occupational risk OR6 (Back injury) is the most serious risk with the lowest utility function value (-0. 324), and OR7 (Tendinitis) is the least severe risk with the highest utility function value (0. 682). Finally, sensitivity exploration and comparative study are implemented to further test the advantages of the developed framework.

NeurIPS Conference 2023 Conference Paper

HeadSculpt: Crafting 3D Head Avatars with Text

  • Xiao Han
  • Yukang Cao
  • Kai Han
  • Xiatian Zhu
  • Jiankang Deng
  • Yi-Zhe Song
  • Tao Xiang
  • Kwan-Yee K. Wong

Recently, text-guided 3D generative methods have made remarkable advancements in producing high-quality textures and geometry, capitalizing on the proliferation of large vision-language and image diffusion models. However, existing methods still struggle to create high-fidelity 3D head avatars in two aspects: (1) They rely mostly on a pre-trained text-to-image diffusion model whilst missing the necessary 3D awareness and head priors. This makes them prone to inconsistency and geometric distortions in the generated avatars. (2) They fall short in fine-grained editing. This is primarily due to the inherited limitations from the pre-trained 2D image diffusion models, which become more pronounced when it comes to 3D head avatars. In this work, we address these challenges by introducing a versatile coarse-to-fine pipeline dubbed HeadSculpt for crafting (i. e. , generating and editing) 3D head avatars from textual prompts. Specifically, we first equip the diffusion model with 3D awareness by leveraging landmark-based control and a learned textual embedding representing the back view appearance of heads, enabling 3D-consistent head avatar generations. We further propose a novel identity-aware editing score distillation strategy to optimize a textured mesh with a high-resolution differentiable rendering technique. This enables identity preservation while following the editing instruction. We showcase HeadSculpt's superior fidelity and editing capabilities through comprehensive experiments and comparisons with existing methods.

AAAI Conference 2023 Conference Paper

RLogist: Fast Observation Strategy on Whole-Slide Images with Deep Reinforcement Learning

  • Boxuan Zhao
  • Jun Zhang
  • Deheng Ye
  • Jian Cao
  • Xiao Han
  • Qiang Fu
  • Wei Yang

Whole-slide images (WSI) in computational pathology have high resolution with gigapixel size, but are generally with sparse regions of interest, which leads to weak diagnostic relevance and data inefficiency for each area in the slide. Most of the existing methods rely on a multiple instance learning framework that requires densely sampling local patches at high magnification. The limitation is evident in the application stage as the heavy computation for extracting patch-level features is inevitable. In this paper, we develop RLogist, a benchmarking deep reinforcement learning (DRL) method for fast observation strategy on WSIs. Imitating the diagnostic logic of human pathologists, our RL agent learns how to find regions of observation value and obtain representative features across multiple resolution levels, without having to analyze each part of the WSI at the high magnification. We benchmark our method on two whole-slide level classification tasks, including detection of metastases in WSIs of lymph node sections, and subtyping of lung cancer. Experimental results demonstrate that RLogist achieves competitive classification performance compared to typical multiple instance learning algorithms, while having a significantly short observation path. In addition, the observation path given by RLogist provides good decision-making interpretability, and its ability of reading path navigation can potentially be used by pathologists for educational/assistive purposes. Our code is available at: https://github.com/tencent-ailab/RLogist.

TMLR Journal 2023 Journal Article

RLTF: Reinforcement Learning from Unit Test Feedback

  • Jiate Liu
  • Yiqin Zhu
  • Kaiwen Xiao
  • Qiang Fu
  • Xiao Han
  • Yang Wei
  • Deheng Ye

The goal of program synthesis, or code generation, is to generate executable code based on given descriptions. Recently, there has been an increasing number of studies employing reinforcement learning (RL) to improve the performance of large language models (LLMs) for code. However, some of the current representative RL methods have only used offline frameworks, limiting the exploration of new sample spaces. Additionally, the utilization of unit test signals is limited, not accounting for specific error locations within the code. To address these issues, we proposed RLTF, i.e., Reinforcement Learning from Unit Test Feedback, a novel online RL framework with unit test feedback of multi-granularity for refining code LLMs. Our approach generates data in real-time during training and simultaneously utilizes fine-grained feedback signals to guide the model towards producing higher-quality code. Extensive experiments show that RLTF achieves state-of-the-art performance on the APPS and the MBPP benchmarks. Our code is available at: \url{https://github.com/Zyq-scut/RLTF}.

NeurIPS Conference 2022 Conference Paper

SCL-WC: Cross-Slide Contrastive Learning for Weakly-Supervised Whole-Slide Image Classification

  • Xiyue Wang
  • Jinxi Xiang
  • Jun Zhang
  • Sen Yang
  • Zhongyi Yang
  • Ming-Hui Wang
  • Jing Zhang
  • Wei Yang

Weakly-supervised whole-slide image (WSI) classification (WSWC) is a challenging task where a large number of unlabeled patches (instances) exist within each WSI (bag) while only a slide label is given. Despite recent progress for the multiple instance learning (MIL)-based WSI analysis, the major limitation is that it usually focuses on the easy-to-distinguish diagnosis-positive regions while ignoring positives that occupy a small ratio in the entire WSI. To obtain more discriminative features, we propose a novel weakly-supervised classification method based on cross-slide contrastive learning (called SCL-WC), which depends on task-agnostic self-supervised feature pre-extraction and task-specific weakly-supervised feature refinement and aggregation for WSI-level prediction. To enable both intra-WSI and inter-WSI information interaction, we propose a positive-negative-aware module (PNM) and a weakly-supervised cross-slide contrastive learning (WSCL) module, respectively. The WSCL aims to pull WSIs with the same disease types closer and push different WSIs away. The PNM aims to facilitate the separation of tumor-like patches and normal ones within each WSI. Extensive experiments demonstrate state-of-the-art performance of our method in three different classification tasks (e. g. , over 2% of AUC in Camelyon16, 5% of F1 score in BRACS, and 3% of AUC in DiagSet). Our method also shows superior flexibility and scalability in weakly-supervised localization and semi-supervised classification experiments (e. g. , first place in the BRIGHT challenge). Our code will be available at https: //github. com/Xiyue-Wang/SCL-WC.

YNICL Journal 2021 Journal Article

A deep learning algorithm for automatic detection and classification of acute intracranial hemorrhages in head CT scans

  • Xiyue Wang
  • Tao Shen
  • Sen Yang
  • Jun Lan
  • Yanming Xu
  • Minghui Wang
  • Jing Zhang
  • Xiao Han

Acute Intracranial hemorrhage (ICH) is a life-threatening disease that requires emergency medical attention, which is routinely diagnosed using non-contrast head CT imaging. The diagnostic accuracy of acute ICH on CT varies greatly among radiologists due to the difficulty of interpreting subtle findings and the time pressure associated with the ever-increasing workload. The use of artificial intelligence technology may help automate the process and assist radiologists for more prompt and better decision-making. In this work, we design a deep learning approach that mimics the interpretation process of radiologists, and combines a 2D CNN model and two sequence models to achieve accurate acute ICH detection and subtype classification. Being developed using the extensive 2019-RSNA Brain CT Hemorrhage Challenge dataset with over 25000 CT scans, our deep learning algorithm can accurately classify the acute ICH and its five subtypes with AUCs of 0.988 (ICH), 0.984 (EDH), 0.992 (IPH), 0.996 (IVH), 0.985 (SAH), and 0.983 (SDH), respectively, reaching the accuracy level of expert radiologists. Our method won 1st place among 1345 teams from 75 countries in the RSNA challenge. We have further evaluated our algorithm on two independent external validation datasets with 75 and 491 CT scans, respectively, and our method maintained high AUCs of 0.964 and 0.949 for acute ICH detection. These results have demonstrated the high performance and robust generalization ability of our proposed method, which makes it a useful second-read or triage tool that can facilitate routine clinical applications.

AAAI Conference 2021 Conference Paper

Diagnose Like A Pathologist: Weakly-Supervised Pathologist-Tree Network for Slide-Level Immunohistochemical Scoring

  • Zhen Chen
  • Jun Zhang
  • Shuanlong Che
  • Junzhou Huang
  • Xiao Han
  • Yixuan Yuan

The immunohistochemistry (IHC) test of biopsy tissue is crucial to develop targeted treatment and evaluate prognosis for cancer patients. The IHC staining slide is usually digitized into the whole-slide image (WSI) with gigapixels for quantitative image analysis. To perform a whole image prediction (e. g. , IHC scoring, survival prediction, and cancer grading) from this kind of high-dimensional image, algorithms are often developed based on multi-instance learning (MIL) framework. However, the multi-scale information of WSI and the associations among instances are not well explored in existing MIL based studies. Inspired by the fact that pathologists jointly analyze visual fields at multiple powers of objective for diagnostic predictions, we propose a Pathologist-Tree Network (PTree-Net) to sparsely model the WSI efficiently in multi-scale manner. Specifically, we propose a Focal-Aware Module (FAM) that can approximately estimate diagnosis-related regions with an extractor trained using the thumbnail of WSI. With the initial diagnosis-related regions, we hierarchically model the multi-scale patches in a tree structure, where both the global and local information can be captured. To explore this tree structure in an end-to-end network, we propose a patch Relevance-enhanced Graph Convolutional Network (RGCN) to explicitly model the correlations of adjacent parent-child nodes, accompanied by patch relevance to exploit the implicit contextual information among distant nodes. In addition, tree-based self-supervision is devised to improve representation learning and suppress irrelevant instances adaptively. Extensive experiments are performed on a large-scale IHC HER2 dataset. The ablation study confirms the effectiveness of our design, and our approach outperforms state-of-the-art by a large margin.

AAAI Conference 2021 Conference Paper

Label Confusion Learning to Enhance Text Classification Models

  • Biyang Guo
  • Songqiao Han
  • Xiao Han
  • Hailiang Huang
  • Ting Lu

Representing a true label as a one-hot vector is a common practice in training text classification models. However, the one-hot representation may not adequately reflect the relation between the instances and labels, as labels are often not completely independent and instances may relate to multiple labels in practice. The inadequate one-hot representations tend to train the model to be over-confident, which may result in arbitrary prediction and model overfitting, especially for confused datasets (datasets with very similar labels) or noisy datasets (datasets with labeling errors). While training models with label smoothing (LS) can ease this problem in some degree, it still fails to capture the realistic relation among labels. In this paper, we propose a novel Label Confusion Model (LCM) as an enhancement component to current popular text classification models. LCM can learn label confusion to capture semantic overlap among labels by calculating the similarity between instances and labels during training and generate a better label distribution to replace the original one-hot label vector, thus improving the final classification performance. Extensive experiments on five text classification benchmark datasets reveal the effectiveness of LCM for several widely used deep learning classification models. Further experiments also verify that LCM is especially helpful for confused or noisy datasets and superior to the label smoothing method.

AAAI Conference 2021 Conference Paper

Minimizing Labeling Cost for Nuclei Instance Segmentation and Classification with Cross-domain Images and Weak Labels

  • Siqi Yang
  • Jun Zhang
  • Junzhou Huang
  • Brian C. Lovell
  • Xiao Han

Nucleus instance segmentation and classification in histopathological images is an essential prerequisite in pathology diagnosis/prognosis. However, nucleus annotations (e. g. , segmentation and labeling) require domain experts, and annotating nuclei at pixel-level is time-consuming and labor-intensive. Moreover, nuclei from different cancer types vary in shapes and appearances. These inter-cancer variations require careful annotations for specific cancer types. Therefore, to minimize the labeling cost, we propose a novel application that considers each cancer type as an individual domain and apply domain adaptation techniques to improve the segmentation/classification performance among different cancer types. Unlike the previous studies that focus on unsupervised or weakly-supervised domain adaptation independently, we would like to discover what kinds of labeling can achieve the most cost-effective domain adaptation performance in nucleus instance segmentation and classification. Specifically, we propose a unified framework that is applicable to different level annotations: no annotations, image-level, and point-level annotations. Cyclic adaptation with pseudo labels and adversarial discriminator are utilized for unsupervised domain alignment. Image-level or point-level annotations are additionally adopted to supervise the nucleus classification and refine the pseudo labels. Experiments demonstrate the effectiveness and efficacy of the proposed framework (jointly using unsupervised and weakly supervised learning) on adapting the segmentation and classification model from one cancer type to 18 other cancer types.

TCS Journal 2020 Journal Article

Distributed representation of knowledge graphs with subgraph-aware proximity

  • Xiao Han
  • Chunhong Zhang
  • Chenchen Guo
  • Yang Ji
  • Zheng Hu

The distributed representation of Knowledge graphs (KGs) aims to embed the original KG into a low-dimensional embedding vector space, so as to facilitate the completion of KGs as well as the application of KGs in other AI fields. Most existing models preserve certain proximity property of KGs in the embedding space, such as the first/second-order proximity and the sequence-aware higher-order proximity. However, the ubiquitous similarity relationship between different sequences has rarely been discussed. In this paper, we propose a unified framework to preserve the subgraph-aware proximity in the embedding space, holding that the sequences within a subgraph generally imply a similar pattern. Especially, according to the composition and structure of KG sequences, we provide three methods for computing the embeddings of KG sequences: 1) Simply adding the involved relations of the KG sequences in a relation subgraph; 2) Recurrent neural network for the KG sequences in a complete subgraph; 3) Dilated recurrent neural network to match the special structure of the KG sequences in a complete subgraph. Empirically, we evaluate the proposed framework on the KG completion tasks of link prediction and entity classification. The results show that our framework outperforms the baselines by preserving the subgraph-aware proximity. Especially, exploring the special structure of KG sequences can further improve the performance.

AAAI Conference 2018 Conference Paper

Geographic Differential Privacy for Mobile Crowd Coverage Maximization

  • Leye Wang
  • Gehua Qin
  • Dingqi Yang
  • Xiao Han
  • Xiaojuan Ma

For real-world mobile applications such as location-based advertising and spatial crowdsourcing, a key to success is targeting mobile users that can maximally cover certain locations in a future period. To find an optimal group of users, existing methods often require information about users’ mobility history, which may cause privacy breaches. In this paper, we propose a method to maximize mobile crowd’s future location coverage under a guaranteed location privacy protection scheme. In our approach, users only need to upload one of their frequently visited locations, and more importantly, the uploaded location is obfuscated using a geographic differential privacy policy. We propose both analytic and practical solutions to this problem. Experiments on real user mobility datasets show that our method significantly outperforms the state-of-the-art geographic differential privacy methods by achieving a higher coverage under the same level of privacy protection.

TIST Journal 2017 Journal Article

SPACE-TA

  • Leye Wang
  • Daqing Zhang
  • Dingqi Yang
  • Animesh Pathak
  • Chao Chen
  • Xiao Han
  • Haoyi Xiong
  • Yasha Wang

Data quality and budget are two primary concerns in urban-scale mobile crowdsensing. Traditional research on mobile crowdsensing mainly takes sensing coverage ratio as the data quality metric rather than the overall sensed data error in the target-sensing area. In this article, we propose to leverage spatiotemporal correlations among the sensed data in the target-sensing area to significantly reduce the number of sensing task assignments. In particular, we exploit both intradata correlations within the same type of sensed data and interdata correlations among different types of sensed data in the sensing task. We propose a novel crowdsensing task allocation framework called SPACE-TA (SPArse Cost-Effective Task Allocation), combining compressive sensing, statistical analysis, active learning, and transfer learning, to dynamically select a small set of subareas for sensing in each timeslot (cycle), while inferring the data of unsensed subareas under a probabilistic data quality guarantee. Evaluations on real-life temperature, humidity, air quality, and traffic monitoring datasets verify the effectiveness of SPACE-TA. In the temperature-monitoring task leveraging intradata correlations, SPACE-TA requires data from only 15.5% of the subareas while keeping the inference error below 0.25°C in 95% of the cycles, reducing the number of sensed subareas by 18.0% to 26.5% compared to baselines. When multiple tasks run simultaneously, for example, for temperature and humidity monitoring, SPACE-TA can further reduce ∼10% of the sensed subareas by exploiting interdata correlations.

YNIMG Journal 2013 Journal Article

Seizure localization using three-dimensional surface projections of intracranial EEG power

  • Hyang Woon Lee
  • Mark W. Youngblood
  • Pue Farooque
  • Xiao Han
  • Stephen Jhun
  • William C. Chen
  • Irina Goncharova
  • Kenneth Vives

Intracranial EEG (icEEG) provides a critical road map for epilepsy surgery but it has become increasingly difficult to interpret as technology has allowed the number of icEEG channels to grow. Borrowing methods from neuroimaging, we aimed to simplify data analysis and increase consistency between reviewers by using 3D surface projections of intracranial EEG poweR (3D-SPIER). We analyzed 139 seizures from 48 intractable epilepsy patients (28 temporal and 20 extratemporal) who had icEEG recordings, epilepsy surgery, and at least one year of post-surgical follow-up. We coregistered and plotted icEEG β frequency band signal power over time onto MRI-based surface renderings for each patient, to create color 3D-SPIER movies. Two independent reviewers interpreted the icEEG data using visual analysis vs. 3D-SPIER, blinded to any clinical information. Overall agreement rates between 3D-SPIER and icEEG visual analysis or surgery were about 90% for side of seizure onset, 80% for lobe, and just under 80% for sublobar localization. These agreement rates were improved when flexible thresholds or frequency ranges were allowed for 3D-SPIER, especially for sublobar localization. Interestingly, agreement was better for patients with good surgical outcome than for patients with poor outcome. Localization using 3D-SPIER was measurably faster and considered qualitatively easier to interpret than visual analysis. These findings suggest that 3D-SPIER could be an improved diagnostic method for presurgical seizure localization in patients with intractable epilepsy and may also be useful for mapping normal brain function.

YNIMG Journal 2009 Journal Article

MRI-derived measurements of human subcortical, ventricular and intracranial brain volumes: Reliability effects of scan sessions, acquisition sequences, data analyses, scanner upgrade, scanner vendors and field strengths

  • Jorge Jovicich
  • Silvester Czanner
  • Xiao Han
  • David Salat
  • Andre van der Kouwe
  • Brian Quinn
  • Jenni Pacheco
  • Marilyn Albert

Automated MRI-derived measurements of in-vivo human brain volumes provide novel insights into normal and abnormal neuroanatomy, but little is known about measurement reliability. Here we assess the impact of image acquisition variables (scan session, MRI sequence, scanner upgrade, vendor and field strengths), FreeSurfer segmentation pre-processing variables (image averaging, B1 field inhomogeneity correction) and segmentation analysis variables (probabilistic atlas) on resultant image segmentation volumes from older (n =15, mean age 69. 5) and younger (both n =5, mean ages 34 and 36. 5) healthy subjects. The variability between hippocampal, thalamic, caudate, putamen, lateral ventricular and total intracranial volume measures across sessions on the same scanner on different days is less than 4. 3% for the older group and less than 2. 3% for the younger group. Within-scanner measurements are remarkably reliable across scan sessions, being minimally affected by averaging of multiple acquisitions, B1 correction, acquisition sequence (MPRAGE vs. multi-echo-FLASH), major scanner upgrades (Sonata–Avanto, Trio–TrioTIM), and segmentation atlas (MPRAGE or multi-echo-FLASH). Volume measurements across platforms (Siemens Sonata vs. GE Signa) and field strengths (1. 5 T vs. 3 T) result in a volume difference bias but with a comparable variance as that measured within-scanner, implying that multi-site studies may not necessarily require a much larger sample to detect a specific effect. These results suggest that volumes derived from automated segmentation of T1-weighted structural images are reliable measures within the same scanner platform, even after upgrades; however, combining data across platform and across field-strength introduces a bias that should be considered in the design of multi-site studies, such as clinical drug trials. The results derived from the young groups (scanner upgrade effects and B1 inhomogeneity correction effects) should be considered as preliminary and in need for further validation with a larger dataset.

YNIMG Journal 2006 Journal Article

Reliability of MRI-derived measurements of human cerebral cortical thickness: The effects of field strength, scanner upgrade and manufacturer

  • Xiao Han
  • Jorge Jovicich
  • David Salat
  • Andre van der Kouwe
  • Brian Quinn
  • Silvester Czanner
  • Evelina Busa
  • Jenni Pacheco

In vivo MRI-derived measurements of human cerebral cortex thickness are providing novel insights into normal and abnormal neuroanatomy, but little is known about their reliability. We investigated how the reliability of cortical thickness measurements is affected by MRI instrument-related factors, including scanner field strength, manufacturer, upgrade and pulse sequence. Several data processing factors were also studied. Two test–retest data sets were analyzed: 1) 15 healthy older subjects scanned four times at 2-week intervals on three scanners; 2) 5 subjects scanned before and after a major scanner upgrade. Within-scanner variability of global cortical thickness measurements was <0. 03 mm, and the point-wise standard deviation of measurement error was approximately 0. 12 mm. Variability was 0. 15 mm and 0. 17 mm in average, respectively, for cross-scanner (Siemens/GE) and cross-field strength (1. 5 T/3 T) comparisons. Scanner upgrade did not increase variability nor introduce bias. Measurements across field strength, however, were slightly biased (thicker at 3 T). The number of (single vs. multiple averaged) acquisitions had a negligible effect on reliability, but the use of a different pulse sequence had a larger impact, as did different parameters employed in data processing. Sample size estimates indicate that regional cortical thickness difference of 0. 2 mm between two different groups could be identified with as few as 7 subjects per group, and a difference of 0. 1 mm could be detected with 26 subjects per group. These results demonstrate that MRI-derived cortical thickness measures are highly reliable when MRI instrument and data processing factors are controlled but that it is important to consider these factors in the design of multi-site or longitudinal studies, such as clinical drug trials.

YNIMG Journal 2004 Journal Article

Cortical surface segmentation and mapping

  • Duygu Tosun
  • Maryam E. Rettmann
  • Xiao Han
  • Xiaodong Tao
  • Chenyang Xu
  • Susan M. Resnick
  • Dzung L. Pham
  • Jerry L. Prince

Segmentation and mapping of the human cerebral cortex from magnetic resonance (MR) images plays an important role in neuroscience and medicine. This paper describes a comprehensive approach for cortical reconstruction, flattening, and sulcal segmentation. Robustness to imaging artifacts and anatomical consistency are key achievements in an overall approach that is nearly fully automatic and computationally fast. Results demonstrating the application of this approach to a study of cortical thickness changes in aging are presented.

YNIMG Journal 2004 Journal Article

CRUISE: Cortical reconstruction using implicit surface evolution

  • Xiao Han
  • Dzung L. Pham
  • Duygu Tosun
  • Maryam E. Rettmann
  • Chenyang Xu
  • Jerry L. Prince

Segmentation and representation of the human cerebral cortex from magnetic resonance (MR) images play an important role in neuroscience and medicine. A successful segmentation method must be robust to various imaging artifacts and produce anatomically meaningful and consistent cortical representations. A method for the automatic reconstruction of the inner, central, and outer surfaces of the cerebral cortex from T1-weighted MR brain images is presented. The method combines a fuzzy tissue classification method, an efficient topology correction algorithm, and a topology-preserving geometric deformable surface model (TGDM). The algorithm is fast and numerically stable, and yields accurate brain surface reconstructions that are guaranteed to be topologically correct and free from self-intersections. Validation results on real MR data are presented to demonstrate the performance of the method.

YNIMG Journal 2002 Journal Article

Automated Sulcal Segmentation Using Watersheds on the Cortical Surface

  • Maryam E. Rettmann
  • Xiao Han
  • Chenyang Xu
  • Jerry L. Prince

The human cortical surface is a highly complex, folded structure. Sulci, the spaces between the folds, define location on the cortex and provide a parcellation into anatomically distinct areas. A topic that has recently received increased attention is the segmentation of these sulci from magnetic resonance images, with most work focusing on extracting either the sulcal spaces between the folds or curve representations of sulci. Unlike these methods, we propose a technique that extracts actual regions of the cortical surface that surround sulci, which we call “sulcal regions. ” The method is based on a watershed algorithm applied to a geodesic depth measure on the cortical surface. A well-known problem with the watershed algorithm is a tendency toward oversegmentation, meaning that a single region is segmented as several pieces. To address this problem, we propose a postprocessing algorithm that merges appropriate segments from the watershed algorithm. The sulcal regions are then manually labeled by simply selecting the appropriate regions with a mouse click and a preliminary study of sulcal depth is reported. Finally, a scheme is presented for computing a complete parcellation of the cortical surface.

v2026.09.13