Arrow Research search

Author name cluster

Han Gao

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

11 papers
2 author rows

Possible papers

11

EAAI Journal 2026 Journal Article

Cascaded U-Net diffusion refiner for deformation prediction in hot strip rolling

  • Han Gao
  • Shanhong Cao
  • Xueqi Dong
  • Xu Li
  • Feng Luan
  • Dianhua Zhang

In the hot strip roughing process, Vertical-Horizontal rolling induces “fishtailing” deformation that leads to yield-reducing profile defects. While finite element method (FEM) simulations accurately model the complex elastoplastic deformation mechanisms underlying this phenomenon, their computational intensity impedes real-time process optimization and large-scale parametric analysis. To bridge this critical gap between accuracy and efficiency, we propose a Cascaded U-Net Diffusion Refiner (CUDR) framework for elastoplastic deformation prediction in hot rolling. The core design of this framework lies in the collaborative operation of two components: the U-Net first performs fast coarse prediction of deformation to provide a “warm start” foundation, and then the diffusion model conducts lightweight denoising refinement on the coarse prediction results. This refinement step primarily aims to suppress unphysical local fluctuations in the U-Net's predictions, thereby further enhancing the overall prediction precision. Validated on three orthogonal-sampled datasets with varying mesh resolutions, the CUDR reduces prediction errors by 26. 1% in Euclidean Mean Absolute Error and 28. 5% in Euclidean Mean Peak Absolute Error compared to the standalone U-Net. Moreover, Fourier-based spectral verification confirms that the framework suppresses unphysical local fluctuations. Critically, for fine-mesh cases, the CUDR achieves a 3900 times speedup over high-fidelity FEM simulations, making real-time deformation prediction feasible. This work demonstrates the substantial potential of generative diffusion models in advancing metal forming simulation, offering a new paradigm for balancing accuracy and efficiency in industrial manufacturing processes.

AAAI Conference 2026 Conference Paper

Diverse Human Driving Vehicle Simulation in Background Traffic for Autonomous Driving Tests

  • Wendi Li
  • Hao Wu
  • Han Gao
  • Bing Mao
  • Fengyuan Xu
  • Sheng Zhong

Realistic background traffic is critical to the simulation platforms for autonomous driving (AD) testing. Given that most vehicles in reality are driven by human beings, introducing human driving (HD) vehicles to the background traffic is necessary to be able to discover more problems of the tested AD vehicle in the simulation stage. However, existing methods rely on ad-hoc rules or data-driven training to mimic partial human driver behaviors, which are not comprehensive and lack transparency. In this work, we design a smart human driving vehicle simulator HDSim which is empowered by cognitively inspired modeling and AI models. HDSim enables diverse, realistic, and scalable HD traffic simulation on AD testing platforms like CARLA in a non-intrusive manner. There are two novel components in HDSim. First, we introduce a driver model to guide the generation of diverse human driving styles by using different combinations of latent cognitive factors in a hierarchy. Second, we design a Perception-Mediated Behavior Influence (PMBI) mechanism to use LLM-assisted perceptual transformations to indirectly fuse driving actions with driving styles. Experiments show that HDSim traffic can help simulation platforms like CARLA to reveal 68% more failures of tested AD vehicles, and the explainability of reported accidents is also improved.

YNIMG Journal 2026 Journal Article

Dynamic reorganization of functional connectome gradients reveals time-specific recovery patterns after stroke

  • Qingwen Chen
  • Tao Zhong
  • Jian Liu
  • Dajun Yan
  • Han Gao

Stroke, a leading cause of death and disability worldwide, severely disrupts brain functional organization and cognitive abilities. Previous research has mainly focused on discrete functional network changes post-stroke, but how stroke affects whole-brain functional hierarchy and its relationship to cognitive recovery remains poorly understood. In this exploratory longitudinal study, we used connectome gradient mapping in 33 patients with first-ever stroke and 21 healthy controls to examine how stroke affects large-scale functional network organization across an early post-stroke(∼2 weeks), a subacute stage (3 months), and a chronic stage (1 year). By projecting functional connectivity patterns onto a low-dimensional gradient space, we found that although overall gradient structure remained relatively stable at the group level, individual patients exhibited significant deviations (EDfunc) from the healthy topology, most prominently at the early post-stroke across visual, somatomotor, ventral attention, and control networks. Furthermore, EDfunc showed time-specific associations with cognitive functions: broad negative correlations with visuospatial attention in the early post-stroke, transitioning to more selective associations with motor and attention measures in the chronic stage. In addition, dynamic interhemispheric functional imbalances emerged in the subacute and chronic stages. Taken together, these findings provide preliminary, hypothesis-generating evidence for dynamic reorganization of whole-brain functional hierarchy following stroke, and suggest that connectome gradient analysis and EDfunc may offer a sensitive framework for monitoring recovery and informing individualized rehabilitation strategies, pending confirmation in larger multi-center cohorts.

IROS Conference 2025 Conference Paper

3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering

  • Rongtao Xu
  • Han Gao
  • Mingming Yu
  • Dong An 0002
  • Shunpeng Chen
  • Changwei Wang 0001
  • Li Guo 0004
  • Xiaodan Liang

With the growing need for diverse and scalable data in indoor scene tasks, such as question answering and dense captioning, we propose 3D-MoRe, a novel paradigm designed to generate large-scale 3D-language datasets by lever-aging the strengths of foundational models. The framework integrates key components, including multi-modal embedding, cross-modal interaction, and a language model decoder, to process natural language instructions and 3D scene data. This approach facilitates enhanced reasoning and response generation in complex 3D environments. Using the ScanNet 3D scene dataset, along with text annotations from ScanQA and ScanRefer, 3D-MoRe generates 62, 000 question-answer (QA) pairs and 73, 000 object descriptions across 1, 513 scenes. We also employ various data augmentation techniques and implement semantic filtering to ensure high-quality data. Experiments on ScanQA demonstrate that 3D-MoRe significantly outperforms state-of-the-art baselines, with the CIDEr score improving by 2. 15%. Similarly, on ScanRefer, our approach achieves a notable increase in CIDEr@0. 5 by 1. 84%, highlighting its effectiveness in both tasks. Our code and generated datasets will be publicly released to benefit the community, and both can be accessed on the https://3D-MoRe.github.io.

ICRA Conference 2024 Conference Paper

ASPIRe: An Informative Trajectory Planner with Mutual Information Approximation for Target Search and Tracking

  • Kangjie Zhou
  • Pengying Wu
  • Yao Su 0001
  • Han Gao
  • Ji Ma 0007
  • Hangxin Liu
  • Chang Liu 0002

This paper proposes an informative trajectory planning approach, namely, adaptive particle filter tree with sigma point-based mutual information reward approximation (ASPIRe), for mobile target search and tracking (SAT) in cluttered environments with limited sensing field of view. We develop a novel sigma point-based approximation to accurately estimate mutual information (MI) for general, non-Gaussian distributions utilizing particle representation of the belief state, while simultaneously maintaining high computational efficiency. Building upon the MI approximation, we develop the Adaptive Particle Filter Tree (APFT) approach with MI as the reward, which features belief state tree nodes for informative trajectory planning in continuous state and measurement spaces. An adaptive criterion is proposed in APFT to adjust the planning horizon based on the expected information gain. Simulations and physical experiments demonstrate that ASPIRe achieves real-time computation and outperforms benchmark methods in terms of both search efficiency and estimation accuracy.

IROS Conference 2024 Conference Paper

Risk-Aware Non-Myopic Motion Planner for Large-Scale Robotic Swarm Using CVaR Constraints

  • Xuru Yang
  • Yunze Hu
  • Han Gao
  • Kang Ding
  • Zhaoyang Li
  • Pingping Zhu
  • Ying Sun
  • Chang Liu

Swarm robotics has garnered significant attention due to its ability to accomplish elaborate and synchronized tasks. Existing methodologies for motion planning of swarm robotic systems mainly encounter difficulties in scalability and safety guarantee. To address these limitations, we propose a Risk-aware swarm mOtion planner using conditional ValuE-at-Risk (ROVER) that systematically navigates large-scale swarms through cluttered environments while ensuring safety. ROVER formulates a finite-time model predictive control (FTMPC) problem predicated upon the macroscopic state of the robot swarm represented by a Gaussian Mixture Model (GMM) and integrates conditional value-at-risk (CVaR) to ensure collision avoidance. The key component of ROVER is imposing a CVaR constraint on the distribution of the Signed Distance Function between the swarm GMM and obstacles in the FTMPC to enforce collision avoidance. Utilizing the analytical expression of CVaR of a GMM derived in this work, we develop a computationally efficient solution to solve the non-linear constrained FTMPC through sequential linear programming. Simulations and comparisons with representative benchmark approaches demonstrate the effectiveness of ROVER in flexibility, scalability, and safety guarantee.

IROS Conference 2024 Conference Paper

SwarmPRM: Probabilistic Roadmap Motion Planning for Large-Scale Swarm Robotic Systems

  • Yunze Hu
  • Xuru Yang
  • Kangjie Zhou
  • Qinghang Liu
  • Kang Ding
  • Han Gao
  • Pingping Zhu
  • Chang Liu

Large-scale swarm robotic systems consisting of numerous cooperative agents show considerable promise for performing autonomous tasks across various sectors. Nonetheless, traditional motion planning approaches often face a trade-off between scalability and solution quality due to the exponential growth of the joint state space of robots. In response, this work proposes SwarmPRM, a hierarchical, scalable, computationally efficient, and risk-aware sampling-based motion planning approach for large-scale swarm robots. SwarmPRM utilizes a Gaussian Mixture Model (GMM) to represent the swarm’s macroscopic state and constructs a Probabilistic Roadmap in Gaussian space, referred to as the Gaussian roadmap, to generate a transport trajectory of GMM. This trajectory is then followed by each robot at the microscopic stage. To enhance trajectory safety, SwarmPRM incorporates the conditional value-at-risk (CVaR) in the collision checking process to impart the property of risk awareness to the constructed Gaussian roadmap. SwarmPRM then crafts a linear programming formulation to compute the optimal GMM transport trajectory within this roadmap. Extensive simulations demonstrate that SwarmPRM outperforms state-of-the-art methods in computational efficiency, scalability, and trajectory quality while offering the capability to adjust the risk tolerance of generated trajectories.

ICRA Conference 2024 Conference Paper

VIDAR: Data Quality Improvement for Monocular 3D Reconstruction through In-situ Visual Interaction

  • Han Gao
  • Yating Liu
  • Fang Cao
  • Hao Wu 0067
  • Fengyuan Xu
  • Sheng Zhong 0002

3D reconstruction based on monocular videos has attracted wide attention, and existing reconstruction methods usually work in a reconstruction-after-scanning manner. However, these methods suffer from insufficient data collection problems due to the lack of effective guidance for users during the scanning process, which affects reconstruction quality. We propose VIDAR, which visually guides users with the streaming incremental reconstructed mesh in data collection for monocular 3D reconstruction. We propose an incremental mesh extraction algorithm to achieve lossless fusion of streaming incremental mesh data via slice-style management for guidance quality. We also design an incremental mesh rendering algorithm to achieve precise memory reallocation by updating the buffer in a fill-in-the-blank pattern for guidance efficiency. Besides, we introduce several optimizations on data transmission and human-computer interaction to improve the overall system performance. The experiment results on real-world scenes show that VIDAR efficiently delivers high-quality visual guidance and outperforms the non-interactive data collection methods for scene reconstruction.

NeurIPS Conference 2023 Conference Paper

Unifying Predictions of Deterministic and Stochastic Physics in Mesh-reduced Space with Sequential Flow Generative Model

  • Luning Sun
  • Xu Han
  • Han Gao
  • Jian-Xun Wang
  • Liping Liu

Accurate prediction of dynamical systems in unstructured meshes has recently shown successes in scientific simulations. Many dynamical systems have a nonnegligible level of stochasticity introduced by various factors (e. g. chaoticity), so there is a need for a unified framework that captures both deterministic and stochastic components in the rollouts of these systems. Inspired by regeneration learning, we propose a new model that combines generative and sequential networks to model dynamical systems. Specifically, we use an autoencoder to learn compact representations of full-space physical variables in a low-dimensional space. We then integrate a transformer with a conditional normalizing flow model to model the temporal sequence of latent representations. We evaluate the new model in both deterministic and stochastic systems. The model outperforms several competitive baseline models and makes more accurate predictions of deterministic systems. Its own prediction error is also reflected in its uncertainty estimations. When predicting stochastic systems, the proposed model generates high-quality rollout samples. The mean and variance of these samples well match the statistics of samples computed from expensive numerical simulations.

NeurIPS Conference 2022 Conference Paper

Multi-Sample Training for Neural Image Compression

  • Tongda Xu
  • Yan Wang
  • Dailan He
  • Chenjian Gao
  • Han Gao
  • Kunzan Liu
  • Hongwei Qin

This paper considers the problem of lossy neural image compression (NIC). Current state-of-the-art (SOTA) methods adopt uniform posterior to approximate quantization noise, and single-sample pathwise estimator to approximate the gradient of evidence lower bound (ELBO). In this paper, we propose to train NIC with multiple-sample importance weighted autoencoder (IWAE) target, which is tighter than ELBO and converges to log likelihood as sample size increases. First, we identify that the uniform posterior of NIC has special properties, which affect the variance and bias of pathwise and score function estimators of the IWAE target. Moreover, we provide insights on a commonly adopted trick in NIC from gradient variance perspective. Based on those analysis, we further propose multiple-sample NIC (MS-NIC), an enhanced IWAE target for NIC. Experimental results demonstrate that it improves SOTA NIC methods. Our MS-NIC is plug-and-play, and can be easily extended to neural video compression.

AAAI Conference 2018 Conference Paper

Ranking Users in Social Networks With Higher-Order Structures

  • Huan Zhao
  • Xiaogang Xu
  • Yangqiu Song
  • Dik Lun Lee
  • Zhao Chen
  • Han Gao

PageRank has been widely used to measure the authority or the influence of a user in social networks. However, conventional PageRank only makes use of edge-based relations, ignoring higher-order structures captured by motifs, subgraphs consisting of a small number of nodes in complex networks. In this paper, we propose a novel framework, motif-based PageRank (MPR), to incorporate higher-order structures into conventional PageRank computation. We conduct extensive experiments in three real-world networks, i. e. , DBLP, Epinions, and Ciao, to show that MPR can significantly improve the effectiveness of PageRank for ranking users in social networks. In addition to numerical results, we also provide detailed analysis for MPR to show how and why incorporating higher-order information works better than PageRank in ranking users in social networks. 1

v2026.09.13