Arrow Research search

Author name cluster

Hao Sun

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

72 papers
2 author rows

Possible papers

72

EAAI Journal 2026 Journal Article

A zero-shot tree-structured multi-objective evolutionary Neural Architecture Search

  • Yan Dai
  • Lixin Wei
  • Ziyu Hu
  • Hao Sun
  • Qianao Xu
  • Kexin Zhang
  • Boya Zhao

Neural Architecture Search (NAS) enables the automated design of high-performance neural networks; however, its practical application is often constrained by substantial computational costs, the limited reliability of single-objective proxy metrics, and insufficient modeling of architectural information flow. To address these limitations, we propose Tree-structured Evolutionary Neural Architecture Search (TreeNAS). The method integrates three components: (1) a tree-structured encoding with refinement to preserve backbone information paths under mutation; (2) a zero-cost multi-objective evaluation that jointly assesses trainability, generalization, and complexity, thereby mitigating instability from single-objective proxy; and (3) a Pareto-dominance-guided evolutionary search to encourage diverse, balanced architectures across objectives. On the standard Neural Architecture Search Benchmark 101 (NAS-Bench-101) and Neural Architecture Search Benchmark 201 (NAS-Bench-201) datasets, TreeNAS achieves state-of-the-art accuracy with a 40 × reduction in search cost. On the ImageNet dataset under strict floating-point operation (FLOPs) budgets, TreeNAS achieves accuracy comparable to training-based NAS methods while keeping the search cost to 0. 45 Graphics Processing Unit (GPU) days. Additionally, TreeNAS generalizes across modalities, from two-dimensional images to medical signals and volumetric imaging, demonstrating its potential in practical medical imaging applications.

AAAI Conference 2026 Conference Paper

Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment Retrieval

  • Haojian Huang
  • Kaijing Ma
  • Jin Chen
  • Haodong Chen
  • Zhou Wu
  • Xianghao Zang
  • Han Fang
  • Chao Ban

In the domain of moment retrieval, accurately identifying temporal segments within videos based on natural language queries remains challenging. Traditional methods often employ pre-trained models that struggle with fine-grained information and deterministic reasoning, leading to difficulties in aligning with complex or ambiguous moments. To overcome these limitations, we explore Deep Evidential Regression (DER) to construct a vanilla Evidential baseline. However, this approach encounters two major issues: the inability to effectively handle modality imbalance and the structural differences in DER's heuristic uncertainty regularizer, which adversely affect uncertainty estimation. This misalignment results in high uncertainty being incorrectly associated with accurate samples rather than challenging ones. Our observations indicate that existing methods lack the adaptability required for complex video scenarios. In response, we propose Debiased Evidential Learning for Moment Retrieval (DEMR), a novel framework that incorporates a Reflective Flipped Fusion (RFF) block for cross-modal alignment and a query reconstruction task to enhance text sensitivity, thereby reducing bias in uncertainty estimation. Additionally, we introduce a Geom-regularizer to refine uncertainty predictions, enabling adaptive alignment with difficult moments and improving retrieval accuracy. Extensive testing on standard datasets and debiased datasets ActivityNet-CD and Charades-CD demonstrates significant enhancements in effectiveness, robustness, and interpretability, positioning our approach as a promising solution for temporal-semantic robustness in moment retrieval.

AAAI Conference 2026 Conference Paper

Differentiable Sparse Identification of Lagrangian Dynamics

  • Zitong Zhang
  • Hao Sun

Data-driven discovery of governing equations from data remains a fundamental challenge in nonlinear dynamics. Although sparse regression techniques have advanced system identification, they struggle with rational functions and noise sensitivity in complex mechanical systems. The Lagrangian formalism offers a promising alternative, as it typically avoids rational expressions and provides a more concise representation of system dynamics. However, existing Lagrangian identification methods are significantly affected by measurement noise and limited data availability. This paper presents a novel differentiable sparse identification framework that addresses these limitations through three key contributions: (1) the first integration of cubic B-Spline approximation into Lagrangian system identification, enabling accurate representation of complex nonlinearities, (2) a robust equation discovery mechanism that effectively utilizes measurements while incorporating known physical constraints, (3) a recursive derivative computation scheme based on B-spline basis functions, effectively constraining higher-order derivatives and reducing noise sensitivity on second-order dynamical systems. The proposed method demonstrates superior performance and enables more accurate and reliable extraction of physical laws from noisy data, particularly in complex mechanical systems compared to baseline methods.

AAAI Conference 2026 Conference Paper

Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing

  • Bingyu Li
  • Haocheng Dong
  • Da Zhang
  • Zhiyuan Zhao
  • Hao Sun
  • Junyu Gao

Open-Vocabulary Remote Sensing Image Segmentation (OVRSIS), an emerging task that adapts Open-Vocabulary Segmentation (OVS) to the remote sensing (RS) domain, remains underexplored due to the absence of a unified evaluation benchmark and the domain gap between natural and RS images. To bridge these gaps, we first establish a standardized OVRSIS benchmark (OVRSISBench) based on widely-used RS segmentation datasets, enabling consistent evaluation across methods. Using this benchmark, we comprehensively evaluate several representative OVS/OVRSIS models and reveal their limitations when directly applied to remote sensing scenarios. Building on these insights, we propose RSKT-Seg, a novel open-vocabulary segmentation framework tailored for remote sensing. RSKT-Seg integrates three key components: (1) a Multi-Directional Cost Map Aggregation (RS-CMA) module that captures rotation-invariant visual cues by computing vision-language cosine similarities across multiple directions; (2) an Efficient Cost Map Fusion (RS-Fusion) transformer, which jointly models spatial and semantic dependencies with a lightweight dimensionality reduction strategy; and (3) a Remote Sensing Knowledge Transfer (RS-Transfer) module that injects pre-trained knowledge and facilitates domain adaptation via enhanced upsampling. Extensive experiments on the benchmark show that RSKT-Seg consistently outperforms strong OVS baselines by +3.8 mIoU and +5.9 mACC, while achieving 2× faster inference through efficient aggregation.

EAAI Journal 2026 Journal Article

Few-shot transfer learning for laser welding prediction

  • Luchen Wu
  • Shijie Wu
  • Hongxin Hu
  • Hao Sun
  • Shuang Ma
  • Zhenya Wang

Laser wire filling welding is a key joining technique in the manufacturing of aluminum battery packs for new energy vehicles, yet predictive modeling of melt pool geometry remains limited by scarce experimental data. A two-stage transfer learning framework is developed by integrating multiphysics numerical simulation, data augmentation, and Bayesian neural network (BNN). High-fidelity multiphysics simulations within the experimental process window are generated to expand parameter space coverage and to provide physics-informed data for model pretraining. Limited experimental samples are augmented using a Wasserstein Generative Adversarial Network with gradient penalty applied to laser power and wire feed speed. Gaussian perturbations on travel speed are introduced to represent measurement uncertainties. A shallow BNN is pretrained on simulated samples and fine-tuned on the augmented experimental dataset using physics-consistent regularization and partial layer-freezing strategies. The augmentation strategy is evaluated through leave-one-out cross-validation on eight experimental samples, and generalization is examined using a separate test under previously unobserved travel-speed conditions. After inverse normalization, the framework achieves root mean square errors of 0. 027 mm for melt pool depth and 0. 025 mm for width, with coefficients of determination of 0. 788 and 0. 741, respectively. Uncertainty-aware quantitative analysis based on Sobol sensitivity indices and reliability assessment is conducted after model validation to characterize dominant parameter influences and to identify high-confidence process windows under limited data conditions. The proposed framework provides a general simulation-informed and uncertainty-aware learning strategy for manufacturing processes with severely limited experimental data.

AAAI Conference 2026 Conference Paper

L2V-CoT: Cross-Modal Transfer of Chain-of-Thought Reasoning via Latent Intervention

  • Yu-Liang Zhan
  • Xinyu Tang
  • Han Wan
  • Jian Li
  • Jirong Wen
  • Hao Sun

Recently, Chain-of-Thought (CoT) reasoning has significantly enhanced the capabilities of large language models (LLMs), but Vision–Language Models (VLMs) still struggle with multi-step reasoning tasks due to limited multimodal reasoning data. To bridge this gap, researchers have explored methods to transfer CoT reasoning from LLMs to VLMs. However, existing approaches either need high training costs or require architectural alignment. In this paper, we use Linear Artificial Tomography (LAT) to empirically show that LLMs and VLMs share similar low-frequency latent representations of CoT reasoning despite architectural differences. Based on this insight, we propose L2V-CoT, a novel training-free latent intervention approach that transfers CoT reasoning from LLMs to VLMs. L2V-CoT extracts and resamples low-frequency CoT representations from LLMs in the frequency domain, enabling dimension matching and latent injection into VLMs during inference to enhance reasoning capabilities. Extensive experiments demonstrate that our approach consistently outperforms training-free baselines and even surpasses supervised methods.

AAAI Conference 2026 Conference Paper

Mitigating Error Accumulation in Knowledge Editing for Multi-Hop Question Answering

  • Jiaxin Guo
  • Hao Sun
  • Wenhao Zhang
  • Xuanbo Fan
  • Yan Zhang

Knowledge editing (KE) has emerged as an effective approach for updating factual information in large language models (LLMs) without the need for full retraining. Most of the existing methods for addressing the "ripple effect" in KE adopt a chain-structured reasoning process, making them vulnerable to error accumulation from early incorrect steps. Moreover, their conflict detection mechanisms are often susceptible to the LLM's inherent confirmation bias, further undermining the reliability of the editing process. To overcome these challenges, we propose Tree of Editing (ToE), a tree-structured, retrieval-enhanced knowledge editing framework designed to support robust reasoning under factual updates. ToE expands reasoning paths using a breadth-first strategy combined with score-guided beam search, enabling diverse and error-tolerant inference. Besides, we introduce an observer to objectively update knowledge, avoiding the bias caused by LLMs' over-confidence. Experimental results on two benchmarks, namely MQuAKE-CF (targeting ripple-aware editing) and DUNE (free-form editing), demonstrate that ToE framework significantly outperforms existing methods.

AAAI Conference 2026 Conference Paper

PIMRL: Physics-Informed Multi-Scale Recurrent Learning for Burst-Sampled Spatiotemporal Dynamics

  • Han Wan
  • Qi Wang
  • Yuan Mi
  • Rui Zhang
  • Hao Sun

Deep learning has shown strong potential in modeling complex spatiotemporal dynamics. However, most existing methods depend on densely and uniformly sampled data, which is often unavailable in practice due to sensor and cost limitations. In many real-world settings, such as mobile sensing and physical experiments, data are burst-sampled with short high-frequency segments followed by long gaps, making it difficult to learn accurate dynamics from sparse observations. To address this issue, we propose Physics-Informed Multi-Scale Recurrent Learning (PIMRL), a novel framework specifically designed for burst-sampled spatiotemporal data. PIMRL combines macro-scale latent dynamics inference with micro-scale adaptive refinement guided by incomplete prior information from partial differential equations (PDEs). It further introduces a temporal message-passing mechanism to effectively propagate information across burst intervals. This multi-scale architecture enables PIMRL to model complex systems accurately even under severe data scarcity. We evaluate our approach on five benchmark datasets involving 1D to 3D multi-scale PDEs. The results show that PIMRL consistently outperforms state-of-the-art baselines, achieving substantial improvements and reducing errors by up to 80\% in the most challenging settings, which demonstrates the clear advantage of our model. Our work demonstrates the effectiveness of physics-informed recurrent learning for accurate and efficient modeling of sparse spatiotemporal systems.

AAAI Conference 2026 Conference Paper

TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding

  • Canhui Tang
  • Zifan Han
  • Hongbo Sun
  • Sanping Zhou
  • Xuchong Zhang
  • Xin Wei
  • Ye Yuan
  • Huayu Zhang

Multimodal Large Language Models (MLLMs) have demonstrated significant progress in vision-language tasks, yet they still face challenges when processing long-duration video inputs. The limitation arises from MLLMs' context limit and training costs, necessitating sparse frame sampling before feeding videos into MLLMs. However, building a trainable sampling method remains challenging due to the unsupervised and non-differentiable nature of sparse frame sampling in Video-MLLMs. To address these problems, we propose Temporal Sampling Policy Optimization (**TSPO**), advancing MLLMs' long-form video-language understanding via reinforcement learning. Specifically, we first propose a trainable event-aware temporal agent, which captures event-query correlation for performing probabilistic keyframe selection. Then, we propose the TSPO reinforcement learning paradigm, which models keyframe selection and language generation as a joint decision-making process, enabling end-to-end group relative optimization for the temporal sampling policy. Furthermore, we propose a dual-style long video training data construction pipeline, balancing comprehensive temporal understanding and key segment localization. Finally, we incorporate rule-based answering accuracy and temporal locating reward mechanisms to optimize the temporal sampling policy. Comprehensive experiments show that our TSPO achieves state-of-the-art performance across multiple long video understanding benchmarks, and shows transferable ability across different cutting-edge Video-MLLMs.

EAAI Journal 2025 Journal Article

A knowledge-guided reinforcement learning method for lateral path tracking

  • Bo Hu
  • Sunan Zhang
  • Yuxiang Feng
  • Bingbing Li
  • Hao Sun
  • Mingyang Chen
  • Weichao Zhuang
  • Yi Zhang

Lateral Control algorithms in autonomous vehicles often necessitates an online fine-tuning procedure in the real world. While reinforcement learning (RL) enables vehicles to learn and improve the lateral control performance through repeated trial and error interactions with a dynamic environment, applying RL directly to safety-critical applications in real physical world is challenging because ensuring safety during the learning process remains difficult. To enable safe learning, a promising direction is to make use of previously gathered offline data, which is frequently accessible in engineering applications. In this context, this paper presents a set of knowledge-guided RL algorithms that can not only fully leverage the prior collected offline data without the need of a physics-based simulator, but also allow further online policy improvement in a smooth, safe and efficient manner. To evaluate the effectiveness of the proposed algorithms on a real controller, a hardware-in-the-loop and a miniature vehicle platform are built. Compared with the vanilla RL, behavior cloning and the existing controller, the proposed algorithms realize a closed-loop solution for lateral control problems from offline training to online fine-tuning, making it attractive for future similar RL-based controller to build upon.

EAAI Journal 2025 Journal Article

A rescheduling strategy based on deep reinforcement learning for new order insertion

  • Hao Sun
  • Zhiwei Hao
  • Ziyu Hu
  • Jin Lv
  • Cong Wang

Manufacturing shops have common dynamic factors that often disrupt pre-established schedules. In production, a buffer time is typically set to account for such dynamic disturbances. However, which increases scheduling complexity and expands the Flexible Job Scheduling Problem (FJSP) into the Flexible Job Shop Rescheduling Problem (FJSRP). Consequently, an efficient method is urgently needed. In this context, the insertion of a new order alters the scheduling state. This study aims to address the objectives of minimizing the earliest completion time of orders and maximizing the utilization of machine resources. For the FJSRP, it is essential to establish both the independent insertion time and pre-processing buffer time while also devising a rescheduling method that balances the number of jobs across orders. To tackle these challenges, the Proximal Policy Optimization for Hybrid Feature Extraction Network (PPO-HFEN) strategy was developed. Innovations in this strategy include an attention mechanism that integrates a Multi-layer Hybrid Feature Extraction Network (HFEN) for feature extraction and job prioritization. Furthermore, a Mixed Reward Function (MRF) is introduced to resolve convergence issues during the training process. Experimental evaluations demonstrated that PPO-HFEN outperformed four state-of-the-art algorithms, yielding superior solutions in 21 out of 40 instances, thereby underscoring the effectiveness of the HFEN and MRF components. Moreover, the validity of the proposed rescheduling strategy is further validated on a real production foundry instance. The primary objective is to analyze the impact of varying order sizes, insertion times, and pre-processing times on the test results of diverse sampling methods.

NeurIPS Conference 2025 Conference Paper

A Set of Generalized Components to Achieve Effective Poison-only Clean-label Backdoor Attacks with Collaborative Sample Selection and Triggers

  • Zhixiao Wu
  • Yao Lu
  • Jie Wen
  • Hao Sun
  • Qi Zhou
  • Guangming Lu

Poison-only Clean-label Backdoor Attacks (PCBAs) aim to covertly inject attacker-desired behavior into DNNs by merely poisoning the dataset without changing the labels. To effectively implant a backdoor, multiple triggers are proposed for various attack requirements of Attack Success Rate (ASR) and stealthiness. Additionally, sample selection enhances clean-label backdoor attacks' ASR by meticulously selecting "hard'' samples instead of random samples to poison. Current methods, however, 1) usually handle the sample selection and triggers in isolation, leading to severely limited improvements on both ASR and stealthiness. Consequently, attacks exhibit unsatisfactory performance on evaluation metrics when converted to PCBAs via a mere stacking of methods. Therefore, we seek to explore the bi-directional collaborative relations between the sample selection and triggers to address the above dilemma. 2) Since the strong specificity within triggers, the simple combination of sample selection and triggers fails to substantially enhance both evaluation metrics, with generalization preserved among various attacks. Therefore, we seek to propose a set of components to significantly improve both stealthiness and ASR based on the commonalities of attacks. Specifically, Component A ascertains two critical selection factors, and then makes them an appropriate combination based on the trigger scale to select more reasonable "hard'' samples for improving ASR. Component B is proposed to select samples with similarities to relevant trigger implanted samples to promote stealthiness. Component C reassigns trigger poisoning intensity on RGB colors through distinct sensitivity of the human visual system to RGB for higher ASR, with stealthiness ensured by sample selection including Component B. Furthermore, all components can be strategically integrated into diverse PCBAs, enabling tailored solutions that balance ASR and stealthiness enhancement for specific attack requirements. Extensive experiments demonstrate the superiority of our components in stealthiness, ASR, and generalization. Our code will be released as soon as possible.

NeurIPS Conference 2025 Conference Paper

Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models

  • Xiaoyu Zhan
  • Wenxuan Huang
  • Hao Sun
  • Xinyu Fu
  • Changfeng Ma
  • Shaosheng Cao
  • Bohan Jia
  • Shaohui Lin

Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved 2D visual understanding, prompting interest in their application to complex 3D reasoning tasks. However, it remains unclear whether these models can effectively capture the detailed spatial information required for robust real-world performance, especially cross-view consistency, a key requirement for accurate 3D reasoning. Considering this issue, we introduce Viewpoint Learning, a task designed to evaluate and improve the spatial reasoning capabilities of MLLMs. We present the Viewpoint-100K dataset, consisting of 100K object-centric image pairs with diverse viewpoints and corresponding question-answer pairs. Our approach employs a two-stage fine-tuning strategy: first, foundational knowledge is injected to the baseline MLLM via Supervised Fine-Tuning (SFT) on Viewpoint-100K, resulting in significant improvements across multiple tasks; second, generalization is enhanced through Reinforcement Learning using the Group Relative Policy Optimization (GRPO) algorithm on a broader set of questions. Additionally, we introduce a hybrid cold-start initialization method designed to simultaneously learn viewpoint representations and maintain coherent reasoning thinking. Experimental results show that our approach significantly activates the spatial reasoning ability of MLLM, improving performance on both in-domain and out-of-domain reasoning tasks. Our findings highlight the value of developing foundational spatial skills in MLLMs, supporting future progress in robotics, autonomous systems, and 3D scene understanding.

YNIMG Journal 2025 Journal Article

Association of spatiotemporal interaction of gamma oscillations with heart rate variability during response inhibition processing in patients with major depressive disorder: An MEG study

  • Junling Sheng
  • Yi Xia
  • Lingling Hua
  • Hongliang Zhou
  • Qian Liao
  • Shui Tian
  • Yishan Du
  • Xiaoqin Wang

BACKGROUND: Impairment in response inhibition function is highly prevalent in patients with major depressive disorder (MDD), yet the spatiotemporal neural activity underlying response inhibition and its relationship with the autonomic nervous system (ANS) remains unclear. METHODS: 35 MDD participants and 35 healthy controls (HC) were included with magnetoencephalography (MEG) and electrocardiogram (ECG) data collecting during a go/no-go task. Heart rate variability (HRV) indices were calculated from the ECG data. Differences in functional connectivity (FC) of gamma oscillations (60-90 Hz) between 0-200 ms, 200-400 ms, and 400-600 ms in the two groups after no-go stimuli were analyzed, and the correlation between FC and HRV indices was examined. RESULTS: The MDD group exhibited poorer task performance and lower HRV indices than the HC group. During the 200-400 ms period, compared to the HC group, the MDD group exhibited decreased FC between the left inferior frontal gyrus (opercular part) and right temporal pole (middle temporal gyrus) (t = 3.62, p < 0.05), and increased FC between the right superior frontal gyrus (orbital part) and right superior occipital gyrus (t = 3.68, p < 0.05). Additionally, a significant positive correlation was found between FC of the left inferior frontal gyrus (opercular part) and right middle temporal gyrus (temporal pole) and the HRV index RMSSD in the MDD group (r = 0.491, p < 0.05). CONCLUSION: Abnormal spatiotemporal interactions in gamma oscillations related to response inhibition are observed in MDD patients and abnormal gamma oscillations showed task-dependent covariation with ANS indices, suggesting their potential interplay in MDD pathophysiology.

ICLR Conference 2025 Conference Paper

CPSample: Classifier Protected Sampling for Guarding Training Data During Diffusion

  • Joshua Kazdan
  • Hao Sun
  • Jiaqi Han
  • Felix Petersen
  • Frederick Vu
  • Stefano Ermon

Diffusion models have a tendency to exactly replicate their training data, especially when trained on small datasets. Most prior work has sought to mitigate this problem by imposing differential privacy constraints or masking parts of the training data, resulting in a notable substantial decrease in image quality. We present CPSample, a method that modifies the sampling process to prevent training data replication while preserving image quality. CPSample utilizes a classifier that is trained to overfit on random binary labels attached to the training data. CPSample then uses classifier guidance to steer the generation process away from the set of points that can be classified with high certainty, a set that includes the training data. CPSample achieves FID scores of 4.97 and 2.97 on CIFAR-10 and CelebA-64, respectively, without producing exact replicates of the training data. Unlike prior methods intended to guard the training images, CPSample only requires training a classifier rather than retraining a diffusion model, which is computationally cheaper. Moreover, our technique provides diffusion models with greater robustness against membership inference attacks, wherein an adversary attempts to discern which images were in the model's training dataset. We show that CPSample behaves like a built-in rejection sampler, and we demonstrate its capabilities to prevent mode collapse in Stable Diffusion.

EAAI Journal 2025 Journal Article

CTPR: Contrastive transition predictive representation for reinforcement learning

  • Hao Sun
  • Changpeng Wang

Learning policies from high-dimensional observations is a challenging problem for pixel-based reinforcement learning. Most existing pixel-based reinforcement learning methods struggle with the inefficiency of extracting meaningful state representations from raw pixel data, lacking temporal correlation and resulting in suboptimal performance. To this end, we propose a innovative method named contrastive transition predictive representation for reinforcement learning (CTPR), which utilizes contrastive learning and a transition model to efficiently extract high-level state representations from raw pixels for sample-efficient reinforcement learning. In the reinforcement learning component, we perform policy control based on the learned contrastive representations. We have evaluated the effectiveness of the proposed method by conducting numerous experiments on DeepMind Control, and the results show that our method has achieve significant improvements over the state-of-the-art methods.

JBHI Journal 2025 Journal Article

EEG-Deformer: A Dense Convolutional Transformer for Brain-Computer Interfaces

  • Yi Ding
  • Yong Li
  • Hao Sun
  • Rui Liu
  • Chengxuan Tong
  • Chenyu Liu
  • Xinliang Zhou
  • Cuntai Guan

Effectively learning the temporal dynamics in electroencephalogram (EEG) signals is challenging yet essential for decoding brain activities using brain-computer interfaces (BCIs). Although Transformers are popular for their long-term sequential learning ability in the BCI field, most methods combining Transformers with convolutional neural networks (CNNs) fail to capture the coarse-to-fine temporal dynamics of EEG signals. To overcome this limitation, we introduce EEG-Deformer, which incorporates two main novel components into a CNN-Transformer: (1) a Hierarchical Coarse-to-Fine Transformer (HCT) block that integrates a Fine-grained Temporal Learning (FTL) branch into Transformers, effectively discerning coarse-to-fine temporal patterns; and (2) a Dense Information Purification (DIP) module, which utilizes multi-level, purified temporal information to enhance decoding accuracy. Comprehensive experiments on three representative cognitive tasksâcognitive attention, driving fatigue, and mental workload detectionâconsistently confirm the generalizability of our proposed EEG-Deformer, demonstrating that it either outperforms or performs comparably to existing state-of-the-art methods. Visualization results show that EEG-Deformer learns from neurophysiologically meaningful brain regions for the corresponding cognitive tasks.

AAAI Conference 2025 Conference Paper

EyEar: Learning Audio Synchronized Human Gaze Trajectory Based on Physics-Informed Dynamics

  • Xiaochuan Liu
  • Xin Cheng
  • Yuchong Sun
  • Xiaoxue Wu
  • Ruihua Song
  • Hao Sun
  • Denghao Zhang

Imitating how humans move their gaze in a visual scene is a vital research problem for both visual understanding and psychology, kindling crucial applications such as building alive virtual characters. Previous studies aim to predict gaze trajectories when humans are free-viewing an image, searching for required targets, or looking for clues to answer questions in an image. While these tasks focus on visual-centric scenarios, humans move their gaze also along with audio signal inputs in more common scenarios. To fill this gap, we introduce a new task that predicts human gaze trajectories in a visual scene with synchronized audio inputs and provide a new dataset containing 20k gaze points from 8 subjects. To effectively integrate audio information and simulate the dynamic process of human gaze motion, we propose a novel learning framework called EyEar (Eye moving while Ear listening) based on physics-informed dynamics, which considers three key factors to predict gazes: eye inherent motion tendency, vision salient attraction, and audio semantic attraction. We also propose a probability density score to overcome the high individual variability of gaze trajectories, thereby improving the stabilization of optimization and the reliability of the evaluation. Experimental results show that EyEar outperforms all the baselines in the context of all evaluation metrics, thanks to the proposed components in the learning model.

NeurIPS Conference 2025 Conference Paper

FlexWorld: Progressively Expanding 3D Scenes for Flexible-View Exploration

  • Luxi Chen
  • Zihan Zhou
  • Min Zhao
  • Yikai Wang
  • Ge Zhang
  • Wenhao Huang
  • Hao Sun
  • Ji-Rong Wen

Generating flexible-view 3D scenes, including 360° rotation and zooming, from single images is challenging due to a lack of 3D data. To this end, we introduce FlexWorld, a novel framework that progressively constructs a persistent 3D Gaussian splatting representation by synthesizing and integrating new 3D content. To handle novel view synthesis under large camera variations, we leverage an advanced pre-trained video model fine-tuned on accurate depth-estimated training pairs. By combining geometry-aware scene integration and optimization, FlexWorld refines the scene representation, producing visually consistent 3D scenes with flexible viewpoints. Extensive experiments demonstrate the effectiveness of FlexWorld in generating high-quality novel view videos and flexible-view 3D scenes from single images, achieving superior visual quality under multiple popular metrics and datasets compared to existing state-of-the-art methods. Additionally, FlexWorld supports extrapolating from existing 3D scenes, further extending its applicability. Qualitatively, we highlight that FlexWorld can generate high-fidelity scenes that enable 360° rotations and zooming exploration. Our code is available at https: //github. com/ML-GSAI/FlexWorld.

NeurIPS Conference 2025 Conference Paper

From Pretraining to Pathology: How Noise Leads to Catastrophic Inheritance in Medical Models

  • Hao Sun
  • Zhongyi Han
  • Hao Chen
  • Jindong Wang
  • Xin Gao
  • Yilong Yin

Foundation models pretrained on web-scale data drive contemporary transfer learning in vision, language, and multimodal tasks. Recent work shows that mild label noise in these corpora may lift in-distribution accuracy yet sharply reduce out-of-distribution generalization, an effect known as catastrophic inheritance. Medical data is especially sensitive because annotations are scarce, domain shifts are large, and pretraining sources are noisy. We present the first systematic analysis of catastrophic inheritance in medical models. Controlled label-corruption experiments expose a clear structural collapse: as noise rises, the skewness and kurtosis of feature and logit distributions decline, signaling a flattened representation space and diminished discriminative detail. These higher-order statistics form a compact, interpretable marker of degradation in fine-grained tasks such as histopathology. Guided by this finding, we introduce a fine-tuning objective that restores skewness and kurtosis through two scalar regularizers added to the task loss. The method leaves the backbone unchanged and incurs negligible overhead. Tests on PLIP models trained with Twitter pathology images, as well as other large-scale vision and language backbones, show consistent gains in robustness and cross-domain accuracy under varied noise levels.

AAAI Conference 2025 Conference Paper

HHAN: Comprehensive Infectious Disease Source Tracing via Heterogeneous Hypergraph Neural Network

  • Qiang He
  • Yunting Bao
  • Hui Fang
  • Yuting Lin
  • Hao Sun

Infectious diseases have historically had profound effects on global health, economies, and social structures. Effective tracing of infectious diseases is essential not only for immediate public health responses but also for shaping future prevention strategies. Traditional tracing methods often emphasize homogeneous networks, overlooking the diverse transmission characteristics of heterogeneous populations. This research addresses two critical challenges: the heterogeneity of transmission across various media and modes, and the significant yet underexplored influence of community structures on epidemic spread and tracing.We propose a Heterogeneous Hypergraph Attention Network (HHAN) modelthat accounts for multiple transmission pathways and patterns within heterogeneous networks. HHAN integrates a heterogeneous graph neural network module to handle the complexity of communication among different populations, and an Agent-Based Modeling Module that combines agent-based ideas to model individual behaviors. This approach effectively captures complex interactions within community structures and addresses individual variability. Experimental results on three real-world datasets demonstrate that the HHAN model significantly outperforms other state-of-the-art methods in tackling the complex challenge of tracing infectious diseases in heterogeneous populations.

NeurIPS Conference 2025 Conference Paper

Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models

  • Haohan Chi
  • Huan-ang Gao
  • Ziming Liu
  • Jianing Liu
  • Chenyu Liu
  • Jinwei Li
  • Kaisen Yang
  • Yangcheng Yu

Vision-Language-Action (VLA) models for autonomous driving show promise but falter in unstructured corner case scenarios, largely due to a scarcity of targeted benchmarks. To address this, we introduce Impromptu VLA. Our core contribution is the Impromptu VLA Dataset: over 80, 000 meticulously curated video clips, distilled from over 2M source clips sourced from 8 open-source large-scale datasets. This dataset is built upon our novel taxonomy of four challenging unstructured categories and features rich, planning-oriented question-answering annotations and action trajectories. Crucially, experiments demonstrate that VLAs trained with our dataset achieve substantial performance gains on established benchmarks—improving closed-loop NeuroNCAP scores and collision rates, and reaching near state-of-the-art L2 accuracy in open-loop nuScenes trajectory prediction. Furthermore, our Q&A suite serves as an effective diagnostic, revealing clear VLM improvements in perception, prediction, and planning. Our code, data and models are available at https: //github. com/ahydchh/Impromptu-VLA

NeurIPS Conference 2025 Conference Paper

LayerNavigator: Finding Promising Intervention Layers for Efficient Activation Steering in Large Language Models

  • Hao Sun
  • Huailiang Peng
  • Qiong Dai
  • Xu Bai
  • Yanan Cao

Activation steering is an efficient technique for aligning the behavior of large language models (LLMs) by injecting steering vectors directly into a model’s residual stream during inference. A pivotal challenge in this approach lies in choosing the right layers to intervene, as inappropriate selection can undermine behavioral alignment and even impair the model’s language fluency and other core capabilities. While single-layer steering allows straightforward evaluation on held-out data to identify the "best" layer, it offers only limited alignment improvements. Multi-layer steering promises stronger control but faces a combinatorial explosion of possible layer subsets, making exhaustive search impractical. To address these challenges, we propose LayerNavigator, which provides a principled and promising layer selection strategy. The core innovation of LayerNavigator lies in its novel, quantifiable criterion that evaluates each layer's steerability by jointly considering two key aspects: discriminability and consistency. By reusing the activations computed during steering vector generation, LayerNavigator requires no extra data and adds negligible overhead. Comprehensive experiments show that LayerNavigator achieves not only superior alignment but also greater scalability and interpretability compared to existing strategies. Our code is available at https: //github. com/Bryson-Arrot/LayerNavigator

AAAI Conference 2025 Conference Paper

MTL-LoRA: Low-Rank Adaptation for Multi-Task Learning

  • Yaming Yang
  • Dilxat Muhtar
  • Yelong Shen
  • Yuefeng Zhan
  • Jianfeng Liu
  • Yujing Wang
  • Hao Sun
  • Weiwei Deng

Parameter-efficient fine-tuning (PEFT) has been widely employed for domain adaptation, with LoRA being one of the most prominent methods due to its simplicity and effectiveness. However, in multi-task learning (MTL) scenarios, LoRA tends to obscure the distinction between tasks by projecting sparse high-dimensional features from different tasks into the same dense low-dimensional intrinsic space. This leads to task interference and suboptimal performance for LoRA and its variants. To tackle this challenge, we propose MTL-LoRA, which retains the advantages of low-rank adaptation while significantly enhancing MTL capabilities. MTL-LoRA augments LoRA by incorporating additional task-adaptive parameters that differentiate task-specific information and capture shared knowledge across various tasks within low-dimensional spaces. This approach enables pretrained models to jointly adapt to different target domains with a limited number of trainable parameters. Comprehensive experimental results, including evaluations on public academic benchmarks for natural language understanding, commonsense reasoning, and image-text understanding, as well as real-world industrial text Ads relevance datasets, demonstrate that MTL-LoRA outperforms LoRA and its various variants with comparable or even fewer learnable parameters in MTL setting.

EAAI Journal 2025 Journal Article

Non-repetitive motion planning for autonomous vehicles under perceptual uncertainty

  • Yuqing Cao
  • Hao Sun
  • Guangjing Yang

Under perceptual uncertainty stemming from occlusion, autonomous vehicles lacking complete information are highly susceptible to collisions with pedestrians suddenly emerging from occluded areas. This problem is often tackled by repetitive training to obtain the distribution of uncertain information, but it is inadequate for addressing untrained new environments. We propose a framework for guiding vehicles in recognizing uncertainty. The framework uses dynamic probabilities to represent pedestrian uncertainty, with different states of pedestrians corresponding to games under different branches of nature choices. The cognitive probabilities will be updated over time until reaching certainty. After establishing cognition at a specific moment, the concept of equilibrium fusion is introduced to address the challenges of strategy-making in the face of multiple potential games under uncertainty. Our proposed framework demonstrates the capability to achieve effective collision avoidance planning in non-repetitive training environments, with behavioral patterns closely resembling those observed in human driving styles, and achieve results comparable to repetitive training.

NeurIPS Conference 2025 Conference Paper

PC-Net: Weakly Supervised Compositional Moment Retrieval via Proposal-Centric Network

  • Mingyao Zhou
  • Hao Sun
  • Wei Xie
  • Ming Dong
  • Chengji Wang
  • Mang Ye

With the exponential growth of video content, aiming at localizing relevant video moments based on natural language queries, video moment retrieval (VMR) has gained significant attention. Existing weakly supervised VMR methods focus on designing various feature modeling and modal interaction modules to alleviate the reliance on precise temporal annotations. However, these methods have poor generalization capabilities on compositional queries with novel syntactic structures or vocabulary in real-world scenarios. To this end, we propose a new task: weakly supervised compositional moment retrieval (WSCMR). This task trains models using only video-query pairs without precise temporal annotations, while enabling generalization to complex compositional queries. Furthermore, a proposal-centric network (PC-Net) is proposed to tackle this challenging task. First, video and query features are extracted through frozen feature extractors, followed by modality interaction to obtain multimodal features. Second, to handle compositional queries with explicit temporal associations, a dual-granularity proposal generator decodes multimodal global and frame-level features to obtain query-relevant proposal boundaries with fine-grained temporal perception. Third, to improve the discrimination of proposal features, a proposal feature aggregator is constructed to conduct semantic alignment of frames and queries, and employ a learnable peak-aware Gaussian distributor to fit the frame weights within the proposals to derive proposal features from the video frame features. Finally, the proposal quality is assessed based on the results of reconstructing the masked query using the obtained proposal features. To further enhance the model's ability to capture semantic associations between proposals and queries, a quality margin regularizer is constructed to dynamically stratify proposals into high and low query-relevance subsets and enhance the association between queries and common elements within proposals, and suppress spurious correlations via inter-subset contrastive learning. Notably, PC-Net achieves superior performance with 54\% fewer parameters than prior works by parameter-efficient design. Experiments on Charades-CG and ActivityNet-CG demonstrate PC-Net’s ability to generalize across diverse compositional queries. Code is available at https: //github. com/mingyao1120/PC-Net.

IJCAI Conference 2025 Conference Paper

PeSANet: Physics-encoded Spectral Attention Network for Simulating PDE-Governed Complex Systems

  • Han Wan
  • Rui Zhang
  • Qi Wang
  • Yang Liu
  • Hao Sun

Accurately modeling and forecasting complex systems governed by partial differential equations (PDEs) is crucial in various scientific and engineering domains. However, traditional numerical methods struggle in real-world scenarios due to incomplete or unknown physical laws. Meanwhile, machine learning approaches often fail to generalize effectively when faced with scarce observational data and the challenge of capturing local and global features. To this end, we propose the Physics-encoded Spectral Attention Network (PeSANet), which integrates local and global information to forecast complex systems with limited data and incomplete physical priors. The model consists of two key components: a physics-encoded block that uses hard constraints to approximate local differential operators from limited data, and a spectral-enhanced block that captures long-range global dependencies in the frequency domain. Specifically, we introduce a novel spectral attention mechanism to model inter-spectrum relationships and learn long-range spatial features. Experimental results demonstrate that PeSANet outperforms existing methods across all metrics, particularly in long-term forecasting accuracy, providing a promising solution for simulating complex systems with limited data and incomplete physics.

JBHI Journal 2025 Journal Article

Retrieval Augmented Thought Process for Private Data Handling in Healthcare

  • Thomas Pouplin
  • Hao Sun
  • Sam Holt
  • Mihaela van der Schaar

Large Language Models (LLMs) have demonstrated the strong potential to assist both clinicians and the general public with their extensive medical knowledge. However, their application in healthcare is constrained due to concerns about the privacy of data used in training, which prevents the integration of private and personal information because of security and ethical issues. Moreover, if their capabilities can be enhanced with information retrieval to access up-to-date knowledge, the current integration of LLMs with Information retrieval lacks robustness to imperfect retrieval, which can hinder their effectiveness and even reduce overall performance. In this work, we address this challenge by introducing the Retrieval-Augmented Thought Process (RATP). Given access to external knowledge, RATP formulates the thought generation of LLMs as a multiple-step decision process. To optimise such a thought process, RATP leverages Monte-Carlo Tree Search and learns a proxy reward function that permits cost-efficient inference. On a private dataset of electronic medical records, deliberately excluded from any LLM training set, RATP achieves 35% additional accuracy compared to in-context retrieval-augmented generation for the question-answering task.

JBHI Journal 2025 Journal Article

S 2 Match: Revisiting Weak-to-Strong Consistency from a Semantic Similarity Perspective for Semi-supervised Medical Image Segmentation

  • Shiao Xie
  • Hongyi Wang
  • Ziwei Niu
  • Hao Sun
  • Shuyi Ouyang
  • Yen-Wei Chen
  • Lanfen Lin

Semi-supervised learning (SSL) for medical image segmentation is a challenging yet highly practical task, which reduces reliance on large-scale labeled datasets by leveraging unlabeled samples. Among SSL techniques, the weak-to-strong consistency framework, popularized by FixMatch, has emerged as a state-of-the-art method in classification tasks. Notably, such a simple pipeline has also shown competitive performance in medical image segmentation. However, two key limitations still persist, impeding its efficient adaptation: (1) the neglect of contextual dependencies results in inconsistent predictions for similar semantic features, leading to incomplete object segmentation; (2) the lack of exploitation on semantic similarity between labeled and unlabeled data induces considerable class-distribution discrepancy. To address these limitations, we propose a novel SSL framework for medical image segmentation, named S 2 Match, powered by two appealing designs from a semantic similarity perspective: (1) rectifying pixel-wise prediction by reasoning about the intra-image pair-wise affinity map, thus integrating contextual dependencies explicitly into the final prediction; (2) bridging labeled and unlabeled data via a feature querying mechanism for compact class representation learning, which fully considers cross-image anatomical similarities. As the reliable semantic similarity extraction depends on robust features, we further introduce an effective Spatial-aware Fusion Module (SFM) to explore distinctive information from multiple scales. Experiments show that S 2 Match yields consistent improvements over the state-of-the-art methods across five public medical image segmentation benchmarks, exhibiting competitive performance on both 2D and 3D tasks.

EAAI Journal 2025 Journal Article

Sparse large-scale multi-objective optimization algorithm based on impact factor assistance

  • Ziyu Hu
  • Xuetao Nie
  • Hao Sun
  • Lixin Wei
  • Jinlu Zhang
  • Cong Wang

In the real world, there exists a special category of multi-objective optimization problems with more than 1000 decision variables. However, only a few decision variables play a crucial role in optimizing the objective functions. Such problems are defined as sparse large-scale multi-objective optimization problems (SLSMOPs). Due to the difficulty in effectively identifying the non-zero positions of decision variables, traditional evolutionary optimization algorithms suffer from slow convergence speed and poor convergence effect, which means it is unable to efficiently obtain the Pareto optimal solution set. To address this challenge, the Impact Factor Assisted Algorithm (IFA) is proposed, which adopts a novel initial population strategy to generate sparse populations. Meanwhile, the impact factor of each decision variable is calculated, serving as a key basis for measuring the importance of each decision variable. During the algorithm’s operation, the impact factors are iteratively updated to rationally group decision variables and guide population evolution. This approach can accurately identify the positions of non-zero decision variables. The experimental results on eight benchmark and real-world problems indicate that the algorithm outperforms several existing sparse large-scale multi-objective optimization algorithms (SLSMOEAs).

IROS Conference 2025 Conference Paper

SparseMeXt: Unlocking the Potential of Sparse Representations for HD Map Construction

  • Anqing Jiang
  • Jinhao Chai
  • Yu Gao 0042
  • Yiru Wang 0001
  • Yuwen Heng
  • Zhigang Sun 0001
  • Hao Sun
  • Zezhong Zhao

Recent advancements in high-definition (HD) map construction have demonstrated the effectiveness of dense representations, which heavily rely on computationally intensive bird’s-eye view (BEV) features. While sparse representations offer a more efficient alternative by avoiding dense BEV processing, existing methods often lag behind due to the lack of tailored designs. These limitations have hindered the competitiveness of sparse representations in online HD map construction. In this work, we systematically revisit and enhance sparse representation techniques, identifying key architectural and algorithmic improvements that bridge the gap with—and ultimately surpass—dense approaches. We introduce a dedicated network architecture optimized for sparse map feature extraction, a sparse-dense segmentation auxiliary task to better leverage geometric and semantic cues, and a denoising module guided by physical priors to refine predictions. Through these enhancements, our method achieves state-of-the-art performance on the nuScenes dataset, significantly advancing HD map construction and centerline detection. Specifically, SparseMeXt-Tiny reaches a mean average precision (mAP) of 55. 5% at 32 frames per second (fps), while SparseMeXt-Base attains 65. 2% mAP. Scaling the backbone and decoder further, SparseMeXt-Large achieves an mAP of 68. 9% at over 20 fps, establishing a new benchmark for sparse representations in HD map construction. These results underscore the untapped potential of sparse methods, challenging the conventional reliance on dense representations and redefining efficiency-performance trade-offs in the field.

AAAI Conference 2025 Conference Paper

SpotActor: Training-Free Layout-Controlled Consistent Image Generation

  • Jiahao Wang
  • Caixia Yan
  • Weizhan Zhang
  • Haonan Lin
  • Mengmeng Wang
  • Guang Dai
  • Tieliang Gong
  • Hao Sun

Text-to-image diffusion models significantly enhance the efficiency of artistic creation with high-fidelity image generation. However, in typical application scenarios like comic book production, they can neither place each subject into its expected spot nor maintain the consistent appearance of each subject across images. For these issues, we pioneer a novel task, Layout-to-Consistent-Image (L2CI) generation, which produces consistent and compositional images in accordance with the given layout conditions and text prompts. To accomplish this challenging task, we present a new formalization of dual energy guidance with optimization in a dual semantic-latent space and thus propose a training-free pipeline, SpotActor, which features a layout-conditioned optimizing stage and a consistent sampling stage. In the optimizing stage, we innovate a nuanced layout energy function to mimic the attention activations with a sigmoid-like objective. While in the sampling stage, we design Regional Interconnection Self-Attention (RISA) and Semantic Fusion Cross-Attention (SFCA) mechanisms that allow mutual interactions across images. To evaluate the performance, we present ActorBench, a specified benchmark with hundreds of reasonable prompt-box pairs stemming from object detection datasets. Comprehensive experiments are conducted to demonstrate the effectiveness of our method. The results prove that SpotActor fulfills the expectations of this task and showcases the potential for practical applications with superior layout alignment, subject consistency, prompt conformity and background diversity.

AAAI Conference 2025 Conference Paper

Trusted Unified Feature-Neighborhood Dynamics for Multi-View Classification

  • Haojian Huang
  • Chuanyu Qin
  • Zhe Liu
  • Kaijing Ma
  • Jin Chen
  • Han Fang
  • Chao Ban
  • Hao Sun

Multi-view classification (MVC) faces inherent challenges due to domain gaps and inconsistencies across different views, often resulting in uncertainties during the fusion process. While Evidential Deep Learning (EDL) has been effective in addressing view uncertainty, existing methods predominantly rely on the Dempster-Shafer combination rule, which is sensitive to conflicting evidence and often neglects the critical role of neighborhood structures within multi-view data. To address these limitations, we propose a Trusted Unified Feature-NEighborhood Dynamics (TUNED) model for robust MVC. This method effectively integrates local and global feature-neighborhood (F-N) structures for robust decision-making. Specifically, we begin by extracting local F-N structures within each view. To further mitigate potential uncertainties and conflicts in multi-view fusion, we employ a selective Markov random field that adaptively manages cross-view neighborhood dependencies. Additionally, we employ a shared parameterized evidence extractor that learns global consensus conditioned on local F-N structures, thereby enhancing the global integration of multi-view features. Experiments on benchmark datasets show that our method improves accuracy and robustness over existing approaches, particularly in scenarios with high uncertainty and conflicting views.

EAAI Journal 2024 Journal Article

A dual strategy of adaptive knee-point guidance and niche sampling for non-cyclic dynamic multiobjective optimization problems

  • Hao Sun
  • Cong Wang
  • Ziyu Hu

Current theoretical studies on dynamic multiobjective optimization problems (DMOPs) are based on periodic problems, while most DMOPs in real life have non-cyclic variations. To better reflect real-life problems, a non-cyclic benchmark test suite was designed to better test the performance of algorithms. The proposed test suite contains 15 problems that introduce difficult and complex geometric features. These features will become increasingly complex over time. Thus, a dual strategy of adaptive knee-point guidance and niche sampling (AKGNS) is proposed. For non-cyclic problems with complex geometric features, knee points are used as feature points to represent the variation of the Pareto front (PF). When faced with increasingly complex PFs, the number of subregions is adaptively adjusted according to the predicted PF complexity such that almost all knee points are identified. To maintain population diversity, new individuals are obtained by sampling around the knee point. The competitiveness of AKGNS was proven by comparing it with four advanced algorithms. AKGNS performed best on 26, 24, and 24 of the 45 test problems in terms of Schott’s Spacing Metric (SP), Hypervolume Difference (HVD), and Inverted Generational Distance (IGD) metrics, respectively.

YNICL Journal 2024 Journal Article

Aberrant high-beta band functional connectivity during reward processing in melancholic major depressive disorder: An MEG study

  • Qiaoyang Zhang
  • Yishan Du
  • Ciqing Bao
  • Lingling Hua
  • Rui Yan
  • Zhongpeng Dai
  • Yi Xia
  • Haowen Zou

OBJECTIVE: To identify the spatial-temporal pattern variation of whole-brain functional connectivity (FC) during reward processing in melancholic major depressive disorder (MDD) patients, and to determine the clinical correlates of connectomic differences. METHODS: 61 MDD patients and 32 healthy controls were enrolled into the study. During magnetoencephalography (MEG) scanning, all participants completed the facial emotion recognition task. The MDD patients were further divided into two groups: melancholic (n = 31) and non-melancholic (n = 30), based on the Mini International Neuropsychiatric Interview (M.I.N.I.) assessment. Melancholic symptoms were examined by using the 6-item melancholia subscale from the Hamilton Depression Rating Scale (HAM-D6). The whole-brain orthogonalized power envelope connections in the high-beta band (20-35 Hz) were constructed in each period after the happy emotional stimuli (0-200 ms, 100-300 ms, 200-400 ms, 300-500 ms, and 400-600 ms). Then, the network-based statistic (NBS) was used to determine the specific abnormal connection patterns in melancholic MDD patients. RESULTS: The NBS identified a sub-network difference at the mid-late period (300-500 ms) in response to happy faces among the three groups (corrected P = 0.035). Then, the post hoc and correlation analyses found five FCs were decreased in melancholic MDD patients and were related to HAM-D6 score, including FCs of left fusiform gyrus-right orbital inferior frontal gyrus (r = -0.52, P < 0.001), left fusiform gyrus-left amygdala (r = -0.26, P = 0.049), left posterior cingulate gyrus-right precuneus (r = -0.32, P = 0.025), left precuneus-right precuneus (r = -0.27, P = 0.049), and left precuneus-left inferior occipital gyrus (r = -0.32, P = 0.025). CONCLUSION: In response to happy faces, melancholic MDD patients demonstrated a disrupted functional connective pattern (20-35 Hz, 300-500 ms), which involved brain regions in visual information processing and the limbic system. The aberrant functional connective pattern in reward processing might be a biomarker of melancholic MDD.

NeurIPS Conference 2024 Conference Paper

Accelerating Non-Maximum Suppression: A Graph Theory Perspective

  • King-Siong Si
  • Lu Sun
  • Weizhan Zhang
  • Tieliang Gong
  • Jiahao Wang
  • Jiang Liu
  • Hao Sun

Non-maximum suppression (NMS) is an indispensable post-processing step in object detection. With the continuous optimization of network models, NMS has become the ``last mile'' to enhance the efficiency of object detection. This paper systematically analyzes NMS from a graph theory perspective for the first time, revealing its intrinsic structure. Consequently, we propose two optimization methods, namely QSI-NMS and BOE-NMS. The former is a fast recursive divide-and-conquer algorithm with negligible mAP loss, and its extended version (eQSI-NMS) achieves optimal complexity of $\mathcal{O}(n\log n)$. The latter, concentrating on the locality of NMS, achieves an optimization at a constant level without an mAP loss penalty. Moreover, to facilitate rapid evaluation of NMS methods for researchers, we introduce NMS-Bench, the first benchmark designed to comprehensively assess various NMS methods. Taking the YOLOv8-N model on MS COCO 2017 as the benchmark setup, our method QSI-NMS provides $6. 2\times$ speed of original NMS on the benchmark, with a $0. 1\%$ decrease in mAP. The optimal eQSI-NMS, with only a $0. 3\%$ mAP decrease, achieves $10. 7\times$ speed. Meanwhile, BOE-NMS exhibits $5. 1\times$ speed with no compromise in mAP.

NeurIPS Conference 2024 Conference Paper

Cross-model Control: Improving Multiple Large Language Models in One-time Training

  • Jiayi Wu
  • Hao Sun
  • Hengyi Cai
  • Lixin Su
  • Shuaiqiang Wang
  • Dawei Yin
  • Xiang Li
  • Ming Gao

The number of large language models (LLMs) with varying parameter scales and vocabularies is increasing. While they deliver powerful performance, they also face a set of common optimization needs to meet specific requirements or standards, such as instruction following or avoiding the output of sensitive information from the real world. However, how to reuse the fine-tuning outcomes of one model to other models to reduce training costs remains a challenge. To bridge this gap, we introduce Cross-model Control (CMC), a method that improves multiple LLMs in one-time training with a portable tiny language model. Specifically, we have observed that the logit shift before and after fine-tuning is remarkably similar across different models. Based on this insight, we incorporate a tiny language model with a minimal number of parameters. By training alongside a frozen template LLM, the tiny model gains the capability to alter the logits output by the LLMs. To make this tiny language model applicable to models with different vocabularies, we propose a novel token mapping strategy named PM-MinED. We have conducted extensive experiments on instruction tuning and unlearning tasks, demonstrating the effectiveness of CMC. Our code is available at https: //github. com/wujwyi/CMC

TMLR Journal 2024 Journal Article

GOPlan: Goal-conditioned Offline Reinforcement Learning by Planning with Learned Models

  • Mianchu Wang
  • Rui Yang
  • Xi Chen
  • Hao Sun
  • Meng Fang
  • Giovanni Montana

Offline Goal-Conditioned RL (GCRL) offers a feasible paradigm for learning general-purpose policies from diverse and multi-task offline datasets. Despite notable recent progress, the predominant offline GCRL methods, mainly model-free, face constraints in handling limited data and generalizing to unseen goals. In this work, we propose Goal-conditioned Offline Planning (GOPlan), a novel model-based framework that contains two key phases: (1) pretraining a prior policy capable of capturing multi-modal action distribution within the multi-goal dataset; (2) employing the reanalysis method with planning to generate imagined trajectories for funetuning policies. Specifically, we base the prior policy on an advantage-weighted conditioned generative adversarial network, which facilitates distinct mode separation, mitigating the pitfalls of out-of-distribution (OOD) actions. For further policy optimization, the reanalysis method generates high-quality imaginary data by planning with learned models for both intra-trajectory and inter-trajectory goals. With thorough experimental evaluations, we demonstrate that GOPlan achieves state-of-the-art performance on various offline multi-goal navigation and manipulation tasks. Moreover, our results highlight the superior ability of GOPlan to handle small data budgets and generalize to OOD goals.

EAAI Journal 2024 Journal Article

Intelligent void identification of particle packing system of caved ore and rock

  • Hao Sun
  • Zongsheng Dai
  • Lishan Zhao
  • Lichang Wei
  • Junze Jia
  • Shenggui Zhou
  • Jianxin Wang
  • Zhen Chi

Analyzing the void structure of the packing system of caved ore and rock and establishing the causal relationship between void structure and macroscopic properties are critical research areas in the field of gravity flow. However, the current method of analyzing void structure primarily involves manual threshold segmentation, leading to low analysis efficiency and the need for improved accuracy. This study introduces a novel approach using a dataset of two-dimensional Computed Tomography (CT) slices of an irregular limestone particle packing system. The primary goal is to enhance existing algorithms by focusing on data and network aspects to address the challenges in identifying this type of image dataset. As a result, a model named Data Attention gate with Recurrent Residual convolutional neural network based on U-net (DA2RU-net) is developed and evaluated for its convergence, accuracy, and robustness. The findings indicate the following: (1) To identify voids in the packing system of caved ore and rock using CT slices, it is advisable to employ 1000 images with dimensions between 300 and 400 pixels. It is essential to ensure that both large and small voids are represented in the image dataset. (2) The DA2RU-net model demonstrates superior performance with an average Dice value of 0. 9859 for identifying large voids and 0. 9654 for identifying small voids, surpassing other iterations of the U-net model and traditional algorithms. (3) The DA2RU-net model shows robustness to variations in brightness levels.

ICRA Conference 2024 Conference Paper

LSSAttn: Towards Dense and Accurate View Transformation for Multi-modal 3D Object Detection

  • Qi Jiang
  • Hao Sun

Fusing the camera and LiDAR information in the unified BEV representation serves as the elegant paradigm for the 3D detection tasks. Current multi-modal fusion methods in BEV can be categorized into LSS-based and Transformer-based in terms of their view transformation. The former leverages inaccurate depth prediction and massive pseudo points for perspective-to-BEV transformation while the latter only fetches sparse image features to the BEV representation. To overcome their shortcomings, an optimized view transformation is proposed, which can be easily modulated into the LSS-based methods. The proposed module capitalizes on the LSS mechanism to establish dense associations between perspective pixels and BEV grids. It utilizes the attention mechanism to compute similarity scores for each associated pair during feature aggregation. Starting from the BEVFusion baseline, we further introduce (1) cross-attention within the associated subsets to transfer image features into the BEV, and (2) a multi-scale feature fusion mechanism for LSS-based view transformation. Extensive experiments on nuScenes validate the effectiveness and efficiency of our proposed module, which achieves an increase of 1. 3% in mAP compared to the baseline model.

AIIM Journal 2024 Journal Article

Non-invasive fractional flow reserve derived from reduced-order coronary model and machine learning prediction of stenosis flow resistance

  • Yili Feng
  • Ruisen Fu
  • Hao Sun
  • Xue Wang
  • Yang Yang
  • Chuanqi Wen
  • Yaodong Hao
  • Yutong Sun

Background and objective Recently, computational fluid dynamics enables the non-invasive calculation of fractional flow reserve (FFR) based on 3D coronary model, but it is time-consuming. Currently, machine learning technique has emerged as an efficient and reliable approach for prediction, which allows saving a lot of analysis time. This study aimed at developing a simplified FFR prediction model for rapid and accurate assessment of functional significance of stenosis. Methods A reduced-order lumped parameter model (LPM) of coronary system and cardiovascular system was constructed for rapidly simulating coronary flow, in which a machine learning model was embedded for accurately predicting stenosis flow resistance at a given flow from anatomical features of stenosis. Importantly, the LPM was personalized in both structures and parameters according to coronary geometries from computed tomography angiography and physiological measurements such as blood pressure and cardiac output for personalized simulations of coronary pressure and flow. Coronary lesions with invasive FFR ≤ 0. 80 were defined as hemodynamically significant. Results A total of 91 patients (93 lesions) who underwent invasive FFR were involved in FFR derived from machine learning (FFRML) calculation. Of the 93 lesions, 27 lesions (29. 0%) showed lesion-specific ischemia. The average time of FFRML simulation was about 10 min. On a per-vessel basis, the FFRML and FFR were significantly correlated (r = 0. 86, p < 0. 001). The diagnostic accuracy, sensitivity, specificity, positive predictive value and negative predictive value were 91. 4%, 92. 6%, 90. 9%, 80. 6% and 96. 8%, respectively. The area under the receiver-operating characteristic curve of FFRML was 0. 984. Conclusion In this selected cohort of patients, the FFRML improves the computational efficiency and ensures the accuracy. The favorable performance of FFRML approach greatly facilitates its potential application in detecting hemodynamically significant coronary stenosis in future routine clinical practice.

NeurIPS Conference 2024 Conference Paper

OneActor: Consistent Subject Generation via Cluster-Conditioned Guidance

  • Jiahao Wang
  • Caixia Yan
  • Haonan Lin
  • Weizhan Zhang
  • Mengmeng Wang
  • Tieliang Gong
  • Guang Dai
  • Hao Sun

Text-to-image diffusion models benefit artists with high-quality image generation. Yet their stochastic nature hinders artists from creating consistent images of the same subject. Existing methods try to tackle this challenge and generate consistent content in various ways. However, they either depend on external restricted data or require expensive tuning of the diffusion model. For this issue, we propose a novel one-shot tuning paradigm, termed OneActor. It efficiently performs consistent subject generation solely driven by prompts via a learned semantic guidance to bypass the laborious backbone tuning. We lead the way to formalize the objective of consistent subject generation from a clustering perspective, and thus design a cluster-conditioned model. To mitigate the overfitting challenge shared by one-shot tuning pipelines, we augment the tuning with auxiliary samples and devise two inference strategies: semantic interpolation and cluster guidance. These techniques are later verified to significantly improve the generation quality. Comprehensive experiments show that our method outperforms a variety of baselines with satisfactory subject consistency, superior prompt conformity as well as high image quality. Our method is capable of multi-subject generation and compatible with popular diffusion extensions. Besides, we achieve a $4\times$ faster tuning speed than tuning-based baselines and, if desired, avoid increasing the inference time. Furthermore, our method can be naturally utilized to pre-train a consistent subject generation network from scratch, which will implement this research task into more practical applications. (Project page: https: //johnneywang. github. io/OneActor-webpage/)

NeurIPS Conference 2024 Conference Paper

Over-parameterized Student Model via Tensor Decomposition Boosted Knowledge Distillation

  • Yu-Liang Zhan
  • Zhong-Yi Lu
  • Hao Sun
  • Ze-Feng Gao

Increased training parameters have enabled large pre-trained models to excel in various downstream tasks. Nevertheless, the extensive computational requirements associated with these models hinder their widespread adoption within the community. We focus on Knowledge Distillation (KD), where a compact student model is trained to mimic a larger teacher model, facilitating the transfer of knowledge of large models. In contrast to much of the previous work, we scale up the parameters of the student model during training, to benefit from over-parameterization without increasing the inference latency. In particular, we propose a tensor decomposition strategy that effectively over-parameterizes the relatively small student model through an efficient and nearly lossless decomposition of its parameter matrices into higher-dimensional tensors. To ensure efficiency, we further introduce a tensor constraint loss to align the high-dimensional tensors between the student and teacher models. Comprehensive experiments validate the significant performance enhancement by our approach in various KD tasks, covering computer vision and natural language processing areas. Our code is available at https: //github. com/intell-sci-comput/OPDF.

NeurIPS Conference 2024 Conference Paper

P$^2$C$^2$Net: PDE-Preserved Coarse Correction Network for efficient prediction of spatiotemporal dynamics

  • Qi Wang
  • Pu Ren
  • Hao Zhou
  • Xin-Yang Liu
  • Zhiwen Deng
  • Yi Zhang
  • Ruizhi Chengze
  • Hongsheng Liu

When solving partial differential equations (PDEs), classical numerical methods often require fine mesh grids and small time stepping to meet stability, consistency, and convergence conditions, leading to high computational cost. Recently, machine learning has been increasingly utilized to solve PDE problems, but they often encounter challenges related to interpretability, generalizability, and strong dependency on rich labeled data. Hence, we introduce a new PDE-Preserved Coarse Correction Network (P$^2$C$^2$Net) to efficiently solve spatiotemporal PDE problems on coarse mesh grids in small data regimes. The model consists of two synergistic modules: (1) a trainable PDE block that learns to update the coarse solution (i. e. , the system state), based on a high-order numerical scheme with boundary condition encoding, and (2) a neural network block that consistently corrects the solution on the fly. In particular, we propose a learnable symmetric Conv filter, with weights shared over the entire model, to accurately estimate the spatial derivatives of PDE based on the neural-corrected system state. The resulting physics-encoded model is capable of handling limited training data (e. g. , 3--5 trajectories) and accelerates the prediction of PDE solutions on coarse spatiotemporal grids while maintaining a high accuracy. P$^2$C$^2$Net achieves consistent state-of-the-art performance with over 50\% gain (e. g. , in terms of relative prediction error) across four datasets covering complex reaction-diffusion processes and turbulent flows.

AAAI Conference 2024 Conference Paper

SAUI: Scale-Aware Unseen Imagineer for Zero-Shot Object Detection

  • Jiahao Wang
  • Caixia Yan
  • Weizhan Zhang
  • Huan Liu
  • Hao Sun
  • Qinghua Zheng

Zero-shot object detection (ZSD) aims to localize and classify unseen objects without access to their training annotations. As a prevailing solution to ZSD, generation-based methods synthesize unseen visual features by taking seen features as reference and class semantic embeddings as guideline. Although previous works continuously improve the synthesis quality, they fail to consider the scale-varying nature of unseen objects. The generation process is preformed over a single scale of object features and thus lacks scale-diversity among synthesized features. In this paper, we reveal the scale-varying challenge in ZSD and propose a Scale-Aware Unseen Imagineer (SAUI) to lead the way of a novel scale-aware ZSD paradigm. To obtain multi-scale features of seen-class objects, we design a specialized coarse-to-fine extractor to capture features through multiple scale-views. To generate unseen features scale by scale, we innovate a Series-GAN synthesizer along with three scale-aware contrastive components to imagine separable, diverse and robust scale-wise unseen features. Extensive experiments on PASCAL VOC, COCO and DIOR datasets demonstrate SAUI's better performance in different scenarios, especially for scale-varying and small objects. Notably, SAUI achieves the new state-of-the art performance on COCO and DIOR.

AAAI Conference 2024 Conference Paper

TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection

  • Hao Sun
  • Mingyao Zhou
  • Wenjing Chen
  • Wei Xie

Video moment retrieval (MR) and highlight detection (HD) based on natural language queries are two highly related tasks, which aim to obtain relevant moments within videos and highlight scores of each video clip. Recently, several methods have been devoted to building DETR-based networks to solve both MR and HD jointly. These methods simply add two separate task heads after multi-modal feature extraction and feature interaction, achieving good performance. Nevertheless, these approaches underutilize the reciprocal relationship between two tasks. In this paper, we propose a task-reciprocal transformer based on DETR (TR-DETR) that focuses on exploring the inherent reciprocity between MR and HD. Specifically, a local-global multi-modal alignment module is first built to align features from diverse modalities into a shared latent space. Subsequently, a visual feature refinement is designed to eliminate query-irrelevant information from visual features for modal interaction. Finally, a task cooperation module is constructed to refine the retrieval pipeline and the highlight score prediction process by utilizing the reciprocity between MR and HD. Comprehensive experiments on QVHighlights, Charades-STA and TVSum datasets demonstrate that TR-DETR outperforms existing state-of-the-art methods. Codes are available at https://github.com/mingyao1120/TR-DETR.

IJCAI Conference 2024 Conference Paper

Vision-based Discovery of Nonlinear Dynamics for 3D Moving Target

  • Zitong Zhang
  • Yang Liu
  • Hao Sun

Data-driven discovery of governing equations has kindled significant interests in many science and engineering areas. Existing studies primarily focus on uncovering equations that govern nonlinear dynamics based on direct measurement of the system states (e. g. , trajectories). Limited efforts have been placed on distilling governing laws of dynamics directly from videos for moving targets in a 3D space. To this end, we propose a vision-based approach to automatically uncover governing equations of nonlinear dynamics for 3D moving targets via raw videos recorded by a set of cameras. The approach is composed of three key blocks: (1) a target tracking module that extracts plane pixel motions of the moving target in each video, (2) a Rodrigues' rotation formula-based coordinate transformation learning module that reconstructs the 3D coordinates with respect to a predefined reference point, and (3) a spline-enhanced library-based sparse regressor that uncovers the underlying governing law of dynamics. This framework is capable of effectively handling the challenges associated with measurement data, e. g. , noise in the video, imprecise tracking of the target that causes data missing, etc. The efficacy of our method has been demonstrated through multiple sets of synthetic videos considering different nonlinear dynamics.

IROS Conference 2024 Conference Paper

VoxelContrast: Voxel Contrast-Based Unsupervised Learning for 3D Point Clouds

  • Yuxiang Qin
  • Hao Sun

The annotation process for 3D point cloud data is more complex than for image data, and training with a small amount of annotated data significantly reduces the performance of deep learning models. Unsupervised learning can better utilize large amounts of unlabeled point cloud data for model pretraining, thereby achieving excellent performance on small-scale datasets. However, many existing 3D point cloud unsupervised learning methods are primarily focused on single-object CAD point clouds and may not be suitable for larger-scale autonomous driving LiDAR point clouds. To address this challenging problem, we propose a voxel contrast-based unsupervised learning method (VoxelContrast), which adapts well to different types of point cloud data through voxelization and can be seamlessly integrated with existing model frameworks. Specifically, we utilize voxelization methods to preprocess point cloud data. Then, we incorporate voxel information into contrastive learning, facilitating the creation of more meaningful positive and negative sample pairs. Finally, we conduct unsupervised training of the model using instance discrimination as the proxy task. Our method was validated in two downstream tasks: point cloud shape classification and 3D object detection. Experimental results demonstrated that models pretrained using a substantial amount of unlabeled data can further enhance the effectiveness of existing supervised learning methods.

NeurIPS Conference 2023 Conference Paper

A Comprehensive Study on Text-attributed Graphs: Benchmarking and Rethinking

  • Hao Yan
  • Chaozhuo Li
  • Ruosong Long
  • Chao Yan
  • Jianan Zhao
  • Wenwen Zhuang
  • Jun Yin
  • Peiyan Zhang

Text-attributed graphs (TAGs) are prevalent in various real-world scenarios, where each node is associated with a text description. The cornerstone of representation learning on TAGs lies in the seamless integration of textual semantics within individual nodes and the topological connections across nodes. Recent advancements in pre-trained language models (PLMs) and graph neural networks (GNNs) have facilitated effective learning on TAGs, garnering increased research interest. However, the absence of meaningful benchmark datasets and standardized evaluation procedures for TAGs has impeded progress in this field. In this paper, we propose CS-TAG, a comprehensive and diverse collection of challenging benchmark datasets for TAGs. The CS-TAG datasets are notably large in scale and encompass a wide range of domains, spanning from citation networks to purchase graphs. In addition to building the datasets, we conduct extensive benchmark experiments over CS-TAG with various learning paradigms, including PLMs, GNNs, PLM-GNN co-training methods, and the proposed novel topological pre-training of language models. In a nutshell, we provide an overview of the CS-TAG datasets, standardized evaluation procedures, and present baseline experiments. The entire CS-TAG project is publicly accessible at \url{https: //github. com/sktsherlock/TAG-Benchmark}.

NeurIPS Conference 2023 Conference Paper

Accountability in Offline Reinforcement Learning: Explaining Decisions with a Corpus of Examples

  • Hao Sun
  • Alihan Hüyük
  • Daniel Jarrett
  • Mihaela van der Schaar

Learning controllers with offline data in decision-making systems is an essential area of research due to its potential to reduce the risk of applications in real-world systems. However, in responsibility-sensitive settings such as healthcare, decision accountability is of paramount importance, yet has not been adequately addressed by the literature. This paper introduces the Accountable Offline Controller (AOC) that employs the offline dataset as the Decision Corpus and performs accountable control based on a tailored selection of examples, referred to as the Corpus Subset. AOC operates effectively in low-data scenarios, can be extended to the strictly offline imitation setting, and displays qualities of both conservation and adaptability. We assess AOC's performance in both simulated and real-world healthcare scenarios, emphasizing its capability to manage offline control tasks with high levels of performance while maintaining accountability.

NeurIPS Conference 2023 Conference Paper

Model-enhanced Vector Index

  • Hailin Zhang
  • Yujing Wang
  • Qi Chen
  • Ruiheng Chang
  • Ting Zhang
  • Ziming Miao
  • Yingyan Hou
  • Yang Ding

Embedding-based retrieval methods construct vector indices to search for document representations that are most similar to the query representations. They are widely used in document retrieval due to low latency and decent recall performance. Recent research indicates that deep retrieval solutions offer better model quality, but are hindered by unacceptable serving latency and the inability to support document updates. In this paper, we aim to enhance the vector index with end-to-end deep generative models, leveraging the differentiable advantages of deep retrieval models while maintaining desirable serving efficiency. We propose Model-enhanced Vector Index (MEVI), a differentiable model-enhanced index empowered by a twin-tower representation model. MEVI leverages a Residual Quantization (RQ) codebook to bridge the sequence-to-sequence deep retrieval and embedding-based models. To substantially reduce the inference time, instead of decoding the unique document ids in long sequential steps, we first generate some semantic virtual cluster ids of candidate documents in a small number of steps, and then leverage the well-adapted embedding vectors to further perform a fine-grained search for the relevant documents in the candidate virtual clusters. We empirically show that our model achieves better performance on the commonly used academic benchmarks MSMARCO Passage and Natural Questions, with comparable serving latency to dense retrieval solutions.

EAAI Journal 2023 Journal Article

Self-adaptive physics-driven deep learning for seismic wave modeling in complex topography

  • Yi Ding
  • Su Chen
  • Xiaojun Li
  • Suyang Wang
  • Shaokai Luan
  • Hao Sun

Solving for the scattered wavefield is a key scientific problem in the field of seismology and earthquake engineering. Physics-informed neural networks (PINNs) developed in recent years have great potential in possibly increasing the flexibility and efficacy of seismic modeling and inversion. Inspired by self-adaptive physics-informed neural networks (SA-PINNs), we introduce a framework for modeling seismic waves in complex topography The relevant theoretical model construction was performed using the one-dimensional (1D) wave equation as an example. Using SA-PINNs and combining them with sparse initial wavefield data formed by the spectral element method (SEM), we carry out a numerical simulation of two-dimensional (2D) SH wave propagation to realize typical cases such as infinite/semi-infinite domain and arc-shaped canyon/hill topography. For complex scattered wavefields, a sequential learning method with time-domain decomposition was introduced in SA-PINNs to improve the scalability and solution accuracy of the network. The accuracy and reliability of the proposed method to simulate wave propagation in complex topography were verified by comparing the displacement seismograms calculated by the SA-PINNs method with those calculated by the SEM. The results show that the SA-PINNs have the advantage of gridless and fine-grained simulation and can realize numerical simulation conditions, such as free surface and side-boundary wavefield transmission.

IROS Conference 2023 Conference Paper

SemanticBEVFusion: Rethinking LiDAR-Camera Fusion in Unified Bird's-Eye View Representation for 3D Object Detection

  • Qi Jiang
  • Hao Sun

LiDAR and cameras are two essential sensors for 3D object detection in autonomous driving. LiDAR provides accurate and reliable 3D geometry information while the camera provides rich texture with color. Despite the increasing popularity of fusing these two complementary sensors, the challenge remains in how to effectively fuse 3D LiDAR point cloud with 2D camera images. Recent methods focus on point-level fusion which paints the LiDAR point cloud with camera features in the perspective view or bird's-eye view (BEV)-level fusion which unifies multi-modality features in the BEV representation. In this paper, we rethink these previous fusion strategies and analyze their information loss and influences on geometric and semantic features. We present SemanticBEVFusion to deeply fuse camera features with LiDAR features in a unified BEV representation while maintaining per-modality strengths for 3D object detection. Our method achieves state-of-the-art performance on the large-scale nuScenes dataset, especially for challenging distant objects. The code will be made publicly available.

YNICL Journal 2023 Journal Article

Spontaneous beta power, motor-related beta power and cortical thickness in major depressive disorder with psychomotor disturbance

  • Yi Xia
  • Hao Sun
  • Lingling Hua
  • Zhongpeng Dai
  • Xiaoqin Wang
  • Hao Tang
  • Yinglin Han
  • Yishan Du

INTRODUCTION: The psychomotor disturbance is a common symptom in patients with major depressive disorder (MDD). The neurological mechanisms of psychomotor disturbance are intricate, involving alterations in the structure and function of motor-related regions. However, the relationship among changes in the spontaneous activity, motor-related activity, local cortical thickness, and psychomotor function remains unclear. METHOD: A total of 140 patients with MDD and 68 healthy controls performed a simple right-hand visuomotor task during magnetoencephalography (MEG) scanning. All patients were divided into two groups according to the presence of psychomotor slowing. Spontaneous beta power, movement-related beta desynchronization (MRBD), absolute beta power during movement and cortical characteristics in the bilateral primary motor cortex were compared using general linear models with the group as a fixed effect and age as a covariate. Finally, the moderated mediation model was tested to examine the relationship between brain metrics with group differences and psychomotor performance. RESULTS: The patients with psychomotor slowing showed higher spontaneous beta power, movement-related beta desynchronization and absolute beta power during movement than patients without psychomotor slowing. Compared with the other two groups, significant decreases were found in cortical thickness of the left primary motor cortex in patients with psychomotor slowing. Our moderated mediation model showed that the increased spontaneous beta power indirectly affected impaired psychomotor performance by abnormal MRBD, and the indirect effects were moderated by cortical thickness. CONCLUSION: These results suggest that patients with MDD have aberrant cortical beta activity at rest and during movement, combined with abnormal cortical thickness, contributing to the psychomotor disturbance observed in this patient population.

YNICL Journal 2023 Journal Article

The relationship between disrupted anhedonia-related circuitry and suicidal ideation in major depressive disorder: A network-based analysis

  • Xiaoqin Wang
  • Yi Xia
  • Rui Yan
  • Huan Wang
  • Hao Sun
  • Yinghong Huang
  • Lingling Hua
  • Hao Tang

BACKGROUND: Several epidemiological studies and psychological models have suggested that major depressive disorder (MDD) with anhedonia is associated with suicidal ideation (SI). However, little is known about whether the functional network pattern and intrinsic topologically disrupted in patients with anhedonia are related to SI. METHODS: The resting-fMRI by applying network-based statistic (NBS) and graph-theory analyses was estimated in 273 patients with MDD (144 high anhedonia [HA], 129 low anhedonia [LA]) and 150 healthy controls. In addition, we quantified the SI scores of each patient. Finally, the mediation analysis assessed whether anhedonia symptoms could mediate the relationship between anhedonia-related network metrics and SI. RESULT: The NBS analysis demonstrated that individuals with HA have a single abnormally increased functional connectivity component in a frontal-limbic circuit (termed the "anhedonia-related network", including the frontal cortex, striatum, anterior cingulate cortex and amygdala). The graph-theory analysis demonstrated that the anhedonia-related network showed a significantly disrupted topological organization (lower gamma and lambda), which the small-world property trend randomized. Furthermore, the anhedonia symptoms could mediate the relationship between the anhedonia-related network metrics (the mean functional connectivity values, the area under the curves values of gamma and nodal local efficiency in nucleus accumbens) and SI. CONCLUSIONS: We found that disruption of the reward-related network in MDD leads to SI through anhedonia symptoms. These findings show the abnormal topological construction of functional brain network organization in anhedonia, shedding light on the neurological processes underlying SI in MDD patients with anhedonia symptoms.

NeurIPS Conference 2023 Conference Paper

What is Flagged in Uncertainty Quantification? Latent Density Models for Uncertainty Categorization

  • Hao Sun
  • Boris van Breugel
  • Jonathan Crabbé
  • Nabeel Seedat
  • Mihaela van der Schaar

Uncertainty quantification (UQ) is essential for creating trustworthy machine learning models. Recent years have seen a steep rise in UQ methods that can flag suspicious examples, however, it is often unclear what exactly these methods identify. In this work, we propose a framework for categorizing uncertain examples flagged by UQ methods. We introduce the confusion density matrix---a kernel-based approximation of the misclassification density---and use this to categorize suspicious examples identified by a given uncertainty method into three classes: out-of-distribution (OOD) examples, boundary (Bnd) examples, and examples in regions of high in-distribution misclassification (IDM). Through extensive experiments, we show that our framework provides a new and distinct perspective for assessing differences between uncertainty quantification methods, thereby forming a valuable assessment benchmark.

EAAI Journal 2022 Journal Article

A deep transfer regression method based on seed replacement considering balanced domain adaptation

  • Teng Zhang
  • Hao Sun
  • Fangyu Peng
  • Shengqiang Zhao
  • Rong Yan

With the development of deep transfer learning, the generalization abilities of models in similar scenarios have been significantly improved. However, for regression tasks, either the marginal distribution or the conditional distribution is usually ignored. In addition, initiative regarding the representation and learning of domain knowledge is lacking due to the reliance on the loss function. A deep transfer regression method based on seed replacement considering balanced domain adaptation, called DTRSR, is proposed in this work. DTRSR is composed of four parts: structure freezing and parameter transfer, deep feature extraction, seed replacement and a fusion loss function. First, domain knowledge is captured at the model level through structure freezing and parameter transfer. Second, seed replacement is used for knowledge learning in the source and target domains at the data level. Finally, a fusion loss function considering balanced distribution adaptation is constructed to acquire domain knowledge at the loss level. In summary, domain knowledge is sufficiently learned through DTRSR. In addition, seed replacement improves the initiative of knowledge learning instead of relying only on the loss function to learn automatically. DTRSR is compared on three datasets, namely, Tool Wear, Battery Capacity and Robot Machining Errors, with nine other methods. The proposed method achieves excellent performance on most tasks, which validates its effectiveness and great potential in regression tasks.

NeurIPS Conference 2022 Conference Paper

A Neural Corpus Indexer for Document Retrieval

  • Yujing Wang
  • Yingyan Hou
  • Haonan Wang
  • Ziming Miao
  • Shibin Wu
  • Qi Chen
  • Yuqing Xia
  • Chengmin Chi

Current state-of-the-art document retrieval solutions mainly follow an index-retrieve paradigm, where the index is hard to be directly optimized for the final retrieval target. In this paper, we aim to show that an end-to-end deep neural network unifying training and indexing stages can significantly improve the recall performance of traditional methods. To this end, we propose Neural Corpus Indexer (NCI), a sequence-to-sequence network that generates relevant document identifiers directly for a designated query. To optimize the recall performance of NCI, we invent a prefix-aware weight-adaptive decoder architecture, and leverage tailored techniques including query generation, semantic document identifiers, and consistency-based regularization. Empirical studies demonstrated the superiority of NCI on two commonly used academic benchmarks, achieving +21. 4% and +16. 8% relative enhancement for Recall@1 on NQ320k dataset and R-Precision on TriviaQA dataset, respectively, compared to the best baseline method.

NeurIPS Conference 2022 Conference Paper

Bayesian Spline Learning for Equation Discovery of Nonlinear Dynamics with Quantified Uncertainty

  • Luning Sun
  • Daniel Huang
  • Hao Sun
  • Jian-Xun Wang

Nonlinear dynamics are ubiquitous in science and engineering applications, but the physics of most complex systems is far from being fully understood. Discovering interpretable governing equations from measurement data can help us understand and predict the behavior of complex dynamic systems. Although extensive work has recently been done in this field, robustly distilling explicit model forms from very sparse data with considerable noise remains intractable. Moreover, quantifying and propagating the uncertainty of the identified system from noisy data is challenging, and relevant literature is still limited. To bridge this gap, we develop a novel Bayesian spline learning framework to identify parsimonious governing equations of nonlinear (spatio)temporal dynamics from sparse, noisy data with quantified uncertainty. The proposed method utilizes spline basis to handle the data scarcity and measurement noise, upon which a group of derivatives can be accurately computed to form a library of candidate model terms. The equation residuals are used to inform the spline learning in a Bayesian manner, where approximate Bayesian uncertainty calibration techniques are employed to approximate posterior distributions of the trainable parameters. To promote the sparsity, an iterative sequential-threshold Bayesian learning approach is developed, using the alternative direction optimization strategy to systematically approximate L0 sparsity constraints. The proposed algorithm is evaluated on multiple nonlinear dynamical systems governed by canonical ordinary and partial differential equations, and the merit/superiority of the proposed method is demonstrated by comparison with state-of-the-art methods.

IJCAI Conference 2022 Conference Paper

Distilling Governing Laws and Source Input for Dynamical Systems from Videos

  • Lele Luan
  • Yang Liu
  • Hao Sun

Distilling interpretable physical laws from videos has led to expanded interest in the computer vision community recently thanks to the advances in deep learning, but still remains a great challenge. This paper introduces an end-to-end unsupervised deep learning framework to uncover the explicit governing equations of dynamics presented by moving object(s), based on recorded videos. Instead in the pixel (spatial) coordinate system of image space, the physical law is modeled in a regressed underlying physical coordinate system where the physical states follow potential explicit governing equations. A numerical integrator-based sparse regression module is designed and serves as a physical constraint to the autoencoder and coordinate system regression, and, in the meanwhile, uncover the parsimonious closed-form governing equations from the learned physical states. Experiments on simulated dynamical scenes show that the proposed method is able to distill closed-form governing equations and simultaneously identify unknown excitation input for several dynamical systems recorded by videos, which fills in the gap in literature where no existing methods are available and applicable for solving this type of problem.

TIST Journal 2022 Journal Article

Efficient and Effective Similar Subtrajectory Search: A Spatial-aware Comprehension Approach

  • Liwei Deng
  • Hao Sun
  • Rui Sun
  • Yan Zhao
  • Han Su

Although many applications take subtrajectories as basic units for analysis, there is little research on the similar subtrajectory search problem aiming to return a portion of a trajectory (i.e., subtrajectory), which is the most similar to a query trajectory. We find that in some special cases, when a grid-based metric is used, this problem can be formulated as a reading comprehension problem, which has been studied extensively in the field of natural language processing (NLP). By this formulation, we can obtain faster models with better performance than existing methods. However, due to the difference between natural language and trajectory (e.g., spatial relationship), it is impossible to directly apply NLP models to this problem. Therefore, we propose a Similar Subtrajectory Search with a Graph Neural Networks framework. This framework contains four modules including a spatial-aware grid embedding module, a trajectory embedding module, a query-context trajectory fusion module, and a span prediction module. Specifically, in the spatial-aware grid embedding module, the spatial-based grid adjacency is constructed and delivered to the graph neural network to learn spatial-aware grid embedding. The trajectory embedding module aims to model the sequential information of trajectories. The purpose of the query-context trajectory fusion module is to fuse the information of the query trajectory to each grid of the context trajectories. Finally, the span prediction module aims to predict the start and the end of a subtrajectory for the context trajectory, which is the most similar to the query trajectory. We conduct comprehensive experiments on two real world datasets, where the proposed framework outperforms the state-of-the-art baselines consistently and significantly.

NeurIPS Conference 2022 Conference Paper

Exploit Reward Shifting in Value-Based Deep-RL: Optimistic Curiosity-Based Exploration and Conservative Exploitation via Linear Reward Shaping

  • Hao Sun
  • Lei Han
  • Rui Yang
  • Xiaoteng Ma
  • Jian Guo
  • Bolei Zhou

In this work, we study the simple yet universally applicable case of reward shaping in value-based Deep Reinforcement Learning (DRL). We show that reward shifting in the form of a linear transformation is equivalent to changing the initialization of the $Q$-function in function approximation. Based on such an equivalence, we bring the key insight that a positive reward shifting leads to conservative exploitation, while a negative reward shifting leads to curiosity-driven exploration. Accordingly, conservative exploitation improves offline RL value estimation, and optimistic value estimation improves exploration for online RL. We validate our insight on a range of RL tasks and show its improvement over baselines: (1) In offline RL, the conservative exploitation leads to improved performance based on off-the-shelf algorithms; (2) In online continuous control, multiple value functions with different shifting constants can be used to tackle the exploration-exploitation dilemma for better sample efficiency; (3) In discrete control tasks, a negative reward shifting yields an improvement over the curiosity-based exploration method.

ICLR Conference 2022 Conference Paper

Rethinking Goal-Conditioned Supervised Learning and Its Connection to Offline RL

  • Rui Yang 0010
  • Yiming Lu
  • Wenzhe Li
  • Hao Sun
  • Meng Fang
  • Yali Du 0001
  • Xiu Li 0001
  • Lei Han 0001

Solving goal-conditioned tasks with sparse rewards using self-supervised learning is promising because of its simplicity and stability over current reinforcement learning (RL) algorithms. A recent work, called Goal-Conditioned Supervised Learning (GCSL), provides a new learning framework by iteratively relabeling and imitating self-generated experiences. In this paper, we revisit the theoretical property of GCSL --- optimizing a lower bound of the goal reaching objective, and extend GCSL as a novel offline goal-conditioned RL algorithm. The proposed method is named Weighted GCSL (WGCSL), in which we introduce an advanced compound weight consisting of three parts (1) discounted weight for goal relabeling, (2) goal-conditioned exponential advantage weight, and (3) best-advantage weight. Theoretically, WGCSL is proved to optimize an equivalent lower bound of the goal-conditioned RL objective and generates monotonically improved policies via an iterated scheme. The monotonic property holds for any behavior policies, and therefore WGCSL can be applied to both online and offline settings. To evaluate algorithms in the offline goal-conditioned RL setting, we provide a benchmark including a range of point and simulated robot domains. Experiments in the introduced benchmark demonstrate that WGCSL can consistently outperform GCSL and existing state-of-the-art offline methods in the fully offline goal-conditioned setting.

ICRA Conference 2021 Conference Paper

Impact Mitigation for Dynamic Legged Robots with Steel Wire Transmission Using Nonlinear Active Compliance Control

  • Junjie Yang 0012
  • Hao Sun
  • Hao An
  • Changhong Wang 0004

Impact mitigation is crucial to the stable locomotion of legged robots, especially in high-speed dynamic locomotion. This paper presents a leg locomotion system, including the nonlinear active compliance control and the active impedance control for the steel wire transmission-based legged robot. The developed control system enables high-speed dynamic locomotion with excellent impact mitigation and leg position tracking performance, where three strategies are applied. a) The feed-forward controller is designed according to the linear motor-leg model with the information of Coulomb friction and viscous friction. b) Steel wire transmission model-based compensation guarantees ideal virtual spring compliance characteristics. c) Nonlinear active compliance control and active impedance control ensure better impact mitigation performance than linear scheme and guarantee position tracking performance. The proposed control system is verified on a real robot named SCIT Dog, and the experiment demonstrates the ideal impact mitigation ability in high-speed dynamic locomotion without any passive spring mechanism.

AAAI Conference 2021 Conference Paper

Knowledge Refinery: Learning from Decoupled Label

  • Qianggang Ding
  • Sifan Wu
  • Tao Dai
  • Hao Sun
  • Jiadong Guo
  • Zhang-Hua Fu
  • Shutao Xia

Recently, a variety of regularization techniques have been widely applied in deep neural networks, which mainly focus on the regularization of weight parameters to encourage generalization effectively. Label regularization techniques are also proposed with the motivation of softening the labels while neglecting the relation of classes. Among them, the technique of knowledge distillation proposes to distill the soft label, which contains the knowledge of class relations. However, this technique needs to pre-train an extra cumbersome teacher model. In this paper, we propose a method called Knowledge Refinery (KR), which enables the neural network to learn the relation of classes on-the-fly without the teacher-student training strategy. We propose the definition of decoupled labels, which consist of the original hard label and the residual label. To exhibit the generalization of KR, we evaluate our method in both fields of computer vision and natural language processing. Our empirical results show consistent performance gains under all experimental settings.

IJCAI Conference 2021 Conference Paper

Physics-informed Spline Learning for Nonlinear Dynamics Discovery

  • Fangzheng Sun
  • Yang Liu
  • Hao Sun

Dynamical systems are typically governed by a set of linear/nonlinear differential equations. Distilling the analytical form of these equations from very limited data remains intractable in many disciplines such as physics, biology, climate science, engineering and social science. To address this fundamental challenge, we propose a novel Physics-informed Spline Learning (PiSL) framework to discover parsimonious governing equations for nonlinear dynamics, based on sparsely sampled noisy data. The key concept is to (1) leverage splines to interpolate locally the dynamics, perform analytical differentiation and build the library of candidate terms, (2) employ sparse representation of the governing equations, and (3) use the physics residual in turn to inform the spline learning. The synergy between splines and discovered underlying physics leads to the robust capacity of dealing with high-level data scarcity and noise. A hybrid sparsity-promoting alternating direction optimization strategy is developed for systematically pruning the sparse coefficients that form the structure and explicit expression of the governing equations. The efficacy and superiority of the proposed method have been demonstrated by multiple well-known nonlinear dynamical systems, in comparison with two state-of-the-art methods.

JBHI Journal 2020 Journal Article

Computer-Aided Diagnosis in Histopathological Images of the Endometrium Using a Convolutional Neural Network and Attention Mechanisms

  • Hao Sun
  • Xianxu Zeng
  • Tao Xu
  • Gang Peng
  • Yutao Ma

Uterine cancer (also known as endometrial cancer) can seriously affect the female reproductive system, and histopathological image analysis is the gold standard for diagnosing endometrial cancer. Due to the limited ability to model the complicated relationships between histopathological images and their interpretations, existing computer-aided diagnosis (CAD) approaches using traditional machine learning algorithms often failed to achieve satisfying results. In this study, we develop a CAD approach based on a convolutional neural network (CNN) and attention mechanisms, called HIENet. In the ten-fold cross-validation on ∼3, 300 hematoxylin and eosin (H&E) image patches from ∼500 endometrial specimens, HIENet achieved a 76. 91 ± 1. 17% (mean ± s. d.) accuracy for four classes of endometrial tissue, i. e. , normal endometrium, endometrial polyp, endometrial hyperplasia, and endometrial adenocarcinoma. Also, HIENet obtained an area-under-the-curve (AUC) of 0. 9579 ± 0. 0103 with an 81. 04 ± 3. 87% sensitivity and 94. 78 ± 0. 87% specificity in a binary classification task that detected endometrioid adenocarcinoma. Besides, in the external validation on 200 H&E image patches from 50 randomly-selected female patients, HIENet achieved an 84. 50% accuracy in the four-class classification task, as well as an AUC of 0. 9829 with a 77. 97% (95% confidence interval, CI, 65. 27%∼87. 71%) sensitivity and 100% (95% CI, 97. 42%∼100. 00%) specificity. The proposed CAD method outperformed three human experts and five CNN-based classifiers regarding overall classification performance. It was also able to provide pathologists better interpretability of diagnoses by highlighting the histopathological correlations of local pixel-level image features to morphological characteristics of endometrial tissue.

IJCAI Conference 2020 Conference Paper

Hierarchical Multi-Scale Gaussian Transformer for Stock Movement Prediction

  • Qianggang Ding
  • Sifan Wu
  • Hao Sun
  • Jiadong Guo
  • Jian Guo

Predicting the price movement of finance securities like stocks is an important but challenging task, due to the uncertainty of financial markets. In this paper, we propose a novel approach based on the Transformer to tackle the stock movement prediction task. Furthermore, we present several enhancements for the proposed basic Transformer. Firstly, we propose a Multi-Scale Gaussian Prior to enhance the locality of Transformer. Secondly, we develop an Orthogonal Regularization to avoid learning redundant heads in the multi-head self-attention mechanism. Thirdly, we design a Trading Gap Splitter for Transformer to learn hierarchical features of high-frequency finance data. Compared with other popular recurrent neural networks such as LSTM, the proposed method has the advantage to mine extremely long-term dependencies from financial time series. Experimental results show our proposed models outperform several competitive methods in stock price prediction tasks for the NASDAQ exchange market and the China A-shares market.

IROS Conference 2019 Conference Paper

A Convolutional Network for Joint Deraining and Dehazing from A Single Image for Autonomous Driving in Rain

  • Hao Sun
  • Marcelo H. Ang
  • Daniela Rus

In this paper, we focus on a rain removal task from a single image of the urban street scene for autonomous driving in rain. We develop a Convolutional Neural Network which takes a rainy image as input, and directly recovers a clean image in the presence of rain streaks, atmospheric veiling effect (haze, fog, mist) caused by distant rain streak accumulation. We propose a synthetic dataset containing images of urban street scenes with different rain intensities, orientations and haziness levels for training and evaluation. We evaluate our method quantitatively and qualitatively on the synthetic data. Experiments show that our model outperforms state-of-the-art methods. We also test our method qualitatively on the real-world data. Our model is fast and it takes 0. 05s for an image of $1024 \times 512$. Our model can be seamlessly integrated with existing image-based high-level perception algorithms for autonomous driving in rain. Experiment results show that our deraining method improves semantic segmentation and object detection largely for autonomous driving in rain.

NeurIPS Conference 2019 Conference Paper

Policy Continuation with Hindsight Inverse Dynamics

  • Hao Sun
  • Zhizhong Li
  • Xiaotong Liu
  • Bolei Zhou
  • Dahua Lin

Solving goal-oriented tasks is an important but challenging problem in reinforcement learning (RL). For such tasks, the rewards are often sparse, making it difficult to learn a policy effectively. To tackle this difficulty, we propose a new approach called Policy Continuation with Hindsight Inverse Dynamics (PCHID). This approach learns from Hindsight Inverse Dynamics based on Hindsight Experience Replay. Enabling the learning process in a self-imitated manner and thus can be trained with supervised learning. This work also extends it to multi-step settings with Policy Continuation. The proposed method is general, which can work in isolation or be combined with other on-policy and off-policy algorithms. On two multi-goal tasks GridWorld and FetchReach, PCHID significantly improves the sample efficiency as well as the final performance.

IROS Conference 2018 Conference Paper

A 3D Convolutional Neural Network Towards Real-Time Amodal 3D Object Detection

  • Hao Sun
  • Zehui Meng
  • Xinxin Du
  • Marcelo H. Ang

We focus on the task of amodal 3D object detection, which is to predict object locations, dimensions, poses and categories in the real world. We introduce a 3D Convolutional Neural Network that takes a volumetric representation of an indoor scene as input and predicts 3D object bounding boxes, object categories, and orientations. Unlike prior state-of-the-arts, our approach does not depend on region proposal techniques to hypothesize object locations. We treat detection and recognition as one regression problem in a single network. Our elegant model is extremely fast and all predictions are reasoned from the global context of a point cloud in a continuous pipeline. We evaluate our approach on two standard datasets: the NYUv2 RGBD dataset and the SUN RGBD dataset. Experiments show that our approach is faster than start-of-the-art 3D detectors by several orders of magnitude towards real-time amodal 3D object detection.

ICRA Conference 2018 Conference Paper

Scene Recognition and Object Detection in a Unified Convolutional Neural Network on a Mobile Manipulator

  • Hao Sun
  • Zehui Meng
  • Pey Yuen Tao
  • Marcelo H. Ang

Environment understanding, object detection and recognition are crucial skills for robots operating in the real world. In this paper, we propose a Convolutional Neural Network with multi-task objectives: object detection and scene classification in one unified architecture. The proposed network reasons globally about an image to understand the scene, hypothesize object locations, and encodes global scene features with regional object features to improve object recognition. We evaluate our network on the standard SUN RGBD dataset. Experiments show that our approach outperforms state-of-the-arts. Network predictions are further transformed into continuous robot beliefs to ensure temporal coherence and extended to 3D space for robotics applications. We embed the whole framework in Robot Operating System, and evaluate its performance on a real robot for semantic mapping and grasp detection.

v2026.09.13