Arrow Research search

Author name cluster

Bo Hu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

28 papers
2 author rows

Possible papers

28

EAAI Journal 2026 Journal Article

A vehicle lateral control algorithm that directly performs safe self-learning in real-world environments

  • Ke Wu
  • Ping Lu
  • Sunan Zhang
  • Bocheng Liu
  • Xin Ye
  • Bo Hu

Lateral Control algorithms in autonomous vehicles that perform well in simulation environments often fail to guarantee performance and safety in real-world scenarios, primarily due to significant differences between the simulated and actual environments. While reinforcement learning (RL) enables vehicles to interact and learn in real-world environments to enhance their performance, ensuring safety during the learning process remains challenging. This paper takes into full consideration the errors between the simulation model and the real environment as well as the propagation of these errors during the prediction process. Through a learning-based Model Predictive Control (MPC) algorithm, the vehicle's lateral control is maintained safely during the online learning process, with the model and the final control performance being continuously refined using data generated from interactions with the environment. Additionally, known priors are utilized to guarantee the initial performance of the policy in real-world environments and accelerate the evolution of the algorithm. Simulation and experiment results show that the proposed algorithm can guarantee safety during the learning process with a high probability and achieve better performance after self-learning. In contrast to the traditional RL paradigm, which learns in a simulation environment before applying its knowledge in the real world, the proposed algorithm stands out by its capability to self-evolve directly within the real environment safely, making it an appealing option for enhancing various model-based methods that require continuous model refinement for further development.

AAMAS Conference 2026 Conference Paper

Node-Level Federated Learning with Adaptive Personalized Aggregation for Spatio-Temporal Traffic Prediction

  • Xiaoying Tu
  • Ying Lin
  • Xingjian Lu
  • Yibing Wang
  • Bo Hu

Accurate and real-time traffic flow prediction is crucial for IntelligentTransportationSystems. Recentadvancesinfederatedlearning and spatio-temporal modeling have improved accuracy and privacy protection. However, existingmethodsoftenrelyonglobaltopology for spatial features, neglecting topology protection, and typically train a generic global model without considering local personalized features, limiting prediction performance. This paper proposes ST-PFLA (Spatio-Temporal Traffic Flow Prediction via Personalized Federated Learning with Adaptive Aggregation), a framework designed for node-level scenarios where clients only have information about their respective connections, to improve prediction accuracy and training efficiency while safeguarding topology privacy. In ST-PFLA, clients conduct prediction by combining spatial and temporal features extracted by the attention mechanism and local datasets respectively. The method aggregates only encoders across clients, retaining decoders locally for personalization. Each client performs an additional local training round to generate a guide model, which is used to inform the calculation of aggregation weights. Experimental results on two public datasets show that ST-PFLA can significantly enhance prediction accuracy while safeguarding topology privacy at lower training costs.

EAAI Journal 2026 Journal Article

Transformer-based offline-to-online reinforcement learning for decision-making and control in autonomous driving

  • Feihong Tan
  • Ping Lu
  • Fulin Zhang
  • Xin Ye
  • Bo Hu
  • Xing Shu

Developing robust decision-making and control systems for autonomous driving in complex, dynamic environments involving multi-vehicle interactions at intersections, roundabouts, and merging ramps remains a significant hurdle. In this context, Reinforcement Learning (RL) emerges as a highly promising approach. The primary methods for applying RL, however, present a core dilemma. On one hand, offline RL cannot adapt well to real-world conditions because it learns from a fixed dataset. On the other hand, online RL requires learning through real-world interaction, which is inherently unsafe for driving. To address these issues, this paper proposes a Transformer-based Offline-to-online Reinforcement Learning (TORL) framework. Firstly, the framework's offline learning paradigm integrates a Transformer architecture with a maximum entropy mechanism. This synergistic approach allows the model to capture long-term temporal dependencies for high-performance decision-making and control while ensuring the initial policy is robust and generalizable. Building on this foundation, the framework employs a trifecta of synergistic mechanisms during online fine-tuning, including Human-in-the-Loop (HITL) safe exploration, a hybrid replay buffer, and a mixed data-source learning approach, to simultaneously mitigate performance degradation from distributional shifts and neutralize the critical safety risks of online exploration. Comprehensive experiments conducted in the MetaDrive simulation environment demonstrate that TORL surpasses baseline methods, achieving an absolute increase of approximately 29. 4% in normalized return and 46. 1% in task success rate, while maintaining a zero-collision record. Furthermore, the framework's real-time feasibility was validated on an experimental autonomous vehicle platform, demonstrating low computational latency suitable for practical deployment. This study demonstrates that the proposed offline-to-online RL paradigm offers a robust and effective solution for developing high-performance decision-making and control systems for autonomous vehicles.

EAAI Journal 2025 Journal Article

A knowledge-guided reinforcement learning method for lateral path tracking

  • Bo Hu
  • Sunan Zhang
  • Yuxiang Feng
  • Bingbing Li
  • Hao Sun
  • Mingyang Chen
  • Weichao Zhuang
  • Yi Zhang

Lateral Control algorithms in autonomous vehicles often necessitates an online fine-tuning procedure in the real world. While reinforcement learning (RL) enables vehicles to learn and improve the lateral control performance through repeated trial and error interactions with a dynamic environment, applying RL directly to safety-critical applications in real physical world is challenging because ensuring safety during the learning process remains difficult. To enable safe learning, a promising direction is to make use of previously gathered offline data, which is frequently accessible in engineering applications. In this context, this paper presents a set of knowledge-guided RL algorithms that can not only fully leverage the prior collected offline data without the need of a physics-based simulator, but also allow further online policy improvement in a smooth, safe and efficient manner. To evaluate the effectiveness of the proposed algorithms on a real controller, a hardware-in-the-loop and a miniature vehicle platform are built. Compared with the vanilla RL, behavior cloning and the existing controller, the proposed algorithms realize a closed-loop solution for lateral control problems from offline training to online fine-tuning, making it attractive for future similar RL-based controller to build upon.

AAAI Conference 2025 Conference Paper

Adaptive Multimodal Fusion: Dynamic Attention Allocation for Intent Recognition

  • Bo Hu
  • Kai Zhang
  • Yanghai Zhang
  • Yuyang Ye

In recent years, deep multimodal learning has seen significant advancements. However, there remains a lack of multimodal fusion methods capable of dynamically adjusting the weighting of information both within and across modalities based on input samples. In the domain of multimodal intent recognition, the text modality often contains the most relevant information for intent detection, while the audio and visual modalities provide comparatively less critical information. There is a significant variation in the density of important information across different modalities and samples. To address this challenge, we propose a Dynamic Attention Allocation Fusion (DAF) method with an adaptive network structure that dynamically allocates attention both within individual modalities and across multiple modalities. This approach enables the model to focus more effectively on the most informative modalities and their respective internal features. Furthermore, we introduce a multi-view contrastive learning framework based on DAF (MVCL-DAF). This framework uses distinct and isolated modules to process information from various modalities, taking inspiration from the way the human brain processes multimodal information. Each modality independently infers intent using its respective module, while DAF integrates the multimodal information to produce a comprehensive global intent prediction. The text modality, functioning as the primary modality due to its rich semantic content, guides the other modules in the multi-view contrastive learning process. Extensive experiments demonstrate that our approach significantly outperforms existing state-of-the-art methods.

EAAI Journal 2025 Journal Article

Adaptive self-evolving extreme learning machine-based terminal sliding mode control with application in retinal vein injection

  • Bo Hu
  • Shiyu Xu
  • Lu Liu
  • Rongxin Liu
  • Mingzhu Sun
  • Xin Zhao

Retinal vein occlusion (RVO) is a serious condition that can lead to blindness. Injecting drugs into the retinal vein is a promising procedure for treating RVO. Due to the fragility of the retinal tissue, maintaining a precise drug flow rate (DFR) with a fast response is critical. Considering the unknown disturbance from piston dynamic and the drug-vein interaction, an adaptive self-evolving neural terminal sliding mode (ASNTSM) controller is proposed for DFR tracking. The integral terminal sliding surface is adopted to track the desired DFR in finite-time. The extreme learning machine (ELM) is utilized to estimate overall disturbances, and the adaptive switching gain is employed to compensate for the estimation error without requiring prior bounds. To achieve a compact ELM structure, a self-evolving mechanism is designed to implement the growth or pruning strategy of the hidden neurons. Theoretical analysis has proven that the ASNTSM controller can guarantee finite-time stability. Comparative experiments are conducted using a silicon phantom with simulated blood flow disturbances. The experimental results illustrate that the ASNTSM controller not only achieves lower transient time and average steady-state error, but also exhibits lower fluctuation and chattering effect. The self-evolving mechanism enhances the practicability of neural network in artificial intelligence-based medical engineering. Therefore, the ASNTSM controller is suitable for retinal vein injection tasks to improve surgical efficiency.

EAAI Journal 2025 Journal Article

An uncertainty-aware safe-evolving reinforcement learning algorithm for decision-making and control in highway autonomous driving

  • Ping Lu
  • Sunan Zhang
  • Feihong Tan
  • Fulin Zhang
  • Yuxiang Feng
  • Bo Hu

Rule-based and optimization-based approaches face challenges in decision-making and control for autonomous vehicles (AVs) in dynamic and complex highway scenarios. In contrast, reinforcement learning (RL) offers a more flexible and adaptable data-driven solution by allowing AVs to learn optimal actions through interactions with the environment, without requiring predefined rules or explicit programming. However, in real-world highway environments characterized by uncertainty, RL algorithm encounter difficulties in ensuring stability and safety. To address these challenges, this paper proposes an uncertainty-aware safe-evolving RL algorithm that integrates internal stability, external stability, and provable safety mechanisms. The internal stability mechanism ensures consistent performance improvements with high probability during policy updates itself, while the external stability leverages a benchmark policy as a reference to ensure the current policy performs at least as well as, if not better than, the benchmark. Furthermore, an action projection mechanism and a mixed learning procedure are incorporated to make minimal modifications to the learned policy, ensuring safety while supporting stable learning from both safe and original actions. The results show that the proposed algorithm maintains stability and safety throughout the learning process, achieves final performance comparable to traditional RL methods, and delivers higher training efficiency in a complex dynamic highway scenario in simulation. This suggests that the algorithm offers a viable solution for self-evolving systems in uncertain real-world environments, where traditional approaches may struggle.

NeurIPS Conference 2025 Conference Paper

EventMG: Efficient Multilevel Mamba-Graph Learning for Spatiotemporal Event Representation

  • Sheng Wu
  • Lin Jin
  • Hui Feng
  • Bo Hu

Event cameras offer unique advantages in scenarios involving high speed, low light, and high dynamic range, yet their asynchronous and sparse nature poses significant challenges to efficient spatiotemporal representation learning. Specifically, despite notable progress in the field, effectively modeling the full spatiotemporal context, selectively attending to salient dynamic regions, and robustly adapting to the variable density and dynamic nature of event data remain key challenges. Motivated by these challenges, this paper proposes EventMG, a lightweight, efficient, multilevel Mamba-Graph architecture designed for learning high-quality spatiotemporal event representations. EventMG employs a multilevel approach, jointly modeling information at the micro (single event) and macro (event cluster) levels to comprehensively capture the multi-scale characteristics of event data. At the micro-level, it focuses on spatiotemporal details, employing State Space Model (SSM) based Mamba, to precisely capture long-range dependencies among numerous event nodes. Concurrently, at the macro-level, Component Graphs are introduced to efficiently encode the local semantics and global topology of dense event regions. Furthermore, to better accommodate the dynamic and sparse characteristics of data, we propose the Spatiotemporal-aware Event Scanning Technology (SEST), integrating the Adaptive Perturbation Network (APN) and Multidirectional Scanning Module (MSM), which substantially enhances the model's ability to perceive and focus on key spatiotemporal patterns. By employing this novel collaborative paradigm, EventMG demonstrates the ability to effectively capture multi-level spatiotemporal characteristics of event data while maintaining a low parameter count and linear computational complexity, suggesting a promising direction for event representation learning.

AAAI Conference 2025 Conference Paper

Graph Mixture of Experts and Memory-augmented Routers for Multivariate Time Series Anomaly Detection

  • Xiaoyu Huang
  • Weidong Chen
  • Bo Hu
  • Zhendong Mao

Multivariate time series (MTS) anomaly detection is a critical task that involves identifying abnormal patterns or events in data that consist of multiple interrelated time series. In order to better model the complex interdependence between entities and the various inherent characteristics of each entity, the graph neural network (GNN) based methods are widely adopted by existing methods. In each layer of GNN, node features aggregate information from their neighboring nodes to update their information. In doing so, from shallow layer to deep layer in GNN, original individual node features continue to be weakened and more structural information, i.e., from short-distance neighborhood to long-distance neighborhood, continues to be enhanced. However, research to date has largely ignored the understanding of how hierarchical graph information is represented and their characteristics that can benefit anomaly detection. Existing methods simply leverage the output from the last layer of GNN for anomaly estimation while neglecting the essential information contained in the intermediate GNN layers. To address such limitations, in this paper, we propose a Graph Mixture of Experts (Graph-MoE) network for multivariate time series anomaly detection, which incorporates the mixture of experts (MoE) module to adaptively represent and integrate hierarchical multi-layer graph information into entity representations. It is worth noting that our Graph-MoE can be integrated into any GNN-based MTS anomaly detection method in a plug-and-play manner. In addition, the memory-augmented routers are proposed in this paper to capture the correlation temporal information in terms of the global historical features of MTS to adaptively weigh the obtained entity representations to achieve successful anomaly estimation. Extensive experiments on five challenging datasets prove the superiority of our approach and each proposed module.

ICLR Conference 2025 Conference Paper

TSC-Net: Prediction of Pedestrian Trajectories by Trajectory-Scene-Cell Classification

  • Bo Hu
  • Tat-Jen Cham

To predict future trajectories of pedestrians, scene is as important as the history trajectory since i) scene reflects the position of possible goals of the pedestrian ii) trajectories are affected by the semantic information of the scene. It requires the model to capture scene information and learn the relation between scenes and trajectories. However, existing methods either apply Convolutional Neural Networks (CNNs) to summarize the scene to a feature vector, which raises the feature misalignment issue, or convert trajectory to heatmaps to align with the scene map, which ignores the interactions among different pedestrians. In this work, we introduce the trajectory-scene-cell feature to represent both trajectories and scenes in one feature space. By decoupling the trajectory in temporal domain and the scene in spatial domain, trajectory feature and scene feature are re-organized in different types of cell feature, which well aligns trajectory and scene, and allows the framework to model both human-human and human-scene interactions. Moreover, the Trajectory-Scene-Cell Network (TSC-Net) with new trajectory prediction manner is proposed, where both goal and intermediate positions of the trajectory are predict by cell classification and offset regression. Comparative experiments show that TSC-Net achieves the SOTA performance on several datasets with most of the metrics. Especially for the goal estimation, TSC-Net is demonstrated better on predicting goals for trajectories with irregular speed.

NeurIPS Conference 2024 Conference Paper

EGSST: Event-based Graph Spatiotemporal Sensitive Transformer for Object Detection

  • Sheng Wu
  • Hang Sheng
  • Hui Feng
  • Bo Hu

Event cameras provide exceptionally high temporal resolution in dynamic vision systems due to their unique event-driven mechanism. However, the sparse and asynchronous nature of event data makes frame-based visual processing methods inappropriate. This study proposes a novel framework, Event-based Graph Spatiotemporal Sensitive Transformer (EGSST), for the exploitation of spatial and temporal properties of event data. Firstly, a well-designed graph structure is employed to model event data, which not only preserves the original temporal data but also captures spatial details. Furthermore, inspired by the phenomenon that human eyes pay more attention to objects that produce significant dynamic changes, we design a Spatiotemporal Sensitivity Module (SSM) and an adaptive Temporal Activation Controller (TAC). Through these two modules, our framework can mimic the response of the human eyes in dynamic environments by selectively activating the temporal attention mechanism based on the relative dynamics of event data, thereby effectively conserving computational resources. In addition, the integration of a lightweight, multi-scale Linear Vision Transformer (LViT) markedly enhances processing efficiency. Our research proposes a fully event-driven approach, effectively exploiting the temporal precision of event data and optimising the allocation of computational resources by intelligently distinguishing the dynamics within the event data. The framework provides a lightweight, fast, accurate, and fully event-based solution for object detection tasks in complex dynamic environments, demonstrating significant practicality and potential for application.

AAAI Conference 2024 Conference Paper

Gradual Residuals Alignment: A Dual-Stream Framework for GAN Inversion and Image Attribute Editing

  • Hao Li
  • Mengqi Huang
  • Lei Zhang
  • Bo Hu
  • Yi Liu
  • Zhendong Mao

GAN-based image attribute editing firstly leverages GAN Inversion to project real images into the latent space of GAN and then manipulates corresponding latent codes. Recent inversion methods mainly utilize additional high-bit features to improve image details preservation, as low-bit codes cannot faithfully reconstruct source images, leading to the loss of details. However, during editing, existing works fail to accurately complement the lost details and suffer from poor editability. The main reason is they inject all the lost details indiscriminately at one time, which inherently induces the position and quantity of details to overfit source images, resulting in inconsistent content and artifacts in edited images. This work argues that details should be gradually injected into both the reconstruction and editing process in a multi-stage coarse-to-fine manner for better detail preservation and high editability. Therefore, a novel dual-stream framework is proposed to accurately complement details at each stage. The Reconstruction Stream is employed to embed coarse-to-fine lost details into residual features and then adaptively add them to the GAN generator. In the Editing Stream, residual features are accurately aligned by our Selective Attention mechanism and then injected into the editing process in a multi-stage manner. Extensive experiments have shown the superiority of our framework in both reconstruction accuracy and editing quality compared with existing methods.

NeurIPS Conference 2024 Conference Paper

Learning Cortico-Muscular Dependence through Orthonormal Decomposition of Density Ratios

  • Shihan Ma
  • Bo Hu
  • Tianyu Jia
  • Alexander K. Clarke
  • Blanka Zicher
  • Arnault H. Caillet
  • Dario Farina
  • José C. Príncipe

The cortico-spinal neural pathway is fundamental for motor control and movement execution, and in humans it is typically studied using concurrent electroencephalography (EEG) and electromyography (EMG) recordings. However, current approaches for capturing high-level and contextual connectivity between these recordings have important limitations. Here, we present a novel application of statistical dependence estimators based on orthonormal decomposition of density ratios to model the relationship between cortical and muscle oscillations. Our method extends from traditional scalar-valued measures by learning eigenvalues, eigenfunctions, and projection spaces of density ratios from realizations of the signal, addressing the interpretability, scalability, and local temporal dependence of cortico-muscular connectivity. We experimentally demonstrate that eigenfunctions learned from cortico-muscular connectivity can accurately classify movements and subjects. Moreover, they reveal channel and temporal dependencies that confirm the activation of specific EEG channels during movement.

YNIMG Journal 2023 Journal Article

Functional to structural plasticity in unilateral sudden sensorineural hearing loss: neuroimaging evidence

  • Yu-Ting Li
  • Ke Bai
  • Gan-Ze Li
  • Bo Hu
  • Jia-Wei Chen
  • Yu-Xuan Shang
  • Ying Yu
  • Zhu-Hong Chen

A cortical plasticity after long-duration single side deafness (SSD) is advocated with neuroimaging evidence while little is known about the short-duration SSDs. In this case-cohort study, we recruited unilateral sudden sensorineural hearing loss (SSNHL) patients and age-, gender-matched health controls (HC), followed by comprehensive neuroimaging analyses. The primary outcome measures were temporal alterations of varied dynamic functional network connectivity (dFNC) states, neurovascular coupling (NVC) and brain region volume at different stages of SSNHL. The secondary outcome measures were pure-tone audiograms of SSNHL patients before and after treatment. A total of 38 SSNHL patients (21 [55%] male; mean [standard deviation] age, 45.05 [15.83] years) and 44 HC (28 [64%] male; mean [standard deviation] age, 43.55 [12.80] years) were enrolled. SSNHL patients were categorized into subgroups based on the time from disease onset to the initial magnetic resonance imaging scan: early- (n = 16; 1-6 days), intermediate- (n = 9; 7-13 days), and late- stage (n = 13; 14-30 days) groups. We first identified slow state transitions between varied dFNC states at early-stage SSNHL, then revealed the decreased NVC restricted to the auditory cortex at the intermediate- and late-stage SSNHL. Finally, a significantly decreased volume of the left medial superior frontal gyrus (SFGmed) was observed only in the late-stage SSNHL cohort. Furthermore, the volume of the left SFGmed is robustly correlated with both disease duration and patient prognosis. Our study offered neuroimaging evidence for the evolvement from functional to structural brain alterations of SSNHL patients with disease duration less than 1 month, which may explain, from a neuroimaging perspective, why early-stage SSNHL patients have better therapeutic responses and hearing recovery.

JBHI Journal 2022 Journal Article

Recursive Decomposition Network for Deformable Image Registration

  • Bo Hu
  • Shenglong Zhou
  • Zhiwei Xiong
  • Feng Wu

Deformation decomposition serves as a good solution for deformable image registration when the deformation is large. Current deformation decomposition methods can be categorized into cascade-based methods and pyramid-based methods. However, cascade-based methods suffer from heavy computational burdens and long inference time due to their structures of repeated subnetworks, while the effectiveness of pyramid-based methods is constrained by their limited numbers of resolution levels. In this paper, to address both the insufficient and inefficient decomposition problems in current deformation decomposition methods, we propose a recursive decomposition network (RDN) to offer a novel solution for deformable image registration. Stage-wise recursion can efficiently decompose a large deformation into different pyramid estimation stages without using repeated subnetworks like in cascade-based methods. Level-wise recursion can sufficiently decompose the deformation inside each resolution level instead of only one-time estimation like in pyramid-based methods. Extensive experiments and ablation studies on two representative datasets validate the effectiveness and efficiency of our proposed RDN.

NeurIPS Conference 2022 Conference Paper

Tenrec: A Large-scale Multipurpose Benchmark Dataset for Recommender Systems

  • Guanghu Yuan
  • Fajie Yuan
  • Yudong Li
  • Beibei Kong
  • Shujie Li
  • Lei Chen
  • Min Yang
  • Chenyun Yu

Existing benchmark datasets for recommender systems (RS) either are created at a small scale or involve very limited forms of user feedback. RS models evaluated on such datasets often lack practical values for large-scale real-world applications. In this paper, we describe Tenrec, a novel and publicly available data collection for RS that records various user feedback from four different recommendation scenarios. To be specific, Tenrec has the following five characteristics: (1) it is large-scale, containing around 5 million users and 140 million interactions; (2) it has not only positive user feedback, but also true negative feedback (vs. one-class recommendation); (3) it contains overlapped users and items across four different scenarios; (4) it contains various types of user positive feedback, in forms of clicking, liking, sharing, and following, etc; (5) it contains additional features beyond the user IDs and item IDs. We verify Tenrec on ten diverse recommendation tasks by running several classical baseline models per task. Tenrec has the potential to become a useful benchmark dataset for a majority of popular recommendation tasks. Our source codes and datasets will be included in supplementary materials.

JBHI Journal 2021 Journal Article

An $L_0$ Regularization Method for Imaging Genetics and Whole Genome Association Analysis on Alzheimer's Disease

  • Xiong Li
  • Yangkai Lin
  • Xu Meng
  • Yangping Qiu
  • Bo Hu

Although the neuroimaging measures build a bridge between genetic variants and disease phenotypes, an assessment of single nucleotide variants changes in brain structure and their clinically influence on the progression of Alzheimer's disease remain largely preliminary. Note that each variant has very weak correlation signal to neuroimaging measures or Alzheimer's disease phenotypes. Therefore, traditional sparse regression-based image genetics approaches confront with unresolvable features, relative high regression error or inapplicability of high-dimensional data. Adopting an $\text{L}_0$ regularization method, we significantly elevate the regression accuracy of imaging genetics compared with group-sparse multitask regression method. With further analysis on the simulation results, we conclude that multiple regression tasks model may be unsuitable for image genetics. In addition, we carried out a whole genome association analysis between genetic variants (about 388 million loci) and phenotypes (cognition normal, mild cognitive impairment and Alzheimer's disease) with using the $\text{L}_0$ regularization method. After annotating the effect of all variants by Ensembl Variant Effect Predictor (VEP), our method locates 33 missense variants which can explain 40% phenotype variance. Then, we mapped each missense variant to the nearest gene and carried out pathway enrichment analysis. The Notch signaling pathway and Apoptosis pathway have been reported to be related to the formation of Alzheimer's disease.

IJCAI Conference 2021 Conference Paper

Domain Generalization under Conditional and Label Shifts via Variational Bayesian Inference

  • Xiaofeng Liu
  • Bo Hu
  • Linghao Jin
  • Xu Han
  • Fangxu Xing
  • Jinsong Ouyang
  • Jun Lu
  • Georges El Fakhri

In this work, we propose a domain generalization (DG) approach to learn on several labeled source domains and transfer knowledge to a target domain that is inaccessible in training. Considering the inherent conditional and label shifts, we would expect the alignment of p(x|y) and p(y). However, the widely used domain invariant feature learning (IFL) methods relies on aligning the marginal concept shift w. r. t. p(x), which rests on an unrealistic assumption that p(y) is invariant across domains. We thereby propose a novel variational Bayesian inference framework to enforce the conditional distribution alignment w. r. t. p(x|y) via the prior distribution matching in a latent space, which also takes the marginal label shift w. r. t. p(y) into consideration with the posterior alignment. Extensive experiments on various benchmarks demonstrate that our framework is robust to the label shift and the cross-domain accuracy is significantly improved, thereby achieving superior performance over the conventional IFL counterparts.

NeurIPS Conference 2021 Conference Paper

On the Provable Generalization of Recurrent Neural Networks

  • Lifu Wang
  • Bo Shen
  • Bo Hu
  • Xing Cao

Recurrent Neural Network (RNN) is a fundamental structure in deep learning. Recently, some works study the training process of over-parameterized neural networks, and show that over-parameterized networks can learn functions in some notable concept classes with a provable generalization error bound. In this paper, we analyze the training and generalization for RNNs with random initialization, and provide the following improvements over recent works: (1) For a RNN with input sequence $x=(X_1, X_2, .. ., X_L)$, previous works study to learn functions that are summation of $f(\beta^T_lX_l)$ and require normalized conditions that $||X_l||\leq\epsilon$ with some very small $\epsilon$ depending on the complexity of $f$. In this paper, using detailed analysis about the neural tangent kernel matrix, we prove a generalization error bound to learn such functions without normalized conditions and show that some notable concept classes are learnable with the numbers of iterations and samples scaling almost-polynomially in the input length $L$. (2) Moreover, we prove a novel result to learn N-variables functions of input sequence with the form $f(\beta^T[X_{l_1}, .. ., X_{l_N}])$, which do not belong to the ``additive'' concept class, i, e. , the summation of function $f(X_l)$. And we show that when either $N$ or $l_0=\max(l_1, .. ,l_N)-\min(l_1, .. ,l_N)$ is small, $f(\beta^T[X_{l_1}, .. ., X_{l_N}])$ will be learnable with the number iterations and samples scaling almost-polynomially in the input length $L$.

AAAI Conference 2021 Conference Paper

Subtype-aware Unsupervised Domain Adaptation for Medical Diagnosis

  • Xiaofeng Liu
  • Xiongchang Liu
  • Bo Hu
  • Wenxuan Ji
  • Fangxu Xing
  • Jun Lu
  • Jane You
  • C.-C. Jay Kuo

Recent advances in unsupervised domain adaptation (UDA) show that transferable prototypical learning presents a powerful means for class conditional alignment, which encourages the closeness of cross-domain class centroids. However, the cross-domain inner-class compactness and the underlying fine-grained subtype structure remained largely underexplored. In this work, we propose to adaptively carry out the fine-grained subtype-aware alignment by explicitly enforcing the class-wise separation and subtype-wise compactness with intermediate pseudo labels. Our key insight is that the unlabeled subtypes of a class can be divergent to one another with different conditional and label shifts, while inheriting the local proximity within a subtype. The cases with or without the prior information on subtype numbers are investigated to discover the underlying subtype structure in an online fashion. The proposed subtype-aware dynamic UDA achieves promising results on a medical diagnosis task.

IJCAI Conference 2020 Conference Paper

Speeding up Very Fast Decision Tree with Low Computational Cost

  • Jian Sun
  • Hongyu Jia
  • Bo Hu
  • Xiao Huang
  • Hao Zhang
  • Hai Wan
  • Xibin Zhao

Very Fast Decision Tree (VFDT) is one of the most widely used online decision tree induction algorithms, and it provides high classification accuracy with theoretical guarantees. In VFDT, the split-attempt operation is essential for leaf-split. It is computation-intensive since it computes the heuristic measure of all attributes of a leaf. To reduce split-attempts, VFDT tries to split at constant intervals (for example, every 200 examples). However, this mechanism introduces split-delay for split can only happen at fixed intervals, which slows down the growth of VFDT and finally lowers accuracy. To address this problem, we first devise an online incremental algorithm that computes the heuristic measure of an attribute with a much lower computational cost. Then a subset of attributes is carefully selected to find a potential split timing using this algorithm. A split-attempt will be carried out once the timing is verified. By the whole process, computational cost and split-delay are lowered significantly. Comprehensive experiments are conducted using multiple synthetic and real datasets. Compared with state-of-the-art algorithms, our method reduces split-attempts by about 5 to 10 times on average with much lower split-delay, which makes our algorithm run faster and more accurate.

YNICL Journal 2019 Journal Article

Disturbed neurovascular coupling in type 2 diabetes mellitus patients: Evidence from a comprehensive fMRI analysis

  • Bo Hu
  • Lin-Feng Yan
  • Qian Sun
  • Ying Yu
  • Jin Zhang
  • Yu-Jie Dai
  • Yang Yang
  • Yu-Chuan Hu

BACKGROUND: Previous studies presumed that the disturbed neurovascular coupling to be a critical risk factor of cognitive impairments in type 2 diabetes mellitus (T2DM), but distinct clinical manifestations were lacked. Consequently, we decided to investigate the neurovascular coupling in T2DM patients by exploring the MRI relationship between neuronal activity and the corresponding cerebral blood perfusion. METHODS: Degree centrality (DC) map and amplitude of low-frequency fluctuation (ALFF) map were used to represent neuronal activity. Cerebral blood flow (CBF) map was used to represent cerebral blood perfusion. Correlation coefficients were calculated to reflect the relationship between neuronal activity and cerebral blood perfusion. RESULTS: At the whole gray matter level, the manifestation of neurovascular coupling was investigated by using 4 neurovascular biomarkers. We compared these biomarkers and found no significant changes. However, at the brain region level, neurovascular biomarkers in T2DM patients were significantly decreased in 10 brain regions. ALFF-CBF in left hippocampus and fractional ALFF-CBF in left amygdala were positively associated with the executive function, while ALFF-CBF in right fusiform gyrus was negatively related to the executive function. The disease severity was negatively related to the memory and executive function. The longer duration of T2DM was related to the milder depression, which suggests T2DM-related depression may not be a physiological condition but be a psychological condition. CONCLUSION: Correlations between neuronal activity and cerebral perfusion maps may be a method for detecting neurovascular coupling abnormalities, which could be used for diagnosis in the future. Trial registry number: This study has been registered in ClinicalTrials.gov (NCT02420470) on April 2, 2015 and published on July 29, 2015.

YNIMG Journal 2019 Journal Article

Neurovascular decoupling in type 2 diabetes mellitus without mild cognitive impairment: Potential biomarker for early cognitive impairment

  • Ying Yu
  • Lin-Feng Yan
  • Qian Sun
  • Bo Hu
  • Jin Zhang
  • Yang Yang
  • Yu-Jie Dai
  • Wu-Xun Cui

Type 2 diabetes mellitus (T2DM) is a significant risk factor for mild cognitive impairment (MCI) and the acceleration of MCI to dementia. The high glucose level induce disturbance of neurovascular (NV) coupling is suggested to be one potential mechanism, however, the neuroimaging evidence is still lacking. To assess the NV decoupling pattern in early diabetic status, 33 T2DM without MCI patients and 33 healthy control subjects were prospectively enrolled. Then, they underwent resting state functional MRI and arterial spin labeling imaging to explore the hub-based networks and to estimate the coupling of voxel-wise cerebral blood flow (CBF)-degree centrality (DC), CBF-mean amplitude of low-frequency fluctuation (mALFF) and CBF- mean regional homogeneity (mReHo). We further evaluated the relationship between NV coupling pattern and cognitive performance (false discovery rate corrected). T2DM without MCI patients displayed significant decrease in the absolute CBF-mALFF, CBF-mReHo coupling of CBFnetwork and in the CBF-DC coupling of DCnetwork. Besides, networks which involved CBF and DC hubs mainly located in the default mode network (DMN). Furthermore, less severe disease and better cognitive performance in T2DM patients were significantly correlated with higher coupling of CBF-DC, CBF-mALFF or CBF-mReHo, especially for the cognitive dimensions of general function and executive function. Thus, coupling of CBF-DC, CBF-mALFF and CBF-mReHo may serve as promising indicators to reflect NV coupling state and to explain the T2DM related early cognitive impairment.

JBHI Journal 2019 Journal Article

Unsupervised Learning for Cell-Level Visual Representation in Histopathology Images With Generative Adversarial Networks

  • Bo Hu
  • Ye Tang
  • Eric I-Chao Chang
  • Yubo Fan
  • Maode Lai
  • Yan Xu

The visual attributes of cells, such as the nuclear morphology and chromatin openness, are critical for histopathology image analysis. By learning cell-level visual representation, we can obtain a rich mix of features that are highly reusable for various tasks, such as celllevel classification, nuclei segmentation, and cell counting. In this paper, we propose a unified generative adversarial networks architecture with a new formulation of loss to perform robust cell-level visual representation learning in an unsupervised setting. Our model is not only label-free and easily trained but also capable of cell-level unsupervised classification with interpretable visualization, which achieves promising results in the unsupervised classification of bone marrow cellular components. Based on the proposed cell-level visual representation learning, we further develop a pipeline that exploits the varieties of cellular elements to perform histopathology image classification, the advantages of which are demonstrated on bone marrow datasets.

KER Journal 2011 Journal Article

A knowledge-rich distributed decision support framework: a case study for brain tumour diagnosis

  • David Dupplaw
  • Madalina Croitoru
  • Srinandan Dasmahapatra
  • Alex Gibb
  • Horacio González-Vélez
  • Miguel Lurgi
  • Bo Hu
  • Paul Lewis

Abstract The HealthAgents project aims to provide a decision support system for brain tumour diagnosis using a collaborative network of distributed agents. The goal is that through the aggregation of the small data sets available at individual hospitals, much better decision support classifiers can be created and made available to the hospitals taking part. In this paper, we describe the technicalities of the HealthAgents framework, in particular how the interoperability of the various agents is managed using semantic web technologies. On the broad scale the architecture is based around distributed data-mart agents that provide ontological access to hospitals’ underlying data that has been anonymized and processed from proprietary formats into a canonical format. Classifier producers have agents that gather the global data from participating hospitals such that classifiers can be created and deployed as agents. The design on a microscale has each agent built upon a generic-layered framework that provides the common agent program code, allowing rapid development of agents for the system. We believe that our framework provides a well-engineered, agent-based approach to data sharing in a medical context. It can provide a better basis on which to investigate the effectiveness of new classification techniques for brain tumour diagnosis.

KER Journal 2011 Journal Article

The design and implementation of a novel security model for HealthAgents

  • Liang Xiao
  • Srinandan Dasmahapatra
  • Paul Lewis
  • Bo Hu
  • Andrew Peet
  • Alex Gibb
  • David Dupplaw
  • Madalina Croitoru

Abstract In this paper, we analyze the special security requirements for software support in health care and the HealthAgents system in particular. Our security solution consists of a link-anonymized data scheme, a secure data transportation service, a secure data sharing and collection service, and a more advanced access control mechanism. The novel security service architecture, as part of the integrated system architecture, provides a secure health-care infrastructure for HealthAgents and can be easily adapted for other health-care applications.

KER Journal 2011 Journal Article

The HealthAgents ontology: knowledge representation in a distributed decision support system for brain tumours

  • Bo Hu
  • Madalina Croitoru
  • Roman Roset
  • David Dupplaw
  • Miguel Lurgi
  • Srinandan Dasmahapatra
  • Paul Lewis
  • Juan Martínez-Miranda

Abstract In this paper we present our experience of representing the knowledge behind HealthAgents (HA), a distributed decision support system for brain tumour diagnosis. Our initial motivation came from the distributed nature of the information involved in the system and has been enriched by clinicians’ requirements and data access restrictions. We present in detail the steps we have taken towards building our ontology starting from knowledge acquisition to data access and reasoning. We motivate our representational choices and show our results using domain examples used by clinical partners in HA.

AAAI Conference 2007 Conference Paper

On Capturing Semantics in Ontology Mapping

  • Bo Hu
  • Paul Lewis

Ontology mapping is a complex and necessary task for many Semantic Web (SW) applications. The perspective users are faced with a number of challenges including the difficulties of capturing semantics. In this paper we present a threedimensional ontology mapping model. This model reflects the engineering steps needed to materialise a versatile mapping system in order to meet the demands on semantic interoperability in the SW environment. We solidify the formalisation with specialised algorithms and we analyse their effectiveness and performance by way of benchmark tests.

v2026.09.13