Arrow Research search

Author name cluster

Ke Hu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

9 papers
1 author row

Possible papers

9

AAAI Conference 2026 Conference Paper

FedTopo: Topology-Informed Representation Alignment in Federated Learning Under Non-I.I.D. Conditions

  • Ke Hu
  • Liyao Xiang
  • Peng Tang
  • Weidong Qiu

Current federated-learning models deteriorate under heterogeneous (non-I.I.D.) client data, as their feature representations diverge and pixel- or patch-level objectives fail to capture the global topology which is essential for high-dimensional visual tasks. We propose FedTopo, a framework that integrates Topological-Guided Block Screening (TGBS) and Topological Embedding (TE) to leverage topological information, yielding coherently aligned cross-client representations by Topological Alignment Loss (TAL). First, Topology-Guided Block Screening (TGBS) automatically selects the most topology-informative block, i.e., the one with maximal topological separability, whose persistence-based signatures best distinguish within- versus between-class pairs, ensuring that subsequent analysis focuses on topology-rich features. Next, this block yields a compact Topological Embedding, which quantifies the topological information for each client. Finally, a Topological Alignment Loss (TAL) guides clients to maintain topological consistency with the global model during optimization, reducing representation drift across rounds. Experiments on Fashion-MNIST, CIFAR-10, and CIFAR-100 under four non-I.I.D. partitions show that FedTopo accelerates convergence and improves accuracy over strong baselines.

EAAI Journal 2025 Journal Article

Fast shallow multi-subnet detector for real-time object detection

  • Yuan Li
  • Mengdie Song
  • Ke Hu
  • Song Chen
  • Yi Kang

Real-time object detection algorithms, underpinned by Deep Neural Networks (DNNs), are extensively applied in fields like autonomous driving and security surveillance. However, current algorithms face issues of low hardware resource utilization and high synchronization delays between network layers when deployed on DNN hardware accelerators, adversely affecting overall performance and efficiency. To address these issues, we have proposed an innovative single-stage object detection framework, the Shallow Multi-Subnet Detector (SMS-Det). SMS-Det adopts a multi-parallel-shallow-subnet architecture, which reduces inter-layer synchronization latency by decreasing network depth. Furthermore, it fully utilizes DNN hardware accelerators by executing convolution operations in parallel, preventing resource underutilization and maximizing throughput. The proposed network is comprised of multiple parallel shallow subnets, each of which processes feature maps of different scales. The Feature Fusion Layer (FFL) ensures seamless information exchange across subnets, significantly improving the detection of small and occluded objects. Finally, we introduce the multi-scale channel attention projections to enhance the feature mapping between the teacher model and the student model in the training process. Experimental results on the Microsoft Common Objects in Context (MS COCO) dataset demonstrate that our model achieves a state-of-the-art mean Average Precision (mAP) of 42. 6%, surpassing You Only Look Once Version 5 Small (YOLOv5-S 37. 4%) with only 19. 4 Giga Floating Point Operations (GFLOPs) and 11. 0 million parameters. Our model obtains 156 Frames Per Second (FPS), achieving a real-time inference acceleration of 51. 4% compared to YOLOv5-S (103 FPS).

NeurIPS Conference 2025 Conference Paper

GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning

  • Shutong Ding
  • Ke Hu
  • Shan Zhong
  • Haoyang Luo
  • Weinan Zhang
  • Jingya Wang
  • Jun Wang
  • Ye Shi

Recent advances in reinforcement learning (RL) have demonstrated the powerful exploration capabilities and multimodality of generative diffusion-based policies. While substantial progress has been made in offline RL and off-policy RL settings, integrating diffusion policies into on-policy frameworks like PPO remains underexplored. This gap is particularly significant given the widespread use of large-scale parallel GPU-accelerated simulators, such as IsaacLab, which are optimized for on-policy RL algorithms and enable rapid training of complex robotic tasks. A key challenge lies in computing state-action log-likelihoods under diffusion policies, which is straightforward for Gaussian policies but intractable for flow-based models due to irreversible forward-reverse processes and discretization errors (e. g. , Euler-Maruyama approximations). To bridge this gap, we propose GenPO, a generative policy optimization framework that leverages exact diffusion inversion to construct invertible action mappings. GenPO introduces a novel doubled dummy action mechanism that enables invertibility via alternating updates, resolving log-likelihood computation barriers. Furthermore, we also use the action log-likelihood for unbiased entropy and KL divergence estimation, enabling KL-adaptive learning rates and entropy regularization in on-policy updates. Extensive experiments on eight IsaacLab benchmarks, including legged locomotion (Ant, Humanoid, Anymal-D, Unitree H1, Go2), dexterous manipulation (Shadow Hand), aerial control (Quadcopter), and robotic arm tasks (Franka), demonstrate GenPO’s superiority over existing RL baselines. Notably, GenPO is the first method to successfully integrate diffusion policies into on-policy RL, unlocking their potential for large-scale parallelized training and real-world robotic deployment.

AAAI Conference 2025 Conference Paper

KD-MSLRT: Lightweight Sign Language Recognition Model Based on Mediapipe and 3D to 1D Knowledge Distillation

  • Yulong Li
  • Bolin Ren
  • Ke Hu
  • Changyuan Liu
  • Zhengyong Jiang
  • Kang Dang
  • Jionglong Su

Artificial intelligence has achieved notable results in sign language recognition and translation. However, relatively few efforts have been made to significantly improve the quality of life for the 72 million hearing-impaired people worldwide. Sign language translation models, relying on video inputs, involves with large parameter sizes, making it time-consuming and computationally intensive to be deployed. This directly contributes to the scarcity of human-centered technology in this field. Additionally, the lack of datasets in sign language translation hampers research progress in this area. To address these, we first propose a cross-modal multi-knowledge distillation technique from 3D to 1D and a novel end-to-end pre-training text correction framework. Compared to other pre-trained models, our framework achieves significant advancements in correcting text output errors. Our model achieves a decrease in Word Error Rate (WER) of at least 1.4% on PHOENIX14 and PHOENIX14T datasets compared to the state-of-the-art CorrNet. Additionally, the TensorFlow Lite (TFLite) quantized model size is reduced to 12.93 MB, making it the smallest, fastest, and most accurate model to date. We have also collected and released extensive Chinese sign language datasets, and developed a specialized training vocabulary. To address the lack of research on data augmentation for landmark data, we have designed comparative experiments on various augmentation methods. Moreover, we performed a simulated deployment and prediction of our model on Intel platform CPUs and assessed the feasibility of deploying the model on other platforms.

EAAI Journal 2024 Journal Article

A novel multi-step ahead prediction method for landslide displacement based on autoregressive integrated moving average and intelligent algorithm

  • Peng Shao
  • Hong Wang
  • Guangyu Long
  • Jianxing Liao
  • Fei Gan
  • Bin Xu
  • Ke Hu
  • Yuhang Teng

Accurate landslide displacement prediction is crucial for prevention and early warning. In this paper, we proposed a novel hybrid multi-step-ahead prediction model (ARIMA-IM) that combines autoregressive integrated moving average (ARIMA) and intelligent models (IMs) for landslide displacement prediction. This model integrates the linear prediction strengths of ARIMA, the signal decomposition capabilities of variational mode decomposition (VMD) and empirical wavelet transform (EWT), and the nonlinear change-capturing ability of IMs. The proposed model cannot only effectively capture abrupt and long-term trends in landslide displacement but also enable high-precision multi-step-ahead prediction. To validate the effectiveness of the proposed model, we predicted the displacement of the Bazimen landslide in the Three Gorges Reservoir by using four IMs, namely long short-term memory (LSTM), bidirectional LSTM (BiLSTM), gate recurrent unit (GRU), and artificial neural network (ANN), in both continuous and jump strategies for multi-step-ahead prediction. Results showed that ARIMA-IM, especially the hybrid model combining ARIMA and deep learning, achieved high accuracy in 1–5-step-ahead prediction. Continuous multi-step-ahead prediction exhibited higher accuracy and provided more prediction information for decision-making. Compared to other models, ARIMA-IM not only achieved multi-step-ahead prediction but also exhibited comparable or better prediction accuracy, which is of high practical significance for landslide disaster early warning.

AAAI Conference 2024 Conference Paper

DALDet: Depth-Aware Learning Based Object Detection for Autonomous Driving

  • Ke Hu
  • Tongbo Cao
  • Yuan Li
  • Song Chen
  • Yi Kang

3D object detection achieves good detection performance in autonomous driving. However, it requires substantial computational resources, which prevents its practical application. 2D object detection has less computational burden but lacks spatial and geometric information embedded in depth. Therefore, we present DALDet, an efficient depth-aware learning based 2D detector, achieving high-performance object detection for autonomous driving. We design an efficient one-stage detection framework and seamlessly integrate depth cues into convolutional neural network by introducing depth-aware convolution and depth-aware average pooling, which effectively improve the detector's ability to perceive 3D space. Moreover, we propose a depth-guided loss function for training DALDet, which effectively improves the localization ability of the detector. Due to the use of depth map, DALDet can also output the distance of the object, which is of great importance for driving applications such as obstacle avoidance. Extensive experiments demonstrate the superiority and efficiency of DALDet. In particular, our DALDet ranks 1st on both KITTI Car and Cyclist 2D detection test leaderboards among all 2D detectors with high efficiency as well as yielding competitive performance among many leading 3D detectors. Code will be available at https://github.com/hukefy/DALDet.

NeurIPS Conference 2024 Conference Paper

Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization

  • Shutong Ding
  • Ke Hu
  • Zhenhao Zhang
  • Kan Ren
  • Weinan Zhang
  • Jingyi Yu
  • Jingya Wang
  • Ye Shi

Diffusion models have garnered widespread attention in Reinforcement Learning (RL) for their powerful expressiveness and multimodality. It has been verified that utilizing diffusion policies can significantly improve the performance of RL algorithms in continuous control tasks by overcoming the limitations of unimodal policies, such as Gaussian policies. Furthermore, the multimodality of diffusion policies also shows the potential of providing the agent with enhanced exploration capabilities. However, existing works mainly focus on applying diffusion policies in offline RL, while their incorporation into online RL has been less investigated. The diffusion model's training objective, known as the variational lower bound, cannot be applied directly in online RL due to the unavailability of 'good' samples (actions). To harmonize the diffusion model with online RL, we propose a novel model-free diffusion-based online RL algorithm named Q-weighted Variational Policy Optimization (QVPO). Specifically, we introduce the Q-weighted variational loss and its approximate implementation in practice. Notably, this loss is shown to be a tight lower bound of the policy objective. To further enhance the exploration capability of the diffusion policy, we design a special entropy regularization term. Unlike Gaussian policies, the log-likelihood in diffusion policies is inaccessible; thus this entropy term is nontrivial. Moreover, to reduce the large variance of diffusion policies, we also develop an efficient behavior policy through action selection. This can further improve its sample efficiency during online interaction. Consequently, the QVPO algorithm leverages the exploration capabilities and multimodality of diffusion policies, preventing the RL agent from converging to a sub-optimal policy. To verify the effectiveness of QVPO, we conduct comprehensive experiments on MuJoCo continuous control benchmarks. The final results demonstrate that QVPO achieves state-of-the-art performance in terms of both cumulative reward and sample efficiency.

IJCAI Conference 2024 Conference Paper

Feature Norm Regularized Federated Learning: Utilizing Data Disparities for Model Performance Gains

  • Ke Hu
  • Liyao Xiang
  • Peng Tang
  • Weidong Qiu

Federated learning (FL) is a machine learning paradigm that aggregates knowledge and utilizes computational power from multiple participants to train a global model. However, a commonplace challenge—non-independent and identically distributed (non-i. i. d. ) data across participants—can lead to significant divergence in model updates, thus diminishing training efficacy. In this paper, we propose the Feature Norm Regularized Federated Learning (FNR-FL) algorithm to tackle the non-i. i. d challenge. FNR-FL incorporates class average feature norms into the loss function by a straightforward yet effective regularization strategy. The core idea of FNR-FL is to penalize the deviations in the update directions of local models caused by the non-i. i. d data. Theoretically, we provide convergence guarantees for FNR-FL when training under non-i. i. d scenarios. Practically, our comprehensive experimental evaluations demonstrate that FNR-FL significantly outperforms existing FL algorithms in terms of test accuracy, and maintains a competitive convergence rate with lower communication overhead and shorter duration. Compared to FedAvg, FNR-FL exhibits a 66. 24% improvement in accuracy and an 11. 40% reduction in training time, underscoring its enhanced effectiveness and efficiency. The code is available on GitHub at: https: //github. com/LonelyMoonDesert/FNR-FL.

YNICL Journal 2021 Journal Article

Multisite schizophrenia classification by integrating structural magnetic resonance imaging data with polygenic risk score

  • Ke Hu
  • Meng Wang
  • Yong Liu
  • Hao Yan
  • Ming Song
  • Jun Chen
  • Yunchun Chen
  • Huaning Wang

Previous brain structural magnetic resonance imaging studies reported that patients with schizophrenia have brain structural abnormalities, which have been used to discriminate schizophrenia patients from normal controls. However, most existing studies identified schizophrenia patients at a single site, and the genetic features closely associated with highly heritable schizophrenia were not considered. In this study, we performed standardized feature extraction on brain structural magnetic resonance images and on genetic data to separate schizophrenia patients from normal controls. A total of 1010 participants, 508 schizophrenia patients and 502 normal controls, were recruited from 8 independent sites across China. Classification experiments were carried out using different machine learning methods and input features. We tested a support vector machine, logistic regression, and an ensemble learning strategy using 3 feature sets of interest: (1) imaging features: gray matter volume, (2) genetic features: polygenic risk scores, and (3) a fusion of imaging features and genetic features. The performance was assessed by leave-one-site-out cross-validation. Finally, some important brain and genetic features were identified. We found that the models with both imaging and genetic features as input performed better than models with either alone. The average accuracy of the classification models with the best performance in the cross-validation was 71.6%. The genetic feature that measured the cumulative risk of the genetic variants most associated with schizophrenia contributed the most to the classification. Our work took the first step toward considering both structural brain alterations and genome-wide genetic factors in a large-scale multisite schizophrenia classification. Our findings may provide insight into the underlying pathophysiology and risk mechanisms of schizophrenia.

v2026.09.13