Arrow Research search

Author name cluster

Lei Zhou

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

12 papers
2 author rows

Possible papers

12

EAAI Journal 2026 Journal Article

Efficient hybrid-strategy Q-learning based power enhancement for dynamic thermoelectric generation systems reconfiguration under heterogeneous temperature distribution

  • Bo Yang
  • Chuanyun Tang
  • Lei Zhou
  • Zijian Zhang
  • Yixuan Chen
  • Hai Lu
  • Hongbiao Li
  • Dengke Gao

The reconfiguration of thermoelectric generation (TEG) systems is a significant advancement in energy conversion efficiency and system optimization. This paper presents an advanced artificial intelligence (AI)-based algorithm, namely the efficient hybrid-strategy Q-learning (EHSQ), designed for the reconfiguration of TEG systems. The primary objective is to mitigate the adverse effects of heterogeneous temperature distribution (HTD) and fully exploit the power generation potential of TEG system. By optimizing the column output power (COP), EHSQ aims to enhance the overall output power and energy conversion efficiency. The COP is a fitting function to achieve this goal. Q-learning (QL) has been enhanced with innovative improvements to improve its global selection and optimization capabilities. Four conventional reinforcement learning (RL) algorithms for AI - dynamic programming Q learning (Dyna-Q), markov decision process (MDP), standard QL, and policy gradient (PG) - are used for comparison. Simulation tests conducted on SimuNPS modeling platform, employing AI techniques, reveal that the EHSQ algorithm markedly enhances the power generation efficiency of both symmetric (15 × 15) and asymmetric (20 × 15) TEG systems. The symmetric configuration achieves a maximum output power of 95. 4 W (W), percentage power increase of 4. 33 percent. while the asymmetric configuration yields 109. 4 W, percentage power increase of 3. 49 percent. Hardware-in-the-loop (HIL) experiments confirm the consistency with simulations, validating the effectiveness of EHSQ in optimizing TEG system performance. These results highlight the significant advantages of EHSQ in enhancing TEG system efficiency.

IROS Conference 2024 Conference Paper

3D Affordance Keypoint Detection for Robotic Manipulation

  • Zhiyang Liu
  • Ruiteng Zhao
  • Lei Zhou
  • Chengran Yuan
  • Yuwei Wu 0002
  • Sheng Guo
  • Zhengshen Zhang
  • Chenchen Liu

This paper presents a novel approach for affordance-informed robotic manipulation by introducing 3D keypoints to enhance the understanding of object parts’ functionality. The proposed approach provides direct information about what the potential use of objects is, as well as guidance on where and how a manipulator should engage, whereas conventional methods treat affordance detection as a semantic segmentation task, focusing solely on answering the what question. To address this gap, we propose a Fusion-based Affordance Keypoint Network (FAKP-Net) by introducing 3D keypoint quadruplet that harnesses the synergistic potential of RGB and Depth image to provide information on execution position, direction, and extent. Benchmark testing demonstrates that FAKP-Net outperforms existing models by significant margins in affordance segmentation task and keypoint detection task. Real-world experiments also showcase the reliability of our method in accomplishing manipulation tasks with previously unseen objects. Our source code and video demo will be public.

IROS Conference 2024 Conference Paper

A Robust and Efficient Robotic Packing Pipeline with Dissipativity- Based Adaptive Impedance-Force Control

  • Zhenning Zhou
  • Lei Zhou
  • Shengxin Sun
  • Marcelo H. Ang

For humans, dense bin packing heavily relies on force perception. However, current robotic packing studies only focus on the visual input or adopt auxiliary push-to-place actions to eliminate gaps, suffering from high time expenditure and poor robustness. To address such limitations, we first introduce a novel external force estimation method based on the generalized momentum observer, which can avoid the influence of joint acceleration noises and achieve real-time high-precision monitoring. Second, to obtain compliant interaction and fine robustness, an adaptive variable impedance policy is developed to track dynamic motion and desired force, and compensate for uncertainties. Meanwhile, we perform dissipativity analysis and a virtual energy supply function is augmented to the system for optimization, providing a solid foundation for stability. Third, we propose an efficient packing methodology with three sub- tasks by considering the distinct interaction and constraint states in different areas. Our packing strategies eliminate the need for subsequent auxiliary actions and are proven to enhance efficiency. We perform quantitative evaluations to verify our external force estimation method, conduct comparison studies with current packing methods, and investigate the contribution of our dissipativity-based adaptive controller. The superior results not only prove the robustness and efficiency of our pipeline, but also pave the way for practical applications of packing.

IJCAI Conference 2024 Conference Paper

CLIP-FSAC: Boosting CLIP for Few-Shot Anomaly Classification with Synthetic Anomalies

  • Zuo Zuo
  • Yao Wu
  • Baoqiang Li
  • Jiahao Dong
  • You Zhou
  • Lei Zhou
  • Yanyun Qu
  • Zongze Wu

Few-shot anomaly classification (FSAC) is a vital task in manufacturing industry. Recent methods focus on utilizing CLIP in zero/few normal shot anomaly detection instead of custom models. However, there is a lack of specific text prompts in anomaly classification and most of them ignore the modality gap between image and text. Meanwhile, there is distribution discrepancy between the pre-trained and the target data. To provide a remedy, in this paper, we propose a method to boost CLIP for few-normal-shot anomaly classification, dubbed CLIP-FSAC, which contains two-stage of training and alternating fine-tuning with two modality-specific adapters. Specifically, in the first stage, we train image adapter with text representation output from text encoder and introduce an image-to-text tuning to enhance multi-modal interaction and facilitate a better language-compatible visual representation. In the second stage, we freeze the image adapter to train the text adapter. Both of them are constrained by fusion-text contrastive loss. Comprehensive experiment results are provided for evaluating our method in few-normal-shot anomaly classification, which outperforms the state-of-the-art method by 12. 2%, 10. 9%, 10. 4% AUROC on VisA for 1, 2, and 4-shot settings.

ICRA Conference 2024 Conference Paper

You Only Scan Once: A Dynamic Scene Reconstruction Pipeline for 6-DoF Robotic Grasping of Novel Objects

  • Lei Zhou
  • Haozhe Wang 0003
  • Zhengshen Zhang
  • Zhiyang Liu
  • Francis E. H. Tay
  • Marcelo H. Ang

In the realm of robotic grasping, achieving accurate and reliable interactions with the environment is a pivotal challenge. Traditional methods of grasp planning methods utilizing partial point clouds derived from depth image often suffer from reduced scene understanding due to occlusion, ultimately impeding their grasping accuracy. Furthermore, scene reconstruction methods have primarily relied upon static techniques, which are susceptible to environment change during manipulation process limits their efficacy in real-time grasping tasks. To address these limitations, this paper introduces a novel two-stage pipeline for dynamic scene reconstruction. In the first stage, our approach takes scene scanning as input to register each target object with mesh reconstruction and novel object pose tracking. In the second stage, pose tracking is still performed to provide object poses in real-time, enabling our approach to transform the reconstructed object point clouds back into the scene. Unlike conventional methodologies, which rely on static scene snapshots, our method continuously captures the evolving scene geometry, resulting in a comprehensive and up-to-date point cloud representation. By circumventing the constraints posed by occlusion, our method enhances the overall grasp planning process and empowers state-of-the-art 6-DoF robotic grasping algorithms to exhibit markedly improved accuracy.

IROS Conference 2023 Conference Paper

DR-Pose: A Two-Stage Deformation-and-Registration Pipeline for Category-Level 6D Object Pose Estimation

  • Lei Zhou
  • Zhiyang Liu
  • Runze Gan
  • Haozhe Wang 0003
  • Marcelo H. Ang

Category-level object pose estimation involves estimating the 6D pose and the 3D metric size of objects from predetermined categories. While recent approaches take categorical shape prior information as reference to improve pose estimation accuracy, the single-stage network design and training manner lead to sub-optimal performance since there are two distinct tasks in the pipeline. In this paper, the advantage of two-stage pipeline over single-stage design is discussed. To this end, we propose a two-stage deformation-and-registration pipeline called DR-Pose, which consists of completion-aided deformation stage and scaled registration stage. The first stage uses a point cloud completion method to generate unseen parts of target object, guiding subsequent deformation on the shape prior. In the second stage, a novel registration network is designed to extract pose-sensitive features and predict the representation of object partial point cloud in canonical space based on the deformation results from the first stage. DR-Pose produces superior results to the state-of-the-art shape prior-based methods on both CAMERA25 and REAL275 benchmarks. Codes are available at https://github.com/Zray26/DR-Pose.git.

IJCAI Conference 2022 Conference Paper

Learning Prototype via Placeholder for Zero-shot Recognition

  • Zaiquan Yang
  • Yang Liu
  • Wenjia Xu
  • Chong Huang
  • Lei Zhou
  • Chao Tong

Zero-shot learning (ZSL) aims to recognize unseen classes by exploiting semantic descriptions shared between seen classes and unseen classes. Current methods show that it is effective to learn visual-semantic alignment by projecting semantic embeddings into the visual space as class prototypes. However, such a projection function is only concerned with seen classes. When applied to unseen classes, the prototypes often perform suboptimally due to domain shift. In this paper, we propose to learn prototypes via placeholders, termed LPL, to eliminate the domain shift between seen and unseen classes. Specifically, we combine seen classes to hallucinate new classes which play as placeholders of the unseen classes in the visual and semantic space. Placed between seen classes, the placeholders encourage prototypes of seen classes to be highly dispersed. And more space is spared for the insertion of well-separated unseen ones. Empirically, well-separated prototypes help counteract visual-semantic misalignment caused by domain shift. Furthermore, we exploit a novel semantic-oriented fine-tuning method to guarantee the semantic reliability of placeholders. Extensive experiments on five benchmark datasets demonstrate the significant performance gain of LPL over the state-of-the-art methods.

IJCAI Conference 2019 Conference Paper

Latent Distribution Preserving Deep Subspace Clustering

  • Lei Zhou
  • Xiao Bai
  • Dong Wang
  • Xianglong Liu
  • Jun Zhou
  • Edwin Hancock

Subspace clustering is a useful technique for many computer vision applications in which the intrinsic dimension of high-dimensional data is smaller than the ambient dimension. Traditional subspace clustering methods often rely on the self-expressiveness property, which has proven effective for linear subspace clustering. However, they perform unsatisfactorily on real data with complex nonlinear subspaces. More recently, deep autoencoder based subspace clustering methods have achieved success owning to the more powerful representation extracted by the autoencoder network. Unfortunately, these methods only considering the reconstruction of original input data can hardly guarantee the latent representation for the data distributed in subspaces, which inevitably limits the performance in practice. In this paper, we propose a novel deep subspace clustering method based on a latent distribution-preserving autoencoder, which introduces a distribution consistency loss to guide the learning of distribution-preserving latent representation, and consequently enables strong capacity of characterizing the real-world data for subspace clustering. Experimental results on several public databases show that our method achieves significant improvement compared with the state-of-the-art subspace clustering methods.

AAAI Conference 2019 Conference Paper

Learning Fully Dense Neural Networks for Image Semantic Segmentation

  • Mingmin Zhen
  • Jinglu Wang
  • Lei Zhou
  • Tian Fang
  • Long Quan

Semantic segmentation is pixel-wise classification which retains critical spatial information. The “feature map reuse” has been commonly adopted in CNN based approaches to take advantage of feature maps in the early layers for the later spatial reconstruction. Along this direction, we go a step further by proposing a fully dense neural network with an encoderdecoder structure that we abbreviate as FDNet. For each stage in the decoder module, feature maps of all the previous blocks are adaptively aggregated to feedforward as input. On the one hand, it reconstructs the spatial boundaries accurately. On the other hand, it learns more efficiently with the more efficient gradient backpropagation. In addition, we propose the boundary-aware loss function to focus more attention on the pixels near the boundary, which boosts the “hard examples” labeling. We have demonstrated the best performance of the FDNet on the two benchmark datasets: PASCAL VOC 2012, NYUDv2 over previous works when not considering training on other datasets.

YNIMG Journal 2008 Journal Article

BOLD study of stimulation-induced neural activity and resting-state connectivity in medetomidine-sedated rat

  • Fuqiang Zhao
  • Tiejun Zhao
  • Lei Zhou
  • Qiulin Wu
  • Xiaoping Hu

Functional magnetic resonance imaging (fMRI) in anesthetized-animals is critical in studying the mechanisms of fMRI and investigating animal models of various diseases. Medetomidine was recently introduced for independent anesthesia for longitudinal (survival) fMRI studies in rats. Since stimulation-induced fMRI signal is anesthesia-dependent and its characteristics in rats under medetomidine are not fully elucidated, the blood oxygenation level dependent (BOLD) fMRI response to electrical forepaw stimulation under medetomidine was systematically investigated at 9. 4 T. Robust activations in contralateral primary somatosensory cortex (SI) and thalamus were observed and peaked at the stimulus frequency of 9 Hz. The response in SI saturates at the stimulus strength of 4 mA while that in thalamus monotonically increases. In addition to fMRI data acquired with the forepaw stimulation, data were also acquired during the resting-state to investigate the synchronization of low frequency fluctuations (LFF) in the BOLD signal (<0. 08 Hz) in different brain regions. LFF during resting-state have been observed to be synchronized between functionally related brain regions in human subjects while its origin is not fully understood. LFF have not been extensively studied or widely reported in anesthetized-animals. In our data, synchronized LFF of BOLD signals are found in clustered, bilaterally symmetric regions, including SI and caudate–putamen and the magnitude of the LFF is ∼1. 5%, comparable to the stimulation-induced BOLD signals. Similar to resting-state data reported in human subjects, LFF in rats under medetomidine likely reflect functional connectivity of these brain regions.

YNIMG Journal 2007 Journal Article

Quantitative basal CBF and CBF fMRI of rhesus monkeys using three-coil continuous arterial spin labeling

  • Xiaodong Zhang
  • Tsukasa Nagaoka
  • Edward J. Auerbach
  • Robbie Champion
  • Lei Zhou
  • Xiaoping Hu
  • Timothy Q. Duong

A three-coil continuous arterial-spin-labeling technique with a separate neck labeling coil was implemented on a Siemens 3T Trio for quantitative cerebral blood flow (CBF) and CBF fMRI measurements in non-human primates (rhesus monkeys). The optimal labeling power was 2 W, labeling efficiency was 92±2%, and optimal post-labeling delay was 0. 8 s. Gray matter (GM) and white matter (WM) were segmented based on T 1 maps. Quantitative CBF were obtained in 3 min with 1. 5-mm isotropic resolution. Whole-brain average ΔS/S was 1. 0–1. 5%. GM CBF was 104±3 ml/100 g/min (n =6, SD) and WM CBF was 45±6 ml/100 g/min in isoflurane-anesthetized rhesus monkeys, with the CBF GM/WM ratio of 2. 3±0. 2. Combined CBF and BOLD (blood-oxygenation-level-dependent) fMRI associated with hypercapnia and hyperoxia were made with 8-s temporal resolution. CBF fMRI responses to 5% CO2 were 59±10% (GM) and 37±4% (WM); BOLD fMRI responses were 2. 0±0. 4% (GM) and 1. 2±0. 4% (WM). CBF fMRI responses to 100% O2 were −9. 4±2% (GM) and −3. 9±2. 6% (WM); BOLD responses were 2. 4±0. 7% (GM) and 0. 8±0. 2% (WM). The use of a separate neck coil for spin labeling significantly increased CBF signal-to-noise ratio and the use of small receive-only surface coil significantly increased signal-to-noise ratio and spatial resolution. This study sets the stage for quantitative perfusion imaging and CBF fMRI for neurological diseases in anesthetized and awake monkeys.

v2026.09.13