Arrow Research search

Author name cluster

Le Zhang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

22 papers
2 author rows

Possible papers

22

AAAI Conference 2026 Conference Paper

Enhancing Generalization of Depth Estimation Foundation Model via Weakly-Supervised Adaptation with Regularization

  • Yan Huang
  • Yongyi Su
  • Xin Lin
  • Le Zhang
  • Xun Xu

The emergence of foundation models has substantially advanced zero-shot generalization in monocular depth estimation (MDE), as exemplified by the Depth Anything series. However, given access to some data from downstream tasks, a natural question arises: can the performance of these models be further improved? To this end, we propose WeSTAR, a parameter-efficient framework that performs \textbf{We}akly supervised \textbf{S}elf-\textbf{T}raining \textbf{A}daptation with \textbf{R}egularization, designed to enhance the robustness of MDE foundation models in unseen and diverse domains. We first adopt a dense self-training objective as the primary source of structural self-supervision. To further improve robustness, we introduce semantically-aware hierarchical normalization, which exploits instance-level segmentation maps to perform more stable and multi-scale structural normalization. Beyond dense supervision, we introduce a cost-efficient weak supervision in the form of pairwise ordinal depth annotations to further guide the adaptation process, which enforces informative ordinal constraints to mitigate local topological errors. Finally, a weight regularization loss is employed to anchor the LoRA updates, ensuring training stability and preserving the model's generalizable knowledge. Extensive experiments on both realistic and corrupted out-of-distribution datasets under diverse and challenging scenarios demonstrate that WeSTAR consistently improves generalization and achieves state-of-the-art performance across a wide range of benchmarks.

AAAI Conference 2026 Conference Paper

Graph Meets Deep Unfolding: An Interpretable Mutual-benefit Multi-view Learning Network

  • Renjie Lin
  • Hongzhi He
  • Yilin Wu
  • Shide Du
  • Le Zhang

Significant efforts have been focused on enhancing the utilization of multiple node features and topological structures in multi-view graph learning through explicit model-driven and implicit deep learning-based methodologies. The former excels in embedding prior knowledge, thereby offering theoretical interpretability but is limited in application flexibility due to manual parameter selection. In contrast, the latter leverages automatic differentiation, providing greater flexibility but lacking theoretical interpretability due to their opaque nature. Motivated by these observations, we propose an interpretable deep unfolding network for mutual-benefit multi-view graph learning, aiming to combine the strengths of both approaches. Specifically, we employ the Alternating Direction Method of Multipliers (ADMM) to solve a multi-view graph learning model with sparse and low-rank constraints. This solution is then integrated into deep unfolding networks to enhance interpretability. Furthermore, we convert optimization conditions into implicit losses and utilize automatic differentiation to update parameters, reducing the need for manual tuning and increasing flexibility. This integration optimizes multi-view learning for a graph representation that balances interpretability and flexibility. Empirical evaluations on six diverse datasets demonstrate the effectiveness and superiority of the proposed method over state-of-the-art approaches.

AAAI Conference 2025 Conference Paper

CharacterBench: Benchmarking Character Customization of Large Language Models

  • Jinfeng Zhou
  • Yongkang Huang
  • Bosi Wen
  • Guanqun Bi
  • Yuxuan Chen
  • Pei Ke
  • Zhuang Chen
  • Xiyao Xiao

Character-based dialogue (aka role-playing) enables users to freely customize characters for interaction, which often relies on LLMs, raising the need to evaluate LLMs’ character customization capability. However, existing benchmarks fail to ensure a robust evaluation as they often only involve a single character category or evaluate limited dimensions. Moreover, the sparsity of character features in responses makes feature-focused generative evaluation both ineffective and inefficient. To address these issues, we propose CharacterBench, the largest bilingual generative benchmark, with 22,859 human-annotated samples covering 3,956 characters from 25 detailed character categories. We define 11 dimensions of 6 aspects, classified as sparse and dense dimensions based on whether character features evaluated by specific dimensions manifest in each response. We enable effective and efficient evaluation by crafting tailored queries for each dimension to induce characters’ responses related to specific dimensions. Further, we develop CharacterJudge model for cost-effective and stable evaluations. Experiments show its superiority over SOTA automatic judges (e.g., GPT-4) and our benchmark’s potential to optimize LLMs’ character customization.

YNIMG Journal 2025 Journal Article

DeepNuParc: A novel deep clustering framework for fine-scale parcellation of brain nuclei using diffusion MRI tractography

  • Haolin He
  • Ce Zhu
  • Le Zhang
  • Yipeng Liu
  • Xiao Xu
  • Yuqian Chen
  • Leo Zekelman
  • Jarrett Rushmore

Brain nuclei are clusters of anatomically distinct neurons that serve as important hubs for processing and relaying information in various neural circuits. Fine-scale parcellation of the brain nuclei is vital for a comprehensive understanding of their anatomico-functional correlations. Diffusion MRI tractography is an advanced imaging technique that can estimate the brain's white matter structural connectivity to potentially reveal the topography of the nuclei of interest for studying their subdivisions. In this work, we present a deep clustering pipeline, namely DeepNuParc, to perform automated, fine-scale parcellation of brain nuclei using diffusion MRI tractography. First, we incorporate a newly proposed deep learning approach to enable accurate segmentation of the nuclei of interest directly on the dMRI data. Next, we design a novel streamline clustering-based structural connectivity feature for a robust representation of voxels within the nuclei. Finally, we improve the popular joint dimensionality reduction and k-means clustering approach to enable nuclei parcellation at a finer scale. We demonstrate DeepNuParc on two important brain structures, i.e. the amygdala and the thalamus, that are known to have multiple anatomically and functionally distinct nucleus subdivisions. Experimental results show that DeepNuParc enables consistent parcellation of the nuclei into multiple parcels across multiple subjects and achieves good correspondence with the widely used coarse-scale atlases. Our code is available at https://github.com/HarlandZZC/deep_nuclei_parcellation.

JBHI Journal 2025 Journal Article

DiffuSeg: Domain-Driven Diffusion for Medical Image Segmentation

  • Le Zhang
  • Fuping Wu
  • Kevin Bronik
  • Bartlomiej W. Papiez

In recent years, the deployment of supervised machine learning techniques for segmentation tasks has significantly increased. Nonetheless, the annotation process for extensive datasets remains costly, labor-intensive, and error-prone. While acquiring sufficiently large datasets to train deep learning models is feasible, these datasets often experience a distribution shift relative to the actual test data. This problem is particularly critical in the domain of medical imaging, where it adversely affects the efficacy of automatic segmentation models. In this work, we introduce DiffuSeg, a novel conditional diffusion model developed for medical image data, that exploits any labels to synthesize new images in the target domain. This allows a number of new research directions, including the segmentation task that motivates this work. Our method only requires label maps from any existing datasets and unlabelled images from the target domain for image diffusion. To learn the target domain knowledge, a feature factorization variational autoencoder is proposed to provide conditional information for the diffusion model. Consequently, the segmentation network can be trained with the given labels and the synthetic images, thus avoiding human annotations. Initially, we apply our method to the MNIST dataset and subsequently adapt it for use with medical image segmentation datasets, such as retinal fundus images for vessel segmentation and MRI images for heart segmentation. Our approach exhibits significant improvements over relevant baselines in both image generation and segmentation accuracy, especially in scenarios where annotations for the target dataset are unavailable during training. An open-source implementation of our approach can be released after reviewing. .

JBHI Journal 2025 Journal Article

Enhancing Ultrasound Scanning Skills in a Leader–Follower Robotic System through Expert Hand Impedance Regulation

  • Baoshan Niu
  • Dapeng Yang
  • Le Zhang
  • Yiming Ji
  • Li Jiang
  • Hong Liu

Traditional breast cancer surgeries require collaboration between ultrasound (US) doctors and surgeons, making the procedure complex and treating physicians prone to fatigue. In leader–follower robotic surgery, a surgeon controls an US robotic arm and an instrument robotic arm with their left and right hands, enabling independent surgical performance. However, the lack of US scanning skills among surgeons, as well as the physical separation in leader–follower operations, can negatively impact both the scanning and surgical outcomes. This paper proposes a robot-assisted scheme based on dynamic arm impedance compensation (IC) that references expert arm stiffness to compensate for novice arm stiffness. The impedance compensator adjusts the compensation strategy according to the scanning area and scanning stage. The impedance force generator estimates the scanning direction via Kalman filtering and applies stiffness and damping forces in the vertical direction to suppress tremors and other involuntary movements. The experimental results revealed that during the coarse and fine scanning phases, the probe position variance decreased by 57. 9% and 73. 6%, the contact force variance decreased by 55. 2% and 42. 5%, and the US image confidence increased by 22. 0% and 23. 8%, respectively. Compared with traditional filtering compensation (FC) schemes, this approach reduces the average position variance and contact force variance by 32. 0% and 25. 3%, respectively, and increases confidence by 7. 3%. In a no-compensation test, the IC training group outperformed the FC group. This scheme can assist leader–follower US scanning and rapidly improve surgical skills.

ICLR Conference 2025 Conference Paper

ThermalGaussian: Thermal 3D Gaussian Splatting

  • Rongfeng Lu
  • Hangyu Chen
  • Zunjie Zhu
  • Yuhang Qin
  • Ming Lu 0002
  • Le Zhang
  • Chenggang Yan 0001
  • Anke Xue

Thermography is especially valuable for the military and other users of surveillance cameras. Some recent methods based on Neural Radiance Fields (NeRF) are proposed to reconstruct the thermal scenes in 3D from a set of thermal and RGB images. However, unlike NeRF, 3D Gaussian splatting (3DGS) prevails due to its rapid training and real-time rendering. In this work, we propose ThermalGaussian, the first thermal 3DGS approach capable of rendering high-quality images in RGB and thermal modalities. We first calibrate the RGB camera and the thermal camera to ensure that both modalities are accurately aligned. Subsequently, we use the registered images to learn the multimodal 3D Gaussians. To prevent the overfitting of any single modality, we introduce several multimodal regularization constraints. We also develop smoothing constraints tailored to the physical characteristics of the thermal modality. Besides, we contribute a real-world dataset named RGBT-Scenes, captured by a hand-hold thermal-infrared camera, facilitating future research on thermal scene reconstruction. We conduct comprehensive experiments to show that ThermalGaussian achieves photorealistic rendering of thermal images and improves the rendering quality of RGB images. With the proposed multimodal regularization constraints, we also reduced the model's storage cost by 90\%. Our project page is at https://thermalgaussian.github.io/.

AAAI Conference 2025 Conference Paper

WiFi CSI Based Temporal Activity Detection via Dual Pyramid Network

  • Zhendong Liu
  • Le Zhang
  • Bing Li
  • Yingjie Zhou
  • Zhenghua Chen
  • Ce Zhu

We address the challenge of WiFi-based temporal activity detection and propose an efficient Dual Pyramid Network that integrates Temporal Signal Semantic Encoders and Local Sensitive Response Encoders. The Temporal Signal Semantic Encoder splits feature learning into high and low-frequency components, using a novel Signed Mask-Attention mechanism to emphasize important areas and downplay unimportant ones, with the features fused using ContraNorm. The Local Sensitive Response Encoder captures fluctuations without learning. These feature pyramids are then combined using a new cross-attention fusion mechanism. We also introduce a dataset with over 2,114 activity segments across 553 WiFi CSI samples, each lasting around 85 seconds. Extensive experiments show our method outperforms challenging baselines.

IJCAI Conference 2024 Conference Paper

DGR: A General Graph Desmoothing Framework for Recommendation via Global and Local Perspectives

  • Leilei Ding
  • Dazhong Shen
  • Chao Wang
  • Tianfu Wang
  • Le Zhang
  • Yanyong Zhang

Graph Convolutional Networks (GCNs) have become pivotal in recommendation systems for learning user and item embeddings by leveraging the user-item interaction graph's node information and topology. However, these models often face the famous over-smoothing issue, leading to indistinct user and item embeddings and reduced personalization. Traditional desmoothing methods in GCN-based systems are model-specific, lacking a universal solution. This paper introduces a novel, model-agnostic approach named Desmoothing Framework for GCN-based Recommendation Systems (DGR). It effectively addresses over-smoothing on general GCN-based recommendation models by considering both global and local perspectives. Specifically, we first introduce vector perturbations during each message passing layer to penalize the tendency of node embeddings approximating overly to be similar with the guidance of the global topological structure. Meanwhile, we further develop a tailored-design loss term for the readout embeddings to preserve the local collaborative relations between users and their neighboring items. In particular, items that exhibit a high correlation with neighboring items are also incorporated to enhance the local topological information. To validate our approach, we conduct extensive experiments on 5 benchmark datasets based on 5 well-known GCN-based recommendation models, demonstrating the effectiveness and generalization of our proposed framework. Our code is available at https: //github. com/me-sonandme/DGR.

NeurIPS Conference 2024 Conference Paper

MSA Generation with Seqs2Seqs Pretraining: Advancing Protein Structure Predictions

  • Le Zhang
  • Jiayang Chen
  • Tao Shen
  • Yu Li
  • Siqi Sun

Deep learning models like AlphaFold2 have revolutionized protein structure prediction, achieving unprecedented accuracy. However, the dependence on robust multiple sequence alignments (MSAs) continues to pose a challenge, especially for proteins that lack a wealth of homologous sequences. To overcome this limitation, we introduce MSA-Generator, a self-supervised generative protein language model. Trained on a sequence-to-sequence task using an automatically constructed dataset, MSA-Generator employs protein-specific attention mechanisms to harness large-scale protein databases, generating virtual MSAs that enrich existing ones and boost prediction accuracy. Our experiments on CASP14 and CASP15 benchmarks reveal significant improvements in LDDT scores, particularly for complex and challenging sequences, enhancing the performance of both AlphaFold2 and RoseTTAFold. The code is released at \url{https: //github. com/lezhang7/MSAGen}.

NeurIPS Conference 2024 Conference Paper

VisMin: Visual Minimal-Change Understanding

  • Rabiul Awal
  • Saba Ahmadi
  • Le Zhang
  • Aishwarya Agrawal

Fine-grained understanding of objects, attributes, and relationships between objects is crucial for visual-language models (VLMs). To evaluate VLMs' fine-grained understanding, existing benchmarks primarily focus on evaluating VLMs' capability to distinguish between two very similar captions given an image. In this paper, our focus is on evaluating VLMs' capability to distinguish between two very similar images given a caption. To this end, we introduce a new, challenging benchmark termed Visual Minimal-Change Understanding (VisMin), which requires models to predict the correct image-caption match given two images and two captions. Importantly, the image pair (as well as the caption pair) contains minimal changes, i. e. , between the two images (as well as between the two captions), only one aspect changes at a time from among the following possible types of changes: object, attribute, count, and spatial relation. These four types of minimal changes are specifically designed to test the models' understanding of objects, attributes of objects (such as color, material, shape), counts of objects, and spatial relationships between objects. To curate our benchmark, we built an automatic pipeline using large language models and diffusion models, followed by a rigorous 4-step verification process by human annotators. Empirical experiments reveal that current VLMs exhibit notable deficiencies in understanding spatial relationships and counting abilities. Furthermore, leveraging the automated nature of our data creation process, we generate a large-scale training dataset, which we use to finetune CLIP (a foundational VLM) and Idefics2 (a multimodal large language model). Our findings show that both these models benefit significantly from fine-tuning on this data, as evident by marked improvements in fine-grained understanding across a wide range of benchmarks. Additionally, such fine-tuning improves CLIP's general image-text alignment capabilities too. All resources including the benchmark, the training data, and the finetuned model checkpoints will be released.

JBHI Journal 2023 Journal Article

Anatomically Guided Cross-Domain Repair and Screening for Ultrasound Fetal Biometry

  • Jun Gao
  • Qicheng Lao
  • Paul Liu
  • Huahui Yi
  • Qingbo Kang
  • Zekun Jiang
  • Xiaohu Wu
  • Kang Li

Ultrasound based estimation of fetal biometry is extensively used to diagnose prenatal abnormalities and to monitor fetal growth, for which accurate segmentation of the fetal anatomy is a crucial prerequisite. Although deep neural network-based models have achieved encouraging results on this task, inevitable distribution shifts in ultrasound images can still result in severe performance drop in real world deployment scenarios. In this article, we propose a complete ultrasound fetal examination system to deal with this troublesome problem by repairing and screening the anatomically implausible results. Our system consists of three main components: A routine segmentation network, a fetal anatomical key points guided repair network, and a shape-coding based selective screener. Guided by the anatomical key points, our repair network has stronger cross-domain repair capabilities, which can substantially improve the outputs of the segmentation network. By quantifying the distance between an arbitrary segmentation mask to its corresponding anatomical shape class, the proposed shape-coding based selective screener can then effectively reject the entire implausible results that cannot be fully repaired. Extensive experiments demonstrate that our proposed framework has strong anatomical guarantee and outperforms other methods in three different cross-domain scenarios.

JBHI Journal 2023 Journal Article

Dual-Stream Contrastive Learning for Channel State Information Based Human Activity Recognition

  • Ke Xu
  • Jiangtao Wang
  • Le Zhang
  • Hongyuan Zhu
  • Dingchang Zheng

WiFi-based human activity recognition (HAR) has been extensively studied due to its far-reaching applications in health domains, including elderly monitoring, exercise supervision and rehabilitation monitoring, etc. Although existing supervised deep learning techniques have achieved remarkable performances for these tasks, they are however data-hungry and hence are notoriously difficult due to the privacy and incomprehensibility of WiFi-based HAR data. Existing contrastive learning models, mainly designed for computer vision, cannot guarantee their performance on channel state information (CSI) data. To this end, we propose a new dual-stream contrastive learning model that can process and learn the raw WiFi CSI data in a self-supervised manner. More specifically, our proposed method, coined as DualConFi, takes raw WiFI CSI data as input and incorporates channel and temporal streams to learn highly-discriminative spatiotemporal features under a mutual information constraint using unlabeled data. We exhibit the effectiveness of our model on three publicly available CSI data sets in various experiment settings, including linear evaluation, semi-supervised, and transfer learning. We show that DualConFi is able to perform favourably against challenging baselines in each setting. Moreover, by studying the effects of different transform functions on CSI data, we finally verify the effectiveness of highly-discriminative features.

AIIM Journal 2023 Journal Article

Least squares support vector regression for complex censored data

  • Xinrui Liu
  • Xiaogang Dong
  • Le Zhang
  • Jia Chen
  • Chunjie Wang

Least squares support vector regression (LS-SVR) is a robust machine learning algorithm for small sample data. Its solution is derived from solving a set of linear equations, making the calculation process straightforward. In order to overcome the difficulties of the regression estimations when the responses are subject to interval censoring or left truncation and right censoring, two LS-SVR methods are proposed. For interval-censored data, one can easily estimate the regression functions by combining the imputation techniques and LS-SVR for right-censored data. For left-truncated and right-censored data, a weight is used to reduce the effects of truncation and censoring on the LS-SVR procedure. Simulation results show that the proposed methods can reduce regression error and yield high accuracy and stability.

NeurIPS Conference 2021 Conference Paper

Learning to Iteratively Solve Routing Problems with Dual-Aspect Collaborative Transformer

  • Yining Ma
  • Jingwen Li
  • Zhiguang Cao
  • Wen Song
  • Le Zhang
  • Zhenghua Chen
  • Jing Tang

Recently, Transformer has become a prevailing deep architecture for solving vehicle routing problems (VRPs). However, it is less effective in learning improvement models for VRP because its positional encoding (PE) method is not suitable in representing VRP solutions. This paper presents a novel Dual-Aspect Collaborative Transformer (DACT) to learn embeddings for the node and positional features separately, instead of fusing them together as done in existing ones, so as to avoid potential noises and incompatible correlations. Moreover, the positional features are embedded through a novel cyclic positional encoding (CPE) method to allow Transformer to effectively capture the circularity and symmetry of VRP solutions (i. e. , cyclic sequences). We train DACT using Proximal Policy Optimization and design a curriculum learning strategy for better sample efficiency. We apply DACT to solve the traveling salesman problem (TSP) and capacitated vehicle routing problem (CVRP). Results show that our DACT outperforms existing Transformer based improvement models, and exhibits much better generalization performance across different problem sizes on synthetic and benchmark instances, respectively.

AAAI Conference 2021 Conference Paper

Two-Stream Convolution Augmented Transformer for Human Activity Recognition

  • Bing Li
  • Wei Cui
  • Wei Wang
  • Le Zhang
  • Zhenghua Chen
  • Min Wu

Recognition of human activities is an important task due to its far-reaching applications such as healthcare system, context-aware applications, and security monitoring. Recently, WiFi based human activity recognition (HAR) is becoming ubiquitous due to its non-invasiveness. Existing WiFibased HAR methods regard WiFi signals as a temporal sequence of channel state information (CSI), and employ deep sequential models (e. g. , RNN, LSTM) to automatically capture channel-over-time features. Although being remarkably effective, they suffer from two major drawbacks. Firstly, the granularity of a single temporal point is blindly elementary for representing meaningful CSI patterns. Secondly, the timeover-channel features are also important, and could be a natural data augmentation. To address the drawbacks, we propose a novel Two-stream Convolution Augmented Human Activity Transformer (THAT) model. Our model proposes to utilize a two-stream structure to capture both time-over-channel and channel-over-time features, and use the multi-scale convolution augmented transformer to capture range-based patterns. Extensive experiments on four real experiment datasets demonstrate that our model outperforms state-of-the-art models in terms of both effectiveness and efficiency 1.

NeurIPS Conference 2020 Conference Paper

Disentangling Human Error from Ground Truth in Segmentation of Medical Images

  • Le Zhang
  • Ryutaro Tanno
  • Mou-Cheng Xu
  • Chen Jin
  • Joseph Jacob
  • Olga Cicarrelli
  • Frederik Barkhof
  • Daniel Alexander

Recent years have seen increasing use of supervised learning methods for segmentation tasks. However, the predictive performance of these algorithms depends on the quality of labels. This problem is particularly pertinent in the medical image domain, where both the annotation cost and inter-observer variability are high. In a typical label acquisition process, different human experts provide their estimates of the ``true'' segmentation labels under the influence of their own biases and competence levels. Treating these noisy labels blindly as the ground truth limits the performance that automatic segmentation algorithms can achieve. In this work, we present a method for jointly learning, from purely noisy observations alone, the reliability of individual annotators and the true segmentation label distributions, using two coupled CNNs. The separation of the two is achieved by encouraging the estimated annotators to be maximally unreliable while achieving high fidelity with the noisy training data. We first define a toy segmentation dataset based on MNIST and study the properties of the proposed algorithm. We then demonstrate the utility of the method on three public medical imaging segmentation datasets with simulated (when necessary) and real diverse annotations: 1) MSLSC (multiple-sclerosis lesions); 2) BraTS (brain tumours); 3) LIDC-IDRI (lung abnormalities). In all cases, our method outperforms competing methods and relevant baselines particularly in cases where the number of annotations is small and the amount of disagreement is large. The experiments also show strong ability to capture the complex spatial characteristics of annotators' mistakes. Our code is available at \url{https: //github. com/moucheng2017/Learn Noisy Labels Medical Images}.

IJCAI Conference 2019 Conference Paper

Trend-Aware Tensor Factorization for Job Skill Demand Analysis

  • Xunxian Wu
  • Tong Xu
  • Hengshu Zhu
  • Le Zhang
  • Enhong Chen
  • Hui Xiong

Given a job position, how to identify the right job skill demand and its evolving trend becomes critically important for both job seekers and employers in the fast-paced job market. Along this line, there still exist various challenges due to the lack of holistic understanding on skills related factors, e. g. , the dynamic validity periods of skill trend, as well as the constraints from overlapped business and skill co-occurrence. To address these challenges, in this paper, we propose a trend-aware approach for fine-grained skill demand analysis. Specifically, we first construct a tensor for each timestamp based on the large-scale recruitment data, and then reveal the aggregations among companies and skills by heuristic solutions. Afterwards, the Trend-Aware Tensor Factorization (TATF) framework is designed by integrating multiple confounding factors, i. e. , aggregation-based and temporal constraints, to provide more fine-grained representation and evolving trend of job demand for specific job positions. Finally, validations on large-scale real-world data clearly validate the effectiveness of our approach for skill demand analysis.

IJCAI Conference 2018 Conference Paper

DEL: Deep Embedding Learning for Efficient Image Segmentation

  • Yun Liu
  • Peng-tao Jiang
  • Vahan Petrosyan
  • Shi-Jie Li
  • Jiawang Bian
  • Le Zhang
  • Ming-Ming Cheng

Image segmentation has been explored for many years and still remains a crucial vision problem. Some efficient or accurate segmentation algorithms have been widely used in many vision applications. However, it is difficult to design a both efficient and accurate image segmenter. In this paper, we propose a novel method called DEL (deep embedding learning) which can efficiently transform superpixels into image segmentation. Starting with the SLIC superpixels, we train a fully convolutional network to learn the feature embedding space for each superpixel. The learned feature embedding corresponds to a similarity measure that measures the similarity between two adjacent superpixels. With the deep similarities, we can directly merge the superpixels into large segments. The evaluation results on BSDS500 and PASCAL Context demonstrate that our approach achieves a good trade-off between efficiency and effectiveness. Specifically, our DEL algorithm can achieve comparable segments when compared with MCG but is much faster than it, i. e. 11. 4fps vs. 0. 07fps.

AAAI Conference 2018 Conference Paper

Kernel Cross-Correlator

  • Chen Wang
  • Le Zhang
  • Lihua Xie
  • Junsong Yuan

Cross-correlator plays a significant role in many visual perception tasks, such as object detection and tracking. Beyond the linear cross-correlator, this paper proposes a kernel crosscorrelator (KCC) that breaks traditional limitations. First, by introducing the kernel trick, the KCC extends the linear crosscorrelation to non-linear space, which is more robust to signal noises and distortions. Second, the connection to the existing works shows that KCC provides a unified solution for correlation filters. Third, KCC is applicable to any kernel function and is not limited to circulant structure on training data, thus it is able to predict affine transformations with customized properties. Last, by leveraging the fast Fourier transform (FFT), KCC eliminates direct calculation of kernel vectors, thus achieves better performance yet still with a reasonable computational cost. Comprehensive experiments on visual tracking and human activity recognition using wearable devices demonstrate its robustness, flexibility, and efficiency. The source codes of both experiments are released at https: //github. com/wang-chen/KCC.

ICRA Conference 2007 Conference Paper

A Reinforcement Learning Based Dynamic Walking Control

  • Yong Mao
  • Jiaxin Wang
  • Peifa Jia
  • Shi Li 0002
  • Zhen Qiu
  • Le Zhang
  • Zhuo Han

A quasi-passive dynamic walking robot is built to study natural and energy-efficient biped walking. The robot is actuated by MACCEPA actuators. A reinforcement learning based control method is proposed to enhance the robustness and stability of the robot's walking. The proposed method first learns the desired gait for the robot's walking on a flat floor. Then a fuzzy advantage learning method is used to control it to walk on uneven floor. The effectiveness of the method is verified by simulation results.

IROS Conference 2007 Conference Paper

Development and depth control of biomimetic robotic fish

  • Le Zhang
  • Wei Zhao 0007
  • Yonghui Hu
  • Dandan Zhang 0004
  • Long Wang 0001

This paper describes the design of a biologically inspired robotic fish which is capable of three-dimensional locomotion, and presents its depth control method. The mechanical design of the pectoral fins improves both the waterproof capability and the control precision. By adjusting the rotation angles of the pectoral fins, the robotic fish can fulfill up-and-down motion. Due to the uncertainty of the depth information measured by the pressure sensor, fuzzy logic method is applied to the depth control of the robotic fish. The experimental results on the designed prototype verify that the given method is effective in design and implementation.

v2026.09.13