Arrow Research search

Author name cluster

Donghao Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

7 papers
2 author rows

Possible papers

7

UAI Conference 2025 Conference Paper

Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis

  • Ruiquan Huang
  • Donghao Li
  • Chengshuai Shi
  • Cong Shen 0001
  • Jing Yang 0002

This paper investigates a hybrid learning framework for reinforcement learning (RL) in which the agent can leverage both an offline dataset and online interactions to learn the optimal policy. We present a unified algorithm and analysis and show that augmenting confidence-based online RL algorithms with the offline dataset outperforms any pure online or offline algorithm alone and achieves state-of-the-art results under two learning metrics, i. e. , sub-optimality gap and online learning regret. Specifically, we show that our algorithm achieves a sub-optimality gap $\tilde{O}( \sqrt{1/(N_0/ \mathtt{C}(\pi^\star| \rho)+N_1} ) )$, where $\mathtt{C}(\pi^\star|\rho)$ is a new concentrability coefficient, $N_0$ and $N_1$ are the numbers of offline and online samples, respectively. For regret minimization, we show that it achieves a constant $\tilde{O}( \sqrt{N_1/(N_0/\mathtt{C}(\pi^{-}|\rho)+N_1)} )$ speed-up compared to pure online learning, where $\mathtt{C}(\pi^-|\rho)$ is the concentrability coefficient over all sub-optimal policies. Our results also reveal an interesting separation on the desired coverage properties of the offline dataset for sub-optimality gap minimization and regret minimization. We further validate our theoretical findings in several experiments in special RL models such as linear contextual bandits and Markov decision processes (MDPs).

JMLR Journal 2024 Journal Article

Random Smoothing Regularization in Kernel Gradient Descent Learning

  • Liang Ding
  • Tianyang Hu
  • Jiahang Jiang
  • Donghao Li
  • Wenjia Wang
  • Yuan Yao

Random smoothing data augmentation is a unique form of regularization that can prevent overfitting by introducing noise to the input data, encouraging the model to learn more generalized features. Despite its success in various applications, there has been a lack of systematic study on the regularization ability of random smoothing. In this paper, we aim to bridge this gap by presenting a framework for random smoothing regularization that can adaptively and effectively learn a wide range of ground truth functions belonging to the classical Sobolev spaces. Specifically, we investigate two underlying function spaces: the Sobolev space of low intrinsic dimension, which includes the Sobolev space in D-dimensional Euclidean space or low-dimensional sub-manifolds as special cases, and the mixed smooth Sobolev space with a tensor structure. By using random smoothing regularization as novel convolution-based smoothing kernels, we can attain optimal convergence rates in these cases using a kernel gradient descent algorithm, either with early stopping or weight decay. It is noteworthy that our estimator can adapt to the structural assumptions of the underlying data and avoid the curse of dimensionality. This is achieved through various choices of injected noise distributions such as Gaussian, Laplace, or general polynomial noises, allowing for broad adaptation to the aforementioned structural assumptions of the underlying data. The convergence rate depends only on the effective dimension, which may be significantly smaller than the actual data dimension. We conduct numerical experiments on simulated data to validate our theoretical results. [abs] [ pdf ][ bib ] &copy JMLR 2024. ( edit, beta )

IROS Conference 2023 Conference Paper

Development of an Autonomous Modular Swimming Robot with Disturbance Rejection and Path Tracking

  • Hankun Deng
  • Colin Nitroy
  • Kundan Panta
  • Donghao Li
  • Shashank Priya
  • Bo Cheng 0008

Here we present the development of an autonomous modular swimming robot. This robot, named µBot 2. 0, was upgraded from our previous robot platform µBot and features onboard computing, sensing, and power. Its compact size and modularity render the robot an ideal platform for studying bio-inspired robot swimming. The robot is equipped with a micro controller in its head that communicates with external computers through Bluetooth Low Energy (BLE) and sends motor commands to the body segments via Inter-Integrated Circuit (I2C) protocol. Each body segment has a customized printed circuit board (PCB) that receives commands and controls the electromagnetic actuator for generating body movements. The robot head is also equipped with an Inertial Measurement Unit (IMU) to measure its heading and a battery for power. In this work, a µBot 2. 0 with three actuators was assembled and the swimming performance was tested. The robot actuators were activated via rhythmic motor input from a central pattern generator (CPG). Experimental results showed that the swimming speed was highly sensitive to the frequency of the motor input, with a maximum swimming speed of 130 mm/s (equivalent to 0. 7 body length per second) at 6 Hz. The robot also had the capability to correct its heading with IMU feedback and follow desired paths using a line-of-sight (LOS) guidance law with an overhead camera. Our results demonstrate the effectiveness of the robot's design and its potential in a variety of aquatic applications.

ICML Conference 2023 Conference Paper

Near-optimal Conservative Exploration in Reinforcement Learning under Episode-wise Constraints

  • Donghao Li
  • Ruiquan Huang
  • Cong Shen 0001
  • Jing Yang 0002

This paper investigates conservative exploration in reinforcement learning where the performance of the learning agent is guaranteed to be above a certain threshold throughout the learning process. It focuses on the tabular episodic Markov Decision Process (MDP) setting that has finite states and actions. With the knowledge of an existing safe baseline policy, an algorithms termed as StepMix is proposed to balance the exploitation and exploration while ensuring that the conservative constraint is never violated in each episode with high probability. StepMix features a unique design of a mixture policy that adaptively and smoothly interpolates between the baseline policy and the optimistic policy. Theoretical analysis shows that StepMix achieves near-optimal regret order as in the constraint-free setting, indicating that obeying the stringent episode-wise conservative constraint does not compromise the learning performance. Besides, a randomization based EpsMix algorithm is also proposed and shown the achieve the same performance as StepMix. The algorithm design and theoretical analysis are further extended to the setting where the baseline policy is not given a priori but must be learned from an offline dataset, and it is proved that similar conservative guarantee and regret can be achieved if the offline dataset is sufficiently large. Experiment results corroborate the theoretical analysis and demonstrate the effectiveness of the proposed conservative exploration strategies.

IROS Conference 2022 Conference Paper

Effects of Design and Hydrodynamic Parameters on Optimized Swimming for Simulated, Fish-inspired Robots

  • Donghao Li
  • Hankun Deng
  • Yagiz E. Bayiz
  • Bo Cheng 0008

In this work, we developed a mathematical model and a simulation platform for a fish-inspired robotic template, namely Magnetic, Modular, Undulatory Robot $(\mu \text{Bot})$. Through this platform, we systematically explored the effects of robot design and fluid parameters on swimming performance via reinforcement learning. The mathematical model was composed of two interacting subsystems, the robotic dynamic model and the hydrodynamic model. The hydrodynamic model consisted of the reactive components (added-mass force and pressure forces) and the resistive components (drag and friction forces). These components were nondimensionalized for deriving key “control parameters” of the robot-fluid interaction. The $\mu\text{Bots}$ were actuated via magnetic actuators controlled with harmonic voltage signals, which were optimized via EM-based Policy Hyper Parameter Exploration (EPHE) to maximize forward swimming speed. By varying the control parameters, a total of 36 cases with different robot template variations (Number of Actuators (NoA) and stiffness) and hydrodynamic parameters were simulated and optimized via EPHE. Results showed that the wavelength of the optimized gaits (i. e. , backward traveling wave along the body) was independent of template variations and hydrodynamic parameters. Higher NoA yielded higher speed but lower speed per body length, suggesting a diminishing gain from added actuators. Body and caudal-fin dynamics were dominated by the interaction among fluid added-mass, spring, and actuation torque, with negligible contribution from fluid resistive drag. In contrast, thrust was dominated by the pressure force acting on the caudal fin, as steady swimming resulted from a balance between resistive force and pressure force, with minor contributions from added-mass force and body drag forces. Therefore, added-mass force only indirectly affected the thrust generation and forward swimming speed via the caudal fin dynamics.

IROS Conference 2021 Conference Paper

Design and Experimental Learning of Swimming Gaits for a Magnetic, Modular, Undulatory Robot

  • Hankun Deng
  • Patrick Burke
  • Donghao Li
  • Bo Cheng 0008

Here we developed an experimental platform with a magnetic, modular, undulatory robot (μBot) for studying fish-inspired underwater locomotion. This platform will enable us to systematically explore the relationship between body morphology, swimming gaits, and swimming performance via reinforcement learning methods. The μBot was designed to be easily modifiable in morphology, compact in size, easy to be controlled and inexpensive. The experimental platform also included a towing tank and a motion tracking system for real-time measurement of the μBot kinematics. The swimming gaits of μBot were generated by a central pattern generator (CPG), which outputs voltage signals to μBot's magnetic actuators. The CPG parameters were learned experimentally using the parameter exploring policy gradient (PGPE) method to maximize swimming speed. In the experiments, two μBot designs with the same body morphology but different caudal-fin shapes were tested. Results showed that swimming gaits with back-propagating traveling waves can be learned experimentally via PGPE, while the shape of the caudal fins had moderate influences on the learned gaits and the swimming speed. Furthermore, robot swimming speed was sensitive to the undulating frequency and the voltage magnitude of the last three posterior actuators. In contrast, swimming gaits and speed were relatively invariant to the variances within the inter-module connection weights of CPG and the voltage applied to the anterior actuator.

ICML Conference 2020 Conference Paper

DessiLBI: Exploring Structural Sparsity of Deep Networks via Differential Inclusion Paths

  • Yanwei Fu 0001
  • Chen Liu 0030
  • Donghao Li
  • Xinwei Sun 0001
  • Jinshan Zeng
  • Yuan Yao 0011

Over-parameterization is ubiquitous nowadays in training neural networks to benefit both optimization in seeking global optima and generalization in reducing prediction error. However, compressive networks are desired in many real world applications and direct training of small networks may be trapped in local optima. In this paper, instead of pruning or distilling over-parameterized models to compressive ones, we propose a new approach based on differential inclusions of inverse scale spaces. Specifically, it generates a family of models from simple to complex ones that couples a pair of parameters to simultaneously train over-parameterized deep models and structural sparsity on weights of fully connected and convolutional layers. Such a differential inclusion scheme has a simple discretization, proposed as Deep structurally splitting Linearized Bregman Iteration (DessiLBI), whose global convergence analysis in deep learning is established that from any initializations, algorithmic iterations converge to a critical point of empirical risks. Experimental evidence shows that DessiLBI achieve comparable and even better performance than the competitive optimizers in exploring the structural sparsity of several widely used backbones on the benchmark datasets. Remarkably, with early stopping, DessiLBI unveils “winning tickets” in early epochs: the effective sparse structure with comparable test accuracy to fully trained over-parameterized models.

v2026.09.13