Arrow Research search

Author name cluster

Hui Yuan

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

10 papers
1 author row

Possible papers

10

EAAI Journal 2025 Journal Article

A diffusion model and knowledge distillation framework for robust coral detection in complex underwater environments

  • Zhaoxuan Lu
  • Lyuchao Liao
  • Chuang Li
  • Xingang Xie
  • Hui Yuan

Coral reefs play a crucial role in marine ecosystems, but their sustainability is increasingly threatened by climate change and human activities. To aid in the protection and monitoring of these ecosystems, developing advanced artificial intelligence (AI)-based automated detection technologies is essential. This paper introduces the MambaCoral-Diffusion Detection framework (MambaCoral-Diffusion Detection, MCDD), an AI-driven approach for robust coral detection, designed to enhance performance in complex underwater environments—a critical challenge in marine engineering. Key AI contributions include integrating a diffusion model to generate realistic and diverse training data from limited and challenging underwater coral datasets, effectively alleviating the issue of data imbalance. Secondly, we adopted an innovative spatial sensing detection mechanism that enhances the accuracy of feature extraction in complex underwater environments. Finally, we introduced an efficient knowledge distillation technique that successfully transfers knowledge from complex models to more lightweight counterparts, thereby reducing computational resource requirements while maintaining efficiency and facilitating practical deployment. Experimental results show that MCDD achieves high performance on the Soft Coral dataset, reaching 91 frames per second (FPS), the mean average precision at 50% Intersection over Union (IoU) threshold of 0. 843, and the mean average precision averaged across IoU thresholds from 50% to 95% of 0. 566, with only 6. 5 million parameters and 13. 6 billion floating point operations per second (GFLOPs). These results demonstrate MCDD’s reliability and efficiency in detecting corals under complex underwater conditions, highlighting its significant potential for advancing marine research and conservation efforts. The code and dataset are available at https: //github. com/RDXiaoLu/MambaCoral-DiffDet. git.

NeurIPS Conference 2025 Conference Paper

Training-Free Guidance Beyond Differentiability: Scalable Path Steering with Tree Search in Diffusion and Flow Models

  • Yingqing Guo
  • Yukang Yang
  • Hui Yuan
  • Mengdi Wang

Training-free guidance enables controlled generation in diffusion and flow models, but most methods rely on gradients and assume differentiable objectives. This work focuses on training-free guidance addressing challenges from non-differentiable objectives and discrete data distributions. We propose TreeG: Tree Search-Based Path Steering Guidance, applicable to both continuous and discrete settings in diffusion and flow models. TreeG offers a unified framework for training-free guidance by proposing, evaluating, and selecting candidates at each step, enhanced with tree search over active paths and parallel exploration. We comprehensively investigate the design space of TreeG over the candidate proposal module and the evaluation function, instantiating TreeG into three novel algorithms. Our experiments show that TreeG consistently outperforms top guidance baselines in symbolic music generation, small molecule design, and enhancer DNA design with improvements of 29. 01%, 26. 38%, and 18. 43%. Additionally, we identify an inference-time scaling law showing TreeG's scalability in inference-time computation.

TMLR Journal 2024 Journal Article

Adversarial Attacks on Online Learning to Rank with Stochastic Click Models

  • Zichen Wang
  • Rishab Balasubramanian
  • Hui Yuan
  • Chenyu Song
  • Mengdi Wang
  • Huazheng Wang

We propose the first study of adversarial attacks on online learning to rank. The goal of the attacker it to misguide the online learning to rank algorithm to place the target item on top of the ranking list linear times to time horizon $T$ with a sublinear attack cost. We propose generalized list poisoning attacks that perturb the ranking list presented to the user. This strategy can efficiently attack any no-regret ranker in general stochastic click models. Furthermore, we propose a click poisoning-based strategy named attack-then-quit that can efficiently attack two representative OLTR algorithms for stochastic click models. We theoretically analyze the success and cost upper bound of the two proposed methods. Experimental results based on synthetic and real-world data further validate the effectiveness and cost-efficiency of the proposed attack strategies.

NeurIPS Conference 2024 Conference Paper

Gradient Guidance for Diffusion Models: An Optimization Perspective

  • Yingqing Guo
  • Hui Yuan
  • Yukang Yang
  • Minshuo Chen
  • Mengdi Wang

Diffusion models have demonstrated empirical successes in various applications and can be adapted to task-specific needs via guidance. This paper studies a form of gradient guidance for adapting a pre-trained diffusion model towards optimizing user-specified objectives. We establish a mathematical framework for guided diffusion to systematically study its optimization theory and algorithmic design. Our theoretical analysis spots a strong link between guided diffusion models and optimization: gradient-guided diffusion models are essentially sampling solutions to a regularized optimization problem, where the regularization is imposed by the pre-training data. As for guidance design, directly bringing in the gradient of an external objective function as guidance would jeopardize the structure in generated samples. We investigate a modified form of gradient guidance based on a forward prediction loss, which leverages the information in pre-trained score functions and provably preserves the latent structure. We further consider an iteratively fine-tuned version of gradient-guided diffusion where guidance and score network are both updated with newly generated samples. This process mimics a first-order optimization iteration in expectation, for which we proved $\tilde{\mathcal{O}}(1/K)$ convergence rate to the global optimum when the objective function is concave. Our code is released at https: //github. com/yukang123/GGDMOptim. git.

AAAI Conference 2024 Conference Paper

Tree Search-Based Evolutionary Bandits for Protein Sequence Optimization

  • Jiahao Qiu
  • Hui Yuan
  • Jinghong Zhang
  • Wentao Chen
  • Huazheng Wang
  • Mengdi Wang

While modern biotechnologies allow synthesizing new proteins and function measurements at scale, efficiently exploring a protein sequence space and engineering it remains a daunting task due to the vast sequence space of any given protein. Protein engineering is typically conducted through an iterative process of adding mutations to the wild-type or lead sequences, recombination of mutations, and running new rounds of screening. To enhance the efficiency of such a process, we propose a tree search-based bandit learning method, which expands a tree starting from the initial sequence with the guidance of a bandit machine learning model. Under simplified assumptions and a Gaussian Process prior, we provide theoretical analysis and a Bayesian regret bound, demonstrating that the method can efficiently discover a near-optimal design. The full algorithm is compatible with a suite of randomized tree search heuristics, machine learning models, pre-trained embeddings, and bandit techniques. We test various instances of the algorithm across benchmark protein datasets using simulated screens. Experiment results demonstrate that the algorithm is both sample-efficient, diversity-promoting, and able to find top designs using reasonably small mutation counts.

NeurIPS Conference 2023 Conference Paper

Global Structure-Aware Diffusion Process for Low-light Image Enhancement

  • Jinhui HOU
  • Zhiyu Zhu
  • Junhui Hou
  • Hui Liu
  • Huanqiang Zeng
  • Hui Yuan

This paper studies a diffusion-based framework to address the low-light image enhancement problem. To harness the capabilities of diffusion models, we delve into this intricate process and advocate for the regularization of its inherent ODE-trajectory. To be specific, inspired by the recent research that low curvature ODE-trajectory results in a stable and effective diffusion process, we formulate a curvature regularization term anchored in the intrinsic non-local structures of image data, i. e. , global structure-aware regularization, which gradually facilitates the preservation of complicated details and the augmentation of contrast during the diffusion process. This incorporation mitigates the adverse effects of noise and artifacts resulting from the diffusion process, leading to a more precise and flexible enhancement. To additionally promote learning in challenging regions, we introduce an uncertainty-guided regularization technique, which wisely relaxes constraints on the most extreme regions of the image. Experimental evaluations reveal that the proposed diffusion-based framework, complemented by rank-informed regularization, attains distinguished performance in low-light enhancement. The outcomes indicate substantial advancements in image quality, noise suppression, and contrast amplification in comparison with state-of-the-art methods. We believe this innovative approach will stimulate further exploration and advancement in low-light image processing, with potential implications for other applications of diffusion models. The code is publicly available at https: //github. com/jinnh/GSAD.

NeurIPS Conference 2023 Conference Paper

Reward-Directed Conditional Diffusion: Provable Distribution Estimation and Reward Improvement

  • Hui Yuan
  • Kaixuan Huang
  • Chengzhuo Ni
  • Minshuo Chen
  • Mengdi Wang

We explore the methodology and theory of reward-directed generation via conditional diffusion models. Directed generation aims to generate samples with desired properties as measured by a reward function, which has broad applications in generative AI, reinforcement learning, and computational biology. We consider the common learning scenario where the dataset consists of majorly unlabeled data and a small set of data with noisy reward labels. Our approach leverages a learned reward function on the smaller data set as a pseudolabeler to label the unlabelled data. After pseudo-labelling, a conditional diffusion model (CDM) is trained on the data and samples are generated by setting a target value $a$ as the condition in CDM. From a theoretical standpoint, we show that this directed generator can effectively learn and sample from the reward-conditioned data distribution: 1. our model is capable of recovering the data's latent subspace representation. 2. the model generates samples moving closer to the user-specified target. The improvement in rewards of samples is influenced by a interplay between the strength of the reward signal, the distribution shift, and the cost of off-support extrapolation. We provide empirical results to validate our theory and highlight the relationship between the strength of extrapolation and the quality of generated samples.

NeurIPS Conference 2023 Conference Paper

Unified Off-Policy Learning to Rank: a Reinforcement Learning Perspective

  • Zeyu Zhang
  • Yi Su
  • Hui Yuan
  • Yiran Wu
  • Rishab Balasubramanian
  • Qingyun Wu
  • Huazheng Wang
  • Mengdi Wang

Off-policy Learning to Rank (LTR) aims to optimize a ranker from data collected by a deployed logging policy. However, existing off-policy learning to rank methods often make strong assumptions about how users generate the click data, i. e. , the click model, and hence need to tailor their methods specifically under different click models. In this paper, we unified the ranking process under general stochastic click models as a Markov Decision Process (MDP), and the optimal ranking could be learned with offline reinforcement learning (RL) directly. Building upon this, we leverage offline RL techniques for off-policy LTR and propose the Click Model-Agnostic Unified Off-policy Learning to Rank (CUOLR) method, which could be easily applied to a wide range of click models. Through a dedicated formulation of the MDP, we show that offline RL algorithms can adapt to various click models without complex debiasing techniques and prior knowledge of the model. Results on various large-scale datasets demonstrate that CUOLR consistently outperforms the state-of-the-art off-policy learning to rank algorithms while maintaining consistency and robustness under different click models.

NeurIPS Conference 2022 Conference Paper

Bandit Theory and Thompson Sampling-Guided Directed Evolution for Sequence Optimization

  • Hui Yuan
  • Chengzhuo Ni
  • Huazheng Wang
  • Xuezhou Zhang
  • Le Cong
  • Csaba Szepesvari
  • Mengdi Wang

Directed Evolution (DE), a landmark wet-lab method originated in 1960s, enables discovery of novel protein designs via evolving a population of candidate sequences. Recent advances in biotechnology has made it possible to collect high-throughput data, allowing the use of machine learning to map out a protein's sequence-to-function relation. There is a growing interest in machine learning-assisted DE for accelerating protein optimization. Yet the theoretical understanding of DE, as well as the use of machine learning in DE, remains limited. In this paper, we connect DE with the bandit learning theory and make a first attempt to study regret minimization in DE. We propose a Thompson Sampling-guided Directed Evolution (TS-DE) framework for sequence optimization, where the sequence-to-function mapping is unknown and querying a single value is subject to costly and noisy measurements. TS-DE updates a posterior of the function based on collected measurements. It uses a posterior-sampled function estimate to guide the crossover recombination and mutation steps in DE. In the case of a linear model, we show that TS-DE enjoys a Bayesian regret of order $\tilde O(d^{2}\sqrt{MT})$, where $d$ is feature dimension, $M$ is population size and $T$ is number of rounds. This regret bound is nearly optimal, confirming that bandit learning can provably accelerate DE. It may have implications for more general sequence optimization and evolutionary algorithms.

AAAI Conference 2020 Conference Paper

Learning Light Field Angular Super-Resolution via a Geometry-Aware Network

  • Jing Jin
  • Junhui Hou
  • Hui Yuan
  • Sam Kwong

The acquisition of light field images with high angular resolution is costly. Although many methods have been proposed to improve the angular resolution of a sparsely-sampled light field, they always focus on the light field with a small baseline, which is captured by a consumer light field camera. By making full use of the intrinsic geometry information of light fields, in this paper we propose an end-to-end learning-based approach aiming at angularly super-resolving a sparselysampled light field with a large baseline. Our model consists of two learnable modules and a physically-based module. Specifically, it includes a depth estimation module for explicitly modeling the scene geometry, a physically-based warping for novel views synthesis, and a light field blending module specifically designed for light field reconstruction. Moreover, we introduce a novel loss function to promote the preservation of the light field parallax structure. Experimental results over various light field datasets including large baseline light field images demonstrate the significant superiority of our method when compared with state-of-the-art ones, i. e. , our method improves the PSNR of the second best method up to 2 dB in average, while saves the execution time 48×. In addition, our method preserves the light field parallax structure better.

v2026.09.13