Arrow Research search

Author name cluster

Ye Yuan

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

49 papers
2 author rows

Possible papers

49

JBHI Journal 2026 Journal Article

An Experience-driven Interpretable Multi-task Model for Segmentation and Classification of Small Cell Lung Cancer and Non-small Cell Lung Cancer from CT Images

  • Zhaoshuo Diao
  • Manyu Cui
  • Taichang Xu
  • Ye Yuan
  • Guoyu Tong
  • Yue Gao

Lung cancer is the leading cause of cancer-related mortality, with small cell lung cancer and non-small cell lung cancer being the primary subtypes that exhibit distinct treatment approaches and prognostic outcomes. Accurate identification of these lung cancer classes holds significant importance in clinical practice. This study introduces an experience-driven interpretable multi-task network to concurrently perform segmentation and classification of small cell and non-small cell lung cancer. The core architecture of this multi-task model is based on StarNet, featuring a shared feature extraction branch and task-specific decoding branches for tumor segmentation and classification. Leveraging clinical knowledge of small cell lung cancer characteristics, such as indistinct edges, tissue invasion, and limited large cavity areas, two auxiliary branches are proposed: edge uncertainty estimation and tumor core area reconstruction. The values from edge uncertainty estimation and reconstruction integrity estimation are utilized in the classification branch to facilitate small cell lung cancer classification. Furthermore, for enhanced interpretability, bottleneck layer features are extracted for comparative learning, and a three-level contrastive loss is proposed to improve the differentiation of disease features. Lastly, an interpretable strategy based on trained feature query matching is presented, providing radiologists with clinical insights and reference images while the model outputs recognition predictions. Experimental results on the public dataset demonstrate that the proposed multi-task model not only outperforms single-task models but also offers a certain level of interpretability, thus enhancing radiologists' clinical decision-making processes.

YNIMG Journal 2026 Journal Article

Experimental evaluation of an integrated focused ultrasound and electroencephalography approach for developing activation-informed neuroimaging

  • Ye Yuan
  • Jun You Li
  • Yongzhi Zhang
  • Sonia E. Vader
  • Tongsheng Zhang
  • Gösta Ehnholm
  • Jue Wang
  • Nathan McDannold

We experimentally evaluated the possibility of using an integrated focused ultrasound (FUS) - electroencephalography (EEG) approach to develop activation-informed neuroimaging for human functional brain mapping, enhancing the power of conventional approaches based solely on recording methods. A brief beam of FUS (50 μs, 4.0 MHz) with a focal transverse profile diameter of 0.4 mm determined in water was applied to the barrel cortex of rat in vivo on the dura, skull and scalp and evoked activity was recorded with EEG. We found using a laminar depth profile analysis and specific channel blockers that the stimulation within a safe pressure range could directly evoke a presynaptic spike (∼1 ms) in layer IV, followed by postsynaptic spikes (2.5-7 ms) in layer V and an excitatory wave (8-12 ms), both regulated by fast inhibition, and a major robust excitatory wave (20-30 ms) regulated by slow inhibition. Stimulus artifact was <1 ms, much shorter than for transcranial magnetic stimulation. All components were reliably detected from the scalp, skull, and dura using the 4 MHz and 2.2 MHz transducers. Thus, the FUS-EEG approach is capable of activating neurons transcranially and noninvasively recording a rich spectrum of excitatory and inhibitory activity starting immediately after stimulation. These characteristics can advance neuroimaging since this approach can identify normal electrophysiology and pathophysiology in known locations anywhere in the brain with precise timing and could possibly enhance the conventional approaches by validating and improving their spatial and temporal accuracy of localization.

AAMAS Conference 2026 Conference Paper

HLSMAC: A New StarCraft Multi-Agent Challenge for High-Level Strategic Decision-Making

  • Xingxing Hong
  • Yungong Wang
  • Dexin Jin
  • Ye Yuan
  • Ximing Huang
  • Zijian Wu
  • Yirui Rao
  • Wenxin Li

Benchmarks are crucial for assessing multi-agent reinforcement learning (MARL) algorithms. While StarCraft II-related environments have driven significant advances in MARL, existing benchmarks like SMAC focus primarily on micromanagement, limiting comprehensiveevaluationofhigh-levelstrategicintelligence. Toaddress this, we introduce HLSMAC, a new cooperative MARL benchmark with 12 carefully designed StarCraft II scenarios based on classical stratagems from the Thirty-Six Stratagems. Each scenario corresponds to a specific stratagem and is designed to challenge agents with diverse strategic elements, including tactical maneuvering, timing coordination, and deception, thereby opening up avenues for evaluating high-level strategic decision-making capabilities. We also propose novel metrics across multiple dimensions beyond conventional win rate, such as ability utilization and advancement efficiency, to assess agents’ overall performance within the HLSMAC environment. We conduct a large-scale evaluation of 21 state-of-the-art MARL algorithms and LLM-based agents, with additional multi-seed analysis for relatively better-performing methods. The results demonstrate that HLSMAC serves as a robust testbed for advancing multi-agent strategic decision-making.

TMLR Journal 2026 Journal Article

Offline Model-Based Optimization: Comprehensive Review

  • Minsu Kim
  • Jiayao Gu
  • Ye Yuan
  • Taeyoung Yun
  • Zixuan Liu
  • Yoshua Bengio
  • Can Chen

Offline black-box optimization is a fundamental challenge in science and engineering, where the goal is to optimize black-box functions using only offline datasets. This setting is particularly relevant when querying the objective function is prohibitively expensive or infeasible, with applications spanning protein engineering, material discovery, neural architecture search, and beyond. The main difficulty lies in accurately estimating the objective landscape beyond the available data, where extrapolations are fraught with significant epistemic uncertainty. This uncertainty can lead to objective hacking (reward hacking)—exploiting model inaccuracies in unseen regions—or other spurious optimizations that yield misleadingly high performance estimates outside the offline distribution. Recent advances in model-based optimization (MBO) have harnessed the generalization capabilities of deep neural networks to develop offline-specific surrogate and generative models. Trained with carefully designed strategies, these models are more robust against out-of-distribution issues, facilitating the discovery of improved designs. Despite its growing impact in accelerating scientific discovery, the field lacks a comprehensive review. To bridge this gap, we present the first thorough review of offline MBO. We begin by formalizing the problem for both single-objective and multi-objective settings and by reviewing recent benchmarks and evaluation metrics. We then categorize existing approaches into two key areas: surrogate modeling, which emphasizes accurate function approximation in out-of-distribution regions, and generative modeling, which explores high-dimensional design spaces to identify high-performing designs. Finally, we examine the key challenges and propose promising directions for advancement in this rapidly evolving field including safe control of superintelligent systems.

AAAI Conference 2026 Conference Paper

Safe Multi-Agent Reinforcement Learning via Distributional Safety Critic and Maximum Entropy Optimization

  • Qiwei Liu
  • Ye Yuan
  • Lingyue Zhang
  • Kaitian Chen
  • Yunkai Lv
  • Sheng Gao
  • Huaicheng Yan

Deploying multi-agent reinforcement learning (MARL) in safety-critical systems faces significant challenges due to insufficient agent exploration and inadequate safety constraint guarantees. Current approaches are constrained by two fundamental limitations: inefficient exploration leading to suboptimal policies, and expected-cost-based constraint frameworks failing to ensure full-process safety. To address these challenges, this paper proposes a novel safety-aware maximum entropy MARL framework using Conditional Value-at-Risk (CVaR) as a joint safety metric, which quantifies constraint satisfaction under worst-case scenarios for multi-agent systems. Moreover, we develop the Worst-Case Multi-Agent Soft Actor-Critic (WCMASAC) algorithm, incorporating sequential update mechanisms and maximum entropy optimization for heterogeneous agents, enhanced with distributed safety critics. Theoretically, we establish the monotonic improvement property, guaranteed constraint satisfaction, and convergence to a generalized Nash equilibrium for WCMASAC. Extensive experiments on Safety-Gymnasium based benchmarks demonstrate that WCMASAC outperforms state-of-the-art baselines in both task reward acquisition and safety constraint violation reduction, while exhibiting superior exploration efficiency and risk-aware control capabilities.

AIIM Journal 2026 Journal Article

Topo-UNet: A topology-aware multi-task network for pulmonary vessel segmentation

  • Lu Liu
  • Ye Yuan
  • Yanxin Ma
  • WEI SHAO
  • Jiahe Song
  • Zhe Wang
  • Ruoyu Wang
  • Wenjun Tan

The precise segmentation of pulmonary vessels is crucial for the early diagnosis and treatment of pulmonary diseases. However, vessel images are frequently compromised by high levels of noise and blurred boundaries, which complicate the extraction of vessel features. Current state-of-the-art (SOTA) methods also encounter challenges such as segmenting fine vessels, interruptions in vessel continuity, and loss of inter-layer information. To address these issues, this study proposes a topology-aware multi-task network called Topo-UNet, which integrates the Bidirectional Slice-wise ConvLSTM (BS-ConvLSTM) module and topology-aware auxiliary task to enhance the accurate capture of vessel structural features. The BS-ConvLSTM module mitigates discontinuities in vessel structures by extracting spatial continuity features. Meanwhile, the topology-aware auxiliary task employs a Gaussian function to simulate the intensity distribution within vessels, improving the network's capability to accurately identify vessel structures. Additionally, this study introduces a joint auxiliary task-based method for vessel refinement that increases the recognition rate of fine vessels while enhancing segmentation continuity. Extensive experiments were conducted on CT and CTA datasets to evaluate the performance of Topo-UNet. Comparisons with various SOTA methods across multiple metrics show that Topo-UNet demonstrates superior performance in the task of pulmonary vessel segmentation. Specifically, it achieved Dice coefficients of 90. 78% and 91. 91%, along with Intersection over Union (IoU) scores of 83. 31% and 85. 09% across two test datasets. Furthermore, the discussion section presents a grouping evaluation strategy to address the segmentation performance of vessels of varying sizes, and explores a quadratic approach for vessel refinement, enhancing the segmentation of fine vessels. The code of the proposed Topo-UNet is publicly available at https: //github. com/liu66-git/Topo-UNet.

AAAI Conference 2026 Conference Paper

TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding

  • Canhui Tang
  • Zifan Han
  • Hongbo Sun
  • Sanping Zhou
  • Xuchong Zhang
  • Xin Wei
  • Ye Yuan
  • Huayu Zhang

Multimodal Large Language Models (MLLMs) have demonstrated significant progress in vision-language tasks, yet they still face challenges when processing long-duration video inputs. The limitation arises from MLLMs' context limit and training costs, necessitating sparse frame sampling before feeding videos into MLLMs. However, building a trainable sampling method remains challenging due to the unsupervised and non-differentiable nature of sparse frame sampling in Video-MLLMs. To address these problems, we propose Temporal Sampling Policy Optimization (**TSPO**), advancing MLLMs' long-form video-language understanding via reinforcement learning. Specifically, we first propose a trainable event-aware temporal agent, which captures event-query correlation for performing probabilistic keyframe selection. Then, we propose the TSPO reinforcement learning paradigm, which models keyframe selection and language generation as a joint decision-making process, enabling end-to-end group relative optimization for the temporal sampling policy. Furthermore, we propose a dual-style long video training data construction pipeline, balancing comprehensive temporal understanding and key segment localization. Finally, we incorporate rule-based answering accuracy and temporal locating reward mechanisms to optimize the temporal sampling policy. Comprehensive experiments show that our TSPO achieves state-of-the-art performance across multiple long video understanding benchmarks, and shows transferable ability across different cutting-edge Video-MLLMs.

IROS Conference 2025 Conference Paper

Awakening Facial Emotional Expressions in Human-Robot

  • Yongtong Zhu
  • Lei Li
  • Iggy Qian
  • Wenbin Zhou
  • Ye Yuan
  • Qingdu Li
  • Na Liu 0007
  • Jianwei Zhang 0001

The facial expression generation capability of humanoid social robots is critical for achieving natural and human-like interactions, playing a vital role in enhancing the fluidity of human-robot interactions and the accuracy of emotional expression. Currently, facial expression generation in humanoid social robots still relies on pre-programmed be-havioral patterns, which are manually coded at high human and time costs. To enable humanoid robots to autonomously acquire generalized expressive capabilities, they need to develop the ability to learn human-like expressions through self-training. To address this challenge, we have designed a highly biomimetic robotic face with physical-electronic animated facial units and developed an end-to-end learning framework based on KAN (Kolmogorov-Arnold Network) and attention mechanisms. Unlike previous humanoid social robots, we have also meticulously designed an automated data collection system based on expert strategies of facial motion primitives to construct the dataset. Notably, to the best of our knowledge, this is the first open-source facial dataset for humanoid social robots. Comprehensive evaluations indicate that our approach achieves accurate and diverse facial mimicry across different test subjects.

TMLR Journal 2025 Journal Article

Design Editing for Offline Model-based Optimization

  • Ye Yuan
  • Youyuan Zhang
  • Can Chen
  • Haolun Wu
  • Melody Zixuan Li
  • Jianmo Li
  • James J. Clark
  • Xue Liu

Offline model-based optimization (MBO) aims to maximize a black-box objective function using only an offline dataset of designs and scores. These tasks span various domains, such as robotics, material design, and protein and molecular engineering. A common approach involves training a surrogate model using existing designs and their corresponding scores, and then generating new designs through gradient-based updates with respect to the surrogate model. This method suffers from the out-of-distribution issue, where the surrogate model may erroneously predict high scores for unseen designs. To address this challenge, we introduce a novel method, Design Editing for Offline Model-based Optimization} (DEMO), which leverages a diffusion prior to calibrate overly optimized designs. DEMO first generates pseudo design candidates by performing gradient ascent with respect to a surrogate model. While these pseudo design candidates contain information beyond the offline dataset, they might be invalid or have erroneously high predicted scores. Therefore, to address this challenge while utilizing the information provided by pseudo design candidates, we propose an editing process to refine these pseudo design candidates. We introduce noise to the pseudo design candidates and subsequently denoise them with a diffusion prior trained on the offline dataset, ensuring they align with the distribution of valid designs. Empirical evaluations on seven offline MBO tasks show that, with properly tuned hyperparamters, DEMO's score is competitive with the best previously reported scores in the literature.

TMLR Journal 2025 Journal Article

Efficient Diffusion Models: A Survey

  • Hui Shen
  • Jingxuan Zhang
  • Boning Xiong
  • Rui Hu
  • Shoufa Chen
  • Zhongwei Wan
  • Xin Wang
  • Yu Zhang

Diffusion models have emerged as powerful generative models capable of producing high-quality contents such as images, videos, and audio, demonstrating their potential to revolutionize digital content creation. However, these capabilities come at the cost of significant computational resources and lengthy generation time, underscoring the critical need to develop efficient techniques for practical deployment. In this survey, we provide a systematic and comprehensive review of research on efficient diffusion models. We organize the literature in a taxonomy consisting of three main categories, covering distinct yet interconnected efficient diffusion model topics from algorithm-level, system-level, and framework perspective, respectively. We have also created a GitHub repository where we organize the papers featured in this survey at github.com/AIoT-MLSys-Lab/Efficient-Diffusion-Model-Survey. We hope our survey can serve as a valuable resource to help researchers and practitioners gain a systematic understanding of efficient diffusion model research and inspire them to contribute to this important and exciting field.

IROS Conference 2025 Conference Paper

EmoRLTalk: Speech-Driven Emotional Facial Animation With Offline Reinforcement Learning

  • Gaofeng Liu
  • Xuetong Li
  • Ruoyu Gao
  • Ye Yuan
  • Jian Liu
  • Hengsen Li
  • Hong Huo
  • Tao Fang

In recent years, significant breakthroughs have been made in audio-guided 3D facial animation. However, existing methods mainly focus on lip shape and audio consistency and still face key challenges to achieve alignment between facial emotions and speech emotions. To overcome this limitation, we introduce EmoRLTalk, a novel framework that integrates offline reinforcement learning to implicitly capture the intricate relationship between 3D facial landmarks and blendshape parameters, thereby enhancing the granularity of emotional expression. Furthermore, we harness the strength of conditional diffusion models to synthesize facial motions that are emotionally coherent with the input speech. Additionally, based on the multi-task learning paradigm, we construct a collaborative training framework of a regression main task and a classification sub-task. Specifically, we use emotion classification of blendshape as a sub-task to further improve the model’s ability to express facial emotions. To further enhance system controllability, we integrate the ControlNet module, allowing users to achieve precise facial expression control. Extensive experiments demonstrate that EmoRLTalk achieves superior emotional expressiveness and lip-sync performance compared to previous approaches.

NeurIPS Conference 2025 Conference Paper

Fira: Can We Achieve Full-rank Training of LLMs Under Low-rank Constraint?

  • Xi Chen
  • Kaituo Feng
  • Changsheng Li
  • Xunhao Lai
  • Xiangyu Yue
  • Ye Yuan
  • Guoren Wang

Low-rank training has emerged as a promising approach for reducing memory usage in training Large Language Models (LLMs). Previous methods either rely on decomposing weight matrices (e. g. , LoRA), or seek to decompose gradient matrices (e. g. , GaLore) to ensure reduced memory consumption. However, both of them constrain the training in a low-rank subspace, thus inevitably leading to sub-optimal performance. To resolve this, we propose a new plug-and-play training framework for LLMs called Fira, as the first attempt to consistently preserve the low-rank constraint for memory efficiency, while achieving full-rank training (i. e. , training with full-rank gradients of full-rank weights) to avoid inferior outcomes. First, we observe an interesting phenomenon during LLM training: the scaling impact of adaptive optimizers (e. g. , Adam) on the gradient norm remains similar from low-rank to full-rank training. In light of this, we propose a \textit{norm-based scaling} method, which utilizes the scaling impact of low-rank optimizers as substitutes for that of original full-rank optimizers to achieve this goal. Moreover, we find that there are potential loss spikes during training. To address this, we further put forward a norm-growth limiter to smooth the gradient. Extensive experiments on the pre-training and fine-tuning of LLMs show that Fira outperforms both LoRA and GaLore. Notably, for pre-training LLaMA 7B, our Fira uses $8\times$ smaller memory of optimizer states than Galore, yet outperforms it by a large margin.

NeurIPS Conference 2025 Conference Paper

Generalization Bounds for Kolmogorov-Arnold Networks (KANs) and Enhanced KANs with Lower Lipschitz Complexity

  • Pengqi Li
  • Lizhong Ding
  • Jiarun Fu
  • Chunhui Zhang
  • Guoren Wang
  • Ye Yuan

Kolmogorov-Arnold Networks (KANs) have demonstrated remarkable expressive capacity and predictive power in symbolic learning. However, existing generalization errors of KANs primarily focus on approximation errors while neglecting estimation errors, leading to a suboptimal bias-variance trade-off and poor generalization performance. Meanwhile, the unclear generalization mechanism hinders the design of more effective KANs variants. As the authors of KANs highlighted, they ``would like to explore ways to restrict KANs' hypothesis space so that they can achieve good performance''. To address these challenges, we explore the generalization mechanism of KANs and design more effective KANs with lower model complexity and better generalization. We define \textit{Lipschitz complexity} as the first structural measure for deep functions represented by KANs and derive novel generalization bounds based on \textit{Lipschitz complexity}, establishing a theoretical foundation for understanding their generalization behavior. To reduce \textit{Lipschitz complexity} and boost the generalization mechanism of KANs, we propose Lipschitz-Enhanced KANs ($\textbf{LipKANs}$) by integrating the Lip layer and pioneering the $L_{1. 5}$-regularized loss, contributing to tighter generalization bounds. Empirical experiments validate that the proposed LipKANs enhance the generalization mechanism of KANs when modeling complex distributions. We hope our theoretical bounds and LipKANs lay a foundation for the future development of KANs.

NeurIPS Conference 2025 Conference Paper

Handling Label Noise via Instance-Level Difficulty Modeling and Dynamic Optimization

  • Kuan Zhang
  • Chengliang Chai
  • Jingzhe Xu
  • Chi Zhang
  • Han Han
  • Ye Yuan
  • Guoren Wang
  • Lei Cao

Recent studies indicate that deep neural networks degrade in generalization performance under noisy supervision. Existing methods focus on isolating clean subsets or correcting noisy labels, facing limitations such as high computational costs, heavy hyperparameter tuning process, and coarse-grained optimization. To address these challenges, we propose a novel two-stage noisy learning framework that enables instance-level optimization through a dynamically weighted loss function, avoiding hyperparameter tuning. To obtain stable and accurate information about noise modeling, we introduce a simple yet effective metric, termed $\textit{wrong event}$, which dynamically models the cleanliness and difficulty of individual samples while maintaining computational costs. Our framework first collects $\textit{wrong event}$ information and builds a strong base model. Then we perform noise-robust training on the base model, using a probabilistic model to handle the $\textit{wrong event}$ information of samples. Experiments on six synthetic and real-world LNL benchmarks demonstrate our method surpasses state-of-the-art methods in performance, achieves a nearly 75\% reduction in storage and computational time, strongly improving model scalability. Our code is available at https: //github. com/iTheresaApocalypse/IDO.

ECAI Conference 2025 Conference Paper

M-Rec: Leveraging Multi-View Feature to Reconstruction Incomplete Localization for Image Manipulation Detection

  • Sifan Wang
  • Yixun Zhang
  • Ye Yuan

Image manipulation detection (IMD) serves as a critical technique for identifying forged images, playing a significant role in safeguarding cyberspace security. In this paper, we focus on a challenge that has been overlooked by previous work: incomplete localization. This issue means the network can identify potential tampered regions but lacks the confidence to decisively classify them as forgeries, resulting in only partial detection of tampered regions. When the detected regions and the undetected regions (due to a lack of confidence) differ significantly in semantics, it may mislead observers into believing that the currently localized regions represent the complete tampered regions. To address this issue, we propose M-Rec, a network designed to correct incomplete localization through a reconstruction strategy. Specifically, we introduce a confidence decoder to identify incomplete localization, and a reconstruction decoder to correct the prediction outcomes. Meanwhile, we design a bidirectional filtering strategy: (a) In the training stage, the proposed strategy is used to remove incorrect labels, ensuring that the reconstruction decoder is supervised with reliable ground truth; (b) In the inference stage, the proposed strategy constrains the reconstruction scope to the network prediction. Extensive experiments on four datasets demonstrate that our method effectively corrects incomplete localization and achieves state-of-the-art performance in manipulation detection.

AAAI Conference 2025 Conference Paper

Prompt Tuning In a Compact Attribute Space

  • Shiyu Hou
  • Tianfei Zhou
  • Shuai Zhang
  • Ye Yuan
  • Guoren Wang

Prompt tuning (PT) has emerged as a key to unlocking the power of visual-language models like CLIP for various downstream tasks. Predominant approaches learn a small set of task-relevant soft prompts by solving an image-class matching problem. Nevertheless, by optimizing merely with respect to class names, they face challenges in learning high performant prompts capable of capturing fine-grained, diverse characteristics of each class, and tends to overfit potentially biased distribution of base classes. In this work, we propose PTinCAS to tackle prompt tuning in a compact attribute space, driven by the premise that attributes offer detailed class interpretations and can facilitate transfer across related categories. Particularly, PTinCAS is grounded in two innovative designs. First, we create a compact attribute space by properly prompting large language models to generate factual descriptions about categories, which are subsequently clustered to form a concise attribute vocabulary. Second, we leverage attributes as a source of supervision in PT to transfer the inherent common sense knowledge in attributes to soft prompts. An object-aware visual prompting mechanism is developed to effortlessly highlight intended regions in the original image, which guides the model towards learning visual attributes associated with object regions rather than the background. We show that PTinCAS not only improves few-shot generalizability compared to existing PT methods, but also provides some level of inherent explainability that helps us understand why a class name is determined based on the attributes activated in an image.

AAAI Conference 2025 Conference Paper

Real-Time Neural Denoising with Render-Aware Knowledge Distillation

  • Mengxun Kong
  • Jie Guo
  • Chen Wang
  • Ye Yuan
  • Yanwen Guo

Real-time Monte Carlo (MC) ray tracing with low sampling rates demands a denoising algorithm that adeptly balances the trade-off between quality and efficiency. Previous works have paid much attention on designing delicate denoising architecture while ignoring model compression. In this work, we present a render-aware knowledge distillation (RAKD) framework, specifically designed for Monte Carlo denoising. We meticulously delineate the Knowledge Distillation (KD) process within RAKD, emphasizing three pivotal techniques: the strategic incorporation of an auxiliary unlabeled dataset, the integration of adversarial learning through generative adversarial network (GAN), and the application of parameter transfer for robust model initialization. These approaches are harmoniously combined to distill knowledge effectively, enabling our student model to adeptly strike a balance between preserving high-frequency details and reducing low-frequency noise. Finally, our results demonstrate that RAKD achieves state-of-the-art quality while upholding real-time performance, successfully tackling the computational constraints faced by resource-limited devices.

NeurIPS Conference 2025 Conference Paper

Towards Multi-Table Learning: A Novel Paradigm for Complementarity Quantification and Integration

  • Zhang Junyu
  • Lizhong Ding
  • Ye Yuan
  • Xingcan Li
  • Pengqi Li
  • Tihang Xi
  • Guoren Wang
  • Changsheng Li

Multi-table data integrate various entities and attributes, with potential interconnections between them. However, existing tabular learning methods often struggle to describe and leverage the underlying complementarity across distinct tables. To address this limitation, we propose the first unified paradigm for multi-table learning that systematically quantifies and integrates complementary information across tables. Specifically, we introduce a metric called complementarity strength (CS), which captures inter-table complementarity by incorporating relevance, similarity, and informativeness. For the first time, we systematically formulate the paradigm towards multi-table learning by establishing formal definitions of tasks and loss functions. Correspondingly, we present a network for multi-table learning that combines Adaptive Table encoder and Cross table Attention mechanism (ATCA-Net), achieving the simultaneous integration of complementary information from distinct tables. Extensive experiments show that ATCA-Net effectively leverages complementary information and that the CS metric accurately quantifies the richness of complementarity across multiple tables. To the best of our knowledge, this is the first work to establish theoretical and practical foundations for multi-table learning.

NeurIPS Conference 2025 Conference Paper

VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning

  • Qi Wang
  • Yanrui Yu
  • Ye Yuan
  • Rui Mao
  • Tianfei Zhou

Reinforcement fine-tuning (RFT) has shown great promise in achieving humanlevel reasoning capabilities of Large Language Models (LLMs), and has recently been extended to MLLMs. Nevertheless, reasoning about videos, which is a fundamental aspect of human intelligence, remains a persistent challenge due to the complex logic, temporal and causal structures inherent in video data. To fill this gap, we propose VideoRFT, a novel approach that extends the RFT paradigm to cultivate human-like video reasoning capabilities in MLLMs. VideoRFT follows the standard two-stage scheme in RFT: supervised fine-tuning (SFT) with chain-of-thought (CoT) annotations, followed by reinforcement learning (RL) to improve generalization. A central challenge to achieve this in the video domain lies in the scarcity of large-scale, high-quality video CoT datasets. We address this by building a multi-expert-driven, cognition-inspired CoT curation pipeline. First, we devise a cognition-inspired prompting strategy to elicit a reasoning LLM to generate preliminary CoTs based solely on rich, structured, and literal representations of video content. Subsequently, these CoTs are revised by a MLLM conditioned on the actual video, ensuring visual consistency and reducing visual hallucinations. This pipeline results in two new datasets, i. e. VideoRFT-CoT-102K for SFT and VideoRFT-RL-310K for RL. To further strengthen the RL phase, we introduce a novel semantic-consistency reward that explicitly promotes the alignment between textual reasoning and visual evidence. This reward encourages the model to produce coherent, context-aware reasoning outputs grounded in visual input. Extensive experiments show that VideoRFT achieves state-of-the-art performance on six video reasoning benchmarks.

NeurIPS Conference 2024 Conference Paper

A robust inlier identification algorithm for point cloud registration via $\mathbf{\ell_0}$-minimization

  • Yinuo Jiang
  • Xiuchuan Tang
  • Cheng Cheng
  • Ye Yuan

Correspondences in point cloud registration are prone to outliers, significantly reducing registration accuracy and highlighting the need for precise inlier identification. In this paper, we propose a robust inlier identification algorithm for point cloud registration by reformulating the conventional registration problem as an alignment error $\ell_0$-minimization problem. The $\ell_0$-minimization problem is formulated for each local set, where those local sets are built on a compatibility graph of input correspondences. To resolve the $\ell_0$-minimization, we develop a novel two-stage decoupling strategy, which first decouples the alignment error into a rotation fitting error and a translation fitting error. Second, null-space matrices are employed to decouple inlier identification from the estimation of rotation and translation respectively, thereby applying Bayesian theory to $\ell_0$-minimization problems and solving for fitting errors. Correspondences with the smallest errors are identified as inliers to generate a transformation hypothesis for each local set. The best hypothesis is selected to perform registration. We demonstrate that the proposed inlier identification algorithm is robust under high outlier ratios and noise through experiments. Extensive results on the KITTI, 3DMatch, and 3DLoMatch datasets demonstrate that our method achieves state-of-the-art performance compared to both traditional and learning-based methods in various indoor and outdoor scenes.

TMLR Journal 2024 Journal Article

AGG: Amortized Generative 3D Gaussians for Single Image to 3D

  • Dejia Xu
  • Ye Yuan
  • Morteza Mardani
  • Sifei Liu
  • Jiaming Song
  • Zhangyang Wang
  • Arash Vahdat

Given the growing need for automatic 3D content creation pipelines, various 3D representations have been studied to generate 3D objects from a single image. Due to its superior rendering efficiency, 3D Gaussian splatting-based models have recently excelled in both 3D reconstruction and generation. 3D Gaussian splatting approaches for image to 3D generation are often optimization-based, requiring many computationally expensive score-distillation steps. To overcome these challenges, we introduce an Amortized Generative 3D Gaussian framework (AGG) that instantly produces 3D Gaussians from a single image, eliminating the need for per-instance optimization. Utilizing an intermediate hybrid representation, AGG decomposes the generation of 3D Gaussian locations and other appearance attributes for joint optimization. Moreover, we propose a cascaded pipeline that first generates a coarse representation of the 3D data and later upsamples it with a 3D Gaussian super-resolution module. Our method is evaluated against existing sampling-based 3D Gaussian frameworks and inference-based pipelines utilizing other 3D representations, where AGG showcases competitive generation abilities both qualitatively and quantitatively while being several orders of magnitude faster.

JMLR Journal 2024 Journal Article

Almost Sure Convergence Rates Analysis and Saddle Avoidance of Stochastic Gradient Methods

  • Jun Liu
  • Ye Yuan

The vast majority of convergence rates analysis for stochastic gradient methods in the literature focus on convergence in expectation, whereas trajectory-wise almost sure convergence is clearly important to ensure that any instantiation of the stochastic algorithms would converge with probability one. Here we provide a unified almost sure convergence rates analysis for stochastic gradient descent (SGD), stochastic heavy-ball (SHB), and stochastic Nesterov's accelerated gradient (SNAG) methods. We show, for the first time, that the almost sure convergence rates obtained for these stochastic gradient methods on strongly convex functions, are arbitrarily close to their optimal convergence rates possible. For non-convex objective functions, we not only show that a weighted average of the squared gradient norms converges to zero almost surely, but also the last iterates of the algorithms. We further provide last-iterate almost sure convergence rates analysis for stochastic gradient methods on general convex smooth functions, in contrast with most existing results in the literature that only provide convergence in expectation for a weighted average of the iterates. The last-iterate almost sure convergence results also enable us to obtain almost sure avoidance of any strict saddle manifold by stochastic gradient methods with or without momentum. To the best of our knowledge, this is the first time such results are obtained for SHB and SNAG methods. [abs] [ pdf ][ bib ] &copy JMLR 2024. ( edit, beta )

NeurIPS Conference 2024 Conference Paper

LaKD: Length-agnostic Knowledge Distillation for Trajectory Prediction with Any Length Observations

  • Yuhang Li
  • Changsheng Li
  • Ruilin Lv
  • Rongqing Li
  • Ye Yuan
  • Guoren Wang

Trajectory prediction is a crucial technology to help systems avoid traffic accidents, ensuring safe autonomous driving. Previous methods typically use a fixed-length and sufficiently long trajectory of an agent as observations to predict its future trajectory. However, in real-world scenarios, we often lack the time to gather enough trajectory points before making predictions, e. g. , when a car suddenly appears due to an obstruction, the system must make immediate predictions to prevent a collision. This poses a new challenge for trajectory prediction systems, requiring them to be capable of making accurate predictions based on observed trajectories of arbitrary lengths, leading to the failure of existing methods. In this paper, we propose a Length-agnostic Knowledge Distillation framework, named LaKD, which can make accurate trajectory predictions, regardless of the length of observed data. Specifically, considering the fact that long trajectories, containing richer temporal information but potentially additional interference, may perform better or worse than short trajectories, we devise a dynamic length-agnostic knowledge distillation mechanism for exchanging information among trajectories of arbitrary lengths, dynamically determining the transfer direction based on prediction performance. In contrast to traditional knowledge distillation, LaKD employs a unique model that simultaneously serves as both the teacher and the student, potentially causing knowledge collision during the distillation process. Therefore, we design a dynamic soft-masking mechanism, where we first calculate the importance of neuron units and then apply soft-masking to them, so as to safeguard critical units from disruption during the knowledge distillation process. In essence, LaKD is a general and principled framework that can be naturally compatible with existing trajectory prediction models of different architectures. Extensive experiments on three benchmark datasets, Argoverse 1, nuScenes and Argoverse 2, demonstrate the effectiveness of our approach.

AAAI Conference 2024 Conference Paper

Preparing Lessons for Progressive Training on Language Models

  • Yu Pan
  • Ye Yuan
  • Yichun Yin
  • Jiaxin Shi
  • Zenglin Xu
  • Ming Zhang
  • Lifeng Shang
  • Xin Jiang

The rapid progress of Transformers in artificial intelligence has come at the cost of increased resource consumption and greenhouse gas emissions due to growing model sizes. Prior work suggests using pretrained small models to improve training efficiency, but this approach may not be suitable for new model structures. On the other hand, training from scratch can be slow, and progressively stacking layers often fails to achieve significant acceleration. To address these challenges, we propose a novel method called Apollo, which prepares lessons for expanding operations by learning high-layer functionality during training of low layers. Our approach involves low-value-prioritized sampling (LVPS) to train different depths and weight sharing to facilitate efficient expansion. We also introduce an interpolation method for stable model depth extension. Experiments demonstrate that Apollo achieves state-of-the-art acceleration ratios, even rivaling methods using pretrained models, making it a universal and efficient solution for training deep models while reducing time, financial, and environmental costs.

NeurIPS Conference 2023 Conference Paper

BCDiff: Bidirectional Consistent Diffusion for Instantaneous Trajectory Prediction

  • Rongqing Li
  • Changsheng Li
  • Dongchun Ren
  • Guangyi Chen
  • Ye Yuan
  • Guoren Wang

The objective of pedestrian trajectory prediction is to estimate the future paths of pedestrians by leveraging historical observations, which plays a vital role in ensuring the safety of self-driving vehicles and navigation robots. Previous works usually rely on a sufficient amount of observation time to accurately predict future trajectories. However, there are many real-world situations where the model lacks sufficient time to observe, such as when pedestrians abruptly emerge from blind spots, resulting in inaccurate predictions and even safety risks. Therefore, it is necessary to perform trajectory prediction based on instantaneous observations, which has rarely been studied before. In this paper, we propose a Bi-directional Consistent Diffusion framework tailored for instantaneous trajectory prediction, named BCDiff. At its heart, we develop two coupled diffusion models by designing a mutual guidance mechanism which can bidirectionally and consistently generate unobserved historical trajectories and future trajectories step-by-step, to utilize the complementary information between them. Specifically, at each step, the predicted unobserved historical trajectories and limited observed trajectories guide one diffusion model to generate future trajectories, while the predicted future trajectories and observed trajectories guide the other diffusion model to predict unobserved historical trajectories. Given the presence of relatively high noise in the generated trajectories during the initial steps, we introduce a gating mechanism to learn the weights between the predicted trajectories and the limited observed trajectories for automatically balancing their contributions. By means of this iterative and mutually guided generation process, both the future and unobserved historical trajectories undergo continuous refinement, ultimately leading to accurate predictions. Essentially, BCDiff is an encoder-free framework that can be compatible with existing trajectory prediction models in principle. Experiments show that our proposed BCDiff significantly improves the accuracy of instantaneous trajectory prediction on the ETH/UCY and Stanford Drone datasets, compared to related approaches.

NeurIPS Conference 2023 Conference Paper

Importance-aware Co-teaching for Offline Model-based Optimization

  • Ye Yuan
  • Can (Sam) Chen
  • Zixuan Liu
  • Willie Neiswanger
  • Xue (Steve) Liu

Offline model-based optimization aims to find a design that maximizes a property of interest using only an offline dataset, with applications in robot, protein, and molecule design, among others. A prevalent approach is gradient ascent, where a proxy model is trained on the offline dataset and then used to optimize the design. This method suffers from an out-of-distribution issue, where the proxy is not accurate for unseen designs. To mitigate this issue, we explore using a pseudo-labeler to generate valuable data for fine-tuning the proxy. Specifically, we propose $\textit{\textbf{I}mportance-aware \textbf{C}o-\textbf{T}eaching for Offline Model-based Optimization}~(\textbf{ICT})$. This method maintains three symmetric proxies with their mean ensemble as the final proxy, and comprises two steps. The first step is $\textit{pseudo-label-driven co-teaching}$. In this step, one proxy is iteratively selected as the pseudo-labeler for designs near the current optimization point, generating pseudo-labeled data. Subsequently, a co-teaching process identifies small-loss samples as valuable data and exchanges them between the other two proxies for fine-tuning, promoting knowledge transfer. This procedure is repeated three times, with a different proxy chosen as the pseudo-labeler each time, ultimately enhancing the ensemble performance. To further improve accuracy of pseudo-labels, we perform a secondary step of $\textit{meta-learning-based sample reweighting}$, which assigns importance weights to samples in the pseudo-labeled dataset and updates them via meta-learning. ICT achieves state-of-the-art results across multiple design-bench tasks, achieving the best mean rank $3. 1$ and median rank $2$ among $15$ methods. Our source code can be accessed here.

ICML Conference 2023 Conference Paper

NeRFool: Uncovering the Vulnerability of Generalizable Neural Radiance Fields against Adversarial Perturbations

  • Yonggan Fu
  • Ye Yuan
  • Souvik Kundu 0009
  • Shang Wu 0003
  • Shunyao Zhang
  • Yingyan Celine Lin

Generalizable Neural Radiance Fields (GNeRF) are one of the most promising real-world solutions for novel view synthesis, thanks to their cross-scene generalization capability and thus the possibility of instant rendering on new scenes. While adversarial robustness is essential for real-world applications, little study has been devoted to understanding its implication on GNeRF. We hypothesize that because GNeRF is implemented by conditioning on the source views from new scenes, which are often acquired from the Internet or third-party providers, there are potential new security concerns regarding its real-world applications. Meanwhile, existing understanding and solutions for neural networks’ adversarial robustness may not be applicable to GNeRF, due to its 3D nature and uniquely diverse operations. To this end, we present NeRFool, which to the best of our knowledge is the first work that sets out to understand the adversarial robustness of GNeRF. Specifically, NeRFool unveils the vulnerability patterns and important insights regarding GNeRF’s adversarial robustness. Built upon the above insights gained from NeRFool, we further develop NeRFool$^+$, which integrates two techniques capable of effectively attacking GNeRF across a wide range of target views, and provide guidelines for defending against our proposed attacks. We believe that our NeRFool/NeRFool$^+$ lays the initial foundation for future innovations in developing robust real-world GNeRF solutions. Our codes are available at: https: //github. com/GATECH-EIC/NeRFool.

NeurIPS Conference 2023 Conference Paper

Reusing Pretrained Models by Multi-linear Operators for Efficient Training

  • Yu Pan
  • Ye Yuan
  • Yichun Yin
  • Zenglin Xu
  • Lifeng Shang
  • Xin Jiang
  • Qun Liu

Training large models from scratch usually costs a substantial amount of resources. Towards this problem, recent studies such as bert2BERT and LiGO have reused small pretrained models to initialize a large model (termed the ``target model''), leading to a considerable acceleration in training. Despite the successes of these previous studies, they grew pretrained models by mapping partial weights only, ignoring potential correlations across the entire model. As we show in this paper, there are inter- and intra-interactions among the weights of both the pretrained and the target models. As a result, the partial mapping may not capture the complete information and lead to inadequate growth. In this paper, we propose a method that linearly correlates each weight of the target model to all the weights of the pretrained model to further enhance acceleration ability. We utilize multi-linear operators to reduce computational and spacial complexity, enabling acceptable resource requirements. Experiments demonstrate that our method can save 76\% computational costs on DeiT-base transferred from DeiT-small, which outperforms bert2BERT by +12\% and LiGO by +21\%, respectively.

ICRA Conference 2023 Conference Paper

RGB-Only Reconstruction of Tabletop Scenes for Collision-Free Manipulator Control

  • Zhenggang Tang
  • Balakumar Sundaralingam
  • Jonathan Tremblay
  • Bowen Wen
  • Ye Yuan
  • Stephen Tyree
  • Charles T. Loop
  • Alexander G. Schwing

We present a system for collision-free control of a robot manipulator that uses only RGB views of the world. Perceptual input of a tabletop scene is provided by multiple images of an RGB camera (without depth) that is either handheld or mounted on the robot end effector. A NeRF-like process is used to reconstruct the 3D geometry of the scene, from which the Euclidean full signed distance function (ESDF) is computed. A model predictive control algorithm is then used to control the manipulator to reach a desired pose while avoiding obstacles in the ESDF. We show results on a real dataset collected and annotated in our lab. Our results are also available at https://ngp-mpc.github.io/.

AAAI Conference 2022 Conference Paper

Boosting Contrastive Learning with Relation Knowledge Distillation

  • Kai Zheng
  • Yuanjiang Wang
  • Ye Yuan

While self-supervised representation learning (SSL) has proved to be effective in the large model, there is still a huge gap between the SSL and supervised method in the lightweight model when following the same solution. We delve into this problem and find that the lightweight model is prone to collapse in semantic space when simply performing instance-wise contrast. To address this issue, we propose a relation-wise contrastive paradigm with Relation Knowledge Distillation (ReKD). We introduce a heterogeneous teacher to explicitly mine the semantic information and transferring a novel relation knowledge to the student (lightweight model). The theoretical analysis supports our main concern about instance-wise contrast and verify the effectiveness of our relation-wise contrastive learning. Extensive experimental results also demonstrate that our method achieves significant improvements on multiple lightweight models. Particularly, the linear evaluation on AlexNet obviously improves the current state-of-art from 44. 7% to 50. 1%, which is the first work to get close to the supervised (50. 5%). Code will be made available.

NeurIPS Conference 2022 Conference Paper

Embodied Scene-aware Human Pose Estimation

  • Zhengyi Luo
  • Shun Iwase
  • Ye Yuan
  • Kris Kitani

We propose embodied scene-aware human pose estimation where we estimate 3D poses based on a simulated agent's proprioception and scene awareness, along with external third-person observations. Unlike prior methods that often resort to multistage optimization, non-causal inference, and complex contact modeling to estimate human pose and human scene interactions, our method is one-stage, causal, and recovers global 3D human poses in a simulated environment. Since 2D third-person observations are coupled with the camera pose, we propose to disentangle the camera pose and use a multi-step projection gradient defined in the global coordinate frame as the movement cue for our embodied agent. Leveraging a physics simulation and prescanned scenes (e. g. , 3D mesh), we simulate our agent in everyday environments (library, office, bedroom, etc. ) and equip our agent with environmental sensors to intelligently navigate and interact with the geometries of the scene. Our method also relies only on 2D keypoints and can be trained on synthetic datasets derived from popular human motion databases. To evaluate, we use the popular H36M and PROX datasets and achieve high quality pose estimation on the challenging PROX dataset without ever using PROX motion sequences for training. Code and videos are available on the project page.

JBHI Journal 2022 Journal Article

Noise-Immune Extreme Ensemble Learning for Early Diagnosis of Neuropsychiatric Systemic Lupus Erythematosus

  • Ye Yuan
  • Tianhong Quan
  • Youyi Song
  • Jitian Guan
  • Teng Zhou
  • Renhua Wu

Early diagnosis is currently the most effective way of saving the life of patients with neuropsychiatric systemic lupus erythematosus (NPSLE). However, it is rather difficult to detect this terrible disease at the early stage, due to the subtle and elusive symptomatic signals. Recent studies show that the $^{1}$ H-MRS (proton magnetic resonance spectroscopy) imaging technique can capture more information reflecting the early appearance of this disease than conventional magnetic resonance imaging techniques. $^{1}$ H-MRS data, however, also presents more noises that can bring serious diagnosis bias. We hence proposed a noise-immune extreme ensemble learning technique for effectively leveraging $^{1}$ H-MRS data for advancing the early diagnosis of NPSLE. Our main results are that 1) by developing generalized maximum correntropy criterion in the kernel extreme learning setting, many types of non-Gaussian noises can be distinguished, and 2) weighted recursive feature elimination, using maximal information coefficient to weight feature’s importance, helps to further alleviate the bad impact of noises on the diagnosis performance. The proposed method is assessed on a publicly available dataset with 97. 5% accuracy, 95. 8% sensitivity and 99. 9% specificity, which well demonstrates its efficacy.

TIST Journal 2021 Journal Article

An Uncertainty-based Neural Network for Explainable Trajectory Segmentation

  • Xin Bi
  • Chao Zhang
  • Fangtong Wang
  • Zhixun Liu
  • Xiangguo Zhao
  • Ye Yuan
  • Guoren Wang

As a variant task of time-series segmentation, trajectory segmentation is a key task in the applications of transportation pattern recognition and traffic analysis. However, segmenting trajectory is faced with challenges of implicit patterns and sparse results. Although deep neural networks have tremendous advantages in terms of high-level feature learning performance, deploying as a blackbox seriously limits the real-world applications. Providing explainable segmentations has significance for result evaluation and decision making. Thus, in this article, we address trajectory segmentation by proposing a Bayesian Encoder-Decoder Network (BED-Net) to provide accurate detection with explainability and references for the following active-learning procedures. BED-Net consists of a segmentation module based on Monte Carlo dropout and an explanation module based on uncertainty learning that provides results evaluation and visualization. Experimental results on both benchmark and real-world datasets indicate that BED-Net outperforms the rival methods and offers excellent explainability in the applications of trajectory segmentation.

AAAI Conference 2021 Conference Paper

AnchorFace: An Anchor-based Facial Landmark Detector Across Large Poses

  • Zixuan Xu
  • Banghuai Li
  • Ye Yuan
  • Miao Geng

Facial landmark localization aims to detect the predefined points of human faces, and the topic has been rapidly improved with the recent development of neural network based methods. However, it remains a challenging task when dealing with faces in unconstrained scenarios, especially with large pose variations. In this paper, we target the problem of facial landmark localization across large poses and address this task based on a split-and-aggregate strategy. To split the search space, we propose a set of anchor templates as references for regression, which well addresses the large variations of face poses. Based on the prediction of each anchor template, we propose to aggregate the results, which can reduce the landmark uncertainty due to the large poses. Overall, our proposed approach, named AnchorFace, obtains stateof-the-art results with extremely efficient inference speed on four challenging benchmarks, i. e. AFLW, 300W, Menpo, and WFLW dataset. Code will be available soon.

NeurIPS Conference 2021 Conference Paper

Dynamics-regulated kinematic policy for egocentric pose estimation

  • Zhengyi Luo
  • Ryo Hachiuma
  • Ye Yuan
  • Kris Kitani

We propose a method for object-aware 3D egocentric pose estimation that tightly integrates kinematics modeling, dynamics modeling, and scene object information. Unlike prior kinematics or dynamics-based approaches where the two components are used disjointly, we synergize the two approaches via dynamics-regulated training. At each timestep, a kinematic model is used to provide a target pose using video evidence and simulation state. Then, a prelearned dynamics model attempts to mimic the kinematic pose in a physics simulator. By comparing the pose instructed by the kinematic model against the pose generated by the dynamics model, we can use their misalignment to further improve the kinematic model. By factoring in the 6DoF pose of objects (e. g. , chairs, boxes) in the scene, we demonstrate for the first time, the ability to estimate physically-plausible 3D human-object interactions using a single wearable camera. We evaluate our egocentric pose estimation method in both controlled laboratory settings and real-world scenarios.

AAAI Conference 2021 Conference Paper

LRSC: Learning Representations for Subspace Clustering

  • Changsheng Li
  • Chen Yang
  • Bo Liu
  • Ye Yuan
  • Guoren Wang

Deep learning based subspace clustering methods have attracted increasing attention in recent years, where a basic theme is to non-linearly map data into a latent space, and then uncover subspace structures based upon the data selfexpressiveness property. However, almost all existing deep subspace clustering methods only rely on target domain data, and always resort to shallow neural networks for modeling data, leaving huge room to design more effective representation learning mechanisms tailored for subspace clustering. In this paper, we propose a novel subspace clustering framework through learning precise sample representations. In contrast to previous approaches, the proposed method aims to leverage external data through constructing lots of relevant tasks to guide the training of the encoder, motivated by the idea of meta-learning. Considering limited networks layers of current deep subspace clustering models, we intend to distill knowledge from a deeper network trained on the external data, and transfer it into the shallower model. To reach the above two goals, we propose a new loss function to realize them in a unified framework. Moreover, we propose to construct a new auxiliary task for self-supervised training of the model, such that the representation ability of the model can be further improved. Extensive experiments are performed on four publicly available datasets, and experimental results clearly demonstrate the efficacy of our method, compared to state-of-the-art methods.

AAAI Conference 2021 Conference Paper

Unsupervised Active Learning via Subspace Learning

  • Changsheng Li
  • Kaihang Mao
  • Lingyan Liang
  • Dongchun Ren
  • Wei Zhang
  • Ye Yuan
  • Guoren Wang

Unsupervised active learning has been an active research topic in machine learning community, with the purpose of choosing representative samples to be labelled in an unsupervised manner. Previous works usually take the minimization of data reconstruction loss as the criterion to select representative samples, by which the original inputs can be better approximated. However, data are often drawn from low-dimensional subspaces embedded in an arbitrary highdimensional space in many scenarios, thus it might severely bring in noise if attempting to precisely reconstruct all entries of one observation, leading to a suboptimal solution. In view of this, this paper proposes a novel unsupervised Active Learning model via Subspace Learning, called ALSL. In contrast to previous approaches, ALSL aims to discover low-rank structures of data, and then perform sample selection based on the learnt low-rank representations. To this end, we devise two different strategies and propose two corresponding formulations to select samples with and under low-rank sample representations, respectively. Since the proposed formulations involve several non-smooth regularization terms, we develop a simple but effective optimization procedure to solve them. Extensive experiments are performed on five publicly available datasets, and experimental results demonstrate the proposed first formulation achieves comparable performance with the state-of-the-arts, while the second formulation significantly outperforms them, achieving a 13% improvement over the second best baseline at most.

NeurIPS Conference 2020 Conference Paper

Beta R-CNN: Looking into Pedestrian Detection from Another Perspective

  • Zixuan Xu
  • Banghuai Li
  • Ye Yuan
  • Anhong Dang

Recently significant progress has been made in pedestrian detection, but it remains challenging to achieve high performance in occluded and crowded scenes. It could be mostly attributed to the widely used representation of pedestrians, i. e. , 2Daxis-aligned bounding box, which just describes the approximate location and size of the object. Bounding box models the object as a uniform distribution within the boundary, making pedestrians indistinguishable in occluded and crowded scenes due to much noise. To eliminate the problem, we propose a novel representation based on 2D beta distribution, named Beta Representation. It pictures a pedestrianby explicitly constructing the relationship between full-body and visible boxes, and emphasizes the center of visual mass by assigning different probability valuesto pixels. As a result, Beta Representation is much better for distinguishing highly-overlapped instances in crowded scenes with a new NMS strategy named BetaNMS. What’s more, to fully exploit Beta Representation, a novel pipeline Beta R-CNN equipped with BetaHead and BetaMask is proposed, leading to high detection performance in occluded and crowded scenes.

IJCAI Conference 2020 Conference Paper

On Deep Unsupervised Active Learning

  • Changsheng Li
  • Handong Ma
  • Zhao Kang
  • Ye Yuan
  • Xiao-Yu Zhang
  • Guoren Wang

Unsupervised active learning has attracted increasing attention in recent years, where its goal is to select representative samples in an unsupervised setting for human annotating. Most existing works are based on shallow linear models by assuming that each sample can be well approximated by the span (i. e. , the set of all linear combinations) of certain selected samples, and then take these selected samples as representative ones to label. However, in practice, the data do not necessarily conform to linear models, and how to model nonlinearity of data often becomes the key point to success. In this paper, we present a novel Deep neural network framework for Unsupervised Active Learning, called DUAL. DUAL can explicitly learn a nonlinear embedding to map each input into a latent space through an encoder-decoder architecture, and introduce a selection block to select representative samples in the the learnt latent space. In the selection block, DUAL considers to simultaneously preserve the whole input patterns as well as the cluster structure of data. Extensive experiments are performed on six publicly available datasets, and experimental results clearly demonstrate the efficacy of our method, compared with state-of-the-arts.

IJCAI Conference 2020 Conference Paper

pbSGD: Powered Stochastic Gradient Descent Methods for Accelerated Non-Convex Optimization

  • Beitong Zhou
  • Jun Liu
  • Weigao Sun
  • Ruijuan Chen
  • Claire Tomlin
  • Ye Yuan

We propose a novel technique for improving the stochastic gradient descent (SGD) method to train deep networks, which we term pbSGD. The proposed pbSGD method simply raises the stochastic gradient to a certain power elementwise during iterations and introduces only one additional parameter, namely, the power exponent (when it equals to 1, pbSGD reduces to SGD). We further propose pbSGD with momentum, which we term pbSGDM. The main results of this paper present comprehensive experiments on popular deep learning models and benchmark datasets. Empirical results show that the proposed pbSGD and pbSGDM obtain faster initial training speed than adaptive gradient methods, comparable generalization ability with SGD, and improved robustness to hyper-parameter selection and vanishing gradients. pbSGD is essentially a gradient modifier via a nonlinear transformation. As such, it is orthogonal and complementary to other techniques for accelerating gradient-based optimization such as learning rate schedules. Finally, we show convergence rate analysis for both pbSGD and pbSGDM methods. The theoretical rates of convergence match the best known theoretical rates of convergence for SGD and SGDM methods on nonconvex functions.

NeurIPS Conference 2020 Conference Paper

Residual Force Control for Agile Human Behavior Imitation and Extended Motion Synthesis

  • Ye Yuan
  • Kris Kitani

Reinforcement learning has shown great promise for synthesizing realistic human behaviors by learning humanoid control policies from motion capture data. However, it is still very challenging to reproduce sophisticated human skills like ballet dance, or to stably imitate long-term human behaviors with complex transitions. The main difficulty lies in the dynamics mismatch between the humanoid model and real humans. That is, motions of real humans may not be physically possible for the humanoid model. To overcome the dynamics mismatch, we propose a novel approach, residual force control (RFC), that augments a humanoid control policy by adding external residual forces into the action space. During training, the RFC-based policy learns to apply residual forces to the humanoid to compensate for the dynamics mismatch and better imitate the reference motion. Experiments on a wide range of dynamic motions demonstrate that our approach outperforms state-of-the-art methods in terms of convergence speed and the quality of learned motions. Notably, we showcase a physics-based virtual character empowered by RFC that can perform highly agile ballet dance moves such as pirouette, arabesque and jeté. Furthermore, we propose a dual-policy control framework, where a kinematic policy and an RFC-based policy work in tandem to synthesize multi-modal infinite-horizon human motions without any task guidance or user input. Our approach is the first humanoid control method that successfully learns from a large-scale human motion dataset (Human3. 6M) and generates diverse long-term motions. Code and videos are available at https: //www. ye-yuan. com/rfc.

NeurIPS Conference 2020 Conference Paper

Scalable Graph Neural Networks via Bidirectional Propagation

  • Ming Chen
  • Zhewei Wei
  • Bolin Ding
  • Yaliang Li
  • Ye Yuan
  • Xiaoyong Du
  • Ji-Rong Wen

Graph Neural Networks (GNN) are an emerging field for learning on non-Euclidean data. Recently, there has been increased interest in designing GNN that scales to large graphs. Most existing methods use "graph sampling" or "layer-wise sampling" techniques to reduce training time; However, these methods still suffer from degrading performance and scalability problems when applying to graphs with billions of edges. In this paper, we present GBP, a scalable GNN that utilizes a localized bidirectional propagation process from both the feature vector and the training/testing nodes. Theoretical analysis shows that GBP is the first method that achieves sub-linear time complexity for both the precomputation and the training phases. An extensive empirical study demonstrates that GBP achieves state-of-the-art performance with significantly less training/testing time. Most notably, GBP is able to deliver superior performance on a graph with over 60 million nodes and 1. 8 billion edges in less than 2, 000 seconds on a single machine.

AAAI Conference 2020 Conference Paper

SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation Guidelines

  • Yinda Xu
  • Zeyu Wang
  • Zuoxin Li
  • Ye Yuan
  • Gang Yu

Visual tracking problem demands to efficiently perform robust classification and accurate target state estimation over a given target at the same time. Former methods have proposed various ways of target state estimation, yet few of them took the particularity of the visual tracking problem itself into consideration. Based on a careful analysis, we propose a set of practical guidelines of target state estimation for high-performance generic object tracker design. Following these guidelines, we design our Fully Convolutional Siamese tracker++ (SiamFC++) by introducing both classi- fication and target state estimation branch (G1), classification score without ambiguity (G2), tracking without prior knowledge (G3), and estimation quality score (G4). Extensive analysis and ablation studies demonstrate the effectiveness of our proposed guidelines. Without bells and whistles, our SiamFC++ tracker achieves state-of-the-art performance on five challenging benchmarks(OTB2015, VOT2018, La- SOT, GOT-10k, TrackingNet), which proves both the tracking and generalization ability of the tracker. Particularly, on the large-scale TrackingNet dataset, SiamFC++ achieves a previously unseen AUC score of 75. 4 while running at over 90 FPS, which is far above the real-time requirement.

IJCAI Conference 2020 Conference Paper

Simultaneous Arrival Matching for New Spatial Crowdsourcing Platforms

  • Boyang Li
  • Yurong Cheng
  • Ye Yuan
  • Guoren Wang
  • Lei Chen

In recent years, 3D spatial crowdsourcing platforms become popular, in which users and workers travel together to their assigned workplaces for services, such as InterestingSport and Nanguache. A typical problem over 3D spatial crowdsourcing platforms is to match users with suitable workers and workplaces. Existing studies all ignored that the workers and users assigned to the same workplace should arrive almost at the same time, which is very practical in the real world. Thus, in this paper, we propose a new Simultaneous Arrival Matching (SAM), which enables workers and users to arrive at their assigned workplace within a given tolerant time. We find that the new considered arriving time constraint breaks the monotonic additivity of the result set. Thus, it brings a large challenge in designing effective and efficient algorithms for the SAM. We design Sliding Window algorithm and Threshold Scanning algorithm to solve the SAM. We conduct the experiments on real and synthetic datasets, experimental results show the effectiveness and efficiency of our algorithms.

JBHI Journal 2019 Journal Article

A Multi-View Deep Learning Framework for EEG Seizure Detection

  • Ye Yuan
  • Guangxu Xun
  • Kebin Jia
  • Aidong Zhang

The recent advances in pervasive sensing technologies have enabled us to monitor and analyze the multi-channel electroencephalogram (EEG) signals of epilepsy patients to prevent serious outcomes caused by epileptic seizures. To avoid manual visual inspection from long-term EEG readings, automatic EEG seizure detection has garnered increasing attention among researchers. In this paper, we present a unified multi-view deep learning framework to capture brain abnormalities associated with seizures based on multi-channel scalp EEG signals. The proposed approach is an end-to-end model that is able to jointly learn multi-view features from both unsupervised multi-channel EEG reconstruction and supervised seizure detection via spectrogram representation. We construct a new autoencoder-based multi-view learning model by incorporating both inter and intra correlations of EEG channels to unleash the power of multi-channel information. By adding a channel-wise competition mechanism in the training phase, we propose a channel-aware seizure detection module to guide our multi-view structure to focus on important and relevant EEG channels. To validate the effectiveness of the proposed framework, extensive experiments against nine baselines, including both traditional handcrafted feature extraction and conventional deep learning methods, are carried out on a benchmark scalp EEG dataset. Experimental results show that the proposed model is able to achieve higher average accuracy and f1-score at 94. 37% and 85. 34%, respectively, using 5-fold subject-independent cross validation, demonstrating a powerful and effective method in the task of EEG seizure detection.

IJCAI Conference 2019 Conference Paper

Metric Learning on Healthcare Data with Incomplete Modalities

  • Qiuling Suo
  • Weida Zhong
  • Fenglong Ma
  • Ye Yuan
  • Jing Gao
  • Aidong Zhang

Utilizing multiple modalities to learn a good distance metric is of vital importance for various clinical applications. However, it is common that modalities are incomplete for some patients due to various technical and practical reasons in healthcare datasets. Existing metric learning methods cannot directly learn the distance metric on such data with missing modalities. Nevertheless, the incomplete data contains valuable information to characterize patient similarity and modality relationships, and they should not be ignored during the learning process. To tackle the aforementioned challenges, we propose a metric learning framework to perform missing modality completion and multi-modal metric learning simultaneously. Employing the generative adversarial networks, we incorporate both complete and incomplete data to learn the mapping relationship between modalities. After completing the missing modalities, we use the nonlinear representations extracted by the discriminator to learn the distance metric among patients. Through jointly training the adversarial generation part and metric learning, the similarity among patients can be learned on data with missing modalities. Experimental results show that the proposed framework learns more accurate distance metric on real-world healthcare datasets with incomplete modalities, comparing with the state-of-the-art approaches. Meanwhile, the quality of the generated modalities can be preserved.

AAAI Conference 2018 Conference Paper

Multi-Modal Multi-Task Learning for Automatic Dietary Assessment

  • Qi Liu
  • Yue Zhang
  • Zhenguang Liu
  • Ye Yuan
  • Li Cheng
  • Roger Zimmermann

We investigate the task of automatic dietary assessment: given meal images and descriptions uploaded by real users, our task is to automatically rate the meals and deliver advisory comments for improving users’ diets. To address this practical yet challenging problem, which is multi-modal and multi-task in nature, an end-to-end neural model is proposed. In particular, comprehensive meal representations are obtained from images, descriptions and user information. We further introduce a novel memory network architecture to store meal representations and reason over the meal representations to support predictions. Results on a real-world dataset show that our method outperforms two strong image captioning baselines significantly.

ICRA Conference 2017 Conference Paper

Computational abstractions for interactive design of robotic devices

  • Ruta Desai
  • Ye Yuan
  • Stelian Coros

We present a computational design system that allows novices and experts alike to easily create custom robotic devices using modular electromechanical components. The core of our work consists of a design abstraction that models the way in which these components can be combined to form complex robotic systems. We use this abstraction to develop a visual design environment that enables an intuitive exploration of the space of robots that can be created using a given set of actuators, mounting brackets and 3d-printable components. Our computational system also provides support for design auto-completion operations, which further simplifies the task of creating robotic devices. Once robot designs are finished, they can be tested in physical simulations and iteratively improved until they meet the individual needs of their users. We demonstrate the versatility of our computational design system by creating an assortment of legged and wheeled robotic devices. To test the physical feasibility of our designs, we fabricate a wheeled device equipped with a 5-DOF arm and a quadrupedal robot.

NeurIPS Conference 2016 Conference Paper

Review Networks for Caption Generation

  • Zhilin Yang
  • Ye Yuan
  • Yuexin Wu
  • William Cohen
  • Russ Salakhutdinov

We propose a novel extension of the encoder-decoder framework, called a review network. The review network is generic and can enhance any existing encoder- decoder model: in this paper, we consider RNN decoders with both CNN and RNN encoders. The review network performs a number of review steps with attention mechanism on the encoder hidden states, and outputs a thought vector after each review step; the thought vectors are used as the input of the attention mechanism in the decoder. We show that conventional encoder-decoders are a special case of our framework. Empirically, we show that our framework improves over state-of- the-art encoder-decoder systems on the tasks of image captioning and source code captioning.

v2026.09.13