Arrow Research search

Author name cluster

Yao Tang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

7 papers
2 author rows

Possible papers

7

NeurIPS Conference 2025 Conference Paper

LEDiT: Your Length-Extrapolatable Diffusion Transformer without Positional Encoding

  • Shen Zhang
  • Siyuan Liang
  • Yaning Tan
  • Zhaowei Chen
  • Linze Li
  • Ge Wu
  • Yuhao Chen
  • Shuheng Li

Diffusion transformers (DiTs) struggle to generate images at resolutions higher than their training resolutions. The primary obstacle is that the explicit positional encodings (PE), such as RoPE, need extrapolating to unseen positions which degrades performance when the inference resolution differs from training. In this paper, We propose a Length-Extrapolatable Diffusion Transformer (LEDiT) to overcome this limitation. LEDiT needs no explicit PEs, thereby avoiding PE extrapolation. The key innovation of LEDiT lies in the use of causal attention. We demonstrate that causal attention can implicitly encode global positional information and show that such information facilitates extrapolation. We further introduce a locality enhancement module, which captures fine-grained local information to complement the global coarse-grained position information encoded by causal attention. Experimental results on both conditional and text-to-image generation tasks demonstrate that LEDiT supports up to 4× resolution scaling (e. g. , from 256$\times$256 to 512$\times$512), achieving better image quality compared to the state-of-the-art length extrapolation methods. We believe that LEDiT marks a departure from the standard RoPE-based methods and offers a promising insight into length extrapolation. Project page: https: //shenzhang2145. github. io/ledit/

EAAI Journal 2025 Journal Article

Loosen Attention: Integrating localized channel and coarse spatial attention for enhanced analysis of complex aurora images

  • Qian Wang
  • Shihao Jing
  • Rui Yang
  • Zhenpei Liu
  • Yao Tang
  • Han Pan

Auroral images present a significant challenge for automated analysis due to their highly variable and dynamic morphology, influenced by complex interactions between the solar wind and Earth’s magnetosphere. These natural phenomena exhibit considerable randomness in shape, brightness, and motion, making them a unique and challenging signal source for artificial intelligence methods. In this work, we propose Loosen Attention (LA), a novel and lightweight attention mechanism tailored to capture the unpredictable and fluid-like nature of auroral patterns. LA integrates localized channel attention with coarse-grained spatial attention, forming a flexible attention framework that enhances the robustness and adaptability of feature extraction in deep learning models. The LA module is engineered around four key strategies: volumetric feature generation, volume-wise attention computation, refinement and reconstruction, and enhanced feature fusion, enabling efficient focus on subtle yet significant auroral structures while tolerating less informative regions. Unlike conventional attention methods that may struggle with complex visual patterns, LA is explicitly designed to process diverse auroral images in a structured and computationally efficient way. We validate the LA mechanism on multiple vision tasks to demonstrate consistent performance improvements over state-of-the-art attention modules. The results highlight the potential of LA to improve deep learning pipelines in geospace imaging and other domains dealing with similarly complex natural signals. This work offers a promising approach for intelligent processing of Earth–space interaction data and other visually ambiguous scientific images.

ICML Conference 2025 Conference Paper

QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search

  • Zongyu Lin
  • Yao Tang
  • Xingcheng Yao
  • Da Yin
  • Ziniu Hu
  • Yizhou Sun
  • Kai-Wei Chang 0001

Language agents have become a promising solution to complex interactive tasks. One of the key ingredients to the success of language agents is the reward model on the trajectory of the agentic workflow, which provides valuable guidance during training or inference. However, due to the lack of annotations of intermediate interactions, most existing works use an outcome reward model to optimize policies across entire trajectories. This may lead to sub-optimal policies and hinder the overall performance. To address this, we propose QLASS (Q-guided Language Agent Stepwise Search), to automatically generate annotations by estimating Q-values in a stepwise manner for open language agents. By introducing a reasoning tree and performing process reward modeling, QLASS provides effective intermediate guidance for each step. With the stepwise guidance, we propose a Q-guided generation strategy to enable language agents to better adapt to long-term value, resulting in significant performance improvement during model inference on complex interactive agent tasks. Notably, even with almost half the annotated data, QLASS retains strong performance, demonstrating its efficiency in handling limited supervision. We also empirically demonstrate that QLASS can lead to more effective decision making through qualitative analysis.

NeurIPS Conference 2025 Conference Paper

Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think

  • Ge Wu
  • Shen Zhang
  • Ruijing Shi
  • Shanghua Gao
  • Zhenyuan Chen
  • Lei Wang
  • Zhaowei Chen
  • Hongcheng Gao

REPA and its variants effectively mitigate training challenges in diffusion models by incorporating external visual representations from pretrained models, through alignment between the noisy hidden projections of denoising networks and foundational clean image representations. We argue that the external alignment, which is absent during the entire denoising inference process, falls short of fully harnessing the potential of discriminative representations. In this work, we propose a straightforward method called $\textit{$\textbf{R}$epresentation $\textbf{E}$ntanglement for $\textbf{G}$eneration}$ ($\textbf{REG}$), which entangles low-level image latents with a single high-level class token from pretrained foundation models for denoising. REG acquires the capability to produce coherent image-class pairs directly from pure noise, substantially improving both generation quality and training efficiency. This is accomplished with negligible additional inference overhead, requiring only one single additional token for denoising (<0. 5\% increase in FLOPs and latency). The inference process concurrently reconstructs both image latents and their corresponding global semantics, where the acquired semantic knowledge actively guides and enhances the image generation process. On ImageNet 256$\times$256, SiT-XL/2 + REG demonstrates remarkable convergence acceleration, achieving $\textbf{63}\times$ and $\textbf{23}\times$ faster training than SiT-XL/2 and SiT-XL/2 + REPA, respectively. More impressively, SiT-L/2 + REG trained for merely 400K iterations outperforms SiT-XL/2 + REPA trained for 4M iterations ($\textbf{10}\times$ longer). Code is available at: https: //github. com/Martinser/REG.

NeurIPS Conference 2024 Conference Paper

Learning Versatile Skills with Curriculum Masking

  • Yao Tang
  • Zhihui Xie
  • Zichuan Lin
  • Deheng Ye
  • Shuai Li

Masked prediction has emerged as a promising pretraining paradigm in offline reinforcement learning (RL) due to its versatile masking schemes, enabling flexible inference across various downstream tasks with a unified model. Despite the versatility of masked prediction, it remains unclear how to balance the learning of skills at different levels of complexity. To address this, we propose CurrMask, a curriculum masking pretraining paradigm for sequential decision making. Motivated by how humans learn by organizing knowledge in a curriculum, CurrMask adjusts its masking scheme during pretraining for learning versatile skills. Through extensive experiments, we show that CurrMask exhibits superior zero-shot performance on skill prompting tasks, goal-conditioned planning tasks, and competitive finetuning performance on offline RL tasks. Additionally, our analysis of training dynamics reveals that CurrMask gradually acquires skills of varying complexity by dynamically adjusting its masking scheme.

AAMAS Conference 2024 Conference Paper

Risk-Aware Constrained Reinforcement Learning with Non-Stationary Policies

  • Zhaoxing Yang
  • Haiming Jin
  • Yao Tang
  • Guiyun Fan

Constrained reinforcement learning (RL) algorithms have attracted extensive attentions nowadays to tackle sequential decision-making problems that contain constraints defined under various risk measures. However, most works only search policies within the stationary policy class and fail to capture a simple intuition: adjust the action-selecting distribution at each state according to the accumulated cost so far. In this work, we design a novel quantile-leveldriven policy class to fully realize such intuition, within which each policy additionally takes the quantile level of the accumulated cost as input. Such quantile level is obtained via a novel Invertible Backward Distributional Critic (IBDC) framework, which utilizes invertible function approximators to estimate the accumulated cost distribution and outputs the required quantile level with their inverse forms. Further, the estimated accumulated cost distribution also helps to decompose the challenging trajectory-level constraints into state-level constraints, and Risk-Aware Constrained RL (RAC) algorithm is designed then to solve the decomposed problem with Lagrangian multipliers. Experimental results in various environments validate the effectiveness of RAC versus state-of-the-art baselines.

JBHI Journal 2022 Journal Article

iCVM: An Interpretable Deep Learning Model for CVM Assessment Under Label Uncertainty

  • Ni Liao
  • Jian Dai
  • Yao Tang
  • Qiaoyong Zhong
  • Shuixue Mo

The Cervical Vertebral Maturation (CVM) method aims to determine the craniofacial skeletal maturational stage, which is crucial for orthodontic and orthopedic treatment. In this paper, we explore the potential of deep learning for automatic CVM assessment. In particular, we propose a convolutional neural network named iCVM. Based on the residual network, it is specialized for the challenges unique to the task of CVM assessment. 1) To combat overfitting due to limited data size, multiple dropout layers are utilized. 2) To address the inevitable label ambiguity between adjacent maturational stages, we introduce the concept of label distribution learning in the loss function. Besides, we attempt to analyze the regions important for the prediction of the model by using the Grad-CAM technique. The learned strategy shows surprisingly high consistency with the clinical criteria. This indicates that the decisions made by our model are well interpretable, which is critical in evaluation of growth and development in orthodontics. Moreover, to drive future research in the field, we release a new dataset named CVM-900 along with the paper. It contains the cervical part of 900 lateral cephalograms collected from orthodontic patients of different ages and genders. Experimental results show that the proposed approach achieves superior performance on CVM-900 in terms of various evaluation metrics.

v2026.09.13