Arrow Research search

Author name cluster

Yuyang Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

11 papers
2 author rows

Possible papers

11

ICLR Conference 2025 Conference Paper

Denoising Autoregressive Transformers for Scalable Text-to-Image Generation

  • Jiatao Gu
  • Yuyang Wang
  • Yizhe Zhang 0002
  • Qihang Zhang
  • Dinghuai Zhang
  • Navdeep Jaitly
  • Joshua M. Susskind
  • Shuangfei Zhai

Diffusion models have become the dominant approach for visual generation. They are trained by denoising a Markovian process which gradually adds noise to the input. We argue that the Markovian property limits the model’s ability to fully utilize the generation trajectory, leading to inefficiencies during training and inference. In this paper, we propose DART, a transformer-based model that unifies autoregressive (AR) and diffusion within a non-Markovian framework. DART iteratively denoises image patches spatially and spectrally using an AR model that has the same architecture as standard language models. DART does not rely on image quantization, which enables more effective image modeling while maintaining flexibility. Furthermore, DART seamlessly trains with both text and image data in a unified model. Our approach demonstrates competitive performance on class-conditioned and text-to-image generation tasks, offering a scalable, efficient alternative to traditional diffusion models. Through this unified framework, DART sets a new benchmark for scalable, high-quality image synthesis.

ICML Conference 2025 Conference Paper

INRFlow: Flow Matching for INRs in Ambient Space

  • Yuyang Wang
  • Anurag Ranjan
  • Joshua M. Susskind
  • Miguel Ángel Bautista 0001

Flow matching models have emerged as a powerful method for generative modeling on domains like images or videos, and even on irregular or unstructured data like 3D point clouds or even protein structures. These models are commonly trained in two stages: first, a data compressor is trained, and in a subsequent training stage a flow matching generative model is trained in the latent space of the data compressor. This two-stage paradigm sets obstacles for unifying models across data domains, as hand-crafted compressors architectures are used for different data modalities. To this end, we introduce INRFlow, a domain-agnostic approach to learn flow matching transformers directly in ambient space. Drawing inspiration from INRs, we introduce a conditionally independent point-wise training objective that enables INRFlow to make predictions continuously in coordinate space. Our empirical results demonstrate that INRFlow effectively handles different data modalities such as images, 3D point clouds and protein structure data, achieving strong performance in different domains and outperforming comparable approaches. INRFlow is a promising step towards domain-agnostic flow matching generative models that can be trivially adopted in different data domains.

NeurIPS Conference 2025 Conference Paper

STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis

  • Jiatao Gu
  • Tianrong Chen
  • David Berthelot
  • Huangjie Zheng
  • Yuyang Wang
  • Ruixiang Zhang
  • Laurent Dinh
  • Miguel Angel Bautista

We present STARFlow, a scalable generative model based on normalizing flows that achieves strong performance on high-resolution image synthesis. STARFlow's main building block is Transformer Autoregressive Flow (TARFlow), which combines normalizing flows with Autoregressive Transformer architectures and has recently achieved impressive results in image modeling. In this work, we first establish the theoretical universality of TARFlow for modeling continuous distributions. Building on this foundation, we introduce a set of architectural and algorithmic innovations that significantly enhance the scalability: (1) a deep-shallow design where a deep Transformer block captures most of the model’s capacity, followed by a few shallow Transformer blocks that are computationally cheap yet contribute non-negligibly, (2) learning in the latent space of pretrained autoencoders, which proves far more effective than modeling pixels directly, and (3) a novel guidance algorithm that substantially improves sample quality. Crucially, our model remains a single, end-to-end normalizing flow, allowing exact maximum likelihood training in continuous space without discretization. STARFlow achieves competitive results in both class- and text-conditional image generation, with sample quality approaching that of state-of-the-art diffusion models. To our knowledge, this is the first successful demonstration of normalizing flows at this scale and resolution. Code and weights available at https: //github. com/apple/ml-starflow.

EAAI Journal 2024 Journal Article

Dual attention transformer network for hyperspectral image classification

  • Zhenqiu Shu
  • Yuyang Wang
  • Zhengtao Yu

Hyperspectral image classification (HSIC) has been a significant topic in the field of remote sensing in the past few years. Convolutional neural networks have shown promising performance in HSIC applications due to their strong local feature extraction ability. However, they struggle to extract global information from HSIs, thereby resulting in classification performance limitations. Recently, vision transformers have been used to solve HSIC problems, and its advantage is to adopt the multi-head self-attention mechanism to explore global dependencies. Nevertheless, the extracted features using MHSA usually exhibit over-dispersion due to the abundance of band information hidden in HSIs. In this work, we propose a novel method, called dual attention transformer network (DATN), for HSIC problems. It consists of two types of modules, namely the spatial–spectral hybrid transformer (SSHT) module and the spectral local-conv block (SLCB) module. Specifically, the SSHT module aims to utilize the MHSA to capture spatial and spectral feature information. Therefore, it can effectively utilize global spatial–spectral features and embed the local spatial information, simultaneously. Besides, we design a SLCB module to extract the local spectral information of HSIs effectively. Then the SSHT and SLCB modules are integrated into an end-to-end framework. Finally, the global and local spatial–spectral features extracted from this framework are input into the fully connected layer, and then classification results of HSIs are obtained. A series of experiments on three HSI datasets have demonstrated that our DATN approach outperforms several state-of-the-art HSIC approaches.

ICLR Conference 2024 Conference Paper

Manifold Diffusion Fields

  • Ahmed A. A. Elhag
  • Yuyang Wang
  • Joshua M. Susskind
  • Miguel Ángel Bautista 0001

We present Manifold Diffusion Fields (MDF), an approach that unlocks learning of diffusion models of data in general non-euclidean geometries. Leveraging insights from spectral geometry analysis, we define an intrinsic coordinate system on the manifold via the eigen-functions of the Laplace-Beltrami Operator. MDF represents functions using an explicit parametrization formed by a set of multiple input-output pairs. Our approach allows to sample continuous functions on manifolds and is invariant with respect to rigid and isometric transformations of the manifold. In addition, we show that MDF generalizes to the case where the training set contains functions on different manifolds. Empirical results on multiple datasets and manifolds including challenging scientific problems like weather prediction or molecular conformation show that MDF can capture distributions of such functions with better diversity and fidelity than previous approaches.

ICRA Conference 2024 Conference Paper

OmniColor: A Global Camera Pose Optimization Approach of LiDAR-360Camera Fusion for Colorizing Point Clouds

  • Bonan Liu
  • Guoyang Zhao
  • Jianhao Jiao
  • Guang Cai
  • Chengyang Li
  • Handi Yin
  • Yuyang Wang
  • Ming Liu 0001

A Colored point cloud, as a simple and efficient 3D representation, has many advantages in various fields, including robotic navigation and scene reconstruction. This representation is now commonly used in 3D reconstruction tasks relying on cameras and LiDARs. However, fusing data from these two types of sensors is poorly performed in many existing frameworks, leading to unsatisfactory mapping results, mainly due to inaccurate camera poses. This paper presents Omni-Color, a novel and efficient algorithm to colorize point clouds using an independent 360-degree camera. Given a LiDAR-based point cloud and a sequence of panorama images with initial coarse camera poses, our objective is to jointly optimize the poses of all frames for mapping images onto geometric reconstructions. Our pipeline works in an off-the-shelf manner that does not require any feature extraction or matching process. Instead, we find optimal poses by directly maximizing the photometric consistency of LiDAR maps. In experiments, we show that our method can overcome the severe visual distortion of omnidirectional images and greatly benefit from the wide field of view (FOV) of 360-degree cameras to reconstruct various scenarios with accuracy and stability. The code will be released at https://github.com/liubonan123/OmniColor/.

ICML Conference 2024 Conference Paper

Swallowing the Bitter Pill: Simplified Scalable Conformer Generation

  • Yuyang Wang
  • Ahmed A. A. Elhag
  • Navdeep Jaitly
  • Joshua M. Susskind
  • Miguel Ángel Bautista 0001

We present a novel way to predict molecular conformers through a simple formulation that sidesteps many of the heuristics of prior works and achieves state of the art results by using the advantages of scale. By training a diffusion generative model directly on 3D atomic positions without making assumptions about the explicit structure of molecules (e. g. modeling torsional angles) we are able to radically simplify structure learning, and make it trivial to scale up the model sizes. This model, called Molecular Conformer Fields (MCF), works by parameterizing conformer structures as functions that map elements from a molecular graph directly to their 3D location in space. This formulation allows us to boil down the essence of structure prediction to learning a distribution over functions. Experimental results show that scaling up the model capacity leads to large gains in generalization performance without enforcing inductive biases like rotational equivariance. MCF represents an advance in extending diffusion models to handle complex scientific problems in a conceptually simple, scalable and effective manner.

AAAI Conference 2022 Conference Paper

Context Uncertainty in Contextual Bandits with Applications to Recommender Systems

  • Hao Wang
  • Yifei Ma
  • Hao Ding
  • Yuyang Wang

Recurrent neural networks have proven effective in modeling sequential user feedbacks for recommender systems. However, they usually focus solely on item relevance and fail to effectively explore diverse items for users, therefore harming the system performance in the long run. To address this problem, we propose a new type of recurrent neural networks, dubbed recurrent exploration networks (REN), to jointly perform representation learning and effective exploration in the latent space. REN tries to balance relevance and exploration while taking into account the uncertainty in the representations. Our theoretical analysis shows that REN can preserve the rateoptimal sublinear regret even when there exists uncertainty in the learned representations. Our empirical study demonstrates that REN can achieve satisfactory long-term rewards on both synthetic and real-world recommendation datasets, outperforming state-of-the-art models.

JMLR Journal 2020 Journal Article

GluonTS: Probabilistic and Neural Time Series Modeling in Python

  • Alexander Alexandrov
  • Konstantinos Benidis
  • Michael Bohlke-Schneider
  • Valentin Flunkert
  • Jan Gasthaus
  • Tim Januschowski
  • Danielle C. Maddix
  • Syama Rangapuram

We introduce the Gluon Time Series Toolkit (GluonTS), a Python library for deep learning based time series modeling for ubiquitous tasks, such as forecasting and anomaly detection. GluonTS simplifies the time series modeling pipeline by providing the necessary components and tools for quick model development, efficient experimentation and evaluation. In addition, it contains reference implementations of state-of-the-art time series models that enable simple benchmarking of new algorithms. [abs] [ pdf ][ bib ] [ code ] &copy JMLR 2020. ( edit, beta )

NeurIPS Conference 2018 Conference Paper

Deep State Space Models for Time Series Forecasting

  • Syama Sundar Rangapuram
  • Matthias Seeger
  • Jan Gasthaus
  • Lorenzo Stella
  • Yuyang Wang
  • Tim Januschowski

We present a novel approach to probabilistic time series forecasting that combines state space models with deep learning. By parametrizing a per-time-series linear state space model with a jointly-learned recurrent neural network, our method retains desired properties of state space models such as data efficiency and interpretability, while making use of the ability to learn complex patterns from raw data offered by deep learning approaches. Our method scales gracefully from regimes where little training data is available to regimes where data from millions of time series can be leveraged to learn accurate models. We provide qualitative as well as quantitative results with the proposed method, showing that it compares favorably to the state-of-the-art.

ICML Conference 2015 Conference Paper

Sparse Variational Inference for Generalized GP Models

  • Rishit Sheth
  • Yuyang Wang
  • Roni Khardon

Gaussian processes (GP) provide an attractive machine learning model due to their non-parametric form, their flexibility to capture many types of observation data, and their generic inference procedures. Sparse GP inference algorithms address the cubic complexity of GPs by focusing on a small set of pseudo-samples. To date, such approaches have focused on the simple case of Gaussian observation likelihoods. This paper develops a variational sparse solution for GPs under general likelihoods by providing a new characterization of the gradients required for inference in terms of individual observation likelihood terms. In addition, we propose a simple new approach for optimizing the sparse variational approximation using a fixed point computation. We demonstrate experimentally that the fixed point operator acts as a contraction in many cases and therefore leads to fast convergence. An experimental evaluation for count regression, classification, and ordinal regression illustrates the generality and advantages of the new approach.

v2026.09.13