Arrow Research search

Author name cluster

Fengyu Yang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

7 papers
1 author row

Possible papers

7

AAAI Conference 2026 Conference Paper

VideoSeg-R1:Reasoning Video Object Segmentation via Reinforcement Learning

  • Zishan Xu
  • Yifu Guo
  • Yuquan Lu
  • Fengyu Yang
  • Junxin Li
  • Lihua Cai

Traditional video reasoning segmentation methods rely on supervised fine-tuning, which limits generalization to out-of-distribution scenarios and lacks explicit reasoning. To address this, we propose VideoSeg-R1, the first framework to introduce reinforcement learning into video reasoning segmentation. It adopts a decoupled architecture that formulates the task as joint referring image segmentation and video mask propagation. It comprises three stages: (1) A hierarchical text-guided frame sampler to emulate human attention; (2) A reasoning model that produces spatial cues along with explicit reasoning chains; and (3) A segmentation-propagation stage using SAM2 and XMem. A task difficulty-aware mechanism adaptively controls reasoning length for better efficiency and accuracy. Extensive evaluations on multiple benchmarks demonstrate that VideoSeg-R1 achieves state-of-the-art performance in complex video reasoning and segmentation tasks.

NeurIPS Conference 2025 Conference Paper

Differentiation Through Black-Box Quadratic Programming Solvers

  • Connor Magoon
  • Fengyu Yang
  • Noam Aigerman
  • Shahar Kovalsky

Differentiable optimization has attracted significant research interest, particularly for quadratic programming (QP). Existing approaches for differentiating the solution of a QP with respect to its defining parameters often rely on specific integrated solvers. This integration limits their applicability, including their use in neural network architectures and bi-level optimization tasks, restricting users to a narrow selection of solver choices. To address this limitation, we introduce dQP, a modular and solver-agnostic framework for plug-and-play differentiation of virtually any QP solver. A key insight we leverage to achieve modularity is that, once the active set of inequality constraints is known, both the solution and its derivative can be expressed using simplified linear systems that share the same matrix. This formulation fully decouples the computation of the QP solution from its differentiation. Building on this result, we provide a minimal-overhead, open-source implementation () that seamlessly integrates with over 15 state-of-the-art solvers. Comprehensive benchmark experiments demonstrate dQP’s robustness and scalability, particularly highlighting its advantages in large-scale sparse problems.

EAAI Journal 2025 Journal Article

Dynamic boundary guided adversarial robust distillation to improve robustness

  • Fengyu Yang
  • Hong Yan
  • Yu Xiong
  • Xiaohui Wei
  • Yihui Zhong

The vulnerability of deep neural networks to adversarial sample attacks seriously affects their reliability and safety in practical applications. By introducing finely designed, imperceptible perturbations to normal samples, the model can be induced to make incorrect judgements, potentially resulting in serious consequences. Existing adversarial defense methods are typically confined to defending against a single specific adversarial attack, and the generated adversarial samples often lack sufficient robustness features. Moreover, enhancing adversarial robustness while maintaining standard accuracy remains a significant challenge. This study meticulously examined the relationship between the state of the decision boundary and model performance. The results revealed that the samples located near the decision boundary are crucial for improving the robustness of the model. On this basis, a novel decision boundary-guided adversarial training method was developed. Further, a dynamic boundary-guided adversarial robust distillation framework was designed based on dynamic boundary-guided distillation. The framework adopts the standard knowledge distillation technique to train the student network at the initial stage, and introduces high-quality boundary samples at the later stage to improve the model’s adversarial robustness and accelerate convergence. Meanwhile, an adaptive strategy for adjusting the decision boundary was devised to dynamically adjust the weight proportion of different loss functions to balance the standard accuracy and adversarial robustness of the deep learning model. Experimental results show that the proposed method improved the average adversarial robustness by 6. 17% on multiple datasets and models. Under the comprehensive AutoAttack method, the proposed method achieved a significant improvement of 5. 98%, demonstrating its superiority in complex scenarios.

AAAI Conference 2025 Conference Paper

TextToucher: Fine-Grained Text-to-Touch Generation

  • Jiahang Tu
  • Hao Fu
  • Fengyu Yang
  • Hanbin Zhao
  • Chao Zhang
  • Hui Qian

Tactile sensation plays a crucial role in the development of multi-modal large models and embodied intelligence. To collect tactile data with minimal cost as possible, a series of studies have attempted to generate tactile images by vision-to-touch image translation. However, compared to text modality, visual modality-driven tactile generation cannot accurately depict human tactile sensation. In this work, we analyze the characteristics of tactile images in detail from two granularities: object-level (tactile texture, tactile shape), and sensor-level (gel status). We model these granularities of information through text descriptions and propose a fine-grained Text-to-Touch generation method (TextToucher) to generate high-quality tactile samples. Specifically, we introduce a multimodal large language model to build the text sentences about object-level tactile information and employ a set of learnable text prompts to represent the sensor-level tactile information. To better guide the tactile generation process with the built text information, we fuse the dual grains of text information and explore various dual-grain text conditioning methods within the diffusion transformer architecture. Furthermore, we propose a Contrastive Text-Touch Pre-training (CTTP) metric to precisely evaluate the quality of text-driven generated tactile data. Extensive experiments demonstrate the superiority of our TextToucher method.

AAAI Conference 2025 Conference Paper

Tri-Ergon: Fine-Grained Video-to-Audio Generation with Multi-Modal Conditions and LUFS Control

  • Bingliang Li
  • Fengyu Yang
  • Yuxin Mao
  • Qingwen Ye
  • Hongkai Chen
  • Yiran Zhong

Video-to-audio (V2A) generation utilizes visual-only video features to produce realistic sounds that correspond to the scene. However, current V2A models often lack fine-grained control over the generated audio, especially in terms of loudness variation and the incorporation of multi-modal conditions. To overcome these limitations, we introduce Tri-Ergon, a diffusion-based V2A model that incorporates textual, auditory, and pixel-level visual prompts to enable detailed and semantically rich audio synthesis. Additionally, we introduce Loudness Units relative to Full Scale (LUFS) embedding, which allows for precise manual control of the loudness changes over time for individual audio channels, enabling our model to effectively address the intricate correlation of video and audio in real-world Foley workflows. Tri-Ergon is capable of creating 44.1 kHz high-fidelity stereo audio clips of varying lengths up to 60 seconds, which significantly outperforms existing state-of-the-art V2A methods that typically generate mono audio for a fixed duration.

NeurIPS Conference 2024 Conference Paper

RSA: Resolving Scale Ambiguities in Monocular Depth Estimators through Language Descriptions

  • Ziyao Zeng
  • Yangchao Wu
  • Hyoungseob Park
  • Daniel Wang
  • Fengyu Yang
  • Stefano Soatto
  • Dong Lao
  • Byung-Woo Hong

We propose a method for metric-scale monocular depth estimation. Inferring depth from a single image is an ill-posed problem due to the loss of scale from perspective projection during the image formation process. Any scale chosen is a bias, typically stemming from training on a dataset; hence, existing works have instead opted to use relative (normalized, inverse) depth. Our goal is to recover metric-scaled depth maps through a linear transformation. The crux of our method lies in the observation that certain objects (e. g. , cars, trees, street signs) are typically found or associated with certain types of scenes (e. g. , outdoor). We explore whether language descriptions can be used to transform relative depth predictions to those in metric scale. Our method, RSA, takes as input a text caption describing objects present in an image and outputs the parameters of a linear transformation which can be applied globally to a relative depth map to yield metric-scaled depth predictions. We demonstrate our method on recent general-purpose monocular depth models on indoors (NYUv2, VOID) and outdoors (KITTI). When trained on multiple datasets, RSA can serve as a general alignment module in zero-shot settings. Our method improves over common practices in aligning relative to metric depth and results in predictions that are comparable to an upper bound of fitting relative depth to ground truth via a linear transformation. Code is available at: https: //github. com/Adonis-galaxy/RSA.

NeurIPS Conference 2022 Conference Paper

Touch and Go: Learning from Human-Collected Vision and Touch

  • Fengyu Yang
  • Chenyang Ma
  • Jiacheng Zhang
  • Jing Zhu
  • Wenzhen Yuan
  • Andrew Owens

The ability to associate touch with sight is essential for tasks that require physically interacting with objects in the world. We propose a dataset with paired visual and tactile data called Touch and Go, in which human data collectors probe objects in natural environments using tactile sensors, while simultaneously recording egocentric video. In contrast to previous efforts, which have largely been confined to lab settings or simulated environments, our dataset spans a large number of “in the wild” objects and scenes. We successfully apply our dataset to a variety of multimodal learning tasks: 1) self-supervised visuo-tactile feature learning, 2) tactile-driven image stylization, i. e. , making the visual appearance of an object more consistent with a given tactile signal, and 3) predicting future frames of a tactile signal from visuo-tactile inputs.

v2026.09.13