Arrow Research search

Author name cluster

Weifeng Chen

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

5 papers
2 author rows

Possible papers

5

ICLR Conference 2025 Conference Paper

IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model

  • Yatai Ji
  • Shilong Zhang
  • Jie Wu 0001
  • Peize Sun
  • Weifeng Chen
  • Xuefeng Xiao 0001
  • Sidi Yang
  • Yujiu Yang 0001

The rapid advancement of Large Vision-Language models (LVLMs) has demonstrated a spectrum of emergent capabilities. Nevertheless, current models only focus on the visual content of a single scenario, while their ability to associate instances across different scenes has not yet been explored, which is essential for understanding complex visual content, such as movies with multiple characters and intricate plots. Towards movie understanding, a critical initial step for LVLMs is to unleash the potential of character identities memory and recognition across multiple visual scenarios. To achieve the goal, we propose visual instruction tuning with ID reference and develop an ID-Aware Large Vision-Language Model, IDA-VLM. Furthermore, our research introduces a novel benchmark MM-ID, to examine LVLMs on instance IDs memory and recognition across four dimensions: matching, location, question-answering, and captioning. Our findings highlight the limitations of existing LVLMs in recognizing and associating instance identities with ID reference. This paper paves the way for future artificial intelligence systems to possess multi-identity visual inputs, thereby facilitating the comprehension of complex visual narratives like movies.

AAAI Conference 2025 Conference Paper

OOTDiffusion: Outfitting Fusion Based Latent Diffusion for Controllable Virtual Try-On

  • Yuhao Xu
  • Tao Gu
  • Weifeng Chen
  • Arlene Chen

We present OOTDiffusion, a novel network architecture for realistic and controllable image-based virtual try-on (VTON). We leverage the power of pretrained latent diffusion models, designing an outfitting UNet to learn the detailed garment features. Without a redundant warping process, the garment features are precisely aligned with the target human body via the proposed outfitting fusion in the self-attention layers of the denoising UNet. In order to further enhance the controllability, we introduce outfitting dropout to the training process, which enables us to adjust the strength of the garment features through classifier-free guidance. Our comprehensive experiments on the VITON-HD and Dress Code datasets demonstrate that OOTDiffusion efficiently generates high-quality try-on results for arbitrary human and garment images, which outperforms other VTON methods in both realism and controllability, indicating a breakthrough in virtual try-on.

AAAI Conference 2021 Conference Paper

Learning to Sit: Synthesizing Human-Chair Interactions via Hierarchical Control

  • Yu-Wei Chao
  • Jimei Yang
  • Weifeng Chen
  • Jia Deng

Recent progress on physics-based character animation has shown impressive breakthroughs on human motion synthesis, through imitating motion capture data via deep reinforcement learning. However, results have mostly been demonstrated on imitating a single distinct motion pattern, and do not generalize to interactive tasks that require flexible motion patterns due to varying human-object spatial configurations. To bridge this gap, we focus on one class of interactive tasks—sitting onto a chair. We propose a hierarchical reinforcement learning framework which relies on a collection of subtask controllers trained to imitate simple, reusable mocap motions, and a meta controller trained to execute the subtasks properly to complete the main task. We experimentally demonstrate the strength of our approach over different non-hierarchical and hierarchical baselines. We also show that our approach can be applied to motion prediction given an image input. A supplementary video can be found at https: //youtu. be/3CeN0OGz2cA.

NeurIPS Conference 2016 Conference Paper

Single-Image Depth Perception in the Wild

  • Weifeng Chen
  • Zhao Fu
  • Dawei Yang
  • Jia Deng

This paper studies single-image depth perception in the wild, i. e. , recovering depth from a single image taken in unconstrained settings. We introduce a new dataset “Depth in the Wild” consisting of images in the wild annotated with relative depth between pairs of random points. We also propose a new algorithm that learns to estimate metric depth using annotations of relative depth. Compared to the state of the art, our algorithm is simpler and performs better. Experiments show that our algorithm, combined with existing RGB-D data and our new relative depth annotations, significantly improves single-image depth perception in the wild.

EAAI Journal 2012 Journal Article

Advanced issues in artificial intelligence and pattern recognition for intelligent surveillance system in smart home environment

  • Seungmin Rho
  • Geyong Min
  • Weifeng Chen

During the last decades, many researchers in image processing and AI community have been focused on developing image and video analysis and understanding. However, despite their extensive efforts on these, there are still several significant challenges such as robust object segmentation and tracking, motion feature extraction, context modeling, and machine learning algorithms. Therefore, more advanced related issues should be taken into account. We have selected nine research papers whose topics are strongly related to the intelligent surveillance system in smart home environment.

v2026.09.13