Arrow Research search

Author name cluster

Yi Rong

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

10 papers
1 author row

Possible papers

10

EAAI Journal 2026 Journal Article

Exploring a double task learning framework for makeup transfer

  • Zhaoyang Sun
  • Shengwu Xiong
  • Yaxiong Chen
  • Yi Rong

Although impressive progress has been made with current makeup transfer methods, they are still facing two main challenges that have not been effectively addressed: (1) Due to the lack of real transferred targets, existing methods attempt to synthesize Pseudo Ground Truths (PGTs) to supervise the model training. Therefore, their performance is heavily dependent on the synthesis quality of PGTs. (2) Most previous works fail to achieve semantic alignment between the high-resolution feature maps of the source and reference images. As a result, some high-frequency makeup details will be lost, limiting their ability to deal with diverse makeup styles. In this paper, we propose a Double Task Makeup Transfer (DTMT) framework to handle these two challenges. Specifically, for the first one, DTMT jointly optimizes an unsupervised main makeup transfer task along with a self-supervised auxiliary reconstruction task to avoid the negative effects of sub-optimal PGTs. For the second challenge, we develop a novel Divide and Conquer Attention (DC-Attention) operation, which semantically aligns the high-resolution feature maps of the input images through a coarse-to-fine procedure. Compared to the traditional cross-attention operation, our DC-Attention has much less computational overhead and thus can be more efficient to process high-resolution feature maps. Extensive experiments on three publicly available datasets indicate that DTMT significantly outperforms eight benchmark methods in both quantitative and qualitative evaluations, promoting the development of artificial intelligence virtual makeup try-on. Moreover, DTMT demonstrates strong generalization and control capabilities for makeup styles, advancing the engineering applications of makeup transfer. Our code is available at https: //github. com/Snowfallingplum/DTMT.

AAAI Conference 2025 Conference Paper

GarFast: Realistic and Fast Garment Transfer with a Simplified Parser-Free Approach

  • Chenghu Du
  • Junyin Wang
  • Yi Rong
  • Feng Yu
  • Shengwu Xiong

A good garment try-on model should learn the transfer between different types of garments while satisfying: 1) high fidelity and 2) low inference speed. Existing methods address either of these two issues, limited processing speed or low generation quality. We directly use a lightweight encoder-decoder, ensuring faster speeds. To tackle the problem of lower image quality typically generated by lighter models, we present GarFast, a simplified, parser-free framework that optimizes the same lightweight network through a two-stage transformation of real data roles (from input to supervision), thereby greatly promoting model convergence. Specifically, first, we propose a correction strategy to prevent the difficulty of convergence caused by the lack of ground truth in the first stage. Second, we propose a fine-grained domain consistency to ensure that the results generated in the unsupervised first stage are highly realistic clothed human images. Finally, we propose a skin-variant refinement loss and a skinMix regularization to amplify texture differences and enhance the realism of skin-variant regions, thereby improving the quality of the generated skin. Extensive experiments thoroughly demonstrate that our method achieves high resolution, near real-time performance, and superior reconstruction quality compared to state-of-the-art approaches, with processing times of less than 0.03 seconds on an Nvidia A100.

IJCAI Conference 2025 Conference Paper

Mask Does Not Matter: A Unified Latent Diffusion-Enhanced Framework for Mask-Free Virtual Try-On

  • Chenghu Du
  • Junyin Wang
  • Kai Liu
  • Shengwu Xiong
  • Yi Rong

A good virtual try-on model should introduce minimal redundant conditional information to avoid instability and increase inference efficiency. Existing methods rely on inpainting masks to guide the generation of the object, but the masks, generated by unstable human parsers, often produce unreliable results with fabric residues due to wrong segmentation. Moreover, large mask regions can lose spatial structure and identity information, requiring extra conditional inputs to compensate, which increases model instability and reduces efficiency. To tackle the problem, we present a novel Mask-Free virtual Try-ON (MFTON) framework. Specifically, we propose a mask-free strategy to eliminate all denoising conditions except for clothing and person images, thereby directly extracting spatial structure and identity information from the person image to improve efficiency and reduce instability. Additionally, to optimize the generated clothing regions, we propose a clothing texture-aware attention mechanism to enable the model to focus on texture generation with significant visual differences. We then introduce a geometric detail capture loss to further enable the model to capture more high-frequency information. Finally, we propose an appearance consistency inference method to reduce the initial randomness of the sampling process significantly. Extensive experiments on popular datasets demonstrate that our method outperforms state-of-the-art virtual try-on methods.

NeurIPS Conference 2025 Conference Paper

Mitigating Occlusions in Virtual Try-On via A Simple-Yet-Effective Mask-Free Framework

  • Chenghu Du
  • Shengwu Xiong
  • Junyin Wang
  • Yi Rong
  • Shili Xiong

This paper investigates the occlusion problems in virtual try-on (VTON) tasks. According to how they affect the try-on results, the occlusion issues of existing VTON methods can be grouped into two categories: (1) Inherent Occlusions, which are the ghosts of the clothing from reference input images that exist in the try-on results. (2) Acquired Occlusions, where the spatial structures of the generated human body parts are disrupted and appear unreasonable. To this end, we analyze the causes of these two types of occlusions, and propose a novel mask-free VTON framework based on our analysis to deal with these occlusions effectively. In this framework, we develop two simple-yet-powerful operations: (1) The background pre-replacement operation prevents the model from confusing the target clothing information with the human body or image background, thereby mitigating inherent occlusions. (2) The covering-and-eliminating operation enhances the model's ability of understanding and modeling human semantic structures, leading to more realistic human body generation and thus reducing acquired occlusions. Moreover, our method is highly generalizable, which can be applied in in-the-wild scenarios, and our proposed operations can also be easily integrated into different generative network architectures (e. g. , GANs and diffusion models) in a plug-and-play manner. Extensive experiments on three VTON datasets validate the effectiveness and generalization ability of our method. Both qualitative and quantitative results demonstrate that our method outperforms recently proposed VTON benchmarks.

AAAI Conference 2025 Conference Paper

RealisID: Scale-Robust and Fine-Controllable Identity Customization via Local and Global Complementation

  • Zhaoyang Sun
  • Fei Du
  • Weihua Chen
  • Fan Wang
  • Yaxiong Chen
  • Yi Rong
  • Shengwu Xiong

Recently, the success of text-to-image synthesis has greatly advanced the development of identity customization techniques, whose main goal is to produce realistic identity-specific photographs based on text prompts and reference face images. However, it is difficult for existing identity customization methods to simultaneously meet the various requirements of different real-world applications, including the identity fidelity of small face, the control of face location, pose and expression, as well as the customization of multiple persons. To this end, we propose a scale-robust and fine-controllable method, namely RealisID, which learns different control capabilities through the cooperation between a pair of local and global branches. Specifically, by using cropping and up-sampling operations to filter out face-irrelevant information, the local branch concentrates the fine control of facial details and the scale-robust identity fidelity within the face region. Meanwhile, the global branch manages the overall harmony of the entire image. It also controls the face location by taking the location guidance as input. As a result, RealisID can benefit from the complementarity of these two branches. Finally, by implementing our branches with two different variants of ControlNet, our method can be easily extended to handle multi-person customization, even only trained on single-person datasets. Extensive experiments and ablation studies indicate the effectiveness of RealisID and verify its ability in fulfilling all the requirements mentioned above.

EAAI Journal 2024 Journal Article

A Fine Rendering High-Resolution Makeup Transfer network via inversion-editing strategy

  • Zhaoyang Sun
  • Shengwu Xiong
  • Yaxiong Chen
  • Yi Rong

While current makeup transfer methods have made progress in realism and color fidelity, they struggle with capturing texture details and producing high-resolution images, limiting their practical utility. To address these challenges, we propose a Fine Rendering High-resolution Makeup Transfer (FRHMT) network, which leverages a powerful style-based generator and introduces a novel inversion-editing strategy tailored for makeup transfer. Concretely, in the inversion phase, considering the semantic decoupling properties in the latent space, we design a Hierarchical Residual Inversion (HRI), which projects the image onto high-dimensional feature maps in coarse layers and low-dimensional style codes in fine layers. This design effectively restores the content information of the image while maintaining flexibility in editing the makeup styles. In the editing phase, the Makeup Modulation Module (MMM) learns two mapping networks to adjust the latent variables of the source image based on those of the reference image. This modification occurs in fine layers to transfer the makeup information and preserve the content information. With new network structures and customized loss functions, our training eliminates cumbersome pseudo-paired data synthesis and unstable adversarial learning. Extensive experiments have demonstrated that our method outperforms existing methods in both image quality and makeup similarity through quantitative and qualitative analysis. Additionally, we address the lack of high-resolution data by collecting a dataset of 9716 face images with a resolution of 1024 × 1024. In conclusion, our framework offers a novel artificial intelligence (AI) implementation of makeup transfer in engineering, with the collected dataset holding substantial value for further advancements in AI research.

AAAI Conference 2024 Conference Paper

CRA-PCN: Point Cloud Completion with Intra- and Inter-level Cross-Resolution Transformers

  • Yi Rong
  • Haoran Zhou
  • Lixin Yuan
  • Cheng Mei
  • Jiahao Wang
  • Tong Lu

Point cloud completion is an indispensable task for recovering complete point clouds due to incompleteness caused by occlusion, limited sensor resolution, etc. The family of coarse-to-fine generation architectures has recently exhibited great success in point cloud completion and gradually became mainstream. In this work, we unveil one of the key ingredients behind these methods: meticulously devised feature extraction operations with explicit cross-resolution aggregation. We present Cross-Resolution Transformer that efficiently performs cross-resolution aggregation with local attention mechanisms. With the help of our recursive designs, the proposed operation can capture more scales of features than common aggregation operations, which is beneficial for capturing fine geometric characteristics. While prior methodologies have ventured into various manifestations of inter-level cross-resolution aggregation, the effectiveness of intra-level one and their combination has not been analyzed. With unified designs, Cross-Resolution Transformer can perform intra- or inter-level cross-resolution aggregation by switching inputs. We integrate two forms of Cross-Resolution Transformers into one up-sampling block for point generation, and following the coarse-to-fine manner, we construct CRA-PCN to incrementally predict complete shapes with stacked up-sampling blocks. Extensive experiments demonstrate that our method outperforms state-of-the-art methods by a large margin on several widely used benchmarks. Codes are available at https://github.com/EasyRy/CRA-PCN.

AAAI Conference 2024 Conference Paper

CycleVTON: A Cycle Mapping Framework for Parser-Free Virtual Try-On

  • Chenghu Du
  • Junyin Wang
  • Yi Rong
  • Shuqing Liu
  • Kai Liu
  • Shengwu Xiong

Image-based virtual try-on aims to transfer a target clothing onto a specific person. A significant challenge is arbitrarily matched clothing and person lack corresponding ground truth to supervised learning. A recent pioneering work leveraged an improved cycleGAN to enable one network to generate the desired image for another network during training. However, there is no difference in the result distribution before and after the clothing changes. Therefore, using two different networks is unnecessary and may even increase the difficulty of convergence. Furthermore, the introduced human parsing used to provide body structure information in the input also have a negative impact on the try-on result. How to employ a single network for supervised learning while eliminating human parsing? To tackle these issues, we present a Cycle mapping Virtual Try-On Network (CycleVTON), which can produce photo-realistic try-on results by using a cycle mapping framework without the parser. In particular, we introduce a flow constraint loss to achieve supervised learning of arbitrarily matched clothing and person as inputs to the deformer, thus naturally mimicking the interaction between clothing and the human body. Additionally, we design a skin generation strategy that can adapt to the shape of the target clothing by dynamically adjusting the skin region, i.e., by first removing and then filling skin areas. Extensive experiments conducted on challenging benchmarks demonstrate that our proposed method exhibits superior performance compared to state-of-the-art methods.

NeurIPS Conference 2024 Conference Paper

SHMT: Self-supervised Hierarchical Makeup Transfer via Latent Diffusion Models

  • Zhaoyang Sun
  • Shengwu Xiong
  • Yaxiong Chen
  • Fei Du
  • Weihua Chen
  • Fan Wang
  • Yi Rong

This paper studies the challenging task of makeup transfer, which aims to apply diverse makeup styles precisely and naturally to a given facial image. Due to the absence of paired data, current methods typically synthesize sub-optimal pseudo ground truths to guide the model training, resulting in low makeup fidelity. Additionally, different makeup styles generally have varying effects on the person face, but existing methods struggle to deal with this diversity. To address these issues, we propose a novel Self-supervised Hierarchical Makeup Transfer (SHMT) method via latent diffusion models. Following a "decoupling-and-reconstruction" paradigm, SHMT works in a self-supervised manner, freeing itself from the misguidance of imprecise pseudo-paired data. Furthermore, to accommodate a variety of makeup styles, hierarchical texture details are decomposed via a Laplacian pyramid and selectively introduced to the content representation. Finally, we design a novel Iterative Dual Alignment (IDA) module that dynamically adjusts the injection condition of the diffusion model, allowing the alignment errors caused by the domain gap between content and makeup representations to be corrected. Extensive quantitative and qualitative analyses demonstrate the effectiveness of our method. Our code is available at https: //github. com/Snowfallingplum/SHMT.

AAAI Conference 2023 Conference Paper

ESPT: A Self-Supervised Episodic Spatial Pretext Task for Improving Few-Shot Learning

  • Yi Rong
  • Xiongbo Lu
  • Zhaoyang Sun
  • Yaxiong Chen
  • Shengwu Xiong

Self-supervised learning (SSL) techniques have recently been integrated into the few-shot learning (FSL) framework and have shown promising results in improving the few-shot image classification performance. However, existing SSL approaches used in FSL typically seek the supervision signals from the global embedding of every single image. Therefore, during the episodic training of FSL, these methods cannot capture and fully utilize the local visual information in image samples and the data structure information of the whole episode, which are beneficial to FSL. To this end, we propose to augment the few-shot learning objective with a novel self-supervised Episodic Spatial Pretext Task (ESPT). Specifically, for each few-shot episode, we generate its corresponding transformed episode by applying a random geometric transformation to all the images in it. Based on these, our ESPT objective is defined as maximizing the local spatial relationship consistency between the original episode and the transformed one. With this definition, the ESPT-augmented FSL objective promotes learning more transferable feature representations that capture the local spatial features of different images and their inter-relational structural information in each input episode, thus enabling the model to generalize better to new categories with only a few samples. Extensive experiments indicate that our ESPT method achieves new state-of-the-art performance for few-shot image classification on three mainstay benchmark datasets. The source code will be available at: https://github.com/Whut-YiRong/ESPT.

v2026.09.13