Arrow Research search

Author name cluster

Jin Yuan

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

5 papers
1 author row

Possible papers

5

AAAI Conference 2025 Conference Paper

DECIDER: Difference-aware Contrastive Diffusion Model with Adversarial Perturbations for Image Change Captioning

  • Guojin Zhong
  • Jinhong Hu
  • Jiajun Chen
  • Jin Yuan
  • Wenbo Pan

Image change captioning (ICC) poses great challenges stemming from describing subtle differences between two similar images in natural language, significantly increasing the complexity of feature extraction and cross-modal learning compared to the image captioning task. Existing ICC methods often suffer from two key challenges: 1) Massive irrelevant information of uni-image features leads to suboptimal visual difference representations; 2) Imprecise inter-modality correspondence degrades the quality of generated captions. This paper proposes a Difference-aware Contrastive Diffusion Model with Adversarial Perturbations (DECIDER) for ICC due to the excellent performance of diffusion models in image/text generation. Technically, difference-aware cross-modal learning is developed to suppress irrelevant information and learn compact yet robust visual difference representations. This is achieved by optimizing a novel objective mathematically derived from the information bottleneck principle that excels in filtering redundant features and highlighting differences. Furthermore, we propose to dynamically generate ``hard'' positive and negative samples via adversarial perturbations, which are involved in contrastive diffusion training with a tighter variational bound. This design encourages our DECIDER to excavate and construct complex correspondences between visual differences and captions, thereby improving generalization performance. Extensive experiments on four datasets demonstrate that DECIDER significantly exceeds state-of-the-art performance.

AAAI Conference 2025 Conference Paper

SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioning

  • Xu Zhang
  • Jin Yuan
  • Hanwang Zhang
  • Guojin Zhong
  • Yongsheng Zang
  • Jiacheng Lin
  • Zhiyong Li

Controllable image semantic understanding tasks, such as captioning or segmentation, necessitate users to input a prompt (e.g., text or bounding boxes) to predict a unique outcome, presenting challenges such as high-cost prompt input or limited information output. This paper introduces a new task ``Image Collaborative Segmentation and Captioning'' (SegCaptioning), which aims to translate a straightforward prompt, like a bounding box around an object, into diverse semantic interpretations represented by (caption, masks) pairs, allowing flexible result selection by users. This task poses significant challenges, including accurately capturing a user's intention from a minimal prompt while simultaneously predicting multiple semantically aligned caption words and masks. Technically, we propose a novel Scene Graph Guided Diffusion Model that leverages structured scene graph features for correlated mask-caption prediction. Initially, we introduce a Prompt-Centric Scene Graph Adaptor to map a user's prompt to a scene graph, effectively capturing his intention. Subsequently, we employ a diffusion process incorporating a Scene Graph Guided Bimodal Transformer to predict correlated caption-mask pairs by uncovering intricate correlations between them. To ensure accurate alignment, we design a Multi-Entities Contrastive Learning loss to explicitly align visual and textual entities by considering inter-modal similarity, resulting in well-aligned caption-mask pairs. Extensive experiments conducted on two datasets demonstrate that SGDiff achieves superior performance in SegCaptioning, yielding promising results for both captioning and segmentation tasks with minimal prompt input.

IJCAI Conference 2024 Conference Paper

CF-Deformable DETR: An End-to-End Alignment-Free Model for Weakly Aligned Visible-Infrared Object Detection

  • Haolong Fu
  • Jin Yuan
  • Guojin Zhong
  • Xuan He
  • Jiacheng Lin
  • Zhiyong Li

Weakly aligned visible-infrared object detection poses significant challenges due to the imprecise alignment between visible and infrared images. Most existing methods explore the alignment strategies between visible and infrared images, yielding unbearable computation costs. This paper first proposes an end-to-end alignment-free architecture Cross-modal Fusion Deformable DEtection TRansformer (``CF-Deformable DETR'') for weakly aligned visible-infrared object detection. Abandoning the traditional image alignment, CF-Deformable DETR introduces a simple yet effective cross-modal deformable attention mechanism to directly implement automatic cross-modal point mapping, generating well-aligned bimodal features with high efficiency. Moreover, we design a Point-level Feature Consistency Loss to guide the cross-modal point mapping, ensuring the consistency of paired features to support the following fusion. Extensive experiments are conducted on three benchmark datasets. The experimental results demonstrate that CF-Deformable DETR achieves close accuracy on weakly aligned and strictly aligned data as well as maintains stable performance to a certain extent against various offset degrees of weakly aligned data. Code is available at https: //github. com/116508/CF-Deformable-DETR.

JBHI Journal 2022 Journal Article

Multi-Discriminator Adversarial Convolutional Network for Nerve Fiber Segmentation in Confocal Corneal Microscopy Images

  • Changqing Yang
  • Xinxin Zhou
  • Weifang Zhu
  • Dehui Xiang
  • Zhongyue Chen
  • Jin Yuan
  • Xinjian Chen
  • Fei Shi

Quantitative measurements of corneal sub-basal nerves are biomarkers for many ocular surface disorders and are also important for early diagnosis and assessment of progression of neurodegenerative diseases. This paper aims to develop an automatic method for nerve fiber segmentation from in vivo corneal confocal microscopy (CCM) images, which is fundamental for nerve morphology quantification. A novel multi-discriminator adversarial convolutional network (MDACN) is proposed, where both the generator and the two discriminators emphasize multi-scale feature representations. The generator is a U-shaped fully convolutional network with multi-scale split and concatenate blocks, and the two discriminators have different effective receptive fields, sensitive to features of different scales. A novel loss function is also proposed which enables the network to pay more attention to thin fibers. The MDACN framework was evaluated on four datasets. Experiment results show that our method has excellent segmentation performance for corneal nerve fibers and outperforms some state-of-the-art methods.

EAAI Journal 2007 Journal Article

Integrating relevance vector machines and genetic algorithms for optimization of seed-separating process

  • Jin Yuan
  • Kesheng Wang
  • Tao Yu
  • Minglung Fang

A hybrid intelligent approach based on relevance vector machines (RVMs) and genetic algorithms (GAs) has been developed for optimal control of parameters of nonlinear manufacturing processes. It concerns the finding of the near-optimal control parameters of the nonlinear discrete manufacturing process with a specific objective. First, the nonlinear process with measurement noise is regressed by the relevance vector learning mechanism based on a kernel-based Bayesian framework. For minimizing the approximate error, uniform design sampling, online incremental learning and cross-validation are used in the learning process of RVMs. Such well-trained models become a specialized process simulation tool, which is valuable in prediction and optimization of nonlinear processes. Next, the near-optimal setpoints of the control system, which maximize the objective function, are sought by GAs from the numerous values of the objective function obtained from the simulation. As a case study, the seed separator system (5XZW-1. 5) is used for evaluating the proposed intelligent approach. The control parameters to reach the maximum weighted objective, which combine the system output and evaluation functions, are optimized. The experimental results show the effectiveness of the proposed hybrid approach.

v2026.09.13