Arrow Research search
Back to IJCAI

IJCAI 2025

Instructing Text-to-Image Diffusion Models via Classifier-Guided Semantic Optimization

Conference Paper Agent-based and Multi-agent Systems Artificial Intelligence

Abstract

Text-to-image diffusion models have emerged as powerful tools for high-quality image generation and editing. Many existing approaches rely on text prompts as editing guidance. However, these methods are constrained by the need for manual prompt crafting, which can be time-consuming, introduce irrelevant details, and significantly limit editing performance. In this work, we propose optimizing semantic embeddings guided by attribute classifiers to steer text-to-image models toward desired edits, without relying on text prompts or requiring any training or fine-tuning of the diffusion model. We utilize classifiers to learn precise semantic embeddings at the dataset level. The learned embeddings are theoretically justified as the optimal representation of attribute semantics, enabling disentangled and accurate edits. Experiments further demonstrate that our method achieves high levels of disentanglement and strong generalization across different domains of data. Code is available at https: //github. com/Chang-yuanyuan/CASO.

Authors

Keywords

  • Computer Vision: CV: Image and video synthesis and generation

Context

Venue
International Joint Conference on Artificial Intelligence
Archive span
1969-2025
Indexed papers
14525
Paper id
299332012358368365
v2026.09.13