EAAI Journal 2025 Journal Article
A visual data unsupervised disentangled representation learning framework: Contrast disentanglement based on variational auto-encoder
- Chengquan Huang
- Jianghai Cai
- Senyan Luo
- Shunxia Wang
- Guiyan Yang
- Huan Lei
- Lihua Zhou
To discover and learn interpretable factors behind the visual data, many approaches use extra regularization terms in learning disentangled representations, which lead to poor results between disentanglement and generative quality. The variational auto-encoder has the ability to learn more semantic information, and the traversal of the generated images along different factor directions shows meaningful and interpretable variations in the latent space. Therefore, we exploit the scalability and training stability of the variational auto-encoder, and focus on meaningful traversal directions in the latent space. We propose contrast disentanglement based on variational auto-encoder, a visual data unsupervised disentangled representation learning framework. Specifically, we explore meaningful interpretable directions in the latent space by constructing the encoder of the typical network module, and then obtain interpretable directions for candidate traversals of target variation images. In this way, the interpretable directions with rich semantic information and disentanglement properties can be obtained. Furthermore, unlike the existing methods, to further improve the ability of interaction and collaborative learning between latent factors and learn more robust and generalized representations, we design the variation space based on disentangled encoders from the contrastive learning perspective to simulate the various variation of the image data, then extract disentangled representations and generate images. Extensive experiments on six disentanglement datasets have demonstrated that the proposed method achieves competing performance on both quantitative metrics and visual quality. Among them, the proposed method achieves better factor variational auto-encoder score and β-variational auto-encoder score (0. 97 ± 0. 04 and 0. 99 ± 0. 01, respectively) on the 3Dshapes dataset compared to existing methods.