Text-to-image synthesis based on conditional augmentation DCGAN network
preprint
OA: closed
Abstract
Abstract In this paper, Aiming at the problems of missing image structure and unreal image generated under the condition of text in the Deep Convolutional Generative Adversarial Network (DCGAN) model, In the proposed conditional augmentation DCGAN (CA-DCGAN) model, the initial step entails utilizing a deep convolution network in conjunction with a cyclic text encoder to encode the input text. This encoding process facilitates the extraction of a feature vector representation of the text, which serves as a fundamental aspect of the model’s architecture. Secondly, a conditional augmentation model (CA) is introduced after the representation of text feature vector by an additional conditional variable is generated through the mean and covariance matrix of text feature vector to replace the original text features with a new vector, combine the conditional variable and random noise for the generator’s input, and add KL loss to its training to prevent overfitting, so that the model can converge better. The experimental results show that the image quality generated by the proposed method on oxford-102-flowers and cub-200 data sets is better than that of align GAN, GAN-CLS and GAN-INT-CLS models, and is closer to the real samples. The Inception Score value is increased by 12% and 9.4% respectively, and the FID value is reduced by 9.6% and 8.7% respectively, which validates the effectiveness of the method proposed in this paper.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00