Aesthetic-Aware Text to Image Synthesis
Document Type
Conference Paper
Publication Date
5-7-2020
Publication Source
2020 54th Annual Conference on Information Sciences and Systems (CISS)
Abstract
Synthesizing an image from natural language description is an important task to many applications such as photo-editing, art generation, and computer aided-design. However, to synthesize an appealing image from the text, image aesthetics criteria should be maintained. In this study, we propose a new framework which first generates a set of mask maps from the input text via mask map generator (MG), and then we compute and rank the image aesthetics score for all generated mask maps via Pre-IG Aesthetic Ranking that contains two composition rules, i.e., the rule of thirds along with the rule of formal balance. At the next stage, we feed the subset of the mask maps, which are the highest, lowest, and the average aesthetic scores, to image generator (IG). The photorealistic images are ranked at the second round through, namely Post-IG Aesthetic Ranking, to determine the lowest aesthetic score and return the most appealing generated image. The experiments on COCO-stuff dataset demonstrate that our framework yields better results compared to previous text-to- image models.
ISBN/ISSN
978-1-7281-4085-8
Publisher
IEEE xplore
Keywords
Text-to-image, mask corpus dataset, image synthesis, central anchor point, image aesthetics
eCommons Citation
Baraheem, Samah Saeed and Nguyen, Tam V., "Aesthetic-Aware Text to Image Synthesis" (2020). Computer Science Faculty Publications. 205.
https://ecommons.udayton.edu/cps_fac_pub/205
COinS

Comments
The first author would like to thank Umm Al-Qura University, in Saudi Arabia, for the continuous support. We also thank NVIDIA Corporation for the donation of GPU.