Aesthetic-Aware Text to Image Synthesis

Document Type

Conference Paper

Publication Date

5-7-2020

Publication Source

2020 54th Annual Conference on Information Sciences and Systems (CISS)

Abstract

Synthesizing an image from natural language description is an important task to many applications such as photo-editing, art generation, and computer aided-design. However, to synthesize an appealing image from the text, image aesthetics criteria should be maintained. In this study, we propose a new framework which first generates a set of mask maps from the input text via mask map generator (MG), and then we compute and rank the image aesthetics score for all generated mask maps via Pre-IG Aesthetic Ranking that contains two composition rules, i.e., the rule of thirds along with the rule of formal balance. At the next stage, we feed the subset of the mask maps, which are the highest, lowest, and the average aesthetic scores, to image generator (IG). The photorealistic images are ranked at the second round through, namely Post-IG Aesthetic Ranking, to determine the lowest aesthetic score and return the most appealing generated image. The experiments on COCO-stuff dataset demonstrate that our framework yields better results compared to previous text-to- image models.

ISBN/ISSN

978-1-7281-4085-8

Comments

The first author would like to thank Umm Al-Qura University, in Saudi Arabia, for the continuous support. We also thank NVIDIA Corporation for the donation of GPU.

Publisher

IEEE xplore

Keywords

Text-to-image, mask corpus dataset, image synthesis, central anchor point, image aesthetics


Share

COinS