Text-to-image via mask anchor points

Document Type

Article

Publication Date

2-13-2020

Publication Source

Pattern Recognition Letters

Abstract

Text-to-image is a process of generating an image from the input text. It has a variety of applications in art generation, computer-aided design, and data synthesis. In this paper, we propose a new framework which leverages mask anchor points to incorporate two major steps in the image synthesis. In the first step, the mask image is generated from the input text and the mask dataset. In the second step, the mask image is fed into the state-of-the-art mask-to-image generator. Note that the mask image captures the semantic information and the location relationship via the anchor points. We also developed a user-friendly interface which helps parse the input text into the meaningful semantic objects. As a result, our framework is able to produce clear, reasonable, and more realistic images. The experiments on the most challenging COCO-stuff dataset illustrate the superiority of our proposed approach over the previous state of the arts.

Inclusive pages

25-32

ISBN/ISSN

0167-8655

Publisher

Elsevier

Volume

133

Keywords

Text-to-image, Mask dataset, Image synthesis, Anchor points


Share

COinS