Journals & Magazines >IEEE Transactions on Image Pr... >Volume: 30

Multi-Sentence Auxiliary Adversarial Networks for Fine-Grained Text-to-Image Synthesis

Download PDF
Download References
Request Permissions
Save to
Alerts

Abstract:

Due to the development of Generative Adversarial Networks (GANs), significant progress has been achieved in text-to-image synthesis task. However, most previous works hav...Show More

Metadata

Abstract:

Due to the development of Generative Adversarial Networks (GANs), significant progress has been achieved in text-to-image synthesis task. However, most previous works have only focus on learning the semantic consistency between paired images and sentences, without exploring the semantic correlation between different yet related sentences that describe the same image, which leads to significant visual variation among the synthesized images. Accordingly, in this article, we propose a new method for text-to-image synthesis, dubbed Multi-sentence Auxiliary Generative Adversarial Networks (MA-GAN); this approach not only improves the generation quality but also guarantees the generation similarity of related sentences by exploring the semantic correlation between different sentences describing the same image. More specifically, we propose a Single-sentence Generation and Multi-sentence Discrimination (SGMD) module that explores the semantic correlation between multiple related sentences in order to reduce the variation between their generated images and enhance the reliability of the generated results. Moreover, a Progressive Negative Sample Selection mechanism (PNSS) is designed to mine more suitable negative samples for training, which can effectively promote detailed discrimination ability in the generative model and facilitate the generation of more fine-grained results. Extensive experiments on Oxford-102 and CUB datasets reveal that our MA-GAN significantly outperforms the state-of-the-art methods.

Published in: IEEE Transactions on Image Processing ( Volume: 30)

Page(s): 2798 - 2809

Date of Publication: 02 February 2021

ISSN Information:

PubMed ID: 33531300

DOI: 10.1109/TIP.2021.3055062

Funding Agency:

Contents

I. Introduction

With the rapid development of computer vision and natural language processing, text-to-image synthesis has recently come to attract considerable attention, which refers to generating a visually realistic image that matches a given textual sentence. Since it is very difficult to parse sentences and bridge the semantic gap between sentence and image, text-to-image synthesis remains an open problem. There are two major challenges associated with the text-to-image synthesis task. One is visual realism, as generating rich yet detailed images using text with a limited number of words is difficult. The other is semantic consistency, as building the relationships between text semantics and visual features is problematic.

References is not available for this document.

Multi-Sentence Auxiliary Adversarial Networks for Fine-Grained Text-to-Image Synthesis

Abstract:

Metadata

Abstract:

ISSN Information:

Funding Agency:

I. Introduction

References

IEEE Account

Purchase Details

Profile Information

Need Help?

Multi-Sentence Auxiliary Adversarial Networks for Fine-Grained Text-to-Image Synthesis

Alerts

Abstract:

Metadata

Abstract:

ISSN Information:

Funding Agency:

I. Introduction

References