BachGAN: High-Resolution Image Synthesis from Salient Object Layout

by   Yandong Li, et al.

We propose a new task towards more practical application for image generation - high-quality image synthesis from salient object layout. This new setting allows users to provide the layout of salient objects only (i.e., foreground bounding boxes and categories), and lets the model complete the drawing with an invented background and a matching foreground. Two main challenges spring from this new task: (i) how to generate fine-grained details and realistic textures without segmentation map input; and (ii) how to create a background and weave it seamlessly into standalone objects. To tackle this, we propose Background Hallucination Generative Adversarial Network (BachGAN), which first selects a set of segmentation maps from a large candidate pool via a background retrieval module, then encodes these candidate layouts via a background fusion module to hallucinate a suitable background for the given objects. By generating the hallucinated background representation dynamically, our model can synthesize high-resolution images with both photo-realistic foreground and integral background. Experiments on Cityscapes and ADE20K datasets demonstrate the advantage of BachGAN over existing methods, measured on both visual fidelity of generated images and visual alignment between output images and input layouts.


page 1

page 5

page 6

page 7

page 12

page 13

page 14

page 15


Generating Multiple Objects at Spatially Distinct Locations

Recent improvements to Generative Adversarial Networks (GANs) have made ...

Semantic Layout Manipulation with High-Resolution Sparse Attention

We tackle the problem of semantic image layout manipulation, which aims ...

Realistic Image Generation using Region-phrase Attention

The Generative Adversarial Network (GAN) has recently been applied to ge...

MC-GAN: Multi-conditional Generative Adversarial Network for Image Synthesis

In this paper, we introduce a new method for generating an object image ...

OneGAN: Simultaneous Unsupervised Learning of Conditional Image Generation, Foreground Segmentation, and Fine-Grained Clustering

We present a method for simultaneously learning, in an unsupervised mann...

FiG-NeRF: Figure-Ground Neural Radiance Fields for 3D Object Category Modelling

We investigate the use of Neural Radiance Fields (NeRF) to learn high qu...

Harvesting Visual Objects from Internet Images via Deep Learning Based Objectness Assessment

The collection of internet images has been growing in an astonishing spe...

Please sign up or login with your details

Forgot password? Click here to reset