NEWS & INSIGHTS

Research Paper Analysis: StarGAN

StarGAN (Unified Generative Adversarial Networks for Multi-Domain Image-to-Image Translation)

StarGAN (Unified Generative Adversarial Networks for Multi-Domain Image-to-Image Translation)
This paper proposes a StarGAN network that enables translation in multi-domain situations. 
Illustration: Research Paper Analysis: StarGAN
If you click on the image, you can see the full text.

[Content Summary]
Earlier approaches such as CycleGAN require separate models for different pairs of domains. The examples below compare their results with StarGAN. 
Illustration: Research Paper Analysis: StarGAN
But... StarGAN can translate multiple attributes at once using only one network. 
The examples show StarGAN translating facial attributes such as age, hair color and gender, including multiple attributes at once. 
For example, the input is a young woman with blonde hair:(INPUT)The combined hair, gender and age transformation produces an older man with black hair: (H+G+A)
Illustration: Research Paper Analysis: StarGAN
StarGAN Examples

[Background Technology Explained]
Generative models
A generative model learns patterns in data and produces new samples resembling that data. It can generate images, text and other data types. Image generation is the focus of this article.
Discriminative and Generative models are as follows.
Illustration: Research Paper Analysis: StarGAN
Generative Model and Discriminative Model
A discriminative model learns to assign input data to a class, often by learning decision boundaries. 
Meanwhile, a generative model learns a data distribution and produces new samplesresembling the training data.Its task is generation rather than classification.Generative modeling can involve learning the joint distribution of relevant features.
For a face model, features such as eye and nose dimensions are related. Learning their joint distribution allows the model to sample combinations that resemble real faces, rather than independently choosing the most frequent value of each feature. 
As a result, the aim is to approximate the underlying distribution represented by the training data. 
One approach to generative modeling is the generative adversarial network (GAN). 

GAN: Generative Adversarial Networks
In my previous post, I’ve written about GANs, but this time I’ll explain them in more detail. [Related GAN article]
GAN utilizes two networks: a generator and a discriminator.After training, the generator can transform a sampled latent vector into a plausible image. 
The generator produces a synthetic sample from a latent vector. The discriminator estimates whether an input sample is real training data or generated data.The two networks are trained together. 
The generator learns to produce samples that the discriminator classifies as real. 
The discriminator learns to distinguish generated samples from real training samples. 
Here, the important thing is the goal of making the generated distribution resemble the real data distribution. 

Conditional GAN(cGAN)
Normally, when you create a GAN, you can change the latent vector to create a new image. However, it can be difficult to create an intended image using a GAN by changing only the latent vector.
A conditional GAN adds information that controls the desired output. For example, a class label can specify the type of image to generate.
The conditional GAN objective includes this additional condition, often denoted y. 
Illustration: Research Paper Analysis: StarGAN
*****
Check the picture below. z(latent vector) and c(conditional vector) are supplied to the generator G. The condition c can specify the target class.The generator uses z together with c to produce a sample for that class. 
Illustration: Research Paper Analysis: StarGAN
*****
The generated sample, G(z, c),is supplied to the discriminator D along with the condition c.Real samples with the corresponding class label, real data(x)are also used in training. This encourages class-consistent generation. 
Illustration: Research Paper Analysis: StarGAN

Pix2pix
Image-to-image translation maps an input image to a corresponding image in another domain, such as translating an outline into a photograph.Pix2pix is an example of an image-to-image translation architecture.It uses a conditional GAN.
In pix2pix, the conditioning input is an image. The generator produces an output consistent with that input, rather than relying only on a class label.
For example, if you have a hand drawn picture (x) that has only a contour, and you want it to be changed to an actual image (G(x)), then the hand drawn picture (x) that only has a contour is a condition here. Here, pix2pix is a form of learning that fills the inside of the contours when the hand-drawn figure (x) is inserted. 
Illustration: Research Paper Analysis: StarGAN
Pix2pix is trained on paired examples from domains X and Y.For example, an outline and its corresponding photograph can form a training pair. 
But... Creating a pair of training data is often difficult. A photograph and a matching painted scene can be difficult to obtain as a pair. Creating such aligned training data may be costly and time-consuming. There is a limit to Pix2pix in this regard.

Limitations of cGAN
You can simply use cGAN to create a zebra image using a specific image during the task to change the horse into a zebra. 
Illustration: Research Paper Analysis: StarGAN
But... Without suitable constraints, a generator may output a plausible zebra that does not preserve the content of the particular horse image supplied. 
Illustration: Research Paper Analysis: StarGAN
A discriminator judging only realism in the target domain may accept such an output. The generator therefore needs a constraint linking its output to the input's content. 

CycleGAN
CycleGAN addresses unpaired image-to-image translation using cycle consistency.A translated image is mapped back toward the original domain. 
Illustration: Research Paper Analysis: StarGAN
Starting with x, the model generates G(x) and then applies the reverse mapping F to reconstruct x. 
Cycle-consistency loss penalizes differences between the original and reconstructed images, encouraging preservation of content while changing domain-specific appearance. 
Two translators are used to perform these tasks. 
One is a generator that performs the transformation from X to Y. (G: X → Y) 
The other is a function that performs an inverse mapping that converts from Y to X again. (F:Y → X)
Training encourages F(G(x)) to approximate x and G(F(y)) to approximate y. 
Illustration: Research Paper Analysis: StarGAN

WGAN-GP
WGAN constrains its critic to be 1-Lipschitz. Weight clipping can cause optimization problems, so WGAN-GP uses a gradient penalty to improve training stability. 
Illustration: Research Paper Analysis: StarGAN

StarGAN
Separate models for different domain pairs can be inefficient. StarGAN instead uses StarGAN one modelto translate images across multiple domains. Sharing a model reduces parameter duplication and allows learning from data across domains. 
*****
Illustration: Research Paper Analysis: StarGAN
The generator takes an input image and target-domain label to produce a translated image. 
*****
Illustration: Research Paper Analysis: StarGAN
The translated image is passed back through the generator with the original-domain label. Reconstruction loss encourages recovery of the input and preservation of its identity. 
*****
Illustration: Research Paper Analysis: StarGAN
The discriminator learns to distinguish real and generated images and to classify the domains of real images. 
The generator uses the discriminator's feedback to produce realistic images matching the target domain. 
It is StarGAN that adds cycle-consistency loss to the cGAN and increases the performance of the generator to be able to operate on multiple domains. 
StarGAN combines adversarial, domain-classification and reconstruction losses. 
*****
 The adversarial loss is shown below. 
Illustration: Research Paper Analysis: StarGAN
Adversarial loss adds the concept of changing the default GAN's loss to a specific domain (c). In the case of Adversarial loss, the idea of WGAN-GP can be used to construct a loss value in the form given by penalty.
*****
Domain classification loss is as follows. 
Illustration: Research Paper Analysis: StarGAN
A fake image's domain classification loss is used for the generator, and the learning is carried out so that the image changed to a specific domain can be classified as a target domain. 
The domain classification loss of the real image is used for the discriminator, and the learning is done to classify the real image as the domain value (c') for the real image when it is entered. 
*****
Reconstruction loss is next. 
Illustration: Research Paper Analysis: StarGAN
Reconstruction loss uses cycle consistency loss so that the image created by the generator can retain the content of the original image during the translation process. Then, when restoring to the original image, configure the reconstruction loss so that the actual original data and the final converted data can be similar. 
*****
The final training objectives are shown below. 
Illustration: Research Paper Analysis: StarGAN
The discriminator objective combines the adversarial objective with classification loss on real images.
The generator objective combines the adversarial objective, target-domain classification loss on generated images and reconstruction loss. 
*****
Joint training across datasets A mask vector identifies which dataset's labels should be used.It enables training with datasets that provide different attribute labels. 
For example, a large dataset may label hair color while a smaller dataset labels facial expressions. Their label vectors are combined with a mask indicating the relevant dataset. 
Illustration: Research Paper Analysis: StarGAN
Training only on the smaller expression dataset, denoted SNG, can limit performance.
Joint training, denoted JNT, also uses the larger face dataset. The mask distinguishes the available labels, allowing shared features to improve translation performance. 
Let’s recap some more specifics.
*****
Illustration: Research Paper Analysis: StarGAN
Joint training procedure
For real images, the discriminator learns the available domain labels as well as real/fake discrimination. Domain-classification loss for the discriminator is computed on real images. 
*****
Illustration: Research Paper Analysis: StarGAN
In the illustrated example, the generator receives the input image, CelebA labels (1 0 0 1 1), zeroed RaFD labels and mask (1 0). The mask selects CelebA attributes. With the stated label ordering, the target is a young man with black hair. 
*****
Illustration: Research Paper Analysis: StarGAN
The generated image is then supplied with the original CelebA labels (0 0 1 0 1), zeroed RaFD labels and the same mask. These labels represent the original young woman with brown hair, enabling reconstruction of the input. 
*****
Illustration: Research Paper Analysis: StarGAN
The generator learns to make translated images appear real and match the target domain, while reconstruction helps preserve the person's identity.

Read the Korean source

This article reflects the information available when it was published. Contact us to discuss your circumstances.
Discuss this topic ↗All articles

Put your IP strategy into practice.