NEWS & INSIGHTS

Understanding the Limitations of GAN Models

An earlier article introduced generative adversarial networks (GANs) in the context of deepfakes. Their operating principles suggest a broad range of generative applications, but GANs also face technical challenges. This article examines those limitations.

An earlier article introduced GANs and their use in deepfakes. Although GANs can generate realistic samples, their training presents several challenges.
Illustration: Understanding the Limitations of GAN Models
If you click on the image, you can see the GAN post.

[The Generator's Problem]
As described above, the GAN model is divided into a generator and a discriminator.
Illustration: Understanding the Limitations of GAN Models
GAN training uses an adversarial minimax objective: the discriminator learns to distinguish real and generated samples, while the generator learns to fool it. 
In practice, neither network is necessarily optimized perfectly at each step. A discriminator that becomes too strong may provide weak or unhelpful gradients to the generator. 
If the discriminator converges reliably toward an appropriate solution, it can guide the generator. Otherwise, the generator may also fail to converge.
The generator relies on useful feedback from the discriminator. If that feedback becomes less informative, the generator's performance may also deteriorate. 

[Oscillation of the model]
GAN training alternates updates to the discriminator and generator in an adversarial process. This interaction can create stability problems. 
Updates to the two networks can counteract each other. The models may repeatedly respond to the other's changes without converging to a stable solution. This behavior is known as oscillation.

[Mode collapse]
Real data often has multiple modes. Mode collapse occurs when the generator represents only a narrow subset of that diversity. 
For example, suppose you have a hand-written image of 0 on paper, a hand-written image of 1 on paper, a hand-written picture of 2 on paper, and a hand-written image of 3 on paper. If you use these images to learn the GAN model, you can see that there are four modes. 
A GAN trained on data with four modes may generate samples from only one of them. This is known as mode collapse. 
For example, training on handwritten-digit data may produce images of the same digit repeatedly, such as 2. The generator finds a narrow type of output that fools the discriminator instead of covering the full data distribution.
This problem can worsen when the discriminator is inadequate, training oscillates or both occur together. The generator then captures only part of the training-data distribution. 

[Decisions]
Balancing the generator and discriminator so that training converges remains a challenge for GANs. Many architectural and training approaches have been proposed to address it. 

Read the Korean source

This article reflects the information available when it was published. Contact us to discuss your circumstances.
Discuss this topic ↗All articles

Put your IP strategy into practice.