This article explains two techniques used in deep learning: mixup training and label smoothing.
If you click on the image, you can see the full text.
Mixup training and label smoothing can improve generalization without changing the network architecture..
[Mixup Training]
Mixup randomly selects two training examples and forms weighted combinations of their inputs and labels. It is a data-augmentation technique.
Example: cat and dog image classificationConsider the following example.
A cat image may have the one-hot label [0, 1], while a dog image has [1, 0].

Common ways to label training data
Mixup TrainingMixup combines the two images and their labels using the same weight.
The resulting mixed example is used during training.
When performing mixup training, not only do you mix two images, but you also mix labels.The mixing coefficient is lambda (λ), which controls the contribution of each example.Depending on the chosen beta-distribution parameter, samples may favor one image or mix them more evenly. Lambda is sampled from a beta distribution.

Mixup training
A 50/50 mixture of a cat and dog image receives the label [0.5, 0.5], teaching the model to interpolate between the examples. Avoid overfittingis one potential benefit. Mixup also creates additional training examples that can improve generalization.
[Label Smoothing]
Label Smoothing is a technique for smoothing labels to improve generalization performance. The big difference with Mixup training is that it smoothes only the label without changing the training image.
Consider an image-classification task with three classes.
A one-hot target assigns probability 1 to the correct class and 0 to the others. This is a hard label.
| Dog | Cat | Bird | |
| Image Data 1 | 1 | 0 | 0 |
| Image Data 2 | 0 | 1 | 0 |
| Image Data 3 | 1 | 0 | 0 |
Hard labels
Label smoothing reduces the target probability for the correct class and assigns a small positive probability to the other classes.This is called a soft label.
| Dog | Cat | Bird | |
| Image Data 1 | 0.9 | 0.05 | 0.05 |
| Image Data 2 | 0.05 | 0.9 | 0.05 |
| Image Data 3 | 0.9 | 0.05 | 0.05 |
Labels after smoothing
Label smoothing acts as regularization by discouraging overly confident predictions.Because training labels may contain errors, forcing extreme confidence in every target can be harmful. Softer targets can reduce that effect. Label smoothing softens training labels to improve generalization.
Both mixup and label smoothing can discourage overconfidence and improve generalization.
This article reflects the information available when it was published. Contact us to discuss your circumstances.

