# Improving Deep Learning Performance: Mixup Training and Label Smoothing

This article explains two techniques used in deep learning: mixup training and label smoothing.

Source: https://www.iplexlaw.co.kr/en/blog/875313

HOME / NEWS & INSIGHTS NEWS & INSIGHTS Improving Deep Learning Performance: Mixup Training and Label Smoothing This article explains two techniques used in deep learning: mixup training and label smoothing. AI & Software 2023.06.30 published IPLEX 4 min read This article explains two techniques used in deep learning: mixup training and label smoothing. If you click on the image, you can see the full text. Mixup training and label smoothing can improve generalization without changing the network architecture. . [Mixup Training] Mixup randomly selects two training examples and forms weighted combinations of their inputs and labels. It is a data-augmentation technique. Example: cat and dog image classification Consider the following example. A cat image may have the one-hot label [0, 1], while a dog image has [1, 0]. Common ways to label training data Mixup Training Mixup combines the two images and their labels using the same weight. The resulting mixed example is used during training. When performing mixup training, not only do you mix two images, but you also mix labels. The mixing coefficient is lambda (λ), which controls the contribution of each example. Depending on the chosen beta-distribution parameter, samples may favor one image or mix them more evenly. Lambda is sampled from a beta distribution. Mixup training A 50/50 mixture of a cat and dog image receives the label [0.5, 0.5], teaching the model to interpolate between the examples. Avoid overfitting is one potential benefit. Mixup also creates additional training examples that can improve generalization. [Label Smoothing] Label Smoothing is a technique for smoothing labels to improve generalization performance. The big difference with Mixup training is that it smoothes only the label without changing the training image. Consider an image-classification task with three classes. A one-hot target assigns probability 1 to the correct class and 0 to the others. This is a hard label. Dog Cat Bird Image Data 1 1 0 0 Image Data 2 0 1 0 Image Data 3 1 0 0 Hard labels Label smoothing reduces the target probability for the correct class and assigns a small positive probability to the other classes. This is called a soft label. Dog Cat Bird Image Data 1 0.9 0.05 0.05 Image Data 2 0.05 0.9 0.05 Image Data 3 0.9 0.05 0.05 Labels after smoothing Label smoothing acts as regularization by discouraging overly confident predictions. Because training labels may contain errors, forcing extreme confidence in every target can be harmful. Softer targets can reduce that effect. Label smoothing softens training labels to improve generalization. Both mixup and label smoothing can discourage overconfidence and improve generalization. Read the Korean source This article reflects the information available when it was published. Contact us to discuss your circumstances. Discuss this topic ↗ All articles TALK TO IPLEX Discuss your IP questions We consider your technology and business needs together. ↗ Contact us Newer Research Paper Analysis: ResNet — Deep Residual Learning for Image Recognition ↗ Older Facebook AI Paper Analysis: DETR and End-to-End Object Detection with Transformers ↗ Related insights AI & Software 2026.10.01 FiX: fine-grained forgetting in softmax attention Yongduck Kim examines FiX’s feature-wise gates, numerical implementation and paged cache, distinguishing reported gains from unresolved limitations. ↗ Read article AI & Software 2026.09.30 MHAR: Reading earlier layers through different feature subspaces Yongduck Kim examines Multi-Head Attention Residuals: depth routing, reported training results, implementation costs and the relationship between technical features and effects. ↗ Read article AI & Software 2026.09.28 Column: Claude Computer Use and the Data That Trains AI Agents Writing for AI Times, IPLEX Managing Partner Yongduck Kim examines the training data behind computer-operating AI agents through U.S. Patent No. 12,585,862. ↗ Read article

- https://www.iplexlaw.co.kr/en
- https://www.iplexlaw.co.kr/en/blog/category/ai
- https://drive.google.com/file/d/1i_WOe3vVLlK1SmPPkpT1jeede0iBBIgx/view
- https://www.iplexlaw.co.kr/forum/view/875313
- https://www.iplexlaw.co.kr/en/contact
- https://www.iplexlaw.co.kr/en/blog
- https://www.iplexlaw.co.kr/en/contact
- https://www.iplexlaw.co.kr/en/blog/876621
- https://www.iplexlaw.co.kr/en/blog/873032
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals
- https://www.iplexlaw.co.kr/en/blog/claude-computer-use-agent-patent
- https://www.iplexlaw.co.kr/en/blog/claude-computer-use-agent-patent
