# Research Paper Analysis: ResNet — Deep Residual Learning for Image Recognition

This article examines the paper “Deep Residual Learning for Image Recognition.”

Source: https://www.iplexlaw.co.kr/en/blog/876621

HOME / NEWS & INSIGHTS NEWS & INSIGHTS Research Paper Analysis: ResNet — Deep Residual Learning for Image Recognition This article examines the paper “Deep Residual Learning for Image Recognition.” AI & Software 2023.07.03 published IPLEX 5 min read This article examines the paper “Deep Residual Learning for Image Recognition.” ( Skip to content ) Presented at CVPR 2016, the ResNet paper attracted attention after residual networks achieved leading results in the 2015 ImageNet competition. The main feature of this paper is that it uses residual learning to deepen the network. Greater depth can improve representational capacity, but simply adding layers can increase training error. The paper addresses this degradation problem through residual learning. [Background] In a Convolutional Neural Network (CNN), each of the different filters performs a convolutional operation for a given layer by input, and extracts the activation map from the output. Here, each of the multiple filters is trained to extract different characteristic values. For an input with three channels, five convolution filters produce five output feature maps. Their spatial dimensions depend on kernel size, stride and padding. Networks often increase channel count while reducing spatial resolution in later stages, but this is an architectural choice rather than an automatic effect of depth. VGG demonstrated the effectiveness of deep networks built from small 3 × 3 convolution filters before ResNet. VGG Network VGG is widely used as a feature-extraction backbone, but deeper plain networks can become difficult to optimize and may have many parameters. ResNet addresses the optimization difficulty through residual connections. [The main idea of the article] In general, the deeper the layer, the higher the learning difficulty of the model, so it is not easy to optimize as intended. The key idea is to replace plain stacks of layers with residual blocks that are easier to optimize. A residual block learns a correction to its input. If the desired mapping is H(x), the block learns the residual F(x) = H(x) − x. The paper argues that this can be easier than directly learning H(x). * * * * * More specifically, when x is entered into a weight layer such as a convolution layer, the feature is extracted, and the feature undergoes an activation function such as a relu. When the network is configured, the entire network performs a non-linear operation. The network is then configured in the form of a continuous convolution layer. Simply adding more weight layers can make optimization harder and increase training error. This motivates the residual formulation. * * * * * The residual branch computes F(x), while a shortcut carries x to the output. Their sum produces H(x) = F(x) + x. In this formulation, the shortcut preserves the input x and the residual branch learns the additional transformation F(x). Rather than learning the full mapping from scratch, the block learns the residual and adds it to x. The result is H(x). This can make deeper networks easier to optimize. By changing the network in this way, learning can be faster and higher performance. If we look at the mathematical formula in more detail, we can first define the F(x) function as follows: For a two-layer residual branch, F(x) can be written as W2·ReLU(W1x). Adding the shortcut gives the output F(x) + x. Residual branches can contain more than two layers. If input and output dimensions differ, a learned projection Ws can align them before addition. Read the Korean source This article reflects the information available when it was published. Contact us to discuss your circumstances. Discuss this topic ↗ All articles TALK TO IPLEX Discuss your IP questions We consider your technology and business needs together. ↗ Contact us Newer Research Paper Analysis: StarGAN ↗ Older Improving Deep Learning Performance: Mixup Training and Label Smoothing ↗ Related insights AI & Software 2026.10.01 FiX: fine-grained forgetting in softmax attention Yongduck Kim examines FiX’s feature-wise gates, numerical implementation and paged cache, distinguishing reported gains from unresolved limitations. ↗ Read article AI & Software 2026.09.30 MHAR: Reading earlier layers through different feature subspaces Yongduck Kim examines Multi-Head Attention Residuals: depth routing, reported training results, implementation costs and the relationship between technical features and effects. ↗ Read article AI & Software 2026.09.28 Column: Claude Computer Use and the Data That Trains AI Agents Writing for AI Times, IPLEX Managing Partner Yongduck Kim examines the training data behind computer-operating AI agents through U.S. Patent No. 12,585,862. ↗ Read article

- https://www.iplexlaw.co.kr/en
- https://www.iplexlaw.co.kr/en/blog/category/ai
- https://arxiv.org/pdf/1512.03385.pdf
- https://www.iplexlaw.co.kr/forum/view/876621
- https://www.iplexlaw.co.kr/en/contact
- https://www.iplexlaw.co.kr/en/blog
- https://www.iplexlaw.co.kr/en/contact
- https://www.iplexlaw.co.kr/en/blog/878875
- https://www.iplexlaw.co.kr/en/blog/875313
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals
- https://www.iplexlaw.co.kr/en/blog/claude-computer-use-agent-patent
- https://www.iplexlaw.co.kr/en/blog/claude-computer-use-agent-patent
