NEWS & INSIGHTS

Research Paper Analysis: ResNet — Deep Residual Learning for Image Recognition

This article examines the paper “Deep Residual Learning for Image Recognition.”

This article examines the paper “Deep Residual Learning for Image Recognition.”
Presented at CVPR 2016, the ResNet paper attracted attention after residual networks achieved leading results in the 2015 ImageNet competition. 
The main feature of this paper is that it uses residual learning to deepen the network. 
Greater depth can improve representational capacity, but simply adding layers can increase training error. The paper addresses this degradation problem through residual learning. 

[Background]
In a Convolutional Neural Network (CNN), each of the different filters performs a convolutional operation for a given layer by input, and extracts the activation map from the output. Here, each of the multiple filters is trained to extract different characteristic values.
Illustration: Research Paper Analysis: ResNet — Deep Residual Learning for Image Recognition
For an input with three channels, five convolution filters produce five output feature maps. Their spatial dimensions depend on kernel size, stride and padding. Networks often increase channel count while reducing spatial resolution in later stages, but this is an architectural choice rather than an automatic effect of depth. 
Illustration: Research Paper Analysis: ResNet — Deep Residual Learning for Image Recognition
VGG demonstrated the effectiveness of deep networks built from small 3 × 3 convolution filters before ResNet. 
Illustration: Research Paper Analysis: ResNet — Deep Residual Learning for Image Recognition
VGG Network
VGG is widely used as a feature-extraction backbone, but deeper plain networks can become difficult to optimize and may have many parameters. ResNet addresses the optimization difficulty through residual connections. 

[The main idea of the article]
In general, the deeper the layer, the higher the learning difficulty of the model, so it is not easy to optimize as intended.
The key idea is to replace plain stacks of layers with residual blocks that are easier to optimize. A residual block learns a correction to its input.
If the desired mapping is H(x), the block learns the residual F(x) = H(x) − x. The paper argues that this can be easier than directly learning H(x). 
* * * * *
More specifically, when x is entered into a weight layer such as a convolution layer, the feature is extracted, and the feature undergoes an activation function such as a relu. When the network is configured, the entire network performs a non-linear operation. The network is then configured in the form of a continuous convolution layer.
Simply adding more weight layers can make optimization harder and increase training error. This motivates the residual formulation. 
Illustration: Research Paper Analysis: ResNet — Deep Residual Learning for Image Recognition
* * * * *
The residual branch computes F(x), while a shortcut carries x to the output. Their sum produces H(x) = F(x) + x. 
Illustration: Research Paper Analysis: ResNet — Deep Residual Learning for Image Recognition
In this formulation, the shortcut preserves the input x and the residual branch learns the additional transformation F(x).Rather than learning the full mapping from scratch,the block learns the residual and adds it to x.The result is H(x).This can make deeper networks easier to optimize. By changing the network in this way, learning can be faster and higher performance. 
If we look at the mathematical formula in more detail, we can first define the F(x) function as follows:
Illustration: Research Paper Analysis: ResNet — Deep Residual Learning for Image Recognition
For a two-layer residual branch, F(x) can be written as W2·ReLU(W1x). Adding the shortcut gives the output F(x) + x. 
Illustration: Research Paper Analysis: ResNet — Deep Residual Learning for Image Recognition
Residual branches can contain more than two layers. If input and output dimensions differ, a learned projection Ws can align them before addition.

Read the Korean source

This article reflects the information available when it was published. Contact us to discuss your circumstances.
Discuss this topic ↗All articles

Put your IP strategy into practice.