ECCV (European Conference on Computer Vision) 2020: An analysis of end-to-end object detection with transformers.
The paper was published by researchers at Facebook AI. It introduced a simpler, competitive object-detection architecturecalled DETR, short for DEtection TRansformer. Its main ingredients are a set-based loss using bipartite matching and a transformer encoder-decoder architecture.
What problem does the paper address?
Many earlier detectors relied on carefully designed components and post-processing steps.
These included anchor-box design and non-maximum suppression to remove overlapping detections. DETR aims to simplify that pipeline.
That is, A simpler architecture was needed to solve the problem that it was quite tricky to use existing object detection methods, and this was solved in this paper.
Overall DETR architecture

1) A backbone network (CNN) is used to extract features for the image, and 2) a positional encoding is used to obtain relative location information in the image.
* * *

CNN features are flattened into a sequence and combined with positional information for the transformer encoder.
* * *

On the other hand, the output of the encoder and the object query (a fixed number of learned location embeddings) are input to the decoder, and the output of the decoder is passed to the feed forward network (FFN) that generates the final output. The final result value is: 1) if a particular object exists, it contains information about the class for that object and the bounding box associated with where that object exists; and 2) if a particular object does not exist, no object is represented as a class, meaning that it does not exist.
* * *
Bipartite matching pairs predictions with ground-truth objects during training. This encourages one prediction per object and reduces duplicate detections.
As a result, it is the whole structure to extract features from the image when the image is entered, and then transform it so that the actual object detection results can be obtained.
Bipartite matching
Many earlier detectors produced numerous candidate boxes and used non-maximum suppression (NMS) to remove duplicates. DETR treats detection as direct prediction of an object set.
DETR instead addresses the set-prediction problem directlythrough a fixed set of predictions and matching-based training..
The model outputs N prediction slots, chosen to exceed the expected number of objects in an image.
That is, One-to-one bipartite matching assigns ground-truth objects to predictions. Unmatched slots are trained as no-object predictions..
For an illustrative N of 6, six slots are predicted. Matching assigns each real object to one slot and labels the others as no object.

Let’s assume that the prediction of two birds performing object detection using the image taken during the training process is as follows:

C = class
b = bounding box information (x coordinates, y coordinates, width, height)
Because N=6, the prediction value is C0 to C5.
Matching considers class predictions and box agreement, selecting the assignment with the lowest total matching cost.
Predictions are paired one-to-one with ground-truth objects. Remaining prediction slots are assigned the no-object target.
Training improves the matched classifications and boxes while teaching unmatched slots to predict no object.
The model thereby learns the number and locations of objects without duplicate detections. The matching is invariant to the order of prediction slots. Swapping two slots does not change the object set, and one-to-one assignment discourages redundant detections..
Why use a transformer?
A transformer uses attention to model relationships among elements in a sequence.
Transformer 1) It is used to understand the contextual information of the entire image through attention, and 2) to facilitate the understanding of the interaction of each instance in the image.
Specifically, the attention mechanism allows each pixel to perform attention by assigning attention scores to each other, in which case the instance is separated, and context is understood for each instance.
Also, Attention can relate distant image features, including parts of a large object.
As with long-range relationships between words, attention can connect spatially distant features without relying only on local convolution operations.
What does the encoder do?

The encoder processes the flattened CNN features together with positional information.Its input has d × HW elements arranged as a sequence of HW feature vectors.
Here, d is the feature dimension and H and W are the feature map's height and width.
In addition, the positional encoding information can be entered to properly encode and process the location information for the actual image data present in 2D.
On the other hand, when the feature information of the image is entered in the encoder, the interaction and association information for each pixel is learned while performing self attention. The encoded information is then used in the decoding part.
After all the encodings have been performed, you can visualize the self-attention map to see that the individual instances are properly separated as follows: In other words, For a selected query location, the attention map shows which other image locations contribute strongly to its representation.

The self-attention map. Four cows are properly separated.
What does the decoder do?

The decoder receives N object queries as an initial input, receives the encoded information from the encoder, and can understand the context of the entire image.Positional information is used alongside encoder features so the decoder can account for spatial relationships.
The decoder allows each object query to output class and bounding box information for a single instance. That is, an architecture is configured to distinguish N different unique instances.
As a result, If the encoder separates an instance through global attention, the decoder extracts the class and boundary of each instance.
The image below shows an attention map of each of the N objects in the decoder part. High attention scores near object boundaries are shown in blue and orange.These features help the prediction heads estimate the bounding boxes.

DETR shows good performance for large objects, but low performance for small objects. The problem with DETR is that it takes quite a long time to learn. To compensate for this, it seems necessary to develop such a technology that can handle small objects well while reducing the training time of the DETR.
This article reflects the information available when it was published. Contact us to discuss your circumstances.

