# Attention Residuals: How AI combines earlier layers, and what matters for patents

An accessible explanation of Attention Residuals, Block AttnRes, the limits of the reported results and the implementation details relevant to AI patent drafting.

Source: https://www.iplexlaw.co.kr/en/blog/attention-residuals-ai-patent-analysis

HOME / NEWS & INSIGHTS NEWS & INSIGHTS Attention Residuals: How AI combines earlier layers, and what matters for patents An accessible explanation of Attention Residuals, Block AttnRes, the limits of the reported results and the implementation details relevant to AI patent drafting. AI & Software 2026.10.03 published IPLEX 4 min read Original Korean-language cover illustrating the combination of information from several stages Attention Residuals changes how information passes between the layers of an AI model. Instead of simply adding earlier outputs with equal weight, it learns how much of each representation to use for a given input. A simple analogy helps explain the mechanism—and the concrete implementation choices that deserve attention when preparing an AI patent application. Paper reviewed: Attention Residuals by the Kimi Team, arXiv v1 released on 16 March 2026. Why reconsider residual connections? A model processes its input through successive layers. Each layer is a processing stage; a residual connection adds a new result to information carried forward from earlier stages and passes the combination onward. With conventional residual accumulation, earlier outputs enter the sum with equal coefficients. The paper examines how an individual layer’s contribution can become diluted as a model grows deeper. Attention Residuals compares earlier representations and combines them with weights that depend on the input. Attention is a mechanism for assigning weights to information. Here, it compares representations of the same token across depth. A token is a small unit into which text is divided for processing; the selection discussed here takes place between layers. An analogy: notes on a report Imagine three colleagues leaving notes on a report. The editor may give more weight to the note that best assists with the sentence currently being revised. Similarly, the model combines earlier processing results in proportions suited to its input. Illustrative example: combine note A at 10%, note B at 20% and note C at 70%, then use the combined result in the next stage. These percentages are invented for explanation; they are not experimental results from the paper. In the actual model, processing results are numerical representations. A learned comparison rule determines the weights; a person does not manually rank the importance of the notes. The mechanism in three steps First, collect earlier outputs and normalise the representations used to calculate their weights. Normalisation puts the values used for comparison on a consistent scale. Second, a comparison criterion learned for each layer produces scores. Softmax converts the scores into weights summing to 100%. Third, the original representations—not merely the normalised comparison values—are combined using those weights to form the next processing input. Keeping every earlier output separately can increase memory and communication costs. Block AttnRes groups several layers: outputs are added within a block, while attention determines the contributions of block-level representations. This reduces the burden of retaining and communicating every layer’s output individually. What the experiments do—and do not—establish The authors report improvements on general knowledge, mathematics, code and other evaluations after training models based on Kimi Linear. These findings depend on the architecture and training and evaluation conditions described in the paper. They do not establish that changing the connections in an existing model, on its own, will reproduce the same gains. Consider a possible application in an enterprise report-summarisation system. An implementation assessment would need to measure summary quality alongside memory for earlier outputs, inter-device communication and processing time. This is an example of potential use, not evidence that a particular commercial product implements the method. From a technical idea to a patent-drafting review This section offers an independent patent-practice perspective on a published paper. It does not construe the scope of a particular granted patent or determine patentability. Claims define the protection sought: the objective of selecting information therefore needs to be examined through the specific processing arrangements that implement it. Selection targets: identify whether the inputs are individual layer outputs or representations of groups of layers. Weight computation: establish how the learned comparison criterion, normalisation and weighting operations are combined, including their order. Storage and communication: describe the arrangements and processing sequence for retaining earlier results and reducing redundant transfers. Technical effects: document the impact on performance, memory and communication together with the conditions used for comparison. For an AI patent review, prepare a processing flowchart, the features that differ from existing approaches, comparative test data and a schedule of planned disclosures. IPLEX draws on its publicly documented AI patent-analysis and IP-R&D work and publications to examine technical distinctions when developing a protection strategy. Yongduck Kim’s professional profile provides further details of that experience. Analysis as of 3 October 2026, based on arXiv v1. We have not independently reproduced the experiments. AI patents and IP strategy Yongduck Kim’s professional profile Read our In-Place Test-Time Training analysis Read the original on Tistory Read the Korean source This article reflects the information available when it was published. Contact us to discuss your circumstances. Discuss this topic ↗ All articles ON THIS PAGE Why reconsider residual connections? An analogy: notes on a report The mechanism in three steps What the experiments do—and do not—establish From a technical idea to a patent-drafting review TALK TO IPLEX Discuss your IP questions We consider your technology and business needs together. ↗ Contact us Older Korean patent scope proceedings after testing and disposal of a product ↗ Related insights AI & Software 2026.10.01 FiX: fine-grained forgetting in softmax attention Yongduck Kim examines FiX’s feature-wise gates, numerical implementation and paged cache, distinguishing reported gains from unresolved limitations. ↗ Read article AI & Software 2026.09.30 MHAR: Reading earlier layers through different feature subspaces Yongduck Kim examines Multi-Head Attention Residuals: depth routing, reported training results, implementation costs and the relationship between technical features and effects. ↗ Read article AI & Software 2026.09.28 Column: Claude Computer Use and the Data That Trains AI Agents Writing for AI Times, IPLEX Managing Partner Yongduck Kim examines the training data behind computer-operating AI agents through U.S. Patent No. 12,585,862. ↗ Read article

- https://www.iplexlaw.co.kr/en
- https://www.iplexlaw.co.kr/en/blog/category/ai
- https://www.iplexlaw.co.kr/en/ai-patents
- https://www.iplexlaw.co.kr/en/professionals/yongdeok-kim
- https://www.iplexlaw.co.kr/en/blog/1525199
- https://iplexlaw.tistory.com/3
- https://www.iplexlaw.co.kr/ko/blog/attention-residuals-ai-patent-analysis
- https://www.iplexlaw.co.kr/en/contact
- https://www.iplexlaw.co.kr/en/blog
- https://www.iplexlaw.co.kr/en/blog/attention-residuals-ai-patent-analysis#attnres-0
- https://www.iplexlaw.co.kr/en/blog/attention-residuals-ai-patent-analysis#attnres-1
- https://www.iplexlaw.co.kr/en/blog/attention-residuals-ai-patent-analysis#attnres-2
- https://www.iplexlaw.co.kr/en/blog/attention-residuals-ai-patent-analysis#attnres-3
- https://www.iplexlaw.co.kr/en/blog/attention-residuals-ai-patent-analysis#attnres-4
- https://www.iplexlaw.co.kr/en/contact
- https://www.iplexlaw.co.kr/en/blog/brunch-case-1787
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals
- https://www.iplexlaw.co.kr/en/blog/claude-computer-use-agent-patent
- https://www.iplexlaw.co.kr/en/blog/claude-computer-use-agent-patent
