# MHAR: Reading earlier layers through different feature subspaces

Yongduck Kim examines Multi-Head Attention Residuals: depth routing, reported training results, implementation costs and the relationship between technical features and effects.

Source: https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals

HOME / NEWS & INSIGHTS NEWS & INSIGHTS MHAR: Reading earlier layers through different feature subspaces Yongduck Kim examines Multi-Head Attention Residuals: depth routing, reported training results, implementation costs and the relationship between technical features and effects. AI & Software 2026.09.30 published IPLEX 15 min read A patent attorney’s reading of the research By Yongduck Kim, Korean patent attorney, IPLEX. Multi-Head Attention Residuals asks how a language model can reuse representations it has already computed. Different feature groups may need information from different earlier layers. The results discussed here were reported by the paper’s authors; they are not experiments conducted by IPLEX. This analysis concerns Cheng Luo, Zefan Cai and Junjie Hu, Multi-Head Attention Residuals, arXiv:2607.27230v2, revised 31 July 2026. The authors list Independent Researcher and the University of Wisconsin–Madison affiliations. The preprint and reproduced research figures are attributed to those authors under CC BY 4.0. From a residual stream to a choice of earlier representations A conventional pre-norm Transformer adds each sublayer’s output to the residual stream. The next sublayer receives the accumulated state, rather than a separately addressable collection of earlier outputs. This does not mean that earlier information disappears; the issue is how it can be selected. Attention Residuals retains the embedding and earlier attention and MLP outputs as individual sources. A learned query scores them, and a softmax over depth supplies a weighted sum. In the single-head version, one depth distribution applies to every feature channel. Feature groups with different preferences must therefore share one selection rule. Single-head and multi-head depth selection — paper Figure 1 What MHAR changes MHAR partitions a query of dimension d and each source representation into H feature subspaces of dimension d/H. Each head scores the available sources and applies its own softmax over depth. The partial weighted sums are concatenated to recover a d-dimensional input. H = 1 recovers single-head Attention Residuals. The distinction is the selection axis: ordinary multi-head self-attention selects among tokens, whereas MHAR selects among earlier layer outputs at the current position. Different learned depth patterns do not establish that particular heads specialise in grammar or mathematics. Nor should the schematic be read as adding a large, separate projection network: the method rearranges the existing query and feature dimensions. Illustrative depth weights, not experimental measurements The processing sequence and its cost Only the embedding and already computed sublayer outputs are available; future layers are not read. RMSNorm supplies normalised keys, while the original source representations supply the values. Queries and keys are split consistently, scoring and softmax are performed independently for each head, and the resulting weighted feature groups are concatenated for the current attention or MLP sublayer. In the principal from-scratch experiments, routing queries start at zero, giving a uniform source distribution before learning. No additional parameters means no increase relative to the single-head routing variant. Against the ordinary Transformer baseline, the paper reports about 0.02% more parameters and 0.5–1.2% more theoretical computation. Reading historical representations still creates memory-traffic costs, motivating a fused Triton implementation. From-scratch training: read the comparisons carefully The main experiments train Qwen3-style 100M, 350M and 1B models for 20,000 steps on the cleaned and deduplicated English anneal_pt_v3 corpus, with substantial synthetic, STEM and code content. Within each size, the data, schedule and batch settings are matched across the baseline, Hyper-Connections, single-head routing and MHAR. Reported losses average 11 evaluations in the final 5,000 steps; that averaging is not itself an average across random seeds. Baseline / single-head / MHAR losses are 3.031 / 2.970 / 2.969 at 100M, 2.997 / 2.876 / 2.848 at 350M, and 2.894 / 2.759 / 2.754 at 1B. At 350M, the improvement against the baseline is 0.149, but the additional improvement over single-head routing is 0.028. The entire gain cannot be attributed to splitting the heads. The paper also contains an Appendix H robustness study with three seeds on FineWeb-Edu under a separate protocol. That additional evidence should be distinguished from the main table’s evaluation average; it does not establish reproducibility in every setting. Validation losses for from-scratch training — paper Table 1 Downstream results are not uniform At 1B, WikiText-2 perplexity falls from 56.4 to 45.4 and LAMBADA accuracy rises from 8.8% to 16.0%, comparing the baseline with MHAR. At 100M, however, HellaSwag falls from 35.0% to 33.0%. The paper notes approximately ±3 percentage points of sampling noise for evaluations with 200 examples. Context lengths also differ across model sizes, so comparisons should first be made within each size. Adapting an already trained 8B model For Marin-8B continued pretraining, the authors retain the original residual stream and add a delta-routing branch. An output gate initialised at zero preserves the original computation at conversion; the reported initial maximum logit difference in fp32 is zero. The new path contributes as its gate learns to open. This variant uses eight block deltas obtained by grouping 32 layers into blocks of four, plus a learned null source: at most nine sources are read by eight heads. It is therefore different from retaining every historical sublayer output. The source data pool is about 1.9 trillion tokens, but the actual continued-training run uses about 10 billion tokens. Plain continued pretraining (Plain CPT) provides a matched schedule, data order and batch comparison. Separate continued training from MHAR’s additional contribution GSM8K accuracy is 19.0% for the original model, 47.0% for Plain CPT and 50.2% for MHAR. MHAR’s additional gain is therefore 3.2 percentage points. GPQA rises from 31.5% with Plain CPT to 34.6% with MHAR, a 3.1-point difference. Much of the improvement from the original checkpoint comes from continued training itself. MATH remains at 19.1% for both methods. Differences on MMLU, HumanEval and MBPP are not described as statistically clear. Paired tests give p = 0.004 for GSM8K and p = 0.038 for GPQA; those tests are not a substitute for repeated-seed evidence across all deployment conditions. 8B continued-pretraining accuracy; compare MHAR with Plain CPT More heads are not necessarily better In the 1B web-corpus experiment with eight KV heads, losses for 1, 4, 8 and 16 routing heads are respectively 3.270, 3.132, 3.129 and 3.173. Four and eight heads perform similarly well, while sixteen performs worse. Too few groups force different preferences together; excessive splitting may separate features that benefit from joint treatment. Eight heads is not a universal prescription. Model width, training stage and data can alter the appropriate partition. The web-corpus experiments also use different initialisation conditions from the main experiments, so their differences cannot be assigned to data distribution alone. Routing-head comparison under a separate experimental protocol Memory traffic matters beyond operation counts Relative to a same-size baseline throughput of 1.00, the ordinary MHAR implementation achieves 0.54 / 0.32 / 0.23 at 100M / 350M / 1B; the fused kernel achieves 0.88 / 0.71 / 0.55. Fusion reduces repeated access and intermediate storage, but the end-to-end training remains slower than the baseline. At 1B, peak memory falls from 47.5 GB to 20.1 GB with fusion, close to the baseline’s 19.0 GB, while throughput remains 55% of baseline. A speed-up of one routing kernel is not the same as a speed-up of the whole model. Equal training steps also do not demonstrate superiority under an equal wall-clock budget. Training throughput and peak memory — table excerpt from the paper Where the evidence ends MHAR offers both a design for new models and an additional internal path for continued training of existing models. The principal evidence nevertheless concerns English data and particular architectures and sizes. It does not by itself establish benefits for Korean-language models, multimodal systems, long-context services or production inference latency and cost. Useful further checks include repeated seeds under the relevant protocol, equal-time or equal-total-compute comparisons, experiments separating data and initialisation effects, and inference memory and latency measurements. Better training metrics should not be presented as proof of faster user-facing responses. Technical features and effects: the patent practitioner’s perspective The technical problem is more specific than improving language-model accuracy: a shared depth distribution cannot fully express different subspaces’ information requirements, while repeated access creates a memory burden. The central relationship is consistent feature partitioning, independent normalisation over sources, weighted aggregation and recombination. Fused execution addresses intermediate tensors and memory access. Block-delta sources and zero-initialised gating address source count and preservation of a pretrained model’s initial behaviour. These features solve different implementation problems and should not be collapsed into one claim of improved performance. Compared with ordinary residual connections, historical outputs become individually selectable; compared with single-head Attention Residuals, selection becomes subspace-specific; compared with token attention, the sources lie along depth. Technical effects should be tied to these relationships and to the actual comparison conditions. This is an analysis of the paper’s technical contribution, not a conclusion that an invention is patentable or that patent rights have been granted. Closing observation MHAR redesigns who reads previously computed information, from which sources and in what proportions. Its contribution is best understood by considering subspace-specific selection, preservation of pretrained computation and real memory-traffic costs together. A promising information path does not automatically guarantee a faster implementation. Original figure labels and forms are preserved; the surrounding text explains their meaning. Related resources: AI patents Read the Korean source This article reflects the information available when it was published. Contact us to discuss your circumstances. Discuss this topic ↗ All articles ON THIS PAGE A patent attorney’s reading of the research From a residual stream to a choice of earlier representations What MHAR changes The processing sequence and its cost From-scratch training: read the comparisons carefully Downstream results are not uniform Adapting an already trained 8B model Separate continued training from MHAR’s additional contribution More heads are not necessarily better Memory traffic matters beyond operation counts Where the evidence ends Technical features and effects: the patent practitioner’s perspective Closing observation TALK TO IPLEX Discuss your IP questions We consider your technology and business needs together. ↗ Contact us Newer When a dosage clarification changes a patent claim: Korea’s Supreme Court on correction ↗ Older Trade mark protection in the UAE: filing routes, timing and practical requirements ↗ Related insights AI & Software 2026.10.01 FiX: fine-grained forgetting in softmax attention Yongduck Kim examines FiX’s feature-wise gates, numerical implementation and paged cache, distinguishing reported gains from unresolved limitations. ↗ Read article AI & Software 2026.09.28 Column: Claude Computer Use and the Data That Trains AI Agents Writing for AI Times, IPLEX Managing Partner Yongduck Kim examines the training data behind computer-operating AI agents through U.S. Patent No. 12,585,862. ↗ Read article AI & Software 2026.09.18 IP Strategy Begins with Understanding Technology: An IPLEX Interview Explore how IPLEX combines AI and software expertise with a practical understanding of business strategy. The interview discusses education for founders, mentoring and communication with clients. ↗ Read article

- https://www.iplexlaw.co.kr/en
- https://www.iplexlaw.co.kr/en/blog/category/ai
- https://www.iplexlaw.co.kr/en/ai-patents
- https://www.iplexlaw.co.kr/ko/blog/mhar-multi-head-attention-residuals
- https://www.iplexlaw.co.kr/en/contact
- https://www.iplexlaw.co.kr/en/blog
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals#editorial-mhar-en-0
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals#editorial-mhar-en-1
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals#editorial-mhar-en-2
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals#editorial-mhar-en-3
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals#editorial-mhar-en-4
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals#editorial-mhar-en-5
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals#editorial-mhar-en-6
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals#editorial-mhar-en-7
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals#editorial-mhar-en-8
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals#editorial-mhar-en-9
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals#editorial-mhar-en-10
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals#editorial-mhar-en-11
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals#editorial-mhar-en-12
- https://www.iplexlaw.co.kr/en/contact
- https://www.iplexlaw.co.kr/en/blog/brunch-case-1781
- https://www.iplexlaw.co.kr/en/blog/uae-trademark-filing-guide
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention
- https://www.iplexlaw.co.kr/en/blog/claude-computer-use-agent-patent
- https://www.iplexlaw.co.kr/en/blog/claude-computer-use-agent-patent
- https://www.iplexlaw.co.kr/en/blog/1539348
- https://www.iplexlaw.co.kr/en/blog/1539348
