# FiX: fine-grained forgetting in softmax attention

Yongduck Kim examines FiX’s feature-wise gates, numerical implementation and paged cache, distinguishing reported gains from unresolved limitations.

Source: https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention

HOME / NEWS & INSIGHTS NEWS & INSIGHTS FiX: fine-grained forgetting in softmax attention Yongduck Kim examines FiX’s feature-wise gates, numerical implementation and paged cache, distinguishing reported gains from unresolved limitations. AI & Software 2026.10.01 published IPLEX 12 min read What information should a model retain? Long conversations require memory, but retaining every feature with equal strength is not always useful. FiX adjusts how individual components of a past token contribute to the current output. This is an analysis of the authors’ research, not a report of a patent granted to IPLEX or its clients. The paper is FiX: Introducing Fine-grained Forget Gate into Softmax Attention, by Runzhong Li, Renjie Liu, Qing Li and Bo Tang. It appears in the ICML 2026 proceedings, PMLR volume 306, pages 68619–68640, identifier pmlr-v306-li26dl. The conference ran from 6 to 11 July 2026. The affiliations are Southern University of Science and Technology and The Hong Kong Polytechnic University; the paper notes an AlayaDB internship for Renjie Liu. The official record does not list a DOI or arXiv identifier. The paper and research figures are attributed to their authors under CC BY 4.0. Why a scalar forget gate is restrictive Attention compares a query Q with keys K to determine weights, then combines the corresponding values V into an output O. Selecting where to look and deciding what information to carry forward are distinct operations. FoX introduces context-dependent forgetting, but its scalar gate applies the same decay rate to all feature components of a past token within a head. A vector-valued gate cannot simply be added to the scalar Q–K dot product. FiX therefore changes where forgetting acts: cumulative, element-wise gates modulate value vectors before their contributions are combined. Some features of the same token can persist while others fade more quickly. Figure 1 illustrates gate values such as 0.2, 0.1 and 0.3; these are explanatory values, not accuracy measurements. The current token has an empty cumulative product, equal to one. Illustration 2 — Why a scalar forget gate is restrictive The role of RMSNorm and RoPE The mathematical bridge is output RMSNorm. A common positive scale factor cancels under ideal RMSNorm, allowing scalar forgetting to be moved from the score path to the value–output path before generalising it to feature-wise gates. The equivalence is exact when the stabilising constant ε is zero; FiX uses ε = 10⁻³⁰ in output normalisation. Softmax remains in FiX’s computation, so this is not a claim that every transformer can dispense with softmax. The layer first derives Q, K, V and gates from the input. Most layers use a low-rank gate projection with intermediate dimension 128 to limit extra parameters. RoPE, when included, operates on Q and K. Cumulative gates act on V, the weighted values are combined, and the output passes through RMSNorm and a linear transformation. This separation permits FiX and RoPE to be used together. The first layer uses learned gate embeddings indexed by token ID because a conventional gate projection there produced unstable gradients. The authors also use a Mamba-style activation. Ablations replacing the small ε, first-layer embeddings or this activation increased training loss relative to the complete design. Illustration 3 — The role of RMSNorm and RoPE Making the calculation numerically practical Products of gates between zero and one can become extremely small over long sequences. Dividing V directly by these products can produce excessively large values. Computing every position separately avoids that division but does not fully exploit GPU matrix multiplication. Flash Fine-grained Attention combines rescaling around block reference points with FlashAttention-style tiling. Ratios of cumulative gates are kept at or below one; most earlier blocks use matrix multiplication, while diagonal blocks receive separate treatment. The implementation also fuses attention, RMSNorm and the following linear layer so sensitive calculations retain float32 precision. Adding a gate alone would not reproduce this combination of numerical and computational measures. Paged VF Cache and its memory trade-off Ordinary autoregressive attention retains past K and V. FiX additionally needs gates F. Under the paper’s precision assumptions, storing every F in float32 adds roughly as much memory as the existing KV cache. Absorbing F into V after every token avoids a separate F cache but repeatedly reads and rewrites all past V, with bandwidth costs and accumulating conversion errors. Paged VF Cache groups tokens into pages. Once a page is full, its cumulative gates are absorbed into V and that page’s values are frozen; only a representative gate product is retained. Values and gates remain separate in the incomplete page. Figure 2 uses a page size of three for illustration; the practical setting discussed is 16. The additional gate storage is approximately 1/p of the standard KV cache. At p = 16, that is about 6.25% additional gate-cache storage relative to the KV cache, not a 6.25% reduction in total GPU memory. Page management also avoids updating every historical value at every decoding step. Illustration 4 — Paged VF Cache and its memory trade-off What the experiments actually compare The main models have approximately 760 million parameters and were trained on 48 billion FineWeb-Edu tokens. Ablations use approximately 340 million parameters and 10 billion tokens. Comparators are RoPE, FoX, FoX-RoPE, FiX and FiX-RoPE, within a shared architecture including SwiGLU, QK-Norm, output normalisation and Short-Conv. The low-rank gate projections add approximately 8–10 million parameters, in addition to the first-layer gate embeddings. FiX has lower training loss than RoPE, FoX and FoX-RoPE, while FiX-RoPE improves further on training loss. The lowest training loss does not, however, identify the winner on every downstream benchmark. Accuracy gains are modest and uneven Table 1 reports an arithmetic average across eight accuracy tasks of 54.82% for RoPE, 54.42% for FoX, 54.67% for FoX-RoPE, 55.24% for FiX and 54.85% for FiX-RoPE. FiX exceeds RoPE by 0.42 percentage points. The two perplexity results are not included in either the accuracy average or its geometric mean. FiX leads on PIQA at 72.58%, HellaSwag at 56.32% and BoolQ at 60.86%. FoX-RoPE performs better on ARC-easy and ARC-challenge and has the lowest Wikitext perplexity. FiX-RoPE does better on LAMBADA but not on average accuracy. These results support a measured improvement under the reported conditions, not a step change across all tasks. Illustration 5 — Accuracy gains are modest and uneven Longer contexts do not eliminate reasoning difficulties Models trained at 4,096 tokens were evaluated without additional fine-tuning up to 16,384 tokens on CodeParrot, PG-19 and NarrativeQA. The tested RoPE baseline’s prediction loss rose beyond the training length; FoX and FiX were more stable, with FiX somewhat better than FoX. The RoPE-combined variants were not compared in this experiment. This finding does not establish that all RoPE models or other context-extension techniques perform poorly. BABILong measures something different: answering questions about facts scattered through long documents. Across five tasks, FiX was particularly competitive at 1k–4k tokens, but its accuracy also dropped substantially at 8k and 16k. Stable next-token prediction and successful long-document reasoning should not be treated as interchangeable results. Illustration 6 — Longer contexts do not eliminate reasoning difficulties Illustration 7 — Longer contexts do not eliminate reasoning difficulties Scope, implementation cost and patent-practice perspective FiX changes the attention architecture and requires training. It is not presented as a setting that can simply be enabled in an existing deployed model. Attention still interacts with past tokens; block processing does not make those interactions linear-time or remove the KV cache. Gate computation and precision management also carry costs. The principal evidence is at approximately 760 million parameters. Larger models, other training data and production throughput or latency require separate evaluation. The small differences in the main table are not accompanied by repeated-run error bars or confidence intervals and should not be elevated into a universal ranking. From a patent-analysis perspective, the technical contribution must be examined as a combination: feature-wise cumulative gates in the V–O path, output normalisation with a very small ε, first-layer gate embeddings, block rescaling, fused precision-sensitive operations and paged storage. The gate’s location and granularity, and its relationship with normalisation and memory, matter more than the mere presence of a gate. The reported effects inform a technical analysis; they do not by themselves determine novelty, inventive step or patentability. What to take from FiX FiX offers a concrete way to make retention selective at the feature level while preserving softmax attention. Its mathematical reformulation is supported by an implementation and cache design addressing the problems that reformulation creates. The promising results at a moderate model size should be read alongside the remaining long-context limitations and the need to establish performance and cost at larger scales. Original figure labels and forms are preserved; the surrounding text explains their meaning. Related resources: AI patents Read the Korean source This article reflects the information available when it was published. Contact us to discuss your circumstances. Discuss this topic ↗ All articles ON THIS PAGE What information should a model retain? Why a scalar forget gate is restrictive The role of RMSNorm and RoPE Making the calculation numerically practical Paged VF Cache and its memory trade-off What the experiments actually compare Accuracy gains are modest and uneven Longer contexts do not eliminate reasoning difficulties Scope, implementation cost and patent-practice perspective What to take from FiX TALK TO IPLEX Discuss your IP questions We consider your technology and business needs together. ↗ Contact us Newer Korean patent scope proceedings after testing and disposal of a product ↗ Older Saudi trade marks: filing, deadlines and the Madrid route ↗ Related insights AI & Software 2026.09.30 MHAR: Reading earlier layers through different feature subspaces Yongduck Kim examines Multi-Head Attention Residuals: depth routing, reported training results, implementation costs and the relationship between technical features and effects. ↗ Read article AI & Software 2026.09.28 Column: Claude Computer Use and the Data That Trains AI Agents Writing for AI Times, IPLEX Managing Partner Yongduck Kim examines the training data behind computer-operating AI agents through U.S. Patent No. 12,585,862. ↗ Read article AI & Software 2026.09.18 IP Strategy Begins with Understanding Technology: An IPLEX Interview Explore how IPLEX combines AI and software expertise with a practical understanding of business strategy. The interview discusses education for founders, mentoring and communication with clients. ↗ Read article

- https://www.iplexlaw.co.kr/en
- https://www.iplexlaw.co.kr/en/blog/category/ai
- https://www.iplexlaw.co.kr/en/ai-patents
- https://www.iplexlaw.co.kr/ko/blog/fix-fine-grained-forgetting-attention
- https://www.iplexlaw.co.kr/en/contact
- https://www.iplexlaw.co.kr/en/blog
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention#editorial-fix-en-0
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention#editorial-fix-en-1
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention#editorial-fix-en-2
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention#editorial-fix-en-3
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention#editorial-fix-en-4
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention#editorial-fix-en-5
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention#editorial-fix-en-6
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention#editorial-fix-en-7
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention#editorial-fix-en-8
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention#editorial-fix-en-9
- https://www.iplexlaw.co.kr/en/contact
- https://www.iplexlaw.co.kr/en/blog/brunch-case-1787
- https://www.iplexlaw.co.kr/en/blog/saudi-arabia-trademark-filing-guide
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals
- https://www.iplexlaw.co.kr/en/blog/claude-computer-use-agent-patent
- https://www.iplexlaw.co.kr/en/blog/claude-computer-use-agent-patent
- https://www.iplexlaw.co.kr/en/blog/1539348
- https://www.iplexlaw.co.kr/en/blog/1539348
