Without the formula, the two projection paths of the gated MLP are combined to create an intermediate activation value Z, and the current Wdown is applied to Z to obtain the output O. Afterwards, the update amount ΔWdown of Wdown is calculated using the target expressions V and Z. The simplified form presented in the paper is ΔWdown = η · Vᵀ · Z.
Figure 1. In-Place TTT overall structure. Attention is maintained and only the Wdown of the MLP is updated with fast weights. Source: Feng et al., paper p. 6, Figure 1.
The practical advantage of this design is that structural changes are small. Instead of inserting a new large memory module with random initialization, we take the output matrix of an MLP that already has language capabilities as a starting point. The paper calls this drop-in enhancement.
One thing to note here is that the expression drop-in does not mean that no additional learning is required. In the Qwen3-4B experiment, continual training was performed on approximately 20B tokens in 32k contexts and approximately 15B tokens in 128k contexts. In other words, the key is that you can start from an existing checkpoint without completely replacing the model structure or pre-training from scratch.
3. Core principle 2: Apply chunks first and update later
In-Place TTT divides the input sequence into multiple chunks. In each chunk, the current fast weights are first applied to create an output, and then Wdown is updated with the intermediate activation and target values obtained from that chunk. The updated matrix is used starting from the next chunk. This sequence is apply-then-update. Paper pp. 4-5.
The order of operation is as follows: The first chunk is processed as a pre-trained initial Wdown, and the next state Wdown is created using Z and target V obtained from that chunk. The second chunk is processed with this updated Wdown and then the update to be used for the third chunk is calculated again. The important thing is that the output of the current chunk is always produced from its pre-update state. Keeping this order avoids leaks where the current chunk pre-reflects its correct answer information.
The existing TTT layer also played the role of a token mixer replacing attention, so small chunks were needed. On the other hand, in In-Place TTT, attention is responsible for detailed token interactions within a chunk, and MLP fast weights are responsible for the adaptation state beyond the chunk. The paper explains that thanks to this division of roles, performance was good even for large chunks such as 512 or 1024.
4. Core Principle 3: Predict next token, not restore current token
Existing TTTs usually use the goal of restoring the representation of the current input token. However, what a language model should actually do well is not reproduce the current token itself, but predict the next token based on the context so far. The paper sees this discrepancy as the core problem.
The target representation is obtained by applying a one-dimensional convolution and a learnable projection Wtarget to token embeddings X0: V̂ = Conv1D(X0) · Wtarget. Adjusting the convolution kernel allows the target to incorporate the next token or a combination of nearby future tokens. The update simplifies to ΔWdown = η · V̂ᵀ · Z. See page 5 of the paper.
For example, in the pattern of A followed by B, A followed by B may appear earlier in the context, and then A may appear again later in the context. Reconstruction targets tend to work by looking at A and storing a representation of A itself, while LM-aligned targets look at A and store a representation of the B that follows it. Therefore, fast weights can work in the direction of increasing the logit of B when A reappears.
5. Theoretical analysis: Why does the logit of the correct answer token rise?
Theorem 1 of the paper analyzes a simplified situation in the form of an induction head. The key token and the next value token appear once in the front, and when the same key appears again in the back, you have to guess the next value token. Under the assumption that token embeddings are nearly orthogonal to each other and that intermediate activations of the same key are aligned, LM-aligned targets significantly increase the expected logit of the correct token and barely change the logit of other tokens. On the other hand, the reconstruction target is not guaranteed to increase the correct logit. Paper pp. 5-6, Appendix A.
These theoretical results should be read limitedly. Theorem 1 is not a guarantee for all LLMs and all inputs, but is an expected value analysis based on assumptions such as approximate orthogonality of embeddings and key-query alignment. Therefore, although it is meaningful as a theoretical basis for explaining experimental results, it should not be interpreted as a general performance guarantee.
6. Context Parallelism: How to calculate sequential updates in parallel
Although chunk-wise updates are conceptually sequential, the update deltas in the paper are additive and associative. The researchers use this property to calculate the ΔW of each chunk in parallel, obtain the cumulative update up to just before each chunk using exclusive prefix sum, and then calculate the effective Wdown and output required for each chunk in parallel. Paper pp. 6-7, Algorithm 1.
The implementation computes each chunk’s intermediate activations and update delta, then uses an exclusive prefix scan to accumulate updates from preceding chunks only. Adding that sum to the initial Wdown gives the fast weights used to compute the current chunk’s output.
Chunk boundaries manage padding and boundaries to ensure that the goal-generating convolution does not pull future information from other chunks. When an independent document is finished, fast weights are reset to the pre-training state to prevent information leakage between documents.
In the long-term evaluation of Qwen3-4B, numerical stability was ensured by limiting the Frobenius norm of the update delta to a threshold of 1e-5. Appendix C.2.
In our implementation, instead of applying TTT to every MLP layer, we apply it to every sixth MLP block to control state size and computational cost. The target expression was constructed to contain the prediction signal of near future tokens using depthwise Conv1D and Wtarget of kernel size 5. Conv1D is initialized to 0 and Wtarget is initialized to a sparse diagonal shape so that the initial updates start at approximately 0. Additionally, at document boundaries, Wdown was returned to the pre-training state to prevent context leakage between independent sequences, and update norm clipping was used to prevent cumulative updates from flooding in long sequences.
7. Experiment results: How much has LLM long-context performance improved?
7.1 Drop-in continual training for Qwen3-4B
The researchers trained Qwen3-4B-Base and a model to which In-Place TTT was applied using the same continual training curriculum. We used about 20B tokens at 32k length, about 15B tokens at 128k length, and extended RoPE with YaRN at 128k steps. The evaluation metric is the RULER average accuracy, which measures the use of long text context. Paper pp. 7-8.
Looking at the RULER results of Qwen3-4B by context length, in 4k, the In-Place TTT was 96.1, 0.5 points lower than the baseline 96.6. However, in 8k it improved by 1.5 points from 94.1 to 95.6, in 16k it improved by 0.6 points from 92.1 to 92.7, and in 32k it improved by 0.6 points from 88.7 to 89.3.
The improvement was greater in longer contexts. The biggest improvement was seen at 64k with a 4.4 point increase from 74.3 to 78.7, at 128k it was 2.2 points higher from 74.8 to 77.0, and at 256k beyond the training length it was 2.2 points higher from 41.7 to 43.9. Therefore, rather than saying that it is always superior in short contexts, it is more accurate to interpret it as a tendency for the advantage to increase as the context becomes longer.
The biggest improvement is +4.4 points at 64k. It maintained +2.2 points at 128k and 256k beyond the training length, respectively. However, in 4K, it was 0.5 points lower. Therefore, rather than summarizing this paper as a method that is always superior even in short contexts, it is more accurate to read it as a trend in which the advantage increases as the context becomes longer.
7.2 Improvements in LLaMA-3.1-8B and Qwen3-14B
When the same idea was applied to other model families and sizes, 64k performance was improved. LLaMA-3.1-8B rose from 81.6 to 83.7, and Qwen3-14B rose from 67.9 to 70.6. Even in the 64k setting where YaRN was applied to Qwen3-14B, it improved from 81.3 to 82.5, showing that it can be used in parallel with the positional embedding expansion technique. Paper p. 8, Table 2.
Improvements were also confirmed in other base models. The RULER 64k score for LLaMA-3.1-8B increased by 2.1 points from 81.6 to 83.7, and Qwen3-14B increased by 2.7 points from 67.9 to 70.6. Even in the 64k setting applying YaRN to Qwen3-14B, there was an improvement of 1.2 points from 81.3 to 82.5.
7.3 Comparison when learning from scratch
At 500M and 1.5B scale, we compared it with Sliding Window Attention, Gated Linear Attention, DeltaNet, and LaCT. The sliding window perplexity in Figure 2 is an indicator of how much the prediction difficulty of the last block decreases when a long preceding context is provided, and the lower the better. In-Place TTT had the lowest perplexity for both model scales from 2k to 32k. Paper p. 9.
In the 4B model, In-Place TTT is added to Full Attention and Sliding Window Attention. Changes on the commonsense reasoning task were generally small, but significant improvements were seen on the long RULER.
An experiment learning from scratch at 4B scale also showed improvements in long RULER performance. In Full Attention, 4k increased by 4.21 points from 45.77 to 49.98, 8k increased by 5.73 points from 38.09 to 43.82, and 16k increased by 13.41 points from 6.58 to 19.99.
In Sliding Window Attention, there were some areas where the improvement was greater. RULER 4k improved by 13.56 points from 14.77 to 28.33, 8k improved by 16.89 points from 9.91 to 26.80, and 16k improved by 2.50 points from 5.07 to 7.57. The key point of this experiment is that although the changes in the commonsense reasoning task were generally small, there was a clear improvement in RULER, which requires long text context.
7.4 Ablation: What Really Mattered
Figure 3. Results of performance comparison while removing state size, chunk size, and components of LM-aligned objective. Source: Paper p. 10, Figure 3.
Ablation results show that the larger the state size of fast weights, the higher the RULER performance, supporting the advantage of the design of reusing a large MLP matrix as a dynamic state. The chunk sizes of 512 and 1024 were the most competitive, and the paper evaluates 1024 as an advantageous choice when considering efficiency. Additionally, the best performance was achieved when using Conv1D, which creates future token information, and Wtarget projection, which converts the representation, together. Conv1D played a particularly important role in long contexts, and projection played a particularly important role in short contexts.
7.5 Efficiency: Is the overhead small compared to the performance improvement?
Figure 4. Comparison of prefill throughput and peak memory in 8k·32k·128k context in H800 environment. The paper evaluates that the practical overhead is small. Source: Paper p. 10, Figure 4.
Looking at the graph, when applying In-Place TTT, throughput decreases slightly and memory increases slightly, but the difference is not large compared to the cost of attention itself, especially in long contexts. However, this conclusion is a prefill measurement under the specific conditions of Nvidia H800, batch size 1, and paper implementation. The results are not generalized to the cost of multi-user serving, decoding steps, and various hardware.
8. What the paper actually proves and what it has not yet proven
It is necessary to distinguish between the scope that the paper has shown and the scope that has not yet been shown. Although we saw improvements in RULER and sliding window perplexity in the long-context, we cannot guarantee that performance will automatically improve in all long-term reasoning or agent tasks. Model compatibility has also been confirmed through continual training at the 4B to 14B scale of Qwen3 and LLaMA, which does not mean that it can be immediately applied to arbitrary checkpoints without additional learning.
Additionally, fast weights keep changing within a document or sequence, but are reset when the document ends, so they do not maintain persistent knowledge across sessions. In terms of efficiency, throughput and memory overhead were small in the H800's prefill environment, but the total cost of the decode stage, large-scale deployment, and multi-tenant environment was not evaluated. Although the numerical stability of long-text updates is secured through norm clipping, there is no verification against malicious prompts, data corruption, or fast-weight poisoning. The theoretical analysis is also at the level of explaining the effect of LM-aligned target on logit in an induction setting, and is not a universal theorem that covers all situations in reality.
The most important critical point is that although the paper aims for continuous learning, the actual implementation initializes fast weights at each document boundary. Therefore, it is more accurate to view In-Place TTT at its current stage as an adaptive memory or context compression device that operates within the session rather than a persistent knowledge update.
9. Potential inventive features: an AI patent practitioner’s perspective
This section offers a patent-practice interpretation of the paper, separate from the authors’ experimental claims. Assessing novelty and inventive step requires a prior art search covering both patent and non-patent literature.
9.1 Combined relationships are more important than single elements
Fast weights, Test-Time Training, chunk-wise update, Transformer FFN as a memory perspective, future token prediction, and prefix scan each have prior research axes. Therefore, focusing patentability on the phrase simply updating weights during inference will likely result in broad but weak claims.
The relatively strong invention concept in this paper is the combination of the following three layers:
In this paper, the relatively strong invention concept can be found in the combination of the three layers. In terms of architecture, the last output projection from the feed-forward block of the pre-training Transformer was selected as the fast weights to enable a warm start without replacing the structure. In terms of the goal function, instead of restoring the current token, we created a target representation containing information about one or more subsequent tokens to align the update direction with the language model's next token prediction. In terms of execution, the pre-update matrix is applied to the current chunk, and the delta calculated in the current chunk is reflected only in subsequent chunks, parallelizing it with an exclusive prefix scan.
9.2 Claim elements and technical effect
When constructing a patent claim, it is important to link the combination of these elements to the technical effect rather than the individual elements. First, by specifying the output projection of the feed-forward block as fast weights, the effect of implementing dynamic adaptation while maintaining the pre-training structure can be emphasized. The Apply-then-update method creates the next state with the delta calculated after output of the current chunk and applies it to subsequent chunks, thus maintaining causality and preventing information leakage.
LM-aligned targets can be organized into constructs that generate a target representation containing subsequent token information from the embeddings of the current chunk, thereby storing useful context for next token prediction. The chunk method, which updates the matrix by multiple tokens rather than by token, is linked to GPU parallelism and improved throughput, and the prefix scan, which calculates the delta of each chunk in parallel and then accumulates it up to the previous chunk, can be seen as a configuration that simultaneously maintains sequential meaning and context parallelism. Combining this with document boundary reset, norm clipping, and near-zero initialization, we can further demonstrate the effects of preventing cross-document leakage, numerical stability, and preserving initial behavior.
9.3 Application strategy and rights exercise perspective
In the application strategy, separately from the independent claim of the inference method, the continual training method itself, which makes fast weights possible to learn, can be reviewed as an independent claim. Context parallel implementation has a separate technical effect of hardware efficiency, so there is room for separation into parallel processing methods or system claims. Bringing together categories of devices, server systems, and computer-readable recording media can complement enforcement entities and enforceability. However, since it is difficult to prove infringement from the outside in a process where the weight changes only inside the server, it is necessary to consider components that are connected to externally identifiable clues such as model files, runtime logs, update status API, open source code, and performance/memory traces. In addition, paper figures such as long-term performance improvement and small overhead can be used as data to explain the technical effect of a software invention, but it is important to leave sufficient reproducible embodiments and parameter ranges in the actual specification.
10. Industrial implications: where it can be useful
Areas where In-Place TTT is particularly attractive are those where the context is long, patterns are repeated within the same document or session, and re-reading all tokens with attention is a costly operation. Examples include analyzing long contracts and patent specifications, understanding large code repositories, logging long agent operations, analyzing streaming logs, and adapting user preferences within sessions.
Conversely, the fact that fast weights are updated directly by input also means that the attack surface can increase. Whether malicious input pollutes temporary memory, how to isolate multiple user requests, when to reset update status, and how to leave audit logs are important design issues in real-world services. These items were not experimentally tested in the paper.
11. Frequently Asked Questions
Q1. Is In-Place TTT the same as general fine-tuning?
It's not the same. Fine-tuning usually involves long-term changes to the model using separate training data and procedures. In-Place TTT temporarily updates some Wdowns while processing the input sequence and reverts them at document boundaries. However, continuous training is required to make this mechanism work well.
Q2. Do all model weights change during inference?
No. The paper fixes Wup and Wgate of gated MLP and uses Wdown as fast weights. In large-scale model experiments, we applied it to every sixth layer.
Q3. Is this a technology that replaces Attention?
No. The key differentiator of this paper is that attention is maintained and MLP adaptation complements it. Because Attention handles token relationships within chunks, In-Place TTT can use larger chunk-wise updates.
Q4. Is the context length virtually unlimited?
You can't be so sure. The 128k trained model was better than the baseline on 256k RULER, but the absolute score was 43.9. This is evidence of improved use and extrapolation of long sentences, but not evidence of perfect memory for unrestricted context.
Q5. What is the most important line from a patent perspective?
It is a combination relationship that uses the existing MLP output projection of the pre-training Transformer as fast weights, applies the state before the update to the current chunk, and reflects the update calculated as the goal, including subsequent token information, from the next chunk.
12. Conclusion
In-Place Test-Time Training does not propose an entirely new architecture to bring dynamic adaptive capabilities to LLM. We reinterpret the existing MLP's Wdown in a rapidly changing state and design a chunk-wise scan suitable for the language model purpose and parallel hardware. This is the most interesting part, both technically and patent-wise.
The experiments show significant improvements in long-form RULER and perplexity, but drop-in should not be misunderstood as a no-learning application. Continual training with a 35B token scale was used, fast weights are reset at document boundaries, and the actual agent, security, and multi-tenant environment has not yet been verified.
Ultimately, this paper is closer to a practical design that converts the internal MLP of a pre-trained LLM into a session-specific adaptive memory rather than completing a permanently self-learning LLM. From a patent perspective, it is appropriate to focus on the relationship that combines Wdown's in-place adaptation, LM-aligned target, apply-then-update causality, and context-parallel scan rather than individual elements.
This article reflects the information available when it was published. Contact us to discuss your circumstances.