# TROLL Paper Analysis: Token-Level Trust Regions for More Stable LLM Reinforcement Learning Beyond PPO Clipping

Reward-based reinforcement learning is increasingly used to improve the reasoning capabilities of large language models. PPO-style policy optimization is a common approach. Methods such as GRPO, Dr.GRPO, GSPO and REINFORCE++ modify advantage estimation or normalization while seeking to avoid excessive policy changes in a single update...

Source: https://www.iplexlaw.co.kr/en/blog/1533996

HOME / NEWS & INSIGHTS NEWS & INSIGHTS TROLL Paper Analysis: Token-Level Trust Regions for More Stable LLM Reinforcement Learning Beyond PPO Clipping Reward-based reinforcement learning is increasingly used to improve the reasoning capabilities of large language models. PPO-style policy optimization is a common approach. Methods such as GRPO, Dr.GRPO, GSPO and REINFORCE++ modify advantage estimation or normalization while seeking to avoid excessive policy changes in a single update... AI & Software 2026.09.08 published IPLEX 23 min read Author: IPLEX IP Law Firm Yongduck Kim, Patent Attorney The number of cases of applying reward-based reinforcement learning to strengthen the inference ability of large language models is rapidly increasing. A method that often appears at this time is PPO-type policy optimization. Methods that improve advantage estimation or normalization methods, such as GRPO, Dr.GRPO, GSPO, and REINFORCE++, generally use PPO-style clipping as a device to prevent policies from changing too significantly at once. The TROLL paper addresses this issue by replacing PPO-style clipping with a token-wise KL trust region projection. It changes how updates are constrained, rather than redesigning the reward function or advantage estimator: an update outside the permitted region is projected onto the nearest admissible distribution. Key finding: TROLL replaces PPO-style clipping with differentiable token-wise trust region projection and sparse probability representations to improve the stability, speed and final performance of LLM reinforcement learning. Paper basic information The title of the paper is ‘TROLL: Trust Regions Improve Reinforcement Learning for Large Language Models’, and it is a published paper at ICLR 2026. The authors are Philipp Becker, Niklas Freymuth, Serge Thilges, Fabian Otto, and Gerhard Neumann, and are affiliated with Karlsruhe Institute of Technology and Microsoft Research. The core topic is replacing PPO-style clipping with differentiable discrete trust region projections, verifying its applicability across mathematical reasoning and code generation, multiple model families, and advantage estimation methods. 1. Question asked by the paper: Why should PPO clipping be reconsidered? In LLM's reinforcement learning post-training, the responses produced by the model are evaluated and rewards are given, and the policy is updated so that highly rewarded responses appear more often in the future. The problem is that if the update rate is too large, the model's output distribution may suddenly collapse. Even small changes to the training data can cause response lengths to spike, specific patterns to repeat, or exploration diversity to quickly disappear. To prevent such rapid changes, PPO clips the probability ratio of the new policy and the old policy so that it does not fall outside a certain range. Although it has been widely used because it is simple to implement and easy to apply to large-scale learning, the authors point out that this method is a crude approximation of the original trust region concept. In particular, because the gradient is truncated for tokens whose probability ratio is outside the critical range, there may not be enough information left about “how to get a policy that has moved in the wrong direction back into a safe range.” If we compare clipping to a car, clipping is similar to a device that suddenly cuts off the accelerator pedal input when the speed limit is exceeded. It can stop you from going too fast, but it doesn't tell you in which direction or how smoothly you should slow the car down. On the other hand, trust region projection is closer to calculating the closest point within the allowable area from the current location and moving the vehicle there. The difference is that the state of the control object is explicitly placed within the safe area. The differences between PPO-type clipping and TROLL are summarized as follows. PPO-style clipping is a method of cutting the policy ratio of a sample token within a critical range, and if it falls outside the range, the objective function and gradient may be cut. In theory, it is a method of heuristically approximating the trust region and does not perform a separate distribution projection. On the other hand, TROLL projects the entire probability distribution for each token position into the KL range, and if it is outside the boundary, calculates the distribution closest to the current distribution within the tolerance area. This is close to convex optimization with explicit KL constraints, and in large vocabularies it uses sparse token distributions to suppress costs. In both methods, there is no additional cost in the inference step. Key distinction: TROLL replaces the clipping mechanism used to update a policy from estimated advantages. It does not replace the advantage estimation methods used by PPO, GRPO, Dr.GRPO or GSPO. 2. Concepts to know first: policy, advantages, KL trust region 2.1 Policy is the probability distribution of the next token In a language model, a policy is a probability distribution assigned to each next token based on the current context. For example, in the same context, tokens such as “correct,” “therefore,” and “but” are assigned different probabilities. Reinforcement learning adjusts model parameters to increase the probability of tokens included in good responses and lower the probability of tokens included in bad responses. 2.2 Advantage measures how much better a choice is than the baseline Advantage is a signal of how much better the response or token selection was than the baseline behavior. PPO can use a separate value model, and GRPO selects multiple responses to the same question and calculates advantage based on relative rewards within the group. Because TROLL is not specific in how it calculates this advantage, it can be combined with multiple algorithms. 2.3 KL divergence measures how much two probability distributions differ. KL divergence is a measure of how far the token probability distribution of the new policy deviates from the distribution of the previous policy. TROLL constrains the KL divergence between the new and old policies at each output location to not exceed a predetermined boundary value ε. The limitation here is not simply looking at the proportion of a single token sampled, but rather the distribution of tokens themselves considered at that location. 2.4 A trust region limits the change allowed in one update The trust region is a safety radius that says, “During this training phase, let’s not let the policy deviate more than this amount.” If the boundary is too small, learning will be slow, while if it is too large, the policy may move too far and become unstable. The paper also confirms that if ε is too conservative, the speed slows down, and if it is too large, the final success rate tends to worsen. Without the formulas: TROLL allows the policy to improve in response to rewards. If an update exceeds the permitted KL radius, it projects the distribution back to the nearest point within the trust region. 3. TROLL’s core idea: Project, not cut. Figure 1. Paper Figure 1. The left side shows the relationship between the old policy, new policy, and projected policy in the three token distributions, and the right side compares the learning speed of TROLL and Clip in math and code evaluation. Source: Becker et al., ICLR 2026, original page 2. How to easily read the left side of a picture The three vertices of the triangle represent an extreme case in which all probabilities are concentrated on the three tokens: cat, troll, and hamster. The red dot is the previous policy used when collecting data, and the blue dot is the new policy created with the current parameter update. If the new policy moves too far toward the hamster and leaves the trust region, TROLL moves the probability distribution from the blue dot to the nearest green dot inside the circle. The important point is that this does not necessarily mean reverting to the previous policy. If the new policy is inside the trust region, it is left as is, and only minimally modified when it is outside. Therefore, it prevents one update from becoming too large while maintaining the direction of learning based on rewards. What the right side of the picture says In an experiment learning Qwen3-14B with GRPO, TROLL with a solid line reaches a high success rate faster than Clip with a dotted line. Especially in code generation, there is a large gap based on the same wall clock time. This result supports the paper's claim that not only did the final score increase, but the efficiency of training time also improved. Key point: Clipping can suppress gradients for some updates outside the clipping range. TROLL retains gradient information through differentiable projection, which the authors argue improves training stability and efficiency. 4. TROLL’s actual processing flow Figure 2. Paper Figure 2. After sparsifying the probability distribution of the current LLM and the probability distribution of the previous policy stored in the rollout buffer, the distribution to be used for learning is calculated through KL trust region projection. Source: Becker et al., ICLR 2026, original page 4. Step 1. Store previous policy Stores the token probability of the policy used when generating the response in the rollout buffer. This will be the benchmark for judging how different the new policy is. Step 2. Recalculate the probability distribution of the current policy. Calculate the token probability of the current model being updated using the same response data. Step 3. Sparse both distributions: Instead of storing the entire vocabulary, we leave only the top tokens that account for most of the probability mass. Step 4. Check the KL bounds Check whether the current distribution is more than ε away from the previous distribution. If it is within the boundary, no projection is performed. Step 5. Only project distributions that have crossed the boundary. If a boundary has been crossed, calculate the distribution that is closest to the current distribution and at the same time is within the trust region around the previous policy. Step 6. Update the policy with the projected distribution. The projected distribution is used to calculate the policy ratio, and a regression term is also added so that the original model output follows the projection result. Why is a regression term needed? If only the projected distribution is used for the policy ratio, the non-projected distribution output by the actual LLM may remain outside the trust region. The paper sets the projection result as if it were the correct answer and adds a regression term to make the original output closer to that, so that the next update continues stably. 5. Trust region projection to understand without math 5.1 The closest compromise between the current policy and the previous policy Simply put, the problem TROLL solves is to find a probability distribution that satisfies two conditions simultaneously. First, your current model should be as close as possible to the new distribution you want to create. Second, it must not go further than the KL boundary ε from the previous policy used for data collection. This problem is written in convex optimization form, and the solution is a geometric interpolation of the log-probabilities of the current policy and the previous policy. 5.2 Projection intensity automatically varies from token position to token position If the current distribution is already within the safe region, the projection strength will be 0 and will not change anything. Conversely, if it deviates significantly from the boundary, it pulls more towards the previous policy. Because the values that determine this intensity are converted into a single scalar optimization problem, they can be computed much more efficiently than solving the entire high-dimensional optimization directly. 5.3 Why differentiable projections are important In reinforcement learning, there must be a gradient from the final loss to the model parameters. TROLL incorporates numerical optimization into the projection process, but uses KKT conditions and implicit differentiation to calculate how the optimal projection strength changes with model output. This allows us to treat the projection itself as a layer in the learning graph. For a hypothetical example, if the probability of a correct candidate token was 0.45 in the previous policy and increased to 0.60 in the current policy, TROLL preserves the reward direction but maintains the change if it is within the KL boundary. If other influential tokens are significantly lowered from 0.35 to 0.15, some of them are restored if the change is excessive, and if the remaining tokens are changed from 0.20 to 0.25, they are normalized together so that the entire probability distribution is valid. Common misconception: TROLL does not independently cap each token probability. It finds a valid categorical distribution whose probabilities sum to one and that satisfies the KL constraint. 6. How to handle over 150,000 token candidates In principle, the probability distribution of the entire vocabulary should be stored for each response token and a KL projection should be performed. The Qwen3 tokenizer used as an example in the paper has 151,936 token candidates. Considering long responses and large batches, storing the entire logit as is is not realistic in terms of memory and computation. 6.1 Most of the probability mass is concentrated in a few tokens The distribution of the following tokens in a language model is usually very sharp. Although there are many possible tokens, only a few candidates actually receive a high probability. The paper uses a method that selects up to K tokens as candidates and leaves them only until the cumulative probability mass reaches 1-δ. At the default setting K=64, δ=10⁻⁵, we report preserving more than 99.999% probability mass with only 5 to 10 tokens on average. 6.2 The token actually selected must be left behind Leaving only the high probability tokens may result in missing the actual response tokens sampled with low probability. However, the policy ratio and gradient of reinforcement learning directly depend on the tokens actually generated. Therefore, TROLL includes the tokens selected by the model in the sparse set regardless of their probability ranking. 6.3 Discarded tokens do not become completely zero Giving probability 0 to tokens excluded from sparsification may make the KL divergence calculation unstable. The paper assigns a small base probability to the excluded tokens and normalizes them again with the remaining tokens. Additionally, sparsification is performed on a chunk-by-chunk basis to prevent large memory peaks in long sequences. Sparsification combines four mechanisms. The top-K limit caps the number of tokens retained in high-entropy situations; the paper uses K=64 by default. Selection stops once the cumulative probability mass is sufficiently preserved, with a default threshold δ=10⁻⁵. The sampled token is always included because it is needed to calculate the policy ratio and gradient. Excluded tokens receive a small default probability pd > 0 rather than an exact zero. From a patent perspective, the relevant technical features may lie in the combination that makes token-wise trust region projection practical: the probability-mass threshold, top-K limit, mandatory inclusion of the sampled token and renormalization. 7. How was the experiment designed? The paper compares TROLL and existing Clip in the RLVR environment for mathematical reasoning and code generation. To confirm that the results were not tailored to a specific model or a specific advantage estimator, we experimented with a wide range of model families, parameter scales, datasets, and reinforcement learning algorithms. Looking at the experimental data and model composition, in the mathematics field, DAPO-Math, MATH evaluation group consisting of MATH500·AMC·AIME·OMNIMATH·OlympiadBench·Minerva, GSM8K, and Eurus-Math were used, and in the code field, Eurus-2-RL-Code was used. The Qwen series includes Qwen3 0.6B·1.7B·4B·8B·14B and Qwen2.5 0.5B·1.5B·3B·7B. In addition, Llama 3.1·3.2, SmolLM3, Apertus, and FineMath-Llama were compared. Policy optimization methods covered GRPO, PPO, Dr.GRPO, GSPO, and REINFORCE++, and the possibility of combining adaptive clipping BAPO and non-clipping GPG was also examined. The experiments test whether TROLL improves learning speed and stability across algorithms and models, as well as measuring the additional computational cost. 8. Performance changes in the Qwen model Figure 3. Paper Figure 3. From Qwen3 600M to 14B, TROLL is indicated by a solid line and Clip is indicated by a dotted line. Overall, the TROLL curve rises faster and higher in DAPO training/evaluation, MATH evaluation, and Eurus-Code training/evaluation. Source: Becker et al., ICLR 2026, original page 7. 8.1 Mathematical reasoning improved by 3 to 10 percentage points. The paper summarizes that in DAPO-Math-based experiments, TROLL showed results that were generally 3 to 10 percentage points higher than Clip, or about 5 to 15% relative value. Improvements in the same direction were seen from the small model to the 14B model, and differences continued not only in the training set but also in the DAPO evaluation and MATH evaluation. 8.2 The gap is even bigger in code generation Eurus-Code reported an improvement of 7 to 18 percentage points, or about 18 to 30 percent in relative terms. The paper presents an example where the Qwen3-1.7B TROLL model showed higher results than the Qwen3-4B Clip model. This result shows that improving the stability of the update method can have a greater effect than increasing the model size. 8.3 Case where a small TROLL model is close to a large Clip model In mathematical experiments, the Qwen3-4B TROLL model showed performance that was close to that of the Qwen3-14B Clip model. Of course, we cannot generalize from this result alone that a smaller model replaces a larger model. However, it clearly shows that the choice of post-training algorithm is as important a performance variable as the parameter size. 9. Is the effect maintained even if the method of estimating the advantage is different? Table 1. Paper Table 1. Comparison of Clip and TROLL versions of GRPO, Dr.GRPO, PPO, GSPO, and REINFORCE++ in Qwen3-8B and Qwen2.5-7B-Instruct. Source: Becker et al., ICLR 2026, original page 8. The most notable part of the table is the GSPO. When combining GSPO and Clip in Qwen3-8B, the success rate collapsed to virtually zero, but using TROLL resulted in a DAPO rating of 0.706. Even in Qwen2.5-7B-Instruct, GSPO Clip showed low results, while TROLL stabilized learning. In other algorithms, TROLL generally increased the success rate. The paper explains that in the tested range, changing Clip to TROLL often gave greater benefits than changing the advantage estimation method. This shows that TROLL is a stabilizer that can be commonly used in the policy update layer, rather than an auxiliary technique for a specific advantage calculation method. Looking specifically at Qwen3-8B's DAPO evaluation, in GRPO, Clip increased by 0.051 from 0.640 to TROLL 0.691, and in PPO, it increased by 0.113 from 0.602 to 0.715. GSPO collapsed in learning to 0.000 in Clip, but recovered learning by recording 0.706 in TROLL. REINFORCE++ also increased by 0.102 from Clip 0.626 to TROLL 0.728. point of interpretation TROLL is not “a better advantage estimator,” but “a more principled constraint method that reflects advantage estimation results in policy.” It is therefore important whether the effects replicate across different advantage estimation methods. 10. Is the effect maintained even if the model series changes? Figure 4. Paper Figure 4. Comparison of the final performance and learning curve of TROLL and Clip in various models such as Llama, SmolLM3, and Apertus and the combination of GSM8K and DAPO. Source: Becker et al., ICLR 2026, original page 9. TROLL's advantages are not limited to the Qwen series. In particular, there was a significant difference in models that showed little learning signal with the existing Clip. The GSM8K results of Llama3.1-8B improved from Clip 0.000 to TROLL 0.759, and Apertus-8B improved from 0.156 to 0.697. Apertus-8B-Instruct also increased from 0.688 to 0.824. Looking at the learning curves, some Llama models start to improve in performance after a significant amount of time in Clip, but learning signals appear earlier in TROLL. This means that the “fast learning” and “stable updating” emphasized in the paper are observed even when the model family is changed. However, not all combinations are necessarily better. The GSM8K results of FineMath-Llama-3B were almost the same or slightly lower, with Clip 0.750 and TROLL 0.746. This exception shows that TROLL is not a one-size-fits-all formula that guarantees performance at all times, but rather a method that becomes particularly valuable in environments where traditional clipping is highly unstable. Criteria for Reading Reliable Papers Instead of just picking out good results, you should also look at cases and exceptions where the amount of improvement is small. The persuasiveness of this paper is not that it wins in every cell, but that improvements in training stability are repeatedly observed across a variety of models and algorithms. 11. Hyperparameter, entropy, and computational cost analysis Figure 5. Paper Figure 5. The left shows performance according to the KL boundary and the number of sparsity tokens, the upper right shows memory and execution time, and the lower right shows the entropy change during learning. Source: Becker et al., ICLR 2026, original page 10. 11.1 KL boundary ε If ε is too small, updates will be overly conservative and training will be slow. Conversely, if it is too large, the policy can move a lot at once, resulting in poor final performance. However, the paper explains that convergence performance is relatively stable within an appropriately small and conservative range. 11.2 Sparsification limit K K=16 performed poorly because it did not sufficiently approximate the entire distribution. K=256 increases computational cost but has no significant performance gain over the default K=64. In other words, there needs to be a balance between preserving enough tokens that are important, but not preserving too many tokens unnecessarily. 11.3 Entropy Conservation Clip showed a tendency for the entropy of token distribution to rapidly decrease as learning progressed. This can be interpreted as a phenomenon in which the model focuses probability on a few patterns that it already knows. TROLL maintains higher entropy while also increasing the success rate. In code generation experiments, we also observed a correlation between entropy preservation and performance improvement. 11.4 Additional costs measured in a controlled environment Looking at the additional costs measured in the control environment, VRAM usage increased from Clip 34.574 GiB to TROLL 36.157 GiB, requiring approximately 4.6% additional memory. The execution time increased from Clip 85.133 seconds to TROLL 92.906 seconds, an increase of approximately 9.1%. There is an additional cost during training. The paper uses sparse distributions to limit that overhead and explains that it depends mainly on vocabulary size rather than model parameter count, so its share of total training cost decreases as models grow. Inference does not use TROLL projections and incurs no additional cost from them. 12. Things to keep in mind when interpreting the paper results 12.1 TROLL is not a way to change the inference structure TROLL is a policy update method in the post-training stage. This is not a method of changing the model architecture or adding new modules during inference. Therefore, the advantage of the paper is that the inference operation and delay time of the trained model remain the same as before. 12.2 It does not solve the reward design problem If the reward function is poorly designed or the verifier is poor, the model may reliably optimize the wrong objective. TROLL is a device that ensures that policy does not change drastically, but does not fix the rewards themselves that define what is a good response. 12.3 The scope of the experiment is focused on mathematics and code RLVR Although the paper covers several model families and algorithms, the task is mainly mathematical reasoning and code generation that can verify correct answers. Further verification is needed to determine whether the same effect can be seen in general RLHF, long-term conversation, tool-using agents, and vision-language models where human preferences play a complex role. 12.4 Verified up to 14B dense model The authors experimented with dense models up to 14B in size. Vocabulary distribution, parallelism, communication cost, and sparsification error may appear different in larger models and MoE structures. The authors also suggest this as a follow-up research topic. 12.5 It is important to assume that the sparse probability distribution is well-fitted The practicality of TROLL relies on the property that a small number of tokens account for most of the probability mass. For special tasks or multimodal token distributions where high-entropy outputs are frequent, an average of 5 to 10 tokens may not be sufficient. In this case, adjusting K and δ or another sparsification strategy is needed. Balanced Evaluation Rather than simply saying that TROLL is “always better than clipping,” it is more accurate to read it as a study that showed through extensive experiments that explicit trust region projection can be advantageous in situations where clipping is unstable or the gradient loss is large. 13. Technology analysis from the perspective of a patent attorney specializing in artificial intelligence patents The evaluation below is a provisional technical analysis based solely on the information disclosed in the paper. To determine the actual novelty/inventive step or scope of rights, a separate prior art search and application progress confirmation are required. 13.1 Response to technical challenges and solutions Looking at the correspondence between technical challenges and solutions, to the heuristic constraints and gradient loss problem of PPO clipping, token-wise KL trust region projection on the categorical token distribution is applied to explicitly limit the update width while maintaining the gradient even in the constrained region. For the memory and computational burden incurred when projecting the entire vocabulary distribution, a sparse distribution combining the probability mass criterion and top K is used, allowing projection to be performed at a practical cost even with the vocabulary size of modern LLM. For problems where numerical optimization breaks the learning graph, we apply KKT condition-based implicit differentiation and differentiable projection to enable end-to-end learning from policy loss to model parameters. Additionally, when there is a gap between the projection distribution and the actual LLM output, a regression term targeting the projection result is used to drive the non-projection policy closer to the trust region in subsequent updates. For problems dependent on a specific advantage estimation method, a policy update structure that does not assume a advantage calculation method is used to ensure compatibility with PPO, GRPO, Dr.GRPO, GSPO, etc. 13.2 Differentiation is more likely to lie in combination than in individual elements Trust region and KL constraints are old concepts in the TRPO family, and there is also prior research on projection-based reinforcement learning, categorical distribution, high-order token sparsification, and OptNet-type differentiable optimization. Therefore, the relatively strong part of the paper is the specific combination of these elements linked to LLM's token-level policy update. In particular, the flow of storing, sorting, and renormalizing the rare token probabilities of the previous policy and the current policy, solving the one-dimensional dual problem by selecting only the boundary violation token positions, and using the projection distribution for the policy ratio and regression objective is an implementation structure that is distinct from simple “adding a KL penalty.” 13.3 Experimental data supporting technological effectiveness is important In this field, training collapse prevention, entropy preservation, wall clock time, memory growth, and reproducibility across multiple algorithms are as important as mathematical constructs. Rather than presenting only accuracy improvements in specifications or technical documents, it is more persuasive to present stabilization under conditions where Clip fails, sparsification error, projection occurrence rate, and convergence characteristics according to boundary value changes. 14. Conclusion: The next task of post-training shown by TROLL TROLL's message is clear. To improve performance in LLM reinforcement learning, we must not only improve the reward model, data, and advantage estimator, but also reexamine the last update rule that actually drives the policy. PPO clipping is a convenient and proven choice, but it can introduce instability and gradient loss in the process of roughly approximating the original trust region principle. TROLL solves this problem with a differentiable KL projection on the categorical probability distribution for each token, and the cost arising from a large vocabulary is reduced by sparsification. Its strength lies in its ability to demonstrate faster learning, higher final success rates, mitigation of training collapse, and preservation of entropy across multiple models and algorithms in mathematical reasoning and code generation. From a patent perspective, it is important to look at the entire processing flow of storage and reduction of the LLM token distribution, boundary violation determination, dual problem-based projection, implicit differentiation, and combination of policy ratio and regression term, rather than individual words such as trust region or sparsification. At the same time, since it is a field with many preceding elements, the actual possibility of obtaining rights must be judged based on specific combination relationships and evidence of effectiveness. In the future, an important verification challenge will be how well the same principles hold in larger models, MoE, general preference learning, tool-using agents, and multimodal models. Nevertheless, this paper is meaningful in that it asks the question, “Is clipping an obvious default?” and presents a realistic path for implementing principled trust region techniques on a modern LLM scale. Summary: TROLL projects token distributions into a permitted KL trust region with minimal adjustment. Sparse representations and a differentiable solver make the method practical for modern LLMs, with more stable and faster training than clipping in the reported mathematics and code experiments. Source and usage information Original text: https://arxiv.org/pdf/2510.03817 Paper citation: Philipp Becker, Niklas Freymuth, Serge Thilges, Fabian Otto, Gerhard Neumann, “TROLL: Trust Regions Improve Reinforcement Learning for Large Language Models,” ICLR 2026. This document is a technical commentary that organizes the attached papers in an easy-to-understand manner. It is not a legal opinion for specific legal judgment or determination of patentability, and actual review requires a separate investigation of related application documents and prior art. Read the Korean source This article reflects the information available when it was published. Contact us to discuss your circumstances. Discuss this topic ↗ All articles TALK TO IPLEX Discuss your IP questions We consider your technology and business needs together. ↗ Contact us Newer Why Trademark Applications Are Refused and How to Respond ↗ Older Russian Trademark Applications: Procedure, Registration and Key Rules ↗ Related insights AI & Software 2026.10.01 FiX: fine-grained forgetting in softmax attention Yongduck Kim examines FiX’s feature-wise gates, numerical implementation and paged cache, distinguishing reported gains from unresolved limitations. ↗ Read article AI & Software 2026.09.30 MHAR: Reading earlier layers through different feature subspaces Yongduck Kim examines Multi-Head Attention Residuals: depth routing, reported training results, implementation costs and the relationship between technical features and effects. ↗ Read article AI & Software 2026.09.28 Column: Claude Computer Use and the Data That Trains AI Agents Writing for AI Times, IPLEX Managing Partner Yongduck Kim examines the training data behind computer-operating AI agents through U.S. Patent No. 12,585,862. ↗ Read article

- https://www.iplexlaw.co.kr/en
- https://www.iplexlaw.co.kr/en/blog/category/ai
- https://www.iplexlaw.co.kr/forum/view/1533996
- https://www.iplexlaw.co.kr/en/contact
- https://www.iplexlaw.co.kr/en/blog
- https://www.iplexlaw.co.kr/en/contact
- https://www.iplexlaw.co.kr/en/blog/1535508
- https://www.iplexlaw.co.kr/en/blog/1533936
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals
- https://www.iplexlaw.co.kr/en/blog/claude-computer-use-agent-patent
- https://www.iplexlaw.co.kr/en/blog/claude-computer-use-agent-patent
