# Why Do LLMs Lose Track as Conversations Get Longer?

Conversational AI can lose track of requirements as a discussion grows longer. A model may repeat a corrected mistake or continue relying on an earlier assumption even though the conversation remains in its context...

Source: https://www.iplexlaw.co.kr/en/blog/1530608

HOME / NEWS & INSIGHTS NEWS & INSIGHTS Why Do LLMs Lose Track as Conversations Get Longer? Conversational AI can lose track of requirements as a discussion grows longer. A model may repeat a corrected mistake or continue relying on an earlier assumption even though the conversation remains in its context... AI & Software 2026.09.01 published IPLEX 13 min read https://arxiv.org/pdf/2505.06120 Author: IPLEX IP Law Firm Yongduck Kim, Patent Attorney There are strange moments when using conversational AI like ChatGPT. The model, which initially followed fairly accurately, goes in the wrong direction as conditions are added one by one. Sometimes you make mistakes again about what you just corrected, and sometimes you hold on to the assumptions you made in the beginning. Even though all the conversation records remain, they act as if they missed some important terms. Simply explaining this phenomenon as “I can’t remember because the context is long” misses the point. “LLMs Get Lost in Multi-Turn Conversation,” published by researchers at Microsoft Research and Salesforce Research, compared when the same task was instructed completely at once and when instructions were divided into several conversation turns. What changed was not the amount of information, but the way in which the information was disclosed. Average performance fell from about 90 to about 65 points. The paper reports an average relative decline of 39% across the evaluated settings. Aptitude, reflecting the model’s best performance, declined by 16% on average, while unreliability increased by 112%. The main problem was therefore a loss of consistency in using the model’s capabilities. 1. The problem is not ‘memory’ but ‘how to handle incomplete requests’ Users often provide requirements gradually: they may first request a report, then specify its length, ask for a table and finally identify the audience and purpose. Many LLM benchmarks instead provide all requirements in the initial prompt. The paper investigates this difference. The first condition distinguished by the researchers is Fully-Specified. Since the purpose of use, constraints, input data, and output format are all included in the first message, the model can immediately produce the final answer. Conversely, in an underspecified multi-turn, only the major purpose is given in the first message, and detailed conditions are sequentially revealed in subsequent turns. At this time, the model should leave blank or ask questions about parts that are not yet known, but in reality, it often guessed the blanks and completed the answer too early. The important point is that the multi-turn format itself is not always a bad thing. In tasks where the results of each turn can be linked independently, such as translation, the performance drop was relatively small. On the other hand, problems emerged significantly in tasks where the entire existing code, SQL, calculation formula, and summary had to be re-edited every time a new condition was introduced. In the end, the difficulty was whether the entire previous dialogue should be reintegrated according to the new conditions. 2. The researchers broke the conversation into small pieces and had them solve the same problem again. The researchers divided the complete instructions into atomic pieces of information, then placed the high-level purpose in the first piece and detailed conditions in the remaining pieces. The user simulator delivered at most one condition that had not yet been revealed at each turn, and the model being evaluated responded as if talking to a regular user without knowing the structure of the experiment. This method is called Sharded Conversation. The model's responses were categorized into questions, holds, debates, rejections, hypothetical candidates, and actual answer attempts. When the model attempted an answer, it was scored using an evaluator appropriate for the task, including code execution, SQL execution, exact match, BLEU, and summary scores with citations. If the answer was incorrect, the next condition was revealed, and if the correct answer was given or there were no further conditions to be revealed, the conversation ended. The comparison conditions were also carefully designed. FULL is the baseline condition for providing the original complete directive in the first turn. SHARDED discloses conditions in several turns. CONCAT combines the sentences divided in SHARDED into one turn and checks whether it is due to simple sentence re-expression or information loss. RECAP rearranges all conditions at the end of the SHARDED conversation, and SNOWBALL cumulatively repeats the conditions so far each turn. By comparing these five conditions, we can to some extent separate “the act of sharing information” and “failure to integrate over multiple turns.” 3. The same phenomenon was repeated with 15 models and 200,000 conversations. The scope of the experiment is not limited to one or two models. The researchers used six types of tasks, including code generation, database tasks that convert natural language into SQL, API call generation, elementary mathematics, tasks that explain tables in sentences, and tasks that cite and summarize multiple documents. For each task, 90 to 120 split instructions were created, making a total of 600 instructions. The comparison target included 15 representative models at the time, including GPT-4.1, GPT-4o, o3, Gemini 2.5 Pro and Flash, Claude 3.7 Sonnet, DeepSeek-R1, Llama series, Phi-4, OLMo 2, and Command-A. By repeating the combination of model, directive, and conversation conditions several times, the total number of conversations exceeded 200,000. The reason why the same problem is repeated is also important. Because LLM generates sentences probabilistically, it is difficult to judge stability with just one success or failure. By running the same problem multiple times, you can see not only the average performance but also the gap between when it worked well and when it failed badly. The key point of this paper is that it presents the “gap” as a separate reliability issue. 4. The number to pay more attention to than the 39% drop is the ‘112% increase’. The paper measures aptitude and unreliability as well as average performance. Aptitude reflects the upper end of performance across repeated runs. Unreliability is the gap between the 90th- and 10th-percentile scores for the same instruction; a larger gap means greater variation between runs. SHARDED performance was lower than FULL across the evaluated models and tasks. Mean scores fell by about 25 points, from approximately 90 to 65; the paper reports an average relative decline of 39% across settings. Aptitude fell by 16%, while unreliability rose by 112%. The gap between the best and worst runs for an instruction averaged about 50 points. The CONCAT results strengthen this interpretation. When the divided conditions were combined once again, performance was maintained at about 95.1% of FULL. In other words, it is difficult to say that performance deteriorated because information was lost or expressions were changed when sentences were broken into smaller pieces. In a situation where the same information had to be received over several turns, the process of the model integrating the conversation state itself was a problem. So the message of this paper is not “Even the latest LLM can’t do multi-turns.” A more accurate expression is “You can do it well, but you cannot reliably reproduce that ability every time.” From a product perspective, it is important to look at how the results fluctuate when the same requirements are presented in different orders and conversation paths, rather than the highest benchmark score. 5. Why do models get lost in conversations? Complete your answer too early In the code and math tasks, when the first answer attempt occurred within the first 20% of the conversation, the average score was 30.9. Conversely, if you waited until the last 20%, the average score was 64.4. If the model creates an answer first even though the necessary conditions have not yet been fully disclosed, there is a greater possibility that the answer will contain assumptions that the user has not heard. Anchored in early assumptions In LLM, there is a tendency to complete answers by filling in plausible values from the context rather than leaving unknown parts as is. The problem is when new conditions come in from behind. Models often modify only part of the structure they have already created rather than discarding and recalculating the entire answer. The initial flawed premise then becomes the framework for the conversation, and subsequent corrections continue to build on top of it. The more you edit, the bigger the answer becomes. The final SHARDED answer was 20-300% longer than the FULL or CONCAT answer. Looking at the correct answers alone, code answers were 27% longer on average and SQL answers were 14% longer on average. Researchers call this Answer Bloat. This is because the model responds by adding exceptions and correction explanations rather than deleting the previous answer and calculating anew. The middle turn becomes weaker and the first and last turns become stronger. In the long summary task, the eighth turn's summary cited 20% of the documents published in the last turn, while the documents published in the second and third turns cited only 8% each. This is a phenomenon in which information at the beginning and end of a conversation is reflected relatively strongly, and conditions entered in the middle are weakened. The researchers named this Loss-in-Middle-Turns. Long, friendly answers can actually become noise. In five of the six tasks, the shortest response group performed 10 to 50 percent better than the longest response group. The longer the description is created in the early turns, the more assumptions and workarounds the user may leave unstated, which become part of the new input in later turns. In a multi-turn situation, the ability to determine “what we don’t know yet” and ask short confirmation questions may be more important than the ability to pour out long answers at the beginning. 6. Longer context windows and more reasoning were not enough The easiest remedy is to repeat the conversation again. If all conditions are rearranged at the end like RECAP, performance is improved over SHARDED. However, in actual services, it is difficult to know when the user stated the last condition. Like SNOWBALL, repeating the previous conditions every turn alleviated about 15-20% of the performance degradation, but the prompt quickly became longer and did not recover to the FULL level. Reducing generation randomness had limited benefits. Lower temperatures substantially reduced unreliability in single-turn tasks, but brought much smaller gains in multi-turn conversations. Even at zero temperature, approximately 30% unreliability remained. Small differences in one turn change the next turn’s input, allowing errors to accumulate. Models that use more inferential calculations also could not be solved automatically. Inference models such as o3 and DeepSeek-R1 also showed multi-turn performance degradation similar to non-inference models. The researchers point out that inference models tend to produce longer answers on average. The longer the answer, the more assumptions the model makes about itself, and the more difficult it can be to distinguish between user needs and the model's estimates in the next turn. 7. Conversational AI products should be designed for ‘state management’ rather than ‘answer generation’ If you read these results from a product design perspective, the direction becomes pretty clear. First, we need a final answer gate that does not produce a complete answer until the prerequisites are met. This means that rather than a simple “if you don’t know, ask” prompt, control logic is needed to determine whether the final creation is allowed in the current state. Second, it is important not to mix user conditions and model assumptions within the same conversation record. Requirements confirmed by the user, items not yet confirmed, values temporarily estimated by the model, and conditions subsequently withdrawn must be stored separately in a structured state. That way, when new conditions come in, you can decide what to keep and what to discard. Third, a procedure is needed to verify whether new information conflicts with existing assumptions and, if so, to invalidate related intermediate results and draft answers. Just telling the model, “Modify,” can create an Answer Bloat that adds to the existing structure while preserving it. If necessary, it is more reliable to start a new inference branch by taking only verified conditions rather than modifying the previous answer. Fourth, rather than summarizing the entire conversation at certain points, it may be useful to reorganize “only verified user conditions” into fully explicit directives. Using signals such as the number of turns, number of tokens, number of condition conflicts, and number of answer modifications to determine the timing of resummarization or restart is also a design point at the product level. Lastly, evaluation should not end with just the average percentage of correct answers. The same requirements should be run repeatedly in different information disclosure sequences to measure performance variance, assumption withdrawal rate, recovery rate after wrong answer, and best-worst gap. The reliability issue that this paper addresses is ultimately not a question of “can you do it well once?” but “can you do it reliably in multiple ways?” 8. In patent practice, the key is ‘dialogue control structure’ Rather than proposing a new base model architecture, this paper is closer to a study that presents an evaluation framework and reliability indicators that reproduce and quantify the vulnerabilities of conversational AI. Therefore, when examining technical differentiation at the product or service level, the specific state management structure and control procedures that implement that goal are more important than the abstract goal of “LLM handles multi-turn well.” For example, unspecified condition detection that determines whether the final answer can be created with only the current input, requirements ledger that stores confirmed conditions, unconfirmed conditions, model assumptions, and withdrawal conditions separately, assumption invalidation that detects conflicts between new statements and existing assumptions and selectively discards intermediate outputs, summary trigger that reorganizes the conversation state into a fully explicit prompt under certain conditions, conversation branching that starts a new inference by inheriting only verified conditions, and a reliability evaluation system that measures recovery rate and volatility by repeating different information disclosure sequences, etc. It can be specified as a technical configuration. However, differentiation is not easy with abstract ideas such as “summarize previous conversations,” “ask questions when necessary,” and “remember conversations.” It becomes easier to explain technical features and effects by specifying what information is stored in what data structure, what signals are used to judge unspecified states, what units the assumptions are connected to and invalidated, when the final model call is blocked or allowed, and how the error rate, recovery rate, token usage, and delay vary as a result. The paper was published on May 9, 2025. Its sharded simulation, FULL/CONCAT/SHARDED comparisons, RECAP/SNOWBALL conditions and definitions of aptitude and unreliability must therefore be considered when assessing later inventions. Patent protection would require a specific contribution, such as product-specific state management, conflict checking, model routing, resource savings or safety controls, assessed against the prior art. 9. Users can also reduce the probability of failure by slightly changing their conversation style. There are lessons that can be applied to ordinary users as well. This does not mean that all conditions must be written perfectly in the first request. Rather, if the conditions are less defined, it is better to specify “If there is insufficient information, ask a confirmation question first” to prevent the model from arbitrarily filling it out. When adding conditions, it is a good idea to clearly state which of the existing conditions to keep and which to change or delete. If the conversation is long, you can ask once in the middle, “Please rearrange only the conditions I have confirmed so far, and indicate separately what you have estimated.” If the model already seems to be deeply anchored in an incorrect structure, it may be more stable to insert only the confirmed conditions into a new dialogue at once and start over, rather than continuously modifying it. For important tasks, a realistic verification method is to run the same fully explicit prompt again in a new dialog and compare the results. Conclusion. A good conversational AI is closer to a ‘system that returns from a wrong path’ rather than a ‘model that gets it right from the beginning’. This paper shows that models with high single-turn benchmarks do not necessarily provide the same quality reliably in actual interactive services. When requirements are released over multiple turns, the model can falter by creating answers too early, holding its assumptions as facts, weakly reflecting conditions in mid-turns, and continuing to append to existing answers. Therefore, the competitiveness of the next generation of conversational AI will likely not be determined solely by more knowledge or longer answers. The ability to wait when information is insufficient, the ability to distinguish between user needs and model estimates, the ability to withdraw existing judgments when new conditions come in, and the ability to start over with only the verified state when going down a wrong path are the core of product reliability. Read the Korean source This article reflects the information available when it was published. Contact us to discuss your circumstances. Discuss this topic ↗ All articles TALK TO IPLEX Discuss your IP questions We consider your technology and business needs together. ↗ Contact us Newer Transferring Patent, Trademark and Design Rights ↗ Older Column: Google’s Generative AI Reasoning Patent — Teaching the Process of Thinking ↗ Related insights AI & Software 2026.10.01 FiX: fine-grained forgetting in softmax attention Yongduck Kim examines FiX’s feature-wise gates, numerical implementation and paged cache, distinguishing reported gains from unresolved limitations. ↗ Read article AI & Software 2026.09.30 MHAR: Reading earlier layers through different feature subspaces Yongduck Kim examines Multi-Head Attention Residuals: depth routing, reported training results, implementation costs and the relationship between technical features and effects. ↗ Read article AI & Software 2026.09.28 Column: Claude Computer Use and the Data That Trains AI Agents Writing for AI Times, IPLEX Managing Partner Yongduck Kim examines the training data behind computer-operating AI agents through U.S. Patent No. 12,585,862. ↗ Read article

- https://www.iplexlaw.co.kr/en
- https://www.iplexlaw.co.kr/en/blog/category/ai
- https://arxiv.org/pdf/2505.06120
- https://www.iplexlaw.co.kr/forum/view/1530608
- https://www.iplexlaw.co.kr/en/contact
- https://www.iplexlaw.co.kr/en/blog
- https://www.iplexlaw.co.kr/en/contact
- https://www.iplexlaw.co.kr/en/blog/1531646
- https://www.iplexlaw.co.kr/en/blog/1530376
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals
- https://www.iplexlaw.co.kr/en/blog/claude-computer-use-agent-patent
- https://www.iplexlaw.co.kr/en/blog/claude-computer-use-agent-patent
