NEWS & INSIGHTS

AgentGym-RL Paper Analysis: Training Long-Horizon AI Agents with Multi-Turn Reinforcement Learning

An AI agent asked to navigate a website, find information and recover from a failed action faces a different task from simply answering a question. It must coordinate a sequence of actions, observations and decisions, including correcting errors along the way...

Illustration: AgentGym-RL Paper Analysis: Training Long-Horizon AI Agents with Multi-Turn Reinforcement Learning
Author: IPLEX IP Law Firm Yongduck Kim, Patent Attorney
The nature of the problem is completely different if you instruct the AI agent not to “say the right answer” but to “navigate the website to find the information you want, press the required button, and if that fails, correct the path.” This is because it does not end with a single response, but rather involves multiple actions, observations, and judgments. During this process, the agent may press the wrong button, repeat the same action, or give up too early.
.
The AgentGym-RL paper targets precisely this multi-turn decision problem. The researchers proposed an integrated reinforcement learning framework that separated the environment, agent, and learning algorithm, and combined it with ScalingInter-RL, which limits the number of interactions briefly at the beginning of learning and then gradually increases them. The core message of the paper is simple. Rather than giving the agent a long freedom of action from the beginning, it may be more stable to let the agent learn the basics through short tasks and then expand the scope of action.
One sentence summary. AgentGym-RL is an integrated framework that trains LLM agents through online reinforcement learning in various real-world environments, and ScalingInter-RL is a method of increasing long-term exploration ability while reducing initial training collapse by gradually increasing the maximum number of interactions.
1. Understand the paper in 3 minutes
First, the starting point of the problem is that existing LLM reinforcement learning is focused on a single response task, and in agent tasks that require multiple actions and observations, it lacks environmental diversity, training stability, and long-term interaction support.
To solve this, AgentGym-RL separates Environment, Agent, and Training, and connects web, search, game, embodied, and scientific environments into one learning pipeline through a standardized server-client structure and parallel rollout.
On the learning side, ScalingInter-RL sets the maximum number of interactions small initially and increases the number allowed as the training phase progresses. A curriculum that reliably learns familiar behaviors and then allows for longer exploration.
Looking at the experimental results, the paper's 7B model recorded BabyAI 96.67, TextCraft 91.00, SciWorld 57.00, WebArena 26.00, and Deep Search 38.25. However, it did not beat the commercial model in all tasks, and the superiority and inferiority of each task were different.
However, generalization was mainly verified in the same series environment, and physical reality environment and multi-agent learning remain as future tasks. Procedural errors and excessive interactivity were consistently observed in some scientific assignments and web navigation.
2. Key terms to know first
An LLM agent is a system where the language model does not simply generate sentences, but also calls tools, observes the environment, and selects multi-step actions.
Multi-turn reinforcement learning is learning that updates the policy based on the entire trajectory consisting of multiple actions and observations rather than a single answer.
The interaction horizon is the maximum number of times an agent can interact with the environment in one episode. For ease of understanding, this article also refers to it as the maximum number of actions or interaction length.
Exploration means trying new action paths; exploitation means using actions already known to work. Their balance affects reinforcement learning stability.
Rollout is the process of putting an agent into a real environment with its current policy to collect trajectories consisting of actions, observations, and rewards.
POMDP is a decision-making model where the agent does not directly see the entire state of the environment and must decide its next action based on only some observations. It works well for situations where you view a web page or game screen step by step.
GRPO is a reinforcement learning method that generates multiple trajectories in the same task and then updates the policy using the relative performance within the group.
A terminology point matters here. In this paper, “training from scratch without SFT” does not mean pre-training a language model from random initialization. The starting point is a pre-trained or instruction-tuned Qwen2.5 model. Agent behavior is then trained using environmental rewards, without first applying supervised fine-tuning to expert trajectories for the agent tasks.
3. Why was reinforcement learning difficult for existing AI agents?
3.1 Differences between single response reinforcement learning and multi-turn agent learning
In tasks where you submit an answer once, such as a math problem or code problem, you may only be rewarded for the accuracy of the final correct answer. On the other hand, in web browsing or scientific experiments, intermediate actions change the next state in a chain. One wrong click can change your entire subsequent observation, and the effects of your initial actions can appear several turns later. This increases the problem of assigning credit to determine which actions contributed to success over long-term trajectories.
As the number of interactions increases, the possible paths of action explode. If you do not explore enough, you will only repeat easy shortcuts, and conversely, if you allow too much exploration from the beginning, learning may collapse due to meaningless actions and high variance. The paper observes these two extremes as a performance upper bound for short fixed horizons and training collapse for long fixed horizons, respectively.
3.2 Without an integrated framework, the experiment itself is difficult
Each environment has different behavioral commands, observation formats, reset methods, and reward calculations. The web browser handles tabs and buttons, TextCraft uses crafting commands, and BabyAI performs spatial movement and door manipulation. SciWorld requires procedural actions like measuring temperatures, connecting circuits, and mixing chemicals. To connect them into a single reinforcement learning code, the environment interface, parallel execution, error recovery, and resource organization must be standardized.
AgentGym-RL addresses not only research ideas but also engineering bottlenecks that arise during long-term rollout. The paper explains that it fixes issues such as multi-browser execution in WebArena, parallel initialization in SciWorld, memory leaks in TextCraft and SciWorld, and full initialization after the end of an episode. In long-term agent learning, this system stability is as important as the algorithm itself.
Illustration: AgentGym-RL Paper Analysis: Training Long-Horizon AI Agents with Multi-Turn Reinforcement Learning
Figure 1. Paper Figure 1: Comparison of performance of five agent tasks and overall accuracy by model size. Source: Xi et al., p. 2.
The left side of Figure 1 shows performance for each of the five environments, and the right side shows the relationship between model parameter scale and overall accuracy. The paper highlights that the 7B-scale reinforcement learning model competes with much larger open source models and, in some environments, scores higher than commercial models. However, it is not possible to generalize all actual abilities from a single point on the graph on the right, and it must be seen that each benchmark's evaluation method and task composition are different.
4. How is the AgentGym-RL framework structured?
Illustration: AgentGym-RL Paper Analysis: Training Long-Horizon AI Agents with Multi-Turn Reinforcement Learning
Figure 2. Paper Figure 2: AgentGym-RL overall structure divided into Environment, Agent, and Training modules. Source: Xi et al., p. 3.
4.1 Environment module
The environment is packaged as an independent server, and agents request observations, check available actions, execute actions, and reset via HTTP-based clients. The advantage of this structure is that even if the web environment, search environment, and game environment have different internal implementations, the agent and learning module can be accessed through a common interface.
The five scenarios included in the paper are: WebArena is for realistic web navigation, Deep Search is for multi-level information exploration using a search engine, TextCraft is for text-based crafting games, BabyAI is for embodied tasks in a grid world, and SciWorld is for performing scientific experiments and procedures.
4.2 Agent module
The Agent module encapsulates an iterative loop that takes observations, makes inferences, and generates the next action. Prompt strategies, sampling settings, reward functions, as well as mechanisms such as long-term planning and self-reflection are designed to be extensible. The key is to separate the agent's reasoning logic from the detailed implementation of a specific environment.
4.3 Training module
The Training module batches collected trajectories in a parallel environment, calculates log probability and reference model probability, and performs advantage estimation and policy update. The paper explains that PPO, GRPO, RLOO, and REINFORCE++ are supported as online reinforcement learning algorithms, and SFT, DPO, and rejection sampling are also supported as auxiliary learning methods.
1. Initialize multiple task and environment clients in parallel.
2. Each agent generates an action based on its current observations.
3. The environment triggers the action and returns new observations and rewards.
4. Record actions and observations as a trajectory until the episode ends.
5. Calculate advantages from the collected trajectories and update the policy parameters.
6. Reset the environment to its initial state and repeat the next rollout.
An easy analogy: a driving school. Environment is similar to roads and signal systems, Agent is similar to driving trainees, and Training is closer to coaches who look at driving records and change the next lesson plan. By keeping roads separate from students and coaches, you don't have to rebuild the entire system if you add a new road, a different student model, or change grading.
5. Key proposition: ScalingInter-RL increases the number of actions step by stepIllustration: AgentGym-RL Paper Analysis: Training Long-Horizon AI Agents with Multi-Turn Reinforcement Learning
Figure 3. Paper Figure 5: ScalingInter-RL, which learns the basics in a short horizon and then expands to a long horizon. Source: Xi et al., p. 8.
5.1 Why not allow long interactions from the beginning
If you give the agent a large number of actions at an early stage when it doesn't yet know the basic rules of the environment, there's no guarantee that those extra computations will lead to useful exploration. You may be pressing the same button repeatedly, going back and forth between pages you've already viewed, or calling up tools that are unrelated to the task. This trajectory increases reward dispersion and credit allocation difficulty, making policy updates unstable.
Conversely, if the number of actions is fixed continuously, there are insufficient opportunities to solve stable but complex tasks. Agents may only adapt to easy paths to success and may not learn long-term behavioral patterns such as planning, reflection, and backtracking.
5.2 Gradual Horizon Expansion
ScalingInter-RL uses a curriculum that monotonically increases the maximum number of interactions as h1, h2, h3. In the early stages, a small interaction budget allows efficient use of basic actions, and after a certain learning period, the maximum number of turns is increased to allow for longer navigation paths.
1. Initial stage: Reliably learn easy tasks and basic tool use with a short horizon.
2. Intermediate stage: Increase the number of interactions to experience recovery after failure, reroute, and further retrieval.
3. Later stage: Enhances complex behaviors such as planning, reflection, and strategic backtracking in the long horizon.
4. The rollout of each stage is generated by the current policy, so the task difficulty and search space scale together as the agent's capabilities grow.
One line from a patent perspective. Curriculum learning in the broad sense is likely to be a known concept. In patent claims, rather than simply expressing difficulty, it is more advantageous to focus on a specific control structure that limits the maximum number of interactions in on-policy multi-turn rollout by stage and suppresses learning collapse by combining it with reward-based policy updates.
6. Training curves: longer interactions accelerate learning but can destabilize it
Illustration: AgentGym-RL Paper Analysis: Training Long-Horizon AI Agents with Multi-Turn Reinforcement Learning
Figure 4. Paper Figure 7: Comparison of rewards and actual number of interactions for fixed 10 turns, fixed 5 turns, and ScalingInter-RL in Deep Search. Source: Xi et al., p. 12.
In the graph on the left, a setting that allows up to 10 turns from the start gives initial rewards that rise quickly, but then quickly collapse after about 150 learning steps. The maximum 5-turn setting is relatively stable, but later performance stagnates. ScalingInter-RL initially maintains short interactions and gradually increases the number of turns, eventually reaching the highest reward.
The graph on the right shows the actual number of interactions used. ScalingInter-RL starts around 4 turns in the beginning and increases to more than 6 turns in the later stages of training. In other words, it is a structure that not only has a large final horizon, but sequentially grants freedom of exploration according to the agent's learning state.
Technical effects claimed by the paper. It reduces the high variance, accumulation of incorrect actions, false tool calls, and repetitive actions caused by early exploration of long horizons, while providing sufficient interaction budget to solve complex tasks in the later stages. The paper explains this as a balance between exploration and exploitation, training stability, and improved computational efficiency.
7. In what environment was the experiment conducted?
The experimental environment consists of five different types, ranging from web browsing to scientific procedures. Each environment has in common that the agent must act multiple times and receive feedback to choose the next action, but their ability to evaluate them differs.
WebArena is a web navigation scenario that evaluates your ability to achieve goals through button clicks, searches, and page navigation across Shopping, Forums, GitLab, Maps, and CMS.
Deep Search is a deep search scenario that repeatedly calls a search engine and evaluates its ability to combine information from multiple sources to answer a question.
TextCraft is a digital game scenario that assesses your ability to gather materials and perform multi-step crafts in a Minecraft-like text environment.
BabyAI is an embodied challenge scenario that assesses your ability to open doors, move around, and reach target objects in a grid world.
SciWorld is a science challenge scenario that assesses your ability to perform procedural experiments such as measuring temperature, testing conductivity, finding organisms, and manipulating materials.
The main backbones are Qwen2.5-3B and Qwen2.5-7B. Comparison targets include commercial models such as GPT-4o, OpenAI o3, and Gemini 2.5 Pro, as well as open source models such as Qwen, Llama, and DeepSeek. The evaluation figures are a result of the paper's chosen task set, maximum number of interactions, and sampling settings, and should not be directly converted to general product performance rankings.
Illustration: AgentGym-RL Paper Analysis: Training Long-Horizon AI Agents with Multi-Turn Reinforcement Learning
Figure 5. Paper Figure 6: Learning reward trends for WebArena, Deep Search, TextCraft, BabyAI, and SciWorld. Source: Xi et al., p. 10.
The learning reward curve shows an upward trend in all five environments, but in different shapes. TextCraft and BabyAI reach high rewards relatively quickly, while WebArena and SciWorld are highly volatile. This is in line with the interpretation of the paper that in a closed environment with clear rules and feedback, it is easier for reinforcement learning to find a success path, and that noise and action space become larger in realistic webs or complex scientific procedures.
8. Read key experiment results without exaggeration
In WebArena, Qwen2.5-7B recorded 9.76, AgentGym-RL-7B recorded 22.00, and ScalingInter-7B recorded 26.00. ScalingInter-7B outperformed the baseline model by +16.24 in absolute score.
In Deep Search, Qwen2.5-7B recorded 18.75, AgentGym-RL-7B recorded 34.00, and ScalingInter-7B recorded 38.25. ScalingInter-7B outperformed the baseline model by +19.50 in absolute score.
In TextCraft, Qwen2.5-7B recorded 42.00, AgentGym-RL-7B recorded 89.00, and ScalingInter-7B recorded 91.00. ScalingInter-7B outperformed the baseline model by +49.00 in absolute score.
In BabyAI, Qwen2.5-7B recorded 66.67, AgentGym-RL-7B recorded 92.22, and ScalingInter-7B recorded 96.67. ScalingInter-7B outperformed the baseline model by +30.00 in absolute score.
In SciWorld, Qwen2.5-7B recorded 1.50, AgentGym-RL-7B recorded 50.50, and ScalingInter-7B recorded 57.00. ScalingInter-7B outperformed the baseline model by +55.50 in absolute score.
The standard model summarized above is Qwen2.5-7B-Instruct described in the paper. The improvement presented in each environment is the absolute score difference between ScalingInter-7B and the baseline model. The biggest improvements were seen in SciWorld and TextCraft, both of which have relatively clear rules and success conditions, with strong signals of learning through trial and error.
8.1 WebArena: Competitive with commercial models, but not the best
ScalingInter-7B recorded 26.00 overall in WebArena, exceeding GPT-4o's 16.00 and coming close to Gemini 2.5 Pro's 28.00. However, it fell short of OpenAI o3's 34.00 and o4-mini's 36.00. Gaps remained, especially in the GitLab and Reddit sub-tasks. Therefore, the statement that “the 7B model beat all commercial models in WebArena” is not accurate.
8.2 Deep Search: Strong, but not as good as o3
ScalingInter-7B was 38.25, exceeding GPT-4o 26.75 and Gemini 2.5 Pro 36.50, but lower than OpenAI o3 49.50 and o4-mini 42.50. It recorded the highest score in NQ with 52.00 and was tied for the best score in TriviaQA with 70.00, but the overall average left a gap with the top reasoning model.
8.3 TextCraft and BabyAI: The Strength of Long-Term Behavioral Learning
In TextCraft, ScalingInter-7B surpassed GPT-4o's 83.00 with 91.00, and was among the few models that recorded 33.33 even in the most difficult Depth 4. BabyAI recorded the highest overall score of 96.67 among all models compared in the paper. This shows that escalation of interaction can work effectively in environments with clear behavioral rules and rewards for success.
8.4 SciWorld: The greatest improvements and most obvious limitations appear simultaneously
In SciWorld, ScalingInter-7B was at 57.00, significantly exceeding OpenAI o3's 41.50. This is a result of an increase of 55.50 points from 1.50 based on Qwen2.5-7B. However, all models scored 0 in the chemical mixing task. The ability to accurately perform scientific procedures and experimentally diagnose the causes of failure remains unresolved.
Things to keep in mind when reading numbers. These results are performance for the paper's specific evaluation sample, tool interface, maximum number of turns, and number of samplings. It is not a figure that directly proves the model's general knowledge, safety, actual web service compatibility, or transferability to new environments.
9. Which reinforcement learning algorithm worked better?
In the 3B model, TextCraft recorded GRPO 75.00 and REINFORCE++ 28.00, BabyAI recorded GRPO 93.33 and REINFORCE++ 70.00, and Deep Search recorded GRPO 25.75 and REINFORCE++ 13.25.
In the 7B model, TextCraft recorded GRPO 83.00 and REINFORCE++ 73.00, BabyAI recorded GRPO 92.22 and REINFORCE++ 84.44, and Deep Search recorded GRPO 34.00 and REINFORCE++ 24.00.
In the paper, GRPO scores higher than REINFORCE++ on all three tasks and both model scales. The researchers interpret that using only full episode returns in long-term trajectories and sparse rewards can result in large variance, and that GRPO, which uses relative performance within groups, provides a more stable learning signal.
However, this comparison is limited to the hyperparameters and environment selected by the paper. Generalizing algorithmic superiority or inferiority requires further verification on the same computational amount, the same number of samples, and different reward densities and model families.
10. Does additional test-time computation improve performance?
Illustration: AgentGym-RL Paper Analysis: Training Long-Horizon AI Agents with Multi-Turn Reinforcement Learning
Figure 6. Paper Figure 8: Change in accuracy when increasing the maximum number of interactions in Deep Search and SciWorld. Source: Xi et al., p. 17.
Accuracy generally increases as the number of interactions increases from 2 to 10 turns in Deep Search, and from 10 to 30 turns in SciWorld. The learned agent outperforms the baseline model even with fewer turns and maintains its superiority as the number of turns increases. This shows that the agent's test-time computation can scale not only with internal reasoning tokens, but also with additional interactions with the real-world environment.
Illustration: AgentGym-RL Paper Analysis: Training Long-Horizon AI Agents with Multi-Turn Reinforcement Learning
Figure 7. Paper Figure 9: Pass@K performance change when increasing the number of parallel samples K. Source: Xi et al., p. 17.
Pass@K, which generates multiple trajectories in parallel and checks whether any one succeeds, also increases as the number of samples increases. The paper reports that in 64 samples, the reinforcement learning model outperformed the baseline model by 5.5 points in Deep Search and 7.05 points in SciWorld. However, this method can significantly increase costs and delays in actual service, so product design requires a balance between success rate and computational budget.
11. Case Analysis: What has Reinforcement Learning Changed?
Illustration: AgentGym-RL Paper Analysis: Training Long-Horizon AI Agents with Multi-Turn Reinforcement Learning
Figure 8. Paper Figure 10: Comparison of action trajectories of basic model and RL model in WebArena. Source: Xi et al., p. 19.
The basic model failed to take advantage of the fact that, in an assignment to subscribe to popular posts on a Pittsburgh forum, the screen did not change despite repeated clicks on the same text element. On the other hand, the reinforcement learning model was successful by navigating to the wrong page and then taking actions in the following order: going back, searching the forum, selecting the search result, and confirming the subscribe button.
The case illustrates an improved action policy that uses environmental feedback to detect failure, change course and verify completion. From a patent drafting perspective, error recovery, suppression of repeated actions and detection of task completion may be described as dependent-claim features or elements of reward design.
12. What the paper actually proved and what it did not prove
Scope directly proven by the paper. A multi-turn online reinforcement learning pipeline can be operated in five types of digital environments. The method of expanding from a short horizon to a long horizon shows a more stable learning curve than the fixed turn setting. The 3B and 7B models significantly improve over the baseline model across multiple agent benchmarks. Test performance can increase as you increase the number of interactions and the number of parallel samples.
Scope not yet fully proven. Generalization and transfer to completely new environments, tools, and reward structures, real-world physical robots or real services with long working hours and safety constraints, multi-agent collaboration/competition, and joint decision-making with people require additional verification. Additionally, the conclusion that it is consistently superior across all commercial models and under all evaluation conditions is not supported by this paper alone.
A point that needs to be read critically. Although the performance comparisons in the paper are very impressive, the number of evaluation samples and interfaces for each environment are different, and some web assignments exclude state change operations. Additionally, the conclusion of “outperforming commercial models” holds strong in certain environments, such as BabyAI and SciWorld, but falls short of the best commercial models in WebArena and Deep Search.
13. Points of invention from the perspective of a patent attorney specializing in artificial intelligence patents
The following is an interpretation from a patent practice perspective derived from the published content of the paper. The actual novelty, inventive step, scope of rights, and possibility of infringement must be determined through country-specific legal principles and a separate search of patent and non-patent literature.
13.1 Dividing the invention concept into five layers
The invention concept can be divided into five layers. The core algorithm is a multi-turn reinforcement learning method that gradually increases the maximum number of agent-environment interactions as learning progresses. The system structure is premised on a modular framework that separates Environment, Agent, and Training and connects them through a standardized server-client interface, and the data flow can be described as an on-policy learning pipeline that collects trajectories in a parallel environment and updates the policy by calculating rewards and log probabilities.
The core of the stabilization technology is control to reduce learning collapse by limiting initial variance and unproductive exploration with a short horizon and then expanding to a long horizon, and the operation technology can be summarized as a method of managing the success rate and computational budget by adjusting the number of interactions and the number of parallel samples at the time of testing.
13.2 ScalingInter-RL as a candidate for patent protection
While AgentGym-RL's modular structure is practical, the concept itself of separating environments, agents, and learners and connecting them over HTTP is likely to be widely used in software frameworks. On the other hand, the structure that increases the maximum number of interactions in conjunction with the reinforcement training stage and suppresses the initial learning collapse of the long fixed horizon is the clearest experimental point of the paper.
However, if you claim the broad curriculum learning concept of “progressing from easy tasks to difficult tasks”, there is a high possibility that it will conflict with prior art. The claims must disclose a combination of a specific control variable called interaction horizon, on-policy trajectory generation, terminal reward or trajectory reward, policy update, step transition condition, and maximum turn increase.
14. Implications from a business and product perspective
AgentGym-RL's direct applications are products that require multiple actions and environmental feedback, such as web task automation, search agents, game agents, virtual robots, and scientific experiment assistants. Especially in services where the number of actions required for success varies depending on the user and task, it is important to design a design that adjusts the interaction budget based on difficulty and reliability rather than fixing it.
During the production phase, you need to optimize not only the highest success rate, but also environment call cost, number of browser sessions, latency, repeat action rate, error recovery rate, and safe shutdown conditions. The paper's test-time scaling results show that more turns and samples can improve performance, but real services require dynamic policies that reflect computational budget constraints.
Practical implications. The competitiveness of an AI agent is not determined solely by the number of model parameters. How it interacts with the environment, the reward design that turns failures into learning signals, and the operating policies that determine when to increase its action budget can be just as important intellectual property points as model size.
15. Limitations of the paper and follow-up research questions
1. Strong in-domain performance does not establish generalization to entirely new environments or tools.
2. Most challenges remain in digital simulation, and long-term challenges remain, including real-world physical constraints and sensor noise.
3. The current framework focuses on a single agent, and multi-agent collaboration and competition are future research topics.
4. All evaluated models failed SciWorld’s chemical-mixing task. Some agents attempted to answer from memorized facts instead of carrying out the required experiments.
5. In WebArena, the problem of excessive interactivity remained, with unnecessary clicking and scrolling even after reaching the target page.
6. The paper presents a plan to disclose code and data, but actual reproducibility requires separate confirmation of the scope of disclosure and execution environment.
7. Interaction horizon transition rules need to be more automated and theoretically justified.
16. Conclusion
AgentGym-RL treats LLM agent reinforcement learning as an integrated system problem that connects various environments, rather than as code for individual benchmarks. The most noteworthy part is ScalingInter-RL. In the beginning, when the agent has not yet acquired the basic skills, freedom of action is limited, and as the policy becomes more stable, the length of interaction is increased to improve long-term decision-making ability.
The experimental results show that even the small 7B model can compete with much larger models with appropriate subsequent reinforcement learning and test-time interactions. At the same time, the gap between WebArena and Deep Search, scientific procedure errors, and excessive interaction clearly reveal that agent intelligence is not yet complete.
From a patent perspective, the key is not the broad idea of “applying reinforcement learning to AI agents,” but the specific combination of solving the technical problems of high variance and training collapse in long-term rollouts with stepwise control of the interaction horizon. In the future, automatically adjusting the horizon based on reward statistics and policy stability and optimizing the cost at the time of testing may become a stronger technical differentiation.
One-line conclusion: Strong AI agents are not built to behave for a long time from the beginning, but can be trained to learn short successes and then explore longer.

Read the Korean source

This article reflects the information available when it was published. Contact us to discuss your circumstances.
Discuss this topic ↗All articles

Put your IP strategy into practice.