NEWS & INSIGHTS

Understanding the Paper “Agentic Reasoning for Large Language Models”

The term AI agent is used widely, but its meaning varies. It may describe a chatbot with search capabilities or a system that plans tasks and operates external tools. Recent research explores these distinctions...

Illustration: Understanding the Paper “Agentic Reasoning for Large Language Models”
Original text of thesis: https://arxiv.org/pdf/2601.12538
Author: IPLEX IP Law Firm Yongduck Kim, Patent Attorney
One of the most frequently heard terms in the field of artificial intelligence these days is AI agent. But when you try to explain what an agent is, the scope is very broad. Some people call a chatbot with a search function an agent, while others see only a system that makes plans and executes various tools as an agent. In addition, recently, mechanisms for updating memories, learning from experience, and sharing roles among multiple agents have emerged.
To clear up this confusion, this paper presents agentic reasoning as an integrated perspective. The key is that the language model does not stop at generating text once, but views the entire process of planning, acting, and learning as inference while continuously interacting with the environment. Therefore, rather than a study proposing a new single model, this paper is closer to a roadmap that systematically classifies the vast amount of research conducted until 2025.
Three points to remember: Agentic reasoning connects reasoning with action. The paper distinguishes foundational, self-evolving and collective reasoning. These capabilities can be implemented through prompts and workflows at inference time or learned through fine-tuning and reinforcement learning.
Paper basic information
The title of the paper is “Agentic Reasoning for Large Language Models,” and it is a large-scale survey paper summarizing the field of agentic reasoning. The authors are Tianxin Wei, Ting-Wei Li, Zhining Liu, and many others, and the University of Illinois Urbana-Champaign, Meta, Amazon, Google DeepMind, UC San Diego, and Yale participated. Released as arXiv v1 on January 18, 2026, we extensively review planning, tooling, search, feedback, memory, self-evolution, multi-agent, applications, benchmarks, and future challenges. However, rather than a paper that proves a new algorithm through experiments, it is closer to a map that synthesizes the flow of existing research and design options.
1. The overall picture presented by the paper
The first figure in the paper summarizes the entire survey in one page. At the top is a flow where the user proposes a task, the agent solves the problem, and the results are generalized to other tasks. In the middle, we see that as traditional LLM inference moves to agentic reasoning, the input changes from static to dynamic context, the computation changes from passive generation to interaction, and the learning changes from a fixed pre-trained state to a continuously changing state.
Illustration: Understanding the Paper “Agentic Reasoning for Large Language Models”
Figure 1. Overall map of agentic reasoning.
Source: Tianxin Wei et al., Agentic Reasoning for Large Language Models, Figure 1, p. 2.
The bottom of the picture is divided into four major areas. The top left is agent-foundational reasoning, which consists of planning, tooling, and web searching. Top right is self-evolving agentic reasoning with feedback, memory, and self-evolution. Bottom left is collective multi-agent reasoning, covering role assignment, collaboration, and co-evolution. The bottom right shows how these features are used and evaluated in fields such as healthcare, finance, law, education, robotics, science, software, games, and the web.
To put it simply, if a general chatbot is a counselor who answers a question immediately, an agent-type system is closer to a project manager who checks the goal, divides tasks, finds necessary data, executes the tool, and plans again if the results are wrong.
The important thing is not to view agents narrowly as simply chatbots that invoke tools. In this paper, inference is not limited to the process of generating internal sentences. What to observe, what to decide on next action, how to interpret failure, what experiences to remember, and what information to exchange with other agents are all seen as part of reasoning.
2. Differences between general LLM reasoning and agentic reasoning
General language model inference involves receiving a set prompt and providing an answer in one or several generation processes. Even if you use chains of thought or techniques to compare multiple candidates, if your model doesn't actually affect the external environment, it's still a relatively closed computation.
In agentic reasoning, the model's output becomes an action that changes the next input. When you send a search term, search results are returned, when you run code, an error message is returned, and when you click on a web page, the screen state changes. The agent takes these new observations, modifies its plan, and acts again. This makes state management and recovery capabilities across multiple stages more important than the quality of a single response.
General LLM inference and agentic reasoning differ in input, computation method, connection to the outside world, state management, error handling, and goals. General LLM inference typically takes a fixed prompt as input and performs a one-time creation or limited internal exploration, has no direct access to the external world, and manages state around the current conversation context. If an error occurs, we respond by regenerating the answer, with the goal being close to generating a plausible response.
On the other hand, agentic reasoning takes as input a dynamic context that continuously changes according to the results of actions, and performs a multi-step process in which thinking, action, observation, and modification are repeated. It is connected to searches, APIs, code, browsers, robots, etc., and manages external memory, task status, and environmental status together, and when errors occur, it verifies, reverts, replans, and recovers by selecting another tool. The ultimate goal is to achieve a stated goal state within the environment.
So agentic reasoning is not synonymous with chain-of-thought reasoning. Chains of thought may be a way to lengthen internal calculations, but agentic reasoning connects internal thoughts to external actions, forming a closed loop that reflects the results of those actions into the next judgment. Having a long plan doesn't automatically make you a good agent. The ability to quickly detect misbehavior and get back to the right spot may be more important.

3. Foundational agentic reasoning: planning, tool use and search
Foundational agentic reasoning addresses the basic capabilities that a single agent needs to achieve a goal in a complex but relatively stable environment. The paper organizes this into three axes: planning, tool use, and search. Although these three elements may seem like independent functions, in real systems they are strongly interconnected. Planning determines what information is missing, search fills in that information, and tooling performs the actual action.
3.1 Planning: The ability to turn goals into actionable steps
A plan is not just a to-do list. A good plan looks at the current state, selects next steps, evaluates intermediate results, and changes the plan itself when necessary. The paper divides the planning approach into workflow design, tree exploration, formalization, decomposition, leveraging external tools, and reward design.
Illustration: Understanding the Paper “Agentic Reasoning for Large Language Models”
Figure 2. Schematic classification of LLM agent planning methods.
Source: Same paper, Figure 2, p. 11.
  • - Workflow design: Divide steps in advance, such as recognition, inference, execution, and verification, and determine what each step should do.
  • - Tree navigation: Spread out multiple options like branches and explore them in depth or breadth while evaluating promising paths.
  • - Formalization: Represent plans as code, state machines, graphs or logical expressions, rather than natural language alone.
  • - Task decomposition: Divide large goals into distinct sub-goals and determine dependencies and execution order.
  • - Leverage external tools: Incorporate knowledge graphs, searches, world models, calculators, code executors, etc. into your planning process.
  • - Post-training: By assigning rewards to desired actions, the planning habit itself is learned to be internalized in the model.
When building a product, rules for handling failure are more important than the grandeur of a plan. For example, if an error occurs just before payment, it should be clear whether you want to start over, return to the last successful state, get user confirmation, or try a different payment method. These state transitions and recovery procedures determine the actual agent quality.
3.2 Tool use: Ability to complement model limitations with external features
Language models have inherent limitations in areas such as up-to-date information, accurate calculations, in-house data, and real-world web operations. Tooling compensates for this limitation by connecting search APIs, databases, calculators, code executors, and business systems. However, as the number of tools increases, new reasoning problems arise such as which tool to choose when, how to call it in the correct format, and how to interpret the results of failure.
Illustration: Understanding the Paper “Agentic Reasoning for Large Language Models”
Figure 3. Difference between existing LLM and agent-type tool use system
Source: Same paper, Figure 3, p. 15.
The traditional LLM on the left lacks external tools, which can limit access to current information and reliable calculation. The agentic system on the right selects a tool, checks the result and chooses another action if needed. The key is the cycle of selection, invocation, observation and reassessment.
Common failures in tool use include creating a non-existent tool name, filling out arguments incorrectly, misinterpreting the tool's error message as success, or repeating the same failed call. Therefore, tool schema validation, pre-invocation checking, execution time limits, number of retries, permission control, and indication of the origin of results must be designed together.
3.3 Search: Ability to find necessary information on your own
In a traditional RAG, when a question comes in, it searches once in a predetermined manner and then answers based on that document. Agent-based search determines at each step when to search, what to search for, whether the results are sufficient, and whether additional search is necessary. That is, the search comes into the inference process rather than the preprocessing stage.
Illustration: Understanding the Paper “Agentic Reasoning for Large Language Models”
Figure 4. Comparison of static RAG and agent-based search systems.
Source: Same paper, Figure 4, p. 18. I have excerpted the schematic part necessary for explanation.
For example, an agent performing patent prior art search should be able to add classification codes or applicants when there are too many results in the first search, and conversely, expand synonyms or higher-level concepts when there are few results. You should check for yourself whether the literature you find shows all the key components. This is different from simply summarizing a few search results.
The paper distinguishes inference-time approaches that guide agentic search through prompts, approaches that train search behavior through fine-tuning or reinforcement learning, and approaches using structured sources such as knowledge graphs. These methods can be combined according to system requirements.
4. Self-evolving agentic reasoning: feedback, memory, and ability expansion
While a base agent focuses on performing one task well, a self-evolving agent aims to use its experience to become better at its next action or next task. I believe that the thesis has feedback and memory at its center. Feedback tells us what went wrong, and memory lets us rewrite those lessons into our next decisions.
4.1 Feedback: How to discover and correct incorrect paths
Illustration: Understanding the Paper “Agentic Reasoning for Large Language Models”
Figure 5. Three forms of agent feedback
Source: Same paper, Figure 5, p. 21.
The paper divides feedback into three types. First, reflection during inference is a method of criticizing and then rewriting the current answer or plan without changing model parameters. Second, parameter adaptation is a method of learning from examined trajectories or preference data and leaving improved behavior in the model weights. Third, verifier-based feedback is where an external verifier, such as a unit test, constraint checker, or simulator, notifies success or failure and lets you try again if it fails.
The difference between the three methods lies in where the feedback is left. Reflection often only has impact within the current conversation, while parameter adaptation leads to long-term behavioral change. Verifier-based methods are strong at picking out passable results, even without detailing the reasons for failure. This is especially useful in areas where verification is cheap and obvious, such as code generation.
4.2 Memory: What you leave behind is more important than accumulating conversation records
Memory can easily be misunderstood as the ability to store long conversations verbatim. However, the key to agentic memory is not storage, but selection and structure. You need to decide what to record, how to summarize it, when to take it out, and when to delete old information.
Illustration: Understanding the Paper “Agentic Reasoning for Large Language Models”
Figure 6. Three design axes of agentic memory.
Source: Same paper, Figure 6, p. 24. I have excerpted the schematic part necessary for explanation.
The first axis is how conversations, summaries, workflows, and past execution trajectories are leveraged in the current prompt. The second axis is structured memory, which stores entity relationships as graphs or handles images, audio, and documents together. The third axis is a method that makes memory writing, updating, deleting, and retrieval itself the object of learning.
The main danger of memory: If incorrect facts enter long-term memory, they can repeatedly affect many subsequent tasks. If user-specific information is mixed, privacy violations or rights leaks may occur. Therefore, it is important to keep the source, creation time, reliability, access rights, expiration conditions, and modification history.
4.3 Self-evolution: the ability to change plans, tools and the search itself
Illustration: Understanding the Paper “Agentic Reasoning for Large Language Models”
Figure 7. How planning, tooling, and search evolve on their own.
Source: Same paper, Figure 7, p. 28. I have excerpted the schematic part necessary for explanation.
Self-evolution is broader than simply revising an answer once. Agents can create their own exercises, change planning strategies by comparing success and failure, and even create new code tools when they encounter problems that are difficult to solve with existing tools. In searches, you can remember frequently used information sources and improve your search path.
Self-evolution should be interpreted within the limits of the research. Many studies update memory or policies using defined feedback in constrained environments; they do not establish that agents choose goals and values as humans do. System designers must specify which components may change, who authorizes changes and how erroneous changes can be reversed.
5. Collective agentic reasoning: roles, collaboration, and collective memory.
In complex tasks, a structure where multiple agents divide roles rather than one agent taking on all tasks is used. For example, one agent might create a plan, another agent might search, and another agent might verify the results. The paper summarizes this structure as collective multi-agent reasoning.
5.1 Role Design: Giving everyone the same prompts is not collaboration
Illustration: Understanding the Paper “Agentic Reasoning for Large Language Models”
Figure 8. General roles and domain-specific applications of multi-agent
Source: Same paper, Figure 8, p. 30.
The representative roles presented in the paper are leader or coordinator, executor, critic or evaluator, memory manager, and communication coordinator. Leaders divide goals and manage progress. Executors do the actual work of searching, running code, and creating documentation. Evaluators look for errors and risks. The memory manager reduces duplicate failures and organizes long-term knowledge. The communication coordinator controls the format and scope of the message.
For example, in legal work, a receiving agent organizes facts, a field-specific analysis agent reviews laws and precedents, and a verification agent checks citations and logical consistency. However, naming specialized roles does not automatically increase accuracy. For each role, the format of input and output, scope of responsibility, stopping conditions, and cross-validation methods must be specifically defined.
5.2 Collaboration structure: sequential, hierarchical, role, automatically generated
Illustration: Understanding the Paper “Agentic Reasoning for Large Language Models”
Figure 9. Inference-time collaboration and collaboration learned through post-training
Source: Same paper, Figure 9, p. 34.
Inference-time collaboration determines the order and roles between agents through prompts or workflows. In the sequential type, the next step inherits the results of the previous step, and in the hierarchical type, a central coordinator manages the subagents. In the role type, specialized functions are divided in advance, and in the automatically generated type, LLM configures the workflow according to the assignment.
Post-training-based collaboration optimizes the prompts for each role, the connection graph between agents, and the policy for selecting the agent to call next using data or rewards. The most difficult problem here is deciding whose contribution is to be attributed to the joint performance. Even if the final answer is correct, it is difficult to isolate how much the search or counterargument in the previous step contributed.
5.3 Joint memory: What will multiple agents share?
Illustration: Understanding the Paper “Agentic Reasoning for Large Language Models”
Figure 10. Four design dimensions of multi-agent memory.
Source: Same paper, Figure 10, p. 40. I have excerpted the schematic part necessary for explanation.
Memory design becomes more complex when multiple agents collaborate. The paper describes this in four dimensions: structure, topology, content, and management. The structure is hierarchical or flat, the topology is a matter of central or distributed storage, the content is fact-oriented or procedure-oriented, and management is a question of how to perform summarization and forgetting or filtering and verification.
The more we share, the richer the information, but also the greater the cost of communication and the risk of contamination. Conversely, too much separation of information by role may not provide needed context. In particular, in a corporate environment, shared memory and private memory must be distinguished considering personal information, trade secrets, and departmental access rights, and the source and viewing history of each information must be recorded.
6. Inference-time design and post-training
The paper distinguishes approaches by when an agent's capabilities are implemented: at inference time or through post-training. These are complementary design choices whose combination depends on cost, flexibility and reliability.
The inference-time approach coordinates behavior with prompts, searches, workflows, and external memory while keeping model weights fixed. It has the advantage of being quick to change and easy to apply to existing models, but is subject to prompt length, call cost, and execution instability. Therefore, it is suitable for products whose business rules change frequently or are in an experimental stage. Typical designs include planning prompts, search loops, memory searches, verification and retry, etc.
On the other hand, post-training methods internalize behavioral patterns into model weights through reinforcement learning, supervised fine-tuning, and preference learning. You can expect consistent and fast behavior in repetitive tasks, but they require training data and reward design, and update costs are high. It is suitable for large-scale repetitive tasks or agents optimized for specific environments, and representative designs include fine-tuning of tool calls, reinforcement learning of search policies, and memory control learning.
A practical approach is to validate the workflow at inference time, then turn repeated successful interactions into training data for post-training. Training every capability from scratch can make later corrections costly, while relying entirely on prompts can increase token costs and latency.
An example of a good mixed design is where tool default selection and call formats are stabilized through learning, up-to-date business rules and user-specific conditions are managed through external memory and policies, and high-risk actions can be subject to verifiers and human approval.
7. Real-world applications and benchmarks
7.1 Where is it applied?
Illustration: Understanding the Paper “Agentic Reasoning for Large Language Models”
Figure 11. Representative application areas of agentic reasoning
Source: Same paper, Figure 11, p. 43. I have excerpted the schematic part necessary for explanation.
Papers cover representative applications of mathematical inquiry and coding, scientific discovery, robots and embodied agents, healthcare, web navigation, and autonomous research. Even with the same technology, important constraints vary from field to field. Execution and testing are important for code agents, evidence and responsible review are important for medical agents, and safe physical behavior and real-time reactions are important for robotic agents.
  • - Math and Coding: Problem decomposition, code execution, testing, error correction, and long-term work at the repository level are key.
  • - Scientific discovery: Generate a hypothesis, plan an experiment, run the tool, evaluate the results, and then design the experiment and iterate.
  • - Robots and embodied agents: Must connect language goals with spatial awareness and actual actions, and respond to changing environments.
  • - Medical: Collaboration between professional agents, searching for evidence, remembering patient conditions, and assisting safe decision-making are important.
  • - Web exploration and independent research: You must go back and forth between multiple sites and tools to verify information and synthesize it into a long-form result.
7.2 What to evaluate
Illustration: Understanding the Paper “Agentic Reasoning for Large Language Models”
Figure 12. Scope of agentic reasoning benchmarks
Source: Same paper, Figure 12, p. 65.
When evaluating an agent, looking only at the final percent correct can miss important failures. Even if you get the same correct answer, the value of the product will vary depending on whether you unnecessarily called the tool dozens of times, left incorrect information in memory, or attempted a risky action. The paper brings together core capability benchmarks such as tool use, memory and planning, and multi-agent, and application benchmarks such as robotics, science, medicine, and the web.
In actual service evaluation, in addition to task success rate, it is better to measure step-by-step accuracy, tool call success rate, number of retries, average time spent, token and API cost, error recovery rate, memory corruption rate, evidence citation accuracy, number of human interventions, and risky behavior blocking rate. Because an agent is a system with multiple components connected, it is difficult to tell which part is problematic based on the final score alone.
8. Perspective of a patent attorney specializing in artificial intelligence patents
The first thing to distinguish when reading this paper from a patent perspective is that the broad concept of agentic reasoning is different from the specific implementation invention. The components of planning, tooling, retrieval, memory, and collaboration have already been addressed in many prior studies and open source frameworks. Therefore, it is difficult to differentiate simply by stating that it includes all of these elements.
Conversely, sufficient intellectual property value can be created in a structure that solves specific bottlenecks encountered while making an actual product. For example, the clearer the technical problem, such as repeated tool errors, incorrect information accumulating in long-term memory, rapidly increasing multi-agent communication costs, user-specific information being exposed to other agents, and intermediate state loss in long tasks, the clearer the direction for securing rights becomes.
8.1 Notable Technology Areas
There are six major technology areas to watch: Planning and state control deal with the problem of error accumulation and high restart costs in long tasks, and require a specific look at state representation, step transition conditions, intermediate checkpoints, revert ranges, and plan reconstruction rules. In tool orchestration, tool selection errors, invalid arguments, repeated calls, and permission misuse are issues, and tool schema normalization, pre-invocation verification, dynamic routing, recovery by error code, permissions, and sandbox are key implementation elements.
In agent-based search, unnecessary searches, lack of evidence, and processing of conflicting literature must be resolved, so determining search necessity, rewriting queries, assessing source reliability, information sufficiency, and search termination conditions are important. Memory control must reduce misremembering, redundant storage, mixing of personal information, and retrieval delays, and key implementation elements include record acceptance, summarization and structuring, reliability and expiration, deletion and correction, access rights, and provenance tracking.
In multi-agent collaboration, role assignment, communication topology, message compression, joint rewards, disagreement reconciliation, and contribution evaluation must be designed to reduce communication costs, role duplication, unclear responsibilities, and collective errors. Lastly, in governance and auditing, accountability tracking and risk control of long-term actions are key, and action logs, policy inspection, approval steps, permissions by risk level, disruption and recovery, and reproducible audit records are important.
8.2 Differentiation lies in the operating conditions rather than the names of the components
Just listing the names agent, planner, memory, and verifier can easily end up being a functional description. The key to differentiation is showing what decision is made at what point based on what input and state, what data structure is updated, and what recovery path is taken in case of failure.
Specify the state representation: explain how the system tracks goals, completed steps, unresolved conditions, tool results, confidence and permissions.
Refine the trigger. It is a good idea to present in numbers or rules the conditions under which further searches, replanning, human approval, and memory updating are triggered.
Refine your data flow. It must be expressed which module receives what information and in what format it passes it to the next module.
Measure technical effectiveness. You need reproducible metrics like call count, latency, memory usage, error recovery rate, and risky behavior blocking rate.
Describe alternative implementations, such as centralized or distributed architectures, rule-based or learned controls, and single-agent or multi-agent systems, to address a changing market.
Conclusion: From language models to action systems
The main trend shown by this paper is clear. The direction of LLM development is not only about creating longer answers. We are moving toward a system that understands goals, makes plans, selects the tools and information needed, remembers the consequences of its actions, recovers from failures, and collaborates with other agents when necessary.
At the same time, as agents become more powerful, their responsibility for system design also increases. This is because a chain of wrong actions can have a greater impact than a single wrong sentence. Therefore, state management, validation, permissions, auditing, memory hygiene, interruption and recovery are as important as performance.
The same conclusion is reached from a patent perspective. Rather than following broad buzzwords, we need to find recurring technical problems in real products and organize specific data flows and control rules that solve those problems. Competitiveness in the agent era will likely come not just from the name of the model, but from the system architecture that allows the model to behave safely and efficiently.

Read the Korean source

This article reflects the information available when it was published. Contact us to discuss your circumstances.
Discuss this topic ↗All articles

Put your IP strategy into practice.