# Preparing Chatbot Training Data: Natural Language Processing Patent Analysis

Shortages of customer-service staff and infrastructure can make it difficult for businesses to meet customer expectations. Chatbots are being introduced to supplement these resources and improve service quality.

Source: https://www.iplexlaw.co.kr/en/blog/864755

HOME / NEWS & INSIGHTS NEWS & INSIGHTS Preparing Chatbot Training Data: Natural Language Processing Patent Analysis Shortages of customer-service staff and infrastructure can make it difficult for businesses to meet customer expectations. Chatbots are being introduced to supplement these resources and improve service quality. AI & Software 2023.06.12 published IPLEX 3 min read Shortages of customer-service staff and infrastructure can make it difficult for businesses to meet customer expectations. Chatbots are being introduced to supplement these resources and improve service quality. Businesses use chatbots to provide around-the-clock customer support and reduce repetitive manual work. However, these chatbots are not simply a service that can be implemented by planting a few lines of code or by providing an API. The introduction of chatbots by certain companies requires processing customer support data, not just service integration, into the data needed to learn the AI module of chatbots. In this post, we will explain a patent related to how to process the training data required for learning the AI module of chatbots. [Patent] Click the image to view the granted patent. This patent is related to a method for augmenting the training data of the chatbot module. Training a chatbot generally requires sets of questions paired with answers. There may be too few examples to meet the intended training needs—the same data-availability problem encountered in supervised learning. Thus, for data augmentation, real-world question data is analyzed to produce similar question statements with the same meaning and used as training data for chatbots. Specifically, if you already have three questions and three answers to increase the amount of data, the patent states that if you have a question set that uses natural language processing to understand the meaning of the sentence, you can generate another tone of questions with the same intent and generate a data set with multiple questions and one answer. If you add data to the pattern, the same effect will occur. First, even if there is very little question-answer data for that domain when constructing a chatbot, the amount of data that can sufficiently run the chatbot's AI module can be generated specifically for the domain. And because the data provided by the client based on the word and the similarity of the phrase can be generated into a cluster of multiple sentences of the same meaning, if the client who wants to develop a chatbot is a startup, it can solve the problem that the amount of data they have is very low. Furthermore, very few data sets can produce meaningful results and even a small amount of testing can maximize the results. This minimizes preprocessing and eliminates the need for customers to label each new question, reducing the workload of both the chatbot company and the customer. Read the Korean source This article reflects the information available when it was published. Contact us to discuss your circumstances. Discuss this topic ↗ All articles TALK TO IPLEX Discuss your IP questions We consider your technology and business needs together. ↗ Contact us Newer Preventing Overfitting in Medical Imaging AI: VUNO Patent Analysis ↗ Older Understanding the Limitations of GAN Models ↗ Related insights AI & Software 2026.10.01 FiX: fine-grained forgetting in softmax attention Yongduck Kim examines FiX’s feature-wise gates, numerical implementation and paged cache, distinguishing reported gains from unresolved limitations. ↗ Read article AI & Software 2026.09.30 MHAR: Reading earlier layers through different feature subspaces Yongduck Kim examines Multi-Head Attention Residuals: depth routing, reported training results, implementation costs and the relationship between technical features and effects. ↗ Read article AI & Software 2026.09.28 Column: Claude Computer Use and the Data That Trains AI Agents Writing for AI Times, IPLEX Managing Partner Yongduck Kim examines the training data behind computer-operating AI agents through U.S. Patent No. 12,585,862. ↗ Read article

- https://www.iplexlaw.co.kr/en
- https://www.iplexlaw.co.kr/en/blog/category/ai
- https://drive.google.com/file/d/17ZPRnXK3pvEbzayiO-5pchD7QIJya8mK/view
- https://www.iplexlaw.co.kr/forum/view/864755
- https://www.iplexlaw.co.kr/en/contact
- https://www.iplexlaw.co.kr/en/blog
- https://www.iplexlaw.co.kr/en/contact
- https://www.iplexlaw.co.kr/en/blog/865384
- https://www.iplexlaw.co.kr/en/blog/862036
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention
- https://www.iplexlaw.co.kr/en/blog/fix-fine-grained-forgetting-attention
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals
- https://www.iplexlaw.co.kr/en/blog/mhar-multi-head-attention-residuals
- https://www.iplexlaw.co.kr/en/blog/claude-computer-use-agent-patent
- https://www.iplexlaw.co.kr/en/blog/claude-computer-use-agent-patent
