AI Times has published Yongduck Kim's patent analysis of OpenAI's graphical user interface interaction and automation technology.
This analysis is about a multimodal-based interaction technology that can understand the user's text and image input at the same time, identify and visually highlight specific elements in an image, or perform real operations on a GUI. With the US patent US 12,051,205 B1 registered by OpenAI, the language model describes a technical structure that directly manipulates and responds to the graphical environment as if it were a user.
In this article, Yongduck Kim patent attorney explained in technical detail the evolution of AI systems that can perform visual and behavioral interactions such as mouse clicks, highlighting, and image editing on the actual GUI. Specifically, we analyzed how this technology can be applied in a variety of industries, including accessibility improvements, image editing and training content, and virtual assistants.
You can find the full article at the link below.
Read the column
This article reflects the information available when it was published. Contact us to discuss your circumstances.
