KAIST Develops Technology to Identify AI Vulnerabilities Seven Times More Effectively

By Na Seon Hye Posted : July 30, 2026, 15:24 Updated : July 30, 2026, 15:24

The Korea Advanced Institute of Science and Technology (KAIST) has developed a safety verification technology that identifies hidden vulnerabilities in artificial intelligence (AI) systems with approximately seven times more diversity than existing methods.


On July 30, KAIST announced that a research team led by Professor Kim Jun-mo of the Department of Electrical and Electronic Engineering has created a new framework called Stable-GFlowNet, which improves upon the limitations of red team technology used to verify the safety of large language models (LLMs).


Red teaming is a safety verification technique that inputs attack prompts to generate AI responses that may be harmful or dangerous, thereby uncovering hidden vulnerabilities in the model. The ability to discover various types of attacks is crucial, as it allows for the preemptive elimination of more vulnerabilities, making both the success rate and diversity of attacks important evaluation factors.


While existing red team technology, based on reinforcement learning, achieved high attack success rates, it faced limitations in identifying diverse attack types. In contrast, the generative flow network (GFlowNet) excelled in attack diversity but struggled with unstable learning. The research team successfully applied GFlowNet to LLMs, enhancing learning stability and effectively uncovering a variety of vulnerabilities.


To address these challenges, the team developed three key technologies: 'Contrastive Trajectory Balancing (CTB)' to stabilize the learning process, 'Noise Gradient Pruning (NGP)' to eliminate unnecessary information, and 'Fluency Stabilization Device (MKS)' to assist in generating natural attack prompts.


As a result, Stable-GFlowNet discovered 134 unique attack types, approximately seven times more than the 17 types identified by previous technologies, while maintaining a high attack success rate of 92%. The defense model trained with this technology effectively defended against various attacks in cross-attack evaluations using different techniques, demonstrating high generalization performance.


The research team anticipates that this technology could also be applied in various AI generative fields, such as designing candidate drug molecules.


Professor Kim stated, "This technology is significant in that it can reliably identify various vulnerabilities in AI, even in environments with limited data and high noise. It will serve as a foundational technology for developing safe and trustworthy AI by identifying and defending against a broader range of risks before deploying generative AI in real-world applications."





* This article has been translated by AI.

Copyright ⓒ Aju Press All rights reserved.