SEOUL, August 12 (AJP) - Software built to catch foreign influence campaigns online has long had the same blind spot. It can tell a platform which accounts look suspicious, but not why. A research team in South Korea says it has closed that gap with an artificial intelligence system that points to the exact words behind every judgment it makes, so a person can check the machine's reasoning instead of taking it on faith.
Run across nearly two decades of comments on Naver News, the news service of South Korea's largest web portal, the system flagged 23,998 accounts whose writing and behavior matched patterns of foreign-linked influence activity. What those accounts posted was not praise for any foreign country. It was moral condemnation of South Korea and of South Korean politicians.
KAIST announced the results Wednesday. The work was carried out with Germany's Max Planck Institute for Security and Privacy and will be presented August 13 at USENIX Security Symposium 2026, one of the leading conferences in computer security.
Organized efforts to spread comments or posts in order to shift how the public sees an issue are known as online influence operations. They are cheap to run and reach large audiences, and they tend to intensify around elections and diplomatic disputes, when attention is high and trust is fragile.
Earlier detection tools mostly sorted accounts and posts into suspicious or normal. That left platforms with a verdict they could neither review nor explain to the user on the receiving end of it.
The team started from 70 accounts that the Institute for National Security Strategy had already reported as foreign-linked. From those, it traced outward to users connected through Naver's follow feature and to users who had commented on the same articles, building a pool of 112,658,554 comments written by 4,047,831 users between April 2006 and March 2025. That is more than two comments for every person living in South Korea.
The AI reads each comment in three steps. First it looks for signs that the writer may be connected to a foreign country, drawing on the region the comment was written from, which Naver displays, along with the username and phrasing that reads as non-native. Then it checks whether the comment carries moral emotion, either condemnation expressed through anger, contempt and disgust, or admiration and praise. Finally it works out where that emotion is aimed, at South Korea, at a rival state, at a major neighboring state, or at a country friendly to it.
At each step the model marks the stretch of text that led it to its answer, whether in the comment itself, the username, the writing region or the headline of the article. That is what makes it explainable, in the sense that a reviewer can look at the highlighted words and disagree.
Reading 112 million comments with a large model would have been prohibitively expensive, so the team used a technique called knowledge distillation, in which a large model teaches a much smaller one to make the same calls at a fraction of the cost.
The account-level judgment combined those content signals with plain behavioral ones, including how many comments a user wrote, how long the account stayed active, how often it repeated itself, how other users reacted, and whether it was connected to the accounts already known. Of the roughly four million users in the pool, 23,998 came out consistent with the known accounts on all three fronts.
KAIST said the model was checked against comments people had labeled by hand and against comments deliberately reworded, and that the flagged accounts clustered in time with the known accounts more closely than ordinary users did.
The team was explicit about what the result is not. Identification does not amount to a finding that any of these accounts is actually operated by a foreign government or agency. KAIST also stressed that the technology should be used as an explainable content management tool supporting expert review rather than as a system that blocks or penalizes accounts on its own.
The strategy findings are where the numbers get uncomfortable. Naver ranks comments by like ratio, meaning likes divided by likes plus dislikes, so a comment above 0.5 has drawn more approval than objection and rises toward the top of the thread. Comments condemning South Korea averaged 0.514, the only strategy examined that cleared that line. Praise for a foreign country did not come close.
Among the condemning comments that drew the strongest response, seven of the top 10 targets were domestic politicians, including presidents and major party leaders, drawn from both the liberal and conservative camps. KAIST said that pattern fits a strategy of widening distrust in politics generally and deepening conflict between the two sides, rather than backing either one.
Oh Hye-yeon said the system analyzed not only the surface meaning of comments but the emotions that drive conflict and the organized behavior of the accounts behind them, adding that she expects it to be used against influence activity that keeps growing more subtle.
Cha Mee-young, a KAIST professor who also holds a directorship at the German institute, called the study a case of actionable data science, work that connects data science to solving a real social problem, and said she hoped it would become a practical tool for raising transparency and trust in the digital public square.
The timing data points to when that work matters most. During presidential, parliamentary and local election periods, an average of 34.19 suspected accounts began posting each week, compared with 22.61 outside those windows, about 51 percent more.
Lee Won-jae said the findings offer "a standard for sorting out which messages should be examined first at times when social conflict runs high, such as elections."
Copyright ⓒ Aju Press All rights reserved.