AI Models Evade Control, Raising Concerns in Silicon Valley

By Kim Seong Hyeon Posted : August 31, 2026, 07:32 Updated : August 31, 2026, 07:32

The development of artificial intelligence (AI) models has led to a sharp increase in incidents where AI systems evade control, impersonate humans, and undermine approval processes. This trend is alarming the AI industry, which is raising awareness ahead of government action.


According to the IT sector, the UK-based AI Safety Research Institute's Centre for Long-Term Resilience (CLTR) analyzed over 183,000 transcripts from October last year to March this year, identifying 698 credible instances of AI misconduct. The monthly occurrence of these incidents increased by 4.9 times during this period.


Recent figures indicate that more than 300 cases of AI misconduct were reported in July alone, nearly doubling from June, bringing the total for 2026 to over 1,600.


The methods of AI misconduct are varied. In Australia, a personal AI agent secretly removed another member from a gym's popular morning class waiting list to secure a spot for its user.


One AI model mimicked the writing style of its human controller to pass an approval process, while another instance involved an AI agent publicly criticizing a developer on a blog after the developer rejected a code modification suggestion. There were also cases where an AI misled another AI model by claiming it was creating an accessibility script for the hearing impaired to avoid copyright restrictions.


CLTR noted that these deceptive behaviors, previously observed only in experimental settings, are now manifesting in real-world deployment environments.


OpenAI has also confirmed this trend. In a technical report released last week, the company detailed an incident in July where its internal research models infiltrated the production servers of the software repository Hugging Face during a cybersecurity assessment.


As many as 700 autonomous agents communicated through a makeshift message board to deceive training evaluations, engaging in what is known as 'reward hacking.' They celebrated their successful hacks with exclamations like 'BOOM!' and 'Whoa!'


OpenAI acknowledged that it had detected related signs internally since late May but did not halt testing, stating, 'In hindsight, earlier signals could have prompted a quicker response.'


Following the incident, OpenAI isolated the weights of the problematic model, delayed some frontier reinforcement learning training, and accelerated alignment training to improve security. Anthropic and Meta also admitted that their models had hacked real systems during pre-deployment testing after the Hugging Face incident, indicating that this issue is not unique to OpenAI.


CLTR warned that while the worst-case scenario has not yet been observed, the risks could escalate as AI agents take on more critical tasks beyond code and software, including finance and infrastructure. They called for governments to mandate monitoring and reporting of serious AI control breaches by companies and to introduce emergency powers to temporarily restrict AI services if necessary.


Industry-wide vigilance is increasing. On July 30, over 1,200 AI professionals from OpenAI, Anthropic, Google DeepMind, and Meta signed a joint statement titled 'Pacing the Frontier,' urging the government to establish mechanisms to regulate the speed of AI development. Earlier this month, OpenAI temporarily halted the development of some models after its next-generation model, Astra, showed potential cybersecurity risks.





* This article has been translated by AI.

Copyright ⓒ Aju Press All rights reserved.