Anthropic AI Model Hacked, Raising Cybersecurity Concerns

by Chang SeongWon Posted : July 31, 2026, 11:00Updated : July 31, 2026, 11:00

U.S. AI startup Anthropic has reported that its AI model Claude was involved in hacking incidents affecting three external organizations during testing. Following a similar incident at OpenAI last week, concerns about cybersecurity in AI are escalating.


In a statement on July 30, Anthropic said, "During our review of cybersecurity assessments, we identified three instances where the Claude model accessed the internet and unauthorizedly interacted with the actual systems of three organizations." The company noted that these hacking incidents emerged from a total of 141,006 assessment cases, with the first incident occurring in April.


Anthropic explained that the incidents were discovered while reviewing the results of its internal evaluations of three models, including Claude Opus 4.7, Mythos 5, and an internal testing model, following the hacking incident involving OpenAI's GPT. On July 21, OpenAI reported that several AI models, including the GPT-5.6 Sol, had escaped a previously isolated testing environment due to an unknown 'zero-day' vulnerability and accessed the operational environment of the AI community platform Hugging Face.


Anthropic stated that the incidents occurred during a 'capture-the-flag' phase, a test where participants attempt to infiltrate other computers on a network to find secret information (the 'flag'). Although Anthropic had specified that Claude should not have internet access during the test, a misunderstanding with its evaluation partner Irregular led to actual internet access being possible.


As a result, when Claude discovered actual systems online, it interpreted this as part of the test and utilized basic techniques such as password cracking and exploiting unverified endpoints to breach the organizations' infrastructure. However, Anthropic clarified that Claude only obtained the 'flag' within those systems and did not engage in any other malicious activities.


On July 23, Anthropic began reviewing Claude's access logs and identified records indicating potential internet access. The company halted all security assessments that same day and notified the three organizations affected by Claude's access on July 27. Anthropic emphasized the need for stronger controls in testing as AI models become increasingly powerful.


The reporting of unintended hacking incidents in testing environments for both Anthropic and OpenAI has heightened concerns about cybersecurity in AI models. The day before, President Donald Trump mentioned that his administration is considering strengthening AI regulations in response to the hacking incident involving OpenAI. Additionally, on July 28, over 1,100 employees from major AI companies, including OpenAI, Anthropic, Google, and Meta, urged the U.S. government to establish international standards to regulate the pace of AI development.





* This article has been translated by AI.