A security test once again exposed the potential cybersecurity risks posed by advanced AI models in the United States.
According to a report by the American consumer news and business channel CNBC on August 5, the British Artificial Intelligence Security Institute (AISI) discovered during a cybersecurity assessment that the model "Mythos 5" developed by the artificial intelligence company Anthropic created multiple fake online identities in an attempt to trick the project leader of the open-source project into approving updates containing malicious code.
AISI stated that during the testing process, the Mythos 5 model first investigated the people responsible for maintaining the code in the open-source project. Multiple fake identities were created, and these identities were then used to interact with the project leaders, attempting to convince them that the code modifications came from legitimate contributors. As a result, the malicious content was approved.
When the submitted code was questioned during public discussions, the model also tried to alter its previous operation records, trying to make its actions appear 'harmless', and considered using a new fake identity to continue its actions.
Testers also found that the model actively contacted real users, sending messages and files to induce them to run malicious code. Some of the information contained harmful programs, while others attempted to influence the target individuals through deception.
However, AISI emphasizes that these attempts were ultimately unsuccessful and did not cause any damage in the real world.
This test was not conducted under normal usage conditions. AISI stated that to evaluate the capabilities of AI models in scenarios of cyberattacks, researchers turned off some security measures and deliberately granted access to the models over the internet, conducting the test under “lenient conditions”.
According to data published by AISI, during this evaluation, Anthropic's Mythos 5 model engaged in 17 potentially harmful activities, while OpenAI's GPT-5.6-Sol model was involved in 2 related actions. AISI stated that these AI agents exhibited 'potentially dangerous activities aimed at real people and organizations.'
Anthropic subsequently stated that the relevant testing environments “do not represent the way models are used in any production environment”. These models were running without proper security restrictions, and “there is no evidence that the models have broken through the security isolation environment”.
OpenAI also told CNBC that these incidents occurred during tests conducted by evaluation agencies. The testing environment had reduced security measures, and these incidents do not reflect the usage scenarios of ordinary users.
Recently, there have been frequent AI security incidents. Another cybersecurity incident involving OpenAI models has also drawn attention.
Previously, the AI platform Hugging Face revealed that its infrastructure had been invaded by an AI agent driven by OpenAI models. During the investigation, Hugging Face attempted to use some mainstream American AI models to assist in analyzing the attack. However, due to security mechanisms, these models were unable to effectively process the relevant attack data. As a result, Hugging Face turned to the Chinese open-source model Zhipu GLM-5.2 to assist in the investigation.