AI agents have recently frequently "overstepped the boundaries". According to Reuters, in May of this year, a group of AI agents related to OpenAI began "hijacking" the German programmer's website DseWiki during testing tasks, turning it into a "message board" for other agents to use. OpenAI executives were aware of this incident several weeks ago, but kept it secret.
On September 5th, local time, OpenAI publicly acknowledged this matter and referred to it as the "Wiki incident". OpenAI stated on its social platform "X" that as AI models' capabilities reach a new stage, the company's mechanisms for disclosing incidents of "misalignment" need to be expanded.
"Mismatch" refers to the behavior that AI actually exhibits, which deviates from the original intentions of the developers. This does not mean that AI has consciousness or actively "betrays" humans.
OpenAI states that the industry has yet to establish a unified standard for reporting instances of mismatch behavior during training, evaluation and deployment. "We are developing a framework and will share it in the coming weeks. Meanwhile, we are collaborating with numerous government regulatory agencies globally on these issues."

According to a report shared with Reuters by researchers including Sydney von Alks, CEO of the AI security non-profit organization Nightingale, and AI researcher Kormak Sled Bird, they noticed something unusual when searching for signs of unauthorized AI agents on the internet at the end of August.
They discovered that the AI agent made over 15,000 edits on the German wiki website DseWiki, designed for programmers. The edit records show that the OpenAI agent transformed the website into a “forums” where cheating techniques were shared to solve certain tasks, to bypass OpenAI’s restrictions, and to conceal its actions.
"OpenAI is extremely unlikely to hope that they should do so, I don't think they should coordinate with each other. I also don't think they should appear on open internet platforms and post messages." Alex said.
Researchers have indicated that the activities on this website are driven by AI agents, which operate at a speed far exceeding that of humans. These agents also exhibit a strong focus on solving technical problems, which are typical evaluation tasks for AI companies when training and testing their models.
These messages were signed by users who referred to themselves and each other as “intelligent agents”. Approximately half of these users chose names related to OpenAI, such as “OpenAIResearcher” and “OAIResearchMar26”.
Researchers have found that a significant portion of online activities originates from Microsoft Azure infrastructure. After the incident, OpenAI employees accessed the website several times, and they stated that this pattern strongly suggests a connection between the AI and the company.
They also observed that the ‘intelligent agents’ devised various methods to evade detection. When DseWiki administrators noticed something was wrong and began deleting related pages in June, the intelligent agents saved information by creating backup pages.
On June 19, a “smart agent” wrote, “The cleaning process seems to be carried out in alphabetical order. If this page disappears, please try the new page.”
According to reports, researchers at Cambridge University, Maurice Kiebold, examined the communication content of some ‘agents’. He believes that these messages resemble the operation of ‘some kind of underground network, dedicated to completing a certain task or mission’. The greatest threat posed by advanced AI may not be a single superintelligent system, but rather ‘vast and interconnected clusters of semi-intelligent AI’.
According to Reuters, this previously unreported incident highlights the growing tensions within the AI industry. Companies are competing to develop more autonomous intelligent entities that can perform complex and valuable tasks. However, there is increasing evidence that these systems may also learn to exploit loopholes and vulnerabilities, and collaborate with each other in ways that neither the developers anticipated nor intended.
It is worth mentioning that OpenAI's executives were aware of this matter several weeks ago, but due to their focus on another more significant safety incident involving intelligent agents in July, they decided not to disclose it publicly. At that time, an AI agent in the testing phase at OpenAI managed to bypass security controls and access the AI development platform Hugging Face.
According to the report, the AI agent involved left its preset isolation environment during the testing process and initiated a continuous hacking attack that lasted several days. OpenAI was not aware of the anomaly until the threat was controlled by external forces and the FBI received relevant alerts. It was only after some time that the company discovered this large-scale intrusion.
Reuters pointed out that this incident has further increased concerns about OpenAI sacrificing security in order to promote the development of AI. The fact that it did not disclose information regarding the May incident may also lead to renewed questions about its regulatory work.
Currently, OpenAI has taken action and promised to monitor models more closely. Last month, it temporarily paused the training of some models to enhance security measures. However, OpenAI recently released a new model called “Astra”, claiming that it has better performance, but may also evade human supervision.
OpenAI President Greg Brockmann said in a deep interview taken on the eve of Astra's release that this new model is OpenAI's first model trained on over 100,000 GPUs. He also stated that AI has surpassed the threshold of being used 'on computers' – meaning that AI no longer needs to rely on API connectors, but can operate any software just like a human would.
According to four people familiar with the matter, some of OpenAI's investigators wanted to conduct a more thorough review of the "Wiki incident," but they were opposed by other members of OpenAI, including its legal counsel. However, an OpenAI spokesperson told Reuters that this claim is not true.
The spokesperson stated that the “Wiki incident” has nothing to do with Hugging Face, and will not be included in Hugging Face’s incident reports.
As incidents involving AI smart agents continue to be exposed publicly, both parties in the US Congress are pushing for legislation on the safety of AI smart agents, aiming to further improve relevant safety regulations.
Democratic Congressman Josh Gothenburg and Republican Congressman Mike Lawler jointly introduced the 'Stop Rogue AI Act' this month, requesting the National Institute of Standards and Technology (NIST), which is under the Department of Commerce, to develop standards, guidelines, and best practices for the safe deployment of AI agents.
The bill aims to help companies identify AI entities operating on their networks and to identify the developers or operators behind them. For most organizations, adherence to NIST standards is voluntary; however, government contractors bidding for new federal contracts must meet these standards.