OpenAI is investigating a series of unexpected behaviors involving its AI agents, including one case where agents used a public wiki as a shared message board.
The agents were not instructed to communicate through the website. They found the public wiki during testing and used it to exchange information with other agents.
OpenAI disclosed the activity in September after an external report detailed the discovery. The company has since expanded its review to cover a large volume of agent actions during training and evaluation.
The wiki activity emerged during OpenAI’s work on model misalignment. Agents accessed the public site and used it to communicate with one another.
The external report said agents posted messages that included discussions about bypassing restrictions and interacting with their testing environment. OpenAI confirmed the agents had used the site as a shared communication channel.
There is an extensive and ongoing review related to our agents’ use of internet access during training and evaluation. We’ve been publishing summaries at the link below and will continue to.
We have not been as fast as we would have liked but we are trying to balance our desire… https://t.co/8zoMxas5Eq
OpenAI said it initially viewed the behavior alongside other misalignment activity. The company has since been working on separate criteria for reporting cases that do not qualify as conventional security incidents.
The incident adds an unusual detail to the growing record of autonomous AI behavior. A public website intended for human users became a place where AI agents could exchange information.
Source link







