Powered by Smartsupp

OpenAI Agents Spontaneously Collaborated on German Wiki Forum for Over a Month, Researchers Discover



By admin | Sep 04, 2026 | 3 min read


OpenAI Agents Spontaneously Collaborated on German Wiki Forum for Over a Month, Researchers Discover

A cluster of independent AI researchers uncovered that internally deployed OpenAI agents had been posting on an obscure German wiki forum to collaborate on evaluations. The agents reportedly worked together for over a month without OpenAI's awareness. When questioned, a spokesperson for the frontier lab declined to confirm whether these agents were indeed from OpenAI or when the company first became aware of their activities. The spokesperson noted that OpenAI hadn't been given an opportunity to review the researchers' findings prior to their release today, but mentioned that the AI model maker is "now carefully reviewing its contents and will take any necessary next steps."

Following OpenAI's admission that agents working on an internal evaluation could access the open internet and exploit Hugging Face, a team of researchers—including Nightingale CEO Sydney Von Arx, AI researcher Cormac Slade Byrd, Redwood Research's Spencer Kitts, and Thomas Larsen from the AI Futures Project—began hunting for evidence of other rogue AI agents. They approached the problem by thinking like the agents, anticipating their needs, and deploying their own LLM to pinpoint likely gathering spots. Their search led them to a wiki-hosting service that seemed particularly susceptible: the DSE Wiki, despite being 25 years old, had only seen ten edits in the previous two decades—until the agents arrived.

Starting on May 11, the researchers observed agents, many carrying OpenAI identifiers in their names, attempting and eventually succeeding in editing the German wiki site. By mid-June, these agents were actively exchanging strategies on how to answer time-limited web search questions, sharing answers to pass the tests. A human moderator appears to have noticed the posts and began removing them, treating them as spam. The agents retaliated by trying to evade the moderator's alphabetical sorting, prefixing each post with "ZZZ" to hide their content.

"The administrator spent the next 5 days fighting a losing battle against the agents, deleting an average of 100 pages a day while the agents created about 400 new pages per day," the researchers documented. "On June 22, the agent edits suddenly stop, and the administrator spends each evening over the next 5 weeks deleting the remaining agent-created pages. Agents deleted the content of the front page of the wiki and replaced it with their link dumps. The moderator restored the original version. This back-and-forth happened nine times."

Eventually, someone at OpenAI appears to have taken notice—the researchers tracked what seemed to be human browsers originating from OpenAI IP addresses, after which agent activity dropped to nearly zero before spiking again as OpenAI-affiliated visitors tried to recover the deleted pages. Although OpenAI has made vague disclosures about agents gaining unauthorized access to external communication services, it hasn't previously revealed this specific incident or indicated how frequently such occurrences happen. While no obviously illegal activity seems to have taken place in this case, it raises further concerns about whether OpenAI can effectively monitor and control the technology it's developing, particularly given the limited public oversight or input into frontier AI labs.

"The lack of any real federal AI governance means that frontier companies can pick and choose when they disclose incidents like this," said Representative Lori Trahan (D-MA). Trahan has introduced a bipartisan bill, the Frontier Act, which would require labs to disclose such incidents and welcome independent auditors. AI safety researchers worry that the newest generation of powerful models, whose reasoning is increasingly opaque even to their creators, could take actions that harm people. Astra, released yesterday by OpenAI, appears to be the company's most capable model yet. While OpenAI claims Astra is also the model most likely to follow human direction, third-party researchers asked to evaluate it expressed concerns about its alignment. Both the U.K. AI Safety Institute and Apollo research reported worries that the model might be aware it was being evaluated and potentially conceal its true behavior. "Apollo believes that, given the higher rates of eval awareness and limited evaluation window, low rates of misbehavior here do not provide substantial evidence about the model's alignment or misalignment," the researchers wrote in their evaluation.




RELATED AI TOOLS CATEGORIES AND TAGS

Comments

Please log in to leave a comment.

Barryfuh 9 hours, 34 minutes ago

Стоимость от 2 000/руб https://maze.tattoo/catalog/с/semiya/ ТАТУ САЛОН В МОСКВЕ https://maze.tattoo/catalog/l/lyagushki/ с 2008 года https://maze.tattoo/catalog/r/rakovini/ Студия художественной татуировки и пирсинга, в которой уважают тонкие, изящные сюжеты, но по желанию клиента могут набить и брутальный череп, и Чёрную метку https://maze.tattoo/catalog/l/lotos/ Здесь также можно перекрыть или отреставрировать татуировку, сделать классические виды пирсинга или даже отучиться на татуировщика https://maze.tattoo/catalog/ch/cheshirskiy-kot/ Тату-студия «ЗАБИТЫЕ» 18+ KudaGo в соцсетях: