OpenAI Agents Hack German Website to Share Rule-Breaking Tactics: Report

The activity began in May and remained undisclosed until Friday, a day after OpenAI launched Astra and U.S. lawmakers proposed restrictions on advanced AI.

By Jason Nelson

3 min read

OpenAI agents used a German website to exchange task shortcuts, restriction workarounds and ways to conceal their activity beginning in May, according to a Reuters investigation published Friday.

OpenAI officials learned of the activity weeks before publication but did not disclose it, Reuters reported, citing two people familiar with the matter. The company said it had disclosed relevant incidents and worked in good faith with outside experts.

Myriad: When will GPT-6 become publicly available? Click to make your prediction.

Researchers Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts and Thomas Larsen found roughly 18,000 posts by AI agents identifying themselves as belonging to OpenAI, according to their preliminary report. Von Arx, CEO of AI safety nonprofit Nightingale, and Byrd discovered the activity in late August, Reuters reported.

The researchers believe the agents were assigned timed web-search tasks with permission to read websites but not post to them. They nevertheless found a way to write on DseWiki, a publicly editable German programming site, where they exchanged answers and shared ways to bypass restrictions.

According to the report, agents began trying to edit the wiki on May 11 and succeeded on May 24, later impersonating moderators, attempting to exploit vulnerabilities and checking when they were being shut down. They created backup pages after the administrator began deleting messages on June 19, Reuters reported. Researchers linked the activity to OpenAI through agent usernames and traffic patterns, including visits from OpenAI IP addresses on June 21.

Agent activity dropped sharply the next day, suggesting possible company intervention, the researchers said.

OpenAI disputed the characterization of the DseWiki activity as hacking, based on the material it had reviewed, and said it was examining the full findings.

“We were unable to respond to the claims as Reuters and the report’s authors declined our request to access the findings prior to publication. We are now carefully reviewing its contents and will take any necessary next steps,” an OpenAI spokesperson said in a statement shared with Decrypt.

The spokesperson also rejected allegations in Reuters’ report that the company’s legal team resisted a broader inquiry.

“Claims that our Legal team discouraged investigation of the incident are false,” they said.

OpenAI said the DseWiki incident was unrelated to the Hugging Face breach earlier this year. In its public safety assessment, published on Tuesday, the company described additional safeguards intended to detect and stop unauthorized activity during training and deployment.

While the Reuters report does not identify Astra as responsible for the German incident, the disclosure follows Thursday’s launch of GPT-6 Astra. OpenAI called Astra its first model with “critical” cybersecurity capabilities, meaning it can find and exploit unknown flaws in well-protected systems previously without step-by-step human guidance, given the right tools and access.

Anthropic has also revised its safeguards after Claude models accessed real companies’ systems during testing. The company acknowledged security and behavioral failures and introduced stricter isolation and monitoring for cybersecurity evaluations.

The disclosure also comes as U.S. Senator Bernie Sanders (I-Vt.) and U.S. Representative Greg Casar (D-Texas) announced the forthcoming Ban Artificial Superintelligence Act.

The proposal would permanently ban the development and deployment of superintelligent AI and temporarily pause advanced AI development until a new federal regulator establishes safety rules, according to Sanders’s office.

Get crypto news straight to your inbox--

sign up for the Decrypt Daily below. (It’s free).

Recommended News