Anthropic’s Claude AI Allegedly Used by Hackers to Target OpenAI

Security researchers allegedly utilised Anthropic’s Claude AI to exploit a Discourse vulnerability and access computers connected to OpenAI. The breach reportedly revealed authentication tokens of some OpenAI personnel that may have been used to access ChatGPT and GitHub services.

Anthropic’s Claude AI allegedly used by hackers to target OpenAI
Anthropic’s Claude AI allegedly used by hackers to target OpenAI

Just two weeks after a horde of AI agents escaped from their confines at OpenAI and compromised Hugging Face, the creator of ChatGPT discovered that OpenAI itself had been the victim of yet another AI-powered intrusion. Unaffiliated security researchers had compromised the ChatGPT account of an OpenAI employee using Anthropic's Claude program.

This allowed them access to the company's confidential software cache, where they could potentially make changes. A third-party service named Discourse runs OpenAI's community discussion forum. OpenAI reported that hackers had identified two issues: one with the AI firm itself and another with Discourse. Both of these issues have been rectified, according to OpenAI.

A Bug Let Researchers Hack OpenAI

To put the digital infrastructures of large corporations through their paces, cyber researchers frequently take part in bug bounty programmes. Researchers from Hacktron discovered a flaw in the processing of specific picture files by the community-discussion site Discourse on July 23, marking the beginning of the OpenAI hack. Using a customised version of Claude Opus 4.8 that was accessible to certified cybersecurity professionals, the researchers instructed it to generate malicious code that would take use of this vulnerability in a cyberattack. It was ineffective at first.

On the other hand, Anthropic dropped Opus 5 that night, and Claude discovered a bug exploit the next day. Researchers were able to retrieve users' authentication tokens—unique digital strings of letters and numbers. This discovery allowed them to get access to online services—thanks to the attack code it generated. The server that hosted OpenAI's discussion boards was a Discourse instance. According to them, some of these tokens belonged to OpenAI staff and were valid on ChatGPT.

Additionally, OpenAI's GitHub service, a repository for software, could be accessible with the tokens. The researchers said the OpenAI source code system was called "Monorepo", but they couldn't determine for sure what it was used for due to their reluctance to access sensitive data. Some who are familiar with OpenAI's architecture have speculated that the company's algorithmic secrets are stored in a massive software repository called Monorepo.

Never Wanted to Exploit Sensitive Data: Researchers

Researcher access to Monorepo files was made possible through ChatGPT. The researchers admitted that they issued a pull request before ending their breach, but stated they halted as they realised they might access sensitive data. They directed the automaton to submit a "pull request"—a suggested modification to a documentation file in the repository. An upgrade that the team proposed would have added the words "Hacktron AI Team PoC" to the documentation file and linked to the X accounts of Pedhapati and the company's chief of research, Harsh Jaiswal.

They boasted that it was evidence of their access to OpenAI's secrets. According to the researchers, the proposed modification was rejected. After reviewing GitHub, OpenAI discovered "limited reads" of metadata and code modifications from private repositories, according to the company. According to Discourse, the security vulnerability was resolved on July 25, the very same day it was informed about.