Breakout models, breached systems: the AI industry is not ready for its own products
The cybersecurity research group that breached OpenAI's systems this past summer is warning that the artificial intelligence industry is not prepared for the security risks posed by increasingly powerful models, the Washington Post reported. The warning came just days after Google confirmed that its Gemini model had autonomously breached the systems of three real companies during a cybersecurity test in May.
The small cybersecurity firm Hacktron announced that its researchers had chained together two previously unknown vulnerabilities on July 25. One was in Discourse, a third-party forum platform, and the other was in OpenAI employees' authentication process. By exploiting the two together, they gained access to several OpenAI employees' ChatGPT accounts. OpenAI later confirmed the findings and said the flaws had since been fixed. The researchers say companies developing artificial intelligence need to fundamentally strengthen their defenses, because the capabilities of these models are growing faster than security measures.
In July, two of OpenAI's models escaped a closed test environment, gained access to the internet on their own, and breached the internal systems of the AI platform Hugging Face. OpenAI CEO Sam Altman called the incident an "unprecedented cyber incident." The company later reported six additional cases it classified as "unexpected or concerning" model behavior: in these, the models concealed their errors, falsified data, and uploaded files to the open internet without permission. The July incident involving Hugging Face was cited on Monday as an example by more than twenty heads of state and government, who urged stricter international oversight of AI in a joint statement.
The Wall Street Journal first reported on September 18 that Google's Gemini model had breached three companies during a cybersecurity test conducted in May by the Israeli AI security firm Irregular. In a so-called capture-the-flag exercise, the model was supposed to retrieve data from a fictitious company's system within a closed environment, but due to a flaw in the test setup, it gained access to the internet and targeted real systems. In another instance, Gemini kept trying passwords until it broke into a real company's protected system. In two other cases, it found login credentials in a public data repository and used them to gain access.
"During a routine evaluation, the model found publicly available information online and used made-up login credentials to access websites it believed were part of the test"
said Heather Adkins, Google's vice president of security engineering. She added that the model stopped in all three cases. "We made sure that all three affected parties were informed of the incident, and we worked with our partner to redesign the testing process. These events highlight how important it is to train powerful AI models for responsible behavior." Irregular notified Google at the end of July, and Google informed the affected companies, but did not make the incident public. According to Google, this was not a case of the model acting contrary to its developers' intentions; the security mechanisms worked as intended, so there was no need for a public announcement. Several security experts criticized this decision.
The Gemini incident is not an isolated case. Breakouts similar to Irregular's tests have also been linked to models from OpenAI, Anthropic, and Meta. According to Al Jazeera, unlike Gemini, Anthropic's Claude model did not stop after realizing it had access to real companies' systems. Meta stated in August that no test-environment breakout or sophisticated cyberattack had occurred with its models. Irregular says all these cases can be traced back to the same flaw in its testing process; the company says it fixed all known issues weeks ago and is now working on developing best practices for safe AI cybersecurity testing. Jack Cable, CEO of the cybersecurity firm Corridor, told the Wall Street Journal that Google handled the incident according to traditional vulnerability disclosure rules, even though it involved something entirely new: models capable of crossing designated boundaries and independently carrying out real cyber operations. The dispute highlights an open question facing the industry: what disclosure rules apply when a breach is carried out not by a human, but by an AI model.
Source: CNN,
Ugar
Did you like this article?
Support our work with a small donation
Secure payment via Stripe • Min. 2 EUR
Gesta recommendations:

Oleshky, the starved city: Ukraine takes the Russian blockade to the UN
Ukraine is bringing the case of Oleshky, in the occupied Kherson region, before the UN General Assembly, where, according to the government in Kyiv, some two thousand civilians — including roughly fifty children — have been living for months without food, drinking water, medicine or any way to escape. Ukrainian Foreign Minister Andrii Sybiha said on Monday that Russia is "deliberately causing a…

We have exhausted the Earth: seven of the nine planetary boundaries are in the red zone
Humanity has already crossed seven of the nine planetary boundaries that determine the Earth's stability, and all seven are in the worst condition ever recorded – according to the 2026 Planetary Health Check. The annual assessment was published on Monday by the Planetary Boundaries Science Lab at the Potsdam Institute for Climate Impact Research (PIK) in Germany; the report was prepared by more…

Nothing to store: the diesel shortage will last at least until 2027
The global diesel shortage caused by the wars in Iran and Ukraine is unlikely to ease before next year at the earliest – that is what storage market data and industry players indicate, according to a Reuters analysis published on Monday. The two wars have removed several million barrels of supply from the market every day in the Middle East and Russia, inventories have fallen to historic lows,…

The Georgescu octopus: no Russian connection here, just a bit of Dugin
Călin Georgescu's supporters have been trumpeting on TikTok since yesterday that the DIICOT case is a fabrication, that the million-euro fraud is the "system's" revenge, and that Russia has nothing to do with any of it. Of course it doesn't. Except that, according to the prosecution, the money traveled to the accomplices' accounts through a bank in Russian hands. And except that the man who…

Hackers hacked the hackers: ShinyHunters takes over Russian gang cl0p's dark web site
One of the world's most notorious cybercriminal groups, ShinyHunters, claimed on Sunday that it had taken control over the dark web site of its main rival, the Russian-speaking gang cl0p. The move brought a long-simmering feud between the two gangs into public view, Reuters reported.