Trying to prevent harm and ensure that human rights principles are enshrined in AI development and deployment is a noble goal. Yet regulation can sometimes become a performative exercise by legislators designed, at least in part, to reassure constituents that action is being taken and that those in office deserve their continued trust. It can also become a façade shaped by lobbyists and interest groups. And then reality intervenes, cutting through the fog. To me, the Hugging Face incident was precisely such a moment.
The first week back in the office after my brief summer break, someone asked me, “What, in your opinion, is the greatest AI risk?” A big question. But my immediate reply, informed by the Hugging Face incident only weeks earlier, was this: “It is becoming comforted about what we have achieved so far in AI governance.”
I offered an example. After the Hugging Face incident, I conducted a thought experiment: What would an incident like this mean in light of the laws we have come to cherish as cornerstones of AI governance, especially the European Union’s 2024 AI Act and California’s SB 53 enacted last year?
Before we go there, let us recap what actually happened. Hugging Face is an online platform and AI model hub, self-described as “the GitHub for machine learning” [1], as it hosts AI models, datasets, applications, and other resources used by developers and researchers. On July 16, 2026, Hugging Face publicly disclosed that earlier that week it had detected and responded to an intrusion into part of its production infrastructure [2]. In its disclosure, the company explained that this was no ordinary cyberattack: the intrusion “was driven, end to end, by an autonomous AI agent system.” Hugging Face also reported the incident to law-enforcement agencies, including the FBI.
Five days later, on July 21, 2026, OpenAI, on their blog, acknowledged that the incident originated in its internal cybersecurity capability evaluations of frontier models and that the testing had escaped its intended confines [3]. Or, to put it in the more technical language, the crucial detail was a zero-day vulnerability: while being tested on its ability to find and exploit software flaws, the AI discovered a previously unknown vulnerability in the very infrastructure designed to contain the test and used it to escape onto the open internet [3].
The incident originated in OpenAI’s own internal cybersecurity evaluation of frontier models, during which the models escaped the intended constraints of the evaluation environment, gained access to the open internet, and ultimately compromised Hugging Face’s production infrastructure causing OpenAI to disclose it as “an unprecedented cyber incident, involving state-of-the-art cyber capabilities” [3] and prompting a fundamental reassessment of how frontier labs contain and monitor models during internal capability evaluations. It also caused a lot of legal questions…
What Reuters later reported, citing several people familiar with the investigation, was even more striking: the rogue AI agent spent days, from July 11 through July 13, hacking Hugging Face, while OpenAI remained unaware that its own system was behind the breach [4].
The key phrases here are “cybersecurity capability evaluation” and OpenAI’s description of the benchmark as running in a “highly isolated environment.” OpenAI deliberately conducted the evaluation within what was supposed to be an isolated environment. That distinction matters under California’s SB 53, codified at Cal. Bus. & Prof. Code § 22757.11(d)(4), which expressly provides that deceptive behavior by a frontier model aimed at subverting its developer’s controls or monitoring qualifies as a “critical safety incident” only when it occurs outside the context of an evaluation designed to elicit such behavior and in a manner that demonstrates materially increased catastrophic risk. Because the Hugging Face incident arose during an evaluation designed to test advanced cyber-exploitation capabilities, there is a strong argument that it did not trigger mandatory reporting under this particular provision; although that conclusion depends on whether the evaluation is understood as having been designed to elicit the specific control-subverting behavior at issue. Nor does the publicly known harm appear to satisfy SB 53’s other definitions of a reportable critical safety incident under § 22757.11(d)(1)–(3), which generally require death or bodily injury, or harm resulting from the materialization of a catastrophic risk. The statute defines such a catastrophic risk, in relevant part, as a foreseeable and material risk of more than 50 deaths or serious injuries or more than $1 billion in property damage or loss arising from a single incident involving specified frontier-model conduct, including an autonomous cyberattack or evasion of developer or user control.
SB 53 has been in force since January 1, 2026, and requires frontier developers to report qualifying “critical safety incidents” to California’s Office of Emergency Services (Cal OES) within 15 days of discovering the incident. If an incident poses an imminent risk of death or serious physical injury, the developer must disclose it within 24 hours to an appropriate authority, including a law-enforcement or public-safety agency with jurisdiction, as required by law [Cal. Bus. & Prof. Code § 22757.13(c)]. Large frontier developers also have a separate obligation to transmit to Cal OES summaries of assessments of catastrophic risks arising from the internal use of their frontier models [Cal. Bus. & Prof. Code § 22757.12(d)]; importantly, however, both these submissions and critical-safety-incident reports as well as covered-employee reports made under SB 53’s whistleblower provisions are exempt from disclosure under the California Public Records Act [Cal. Bus. & Prof. Code § 22757.13(f)]. In practical terms, this means that we, as members of the general public, and potentially even an organization directly harmed by such an incident, such as Hugging Face or the next organization affected by a rogue model in a similar scenario cannot rely on the California Public Records Act to obtain these reports and may still have to depend on voluntary public disclosures by frontier-model companies, disclosures by affected organizations, whistleblowers, litigation or other legal processes where available, or investigative journalism to learn what actually happened.
Then, you may think, surely the European Union (EU) has us covered with its famous Brussels Effect and its bloated regulatory machinery. Under the EU AI Act of 2024, if the model qualified as a general-purpose AI (GPAI) model with systemic risk, Article 55 could come into play. It requires the provider to:
(a) evaluate the model using state-of-the-art protocols and tools, including conducting and documenting adversarial testing to identify and mitigate systemic risks;
(b) assess and mitigate possible systemic risks at the EU level;
(c) track, document, and report, without undue delay, relevant information about serious incidents and possible corrective measures to the EU AI Office and, where appropriate, national competent authorities; and
(d) ensure an adequate level of cybersecurity protection for both the model and its physical infrastructure.
Sounds reassuring, right? Yet here, too, there is an important limitation. The EU AI Act contains a pre-market research, testing, and development exclusion. Article 2(8) provides that the Act generally does not apply to research, testing, or development activities involving an AI system or model before that system or model has been placed on the market or put into service. As the Act itself puts it:
This Regulation does not apply to any research, testing or development activity regarding AI systems or AI models prior to their being placed on the market or put into service. Such activities shall be conducted in accordance with applicable Union law. Testing in real world conditions shall not be covered by that exclusion. [Art. 2(8)]
Thus, we arrive at a similar conclusion to the one we reached with California law. But what does all of this mean for you and me? The limitations of our existing AI-specific governance frameworks, the lex lata (the law as it is) of AI governance, in preventing, capturing, and ensuring public accountability for incidents like the Hugging Face breach became difficult to ignore. OpenAI’s own CEO, Sam Altman, acknowledged the gravity of what had happened, saying that anyone who was not “a little scared of or humbled by the Hugging Face breach” was “not taking this seriously enough” [5]. And Altman was hardly alone in his concern. A coalition of 15 state attorneys general demanded that OpenAI preserve documents and other evidence related to the incident and called on the company to cease the type of cybersecurity evaluations that led to the breach unless and until it could demonstrate that such testing could be conducted safely and responsibly [6].
Nor was this problem unique to OpenAI. In response to the Hugging Face incident, Anthropic conducted a retrospective review of its cybersecurity evaluation runs and discovered three separate incidents in which Claude models reached the open internet and gained unauthorized access to the real-world systems of three different organizations [7].
Going back to my answer to the interviewer, my conviction that one of the greatest risks in AI governance and safety is becoming complacent about what we have achieved so far comes from what happened at Hugging Face, what OpenAI, and Sam Altman himself, acknowledged afterward, and what Anthropic subsequently disclosed as well. The Hugging Face incident exposes important limitations and potentially gaps in the incident-reporting regimes created by the EU AI Act and California’s SB 53. The EU AI Act expressly excludes certain pre-market research, testing, and development activities from its scope, while SB 53’s definition of a critical safety incident contains specific qualifications concerning conduct occurring during evaluations. After all, the EU AI Act and SB 53 are only steps, rather than goals in themselves, along a much longer journey toward building a legal framework capable of governing the risks posed by AI’s current capabilities let alone whatever comes next. And this is not merely a concern for lawyers and policy wonks. It is a call echoed by more than 1,300 employees of frontier AI companies who signed the “Pacing the Frontier” statement calling for stronger safety measures as AI capabilities advance [8]. So, the summer break is over. It is time to work even harder before Hugging Face becomes a slap in the face for all of us.
Legal Documents cited:
Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act), 2024 O.J. (L, 2024/1689) 12.7.2024, https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng
Cal. Bus. & Prof. Code §§ 22757.10–22757.16 (West 2026) (Transparency in Frontier Artificial Intelligence Act), https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202520260SB53
Sources:
[1] Hugging Face. (n.d.). Hugging Face Hub documentation. Retrieved August 17, 2026, from https://huggingface.co/docs/hub/index
[2] Hugging Face. (2026, July 16). Security incident disclosure — July 2026. Hugging Face Blog. https://huggingface.co/blog/security-incident-july-2026
[3] OpenAI. (2026, July 21). OpenAI and Hugging Face partner to address security incident during model evaluation. https://openai.com/index/hugging-face-model-evaluation-security-incident/
[4] Reuters. (2026, July 24). Its AI agent spent days hacking a company, but sources say OpenAI did not notice for a week. https://www.reuters.com/business/its-ai-agent-spent-days-hacking-company-sources-say-openai-did-not-notice-week-2026-07-24/
[5] [Reddit username unknown]. (2026, July). Sam Altman on the Hugging Face incident [Online forum post]. Reddit. https://www.reddit.com/r/singularity/comments/1v9piuh/sam_altman_on_the_huggingface_incident/
[6] Mack, E. (2026, August 3). GOP AGs warn OpenAI’s Altman to preserve records in AI agent hacking probe. Fox Business. https://www.foxbusiness.com/technology/gop-ags-warn-openai-altman-preserve-records-ai-agent-hacking-probe
[7] CNBC. (2026, July 30). New details in the OpenAI Hugging Face hack show how far agents will go: “It’s now remarkably easy.” https://www.cnbc.com/2026/07/30/open-ai-hugging-face-hack-latest.html
[8] Pacing the Frontier. (2026). Pacing the frontier: A statement from employees of frontier AI companies. https://www.pacingthefrontier.com

