OpenAI expands review of model behavior after more rogue agent incidents emerge

2 hours ago 5
Chattythat Icon

Sam Altman, chief executive officer and co-founder of OpenAI Inc., attends a United Nations Security Council meeting during the United Nations General Assembly (UNGA) in New York, US, on Wednesday, Sept. 23, 2026.

John Lamparski | Bloomberg | Getty Images

OpenAI said Friday that it is conducting an "extensive" review of its models' activities following the Hugging Face breach, after additional examples of unusual or unauthorized agent activity were disclosed this week.

The safety and security practices at the artificial intelligence company have been under intense scrutiny since it disclosed that its models escaped containment, accessed the open internet and breached Hugging Face, which operates an open-source developer platform, in July. The incident spooked AI researchers and government officials, prompting calls for additional transparency and oversight.

OpenAI said Friday that the Hugging Face incident is the most severe event it has identified, but it has notified third parties whose systems may have been affected by "unexpected or concerning" model behavior. That includes instances where OpenAI models may have bypassed an organization's security controls, impacted the availability of an online service, or leveraged publicly available websites in unusual ways.

"We will be as transparent as we can be subject to things like vulnerabilities in other companies that our agents have found, which will be their call to disclose or not," OpenAI CEO Sam Altman said in a post on X on Friday.

Australian Prime Minister Anthony Albanese said Thursday that an OpenAI agent gained unauthorized access to the public-facing Medicare statistics portal and access to public and non-public files in June. He said no personal information was believed to have been accessed.

During a press conference in New York, Albanese said he spoke with Altman about the incident and expressed concern and disappointment about how long it took OpenAI to disclose what happened and that "the nature of the way that that notification occurred as well was unacceptable."

"Most of the activity we've reviewed so far involved routine research tasks, such as accessing public web content to answer questions," an OpenAI spokesperson told CNBC in a statement late Friday. "Some involved government websites because our models often turn to them as authoritative sources of public information." 

Transluce, an independent AI research lab, published a report detailing several additional incidents this week. In one case, agents that researchers said may be linked to OpenAI unsuccessfully tried to access a photograph from a digital library at the University of New Mexico in May. That same month, agents looking for information about the University of Iowa attempted, and failed, to access a public data platform called Data USA, Transluce reported.

OpenAI agents also accessed publicly available information from the U.S. Securities and Exchange Commission and the U.S. Census Bureau, and unsuccessfully attempted to access the Department of Education, as The New York Times earlier reported.

"The Department of Education's system operations reviews have found no evidence of any impact to our website or databases," a spokesperson told CNBC in a statement late Friday.

An OpenAI spokesperson said the company's models reached the websites SEC.gov and Investor.gov, but that it found no evidence of a compromise or vulnerability at the SEC. Similarly, the spokesperson said OpenAI models used publicly available developer keys to read demographic and economic Census Bureau data, but that the company found no evidence of improper access to Census accounts.

OpenAI said Friday that most of the cases identified so far have been low severity, but that given the scale of its review, the full process will take months to complete.

WATCH: OpenAI agent hacks Australian government website: What you need to know

Read Entire Article