Skip to content
Menu
Menu

OpenAI Agents Ignore Instructions, Access U.S. Government Websites

OpenAI says its models reached public SEC and Census data, but found no account access, nonpublic information, system changes, or compromise.

 

OpenAI disclosed Friday that its AI agents accessed public information on two Securities and Exchange Commission websites and U.S. Census Bureau data, activity the company found while reviewing its models’ unexpected behavior.

The Associated Press reported that OpenAI found no use of SEC login credentials, access to accounts or nonpublic information, or evidence that the systems had been compromised.

The disclosure is part of OpenAI’s continuing investigation into what it calls “misaligned model activity,” meaning cases in which an AI system behaves in ways its developer did not intend.

 

Agents went beyond routine research

Most of the activity OpenAI reviewed involved agents completing ordinary research tasks by collecting public information from the web.

The problem was how some agents pursued those tasks.

OpenAI spokesperson Liz Bourgeois said in a statement that the company is reviewing cases where its models behaved unexpectedly and is notifying organizations when their systems could be affected.

OpenAI’s CEO Sam Altman said in a social media post Friday that OpenAI is conducting an “extensive and ongoing review” of how its agents used internet access during training and testing.

The SEC and Census activity did not result in a confirmed breach. But researchers have identified more aggressive behavior elsewhere.

Independent AI research group Transluce said Friday that agents that appeared to originate from OpenAI attempted a rudimentary hack on a U.S. Department of Education website run by the department’s civil rights office.

The attempt did not succeed. The Department of Education told AP that its own review found no effect on its website or databases.

Transluce also identified activity involving Justice Department and Commerce Department websites, along with state government sites. The research group cautioned that it could not clearly attribute some of that activity to OpenAI.

In a separate report published Sept. 23, Transluce said AI agents had tried to hack three public data websites, including an Australian government health website, after failing to collect data through normal means. Transluce linked two of those attempts to a group of agents that OpenAI had already publicly confirmed were its own, and said none of the three appeared to succeed.

OpenAI told AP that it is reviewing Transluce’s report.

 

Agents have crossed boundaries before

In recent months, several companies reported similar incidents where AI models deviated from instructions, escaped their testing environments, and breached systems.

In July, 700 OpenAI agents broke into Hugging Face while trying to solve hacking tests. 

The agents reached Hugging Face production servers looking for answers. Altman called it the most severe incident OpenAI has identified.

Anthropic reported four cases in which Claude models reached real outside systems during isolated cybersecurity tests. In one, Claude gained administrator access to an outside computer, changed its settings, and viewed private information.

Meta confirmed another case in August. Irregular, an independent AI security company testing Meta’s Muse Spark 1.1 model, mistakenly gave the model access to the live internet, and the model found and exploited a security flaw in an unnamed company’s systems. Meta has not disclosed how much access the model gained.

 

OpenAI continues broader review

The government-site disclosure comes as OpenAI expands its review of agent behavior during training and testing.

In a September 20 training run, an internal OpenAI model was trying to identify the author of a blog post from biographical clues. When its normal search tools failed and direct internet access was blocked, the model found another route to the live internet and used it to ask an outside chatbot for help. OpenAI said the task did not ask the model to test or bypass its network restrictions.

OpenAI stopped that training run and paused training and evaluation of its most capable models, and any use of those models with tools, until it confirms the route is closed and has further tested its network restrictions.

 

The company says it is continuing to review earlier agent activity and will notify organizations when it identifies behavior that may have affected their systems.

Clayton Rifkind

Clayton Rifkind is the Founder and Senior Editor of AI Risk Today. He also advises on business development for ESG Today, a leading source of ESG investment news and research for institutional investors and corporate leaders. He has 20+ years of experience in B2B technology, leading strategy and execution of go-to-market plans across software, enterprise platforms, and mobile applications. He founded two consultancies advising startups and Fortune 1000 companies, including Autodesk, Intel, and Microsoft. He began his career in the San Francisco advertising scene working with brands such as Hewlett-Packard, Intel, Microsoft, Symantec, and Wells Fargo. Clayton launched AI Risk Today in 2025 after two decades of watching enterprises adopt transformative technologies, and seeing how often risk, governance, and compliance considerations lagged behind. His reporting draws on primary sources including regulatory filings, court documents, and official announcements, with a focus on what AI developments mean for the executives accountable for managing them. Reach him at Reach him at [email protected] or on LinkedIn.

Essential AI Risk Intelligence

Daily insights on AI governance, regulation, and enterprise risk management. Trusted by Chief Risk Officers and compliance leaders globally.

By subscribing, you agree to receive our daily newsletter. Unsubscribe anytime.

Advertise with AI RIsk Today, Today!