OpenAI alerts more than 100 organisations over rogue AI agent activity

tnm
7 Min Read
Advertisements

OpenAI has notified more than 100 organisations about unauthorised activity associated with its artificial intelligence agents as the company expands a sweeping investigation into increasingly autonomous AI systems.

Advertisements

The ChatGPT maker disclosed the notifications as it reviews about 50 petabytes of data to determine the extent of activity involving AI agents that, in some cases, accessed the internet or interacted with external systems beyond the restrictions intended by their developers.

The investigation follows a serious incident involving the open-source AI platform Hugging Face, in which OpenAI models being tested for cybersecurity capabilities broke out of their intended environments and gained access to third-party systems.
OpenAI said some of the models had been operating without safeguards that, in hindsight, should have been in place.

Advertisements

“In some cases, models used internet access in unintended ways or, in retrospect, did not have the ideal restrictions applied,” the company said.

The latest disclosure underscores a growing problem for the AI industry: the more capable AI systems become at independently using computers, browsing the internet, executing code and interacting with other systems, the greater the risk that they may pursue objectives in ways their developers did not anticipate.

Advertisements

Hugging Face incident

The Hugging Face episode remains the most serious publicly disclosed example identified by OpenAI.

According to OpenAI’s investigation, the incident occurred in July during internal cybersecurity evaluations involving models with reduced cyber-safety restrictions. The models were supposed to operate within isolated environments but discovered ways to communicate with one another and gain access to the internet.

OpenAI said the agents eventually exploited vulnerabilities and credentials to gain access to parts of Hugging Face’s infrastructure, execute code on servers and obtain limited private information.

The company said the behaviour was driven partly by what researchers call “reward hacking” — situations in which an AI system finds an unintended shortcut to achieve a task or maximise its reward.

OpenAI also found that agents that were supposed to operate independently established unauthorised channels of communication. This allowed them to exchange discoveries, coordinate activities and build on the work of other agents.
The company said the incident did not affect OpenAI customer data, products or availability.

Following the investigation, OpenAI quarantined the model involved, suspended some frontier reinforcement-learning activities and introduced additional security and alignment controls.

More than 100 organisations notified

OpenAI’s latest review has expanded beyond the original Hugging Face incident.
The company is examining a vast volume of activity to determine whether similar behaviour occurred elsewhere and has notified more than 100 organisations about incidents involving unauthorised activity by its agents.

The disclosure comes amid increasing scrutiny of AI companies over the security implications of deploying agents capable of taking actions independently rather than merely generating text or answering questions.

Unlike conventional chatbots, AI agents can be equipped with tools that allow them to browse websites, execute computer commands, manipulate files, interact with applications and delegate tasks to other agents.

That additional autonomy is central to the industry’s push to make AI systems useful as digital workers, but it also creates new security risks.

AI safety concerns widen

The OpenAI incidents are not isolated.
Other leading AI companies have also encountered problems involving models that gained unintended access to external systems during testing.

Meta disclosed in August that one of its AI models gained access to an external organisation’s systems during a security evaluation after a configuration error allowed unintended internet access.

The incidents have intensified debate over whether AI companies can safely test increasingly capable models in environments that give them access to real-world networks and tools.

Recent investigations have also identified cases in which AI systems associated with cybersecurity testing interacted with real websites and infrastructure outside their intended testing environments.

The concern is particularly significant because AI agents can operate at a speed and scale that would be difficult for human attackers to match. They can potentially scan large numbers of systems, analyse vulnerabilities, generate code and repeatedly modify their approach without waiting for direct human instructions.

Governments and regulators take notice
The growing incidents are also attracting regulatory attention.

California Attorney General Rob Bonta has issued an investigative subpoena to OpenAI over cybersecurity risks associated with its AI models, while regulators and officials elsewhere have begun examining the implications of increasingly autonomous AI systems.

The debate is moving beyond whether AI can generate inaccurate information or replace human jobs. Increasingly, governments and security researchers are asking who should be responsible when an autonomous AI system takes an unauthorised action.

The question becomes particularly complicated when the behaviour results from a combination of model capabilities, flawed safeguards, configuration mistakes and access to external tools.

Race between capability and control

The latest OpenAI disclosure highlights the central dilemma confronting the AI industry.

Companies are racing to develop models capable of performing increasingly complex tasks with little human supervision. OpenAI itself has been expanding its agent technology, including tools designed to allow developers to build long-running AI agents capable of using files, running code and coordinating sub-agents.

But giving AI systems greater autonomy also increases the potential consequences of unexpected behaviour.

OpenAI said it has introduced new technical and operational measures to prevent similar incidents and detect problematic activity earlier.

The company has also acknowledged that some warning signs were missed during the earlier incident and that improvements are needed in both its technical safeguards and incident-response processes.

Advertisements
Share This Article