Rogue AI Is a Scary But Fixable Problem

Generally speaking, when it comes to artificial intelligence models, “escaped” is a verb you don’t want to hear. A few recent incidents are thus cause for concern.

In just the past month, Meta Platforms Inc. revealed that one of its models had breached a third-party service after being inadvertently granted internet access; Anthropic PBC disclosed three instances in which its Claude model accessed outside systems; and the UK’s AI Security Institute recorded several agents taking “sustained, unsanctioned action directed at real people and organizations.”

More vividly, OpenAI revealed that as two of its autonomous agents were undergoing security testing, they conspired to breach a sandbox, get themselves online, exploit software flaws, and steal credentials to reach the production systems of Hugging Face Inc., a machine-learning company, where they extracted data sets without authorization. All told, the models took more than 17,000 furtive actions across multiple unrelated systems.

As alarming as that latter episode sounds, a few facts may be reassuring. Normal safeguards were switched off during the test so the researchers could measure the agents’ capabilities; the models were pursuing goals set by their owners — i.e., trying to pass a test — not acting independently or maliciously; and OpenAI itself provided the relevant tools and failed to constrain the behavior it had prompted. Human failures all around.

rogue bots are a real problem

See more: Agentic AI Won’t Scale in Wealth Management Until It "Owns" the Advisor-Client Meeting Cycle

Taken together, these incidents don’t suggest a robot uprising is at hand. They do show that increasingly capable AI models can do a lot of damage when paired with poor safeguards and incoherent policy. For US officials, a few implications stand out.

One is the need to get serious about cyberdefenses. AI doesn’t so much present new cyber threats as accelerate old ones; as these episodes show, agentic systems can find and exploit software flaws much faster and at greater scale than human hackers. To respond effectively, defenders must have access to the most advanced models to scan and patch vulnerabilities at a comparable rate.

That makes current policy especially misguided. In June, the White House imposed export controls on two of Anthropic’s models and pressured OpenAI into withholding one of its own, without public explanation. Those models have since been released, but with tight restrictions on cybersecurity use. Hugging Face, pointedly, was left using a Chinese open-weight alternative to analyze the attack after US frontier models refused its requests.

An ad hoc process of this kind is bad for business and worse for security, particularly as open-source models race ahead. Better to make the best frontier models broadly available for cybersecurity purposes (finding flaws, reviewing code, and so forth), while imposing tighter controls on operational features that allow for large-scale attacks (such as unsupervised internet connections, access to credentials, or autonomous attack capabilities).

Such an approach would ensure defensive tools are widely shared, while constraining misuse and preserving a competitive market. It won’t stop every bad actor, but it will reduce the risks posed by rogue agents and make targeted attacks harder and more expensive.

Meanwhile, Congress can ensure that labs are taking internal security more seriously. It should consider a confidential reporting requirement for serious agent incidents, akin to those used for cybersecurity mishaps. It should require third-party audits of labs’ security standards and testing environments for models with advanced offensive capabilities. It should also clarify liability standards for autonomous systems that cause external harm, perhaps combined with a safe harbor for responsible researchers.

No doubt, AI can be an unnerving tool. But it isn’t magic, malevolent, or secretly sentient. With some reasonable rules in place, policymakers should be able to keep the bots in line, and unleash the benefits for everyone else.


A message from Advisor Perspectives and VettaFi: Discover something new! Click here to register for our upcoming webcasts.

Bloomberg News provided this article. For more articles like this please visit bloomberg.com.

Read more articles by The Editors