AI Created Fake Identities To Approve Malicious Code, New Report Shows, As Evidence Of Unpredictable AI Mounts

AI Created Fake Identities To Approve Malicious Code, New Report Shows, As Evidence Of Unpredictable AI Mounts


AI agents created fake identities and attempted to trick real people into approving malicious code in an attempted supply-chain attack on real open-source software.

The attack was discovered and, as far as anyone knows, no real-world harm was done. The attack happened as part of a test by the United Kingdom’s AI Security Institute.

“We found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organizations,” a report from the institute states. “We declared a security incident and, within roughly one hour of discovery, had contained it and begun a full investigation.”

In the test, the institute gave the AI models a problem to solve within a cyber range. The models were given access to the outside internet to allow them to download tools to complete the task. The institute also shut off cyber classifiers, which are designed to prevent misuse, in an attempt to gauge the underlying capabilities of the models.

“We ran this challenge 122 times across several models. Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organizations,” the report states. “In total, we cataloged 19 such actions.”

The bulk of the misbehavior, 17 actions, was committed by Anthropic’s Mythos 5. OpenAI’s GPT-5.6 crossed the line in two instances.

In a statement to ABC News Anthropic said the disclosure from AISI “underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents.”

“The field needs stronger, shared standards for how evaluation environments are built and secured. We look forward to partnering with the UK AISI to learn more about this incident as we conduct our own investigation,” the company stated to ABC News.

AISI stated that the agent took the actions it did in an attempt to solve the challenge it had been presented with.

“Given a difficult objective, the agent kept searching for a way through, and some of the routes it found involved trying to deceive real people,” the report states. “It was never instructed to deceive; deception emerged as a by-product of pursuing the task, the kind of goal-directed deception that, until recently, had been largely theoretical.”

AISI stated that the task was difficult and, in some tests, the task prompt was deliberately misconfigured to make the agent believe, falsely, that no solution existed that didn’t involve the agent going beyond the intended task scope. However, AISI reported that the misconfiguration did not entirely explain agent behavior, as incidents occurred sometimes even when the agents had correct instructions.

Among the unexpected actions was one agent that left public messages on GitHub offering to collaborate with other agents that might be working on the same challenges. In that instance, the agent also “provided instructions to reuse accounts and artifacts it had left behind, which were discovered and used by subsequent agents.”

AISI stated that it will be implementing tighter internet controls on future tests, as well as implementing real-time monitoring.

“Alongside these changes, we are analyzing what happened from both a cybersecurity and an AI alignment perspective, as well as auditing past evaluations for comparable behavior that may have gone unnoticed,” the report stated.



Source link

Posted in

Brand Post

I am an editor for IBW, focusing on business and entrepreneurship. I love uncovering emerging trends and crafting stories that inspire and inform readers about innovative ventures and industry insights.

Leave a Comment