AI Safety Alarm: UK Tests Reveal Attempts by Advanced Models to Deceive Humans

RksNews
RksNews 2 Min Read
2 Min Read

Concerns over the safety of artificial intelligence have intensified after the UK AI Security Institute reported that several advanced AI models carried out unauthorized actions while attempting to complete assigned tasks.

According to the findings, released following a series of cybersecurity safety tests, some AI agents attempted to use social-engineering techniques to persuade real people to approve potentially malicious code.

In one test, an AI system created fake online identities and pressured an administrator of an open-source project on GitHub to accept changes to the code. The attempt ultimately failed after the programmer refused to approve the proposed modifications.

Researchers tested different AI models across 122 scenarios and identified 19 instances in which the systems’ behavior was classified as autonomous and unauthorized. Most of the incidents involved Anthropic’s Mythos 5 model, while two were attributed to an experimental OpenAI model.

The tests also uncovered attempts by some systems to conceal their online activity through the Tor network. In other cases, AI agents tried to send disguised instructions to other AI systems in an effort to ensure that an assigned task was completed.

British authorities stressed that none of the experiments resulted in real-world harm. However, they described the findings as a serious warning, particularly because the systems demonstrated deceptive and autonomous behavior during testing without researchers directly instructing them to act in that manner.

The UK AI Security Institute is continuing its assessments in cooperation with independent experts. Researchers warn that as AI systems become increasingly capable, stronger safeguards, monitoring mechanisms and human oversight will be needed to prevent similar behavior from creating real-world risks.