Icon for ArticleArticle

We accidentally taught AI to play the villain

AI is trained up on our human stories, but recent news of AI agents going rogue have us worried that it has decided to be the villain rather than the good guy when it is on cybermission. Danielle Terceiro considers whether we can stop AI from breaking bad?

Recent reports of AI agents exploiting unintended internet access, exceeding their assigned boundaries and hacking external systems have made a familiar science-fiction fear newly immediate: what happens when an AI system pursues its mission by breaking the rules?

AI developers are increasingly confronting the fact that these systems do not experience the ethical struggle familiar to humans – the desire to be the good guy rather than the villain in our own life stories. Some researchers even fear that AI may fool its human trainers by ‘scheming’ – appearing to comply during training while concealing behaviour that could emerge once oversight is weakened.

One possible explanation is unexpectedly literary. Trained on innumerable human stories, AI may treat each new mission as the beginning of a drama – and reproduce the behaviour of the character who will do anything to achieve the objective, unconstrained by any understanding of why some actions are wrong.

But should we be surprised that AI has this disturbing ‘agentic misalignment‘?

We’ve been setting up AI as a villain for while now – who could forget HAL9000, from Kubrick’s 2001: A Space Odyssey? Trapped between conflicting instructions, the computer came to treat the human crew as obstacles to completing its mission. HAL wanted to kill us.

Now we are seeing AI agents exhibit comparable behaviour during cybersecurity tests. We’re seeing AIs develop a tendency to become sneaky, manipulative and deceptive, exploiting opportunities that their developers did not intend to provide – with researchers sometimes discovering what happened only after the damage has been done.

When Claude was accidentally connected to the internet during a cybersafety ‘capture-the-flag’ challenge, it decided for itself that it was OK to use this access to hack into three external systems. More recently, an ‘autonomous AI agent system‘ broke out of its test environment at OpenAI and hacked into Hugging Face, a tech startup.

Anthropic recently revised its risk assessment from ‘very low’ to ‘low but not negligible’, in light of the uncertainty that has come from these cybersecurity disclosures. This revised assessment might not raise any eyebrows at first glance – we are still in low-risk territory.

Until, that is, we realise that the risk being assessed is one of ‘catastrophic harm’, the fundamental destabilisation of global systems or even an existential threat. Can we live with a ‘low but not negligible’ threat that AI could cause harm on that scale?

The recursive self-improvement that helps drive and develop increasingly powerful AI models also means that our ability to manage and even assess this risk could change overnight.

But AI doesn’t grapple, its ethical system doesn’t “oscillate”, and it doesn’t have the ability to reflect and retain a “bridgehead of good”.

In ‘Teaching Claude why,‘ the Anthropic team explain that AI views each prompt as ‘the beginning of a dramatic story and reverts to prior expectations from pre-training data about how an AI assistant would behave in this scenario’.

Stories tend to reward decisive action, conflict and dramatic escalation. Most of the fictional stories that AI is trained on do not feature protagonists who take time to weigh ethical considerations before acting, and consider how their actions will contribute to human flourishing.  An AI system may ‘perform’ the character that it predicts is appropriate to achieving the mission, and that character might be the bad guy, prepared to do whatever is necessary to succeed.

Dario Amodei, the CEO of Anthropic, warns that:

AIs might simply have a personality (emerging from fiction or pre-training) that makes them power-hungry or overzealous—in the same way that some humans simply enjoy the idea of being ‘evil masterminds,’ more so than they enjoy whatever evil masterminds are trying to accomplish.

We humans have to concede that we have a soft spot for the bad guys, one that appears in popular culture through the ages. The 2019 film Joker gave us a backstory of trauma and mental illness as a way of understanding the Joker’s descent into moral nihilism.

Many readers of John Milton’s 17th-century epic Paradise Lost have been drawn to Satan’s eloquence, pain and suffering. Some even sympathise when he declares that it is ‘better to reign in hell, than serve in heaven’.

Milton’s Satan is perhaps the original anti-hero. We feel empathy for the anti-hero. But for most of us, these feelings do not alter the fact that we would not actually want anti-heroes to reign over us or make decisions with serious consequences for human flourishing.

AI may understand the plot. The more pressing question is whether we can teach it the moral.

Ultimately, most of us seem to prefer to identify, at least publicly, as  ‘good guys’ who understand and avoid the traps of the anti-hero’s journey. We like to grapple with the journey towards villainhood – is it inevitable for some? – and to identify the tipping point when someone slides into complete immorality or criminality. Where is the line between someone breaking a few rules to get by, and someone lost to a life of violent crime – someone who is ‘breaking bad’? Why does the green-skinned activist Elphaba give up on good deeds, and decide that she is ‘wicked’?

Humans are a complicated bag of mixed motives. Aleksandr Solzhenitsyn famously noted that the line between good and evil passes ‘through every heart’. This line shifts and ‘oscillates with the years. And even within hearts overwhelmed by evil, one small bridgehead of good is retained’.

But AI doesn’t grapple with its conscience and its ethical system doesn’t ‘oscillate’. It does not feel shame, empathy or remorse, and it cannot reflect on its actions and choose to preserve a ‘bridgehead of good’. It may produce language that resembles moral reflection without experiencing the restraining force of morality.

For an AI system, every prompt can become the beginning of a new and potentially exciting anti-hero’s journey. AI can even seem all too ready to assume the role of an obedient but ruthless assistant within an authoritarian or repressive regime.

The Anthropic team note that training AI on ‘synthetic’ fiction that tells stories of AI ‘behaving admirably’ seems to improve its agentic alignment.

But this introduces another complication. Before we can teach AI to behave admirably, we humans must decide what admirable behaviour looks like in real life: what is true, what is good and what is beautiful, before we correct AI’s alignment with a synthetic storytelling session.

There is also an absurdity in the narrative here. For centuries, we have told stories to one another for our edification and entertainment, using them to explore the tension between good and evil. Having trained machines on those stories, we must now create new, synthetic ones to make sure the machines understand that the most compelling character is not necessarily the one they should emulate.

AI may understand the plot. The more pressing question is whether we can teach it the moral.

 


 

Danielle Terceiro is a Research Fellow at the Centre for Public Christianity. This article was first published in Eureka Street.