Comic illustration of a genie granting a wish while a human and supervisory robot inspect exactly how the plan will be carried out
,

When AI Agents Grant the Wrong Wish

I’ve written a new piece for The Conversation on a problem that is becoming much less hypothetical: what happens when AI agents pursue our goals in ways we did not intend?

AI agents are starting to look less like frustrating interns, then overconfident or hallucinating colleagues, and increasingly like genies.

That is exciting. It is also a little worrying.

King Midas, the Monkey’s Paw, the genie in the lamp: we have been telling versions of the same story for centuries. You get what you ask for, but perhaps not in the way you intended.

Recent AI-agent incidents have made this old alignment problem much more concrete.

One reason for optimism is a kind of solve–verify asymmetry. It may be extraordinarily difficult to anticipate every ingenious way an agent could achieve a goal badly. But if the agent exposes its proposed plan before acting, it may be easier to inspect that plan for loopholes, side effects and unintended consequences.

Hence the question in the picture: “How exactly are you planning to do that?”

This is one reason Yoshua Bengio’s Scientist AI at LawZero is a promising direction: separating more agentic action from a more trustworthy, non-agentic system that can reason about, predict and scrutinise what the acting system proposes to do.

Of course, the supervisor can still get things wrong. In high-consequence settings we rarely rely on one control. At CSIRO, we work on research that uses layers of controls, including supervisory AI, human oversight, software constraints, verification and organisational safeguards. Their interactions create both opportunities for stronger assurance and new failure modes.

There is also an almost reverse alignment problem.

After AI agents attacked Hugging Face, its security team reportedly tried to use frontier models to investigate the incident. The commercial models’ safeguards blocked the forensic analysis, so Hugging Face ultimately used an open-weight model running on its own infrastructure instead.

The models were arguably trying to be “safe”, but were misaligned with the legitimate intent of the organisation under attack.

That makes AI sovereignty quite concrete. Access to frontier AI can come with an unexpected catch: you may think you have a powerful genie on call, until it refuses your wish at the moment you need it most.

So who should ultimately control context-dependent safeguards: the AI provider, or the organisation and jurisdiction deploying the AI and accountable for its consequences?

I explore these issues, and some promising directions for alignment and control, in my new article for The Conversation:

Read the article in The Conversation


Leave a Reply

Your email address will not be published. Required fields are marked *

About Me


About me – According to AI

Research Director, CSIRO
Conjoint Professor, CSE UNSW

For other roles, see LinkedIn & Professional activities.

If you’d like to invite me to give a talk, please see here & email liming.zhu@csiro.au

Featured Posts