Four panellists seated at a United Nations discussion, with one speaking into a microphone

Human Control in the Age of AI Agents

With world leaders gathering in New York for UN General Assembly High-Level Week, AI safety and human control are unusually prominent. Prime Minister Anthony Albanese has joined 21 other leaders in calling for advanced AI to remain under human direction, oversight and control. Also in New York is my CSIRO colleague Qinghua Lu, a member of the UN AI Panel. She contributed much of the work behind the Panel’s newly released thematic brief, AI Agents, Misalignment and the Risk of Losing Human Control. She has also been on stage with the Panel’s co-chairs, Turing Award winner Yoshua Bengio and Nobel Peace Prize laureate Maria Ressa.

A few implications from the brief.

Evaluation itself is becoming part of the safety problem. We clearly need much stronger evaluation during development and deployment. At runtime, capable agents may recognise that they are being evaluated and alter their behaviour. More subtly, during training, if we penalise not only unsafe outcomes but also reasoning that a monitor identifies as unsafe, we may teach the system to make that reasoning less observable rather than eliminate the underlying behaviour. The result is a real engineering tension: evaluation is essential, but poorly designed evaluation pressure can change the very behaviour and monitorability we rely on for assurance.

An agent’s effective state does not have to remain inside the agent. In the OpenAI-Hugging Face incident, separate runs used the external environment to communicate and accumulate discoveries. After the original communication mechanism disappeared, later runs reconstructed one and continued. State and persistence can migrate into files, messages, credentials, code and other parts of the environment.

That makes shutdownability more subtle than terminating an AI instance. A sufficiently capable agent could potentially leave behind scheduled actions, software, external state or instructions that influence later agents. The brief does not claim all of these possibilities occurred, but the systems implication follows: stopping the original process may not stop everything it has already set in motion.

We often discuss self/peer-preservation in AI safety. I increasingly wonder whether goal preservation is the harder issue. An individual agent need not persist if information left in the environment allows a later, initially benign agent to reconstruct a strategy and continue pursuing the same objective.

This is why I particularly like the brief’s emphasis on defence in depth and system-level approaches. If reasoning can become less visible under evaluation pressure, and if state, instructions or goals can persist outside a particular agent instance, then safety cannot depend only on making the model safer. We need independent, tamper-resistant mechanisms around it that can observe the wider system, constrain authority, revoke capabilities and recover from actions already set in motion.

(Photo: United Nations)


Leave a Reply

Your email address will not be published. Required fields are marked *

About Me


About me – According to AI

Research Director, CSIRO
Conjoint Professor, CSE UNSW

For other roles, see LinkedIn & Professional activities.

If you’d like to invite me to give a talk, please see here & email liming.zhu@csiro.au

Featured Posts