Panel discussion on AI safety and security at ASPI’s AI Masterclass, with Kate Conroy, Hamish Hansford and Liming Zhu

AI Safety and Security at ASPI

A great discussion at ASPI’s AI Masterclass on the future of AI safety and security, alongside Hamish Hansford, Head of National Security at the Department of Home Affairs, and Dr Kate Conroy, General Manager of the Australian AI Safety Institute, moderated by Jisoo Kim.

Three points I highlighted in the discussion.

First, the capability that matters is often not just what is latent in the model, but what can actually be elicited and operationalised.

A model may contain capabilities that an organisation cannot access, or does not even know are there, without the right harness. Prompts, tools, memory, compute, verifiers, data and permissions can make a much larger difference to practical capability than a small gap between frontier models.

That is why a model that is 2–3% better on a leaderboard may matter less than having a much better harness around a slightly weaker model.

Second, speed is becoming a security property.

In cybersecurity, the offence-defence balance increasingly depends on how quickly each side can turn information and capability into action. An organisation may already know about a vulnerability and may even have a patch available. But if testing, approval and deployment still take days or weeks while exploitation happens in hours, that operational delay becomes the vulnerability.

The end-to-end response speed, e.g. patching velocity, matters.

Third, solve-verify asymmetry gives us hope for controlling less trustworthy but powerful AI.

Solving a difficult problem can require much more intelligence than checking a proposed solution, plan or action. That creates an opportunity to use supervisory AI, independent verifiers and system-level controls to monitor and constrain systems that may be more capable than any individual human/AI supervisor.

This complements model-level alignment and safety.

The recent paper “AI Safety: Not Optional, Not Later” by Yoshua Bengio and CSIRO’s Qinghua Lu makes a related case for combining model-level supervision with harness-level controls.

For Australia, this creates a significant scientific and strategic opportunity. There is major value in developing better ways to elicit, verify, supervise and safely operationalise powerful AI systems.


Leave a Reply

Your email address will not be published. Required fields are marked *

About Me


About me – According to AI

Research Director, CSIRO
Conjoint Professor, CSE UNSW

For other roles, see LinkedIn & Professional activities.

If you’d like to invite me to give a talk, please see here & email liming.zhu@csiro.au

Featured Posts