When AI Agents Grant the Wrong Wish
What happens when AI agents pursue our goals in ways we did not intend? A new piece on alignment, supervisory AI, layered controls and AI sovereignty.
What happens when AI agents pursue our goals in ways we did not intend? A new piece on alignment, supervisory AI, layered controls and AI sovereignty.
A framework for where Australia can lead in the AI stack, from post-training and domain grounding to supervisory AI and trusted infrastructure.
Final reflections from the AI Impact Summit on verification-first workflows, post-training learning, cross-domain synthesis and Australiaโs opportunities in safe, system-level AI.
Day-one reflections on sovereign AI strategy, Indiaโs layered digital infrastructure and the durable meta-skills needed beyond short-lived prompting techniques.
Questions brought to the AI Impact Summit about system-level capability and risk, the durability of AI engineering advantages, and access to inference-time compute.
Why the real leverage of coding agents comes from redesigning work around automated, scalable verification rather than accelerating workflows that leave humans as the bottleneck.
As part of the international expert advisory group representing Australia, Iโm pleased to see the second Key Update of the AI Safety Science Report, led by Yoshua Bengio, released (link in the comment). The official update captures the broad progress: more robust adversarial training, better AI-generated content tracking, wider adoption of frontier safety frameworks, and…
I often share missed or surprising observations in my talks. Here are five of them, summarised as food for thought. ๐ ๐๐๐ต ๐ญ: ๐ข๐๐ฒ๐ฟ๐๐ฒ๐ฎ๐-๐๐ฟ๐ฎ๐ถ๐ป๐ฒ๐ฑ ๐๐ ๐๐ผ๐ปโ๐ ๐ฎ๐น๐ถ๐ด๐ป ๐๐ถ๐๐ต ๐๐๐๐๐ฟ๐ฎ๐น๐ถ๐ฎ๐ป ๐๐ฎ๐น๐๐ฒ๐.Itโs intuitive to think globally trained models canโt reflect our norms. Yet when frontier models answer the same cultural questions posed to national cohorts, they align most…
Every confidential message youโve sent may already be stored โ waiting to be unlocked. This is the essence of the โ๐ต๐ฎ๐ฟ๐๐ฒ๐๐-๐ป๐ผ๐, ๐ฑ๐ฒ๐ฐ๐ฟ๐๐ฝ๐-๐น๐ฎ๐๐ฒ๐ฟโ problem: data intercepted today could be decrypted once quantum computers mature. This week CSIRO’s Data61 released our new report, ๐๐ช๐๐ฃ๐ฉ๐ช๐ข ๐๐๐๐ ๐๐ง๐๐ฃ๐จ๐๐ฉ๐๐ค๐ฃ: ๐๐๐๐ก๐๐ฉ๐ฎ, ๐๐ช๐ง๐๐ก๐๐จ ๐๐ฃ๐ ๐๐๐ฉ๐๐ฌ๐๐ฎ๐จ, mapping the scale of the challenge and…
What counts as a ๐ฑ๐ฒ๐ฐ๐ถ๐๐ถ๐ผ๐ป, really? From the strict sense in ๐ฎ๐ฑ๐บ๐ถ๐ป๐ถ๐๐๐ฟ๐ฎ๐๐ถ๐๐ฒ ๐น๐ฎ๐โwhere a decision formally alters someoneโs rights or entitlementsโto the everyday sense where almost everything, from selecting a movie to choosing a candidate, feels like a decision. Yet across these very different contexts, the underlying anatomy looks surprisingly similar. In CSIRO’s Data61 new…

Research Director, CSIRO
Conjoint Professor, CSE UNSW
For other roles, see LinkedIn & Professional activities.
If you’d like to invite me to give a talk, please see here & email liming.zhu@csiro.au