Cover of Epistemic Safety in Scalable Oversight beside a photograph from the report launch at CSIRO’s Mixed Reality Lab.

Epistemic safety in scalable oversight

Getting the right answer is not enough. Recent AI safety incidents have highlighted one side of this: even when an AI is pursuing a benign goal, the actions it takes to get there can still be unsafe.

Our new work looks at a quieter version of the same problem. What if an AI reaches the correct result, but the way it selects evidence, explains its reasoning or frames the answer influences the person, or another AI system, overseeing it?

We worked with the Australian AI Safety Institute to explore this question in our new report, Epistemic Safety in Scalable Oversight, launched yesterday by Senator the Hon Tim Ayres, Minister for Industry and Innovation and Minister for Science, at CSIRO‘s Mixed Reality Lab.

The research looks beyond accuracy alone. It develops methods, benchmarks and research tools to detect hidden influence, and explores how oversight protocols and AI interactions can be designed so that systems remain useful while making undesirable influence harder to sustain.

For me, the broader point is that trustworthy AI is not only about what outcome it reaches, but also how it gets there and how it influences others along the way.

We are keen to work with government, industry and research organisations interested in applying these ideas to real-world AI assurance and oversight.

report link: https://www.csiro.au/en/research/technology-space/ai/Scalable-AI-Oversight


Leave a Reply

Your email address will not be published. Required fields are marked *

About Me


About me – According to AI

Dr Liming Zhu FTSE
Research Director, CSIRO
Conjoint Professor, CSE UNSW

For other roles, see LinkedIn & Professional activities.

If you’d like to invite me to give a talk, please see here & email liming.zhu@csiro.au

Featured Posts