Opening Keynote on International AI Safety Report
7 July 2026 · 9:45 am–10:00 am · Refectory
The latest international scientific assessments suggest that AI safety is no longer simply about preventing model failures, but about understanding and governing increasingly complex AI systems. This keynote explores emerging evidence on AI capabilities, risks and risk management, highlighting why evaluation, real-world evidence and runtime control are becoming central to safe adoption. It concludes with opportunities for Australia to build sovereign capability through AI safety science, evaluation infrastructure and trusted operational evidence.
Recording
Audience Q&A
Ask a question or upvote others.
Loading questions…
Transcript
Liming Zhu
Thanks very much for having me deliver this keynote. So I think rather than just giving an overview of the report itself, I would also look beyond what's after the report. And I had the privilege to serve as the expert, advisory panel members representing Australians for this report. And many of the experts in this room have contributed to that discussion. Feedback through reviews and Australian scientific contribution has been reflected in this report significantly. And Australia is one of the most active countries in delivering input into this report.
So the report is 220 pages or so. If I'm going to use the next 15 minutes to summarise, it probably is going to be unsafe for use of me, not of AI. So I'm going to just make a few remarks on the important things I think are interesting and identify what it means to Australia. So the first thing is this report scope is strictly, general purpose AI. However, I think there is a misunderstanding of general purpose AI versus the traditional general purpose technology, traditional general purpose technology.
You need to do additional domain engineering to make it useful for some use cases. But the general purpose AI we are seeing today is actually an aggregation of all the narrow purpose AI, which seemingly works out of the box. So it's an aggregation of narrow purpose AI rather than a general purpose AI which requires additional engineering. Additional engineering gives us additional understanding of the context, risk controls. Because of the lack of that, there are significant risks that the general purpose AI has compressed. The safety, control and understanding into something seemingly working out of the box.
Another one is the capability is really a jagged frontier that people are looking at. We see very rapid advance in domains that we know the answer once we see it, for example, you can ask it's extraordinarily difficult to factor a very large number into factors, but once you see the two numbers, you multiply them back. You immediately know whether the answer is correct or not. So there are many domains like coding, mathematics and some of the science problems in that domain. We see significant progress in those domains.
But in other problems that it is not very clear to know what is correct or not, whether it's aligned or not, it's still got a lot of jagged frontier, and the risks involved. The next one is really there's a lack of evidence. We are already seeing a lot of systems being deployed, but the data on the ground is not coming back fast enough, secure enough for us to understand the risks, to improve the systems. And finally, the risk management is moving towards runtime control pre-deployment assessment and evaluation is important, but it is very limited.
Because it is a highly intelligent system that once it is released into the real world there are things: context can shift, things can happen. So more runtime monitoring, control and the safeguards is going to be very important forAI safety rather than pre-release evaluation. So I'm going to talk about four things that I think what's emerging out of the report. The first is very clear. If the capability is moving outside the model capability increasingly above the model layer and many of the control and the risks, in the same way.
So first post-training (after the model being pre-trained). Post-training: give them system, give them, tools, additional data, the so-called harness around a model. It's really delivering the system capability boundary. The question is not is this model capable of pursuing something? But what are the systems capable of doing? This elicitation becomes very important. And so the opportunity actually for a lot of organisations and for a lot of countries, it's shifting from building the smartest model or the smartest the system to building the smartest system that elicits, verifies, and to control these capabilities.
The second interesting question I think emerging from this report is evaluation is becoming elicitation. In the traditional mindset, evaluation means scoring whether something is correct or not. But elicitation is the new form of evaluation. It is to discover what a powerful AI system or model can do. So the old question was often what benchmark score did this model get? But the new question is, under what special conditions can these dangerous or beneficial capabilities be elicited? This is a very hard science problem to discover the boundaries of very reliable capabilities, failures, or misuses.
So it is actually adversarial science. I am often called it calling it, because you have to actually design the benchmark in a unique way to test the reliable boundary, to deal with some of the potential deceptions that Minister Charlton mentioned earlier from the model. So elicitation methods are becoming actually both a source of capability and safety. And it could be Australia leading in this space. Again to this point: it is not a trade off between stronger beneficial capability versus safety. The same underlying methodology elicits very beneficial capability but also some dangerous capabilities.
We need to control. The third thing I want to talk about is trust will not come from transparency alone, and it needs to come from verification. What do I mean by that? I'm not talking about institutional trust, governance. I think, Kimberlee, we will talk about those things later. There are experts in governance to gain that trust in this room. What I'm talking about is exactly what Minister Charlton mentioned. To do proper governance, to have confidence, it relies on us seeing what the model is doing. So in the past we have been relying on looking into the model, asking the model to do a chain of thought.
Tell me what you are doing. Observing the reasoning traces the execution traces. However, we have seen strong evidence that many of the behaviours are shaped by evaluation. Once you put pressure onto an AI system to evaluate them, their behaviour changes. They are starting to hide their reasoning traces. They are going to tell you something, but they are actually doing another thing. So this is really becoming one of the fundamental challenges left up to the current evaluation and trust. By looking into the models superficially on what they do, can no longer give us the trust that is actually doing what they are doing.
So alignment deception is happening, and if you put too much evaluation and the stress testing showing your working, those workings are becoming a performance rather than what's actually happening. So verification, you know, trust but verify. Verification on what exactly has been achieved through a triangulation of methods is going to be increasingly important. Finally, as I mentioned earlier, safety is a runtime control and under partial understanding we do not fully understand it, but we can put things around the model to understand. So, a lot of the safeguards we put into the model itself are actually very easy to remove.
So there's this asymmetry between building model level safeguards versus removing them cheaply later. So a lot of the safeguards need to be around the model rather than inside the model. And we also need operational evidences coming back to inform the best way of doing evaluation, monitoring, observing control, recover and eventually learn from it. So AI safety requires very trusted access to data, to operational data, to operational evidence in the real world throughout the deployment lifecycle. One thing, this report was finished in February and the AI field is moving very fast.
For people who are watching the news, the UN has recently released a report two days ago, and there is a section on AI security and the system implications. And there is no AI safety working group or sections. So my colleague, Dr. Qinghua Lu from CSIRO is the only Australian expert serving as an independent expert on the UN panel. She would have delivered, a talk in this forum about the UN report, but she's actually in Geneva right now joining the Global Dialogue of this report. She leads the working group on AI security and safety, which demonstrates Australia's expertise in respect to this very important field.
So very briefly in the report, it highlights some of the key challenges: critical infrastructure and cybersecurity is very important, detection and verification to defeat the deception is very important, synthetic media and deepfakes, and also system resilience. Acknowledging that sometimes AI will behave, in the wrong way. But how do we recover from that? So the challenge is really about governing interacting AI systems in the real world rather than evaluating isolated models. So this is my final slide. And so what that leaves Australian was Australian's role in this.
As I mentioned earlier, I think there is an opportunity for AI safety to control above the model layer because it's always about the system around the AI system or AI model that gives you the assurance, the guardrails, the monitoring. As Minister Charlton said earlier, we need to build a national evaluation capabilities in this space AISI, ASD, CSIRO, many of the research organisations, industry and the government agencies are working collectively to have these evaluation protocols, safeguards, standards and some of the harness designs to control AI is undergoing work right now.
And finally, I think, create the sovereign evidence data sets. We need data, we need evidence to assure that they are safe. And those data are coming from real world deployment. Finally, I want to end by saying capability and safety now it really shares the same infrastructure. That's no trade-off from a scientific point of view. Discovering dangerous capabilities beneficial comes from the fact that it uses the same technology and techniques. So maybe the opportunity for Australia is to lead in the field of AI evaluation, elicitation and control.
Thank you very much.
