Skip to main content
AI Safety Forum Australia
keynoteFraming & overviewCurrent capabilities (incl. agents)

Opening Keynote on AI Science

7 July 2026 · 10:00 am–10:15 am · Refectory

This session presents key facts about the pace of AI capability progress, to add to the background context for the forum's discussions.

Recording

Speaker

Audience Q&A

Ask a question or upvote others.

Loading questions…

Transcript

0:04

Tiberio Caetano

Before I start, I have a question. And the question is. And of course, before I ask my question, I want to start by flagging that this presentation is related to the specific question around the current capabilities of powerful models. Okay. So that's going to be the scope of what we are going to be talking about and the speed of the acceleration and how is that is all going. Now the question I have for the audience is, you know, show of hands, if anyone here who was present at the 2024 AI Safety Forum, okay, that's a good number of hands up.

0:46

You might remember if you attended the talk I gave then that I had this slide up at the time. And, and I was thinking when I was preparing yesterday, very depressed because as a Brazilian, Brazil lost to Norway and, and I couldn't do any work yesterday, but I still had to prepare this slide. So at night, what do I do here? And then I thought, okay, so what is the corresponding version of this for where we are now. And this is where I came up with.

1:27

I well, look, if you have one thing to take away from my talk, just take this, okay? But now let's go to the serious stuff. Well, as we've seen, the event is grounded, is framed around the International AI Safety Report. This is the 2025 report. And look at the phrasing there. Look at the words there okay. So experts disagree about the pace of future progress. Companies are exploring an additional new type of scaling. This was about, you know, 18 months ago. It was early, early days of reasoning models when OpenAI put out o3 and things started to shift.

2:15

Now, AI Safety Report this year from February. Okay. Look at the language. General purpose AI capabilities have continued to improve new techniques that enhance performance after initial training. Okay. So we are going to look into that now. What happened in that period was very significant. As Liming just pointed out we've got a gift from the UN just a few days ago and a new report, very legitimate, it's from the United Nations. You know, our friend Qinghua is there helping ensure that this is actually done well.

3:01

So here we have it. Look at the language. Now the field is advancing rapidly towards orchestrated agentic networks and self-improving systems. Now I just want you to, before you look at the numbers that I'm going to show, just look at the change in the rhetoric here, changing the language from just like 18 months ago, a bit more than that to where we are now. So what this talk is about is providing a few numbers and graphs to actually ground us on the science a little bit.

3:35

It's too big. I don't have time to go through the details, but I just want you to give a sense of the speed and the acceleration of what's going on. And in order to do that, it will be helpful to source a few other highly legitimate data sources here. I'm using no direct data from any lab. I'm using data from the international reports, from the UN report, as well as from the UK AISI here. Well, let's look at the evidence of capabilities. This here shows you that the evidence there is evidence, concrete evidence of capabilities have accelerated

4:16

over the past three years, up to about two years ago and up until about the end of 2024, which was when we held the first event here, the first forum. Capabilities were advancing predominantly due to pre-training, increasing compute in pre-training, starting around, you know, almost two years ago, things changed, okay. And the rate of progress changed. So it's been accelerating for a long time. This is an exponential growth. But the exponent now is different. So you have evidence on the left hand side for a collection of types of capabilities.

4:55

This is an index that Epoch AI (if you're not familiar with this nonprofit you should check it out, it's an absolutely extraordinary work that they are doing) and what on the left hand side you can see, is the result of an index that is combining, results and capabilities and performances across a range of different types of capabilities, at least on the left hand side there. On the right hand side. I will tell you a little bit more about this, because it is super interesting. It’s basically trying to assess, you know, for how long, an agentic software system can successfully accomplish a certain task at the similar level as an expert or a human right.

5:37

So, that type of measure is also increasing. That's the number of time, the number of hours that it takes to complete successfully to, where success is defined at 80% of in this case. So bottom line of this slide: so over the last 18 months there was an acceleration in the acceleration okay. So now let's again, this is spanning the entire period roughly from since the first forum. What did we have? We have that about what you know, the estimate from February this year, both by international report and by the AISI, was that these horizons (with which you could run these automated agentic workflows and they would be successful according to a given, the threshold of, of success, in this case, 80%) have been increasing, increasing exponentially is being doubling every certain number of months.

6:48

And this doubling period was estimated around eight months in the beginning of this year. And just a few months ago this doubling period is being sort of re-evaluated with data only from the past year, you know, including the latest models that came (but this was still before Fable and so on). And that shrank to 4.7. Okay. So if you know anything about exponents: so this is an exponent. Okay, so, here's another data point. This relates very nicely with what Liming said before. There is evidence not only that acceleration is

7:36

accelerating, so to speak, but that the measurement of the acceleration is actually underestimating the real acceleration. And that's related to elicitation. Here's what the UK has done and gave us this gift just on the 2nd of July, which I believe was five days ago. And, well, since I did this last night anyway, I might just as well include this one. So, here it is. Look at what's happening. It's basically saying that if you use 2.5 million tokens here in the evaluation of your capability, you're going to see an estimated doubling period that's going to be around, you know, between 2 and 3 months.

8:32

But if you actually use more compute for the same model (so you're just basically putting more money and more compute into your evaluation) you are going to discover that you are going to end up with, faster, doubling times. So the takeaway message from this slide is very important. And maybe one way to read it is from what's on the left hand side here. If we keep treating capability as a fixed score rather than a curve over compute, we will keep being surprised by what these systems can do.

9:19

Okay. Another data point that's showing that not only things are advancing very fast and accelerating, but also, our very assessment of how much they are accelerating may actually be an underestimate. Okay, look, enough numbers. Just a few more things to say here. Anyone seen the results of solving these open conjectures by Erdős? Yes, some people have seen. So basically what we have in terms of capability so far, we just talking about advancement in capability is all how fast these, especially agentic systems with verifiable feedback loops, as Liming alluded to before are actually advancing:

10:19

it's extraordinarily fast. And this has completely astonished mathematicians, you know, in. Whoo hoo hoo hoo, who's watching the Tim Gowers video, which was very sort of impressive, and so on. Well, look at people. You know, serious mathematicians, seriously astonished by what's happening. So the level of capability here is, is something that is really at the very edge and exceeding the even the most aggressive estimates, even from the previous reports. So that's all good and fine, you know, lots of capability. You know, the word capability. Usually you go out in the street and you talk to someone, you know, capability I mean, it's a good thing, right?

11:08

You're capable person. You know, we are building capability. You know enterprise capabilities are good. But of course, I mean, in this audience, we all know that this coin has two sides, right. So, what does this all have to do with safety? And it's all very simple. Capabilities relate to safety in this way. So you start with capability and then you go all the way to harm. And then there are safety measures that you can put in between. You can design better capabilities. You can intervene in model training, but you can also intervene later on in the process.

11:45

So it's key to understand that capability is a leading indicator of harm, as was very well put earlier by Minister Charlton. You know, we cannot wait for many of the harms to materialise. We need to do preemptive work and start working now on the leading indicators. Harm is a lagging indicator of capability. Okay. By the time you have harm, you know, it's kind of of course for many things already too late. Of course not yet too late because you have to you can still do something,

12:16

but you don't want to wait for the harm to turn up. And what is risk? Risk is the realised harm of the compounded probability of progression throughout this entire chain. Okay. And that involves not only the extent to which the capabilities are strong and in progressing, but also how good are the safety measures to block, you know, the flow of, you know, dangerous capabilities forward. So to be a I'm just putting an example here on voice cloning, you know, thanks to Greg Sadler for pointing this out to me the other day.

12:58

I mean, if you look at voice cloning, like, what is the capability of a voice cloned from a few seconds of public audio? Okay, what is the threat? Fake emergency and kidnaping scams targeting families. What is the incident? Mother hears child sobbing. Ransom demand believed. What's the harm? Well, in this particular case, the payment was averted. But there was still lasting trauma. And, you know, lack of trust and all this. What are we doing here? You know these things are coming so fast and we don't know what to do.

13:37

Now, the bottom line here is what you've got in green here. We already knew the capabilities that when actually put together would end up with a voice cloning capability. And here are they, you know, look at the papers there to check for yourself and then scams at scale. So the lesson clearly is that we need to avoid making the same mistake we did with synthetic media generation here. And this caught us by surprise. So those long term risks or concerns and so on, you know we need to be careful.

14:13

Let's work towards preventing those harms. They are not yet here, but they are the leading indicators. Just my final slide here with some agentic risks that we want to prepare for. And as Minister Charlton mentioned, the report we are working with, the Gradient Institute is working with the AISI, and should be out soon. And we touch on this a little bit there as well. You know, there are a wide range of risks that are very much directly enabled by these accelerations in the agentic software capability

14:49

I described before. And we need to start to prepare seriously for those. And this is not a matter of, let's sit and wait for something bad to happen. We need to start acting yesterday on these things, you know, from society. We need to act as a group, we need to get our act together on this. And just to finalise, there are increasing rates of acceleration. It's not just acceleration. It's important to understand that. So if we don't bring the right safety measures (and when I talk safety measures it’s not just technical, it's institutional, it's human, we need human preparedness as well) then, you know, it's certain that we are going to be faced with surprises.

15:50

So the future is shrinking. Okay? The future: don't think in terms of years or months or think in terms of, you know, now. Now. There's only now. So I am a Buddhist practitioner. And that's one thing that we are learning in Buddhism is that there's only now and it's always now. So, we need to act now.