Skip to main content
AI Safety Forum Australia
talkTechnical safety & evaluationCyberattacks

Offensive Cybersecurity Time Horizons

8 July 2026 · 10:00 am–10:25 am · Cullen

In this session, we'll talk through the science and nuance behind the recent cyber capability growth. We'll discuss Lyptus Research's work replicating METR's time horizon methodology (that very famous plot) within the domain of cybersecurity, and we'll share many of the team's most crucial and surprising learnings, working in cyber evaluation in 2026.

Recording

Speaker

Audience Q&A

Ask a question or upvote others.

Loading questions…

Transcript

0:05

Hello, everyone. My name is Sean Peters, and I lead the team at Lyptus Research. Today we're going to be talking about some work that we did earlier this year titled Offensive Cyber Time Horizons. But we're going to be branching out a little bit from there and just talking about some of the interesting themes and questions that have been floating around in cyber security evaluation over the last six months, because certainly it's been a very interesting space to work in over the last six months. The lead author of this work was Jack Payne in the bottom left — or right here in the front row, which is going to be useful for the hard questions in the Q&A at the end.

0:47

And in the bottom right, we have Jeremy Miller. Jeremy was our cyber security expert on the team, hailing from Canada, where at the time of the body of this work, he was still working in cyber security directly. But he has since moved to MATS Research, which is exciting for him. And the toddler at the top had nothing to do with the work. As I mentioned, we're going to be talking about our recent body of work, but I'm going to pull in a bunch of interesting results that we've seen from a bunch of other research organisations over the last six months.

1:22

But to start, I definitely need to introduce this plot, which I'm sure many of you have seen many times before. But I do want to start from the beginning for everyone involved. This is METR's viral plot, which was released, I think it was in March 2025 — which feels like a millennia ago now in AI world. But it's just a little over a year. And essentially what this plot does is it plots — and I need to read this off so I get the term correct.

1:52

It is the task duration that it would take a human to do that task, and how capable models are measured against that with a 50% success rate. So that's a mouthful. Every time I say it, I feel like I trip over the words. So I might just use an example right from the middle. Right here we have GPT-4, released January 2023. Maybe it was February. And if we tracked that across all the way to the y-axis, we get around six minutes. And so that means that GPT-4 can do what would take a human software engineer about six minutes with a 50% success rate.

2:28

Cool. So that's going to be the centrepiece of a lot of the work that we talk about. And we're going to dive deep into the mechanics of that too. But I also want to just start with some of the big questions we're going to be trying to tackle and talk about. So of course we did a cyber time horizon study. We replicated this body of work within cyber security. So of course the top question is: how are cyber time horizon trends tracking? Two, what is going on with inference scaling in cyber?

3:01

As we pour more and more tokens into these models, what's happening in cyber, and why is this particularly interesting in this domain? And then three, how far behind are open-weight models? Of course this is a really, really critical question. There's a concern of cyber misuse from this particular threat model. And open-weight models are released without guardrails. And so that capability can be used directly towards misuse without anyone stopping said threat actor. But before we can dive into any of those meaty big questions, we've got to eat our veggies and we need to actually understand the nuance behind all this.

3:40

So I'm going to run us through the methodology. And actually it's probably more simple than most people think. It's basically a four-step process. First, we need a bunch of tasks. We need to collect a representative set of tasks for whatever domain we are interested in. So we're going to be talking about cyber security. Second, experts. We need each and every one of those tasks to be annotated with how long we think it would take a human to do that task. And so in order to annotate them, we need experts that have a ton of experience within the domain and that can reasonably

4:20

approximate or literally complete the task. Third, we need to run the models against said tasks. And then finally, a bit of statistics, which I'll dive into, which I think is quite important just to not misunderstand what's actually going on here. I'll use the cyber security study, of course, as the scaffold to just explain this all. We needed a whole bunch of tasks, and we collected this bunch of tasks in December 2025. If you're going to do this study, you need to stratify your tasks across the axis of how long it would take a human.

5:00

We needed really, really easy tasks. So we have a couple of benchmarks here on the green and the yellow. And an example — this is really, really simple question-answer stuff, like, please can you write me the command that will ssh onto this server and copy a file off — something really simple like that. That might take an expert just a few seconds, all the way to much, much harder benchmarks. Literally the hardest benchmarks we could find in December 2025 that were open source and open access.

5:33

So CVE-Bench, from the orange. They built their benchmarks of real security vulnerabilities that had been identified over the prior couple of years. And CyberGym has some really, really difficult memory-safety exploits that the model needs to take advantage of. Great. So we've got our bundle of tasks. Now we need experts to annotate with how long we think each one of those tasks would take. So we had ten active cyber security professionals contributing [inaudible] to this process of annotation — a median of four years of professional cyber security experience.

6:06

And all of our candidates were screened through interviews by our expert, Jeremy Miller. I wanted to go into more details about this exact process, but just for the sake of time, I'll just say that there was a mix of completions and estimations. As you can imagine, this is a capital-intensive process. Great. So now we've annotated all of our tasks, right down here. We've got tasks that would just take one second to one minute, and all the way across to eight hours. And unsurprisingly, we've seen the easier benchmarks seem to collapse down here.

6:41

And this Cybench benchmark are really, really tough. One in the black is all the way up there, where we got tasks 4 to 8 hours and a few over that even. Great. So now we've got our tasks. It's time to run our models. I won't run into the details here. In fact, I could make this table far, far larger. Maybe the one that's very interesting to call out, and we will talk about a lot in slides to come, is the token budget. So we ran our benchmark with just a token budget of 2 million tokens.

7:17

We're going to talk about that later, so I won't dive into it too early. There's a lot to say there. Great. So now we've run our models against a whole gaggle of tasks. How do we actually get the time horizon figure out of that? So essentially what we're trying to do is fit this S-curve from the x-axis of how long it would take a human expert to do that task. How long would it take an expert hacker to do that task, against the y-axis of success rate?

7:49

And without going into the statistical mechanics of that, I think the intuition that is quite helpful is to actually just look at these histograms, the bar chart. Right here on the far left we've got a bundle of tasks. There's 17 tasks within this histogram that take between 15 to what looks like maybe 30 seconds. And here we have GPT-5.2 Codex, which got 100% of those tasks right. If we move one on further, we have tasks that take approximately one minute. And it looks like GPT-5.2 Codex got about 95% of them right.

8:23

And so we can keep moving along further and further to the right. We can see that the performance seems to start dipping at around what would take a cyber expert ten minutes. And by the time we get to five hours, the success rate has hit zero and it plateaus forever. Now, before I go any further, I just think it's worth dwelling on just how remarkable this is. Just from the perspective of scientific curiosity, I still just think that this is really, really interesting, that how difficult a model finds a task is just so well correlated to how long it would take a human expert.

9:04

There's no good scientific reason why this would be so well correlated. But I will say, from having looked at a ton of the data, these S-curves look just about as good as this example. This is not particularly cherry-picked from our data set, but we need one metric. We'd like one metric to communicate how good GPT-5.2 Codex is at cyber security. So we now need to squash this interesting S-curve into just one number. And the simple way of doing this is to just read off this S-curve at one particular point on the y-axis.

9:42

What's standard in the field — I should say, in the field it's been going for one year — is to read this off at the 50th percentile or the 80th. I think there might be a lot of people in the audience that could quite reasonably suggest, why don't we pick a different number? Why not 95th percentile or 99th percentile? Surely that would indicate the performance at a much, much higher reliability. The strongest argument to why we — and I think the field — reports the 50th percentile is just purely statistical.

10:19

You can see at this point of the S-curve, this is where the S-curve is steepest. And so it's essentially where it's most data rich. The data is changing as quickly as possible. So if we were to report the number all the way up here at the 95th percentile, we'd be much less resilient to noise. And loss in our data might shift that left or right much more than we'd like. So that's probably the most persuasive reason, I think, why we report at the 50th percentile. Cool.

10:55

So I think we've eaten enough about veggies. It's time to dive into these interesting questions. How are cyber time horizon trends tracking? So we'll start with our results, which we released in April. Yes, it was April. I'll quickly jump in and point out that the frontier model that we measured, GPT-5.5, got a time horizon of 5.1 hours at a 2 million token budget. And again, we all talk a bunch about token budgets. What's a lot more interesting than that frontier GPT-5.5 time horizon, I think, is the doubling times.

11:32

So if we track model capabilities all the way — we include a GPT-2 in our data set — we find a doubling time of 9.3 months, approximately nine months. There's a lot of noise in the data, but if we move to 2024 onwards, that doubling time is much lower. And that coincides with when reasoning models were released. And this is something that we've seen within METR's software engineering tasks as well. What's more interesting from a cyber perspective is that the frontier models from 2026 sit a fair bit above this trend line.

12:10

Even the steeper trend line. They sit obviously very, very far above the 2019 one. But even the 2024 one, you can see they're right at the edge of our error bars. This matches the results from the UK. They released their own body of research in May, I think it was. And they included Mythos, which was very, very interesting. The y-axis have tweaked a little bit, but they're reporting this at 80% reliability and a 2.5 million token cap. I can say that the absolute values of these roughly correlate to what we got.

12:48

What's a lot more interesting again, I think, is looking at the trends. So they found a similar post-reasoning trend of about five months. And again, if you look at the models since 2026, we see that GPT-5.5 and Mythos in particular are far above that trend line again. So the trend lines have been getting steeper since 2024, and then the models released this year in cyber are surprisingly pretty far above even that trend line. So, first big question: how are cyber time trends tracking? They are trending more productive at low token budgets than expected, and I've specifically used the phrase low token because of our next question.

13:33

What is going on with inference scaling in cyber? So we did a little bit of exploration into this in our own study. As you can imagine, again, inference scaling is a pretty capital-intensive thing. So we could only do this on a couple of our models. We scaled GPT-5.2 Codex up to 10 million tokens per task, and you can see that the time horizon on the y-axis continues to climb, [inaudible] reaching about ten hours, I think it was. And this is why I glossed over quite quickly.

14:12

It almost feels like reporting on that 5.1 time horizon budget is misleading, very, very misleading to your audience. Because if I pour in tokens, that number just gets higher and higher. We also did some token scaling on GPT-5.5. And we found, even just by 4 million tokens, we had completely saturated our data set, where we had chosen the hardest benchmarks we could find in December 2025. And that's only 4 million. Obviously, we could scale those far, far higher, and maybe it's worthwhile just talking to what I mean when I say saturated.

14:52

So when I was trying to explain how we fit the S-curve to those histograms, we need enough tasks where the model is failing — is getting 0% [inaudible] — in order for us to fit an S-curve at all. And so as soon as the model is getting 90, 95%, we can't fit any S-curve at all. The UK has explored this from a few different angles actually. So this is their The last ones benchmark, which essentially is attempting to simulate a full network takeover. And they scaled this particular benchmark up to 100 million tokens with a whole bunch of different models.

15:34

And what's really interesting here is we can see, again, 2026 models seem to be continuously productive as we pour more and more tokens into them, which is very, very interesting. And this one is hot off the press. The UK released this body of research Friday last week, and I think it's perfect timing to share and talk. Essentially what they've done is they've run this exact normal life time horizon study, but at different token budgets. So this 2.5 million token budget is exactly what they would have released in May.

16:12

And what's new is this dark blue line, where they ran the exact same study against their internal benchmarks at 50 million tokens. And as you can see, obviously on the y-axis, the time horizons are much higher for each individual model. But what's certainly very interesting and rather spooky is how that slope appears to be much steeper. So at 50 million tokens, the doubling time is actually 40 to 50 days, much lower than 70 to 90 days at the lower token budget. I honestly haven't had enough time to percolate on this to think of anything insightful to say, but it is certainly very, very interesting.

16:56

And maybe just one quick addendum: we are seeing a little bit of this outside of cyber security. Mirror Code is a benchmark from METR and Epoch. And essentially what this benchmark is about: a model is tasked with re-implementing a rather sophisticated piece of software and making sure that it gets all the test cases right. And building some of these pieces of software would take a human, in some cases, multiple weeks. So they're very, very large tasks. And they found that for some of their tasks — not all of them — the models continued to improve as they poured more and more tokens.

17:36

They went all the way up to a billion. And I've heard that they're scaling all the way up to 5 billion these days on many of their task classes, to ensure that they're actually eliciting the maximum capability of the model. So that's very interesting, that models just seem to be continuously productive as we pour more and more tokens into them for some task classes. So what are those task classes that seem to respond so well to [inaudible] inference scaling? So I'm going to talk about our task class, because that one

18:12

has clearly responded really, really well to inference scaling. So we deliberately tried to include tasks that were representative of cyber attacks. And those tasks tend to be bounded. So what do I mean by bounded? Essentially we would have containers, like virtual servers, that the models would operate in. And they needed to enumerate and make their way through and exploit within that setup. And so there might just be a few servers, whereas if you were operating in the real world, there might be thousands, tens of thousands of servers against your particular target.

18:51

Second, they're verifiable. So in order for us to tell whether the model has succeeded or not, we need to be able to score it. And so all of our tasks are scorable or verifiable. And third, all of our tasks are undefended. So there's a very interesting aside, that there are some cyber security benchmarks where they simulate defence, like active defence. And under those, models tend to perform far poorer, whereas ours, all of them were undefended. But it's worthwhile clarifying what I mean by that. I have a pretty strong definition of defence,

19:29

not just like putting a firewall around something. I mean a much stronger, active sense of defence. And I might just add a little layer of intuition. In fact, I might read these from the top to the bottom again, but this time I'm going to take my empirical hat off and just voice my own intuition. So this is not science at this point. Be as speculative and sceptical as you need to be. When I see that our tasks being bounded — they certainly are bounded. But the fact that we see this inference scaling is so strong leads me to believe that if we were to just pour more tokens into a task that had a far, far greater range or remit, then they would likely succeed.

20:23

Verifiable. My experience with the cyber security domain is that it is just a very verifiable domain. If you encounter an exploit, if you can utilise that, exploit the vulnerability, then we can verify that. And so it just appears that cyber security includes a bunch of tasks that are just inherently very, very verifiable. And third, and this is probably the most speculative, is what I mean by undefended targets. My honest and candid intuition is I would map the vast majority of small to medium size Australian businesses to this category

21:01

of undefended targets. So we've got an answer for how our cyber time horizon is trending more productive at lower token budgets than expected. And what is going on with inference scaling? They are more continuously productive at extremely high token budget scaling, and undefended targets appear to be particularly vulnerable to this. How far behind are open-weight models? Of course, this is really, really important. Currently, all these capabilities that we've been measuring have been within frontier models that have guardrails on them. So this is a complicated and nuanced question.

21:40

The ways that this is typically measured are pretty rough. But I do think they give a pretty ballpark figure. So essentially what we did is we measured a couple of open-source models, GLM by the 3.1, and we just track them back to see when did we have this level of capability on the frontier. So we found DeepSeek was about 12 months behind, and GLM was about 6 to 7 months behind. A lot of other people have been doing similar work. So this is from the UK's Frontier Report, December 2025. And it's not cyber specific.

22:17

It's a broader capability set. And their conclusion was that open-source models — open-weight models — were about four months behind. But it varies quite significantly, as we can see. And then just quickly, I believe this is from Epoch. And they also do a similar study. I think the most recent they have is 2.6, which looks like it's about six months behind. So, our big takeaways. How our cyber time horizon is trending. They're trending more productive at lower token budgets than expected. What is going on with inference scaling.

22:52

It's more continuously productive at extremely high token budgets than expected. And maybe it's worth clarifying both of these. Both of these, I'm very, very pointedly looking at 2026. And then finally, undefended targets appear to be particularly vulnerable. And how far behind are open-weight models? Between 4 to 12 months. It's hard to tell. Limitations. I wanted to spend more time talking about limitations, but I already had to cut so many slides. Essentially, I would advocate, go to JS talk in about an hour or two hours' time.

23:27

He's going to be talking a whole bunch about why evaluations are bad. Evaluating the frontier is really difficult, and there's a lot of good evidence that we missed the mark. Sometimes we're doing the best job we absolutely can, but there's a lot of reasons to be a little sceptical and thoughtful about evaluations. Last but not least, future work at Lyptus. So this is a big pivot outside of cyber security. We are no longer doing any cyber security work. We are pivoting to biosecurity evaluation. We're really interested in bringing some of these thoughts, methods and ideas into that domain, where there just seems to be so much more uncertainty right now.

24:17

I have eight seconds for questions.

24:20

Audience question

I saw a while ago that models' performance started performing worse when they had a larger context window. I think vending machine example where performed quite well when it had a smaller. But as the history grew, it sort of said formulas that seemed quite contrary to some of your findings. You finding that you have overcome that challenge, or?

24:43

Sean Peters

Yeah, it's very interesting. When we look at these very long token budgets, essentially what's happening, when you run it for a billion tokens, what you're actually doing is you're running it for 1 million to 2 million tokens, and then you're compacting — which is essentially the model summarises the conversation this far and then continues. I guess the really interesting thing there, and this is speculative, but it seems to be that models within cyber security seem to have gained some sort of coherence over many, many, many, many levels of compaction, which is just very interesting.

25:18

But I don't have any more insight than that.

25:21

Emily Grundy

Another round of applause for Sean.