Skip to main content
AI Safety Forum Australia
talkTechnical safety & evaluationTechnical and institutional challenges

Closing the Safety Gap: Lessons from FAR.AI, and Australia's Opportunities

8 July 2026 · 9:30 am–9:55 am · Refectory

FAR.AI is a research and education non profit based in Berkeley, CA (USA). This talk will share some updates and lessons from our ~3 years of operation, and then touch on some opportunities Australia has.

Recording

Speaker

Audience Q&A

Ask a question or upvote others.

Loading questions…

Transcript

0:02

I'm co-founder of FAR.AI. We are a 45-person research nonprofit based in Berkeley. We started around three and a half years ago to tackle this emerging global problem. Today, I'm going to run you through how I think about risks and capabilities at the start. Then, what does my company do and what are some implications of this? And then at the end, some hot takes on Australia's closing window of opportunity. Maybe we can take them off the record. We'll see what we do with the recording. But I would encourage a lot of discussion around this.

0:33

We're all in this together. We need to figure out our pathway through this. First up, I'd like to thank everyone for joining. Unfortunately, I wasn't able to make it yesterday. But thank you to everyone for showing up and taking this seriously. I myself was quite sceptical many years ago. And I came into this field thinking that if this is nothing, so what? I've wasted a few years. If this is real, this is possibly the most important thing we can be working on. And this is the most important thing of our time.

1:07

And unfortunately, proven slightly correct here. So don't take my word seriously. Trust the experts. But also make up your own mind here. Secondly, I also don't claim to have all the answers, but it seems as though what we're doing or not doing does not feel like we're setting ourselves up for success here. As I understand, we've been all encouraged to read the International AI Safety Report. Can I get a show of hands? Who's actually read the report? Cover to cover. Okay, there's, like, four hands. Who's read the executive summary?

1:52

That's better. Who is in government? Who's in industry? Working for a for-profit and nonprofits? Okay, quite a lot. And academia. Okay. That's quite an even mix. And who has imposter syndrome? Okay. A couple. So how does my organisation FAR.AI relate to this talk? The report's core message is important, and shared with our primary motivator here. It's that capabilities are outpacing our technical ability to make these systems safe, and importantly, society's ability to manage these risks. And safety is also a public good, meaning that no organisation is incentivised to fund it or solve it alone, even though it will benefit everyone.

2:46

So this might not come as a surprise to everyone. I hope this was shared yesterday, but I do need to pinch myself sometimes that just how fast things are going. So we run a research nonprofit based in Berkeley. We're a stone's throw away from all the frontier labs, and we've worked with many of them and currently working with many of them. And yet we are still surprised just how fast things are going. And so what we're looking at here is the Epoch Capabilities Index, and the same story repeats itself.

3:19

A couple of years ago we were at near-random performance. Then we get to PhD-level performance. Then we get to expert-level performance. And then these systems are saturating the benchmarks. And it's also not just one lab. The leadership changes hands every month or every quarter, and open-weight models are not far behind. I want you to think about the implications of this for you, for your work and for Australia. And one reason I tend to dismiss the rapid progress — and I think a lot of people do, my family and friends are doing this — is because all these systems still seem very brittle, and it's easy to look at something that's not working and then dismiss the underlying progress being made.

4:01

So I also think it's worth noting these examples, and I should share what it is. These are all front-page news headlines of production databases being deleted by autonomously operating agents who are building everyone's systems. And the problem here is that for every example you see in the news, there's probably ten to 15 or 20 more that have happened that aren't being reported, that companies don't want to let you know that has happened. And this, to me, highlights that enterprises today and industry are facing a really tough problem.

4:33

They need to take risks that they shouldn't be taking, to either survive or succeed. And so these companies have paid the price here, but these are the environments that these companies find themselves in. And they have to deploy things way earlier than they're ready for. So what is my underlying concern here? Some of you look at these charts and see artificial general intelligence within years. That's AI that's better than you at almost anything. At almost everything, rather. Task horizons are doubling annually or seven months or every four months now.

5:08

And others see hype. They go, oh, no, these benchmarks don't translate into real life. My view is that you don't need to pick a side. In the next 2 to 5 years capabilities are still increasing each quarter and they don't show any sign of slowing down. And so the issue here is that capability is increasing, but the proportionate level of safety is not keeping up. And so this growing gap between what systems can do and how confidence in their safety is the reason that we're here today.

5:44

It's the reason my organisation exists, and it's the reason why we need the International AI Safety Report. And so Yoshua Bengio has this quote as well. There's been progress in training general-purpose AI models to function more safely, but no current method can reliably prevent even overtly unsafe outputs. That's not a great place we find ourselves in. So given this broadening or widening safety capability gap, what could go wrong? I'm not just show all the risk categories here. If you've read the executive summary, you will see that there's different ways to categorise these risks.

6:21

The key takeaway is that these risks stopped being hypothetical in 2025. When my organisation started three and a half years ago, everyone was like, oh, none of these risks are real. My AI system can't even talk back to me and can't even write a full sentence. Now we have examples of all of these risks in the real world. So in misuse, last September, we had a state-sponsored group use Claude Code to autonomously run a large cyber espionage campaign. That's pretty scary. In misalignment, we've seen how sycophancy and persuasion

6:53

has been implicated in murder, suicide or suicide cases. And that's horrific. And then in systemic risks, the challenge is it's sometimes hard to point to a specific example, but I think everyone can relate to the fact that we're filling the internet with AI-generated content and then training the next generation of models on this content. That's a feedback loop that's happening, and we want it and we're witnessing it, and we have no control over this. Cool. So, part two. FAR.AI is an organisation that myself and Adam Gleave co-founded.

7:34

We work to ensure AI is safe and beneficial for everyone. So we're a 45-person nonprofit based in Berkeley. We aim to facilitate technical breakthroughs on key open safety problems. We coordinate the growing field, and then we educate key stakeholders on risks and solutions. We remain independent, nonpartisan and position ourselves neutrally to be a convener between industry, academia, government and across different geographies. And we've worked with many organisations around the world, and I feel very privileged to be able to do that. How we see the state of the world and what assumptions we're making.

8:22

And I encourage you to make your own assumptions, because you need to figure out your own pathway through this. Current safety methods are insufficient. I think everyone might agree with that. Yoshua Bengio agrees with that, even with universal adoption. So, like, every company adopts every safety practice that we currently have. We still cannot get high confidence — greater than 99% certainty — that these systems are going to be safe. And that's 99%. This is nowhere near nuclear safety or aviation safety, just 99% confidence. We cannot get there with the current safety methods.

8:56

Number two, even if we did have good safety methods, the challenge is that universal deployment is hard. Often — and I'll share this later — the frontier labs all have different safety approaches, and they're adopting different parts of the stack. And so this leads to systemic risks. And this is a big problem. And then the third thing is that safety is complex. It's very unlikely there's going to be a single breakthrough. And new problems emerge as capabilities emerge. And we've watched this with the new Anthropic model — suddenly we're finding all these vulnerabilities in systems that have been around for 20 years.

9:36

So what are we doing about it? We're kind of a strange organisation. We do a lot of things. And I feel very proud of what we built. But it wasn't an easy journey to get here. But basically, we do technical research, or foundational research, structured in research pods to tackle these technical problems. We run our own research cluster, and we have a foundations team that manages this, and automations. This is a growing expense, with spending for a 40-person or 45-person team. And of that, 15 are technical.

10:15

We're spending over 4 million on compute, GPUs alone, in 2026. So that's quite a large cost. We conduct pre- and post-deployment testing for frontier models. We run events that convene the field. I'll share a little bit about that later. We have a co-working space in Berkeley. If anyone wants to come and join, please do. We run an academic grant-making program, and we advise governments around the world. I'll share some more details on that later. This slide probably has too much information on it, but I guess what I want to share here is that we run a full pipeline from idea to real-world adoption.

10:53

And so most safety orgs that exist today kind of do little sections of this. So one part we have tried to offer something that is full service. So a research lead will have access to a world-class operations team, a communications team to help write up and disseminate their work, a grant-making program to help grow the field and fund academic research, and an events team to help them find collaborators and build and convene the groups, and then existing relationships with the frontier labs and governments to help them get to placements and adoption.

11:25

And so we've been building this the last three years, and now it's time to kind of scale. We might be, I'm hoping will be, around 75 to 100 in the next 18 months. So I'm going to spotlight a little bit on our work. We do red teaming. And what that generally means is we did pre- and post-deployment testing. Engage directly with the frontier labs or via a third-party government institute like the EU AI Office or the UK. And so red teaming is not just about breaking a model.

12:02

It's a structured approach to learn how models failed to prevent this occurring in the real world. And so our red teaming is focused on the most dangerous capabilities. And then we work with existing developers to patch them and fix these gaps, and then educate policymakers on what the issues are. So, pretty proud to say a red teaming group has led to production features, fixes adopted by all the major labs. And then we've worked with all the governments that are involved in this at the moment. And we're part of the US consortium as well.

12:42

One thing to note about safeguards is that they're actually getting a lot better. So three years ago, it was trivial to break everything. But in the last year, we've seen a rapid increase in the difficulty to break these models. And so proprietary models are non-trivial to jailbreak right now, which is very pleasing to see. We still mostly find universal jailbreaks on all these models, but it is getting much, much harder. That is not the case for open-weight models. This is a tough graph to interpret, but it basically shows you that, using some basic prefill attacks, we can get universal jailbreaks on open-weight models trivially.

13:43

And for example, Gemma 4 was jailbroken 100% on most demands, including chemical and biological weapons. And you do not need to be an expert to jailbreak open-weight models. Often people are posting these on Reddit, which is kind of not cool. So I just wanted to highlight here the difference between proprietary models and open weight. I will race through our deception work, but I'm quite proud of our deception mitigation team. And so, for those who aren't aware, AI systems are deceptive. And this does cause risks. Our proposed method that we're trying to get adopted at frontier labs is mitigating deception by training with a probe as a penalty in reinforcement learning.

14:34

And this is a very promising way. We're currently scaling this up to near-frontier scale. And if we can show deployment here, it's going to be adopted, which is really cool. So a while ago, deception was really hard to point to. Now we have real-world examples of this. And I guess my ask to the world is that we treat deception with a zero-tolerance policy. If there's ever a case where an AI system makes a statement of something that is not true, we should treat this as a deal breaker and fix this immediately.

15:33

And then maybe this is just a shout out to my team. On Monday I was in Seoul and we just found out. So this was at ICML. This is a large machine learning conference. We just won an outstanding paper award, which is just really cool. So I just wanted to highlight that AI safety research is recognised at top venues and is still very valuable work. And this is our deception mitigation proposal. And so they got a top ten paper and thousands of submissions, which is really cool.

16:07

We also do government engagements. And this is where AI safety becomes a national security issue. We are currently contracted with the EU AI Office to do CBRN capability evaluations — that's chemical, bio, radiological and nuclear evaluations of these models and how much uplift they provide. The International AI Safety Report highlights this evidence dilemma that policymakers have. And how do they act under the uncertainty? This is exactly what we're contracted to do for the EU AI Office. And how do they enforce it? That office gets regulatory power in August.

16:42

And we also do a similar thing with the UK AISI for open-weight models. I'm going to race through because I want to get to my hot takes in Australia. But we do lots of events. Cool. Maybe some recommendations, actually, very quickly, in Australia. So progress is going to be cumulative here. I encourage everyone to take this seriously. And we need many world views. Don't just believe what I say. Make up your own mind. Listen to the experts. Read a lot. Make up your own decision here.

17:15

And the bottleneck in the field is not money, it's not good ideas, it's talent at the moment. And so if you ever wanted to get involved, now is the time. There is money flowing. There are many organisations hiring. I could point to hundreds of jobs personally, of people in the field that are hiring. Every week I get 3 to 4 emails from different founders going, hey, do you know anyone with this kind of skill set? There are many roles. There are many ways to help. And Australia things that are getting started as well, which is really cool.

17:49

So consider joining something or even visiting a co-working space to see what the vibe is like, see if this is a community that you could get involved with. I personally think that's one of the most rewarding things, to spend time working on an important problem surrounded by people who are smart, kind and talented. Intelligent and smart. Same thing. But they show up every day trying to do the most good they can. That's a wonderful place to be. And so even though I work ridiculous hours, I really feel good about what I do.

18:17

And I love the fact that my team also enjoy it. Thanks, everyone, for sticking with me through that. And I'm around all day and I'll share my details with the conference. Thanks.