Safety for the First Billion Agents
8 July 2026 · 11:49 am–11:54 am · Cullen
Current AI safety is not enough to prevent extreme power concentration, gradual disempowerment, and resource contention with social-scale agentic systems. Safety needs a new approach which acknowledges the unique challenges and tooling of large agent systems, from efficiently scaling existing techniques to developing infrastructure which adapts to its participants. I will discuss new demands of large agent systems, compare nascent approaches, and note barriers to progress. I will conclude by sharing a community and centralised resources for interested researchers.
Recording
Audience Q&A
Ask a question or upvote others.
Loading questions…
Transcript
I'm here to talk to you today about AI safety for the first billion agents. So I guess the context of this is the Remote Labor Index is a good example of improving AI capabilities on real economic tasks. Gemini 2.5 Pro in March of 2025 did 0.8% on this benchmark, and it measures general ability to do remote work tasks. And Fable 5 that just came out in June did 16%. So that's a 16 times improvement in 15 months, which is pretty crazy. Looking forward, we can imagine a world where agents are doing a lot of economically productive work, maybe a lot sooner than we think.
So humans currently have designed a lot of systems — like economies and financial markets, I guess social networks as well, to some extent — to aggregate people who otherwise wouldn't be able to connect and cooperate at scale. As agents start to do a lot of functions that humans were previously doing, whether that's economic or social or just things that were productive for humans, we're going to need systems for agents to cooperate at scale as well. Safety needs a systemic approach is essentially the change here. We're moving from single- and few-agent safety, which is the traditional concerns, to something more like a macroeconomics or a financial systemic stability of AI agents.
We can imagine very easily coming from these systems, like financial crises or unstable markets, where the production of goods and supply chains dry up due to volatility in agent behaviour. I think we all understand. The current generation of agents are just not prepared for economically valuable tasks, and if we exist in a world where our systems are totally integrated with agents, there are agents next to humans doing work, this is quite a scary prospect, I think, given the current quality of the technology. So, my call to everyone here is to consider this new large agent system problem — or just large multi-agent system, but we remove the “multi-” so as to avoid confusion with a few-agent case.
Large agent systems. We're talking now about mechanisms of interaction for scalable cooperation. We're talking about emergence of behaviour at the scale of 1 million or 1 billion agents, which goes far beyond the problem of a few agents or a single agent in complexity. It builds on this single- and few-agent work, but it's concerned with things like macro policy and macro models, the types of causes of agent behaviour that only emerge from very large collections of agents. And the third pillar of that, alongside mechanisms and emergence, is regulation.
So there are very few system operators, right. There's just, point blank, very few human systems at the scale of a billion agents or even tens of millions, hundreds of millions of agents. There's economies and markets on those economies, and there's financial systems and there's social networks essentially. So there are only a few operators here who have leverage over these systems. And I think the sort of very fast emergence of this problem means that we need to start talking to those people, whether that's financial market operators or regulators, and bring this to their attention.
The field also needs to move beyond what it has now in its strong conceptual work, in things like gradual disempowerment and power concentration, into measures and controls or steering mechanisms for these large systems. And we also need to start thinking about how to bring talent into the area. So naturally, a lot of these problems intersect with social sciences like economics or market design. And I think it's an extension of those areas or an augmentation of those areas. But fundamentally the concerns are very similar. So we need to start thinking about how do we engage with academics, not only regulators but also academia, to bring that talent in and communicate the risk and understand how it changes.
So the first step towards that, I guess, that I'm taking is the foundation of my group, where we're working on empirical problems like simulations of gradual disempowerment processes. But for the community more generally, I've started a Slack community. So if this is something that you're interested in, I'm going to put the link (https://largeagentsystems.org/) in my description on my profile, and I would love for you to join. There's a bunch of resources, organisations working on this group, and we're bringing in more and more people who have done work in the area so that we can get everybody in one place and start talking about these problems.
Thank you.
