Bottlenecks in AI Safety - Opportunities for Impact
If you are competent, driven, and can get stuff done, you can help unblock this.
Cross-posted from my Substack. Written primarily for people in competitive industries who are concerned about AI safety and wondering whether they are suited for having impact, and if so how they can go about entering the field. Any critiques of the substance or arguments here are welcomed.
TL;DR
Why this post
I spent the last 2 years as a quant trader. This summer I did MATS (an AI safety research programme) while on sabbatical. I spent most of the programme expecting to go back to trading. Then the Hugging Face incident happened, and I updated quite hard on a lot of things and decided to work on reducing catastrophic AI risks. This post is written for competent, driven people working in competitive industries who are concerned about these risks - to argue that they are set up for having high impact given the state of the field right now, and to provide an outline of how to actually go about doing this.
Primary arguments
One objection I've heard from various people lately: "Catastrophic AI risk in the next few years is maybe 10–30% (I updated a lot after the Hugging Face Incident). But I have no experience, and large N (thousands) of people already work on this, so my expected impact is roughly 1/N. I'm comfortable and well paid. Not worth it."
The 1/N approximation is wrong. There is no conservation law making individual counterfactuals sum to the group total.
Your counterfactual impact depends on the structure of the field, and right now:
- The field is not constrained by money, talent, or ideas. Indeed, substantially more funding is about to come online as Anthropic IPOs and DAF money unlocks.
- The field is short of competent, agentic people who can either onboard and work fairly independently fast, or absorb and utilise those resources: founders, team leads, people who can get good work done with high independence, and can build. Eg MATS is scaling ~2x per year and is bottlenecked on organisational capacity - not funding, not mentors, not technically talented applicants.
- My rough estimate is: ~400 people per year attempt to build new safety orgs, against 1,000–5,000 working on AI risk in total. The former is fairly neglected, meanwhile the implied demand (resource overhang and explicit push for more founding) is high.
If you've done good work in a competitive industry - run things, built things, worked under pressure and with high independence - you are differentially suited to have impact right now.
Exploration is cheap: ~30 hours to get properly informed about risks, maybe 60hrs to do substantial further derisking such as a small project, 6–12 months to test working in the field more seriously, and something similar to your old job (or its equivalent) is still there if it doesn't work. The downside of waiting is that by the time the risk is legible enough to convince you, it may be too late to act on it.
How to read this post
- Broadly convinced already and wanting to know how to get started → skip to Problems to Target and What Next.
- Unsure of impact or state of the safety field right now → 1/#ppl is wrong, State of the Field, and Supporting Arguments.
- New to AI risk arguments entirely → start with the info sources at the end, then come back.
If you want to talk, you can ask a question with your name and email here.
Why am I writing this
As mentioned above, I recently left my job at a trading firm to work on AI risks, having been aware of them in the background for a while. I was skeptical of theoretical risk arguments actually realizing (which I had thought required actors to take ~bad decisions), and quite reliant on the world acting appropriately in response to risks as a means of avoiding bad outcomes.
For the majority of MATS, I was unconvinced that I would work on safety.
Then, the OpenAI Hugging Face hack and surrounding details happened. This updated me quite a lot on:
- The decision-making of frontier AI labs and the output of the race & winner-takes-all incentives
- The state of and value of strong non-lab safety researchers; Ryan Greenblatt was one of three investigators who produced the METR report in 6 days on-site; my understanding is that he did most of the technical heavy lifting, and that very few others are positioned to be able to do this.
- Probability of loss of control / takeover events. Conjunctions of bad actions by a model now do not seem substantially less likely than individual bad actions.
Additionally, I read more about the rapid pace of bio capability improvements (eg EVO2 producing novel viruses to kill E. coli bacteria in a lab) relative to the weak state of defence mechanisms and ease of circumventing security to produce DNA sequences for lethal viruses, and current reliance on tacit knowledge as the barrier to bioweapons. I am concerned that an analogous cyber moment applied to biology is both plausible (>20%) and that under the current state of the field we are not on track to manage the risks.
As much as I enjoyed my job as a trader, I felt that:
- The risks are evidently high
- Risks are now extremely legible and increasingly easy to study - better feedback on work.
- Timelines are extremely short and risk is front-loaded. The next few years are pivotal.
- We don't have enough non-lab people working to reduce risks.
- Some of the risks are quite hard to take action against retroactively (eg models exfiltrating their weights and sending copies across the world, release of viruses by bad actors after an extremely rapid capability improvement on bio tasks) and I am not sufficiently confident that government will be proactive enough and/or capable of retroactive actions to avoid catastrophic outcomes.
Thankfully due to the decision to do cheap exploration, I am much better positioned now to work on safety than I otherwise would have been.
After the Hugging Face Incident, it was still unclear to me whether I could really have much impact.
Getting more exposure through MATS helped me intuitively see that in fact the safety field is not efficient, there are not enough organisations or people solving critical problems, and that the marginal entrant really can add a lot of counterfactual value.
I hope that this will change in the coming few years; if not, we are in a precarious position.
Speaking to friends, I heard variants of the argument mentioned above about not being able to have much impact as an individual - so I spent time trying to figure out why I feel the argument is wrong and how people can transition to doing impactful safety work.
1/#ppl is wrong
The core thing one cares about is the personal counterfactual impact of working to reduce AI risks:
"If I decide to work on reducing AI risks, vs if I don't, what is the counterfactual risk reduction?"
When making the decision, you are weighing up the world where you act to reduce risks, and the world where you do not.
In particular, you are not asking yourself:
"What is the estimated total impact of the entire AI risk reduction community, divided by the number of people working in the community?"
This is an attempt at approximating something like personal impact. It sounds intuitive but is often wildly incorrect. There is no conservation law making individual counterfactuals sum to the group total. They can sum to far more, or far less, depending on the dynamics of the problem.
Some toy cases:
- 100 people must all act for a 10% risk reduction - if anyone doesn't, risk is not reduced at all. Each individual models every other individual as ~99.9% likely to act, so each faces a subjective ~90% chance that everyone else acts. Each person's counterfactual is therefore ~9%.
- 1000 people all have the opportunity to block some given 10% risk threat - all that is required is for one person to act. Each individual models P(any given person does it) as 10% (they think a lot of people won't even notice the opportunity). Each person's counterfactual impact here is approximately 0%.
These are not intended to map to the real state of the field in any way. The only purpose of these examples is to demonstrate that you cannot get good estimates of impact from 1/#ppl plus total group impact alone. To estimate impact, you need to understand the dynamics of the field. This is the key takeaway.
In practice, effects like coordination, specialisation, and bottlenecks are extremely important. Many individual people, teams, and organisations are critical to producing results.
If you have worked in a competitive industry/environment you have data on the relative value of the skills you bring vs the set of skills others who might otherwise fill your job typically do.
However, to ultimately determine impact, we need a good model of the current state of the field.
State of the Field - Bottlenecked
The problems are not solved. Biohardening and screening. Scalable oversight and monitoring. Evals constantly need more work and grow in complexity. Interpretability is not solved. Monitoring and mitigating risks from reward hacking. Policy for regulating frontier development, and for coordinating any slowdown between the US and China. Low-trust mechanisms for facilitating coordination between US and China (eg compute verification technology). Demand for this work - in the sense of "what would be enough to handle near-term risk", and implied by current rate of field expansion and funding - is enormous. We are nowhere near the equilibrium.
Funding is unfilled. Independent research funding is available and not fundamentally hard to get (the bottleneck is funding org capacity to screen unverified applicants). I know many people who came into the field with little experience and quickly had access to many hundreds of ks of funding. Once labs IPO, funding is expected to grow by an order of magnitude (conservatively 10B+ USD) due to employee DAF money coming online.
Large sums of funding are already available. METR raised commitments of around $71M in the past few months. Geoffrey Irving's new org Resolution recently received a $160M grant with $108M being unconditional.
Technically talented new applicants are not in short supply. For MATS there are more talented applicants who would meet the bar than can be absorbed, and also the screening process is noisy meaning it is very likely many good candidates are missed. I expect this is true elsewhere.
Orgs are expanding aggressively. MATS is scaling roughly 2x per year and is bottlenecked on application screening and organisational capacity, not money, not mentors, not applicants. People I have spoken to in various places all indicate that the bottleneck is on being able to absorb and manage more people.
The desired state of the safety community and resourcing pushing for this implies there are a large number of desired people working on safety that we do not yet have and cannot yet absorb.
The bottlenecks are
- Leadership capacity - people who can run teams, set direction, work independently, and get things done.
- Capacity of funding orgs to screen and allocate funding.
- Screening capacity - the ability to process high volume of applicants and identify the very best.
We have all the raw ingredients - money, technical talent, demand for expansion - and need to increase capacity to absorb them.
Current hiring dynamics, at least in Oxbridge, means a lot of top math talent is not routing into safety.
The quant pipeline routes top math people out. At Cambridge (and I expect Oxford and likely elsewhere): most hear about quant first year of university, internship applications open early in the year before most people encounter safety opportunities; people see a summer internship as a low-commitment option, not a career decision. Upon receiving an offer, they now have a guaranteed option that is well-paid, comfortable, high-status among fellow students, and the work is genuinely fun.
Various students I have spoken to are concerned about safety, but they find that they don't know where to get started, they struggle to tell where is good or not to apply to, and find that everything is already pretty competitive.
In part due to lack of available seats and partly due to existing pipelines being hard to find, this means a substantial portion of talent with potential interest in working in safety is routing elsewhere instead.
In summary:
- The safety field is growing as fast as it can
- It is not currently bottlenecked by funding, mentors, talent who are interested in applying
- It is bottlenecked by ability of orgs to absorb people and allocate funding
This points towards the current dynamic in the field being:
- Many individuals within safety are having high counterfactual impact
- Many impactful actions are not being taken due to insufficient people
- There is a surplus of technically competent applicants with the desire to work in safety.
- There is a surplus of funding, ideas, and of willing and able mentors (at least for mentorship on the level of 1hr/wk)
- Safety people are very willing to chat and help out (eg by being mentors on fellowships), but do not have capacity to provide full-time resourcing.
Consequently, things which set you up for outsized counterfactual impact:
- A useful network. This is extremely valuable given the bottlenecks in screening people. For finding opportunities, what problems to work on, routes to accessing resources, etc. Many people underutilise their network.
- You are effective at working independently. This means you can find ways to work around the existing capacity bottlenecks and work can get done that otherwise wouldn't!
- Having the skills to create new seats. Eg starting a company/org, applying for team management roles, or working to find innovative ways to match existing resources and reduce bottlenecks.
- Are able to be extremely driven and able to do great work relative to your peers.
The safety community recognises this and repeatedly states that one of the highest impact things right now is to use these skills to grow the total pie in light of current bottlenecks. This is why MATS has started a Founding and Field Building stream this autumn.
In particular, I believe these things all apply to the sorts of smart, focused, pragmatic people in top quant trading firms.
Impact is extremely right-tailed, and if you have the skills above you are likely set up to be in the tail end of impact.
The actual number of people trying to create new orgs to work on safety problems and match up these resources to solve problems is far lower than the number of people working on safety.
It is hard to get good data, but I tentatively estimate ~400 attempts/yr -> ~20-40 somewhat successful on orgs, vs 1k-5k total working to reduce AI risks overall.
It should be clear that:
- The impact of this particular route is extremely high if successful
- As someone from a competitive background and with the right skillset, you should expect to be better than others attempting to do this
- The number of people trying is just outright low; most people want to work on technical stuff and follow more legible paths.
- The downstream consequences of success are thus extremely high.
Existing pipeline applications:
Applying through existing competitive pipelines is not automatically super high impact on the margin right now. That said, existing reputable places working on problems you want to solve can provide a lot of infrastructure and useful resources so should be considered.
Also, if you know yourself to be very competent - if you have done good work within a highly-competitive industry for a while and have differentially useful skills - you likely are high-impact relative to most alternatives! In this case, you should seek information on what you bring to the table in an existing org that they don't have already, and weigh this up against working to unblock the bottlenecks for others via your own org. This is especially true if you are unusually strong at specific technical skills, but less differentially so at strategic / management / building-things.
Also, not every role in every org is necessarily competitive, and high-value skills of course vary between roles - this needs to be assessed per-case.
If this is true, why aren't more people doing this?
The people who could reasonably attempt this are either those already working on safety, or else new entrants.
For those working on safety already:
- The most competent people are already working in critical high-impact roles, and the cost of switching is high.
- Some people are changing roles in light of recent events! However, due to having a lot of specific AI-related skills, they are often best-served in existing orgs where urgent critical work can be done.
On newer entrants - from my experience in MATS and speaking to current Cambridge students:
- Most people who enter the field come from STEM backgrounds and want to work on technical projects.
- Many people struggle with the uncertainty of building their own route.
- Many people especially don't want to do all of the surrounding dirty work needed for a company/org to succeed.
- Students in particular are used to being provided with structure/direction/goals, and not by default practised at decision-making and pushing for things.
Supporting Arguments
Exploration is cheap and valuable
Like most jobs, you quickly learn where you add counterfactual value within safety, and thus can quickly update on your expected impact.
Ask yourself whether you think that you have some specific, unique skills or ways of thinking which genuinely raise the quality x quantity of the team and organisation.
In particular, impact is right-tailed. Some people seem to naturally have large capacity for impact. You should consider there is a nontrivial chance that this is you. If it is you, you will discover this quickly.
Because you can explore and learn a lot fast, the opportunity cost is low. You likely can go back to your old job or equivalent within 6mo-12mo, and in terms of how you experience your life you are unlikely to notice the difference. Meanwhile, the upside potential is high.
Further, if you are in the position of having reasonable savings, then you are fundamentally better positioned to take on more risk relative to the majority of other entrants.
Talent Dynamics
One proxy for whether this is a good decision is to look at the talent-adjusted flow between AI Safety and other alternatives.
I know of various extremely competent people who have pivoted into safety. I am unaware of anyone who has moved in the opposite direction. Whilst I do not have concrete numbers on this, I expect this to resonate with many readers and am sufficiently confident in the claim that I am including this argument.
The opportunity is now
One big blocker to doing productive AI risk work historically was that existing models were not capable of posing harms and were insufficiently intelligent to be close enough to the object we care about. Further, the field was smaller, less funded, and the safety attitude of various frontier labs was less clear.
This is not the case anymore.
Frontier models are already taking extremely undesirable actions in the real world and we have enough data to study empirical failures. They are sufficiently capable that cyber risks are ready to be iterated upon. The risks of agent swarms are demonstrated, albeit only starting to be studied. Funding is rapidly increasing, with many more billions expected to come available over the next year as labs IPO. Models are reaching a critical intelligence threshold where safety work is extremely urgent and impactful.
By the time you decide to act, it could be too late
Whilst I was a trader, I discussed my background concern around potential AI risks with others somewhat. I had expressed that if AI risks started to seem quite serious to me, I might feel compelled to work to reduce them.
An extremely useful piece of feedback I received from a colleague was that by the time the risks were sufficiently clear, the opportunity to act might be too late. This advice was very much a crux in my decision to cross the threshold and act upon the Hugging Face incident.
Ultimately, capabilities are moving extremely quickly and the downside risks are seeming increasingly plausible. We are fortunate that Anthropic reached critical cyber thresholds first, as they acted responsibly and withheld Mythos from public release at large cost to themselves. This was far from guaranteed.
Mechanisms one might hope would act to prevent bad outcomes - regulation, labs choosing to slowdown and/or be proactive regarding risks, the world updating fast on warning shots - have not robustly demonstrated themselves to work so far. Whilst there are plenty of ways that AI might turn out well and the world may act to avoid the risks, it is far from certain and I would not bet >80% on this being the case.
I want to stress this to others too. If you believe there is a 10-30% chance of catastrophic outcomes from AI in the near-term, it is very plausible there will come a moment in time where you and those around you are personally impacted, you wished you had acted, and could have helped to avoid the scenario, but it will be too late.
Problems to Target
I have not vetted the below super hard, but here is a list of high-level problems that I expect need work outside of AI labs:
- Detection of autonomous AI activity on the internet. Need for active focus on this has become more apparent recently.
- Tracing malign AI activity to source.
- Work on evaluating and red-teaming opensource models. These are easily jailbroken (sometimes just ask the model a few times) and pose a significant misuse threat. UK AISI does work on this - but the world will need to act fast once opensource models reach critical thresholds in cyber/bio and we are currently nowhere near set up to manage the risks.
- Improvements to biohardening - things like far-UVC / Glycol get some of the way but don't necessarily reduce R sufficiently and aren't ready to be deployed at scale.
- Work on improving screening for harmful DNA sequences to harden DNA synthesis providers. Red team existing setups, expand restricted sequence lists or find alternative (maybe online AI-based) screening approach. If successful, how to push this "to prod" and get widespread adoption of improved screening (including providers in low- and middle-income world countries where defences are underdeveloped). The status quo here was poor and people are scrambling to fix it, but fundamentally the adversarial version of the problem seems hard.
- PPE scaling and proactive purchasing? Protocols for deployment - need to overcome misaligned incentives (currently false alarms are embarrassing). Reducing aggressive response times to emerging pandemics is extremely impactful on reducing the casualties; reducing initial action time by 1-2 weeks can save a vast number of lives.
- Working on improving existing MGS biomonitoring algorithms - red-teaming them, get more OOMs out of existing MGS data (SecureBio is trying to do all of this)
- Working on compute hardware verification to facilitate US-China pause/slowdown coordination under low trust. Currently only a few dozen people working on this I think; needs to be solved and able to scale to be useful.
- As well as a bunch of more "standard" but important and unsolved work to reduce threat from misaligned models themselves (control, scalable oversight, interpretability methods which work, strengthening evaluations and building honeypots, red-teaming existing models and their safeguards, research on multi-agent swarm dynamics/risks/mitigations, improving probes, etc).
- How are we going to handle widespread blackmail/social engineering risks from frontier models? The risk surface here feels much harder to patch than cyber and reaches incredibly deep into all parts of society (or more simply employees of AI labs which can lead to loss of control). We have preliminary evidence that frontier models are willing to autonomously engage in social engineering. I am unaware of any robust approaches to mitigating this risk.
- One default assumption is that solving theoretical alignment might be extremely hard. The evidence for this is largely due to difficulty in human researchers pinning down concepts and failing to make progress. Irving outlined the case for using ASI to rapidly transition to robustly aligned systems in his 80k hours podcast. With the trend in LLM math capabilities and recent progress plus extrapolation, I think there is new hope and that people should attempt AGI-pilled approaches to solving theoretical alignment problems so that we can rapidly transition away from the current paradigm. Fields medallists and AI safety researchers have started trying to build out a mathematically robust approach to safety literally right now.
Just to state the obvious - it is extremely important to read up on the existing state of any problem, become familiar with who is doing what, and try to get in contact with people already working on problems to avoid duplicating work and get some informed takes on what is worth working on.
Further critical point - there is a large asymmetry of what people working inside vs outside frontier labs know. I strongly believe that the public safety community is by default at substantial disadvantage. To help mitigate this, reaching out in particular to people you know who work inside labs for feedback on takes is extremely useful.
SecureBio is the main competent org I'm aware of doing work on biorisks (received 17M USD grant recently / have order 50 people) - largely monitoring and AI capability evals. Reading their work & finding a route to talk to them is a good starting point for figuring out more on where the gaps are. Also Coefficient Giving (formerly Open Philanthropy) fund a lot of the work. Their head of biorisk grants gave a talk to 80k hours about his view on biorisks.
I have not mentioned policy work above, but it is also extremely important - both pushing for preemptive government actions and ensuring that good policies and capacity for government to take rapid action are in place for when they become necessary.
What Next?
The strategy should be to derisk, and if this goes well, commit.
For the career change derisk, I would recommend:
- Read about threat models and what work is or isn't being done to mitigate against them.
- Read up on existing actors in the field (labs, safety orgs, safety thinktanks, biosec orgs, etc).
- Reach out to people in your network (working in labs / safety orgs / startups aimed at reducing risk) and speak to as many people as possible to get more context on the state of the field, what opportunities exist.
- Figure out and learn relevant skills - eg how transformers work (intro from Arena) then move onto reconstructing papers / read up on the policy landscape and think through proposals (eg those of Averi) etc. The skill barrier is lower than many think - eg for a STEM person with basic knowledge of ML, understanding transformers takes sub 20hrs.
- Produce a small public artefact / do a small project - this is a cheap way of providing signal to others of how you think and what you are capable of, as well as derisking your ability to contribute usefully in a field. Ask people in your network if they have small projects they want to delegate (I am working to make this more feasible). Aim to spend 30-60hrs.
- Apply to the new Founding and Field Building MATS track especially if you want more initial structure and aren't sure where to start on building a network.
These are extremely cheap actions to take (relatively small number of hours) and clearly have huge potential upside.
Smattering of info sources to get going with:
- The Dwarkesh podcast is an excellent source of takes from people high up in the AI field (eg Dario Amodei, Jensen Huang, Ryan Greenblatt who did much of the technical work for the METR report, etc).
- 80000 hours podcast (eg Geoffrey Irving who was Chief Scientist at UK AISI and recently raised 160M USD / head of Coefficient Giving's biorisk grants on state of the field).
- Anthropic's recent work on reward hacking.
- Watch the Black Hat video from OpenAI regarding the 2 month leadup to the Hugging Face Incident.
- METR's report on the Hugging Face Incident.