An AI4Chemist’s perspective on biosecurity
Over the past couple of months I’ve been exploring more and more opportunities in the AI sector, with a focus on the safety and security side. As you can imagine, AI safety (AIS) is an entirely new field, which emerged with the need to make sure we as a society use this (oftentimes overwhelming) technology in the right way. Every day I’m seeing big dogs leaving frontier labs and large corporations to go work in AIS, ultimately driven by the same purpose - make sure we get AI right for our society, and not let it become the next Terminator (I’m joking but also I’m not…) While talking to some of these people, and mentioning that my background is in Chemistry, I’ve more often than not gotten the same reaction - ‘oh so you’re doing biosecurity, right?’ What now?
AI - sure, I have experience working for some AI startups, from ML to LLMs and agent evaluations; I ultimately remain a computational AND experimental chemist at heart - all my PhD revolved around how to build and deploy data pipelines, using models to predict the best materials, think about how to validate these predictions experimentally, then going to the lab and actuary collect this data myself. So it really intrigued me that according to more than a couple of experts in their fields, biosecurity seemed like the right adjacent field to translate my skills to.
I was therefore recommended to the BlueDot Biosecurity Course, a pay-as-you-wish intensive run through biosecurity. Over six days, we focused on biosecurity in the context of pandemic preparedness - moving from the basics through engineered pandemics, early detection, non-pharmaceutical interventions, medical countermeasures, and finally the broader question of how people from different backgrounds can contribute to biosecurity.
I tend to think about biosecurity more through the lens of AI4X - questions of capabilities, misuse, safeguards and emerging technologies. Pandemics are more than this though, as we all remember from COVID times - many aspects of infrastructure, coordination, information sharing, manufacturing, supply chain access and politico-economy are central to how they play out.
AI is just starting to be embedded in these. Here are my main takeaways from this course, either related to AI or 'traditional' biosecurity. The course structure is freely available on their website, with many extra resources - my views are mainly inspired by going through these, as well as the facilitated discussions with the other course attendees.
By 2050, pandemic preparedness should be boring
One idea I kept returning to throughout the course was that, given the context of recent events, I hope that by 2050 pandemic preparedness is treated more like fire safety or cybersecurity. Think of it as a permanent layer of infrastructure rather than something governments scramble to build after an emergency. We don't wait for a building to catch fire before deciding whether it should have smoke detectors, sprinklers and evacuation plans, and we don't think of cybersecurity as something that should be constructed only after a major breach.
So, why should pandemic preparedness be different?
The goal isn't necessarily to create a world in which pandemics are impossible. That is probably unrealistic. The goal is to create a world in which a novel pathogen is detected early, transmission is reduced where possible, countermeasures can be developed quickly, and societies have already rehearsed what to do when something goes wrong. And to achieve all this, we need a world where scientists, bioengineers, health organisations, politicians and policy makers are ALL working together - everywhere, not just in one country.
Prevention methods to make resilience part of the environment
Some of the infrastructure needed for this pandemic-proof world is surprisingly mundane.
Pandemic-proof buildings
Think about how obsessed we are with fireproofing in the UK. Ventilation, filtration and potentially far-UVC could become standard features of buildings, incorporated into building codes in much the same way that sprinklers and fire alarms are today. A building that is well ventilated protects its occupants whether or not they are thinking about respiratory disease.
This is one reason I became interested in far-UVC during the course. It seems promising precisely because it could be a relatively passive intervention, and I’ve seen it gaining popularity in recent years throughout the science community.
I would argue that there are still significant questions around inconsistent efficacy data, outdated standards, limited long-term safety data, regulatory risk aversion and the cost of deployment. More fundamentally, there seems to be a scale-up problem: we have promising evidence from individual case studies and proof-of-concept work, but much less evidence on what happens when these interventions are tested, deployed and evaluated at scale. Moving from a promising PoC to large-scale, real-world deployment introduces additional technical and operational challenges, including reproducibility - an issue that will probably sound familiar to anyone working across AI4X pipelines.
PPE is also a design problem
As a materials scientist, I was very excited about this part of the course, as in my mind the whole issue with PPEs efficacy is how resistant, sustainable, reusable and cost effective it is. But one observation from the discussion particularly stuck with me: existing PPE has often effectively adapted people to the equipment rather than adapting the equipment to the requirements of the people who use it. Women usually have different fit requirements, yet the "standard" design has often effectively treated one body type - the men’s - as the default. Pandemic preparedness should therefore include diversity by design, rather than treating adaptation for different populations as an afterthought. In practice, it's much easier to start off with female PPE and adapt it to male than vice versa.
There is also a sustainability question. PPE requirements can become enormous during a pandemic, and disposable equipment creates obvious material and waste challenges. There is ongoing effort in this field as we speak - as with all packaging, given the large scale and quantities needed, even an improvement in reusability or recyclability by 1% can be significant.
Screening, surveillance and governance
On top of making buildings safer, prevention also means reducing the probability that dangerous biological capabilities become too accessible in the first place. DNA/RNA synthesis screening, zoonotic surveillance and stronger governance of high-risk dual-use research are all pieces of this.
This is also where AI starts to complicate things. AI could make biological research substantially faster and more accessible, which is obviously exciting from a scientific perspective, but it also changes the distribution of biological capabilities. The question therefore isn't simply “Can AI make biological research more powerful?” - we're already seeing first hand that this is indeed the case. In the context of safety, however, “Who gets access to those capabilities, under what conditions, and what safeguards exist around them?” is critical - and not at all straightforward to answer. It's quite interesting that during this course somehow all discussions around this always started with where the R&D is currently at, and finished with ideas on how to improve current policy making.
Detection, or building a smoke-detector network for pathogens
Our current biosurveilance systems often depend on people becoming sick, recognising their symptoms, seeking healthcare and eventually being diagnosed. There is an obvious delay built into that process. A more resilient system could bring together wastewater, air, metagenomic, hospital, genomic, pharmacy and ecological data - essentially creating a distributed smoke-detector network for pathogens. Metagenomic sequencing is particularly interesting because it allows us to look for novel pathogens without necessarily knowing what we are looking for in advance. And once we start collecting enormous amounts of data, AI anomaly detection becomes a great platform for more accurate predictions.
But, as with far-UVC, how can we do it cheaply, quickly, consistently and at the scale required? There are still major bottlenecks around cost, turnaround time, bioinformatics, standardisation and coverage.
The response is only as good as its supply chain
We’ve all experience infinite rabbit holes during the pandemics looking up how quickly can we *actually* develop a vaccine. But there are another dozen questions hiding behind this one: can we manufacture enough of it, do we have the raw materials and workforce, can clinical trials adapt quickly enough, can regulators approve it, and can we actually get the doses to the people who need them?
The COVAX program was a particularly interesting example because it showed that even when we have the scientific and manufacturing capability, there is still a huge allocation and coordination problem. During a global shortage, who gets vaccines first? How much should wealthy countries be allowed to buy? How do we prevent manufacturing capacity from becoming concentrated in only a few countries?
A 2022 report from the US I looked at included case studies involving fentanyl, injectable dexamethasone and albuterol inhalers, alongside a much longer list of essential medicines. What surprised me was how many shortages involved seemingly mundane preservatives and additives (e.g. sodium chloride, potassium bicarb - you know, salts that are relatively easy to procure and purify for the high pharmaceutical grade), rather than just the active pharmaceutical ingredients (the ones that take ages to make even on miligram scale). The core of the report was centred on other meds, but I'm wondering if the lack of availability of these additives can be correlated to any blockers in administration flow.
The AI/biosecurity boundary is probably the wrong boundary
I came into biosecurity thinking partly in terms of a distinction between traditional biosecurity and AI biosecurity (more related to AI4Bio). I'm increasingly convinced that this is the wrong boundary. AI is becoming embedded in surveillance, drug discovery, biological research, diagnostics, manufacturing and information access. At the same time, many of the risks we care about aren't uniquely AI risks. A dangerous biological capability doesn't really exist in isolation. What matters is who can access it, what infrastructure they have around them, how information moves, whether the safeguards actually work, and what incentives people have.
This is where my existing interest in AI evaluations came back into the picture. Frontier AI labs are increasingly benchmarking models for biological capabilities and misuse potential, using refusal training, classifiers, monitoring and access restrictions. These are useful interventions, but there is a fundamental limitation, persistent throughout the whole field of evals:
Benchmark performance is not the same thing as real-world biological (or, more generally, scientific) capability.
Biological (and in general, scientific) expertise is often tacit (here meaning that it can't be fully grasped through text representation - such as when deciding which glassware is appropriate for an experiment, a decision informed mainly by intuition rather than a book) and distributed across multiple steps. A model might perform badly on a benchmark but become much more useful when combined with search, external databases, tools, human expertise or iterative experimentation.
And we need to distinguish between what a model can do and what it will do. A model refusing a harmful request doesn't necessarily mean that the underlying capability is absent, just as demonstrating a capability doesn't necessarily imply malicious intent - and unfortunately, models are becoming better and better at pretending they don’t have certain capabilities.
Ultimately, what I actually care about is much closer to real-world biological impact: does an AI system meaningfully reduce the expertise, time or resources someone needs to carry out a dangerous biological task? That is much harder to evaluate, and unfortunately directly testing it can itself create risks.
The hardest problems aren't necessarily the most futuristic
Our society is in an incredibly privileged position - we reached the highest rate of scientific discovery in the history, which will only increase in due time (or will it? I’ve been exploring the option that we’ll soon experience a plateau given how AI is directly influencing our cognitive abilities, especially of younger people, and how AI slop is hurting the scientific community). We already have plans in place to develop technologically sophisticated interventions: metagenomic surveillance, AI anomaly detection, universal vaccine platforms, automated synthesis screening.
Still, some of the most important bottlenecks are much less glamorous. Aside from many politicised areas, I’m even just thinking about aspects like:
- Do we have enough N95s? Do they fit different bodies?
- Can we manufacture essential medicines?
- Do we have enough workers?
- Can we distribute stockpiles?
- Can countries share information? Is it even ethical to collect so much information?
- Do we have the political will to pay for preparedness when there is no pandemic?
- And can we keep doing all of this for decades, just like we currently d with fire protection?
The technology obviously matters tremendously, but it is only one piece of the puzzle. We still need the infrastructure, incentives, governance, manufacturing, distribution and coordination before you desperately need them.
My perspective on this
I am now more certain that AI will play a key factor in how the next pandemic will play out (backed by data rather than common vibes) - it remains to be seen if for the better or worse. This is why, as a scientist, I am so interested in how we make sure the current - and future - capabilities are used towards good purposes. I think there are a lot of areas where actual subject-matter expertise is going to be essential for building better benchmarks. I’ve written before about self-driving labs, which I think have the potential of becoming the next ‘big’ think in science - perhaps bigger than AI4Science, if it plays out well. I feel like before this course my perspective was mostly focused on Materials and Chemistry, and after so many wonderful discussions I am ready to dive deeper into considering more Bio-related threats. I’ve been developing my own small benchmarks and evaluation pipelines and explored different aspects of how LLMs and agents quantify existing and potential risk (always looking for collaborators and brutally honest perspectives).