ARENA 8.0 Impact Report
The impact report from ARENA’s previous iteration, ARENA 7.0, is available here.
Summary:
ARENA 8.0 took place at the London Initiative for Safe AI (LISA) between May 25th and June 26th, 2026. The purpose of this report is to evaluate ARENA 8.0’s impact according to ARENA’s four success criteria:
- Source high-quality participants;
- Upskill these talented participants in ML skills for AI safety work;
- Integrate participants with the existing AI safety community;
- Accelerate participants’ career transition into AI safety.
Overall, we feel that ARENA 8.0 was successful in meeting these criteria. We are delighted that our 26 in-person participants rated their overall enjoyment of ARENA 8.0 at 8.8/10 on average, with 10/26 respondents giving a perfect 10/10 score.
Criterion 1: Our participants were of a strong calibre, coming from diverse backgrounds and bringing a wealth of different expertise with them. Notably, twelve participants either held or were pursuing doctoral degrees, and eight arrived with prior work or research experience in AI safety. Other participants came from backgrounds including software engineering, cybersecurity, postdoctoral research across academia, and full-time study. This was our most academically senior cohort to date: twelve participants held or were pursuing PhDs, eight held or were pursuing master’s degrees, and six held or were pursuing bachelor’s degrees. In terms of applications received, ARENA 8.0 attracted 571 applications; we made 30 offers, and had an initial intake of 27 participants. 26 participants completed the programme; one left partway through the programme. This was comfortably our largest applicant pool to date, and a substantial increase on ARENA 7.0's ~370. Seven of our 26 final participants were female. These factors combined to make the ARENA 8.0 cohort a strong and highly competitive intake.
Criterion 2: At the beginning and end of the programme, we gave our participants surveys in which they rated their confidence and self-assessed competence in the skills taught by ARENA. Our results suggest that ARENA 8.0 delivered significant upskilling gains for our participants across the domains we teach. When asked to rate out of 10 how satisfied they were that they had achieved their pre-programme goals, participants responded with an average of 8.2/10. After the programme’s conclusion, participants were significantly more confident in mechanistic interpretability (improving from 3.7 to 6.4 self-reported competence on average), reinforcement learning (3.0 to 6.0 on average), and LLM evaluations (3.7 to 7.3 on average). The gains recorded on our specific, task-anchored questions were larger still; in RL, confidence in building a DQN agent rose from 3.3/10 to 7.9/10, and confidence in building a PPO agent from 2.9/10 to 7.6/10. This iteration, we again asked participants how accurately they felt they had assessed their own knowledge when filling out the pre-programme survey five weeks previously, in an attempt to account for the Dunning-Kruger effect. Here, most participants judged their initial self-assessment to have been broadly accurate, though a significant minority were explicit about having underestimated themselves at the start of the programme, and two felt they had overestimated their skills when first surveyed.
Our in-person taught programme lasts 5 weeks. On average, participants estimated their counterfactual time to learn the full ARENA content unsupervised would have been ~9.6 weeks. Six participants felt the programme was too short, and the remaining twenty stated that the duration was 'just right'. The reasons given by those who wanted more time centred on the pace of consolidation rather than the volume of content: participants described wanting more room to look back over material, organise what they had learnt, and recover if they fell behind. Several suggested extending the capstone or project components specifically. One participant stated: ‘the main benefit of ARENA for me is not that it taught me the material that much faster than I could have taught it to myself, but rather that it provided the structure and motivation for it to actually happen’. We recognise this phenomenon as a key value-add of our in-person programmes, and is a key motivation for us to scale up in the near future.
Criterion 3: Participants rated the value of being in the LISA environment at 9.4/10, with 21 of 26 respondents giving a perfect 10/10 score. This underscores the value of hosting ARENA at LISA for the sake of community integration and relationship-building, and we are grateful for the continued opportunity to host our programmes here. Participants’ post-programme feedback consistently highlighted their enjoyment of LISA as a lively hub of AI safety activity, enabling them to make serendipitous connections with professionals and organisations across the AI safety community in an organic, informal setting. Mealtimes emerged as a particularly strong theme, with participants repeatedly identifying shared lunches and dinners as the setting in which they learnt most about the wider field. Conversely, one recurring issue this iteration was crowding: four participants noted that the workspace had become congested to the point of affecting their concentration, although others mentioned this same liveliness as a positive feature of their experience on the programme. For others, it presented no issues at all, with one participant responding – in line with feedback from our past few iterations – that ‘LISA is really good at getting all the distractions and inconveniences out of the way so you can focus on work’. We put such divergences down to individual preferences.
Criterion 4: Participants’ confidence that technical AI safety is the right career path for them increased on average from 7.4/10 to 8.7/10, a larger swing than we recorded in ARENA 7.0. Participants reported an average of 8.6/10 agreement that the programme left them in a stronger position to contribute to technical AI safety, indicating ARENA’s impact on career acceleration. At the end of the programme, six participants had confirmed full-time positions in AI safety or alignment due to start within four months (all of these offers were confirmed before the start of ARENA), and a further six stated that they were currently involved in one or more hiring processes for a role in AI alignment (two of which directly came about as a result of being based at ARENA/LISA over these five weeks). Beyond this, a further eight participants were actively applying for AI alignment roles and five stated their plans to do so in the future. One participant stated that they were not planning to pursue a role directly in AI alignment, but rather in an adjacent field (i.e., the intersection of AI safety and biosecurity).
Programme Information
First, we outline when the programme took place, what topics were covered, and the main changes made to the programme in contrast to previous iterations. For more information about our curriculum’s content, see our website.
ARENA 8.0 Programme
ARENA 8.0 ran from May 25th to June 26th 2026. The schedule of the programme was as follows:
- Neural Networks Fundamentals (optional): May 25th – May 29th
- Transformers & Mechanistic Interpretability: June 1st – June 5th
- Reinforcement Learning: June 8th – June 12th
- LLM Evaluations: June 15th – June 19th
- Capstone Projects: June 22nd – June 26th
Main Changes:
ARENA 8.0 delivered a significant expansion to the curriculum, described in full in the ARENA 7.0 Impact Report. The headline additions were a new chapter on Alignment Science authored by our founder Callum McDougall, new interpretability material on linear probes, activation oracles and attribution graphs, new RLVR material in Week 2 authored by David Quarel, and new AI control content supplementing the evals material in Week 3 authored by James Hindmarch.
Staff: James Hindmarch and Joly Scriven acted as Programme Lead and Operations Lead respectively for ARENA 8.0. David Quarel resumed his role as Head TA, while Nicky Pochinkov, Callum McDougall, and Bart Jaworski acted as a rotating cast of TAs during the programme. Beyond being present to TA each week of this iteration, Nicky Pochinkov was also responsible for the compute infrastructure, for which we are extremely grateful. James Fox adopted an advisory role for this iteration.
Speakers: We hosted at least two guest speakers a week starting from Week 1 of the programme. In Week 1 we welcomed Callum McDougall, Jordan Taylor and Arthur Conmy; in Week 2, Liza Tennant and Roberto-Rafael Maura-Rivero; in Week 3, Justin Olive and Marius Hobbhahn; and in Week 4, Matthew Wearden and Marius Hobbhahn. The speaker roster drew consistently strong feedback from participants, several of whom identified the evening talks as among the most valuable parts of the programme for understanding how different organisations conduct safety work and hiring.
Career Focus: In keeping with our approach for the previous three iterations, our Capstone Week doubled up as a career-oriented week. When participants were not working on their Capstone Projects, we guided them to develop professional connections in the field and plan for their next steps post-ARENA. To this end, we hosted two careers talks – delivered by Matthew Wearden (MATS) and Marius Hobbhahn (Apollo Research) – which included advising and Q&A sessions to address participants’ questions ahead of navigating the AI safety job market and their next steps in the field. We also ran a hackathon during this iteration, which was very well received; participants valued having two distinct opportunities to run self-contained research projects.
Additionally, as has been standard practice since ARENA 5.0, we gave participants the opportunity to have 1-1 discussions with ARENA alumni now working in AI safety, giving them further support and insight into what doors ARENA might open for them. These efforts were met with positive feedback, and we plan to continue this provision in future iterations.
Accommodation:
For ARENA 8.0, participants were housed at ARK Canary Wharf. The venue’s facilities were well regarded, with participants making particular mention of the gym, pool, communal study space and 24-hour reception. However, feedback on the accommodation was mixed overall, and two issues recurred consistently enough to warrant attention. The first was temperature: ARENA 8.0 ran during a series of summer heatwaves in London, and the rooms had no air conditioning which made for difficult sleeping conditions. This was the most frequent complaint in relation to accommodation. The second was the commute, which lasted 40 minutes door-to-door to LISA. Beyond the direct cost in time and energy, several participants observed that a closer venue would likely have improved cohort bonding, since proximity would have allowed participants to travel together and spend more informal time in one another’s company.
We take both points seriously, and plan to adjust our approach accordingly in the future. In summer, demand is especially high for group accommodation in London, which can drive up prices and limit the set of possible options for ARENA accommodation. We thank participants for their patience and understanding.
Automation of Processes:
Our processes aim to ensure that our participants make the most of their day-to-day time on the programme. To this end, we have implemented a light-touch daily feedback form for participants to fill in at the end of each day. This form is designed to be minimally intrusive for our participants (readily accessible with a QR code, and taking ~2 minutes to fill in) and maximally informative for us. We use the data collected here to provide tailored support to participants; most notably, their responses about their pair-programming experience are fed into a purpose-built algorithm that aims to pair them with fellow participants with whom they work best. This form also provides a space where participants can give free-form feedback to raise concerns or make specific requests. This mechanism worked well here as it has done in previous iterations. One suggested addition from this iteration was to give participants the option of booking 1-1’s with TAs, which might be especially helpful for the curriculum’s more challenging days; another was to pre-empt participant demand for clarification in the middle of the day on our most demanding content (e.g., the start of interpretability week), thus relieving pressure on participants’ keeping pace with the morning lecture. We think that such changes represent low-hanging fruit for improving future iterations, and will adjust accordingly.
Methods of Impact Assessment:
We surveyed all of our 26 participants at the programme’s start (prior to day 1) and at the end (on the last day). The surveys were intentionally constructed to test the same domains, enabling a detailed ‘before and after’ analysis to gauge the programme’s impact. We also conducted 2-minute daily feedback surveys at the end of each day of the programme, but the bulk of this impact report relies on the pre- and post-programme survey data.
We collected three types of responses:
- Numerical ratings (out of 10);
- For skills-based questions, ‘1’ designates a complete beginner and ‘10’ a complete expert, or ‘totally disagree’ and ‘totally agree’ as appropriate.
- Multiple choice;
- Open-ended questions and responses.
We evaluated open-ended responses using thematic analysis. We highlighted key words in each response, identified recurring themes and patterns across responses, reviewed the themes, and then counted the frequency of each theme across participant responses. Each count comes from a different participant, but each participant can add to multiple theme counts if their response mentions them.
Criterion 1: Sourcing high-quality participants
ARENA’s primary selection criteria for participants remain (1) their likelihood to pursue AI safety rather than general AI development, despite teaching skills applicable to both, and (2) their technical aptitude relevant to AI safety research and engineering (especially their skill in Python programming). We continued to advertise in relatively narrow channels – including the 80,000 Hours job board, established AI safety communication channels, and our ever-expanding mailing list – to ensure that we primarily targeted an audience who had some degree of prior familiarity with AI safety. In terms of age and gender distribution, ARENA 8.0’s intake was broadly similar to ARENA 7.0’s; however, this iteration’s participants in general had a higher degree of academic seniority, and less professional software engineering experience. Four participants had over a year’s experience as professional software engineers, and twelve had either completed or were pursuing doctoral degrees. Elsewhere, eight candidates either held or were completing master’s degrees, and six either held or were pursuing bachelor’s degrees.
Selection process
Initial applications for ARENA 8.0 opened on February 24th and closed on March 8th 2026. The coding test ran from March 19th to April 5th 2026. Interviews ran from April 6th to April 10th 2026. Final decisions were communicated on April 15th 2026, which gave participants approximately 6 weeks’ notice between receiving their offers and joining us in person at LISA.
Who we selected:
We selected 35 participants (of whom 8 were eventually unable to attend, due to a mixture of personal circumstances) from ~570 applications. ARENA 8.0 had a geographically diverse cohort of participants, with participants coming from the UK, EU, USA, Canada, India and Argentina. A large proportion of this cohort was comprised of PhD and postdoctoral academics hoping to enact career transitions away from academia and into AI safety. This ensured that this intake of participants brought notably deep expertise across a diverse range of fields – spanning epidemiology, neuroscience and particle physics among others – into ARENA’s pair-programming environment for these five weeks. We are happy that ARENA represents an attractive proposition for such candidates, who are ready to take calculated risks. This made this intake unique for the range of participants’ backgrounds, and we hope to witness the continuation of such diversity in the future.
At the point of joining the programme, participants were distributed by occupation as follows:
- AI safety work or fellowships: 8
- Current doctoral students: 8
- Full-time students (undergraduate or master's level): 2
- Software engineers: 1
- Cybersecurity practitioners: 1
- Other: 6
Where previous iterations have been weighted towards mid-career software engineers seeking a transition into the field, ARENA 8.0 drew more heavily on people from academic pathways and those already working on AI safety seeking to deepen their ML skills. It is worth mentioning that the label is agnostic on our participants’ seniority and the nature of their positions; included in this count are participants conducting independent research, or who had recently concluded their time on another AI safety programme or research fellowship. We find that ARENA has value to offer to candidates both from inside and outside the AI safety ecosystem, and the mix within this cohort was itself frequently cited by participants as a strength.
Seven of our 26 participants were female, compared to ten of 29 in ARENA 7.0.
Participant Quality Indicators
The quality of our participant selection was evidenced by their pre-programme technical competencies, first assessed by us during our coding test and interviews. Participants entered the programme with solid foundations across key domains:
- Neural Network Fundamentals: Average pre-programme self-assessment of 5.7/10, indicating participants already possessed core knowledge of fundamental ML concepts;
- PyTorch Proficiency: Average pre-programme self-assessment of 4.3/10, demonstrating some familiarity with essential ML tooling yet substantial room for improvement.
These pre-programme scores reflect our identification of participants who possessed the technical aptitude necessary to engage meaningfully with ARENA’s curriculum whilst having room for substantial skill development with us over 5 weeks.
Selection Process Effectiveness
On the whole, the technical baseline of our participants met expectations given our selection methodology. Unlike complete beginners, our cohort possessed foundational knowledge necessary to tackle technical AI safety concepts immediately, whilst still recognising their own room for growth in specialised ML domains relevant to AI safety:
- In mechanistic interpretability, participants self-assessed their skills at 3.7/10 pre-programme, indicating appropriate entry-level knowledge for intensive upskilling;
- In reinforcement learning, participants self-assessed their skills on average at 3.0/10 pre-programme, positioning participants well for comprehensive RL upskilling;
- In LLM evaluations, participants self-assessed their skills at 3.7/10 pre-programme on average, again indicating substantial room for upskilling.
This skill distribution – showing strong programming foundations with targeted, safety-relevant knowledge gaps – approximately represents the profile that ARENA targets in our selection processes. We are not a true entry-level AI safety programme, and we select participants who combine a strong fundamental level of AI safety engagement and technical skills with abundant potential for upskilling in AI safety-specific ML domains. It should be noted that, given the diversity of participants' backgrounds when they join ARENA, the variance in participants' self-reported skills is large; across the topics we teach, some participants may arrive with a strong grounding in certain areas of the curriculum and no experience in others.
We note that pre-programme self-assessments across RL, interpretability and evals were lower for this cohort than for ARENA 7.0. Given that this was our most competitive application round and most academically senior intake, we do not read this divergence as a meaningful indication of a weaker cohort. More likely, it reflects the significant proportion of doctoral candidates from outside AI safety who we accepted, who were especially conscious of their room for growth in safety-relevant ML domains.
Criterion 2: Upskilling
Our core goal is to upskill participants to tackle technical problems in AI safety. The first four weeks of the ARENA in-person programme cover four technical topics (more detail on each topic is provided in the relevant sections):
- Neural Networks Fundamentals (optional): Deep learning foundations and PyTorch proficiency;
- Transformers and Interpretability: Transformer analysis, superposition, circuit identification, and other mechanistic interpretability techniques;
- Reinforcement Learning: Classical and deep RL methods, policy optimisation, and RLHF implementation;
- LLM Evaluations: Eval design and threat modelling, creating evals, infrastructure for evals and agent evaluations.
Week 0: Fundamentals]
The aim of this week is for participants to reinforce basic deep learning concepts. This week had 25 participants, as it was optional for those who felt they had sufficient deep learning experience (just two participants opted to skip this week). Topics covered included PyTorch, basics of neural networks, residual neural networks, CNNs, weights and biases, optimisation, and backpropagation.
At the end of the programme, participants self-assessed as having strengthened their foundational ML skills significantly:
- NN Fundamentals Confidence: Improvement from 5.7/10 on average to 8.0/10;
- PyTorch Confidence: Improvement from 4.3/10 on average to 7.0/10.
Participants felt that they had upskilled substantially across both of these domains. We typically expect the improvements in Week 0 to be more modest than in other areas of the curriculum; this week is an optional refresher course in the fundamentals of ML, and aims to ensure participants have the requisite toolkit to complete the rest of the ARENA material over the 4 subsequent weeks. Consequently, participants usually arrive with some familiarity with Week 0’s material, and our selection process aims to ensure that this is the case.
The Fundamentals material also produced one of the most striking individual comments of the iteration, from a participant describing the day spent building GPT-2 from scratch: 'I walked away from that day feeling like I learned months of work in just a few hours, which was an incredible feeling.'
If left to their own devices, participants estimated that self-studying the Fundamentals content to the same degree of proficiency would have taken them 2.6 weeks on average, compared to the 1 week spent on the content with us.
Week 1: Transformers and Mechanistic Interpretability
The aim of this week is for participants to understand some of the methods that can be used to analyse model internals, and replicate the results from key interpretability papers. Topics covered include the following: GPT models, training and sampling from transformers, TransformerLens, induction heads, indirect object identification, superposition, linear probes, activation oracles, attribution graphs, inference-time intervention, and sparse autoencoders. We had three speakers to supplement this week's taught content: Callum McDougall, Arthur Conmy (both of Google DeepMind) and Jordan Taylor (UK AISI).
Participants showed strong progress in our interpretability section, one of the most technically demanding areas of AI safety research:
- Domain Confidence: Improvement from 3.7/10 on average to 6.3/10;
- Practical Application: Confidence in circuit identification tasks increased from 4.3/10 on average to 7.9/10;
- Transformer Implementation: Confidence in building a toy one-layer transformer from scratch increased from 6.6/10 to 9.2/10, with 16 of 26 participants stating their confidence at 10/10 post-programme.
The transformer implementation responses represent one of the strongest single upskilling outcomes from this iteration, and the material underpinning it was well received. One participant described the transformers content as ‘among the best of the entire material’.
Week 1 was also the most popular week of the programme, chosen by 9 of 26 participants as their favourite part of the curriculum. However, several participants found the pacing of this week’s chapters more intense than elsewhere in the curriculum, with some reporting that they and others had only reached around halfway through the intended material. One participant queried whether the emphasis on mechanistic interpretability should be reduced in light of shifting industry priorities. We take this feedback seriously and are considering how to improve this week’s learning experience for future participants.
If left to their own devices, participants estimated that self-studying the Transformers and Interpretability content to the same degree of proficiency would have taken them 2.4 weeks on average, compared to the 1 week spent on the content with us. This is significantly less than in previous iterations (3.2 weeks in ARENA 7.0, and 3.5 weeks for ARENA 6.0).
Week 2: Reinforcement Learning
This week's core aim is for participants to understand classical and deep RL methods, and how RLHF is implemented on LLMs as the dominant alignment method used today. Topics covered include the following: fundamentals of RL, gym and gymnasium environments, policy gradient optimisation, PPO, deep Q-learning, RLHF, RLVR, HuggingFace, and fine-tuning LLMs. We had two guest speakers during this week: Liza Tennant (Google DeepMind, also an ARENA 5.0 alumna) and Roberto-Rafael Maura-Rivero (Meta).
Our participants’ self-assessment suggested that they had upskilled substantially in RL over the course of Week 2 with us:
- Domain Confidence: Improvement from 3.0/10 to 6.0/10 on average;
- Technical Definitions: Self-assessed understanding of core RL concepts (MDPs, policies, Bellman equations, etc.) increased from 4.0/10 to 8.1/10 on average;
- Practical Implementation: Confidence in building DQN agents improved from 3.3/10 to 7.9/10 on average, and PPO agents from 2.9/10 to 7.6/10;
- RLHF: Confidence in implementing RLHF given a pre-made PPO implementation improved from 4.6/10 to 8.5/10.
These results represent the largest gains across this iteration. Participants entered this week with the lowest baseline confidence of all domains we teach, and left with confidence in RL methods on par with that which they held in the rest of our curriculum’s subject matter. RL was the second most popular week of the programme, selected by 8 of 26 participants as their favourite.
Participant feedback on the teaching in this week was particularly strong, with several participants singling out the clarity of the RL lectures and the depth of the TAs’ command of the material. Some noted that the RL lectures felt too condensed.
If left to their own devices, participants estimated that self-studying the RL content to the same degree of proficiency would have taken them 3.3 weeks on average, compared to the 1 week spent on the content with us. This is the highest counterfactual estimate of any week in the curriculum, and we are pleased that our participants feel we accelerated their learning in RL so dramatically.
Week 3: LLM Evaluations
ARENA’s evals content aims for participants to build alignment and dangerous capability evaluations in multiple-choice and agentic settings, and understand how to use these evaluations to gain information about current frontier LLMs. Topics covered include the following: threat-modelling, using LLM APIs, implementing a pipeline to generate questions using LLMs, UK AISI’s inspect library, implementing LLM agents, scaffolding LLM agents, and new material on AI control. We had two guest speakers this week: Justin Olive (Arcadia Impact) and Marius Hobbhahn (Apollo Research).
The quantitative results for this week were strong. Evals produced the highest post-programme domain confidence score of any technical week:
- Domain Confidence: Improvement in participants’ confidence in evals from 3.7/10 on average to 7.3/10;
- Eval Design: Confidence in designing LLM evaluations increased from 5.0/10 to 8.1/10;
- Agent Building: Confidence in designing and building an LLM agent increased from 4.5/10 to 8.3/10.
Alongside these strong outcomes, participants offered constructive feedback on the evals curriculum that pointed towards a few clear areas for refinement: often, they felt that the exercises were less polished than those in earlier weeks, that it offered less immediate feedback on whether their work was on track, and that it could be tightened for pace in certain areas. Some also wanted more explicit guidance on what distinguishes a strong eval from a weak one, including how to think about baselines and statistical significance.
The domain confidence scores indicate that this week still delivered strong learning outcomes, and several participants who offered criticism were explicit that ARENA should keep the week. The AI control material introduced this iteration was received positively.
If left to their own devices, participants estimated that self-studying the evals content to the same degree of proficiency would have taken them 1.3 weeks on average, compared to the 1 week spent on the content with us. This is the lowest counterfactual estimate in the curriculum, which is consistent with the feedback above: participants felt that a good deal of this week’s material was of a kind they could have picked up independently.
Overall Learning Experience
Finally, we asked participants how they found the ARENA materials overall. This helps us assess participant experience across different ARENA cohorts and gather feedback on the quality of our teaching methods.
We asked all 26 participants what their favourite aspect of the ARENA curriculum was. Below is a tally of how these 26 responses were distributed:
- Neural Network Fundamentals: 5
- Mechanistic Interpretability: 9
- Reinforcement Learning: 8
- LLM Evaluations: 1
- Capstone Projects and Careers: 3
Participants provided positive feedback on the programme's educational components:
- Exercise Enjoyment: Average rating of 8.1/10;
- Teaching Quality: Average rating of 8.7/10, with 12 of 26 respondents awarding a perfect 10/10;
- Goal Achievement: Average rating of 8.2/10 for whether participants felt they achieved their personal learning objectives at the start of the programme;
- Overall Enjoyment: Average overall enjoyment of 8.8/10, with 10 respondents giving a perfect 10/10 score.
Teaching quality drew the most positive qualitative feedback of any aspect of the programme. Participants described the TAs as patient, knowledgeable and willing to acknowledge the limits of their own understanding, and several identified the teaching (including our instant ‘summon TA’ mechanism in Slack) as one the best things about ARENA.
The most substantive teaching criticism concerned the morning lectures, raised in some form by around ten of our 26 participants. The most common thread was pace: several found the lectures rushed, particularly in RL, and wished for more time when being introduced to a new topic. However, this was not a universal view: a comparable number of participants praised the lectures as high-quality, comprehensive or a highlight of the programme, and one named the morning lectures as the most useful part of their day. Suggested remedies included lengthening the lectures at the expense of pair-programming time, adopting a flipped model where participants read the material first and have a later session for consolidation and questions, and providing a short scaffold or concept map the evening before each day. A few participants felt that the lecture slides were too congested with technical material, a format which did not lend itself to presentation on LISA’s TV screen.
Pair programming remained popular with most participants but produced the widest divergence of opinion of anything in the programme. Many described it as the single most valuable format for this curriculum, crediting it with both technical learning and the informal transfer of skills outside the curriculum.
Criterion 3: Integration
Our participants spent 4 to 5 weeks full-time in the LISA office in London. When asked to rate out of 10 how valuable it was to be based in LISA for the programme, our participants responded with an average score of 9.4/10, with 21 of 26 respondents giving perfect 10/10 responses. This is an excellent score and highlights the value of LISA – with its speakers, researchers, events and non-stop activity – as a home for the ARENA programmes. We are grateful for our relationship with LISA and hope to maintain this successful partnership long into the future.
The one consistent criticism of the LISA environment this iteration concerned capacity. Several participants reported that the workspace had become noisy and congested enough to affect their ability to concentrate, with specific mention of talks taking place adjacent to the main working area and of circulation space around desks. One participant described the noise as making sustained focus difficult throughout the day. Others raised the access control system, noting delays with the app-based entry and difficulties with the newly introduced turnstile. We raise these points not as criticisms of LISA, which continues to provide an outstanding environment for our programmes, but because they reflect the practical consequences of a workspace whose popularity has grown considerably. We will continue to work with the LISA team on managing the effects of this for future cohorts.
Similarly, when asked to rate the quality of ARENA's operations out of 10, participants responded with an average score of 9.4/10, with 15 of 26 respondents giving perfect 10/10 responses. We are delighted that our participants felt that we provided them with the support they needed to concentrate on their learning outcomes, connect with their peers and integrate into the AI safety community over these five weeks.
Beyond quantified scores, some of our participants' comments provide more detail on specific themes.
Connections to and feeling part of the AI safety community
- ‘I love LISA! I love how bustling it is. I love all the people in and out of here.’
- ‘[I enjoyed] being surrounded by AI safety researchers, other like-minded people, [and] being able to interact with not only technical researchers but also people in the broader AIS ecosystem.’
- ‘Super valuable exchanges over lunch. Everyone is so nice and open to talking, I met types of people I never would have met otherwise.’
- ‘[I loved] meeting and making friends with random people outside ARENA. [I also liked] talking with and hearing people’s chats over lunch and dinner on their research and hot takes.’
- ‘I really enjoyed that there were so many different people working on different topics and in different fellowships I was able to talk to and connect with. It was super interesting to see what other people besides ARENA people were working on. And also the events and talks that were happening at LISA.’
Access to researchers and the wider field
- ‘[I liked] being able to meet with researchers and alumni from other fellowships constantly and easily.’
- '[My most valuable gain were the] insights and ideas from talking to people; also generally orienting myself to what people seem to think is important and whatnot. I think the latter is especially important for continuing to learn independently.’
- ‘I think eating with different people over lunch and dinner is very valuable and just gives a lot of insight into different fields. It also decreases the threshold for getting help or opinions from others.’
- ‘I became aware of opportunities that I wouldn’t have otherwise.’
Upskilling and having a suitable learning environment
- ‘In person is soooo much better. It was immersive, no distractions. [I liked] chatting with other participants, exchanging about career plans and opportunities. And also meeting other people in the field during lunch breaks was really cool. I hope to come back.’
- ‘It was great to be able to be surrounded by others trying to work towards a similar goal, learning from each other, and making strong connections. I seriously do not think I would have learned even 10% of any of this if I tried to pursue the ARENA curriculum by myself!’
- ‘Pair-programming is a fantastic format for this curriculum, and it would've been much much harder to do alone. Also being in a new setting for the 5 weeks creates an atmosphere of purpose and focus.’
- ‘LISA is really good at getting all the distractions and inconveniences out of the way so you can focus on work.’
- ‘It was a lot of fun spending so much time immersed with a big group of value-aligned people. Also, since food and accommodation was cared for, I could focus on ARENA completely.’
Immediate access to expert help
- ‘The TAs honestly made the programme. All of the TAs were incredibly patient, calm, and so so so informative. I loved working with each and every one of them! This program would not be the same without these TAs.’
- ‘Talking with the TAs about my confusions was one of my favorite parts of the programme.’
- ‘Loved the TAs’ enthusiasm about the materials, and their patience to go over the math parts again and again.’
- ‘Nicky always starts running to come and answer a question we had. And generally all TAs were very patient with my questions.’
- ‘The TAs all really know their stuff. Also, they really focus on the essentials and don't bother us with stuff that's not really relevant. So much better than university.’
- ‘The TAs were wonderfully patient and sweet. They are the best part of ARENA.’
Criterion 4: Career Acceleration
Finally, ARENA aims to accelerate participants’ AI safety careers. We’re really excited about this cohort’s next steps in the world of AI safety work. At the end of the programme, our 26 participants fell into the following camps with respect to pursuing careers in AI safety:
- Holding confirmed full-time positions in AI safety starting within four months: 6
- Currently involved in one or more hiring processes for a role in AI safety: 6
- Actively applying for AI safety roles: 8
- Planning to apply for AI safety roles in the future: 5
- Not planning to pursue AI safety roles in the future: 1
These outcomes represent an exceptional return into the AI safety talent pool, and they make us optimistic about ARENA’s impact along these lines in the future. It is rewarding to witness the programme’s impact in securing one of ARENA’s core goals: to provide talented individuals with the skills to go directly into technical AI safety work. The two participants whose hiring processes began as a result of ARENA represent counterfactual placements, and we expect the fuller picture of our contribution to emerge over the coming months as those currently applying to AI safety positions receive outcomes.
We are pleased that our participants agreed with 8.6/10 confidence, on average, that they felt in a stronger position to contribute to technical AI safety work after completing the ARENA programme with us, with 8 of 26 respondents answering this question with a 10/10 agreement rating.
We also saw a marked difference in participants’ confidence in AI safety being the right field for them. Prior to the programme, when asked to score out of 10 ‘How confident are you that AI safety is the right field for you?’, participants gave an average response of 7.4/10. By the end of the programme, the same question was met with an average response of 8.7/10. This increase of 1.3 points is larger than we recorded in ARENA 7.0 (which saw a swing of +0.4), and suggests that ARENA functions effectively as a means for participants to test their fit for AI safety work. Eighteen of 26 participants reported that their career goals had shifted at least partly over the course of the programme, with four describing this change as significant. While our participants have mostly grown in confidence regarding their suitability for work in AI safety this time, we recognise the possibility that the programme may have the opposite effect on others, and we regard that as an equally useful outcome where it occurs.
Overall Programme Experience:
We asked the participants some key questions to gauge ARENA's counterfactual impact and participants' overall appreciation of the programme.
- To what extent do you feel the ARENA programme has left you in a stronger position to work in technical AI safety?
- Average response: 8.6/10.
- How much did you enjoy the programme overall?
- Average response: 8.8/10.
- Do you feel you achieved your goals?
- Average response: 8.2/10.
- To what extent do you feel being in-person at the LISA offices contributed positively to your experience on ARENA?
- Average response: 9.4/10.
- How much did you enjoy the exercises, overall?
- Average response: 8.1/10.
- How did you find the quality of the teaching?
- Average response: 8.7/10.
- Did you feel ARENA was well run, logistically?
- Average response: 9.4/10.
Most Valuable Gain
We asked participants ‘What was the most valuable thing you gained from the programme?’ and thematically analysed their open-ended responses. All 26 participants gave responses for this question; as their responses were long-form, they often addressed more than one of the categories outlined below. In total, we counted 39 distinct tokens from the 26 responses given:
- ML skills and knowledge: 18 (46%)
- Making personal connections in the AI safety field: 12 (31%)
- Confidence to take on technical AI safety work: 5 (13%)
- Improved career prospects and direction: 4 (10%)
ARENA's core mission is to ‘provide talented individuals with the skills, community, and confidence to contribute directly to technical AI safety’. With this in mind, it is encouraging to note that 35/39 (90%) of the tokens we counted in response to this question explicitly identified the acquisition of new skills, valuable connections or confidence to contribute to technical AI safety as the most valuable asset that our participants took away from the programme. These categories need not be seen as mutually exclusive – improvements in people's skills and connections likely contribute to enhanced career prospects, for example – but these tokens were counted according to what participants made most salient in their responses. Adjusting for these different interpretations makes no significant difference to the conclusions that we draw from these responses.
Improvements
As a team, we endeavour to use feedback to improve the quality of ARENA for participants. Each iteration, we learn how better to run the programme to scale its impact. Although this programme was successful according to ARENA’s success criteria, we noticed several improvements that would benefit future iterations.
Development of New Materials
Due to the nature of the AI safety ecosystem, our materials need to be reviewed and updated regularly to retain their relevance. Techniques, methodologies and disciplines that are considered state-of-the-art in AI safety can become obsolete within a matter of months, such is the rate at which AI innovation progresses; as such, technical safety work also shifts to keep pace with this.
ARENA 8.0 delivered a significant expansion to our taught material, including the new Alignment Science chapter, new interpretability content on linear probes, activation oracles and attribution graphs, new RLVR material, and new AI control content. Participant feedback on these additions was positive; the Alignment Science exercises and the control material were both singled out favourably, with the caveat that some of the Alignment Science exercise sets run longer than a single day comfortably allows.
Participants also identified areas they would like to see covered in future. The most frequently requested were agentic coding workflows, red teaming and scalable oversight, and deeper treatment of AI control. Requests for material on model organisms, jailbreaks, debate, and cognitive-science-inspired evals were also made. We will weigh these against our existing commitments as we plan updates to the curriculum for future iterations.
Teaching and Support
While the quality of ARENA’s teaching received overwhelmingly positive feedback this iteration, we are constantly receptive to new ideas for improving the support we provide to participants. It has been brought to our attention that certain aspects of the curriculum – notably the conceptually challenging days of weeks 1 and 2 – might benefit from proactive, additional TA support. We have a few ideas for how to address this in future iterations, including providing participants with the option to pre-book 1-1’s with TAs to clarify elements of the material, and scheduling afternoon group drop-in sessions for days where learners typically encounter the most difficulty.
Accommodation
As discussed earlier in this report, our two clear lessons from ARK Canary Wharf are that air conditioning must be treated as a requirement for summer iterations, and that proximity to LISA should be weighted more heavily in venue selection in the future. We appreciate our participants’ patience with the challenges posed by the weather and commute in this iteration, and thank them for their understanding.
Acknowledgements:
This report was produced by @JScriven at ARENA, and was reviewed with comments by @JamesH. We thank David Quarel, Nicky Pochinkov and Bart Jaworski, who acted as TAs during this iteration of the programme. We also thank Callum McDougall, Jordan Taylor, Arthur Conmy, Liza Tennant, Roberto-Rafael Maura-Rivero, Justin Olive, Marius Hobbhahn and Matthew Wearden for speaking to our participants during ARENA 8.0. We thank Coefficient Giving for their generous support of the ARENA programme.
- These educational statistics should be interpreted as each individual’s highest level of academic achievement; all participants with master’s degrees should be understood to hold bachelor’s degrees, etc.
- All of these offers were obtained prior to the start of ARENA 8.0.
- Two participants stated that their involvement in such processes was a direct result of being based at ARENA/LISA during the ARENA 8.0 programme.