An update on Generality Labs

This is an update on what we’re doing at Generality Labs, and a quick introduction for those who aren’t familiar with our work

Written by Justin Olive, co-founder & CEO

Tl;dr

I started Generality Labs because, when the time comes, I don’t think we’re going to be able to accurately measure the safety of dangerously capable AI systems.

We're focused on AI evaluation, and our goal is to ensure that the best available tools and practices are adopted at the places where it matters most - AI safety labs, governments, 3rd party evaluators, frontier AI labs.

We have 2 work streams:

Services:

  • Independent evaluation audits: we provide 3rd party audits for evaluation developers who want to catch flaws in their evals, and build trust with stakeholders through independent quality assurance
  • Advising: we provide expert input into internal processes and standards relating to AI evaluation
  • Research services: our forward deployed scientists & engineers help organisations run experiments that use our tools and infrastructure

Tools/infra:

  • Inspect Evals: our platform for sharing and running evals
  • Inspect Auditor: runs experiments to stress-test an eval, and compiles findings into an audit report

Here's how I see the next 1-2 years playing out:

  • Our research and tooling will help expose and quantify substantial gaps in the evaluation ecosystem
  • This will motivate development and adoption of better tools and practices (which we're also developing)
  • The result will be that risk assessments and AI safety research are much more accurate, helping us avoid incidents through interventions like deployment pacing and better alignment training methods

We're hiring!

Who are we?

  • Generality Labs is an AI safety lab based in London
  • We were founded in July 2026, spinning off from Arcadia Impact with $4.4M in seed funding from Coefficient Giving
  • Prior to spinning off, we delivered >$1M in contracts with the UK AISI, Epoch AI and METR. This included:
    • Co-founding and building Inspect Evals with the UK AISI
    • Designing and building 20+ evals across AI R&D and loss of control for the UK AISI internal suite
    • Eval quality control projects for METR and Epoch
  • We’re a team of 2 cofounders (James and Justin), and 8+ contractors who have previously worked for places like METR, CERN and Anthropic.

Our plan to make AI safer

High level:

  • Our mission is to ensure the field can accurately evaluate the safety of advanced AI
  • We’re doing this by building tools and standards which address common problems, and then driving adoption of these solutions where they are most needed - governments, AI safety labs, evaluators, and frontier labs

Examples of what we did in Q3 2026:

  • Conducted audits of open-source evaluations used to assess misuse risks, including ExploitBench, LabBench2, Bioinformatics(Bix)Bench, FORTRESS
    • You can find published audits on our blog
  • 3rd party auditing: we conducted an independent audit of Epoch’s internal benchmark for continual learning (EBR-Bench)
  • Built Inspect Auditor - a tool which automates quality checks for common issues like reward hacking vulnerabilities and false ground-truth labels
  • Published novel findings on:
    • Open-weight cyber risks: GLM-5.3-Flash achieves Mythos-level performance on ExploitBench (link)
    • Agent harness variability: performance on long-horizon cyber tasks diverges across agent harnesses (link)
  • Expert advising & information sharing: we’ve had opportunities to provide input into the work of highly impactful organisations like the EU AI Office, SaferAI, SecureBio, Epoch AI, the UK AISI and US CAISI
  • Overhauled the Inspect Evals documentation (link) and contribution standards
    • If you use Inspect Evals in your research, let us know so that we can include it in our featured research :-)
  • Released Inspect Evals Lint (link) - a library of automatic quality checks that we’ve built throughout 2 years of maintaining Inspect Evals

What we have planned for Q4:

  • Scale up our 3rd party auditing work, with an aim to publish at least four independent audits
  • Release Inspect Auditor for beta testing with partner organisations
  • Upskill 10+ ARENA graduates in AI evaluation best practices through our 4-week ARENA extension program
  • Roll out Inspect Auditor across all evaluations in Inspect Evals

What we have planned for 2027:

  • Independent audits of >50% of evaluations used in risk assessments of frontier AI models
  • Inspect Auditor is adopted by all major evaluators and eval developers
  • Provide access to standardised, verified implementations of all the best safety-relevant evals through Inspect Evals
  • Develop an Inspect Evaluation Standard, which eval developers can use to help them in building high quality evaluations

We're also thinking about:

  • A certification program for evaluation auditors, to help professionalise and scale up the area
  • A live or recurring systematic review which summarises the quality and coverage of evaluations across different risk areas - e.g. bio, cyber, deception, etc

Why is this work necessary?

AI safety requires us to define and measure safety

Without accurate measurement, you can’t detect problems, you can’t validate solutions, and you don’t have an objective basis for implementing governance frameworks.

Currently, measuring safety is unsolved, and getting worse

Despite 2 years of busily evaluating propensities and capabilities, we’re still often surprised by apparent step-changes and unforeseen possibilities.

  • For example, we failed to accurately / precisely predict:
    • The Opus 4.5 step-change in agentic capabilities
    • The recent paradigm shift in cyber
    • Multi-agent capabilities and alignment challenges
  • I’m particularly concerned by the egregious levels of misalignment we've witnessed in real-world settings, despite anecdotes about alignment evaluators struggling to detect such behaviours (while often trying very hard to elicit them)

In 2024, Apollo authored the Evals Gap. Unfortunately, the situation appears to have become worse since then.

To add more colour to the current situation:

  • Evaluations for tracking dual-use capabilities (e.g. cyber, bio) are mostly saturated or low quality. Coverage across risk models is very sparse.
  • Validity of alignment evals is highly suspect, with realism/evaluation awareness being an obvious confounding factor, and construct validity & predictive validity being a problem more broadly.

As capabilities continue to advance, we simply cannot afford to keep being surprised. Increased interest in monitoring and incident analysis are symptoms of our evaluations failing to predict important downstream behaviours.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论