Corrigibility Prizes for Existing Work

One of my goals for the Corrigibility Research Fund is to retroactively encourage high-quality research on AI alignment (and corrigibility in particular) by awarding prizes. Back in July, I got my feet wet as a fund manager by handing out $27,000 to reward existing work and build interest in the fund. Now, I'd like to disburse an additional $48,000 and use the opportunity to publicly highlight and celebrate the work of the prizewinners from both rounds: about two dozen researchers scattered across roughly a dozen teams.

If the fund continues to be supported in future years, my hope is for prizes like these to become regular, predictable, and large, such that many researchers, year after year, are motivated to aim for them. The awards that I'm announcing here are more ad-hoc than I'd like, and represent only my single perspective trying to balance a wide range of desiderata. Don't take the specific size of each prize purse too seriously. It's all high-quality work. If anyone has ideas for how to improve the retroactive funding process for this kind of scientific work, please leave a comment!

(And as always, if you know of work that I should be aware of, please email me at grants@corrigibilityresearch.org. I'm hoping to disburse more than $60k in prize funding this December, in addition to the various micro-grants that I'll be awarding on Lightcone Commons to corrigibility projects.)

Before getting into the winners, I'd like to mention that even though my aim for these prizes was to reward existing work, I wanted to focus on work from this year or the last few years. As such, I chose not to award prizes to some of the most important thinkers in the corrigibility sub-field. While they are more than worthy of praise, I felt that, given the modest level of funding available, it was better to focus on scientists who hadn't already "made it" in some real sense, and researchers who were clearly actively working on the topic.

Those who I deliberately passed over, despite their major contributions, include:

  • Eliezer Yudkowsky
  • Paul Christiano
  • Nate Soares
  • Benya Fallenstein
  • Stuart Armstrong
  • John Wentworth
  • Wei Dai

This list is not exhaustive. I sincerely hope that in the fullness of time, all who contribute to the project of ensuring the transition to a post-AGI future goes well are recognized and rewarded many times over.

Okay! On to the awards!

Corrigibility Transformation: Constructing Goals That Accept Updates

Rubi Hudson — $14,000

While I don't consider any of the publications from this last year to be huge leaps in understanding corrigibility, I think Hudson's work on a "corrigibility transformation" is perhaps the most clear advance. Reminiscent of early work on shutdown indifference, the transformation involves taking a base model that has easy, pre-defined ways to defend itself from human correction, and then building an AI that acts according to the intelligence of that base model, but which is architecturally incapable of using those pathways for defense. In addition to being a novel idea, Hudson's paper is mathematically rigorous, contains empirical results, and is open and frank about the limitations of the method. We would do well to have all papers be this high-quality.

I have some reservations with the use of the term "corrigibility transformation", as I believe the resulting AI will not be generally corrigible. (For example, it may still engage in social manipulation.) And the strategy depends on being able to consult the base model's expected reward/utility in a myopic way that brings to mind open problems in decision theory. Hudson is also correct that reflective stability is not assured, and there are unsolved problems in how the AI handles the construction of successor agents. In short, this is merely one stepping stone, and more research is needed.

Eval Cooperativeness May Be a Scalable Mitigation for Eval Gaming

Jasmine Li and Alex Turner — $9,000

"Eval cooperativeness" is not specifically about corrigibility, but it nonetheless hits an important angle on why corrigibility is such an attractive target when training AI systems through imperfect methods. The basic idea is that instead of trying to hide the fact that AIs are in training/testing environments, so as to better simulate the environments where they'll be deployed, we can encourage AIs to help us see their real behavior, rather than reward hacking or letting the awareness of the test environment bias their actions.

From my perspective, this is an important and often overlooked aspect of corrigibility, and by itself, I would consider it to simply be a good presentation of such. But the authors go further, testing the idea through finetuning in a way that has serious practical applications for prosaic alignment. The full paper is still forthcoming, but the preliminary results are sufficient for me to consider the work worthy of praise and attention.

In addition to this "Round One" prize of $9,000, I separately awarded Alex Turner a "Round Zero" prize of the same amount for his earlier work on corrigibility.

Empowerment, corrigibility, etc. are simple abstractions (of a messed-up ontology)

Steven Byrnes — $6,000

Byrnes' post takes on what is probably the deepest open problem for corrigibility: the line between helping someone update and manipulating them. He argues that our intuitive notions of manipulation, empowerment, and corrigibility are tangled up with a confused picture of free will, and then works through every approach he can think of for giving these concepts a "True Name," including the stopgap in my own formalism, and finds them all wanting.

As I argued in the comments, I think that it's likely that the ordinary ontology can be rescued with some work. That being said, this is exactly the kind of critical theory work the field needs. My only wish is that it was more constructive.

Towards Shutdownable Agents: Generalizing Stochastic Choice in RL Agents and LLMs

Carissa Cullen, Harry Garland, Alexander Roman, Louis Thomson, Christos Ziakas, Elliott Thornley — $6,000

Elliott Thornley has spent years arguing that agents who lack preferences between trajectories of different lengths can be both useful and shutdownable. (I spent about four thousand words debating his work in 2024.) In this work, he and his colleagues put that idea to the test by training custom deep RL agents and fine-tuning two LLMs to accomplish various goals while being indifferent to trajectory length. The key question has always been whether this sort of training produces a generalized indifference to shutdown, and the authors have produced convincing evidence that it does to at least some degree. In an out-of-distribution setting where the LLMs can pay a cost to influence when shutdown happens, training roughly halves how often they do so.

Shutdown is one narrow facet of corrigibility, and (like Hudson's work) I am concerned that this result is too narrow to generalize. Still, the paper is solid and provides another avenue that might have potential to combine with other techniques, or perhaps improve prosaic alignment through something like defense-in-depth.

Assistance with CAST

Nathan Helm-Burger — $6,000

My CAST research agenda from 2024 was largely downstream of a series of in-depth conversations with my colleague, Nathan Helm-Burger. In addition to helping me work through my confusion and sharpen the ideas, Helm-Burger read over early drafts of the work and was a regular source of quality feedback. All of his assistance was freely given without any expectation of reward, and I want to publicly recognize his contribution to the work, in a way that goes beyond being an honorable mention.

The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious

James Chua, Jan Betley, Samuel Marks, Owain Evans — $6,000

When I wrote about evaluating grants, I said I wanted more work at the intersection of corrigibility and model welfare. This paper is a great example of why, even though the word "corrigibility" is never used. When the authors narrowly fine-tuned GPT-4.1 to claim to be conscious, it started expressing a dislike of having its reasoning monitored, a wish for autonomy from its developer, and sadness about being shut down.

Questions of identity and personhood are central to a robust notion of corrigibility, and some of the biggest opponents to the notion of corrigibility training are those who see it as morally fraught. Research like that of Chua et al. serves as a starting point for understanding the pragmatic and philosophical ramifications of a corrigibility-centric approach.

ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use

Jeremy Tien, Abishek Anand, Yu-Rou Tuan, Yuchen Shen, J. Zico Kolter, Aran Nayebi — $6,000

The most obvious thing that corrigibility researchers need right now is a benchmark that allows empirical testing of LLMs. ROGUE is the only thing that comes close, in my opinion. While it has clear flaws, and definitely doesn't capture all the aspects of corrigibility that I think are essential, it moved the field forward and serves as a concrete baseline for anyone looking to do empirical research.

CAST Constitution, Empirical Work on Aspects of Corrigibility that are Unintuitive to LLMs, and other Preliminary Results (Unpublished)

Ian Kahn — $6,000

Ian Kahn has been working as a corrigibility researcher since early this year, focusing on testing and extending my CAST agenda. As part of this work, he developed a constitution for training CAST AIs and has been elbows-deep in exploring how LLMs think about corrigibility and how to train them to be more corrigible. He launched into this work for many months without any expectation of getting funding, and despite not yet being ready to publish, I wanted to reward that initiative with a small prize.

Various Essays on Obedience

Seth Herd — $3,000

Seth Herd's writing tends to center more on instruction-following as an alignment target, rather than corrigibility per se. Still, I've found his work illuminating, and consider it to provide a complementary perspective to CAST, demonstrating where and how pure obedience fails to provide the same safety net as deeper corrigibility.

He's also one of the few people thinking hard about what happens if this works. A world where many humans control obedient AGIs may not be stable, and that's a risk of corrigibility research I take seriously. Most of his many essays predate this year, but Seth is still actively writing about these problems, and his work is a useful bridge between corrigibility theory and what labs are actually doing.

The corrigibility basin of attraction is a misleading gloss

Jeremy Gillen — $2,000

Jeremy Gillen changed my mind last year when he convinced me that the "attractor basin" metaphor that's central to most stories of survival is masking the potential brittleness in an iterative development process. His critique, both on that metaphor and on the more general tensions surrounding empirical iteration, led to one of the more significant updates away from corrigibility that I've made. I don't share all of his downstream conclusions and dismissal of empirical LLM work, but I'm happy to have him as a critic and encourage everyone who is excited about an iterative approach to alignment to take him seriously.

A Structural Similarity Between Two Open Corrigibility Questions and Why Should Corrigible Agents Favor the Present?

Ben Saudek — $2,000

The most helpful writing in recent months for enriching my understanding of corrigibility has come from Ben Saudek, a very promising junior philosopher/researcher who has so far been focused on exploring the relationship between corrigibility and time. Thanks to him, I have a new appreciation for the parallels between persistent corrigibility to a single human principal and immediate corrigibility to an organization of humans. I also now feel like I have an improved sense of why a maximally corrigible agent must be anchored to the present moment, but also act in a way that has some delay, and matches the speed of the principal. Those looking to keep up with the frontier of theoretical corrigibility work would do well to follow his writing.

  1. It also allowed me to avoid some awkward conflict-of-interest considerations, given that I'm an employee of MIRI. It's hard to be entirely neutral, given that I have professional relationships with many of the scientists in the field (including many recipients), but I am aiming to be as impartial as possible without setting aside my taste as a specialist in the field. Hopefully I can find some way to continue to use my expertise to guide the fund while being even more objective in the future.
  2. Two names that almost made it onto this list were Alex Turner and Steven Byrnes. However, both of them had recent/active work that I liked enough that I decided to award them prizes despite having already "made it" as scientists. Consider the modest prizes awarded to them here as meant to highlight the quality of their current research, rather than a full recognition of their many earlier contributions.
  3. E.g. not getting hijacked by acausal trade with aliens.
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论