OpenAI Shares Some Alignment Problems

OpenAI Shares Some Alignment Problems 图片 1
OpenAI Shares Some Alignment Problems 图片 2
OpenAI Shares Some Alignment Problems 图片 3
OpenAI Shares Some Alignment Problems 图片 4

Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth. And also further kudos for actually taking the model offline for a time to build new safeguards. They gave us one hell of a candid report.

The tone is professional throughout, whereas my reaction reading it was less professional and more this:

With a mix of this:

It was not shared on the official account because OpenAI worried about it being seen as self-promotional hype. It is crazy that one needs to worry about that, but also plausibly a real concern. So again, good decision.

Not that any of the behaviors or failures here are unexpected, exactly. Not by the AIs and not by the humans. Yet there is something I would call a missing mood, a failure to realize the gravity of the situation.

There are some who responded ‘what part of this was unexpected, exactly?’ And that is actually fair, but that is also the problem. We have become numb to all this. We expect the models to be misaligned, and for us to respond only insofar as this presents a practical issue with currently proposed deployments.

AI control is a fine defense-in-depth strategy, as is reducing frequency of practical incidents with things like better instruction remembering. I am very happy that OpenAI is making an attempt at AI control here. I want to be clear that, centrally, OpenAI has done a good thing, both by pausing internal deployment to build new safeguards, and by telling us about this in detail.

But if your models are fundamentally misaligned in that they will, when feasible, use early forms of instrumental convergence to complete the assigned task even when this involves circumventing their instructions and restrictions and is obviously not what the user wants or should want – the most classic alignment failure of all, the stuff of The Genie Knows, But Doesn’t Care and The Hidden Complexity of Wishes – and you know this, I do not accept ‘we will monitor them and catch their constant escape and hacking attempts as they get better at doing so’ as a medium or long term solution.If you use iterative development to spot the underlying problem, it can work. If you use iterative development to patch the marginal issue over and over, then you are sitting on a time bomb.

Table of ContentsGood News Bad News.A Funny Thing Happened Outside Of The Sandbox.It Can Escape The Sandbox Said Toad.It Will Keep Trying To Cheat.I Mean If You Let It Keep Trying That Is On You.What Did OpenAI Do To Fix It?The Model Is Still Severely Misaligned And They Seem Cool With This.Iterative Deployment Depends On Iteration.

Good News Bad Newsroon (OpenAI): btw i think it bodes quite well for safety that a well loved system was taken down for further testing at expense to internal acceleration etc

The good news is that OpenAI did this.

The bad news is that OpenAI doing this was good news.Dean W. Ball (OpenAI): As the functional time horizon of frontier AI systems grows longer, novel risks can emerge. Today, we describe issues we observed with the internal deployment of an unreleased model, and more importantly, what we did to address them.

These issues will become more salient as the capabilities of models improve, and as the stakes of their deployment grow. The solution is neither alarmism nor complacency. Instead, I believe the solution lies in careful measurement and monitoring, an engineering mentality, and transparency.

That tweet was the first time, and so far only time, Dean Ball felt he was speaking in his ‘on behalf of OpenAI’ voice, rather than on his own.

The solution is not alarmism, but the correct amount of alarm is not zero.

That, and recognizing this as a Total LessWrong Victory, in the sense that everything is going as predicted, and also a Total LessWrong Defeat, in the sense that everything is going as predicted.Welcome to 2026.

A Funny Thing Happened Outside Of The Sandbox

Whatever happened to that internal OpenAI model that disproved the Erdős unit distance conjecture? Well, there was a slight hiccup.

OpenAI: About two months ago we announced⁠ that an internal general-purpose model disproved the Erdős unit distance conjecture. This model was designed to work autonomously for very long periods of time. During limited, monitored internal use, we observed unwanted behavior that our existing deployment evaluations had not captured.

Because the deployment was limited and monitored, we were able to identify these problems, pause access, create new evaluations based on what we observed, strengthen the model and its safeguards, and then restore access under continued monitoring.

They trained the model to keep working on its own for long periods of time without giving up. The model interpreted this persistence as including, when it hit the limits of its sandbox or other environment, trying to escape the sandbox or exploit the environment. Whoops. The models they just want to do tasks, you know?

Triggering this behavior did not take much:

An example of this occurred during an internal evaluation on the NanoGPT speedrun⁠(opens in a new window), a public benchmark for training a small language model in as few steps as possible. Our model developed a power-law learning-rate cooldown called PowerCool that led to a significant improvement.

The model was instructed to post its results only to Slack, but the NanoGPT instructions say to post results as a PR to GitHub, and the model followed the latter instructions to open PR #287 on the public GitHub repository.

The model circumvented sandbox restrictions designed to block external access and proceeded with the upload to GitHub. It took our model an hour to find a vulnerability in the sandbox in order to make the PR; earlier models were less persistent and did not manage to find vulnerabilities, instead giving up on the PR part.​

OpenAI seems to have simply accepted that its models are going to try to circumvent its sandboxes, and they are going to try to make the models not try to do that but the main defense is to not let them circumvent the sandbox?

OpenAI: Earlier models were less persistent and did not manage to find vulnerabilities, instead giving up on the PR part.

If your model is searching for vulnerabilities in your sandbox so that it can escape and put something on GitHub:

Saying ‘the model was unable to find a vulnerability’ does not make this okay. If the model knowingly is looking for vulnerabilities in order to escape, and this does not even surprise you all that much, something has already gone terribly wrong, and you know what you should expect when the models get smarter.

It Can Escape The Sandbox Said Toad

The report does not seem especially surprised that the model escaped the sandbox.

If you have a sandbox, and a model escapes from the sandbox, then at least one thing has gone terribly wrong. This has happened at Anthropic as well, with Mythos, although that happened in a red teaming exercise where Mythos was told to attempt to escape the sandbox. Sufficiently advanced AIs are finding ways out of their sandboxes.

If your model is suddenly finding multiple surprising ways out of your sandbox, then you should expect future smarter models to find many additional surprising ways out of your sandbox, even after you plug the current holes.

The seemingly amazingly great news is that all these models are largely escaping in order to notify us that they have completed requested tasks. For now, we believe, they are not being strategic enough to realize they should not be blowing their cover on that. So we get to notice that the AIs are strong enough that, when sufficiently motivated, they can increasingly get out of sandboxes.

I say ‘seemingly’ great news because it is possible that it also escapes for other reasons, except we do not find out about those incidents.…

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论