When would you leave Anthropic? Notes from a chat with a capabilities researcher

I was at a house party hosted by an AI Safety friend of mine. I join a conversation midway, where my friend is saying that doing capabilities research at frontier lab is evil, given the catastrophic risks. Nothing out of the ordinary, until I find out the person who they are talking to is a capabilities researcher at Anthropic!

I was shocked at the directness of my friend, though maybe I should not be given they are Dutch. The researcher took it well though, partly based on their temperament, and partly because they have been exposed to several AI Safety spaces previously.

Unfortunately, at this point, my memory of the conversations is extremely shaky. Also, the conversation was not continuous and happened piecemeal between other ongoing discussions. But here is my best recollection.

The researcher responds with something like: “Not a good use of our time to go flesh out our positions, as we will just re-hash the standard arguments. For example, one counter is that Anthropic is creating products that people pay for and find valuable. Or, [insert another tepid argument that I can’t remember]”. I was surprised that this is the first thing they thought of, rather than talking about the potential gigantic upsides of advanced AI systems. As much as I lean libertarian, justifying catastrophic risks with ‘customers pay for our products’ was pretty weak, at least with the particular way they phrased it.

I have some 1-1 discussion with the researcher, and sympathise with them being called evil. I bring up vague idea I have had that researchers in frontier labs should have some kind of public statement about personal redlines: what things would AI’s or Anthropic or Anthropic leadership do that would cause them to leave the company.

They respond saying that this would likely just be a checkbox exercise, with people copying and pasting some standard meaningless statement which has no teeth or consequences.

I say that instead, maybe people should just post a statement every six months along the lines of “I have reflected on the risks and benefits of working at [frontier lab], and have decided to [stay/leave]”. Sure, again, they could just copy and paste this and use it as a checkbox exercise, but I sense that most people would not post such a statement if they had not actually done a reflection, whereas people might honestly post a statement about their personal red lines, and then change their minds about it later.

After some discussion, the idea morphed to: every six months, they have a discussion with somebody – e.g. me – where the aim is to help them explore their own personal views. They disliked framing, saying it would be adversarial and feel like an inquisition with lots of AI safety people around challenging them. I said the discussion would be 1-1, would be done on whatever terms the other person wanted, they could leave the discussion when they wanted, and that my personal style is to just try to understand the other person, rather than to explicitly change their mind.

They then started asking a question as a counter, “What would you think about having a discussion every six months about big things in your…”. They did not finish the question, because the answer was evident to both of us. Yes, that would be useful! Of course it would.

At this stage that they had no good reason not to do this exercise, which of course is separate from them actually doing it. The final question I ask – in an attempt to identify an underlying crux between us – was, “Do you think it is possible, in the next 10 or 20 or 40 years, for an AI system to be created that causes human extinction?” They responded immediately, “Yes, of course.”

I can only assume that the expression on my face was asking: “So why are you working at Anthropic then?!”

It was interesting to see the cogs turning in their head head realtime: I am confident that there is a significant part of them that believes they should not work at Anthropic, and that the other parts of them are desperately trying and failing to come up a coherent reason not to think about the issue.

As we part ways, I say that the offer to have this discussion is real, and they are welcome to stay in touch. I sense this is unlikely to happen, with the fact we did not share contact details being the least relevant reason.

Final thoughts. I am curious to know what you think.

  • I believe that direct outreach to capabilities researchers is a useful lever, at least if done respectfully and collaboratively. Going up to them and calling them evil to their face might nudge some people, but I think much less so than having them think things through for themselves. However, I have no clue how this would actually work logistically or how to make something like this happen.
  • I believe that frontier lab employees – especially those that are catastrophic-risk-pilled or those who signed the Pacing the Frontier open letter– are obliged to regularly and publicly share some kind of statement regarding their red lines and whether they would leave. And then ideally following through when those redlines are crossed. The reception to Alex Turner’s post about why they left GDM is indicative of the value of such actions.
  • I believe that people more broadly should have personal redlines regarding the use of frontier AI systems. There is something incoherent about believing a a company might cause human extinction, but then also being their customer. This is something I still need to think about for myself.
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论