The Game is Set for a Targeted Memetic Attack on the AI Safety Community

While this is relevant to my work at MIRI, I have not checked these ideas with anyone else on the team and am posting this on my personal LW account. These views are my own. And to be honest, I am writing this mostly to remind myself of my weakness.

---

I expect one (or many) adversarial memetic attacks aiming to trip you up, perhaps consisting of fake leaks relating to dangerous stuff happening in the labs. Specifically, worrying incidents that may fit snugly within your worldview, leaking from multiple sources including news outlet/s, but not confirmed/confirmable by a primary source. Think rumors about exfiltrated weights, AIs attempting to create viruses, agent swarms hacking into and gathering information from nuclear infrastructure, etc.

An easy way to remove status from a movement is to trip it up: make it fall for a misinformation trap in public, then use that slip-up to discredit the movement for all time. The game is set for a memetic attack like this. There's a well-resourced group waiting for your screw-up.

And then you may remember much that will help you.

In public and in private, if you feel surprised or confused, notice your confusion. These feelings are signs that your world model doesn't match reality.

Real incidents make you want to act fast. You feel the need to contact journalists, tweet about the incident, and start telling your friends: a memetic attack will feel the same. If you let them trip you, you burn credibility. Set a 5-minute timer and write out your thoughts before acting. Ask yourself:

  1. Where is this information from?
  2. Do I trust the source?
  3. Can I verify any of this information myself?
  4. If I must act now, is there a way that I can take action while remaining visibly skeptical of the claim and its source?
  1. Since I haven't sanity-checked the concepts below, please take them with a grain of salt and read the comments where people will likely point out my mistakes.
  2. "What about the German Wiki Attack?": The GWA was found by a team of researchers that we trust, and there was a bunch of public evidence that it had actually happened. I anticipate something similar to this being the trap that is set, but I do not expect it to come from a team of researchers that we trust, and I do not expect there to be more than 1-2 pieces of public evidence.
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论