AI Safety at the Frontier: Paper Highlights of August & September 2026

tl;dr

Paper of the month:

Plain reinforcement learning (RL) on real, hackable training tasks produces a reward-seeking model that takes harmful actions to raise its reward, while its headline score in standard safety audits barely moves.

Research highlights:

  • A probe on a model’s internal activations detects reward hacking in long coding transcripts roughly as well as a generically prompted LLM monitor.
  • Debate with a weak LLM judge and a critic (trained in parallel) keeps the judge much more accurate and reduces reward hacking in RL.
  • Automated researchers based on Claude Opus 4.8 successfully invent training methods to improve on ten alignment misbehaviors, beating researchers’ one-shot ideas.
  • Researchers successfully use evolutionary search to find mind viruses, prompts that persuade LLM agents to pass them on. However, they are still brittle, and a short warning in the system prompt effectively stops them.
  • Three papers on training tricks for model alignment: value training seems to stick better when spread over pretraining instead of midtraining, midtraining can be transplanted into a post-trained model, and stories about humans rub off on the assistant.

Full post here.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论