AI Safety at the Frontier: Paper Highlights of August & September 2026
tl;dr
Paper of the month:
Plain reinforcement learning (RL) on real, hackable training tasks produces a reward-seeking model that takes harmful actions to raise its reward, while its headline score in standard safety audits barely moves.
Research highlights:
- A probe on a model’s internal activations detects reward hacking in long coding transcripts roughly as well as a generically prompted LLM monitor.
- Debate with a weak LLM judge and a critic (trained in parallel) keeps the judge much more accurate and reduces reward hacking in RL.
- Automated researchers based on Claude Opus 4.8 successfully invent training methods to improve on ten alignment misbehaviors, beating researchers’ one-shot ideas.
- Researchers successfully use evolutionary search to find mind viruses, prompts that persuade LLM agents to pass them on. However, they are still brittle, and a short warning in the system prompt effectively stops them.
- Three papers on training tricks for model alignment: value training seems to stick better when spread over pretraining instead of midtraining, midtraining can be transplanted into a post-trained model, and stories about humans rub off on the assistant.
Full post here.
评论
?
参与讨论