I would love to post things like: [thing i'm wondering about] [here’s my initial stab at it] [but i have no idea if this is right or wrong] [i'm sure people on LW would love to tell me where i can fin...
TL;DR Google DeepMind's recent model DiffusionGemma (DG) generates text via diffusion, meaning many diffusion steps happen before generating the final output. In particular, these diffusion steps carr...
Thanks again to everyone who submitted an essay to the contest. Without further ado, the winner of the top prize of $1000 is: The Uncertainty That Matters Isn’t Fundamental by Jimmy What I like about ...
Some multi-agent training set-ups could make language models more sympathetic to causal decision theory (CDT), even in abstract discussion.[1] We give an initial empirical demonstration of this effect...
Hi! This is my first attempt at an AI safety experiment. The full record is on GitHub and comments are very welcome. I got the idea after reading Seth Herd's post LLM AGI will have memory, and memory ...
Someone asked me "how do you get into so many things?" after seeing the random mix of stuff I'm involved in. There's no easy answer to this question, but I'll try my best to write something that sound...
Advice String Distillation This post will advocate for a fine-tuning methodology that I think is currently extremely under-rated for alignment and interpretability. It uses Context Distillation, but I...
Summary: Monitoring long transcripts in chunks, rather than all at once, catches behaviors previously missed. We pose chunked monitoring as an effective means of finding the needle in the haystack, wh...
Alex Zhao用曼哈顿计划科学家担心核爆点燃大气层来类比AI安全风险,这篇文章逐条核查了这个历史类比的准确性。第一个前提——科学家确实意识到点燃大气的可能性——是确凿的。第二个前提——奥本海默对格罗夫斯说"概率接近零"——最多是戏剧化演绎。第三个前提——曼哈顿计划在不确定中依然进行了试验——完全错误。文章指出 Nolan 电影等流行文化不断传播不符合史实的叙事。
This is a crosspost from my blog post. I recently had the great joy of meeting a group of utilitarians, but, to my complete horror, out of the twelve of them, not a single one was a classical utilitar...
Some podcasts are self-recommending enough that I look to break them down if I have the chance. This, as a debate about recursive self-improvement, was one of those. So here we go. As usual for podcas...
tl;dr: Some important AI safety research is never rerun on the newest models. There are probably cases where this would be valuable and a single well-positioned researcher could likely do this with su...
1.4 Two philosophical methods 1.4.1 Philosophy consists of updating the highest-level concepts of the mind. As discussed above, this 'updating' process can ultimately involve anything in the mind. Tha...
🔵 The blue team proposes an evaluation protocol for some capability/propensity of interest. This consists of a suite of measurement tasks, together with a preregistered decision-making process they w...
I completed this work as part of the BlueDot Impact Technical AI Safety Project. This linkpost is a somewhat condensed version of the writeup on my blog. Training against probes is considered a forbid...
Work done as part of MATS 10 with Maksym Andriushchenko TLDR: We had two CLI agents, Claude Code and Codex, predict, execute and then retrospectively estimate their own wall-clock runtime. We ran expe...
Timelines are short. Given that, the sooner we can onboard people into the alignment field, the better. In that spirit, and in light of our current applicant count and quality, Iliad is launching thre...
TL;DR: We fine-tune a judge LLM on our conceptual reasoning dataset to output a critique rating in a single forward pass. This method provides significant uplift in performance on held-out critiques, ...