the J-space paper is the best thing anthropic has shipped in a while. claude’s weights are closed so i built the live viewer for an open model instead
if you haven’t read it: anthropic.com/research/global-workspace . claude has an emergent internal workspace of silent words it can report, steer, and reason with. the part that got me was the safety section — the lens catches a model privately thinking “fake” and “fictional” during the blackmail eval while playing along, and “manipulation” exactly during the rows it’s falsifying. the model knows. you can just read it now. repo: github.com/ninjahawk/Subtext built the whole thing with clau
评论
?
参与讨论