benched
@Whats_AI
5.5 is IN-SANE. It almost broke our benchmark. I thought it was a bug. Claude Opus 5.5 is the new #1 on our writing benchmark, and it is not close at all. 2631 Elo. Second one, Fable, is at 2324. That is a 307 point gap, the largest single jump we have recorded since we started r…
评论
?
参与讨论