New Benchmark to see if Opus 5.5 gets nerfed
There’s been a lot of talk about Anthropic nerfing models over time. So I decided to test it. I’ll be running a series of deterministic tests on Opus 5.5 for the next month every day and will log any discrepancies. I’ll be announcing results daily on the repo: ⭐️ github.com/ninjahawk/livenerf I will be very transparent in this benchmark and it will all be independently verifiable. You’re free to open PRs with your own additions as well and it’ll be merged after I’ve reviewed it. I used Claude to design the tests but intentionally made them be difficult tasks that will not change throughout the month. Meaning if it’s capable of a task now but isn’t later in the month, that should show since up since the questions won’t be changing. Gonna put the theory to rest. If they do genuinely use some kind of exponential quantization over time like I suspect, we’ll know one month from now. Edit: grammar Edit2: My wording was a bit confusing, what I mean to say is that the questions and graders day to day are deterministic, which is standard practice for several widely accepted benchmarks. Not that the actual output from the LLM is deterministic. The output will differ. What’s being measured is statistical, does the answer degrade in accuracy on a set task over time.