When people say Opus 5.5 never made a mistake on launch, I get scared.

So we’re back at nerf discussions again and as much as I think variation in performance is a real thing. I can’t help but think people who say opus 5.5 never made mistakes on launch, are not paying attention to what the model is doing behind the scenes. These models are great, especially if your use case is coding but to kind of think they judged on how well they one-shot prompts is insane to me. 5.5 definitely made mistakes or introduced bugs in my use case from launch. Not because it was being dumb but because bugs are a part of software development. You have to ask, why do bugs happen? Usually it’s because a requirement was missed, logic wasn’t fully fleshed out, a use case wasn’t considered to be plausible until a user did it or just bad code. The reason I struggle to believe a nerf has happened is that I more or less see the same performance and more importantly I never trusted 5.5 to not make mistakes, that just seems silly. People who are saying it didn’t either never revised its workings beyond the front end. Even if you’re not a developer, you can learn to write tests for Claude and measure its output against those and by tests I mean feature tests. Another thing that worries me is that people seem to never think of user error when they have frustrating sessions with Claude. How are you prompting? How well is your Claude.md file written to handle the use case you’re asking it to do? How well do you understand the problem you’re asking Claude to solve? I could go on but you get my point. Idk man, I just feel like AI is making us lazy and entitled. Especially when most of us are using this thing for hobby projects that will never turn into anything.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论