I created a new benchmark and it interestingly showed the regression from Opus 4.6 -> 4.7

I originally created ObviousBench to measure the performance of small and low reasoning model's exposures to making 'dumb' mistakes, like not being able to spell Google, or walking to the car wash etc. By its nature, the benchmark is designed to be saturated by the top configurations, ie GPT-5.5 on medium+, Gemini 3.1 Pro from Low+. What I was not expecting to see was how obvious a regression Opus 4.7 was. The minimum configuration to hit >95%+: Opus 4.5: Low for $1.60 Opus 4.6: High for $0.65 Opus 4.7: xHi

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论