3 models voted to drop GLM-5.2 (including GLM-5.2). A blind test disagreed.
Follow-up to a benchmark I ran a while back - 8 models, 4 strategic tasks, blind-scored. Three models ended up clustered at the top within 1.45 points: Opus 4.8, GLM-5.2, DeepSeek V4 Pro. (GPT-5.5 scored in the same band but was already covered by a separate route, so it is out of scope here.) A near-tie is not a decision. This is what happened when we tried to break it. reddit.com/r/opencode/comments/1u95pph/...rategic_tasks_blindscored_the_top_tier The consensus We gave all three models t
评论
?
参与讨论