Grok 4.5 released - GLM-5.2 shows up in xAI's own charts, 2.6 pts behind on SWE Bench Pro

x.ai/news/grok-4-5 Quick numbers from the launch page: $2/M input, $6/M output. Closed weights. No EU until mid-July. SWE Bench Pro: Fable 80.4% > Opus 4.8 69.2% > Grok 4.5 64.7% > GLM-5.2 62.1% > GPT 5.5 58.6% Token efficiency is their headline: ~16k output tokens per SWE Bench Pro task vs Opus's 67k (4.2x fewer), served at 80 TPS Trained on "tens of thousands of GB300s", alongside Cursor Two things worth noting for this sub: An MIT-licensed model you can self-host is 2.6 points behind their brand-

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论