Yeltsin in the AI Aisle

A year-old GPT-OSS-120b still serves 36% of Claude Opus 4.8's daily token volume on OpenRouter because the inference market has segmented across cost, speed, & accuracy. GLM 5.2 now out-serves the frontier at 495B tokens a day, & Opus 5 shipped to contest the open-weight field. Mixture-of-experts lets a 118B model decode at the cost of a 26B model, moving the boundary of the local segment.
评论
?
参与讨论