Wafer pushes GLM-5.2 Fast onto AMD as Nvidia inference costs bite

The six-person inference startup says MI355X tuning got GLM-5.2 to 2626 tok/s/node, with Vercel and OpenRouter now carrying the model.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论