Wafer pushes GLM-5.2 Fast onto AMD as Nvidia inference costs bite
The six-person inference startup says MI355X tuning got GLM-5.2 to 2626 tok/s/node, with Vercel and OpenRouter now carrying the model.
评论
?
参与讨论
The six-person inference startup says MI355X tuning got GLM-5.2 to 2626 tok/s/node, with Vercel and OpenRouter now carrying the model.