Don't look at generation tokens per second as the main metric for local inference. Soon Cinese open model providers will realize that the game is all about making the thinking phase as short as possible, and 20/30 t/s will be enough if you have fast prefill.
评论
?
参与讨论