llm performance community metric
my question about LLM performance We see a lot of posts about token prediction, token generation per second, etc. But is it really the metric? I can see that DeepSeek V4 Flash 0731 (with DSPark; mac studio + llama.cpp) produces about 22–28 TPS, but I also see that the LLM does a lot of reasoning. And this relates to others. So maybe the correct way is not to check TPS or other metrics, but to check execution: task complexity/second. I don't know if such a metric already exists and if it exists why community doesn't use it by default
评论
?
参与讨论