Prompt Processing vs Generation: Why Your Box Is Fast at One and Slow at the Other
The same machine can rip through generating tokens yet crawl when it reads a long prompt — or the reverse. That's because local LLMs run in two phases with opposite bottlenecks. Understand them and you'll know exactly which hardware spec to buy for.
评论
?
参与讨论