Squeezing Open Models At Scale
This is the kind of inference optimisation work I have been doing in go-pherence and my personal llama.cpp fork, but at Cloudflare scale. I have been trying to squeeze more useful models into the hardware I already own; they are trying to squeeze more requests into a global GPU fleet, and many of the trade-offs are the same.
I expect “mid-tier” open-weight models to become much more popular throughout the rest of the year as optimisations like these make them cheaper to serve, hopefully putting a ceiling on OpenAI and Anthropic pricing and forcing some competition. I would like this to help burst the AI bubble, but OpenAI and Anthropic’s fantasy pricing is hardly the only thing keeping it aloft.
评论
?
参与讨论