Ninfer Stahp

GPU: Blackwell MaxQ Model: qwen3.8-27b-nvfp4 (Orcarouter) - Dflash2 - MTP-7 Running this on my Blackwell MaxQ. I can't believe how fucking fast it is compared to llama-server . It easily reaches speeds between 130-180 t/s, even in parallel decoding I hardly see any slowdowns at all! Its concurrent capabilities are on par with vLLM in my experience. Spent this afternoon installing it on Windows 10, trying to migrate my daily driver from llama-server to something else. When I found out that I could actually run this backend on Windows 10 I gave it a try. Several hours later, this is the result. The hype is real for once. I love the kind of work put in here. This backend has such a niche level of specialization it becomes extremely powerful when used for agentic vibecoding. Props to the author for building this from scratch. I love llama-server for its versatility and level of control, but this backend is like a crotch rocket in comparison. Only thing I can say is that since its still very new, it is going to be initially missing a lot of features like advanced sampling params, broader model support, etc. but it already has all I need to do my work. Really mindblown by this thing, seriously.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论