Spiritbuun's VBR (Variable Bit Rate) KV cache — first impressions

Well, this is an appreciation post on the spiritbuun llama.cpp fork. Some weeks ago I was testing forks and configurations to find which one was the best for my secondary model on my 3060, and the winning combo turned out to be Spiritbuun's fork + CUDA + mudler's Apex I-Compact quantization for the Qwen3.6-35B-A3B model. See reddit.com/r/LocalLLaMA/comments/1tq0h1...3bapex_128k_ctx_on_rtx_3060_12gb_37_ts Then Spiritbuun shipped turbo8, and I switched my key cache to it. Same speed, dou

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论