Need help tuning cache in llama-server

Hey I am running a few models on a strix halo box. Especially for the larger models (like Qwen 3.5 122B) they work okayish performance wise if the cache is utilised properly but a full cache miss at 100k context causes roughly 10-20 minute of PP time - which is extremely annoying. I will first show what I have already configured (and it helps!), then describe what is still not working well. I am very interested on your input what I could still fine tune. What I've configured so far (and what I understand it

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论