A brief history of KV cache compression developments

How KV cache compression - from MQA and GQA to MLA and linear-attention hybrids - quietly unlocked the long context windows that make modern agentic LLMs possible.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论