Papers for Stable LatentMoE and Gated MLA?

Four technologies used by Kimi K3 to make it SOTA: KimiDeltaAttention (used in Kimi Linear, essentially a more general gated delta net) AttnRes - described in Kimi's own publication: arxiv.org/pdf/2603.15031 Stable LatentMoE - LatentMoE was introduced by Nvidia first that allows sparser MoE (ie more experts per routed experts): arxiv.org/html/2601.18089v1 Stable LatentMoE is supposedly even more sparse. Supposedly it is LatentMoE with Quantile Balancing according to Kimi Blog but where is the paper? Gated MLA - MLA is the KV cache compression method introduced by DeepSeek. But what is Gated MLA? Is it Embedding Gated MLA? arxiv.org/abs/2509.16686 or something else? Thanks a lot in advance.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论