Lesson to take home: DeepSeek v4.1 Flash encoder/decoder with shared kv cache architecture does not just remove compute during prefills (when the LLM reads), it also allows, in local low-mem inference setups, to have very fast big prefills.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论