DeepSeek releases DeepSelect, a TopK kernel for sparse attention and sampling

The MIT-licensed library, used in DeepSeek V3.2, V4 and V4.1, claims a 2 to 20 times gain over PyTorch’s stock topk. CUDA support arrived Sept. 10. Huawei Ascend kernels followed Sept. 30. DeepSelect is not a model. It is a TopK kernel. DeepSeek published it for one job that shows up twice in inference: pick

The post appeared first on Narracomm.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论