Made theoretical implementation of v4 architecture

Compact PyTorch implementation of the DeepSeek-V4 architecture (from arXiv 2606.19348) — hybrid attention, MoE, MTP heads, Muon optimizer, configs from 50M to paper-scale. Research code, not official. Made it to understand the architecture better. github.com/Likara789/DS-v4 (Not an ad lol) submitted by /u/HolidayResort5433 [link] [comments]

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论