I asked Codex to optimize DeepSeek V4 Flash 8-bit MLX on oMLX. Got ~1.6x prefill and ~3x decode speedup.

Follow-up to my earlier posts: Should I sell my Mac Studio? reddit.com/r/MacStudio/s/GK7QP8Lg87 Kimi benchmark: reddit.com/r/LocalLLaMA/s/ujBsYLYmpd Short version: my Mac Studio was sitting mostly idle, and from those Reddit threads I learned about DS4 and then oMLX. DS4 got me running DeepSeek V4 locally, but I wanted the 8-bit MLX version because I worry about accuracy loss in 4-bit variants. So I tried mlx-community/deepseek-ai-DeepSeek-V4-Flash-8bit , the 302GB 8-bit affine MLX m

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论