DeepSeek V4.1 Flash in 3 charts: vs its predecessor, a top open-weight rival, and Claude Opus 5

DeepSeek published a pretty large benchmark table for V4.1 Flash, but I found it hard to see the overall capability pattern from the raw numbers. So I grouped the shared fixed-scale benchmarks by domain and made three comparisons: V4.1 Flash vs V4 Flash — the previous generation The improvement looks broad rather than incremental, especially in coding, cybersecurity and productivity. V4.1 Flash vs Kimi K3 — a top open-weight rival V4.1 Flash comes out ahead in coding, multimodal and productivity in the shared domain averages, while K3 is slightly ahead in science/health. V4.1 Flash vs Claude Opus 5 — a frontier proprietary model This is probably the most interesting comparison. Opus 5 still leads in coding, science/health and multimodal overall, but V4.1 Flash is surprisingly competitive, and actually comes out ahead in the shared productivity benchmark. The thing that stands out to me is that V4.1 Flash looks much more like an agent/coding upgrade than a simple reasoning upgrade. These aren't universal capability scores or controlled head-to-head reruns. Each chart averages only the shared fixed-scale published benchmarks available in that domain, so missing domains are omitted rather than treated as zero. I put the underlying benchmark rows and sources here: DeepSeek V4.1 Flash vs Kimi K3 llmlearner.com/compare/deepseek-v4-1-flash-vs-kimi-k3 DeepSeek V4.1 Flash vs Claude Opus 5 llmlearner.com/compare/deepseek-v4-1-flash-vs-claude-opus-5 Curious whether people actually running V4.1 Flash in coding agents are seeing the same pattern.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论