DeepSeek V4.1 Flash is legit... at about 1/100th the price of Fable
Big news (and confirmation) from our internal writing benchmark (early results): DeepSeek V4.1 Flash is legit... at about 1/100th the price of Fable. I did not expect that from a "Flash" model. DeepSeek V4.1 Flash lands at #7 for writing in our editorial voice, at 2085 Elo. Its predecessor, V4 Flash 0731, sat at #20 with 1778. It runs at about $0.012 per task. Kimi K3, the only open model above it, costs 21x more. Fable (max) costs 258x more. preview.redd.it/v8hqd375axrh1.png Definitely a good model. Between #5 and #10 on all our individual metrics. But the cost is insane. $0.0077 per script at the default effort, $0.012 at max. GLM-5.3 Flash was the only good writer in that price bracket. V4.1 Flash beats it by ~100 Elo and takes 43 seconds per script instead of 7 minutes. Against the closed models: GPT-5.6 Sol (ultra) costs 29x more for +60 Elo. GPT-6 Astra (max) costs 69x more and scores lower. preview.redd.it/6nd80n47axrh1.png One downside to highlight. Is it one of the only models that leaves "TBD" in parts of our scripts? For example, when it needs an image idea, it writes "image TBD" instead of adding one. It thought to come back later or something? It got some worse results because of these weird artifacts. Another thing to highlight: its weakness seems to be on "slop sounding" based on how we calculate it + using sam_paech 's slop EQ Bench. preview.redd.it/pctsq179axrh1.png Practical takeaway from these results: If you want the best open-weight writer, Kimi K3 still holds it (#5, 2148 Elo), at $0.26 a script and four minutes per draft for us. DeepSeek V4.1 Flash lands 63 Elo behind it for a twentieth of the price, in under a minute. That is the new default for drafting at volume. It is a really good model. Really. If the words are the product and the voice has to be yours, the Claude models still win by a wide margin in our benchmark. But for scale or for first drafts, and anything a human rewrites anyway, this is the cheapest good writing we have ever measured.