Fastest deepseek v4 flash setup

I built an inference engine in three days. It runs a 284B model at nearly 3× a published same-GPU benchmark. One RTX PRO 6000. A desktop Ryzen. 96 GB of system RAM. Windows 11. The result: 91.6 tokens per second across two complete nine-turn coding sessions, compared with 32.0 tokens per second for the best llama.cpp + DSpark run in the published benchmark I used. 2.86× the decode throughput on the same GPU class. This is ZLE, my private inference engine, developed as part of Zerolab. It is not a collection of launch flags or a wrapper around llama.cpp. I built the engine and its kernels using my own language. The benchmark—not a peak screenshot I used the published nine-turn DeepSeek-V4-Flash coding harness unmodified. The model builds an application and extends it over successive turns, carrying the conversation forward. Both completed runs used temperature 1.0, top-p 0.95, thinking enabled, and one request at a time. Reasoning tokens count toward generation. Decode throughput is total generated tokens divided by total generation time—not an average of selected fast moments. Complete nine-turn run Decode throughput Generated tokens Published llama.cpp + DSpark comparison 32.0 tok/s 47,977 ZLE — seed 0 90.5 tok/s 45,515 ZLE — seed 1 92.8 tok/s 47,434 That is 92,949 generated tokens across two completed sessions. In the first run, ZLE was 2.73–3.11× faster on every corresponding turn than the comparison run—not just ahead on one favorable prompt. My rule: speed does not excuse changed output The requirement is not “the answers look close.” ZLE’s speculative path must produce the identical committed token sequence as its own one-token decoding path, with the same seed and sampling settings. The validation covers greedy and sampled generation, incorrect as well as correct drafts, full-model comparisons, and continuation from saved context state. The weight representation is also checked for lossless reconstruction. The tested speculative paths reported zero differing committed tokens against that reference. That distinction matters: I am claiming exactness within ZLE’s tested execution paths, not identical output to another engine running a different model file. There are faster results—but they belong in a different category Separate short-form tests reached 162 tok/s on code, 117–120 tok/s on prose, and 97–101 tok/s on stories. Plain, non-speculative decoding measured 81–83 tok/s. Those are separate workloads. I am not presenting the 162 tok/s result as the average of the nine-turn benchmark. The hardware and the limits My benchmark machine is an RTX PRO 6000 Blackwell Workstation Edition with 96 GB VRAM, Ryzen 9 9950X3D, and 96 GB DDR5-5600, running Windows 11 Pro. ZLE was configured for 64K context; that is the configured capacity, not a claim that every turn began with 64K tokens. The comparison used the same GPU class, a Ryzen 9 9950X, Ubuntu/Docker, and a larger configured context. The model files also differ: both are DeepSeek-V4-Flash-0731 variants, but mine is the huihui-ai abliterated version and the comparison uses an unsloth quantization. This is a shared-harness, same-GPU-class comparison—not a perfectly controlled, identical-checkpoint experiment. The external results are the benchmark author’s published measurements, not runs I performed on their machine. Prefill is still a weakness. Including prompt processing and the other wall-clock costs, my two completed runs measured 61.9 and 59.0 tok/s end-to-end, versus 30.8 tok/s for the published comparison. That is why the headline explicitly says decode throughput. These are developer-reported results from one workstation and one model family—not an independently certified world record or a claim that the engine wins every workload. What I’m sharing The measurements, benchmark conditions, and validation results. Not the implementation. The source, internal designs, and proprietary methods remain private. I started this project to find out how much performance was being left on the table. Three days later, I have a working engine, completed public-harness runs, and a result substantial enough to put numbers behind. I didn’t need a second GPU to get this result. I needed different software.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论