DeepSeek V4.1 Flash vs V4 Flash Vision Exp: 38% faster and 43% fewer tokens in my quick coding test
I do quicktest for DeepSeek V4.1 Flash and DeepSeek V4 Flash Vision Exp using the same prompt. Both models ran on the latest DeepSeek Harness. I asked each model to build the same app while following one skill file and an app specification. Results: DeepSeek V4.1 Flash - Duration: 18m 49s - Total usage: 11.59M tokens - Output: 126K tokens - LLM time: 7m 48s - Generation speed: 361 tok/s - Cache hit: 99.5% DeepSeek V4 Flash Vision Exp - Duration: 30m 11s - Total usage: 20.31M tokens - Output: 154K tokens - LLM time: 25m 10s - Generation speed: 120 tok/s - Cache hit: 99.7% In this run, V4.1 Flash finished about 38% faster and used roughly 43% fewer total tokens. Its reported generation speed was also around 3x higher. The V4.1 Flash session still showed one task in progress and one pending when I took the screenshot, while Vision Exp had completed all 11 tasks. So this is not a controlled benchmark, just an early quick test. I’ll put the full prompt and specification in the comments. The 3rd image is the result for V4.1 Flash, the 4th image is the result for V4 Flash Vision Exp.