We stress-tested DeepSeek V4.1 Flash. Now it’s $0.20/hour.
DeepSeek V4.1 Flash 262K context window, full weight $0.20/hour reserved lane 2 concurrent requests No token limits. No cache pricing. You get the lane for the hour and can use it as much as you want within that period for a flat $0.20. Over the last few sessions, we opened our DeepSeek V4.1 Flash inference lane to the community for free and basically told people to hammer it. Across those sessions, we processed billions of tokens, served thousands of requests, and maintained ~98–99% prefix-cache hit rates while sustaining throughput on a single reserved lane. The whole experiment was meant to answer one question: Can we make long-running Flash inference cheap enough that you stop thinking about token costs altogether? The results gave us enough confidence to take the next step. So, what exactly is the $0.20/hour lane? It’s a reserved DeepSeek V4.1 Flash inference lane with a 262K context window and 2 concurrency. Instead of paying per input token, output token, or cached token, you reserve the lane for $0.20/hour. During that hour, there’s no token meter running in the background. It’s mainly designed for workloads where prefix caching stays reasonably high — things like coding agents, long-running agent sessions, large-repo workflows, repeated long-context requests, and other workloads that reuse a substantial portion of the same context. That’s where this model really starts to make sense: if your workload has a healthy cache rate, you can keep a long-running session going without constantly calculating token costs. The idea is pretty simple: predictable hourly pricing for cache-friendly inference workloads. We’re now opening the beta to the general public gradually. Apply / join the waitlist here: singularityapi.tech Your first 5 hours of DeepSeek V4.1 Flash are on us ;) For some additional context, here are our previous Reddit posts: Would love to see what kinds of workloads people end up running on it. I'll be in comments to answer any questions!