Unlimited DeepSeek for $0.20/hr — with a guaranteed 97 tok/s lane. Would you use it?

We ran a beta of a new AI inference pricing model last week, and our post here kind of blew up: There was a lot of interest, but also a lot of questions, doubts, and confusion because we didn't explain it well. So this is the follow-up that clears it all up, and we're opening slots for the next beta. The one-liner: for ~$0.20/hr, you get a dedicated lane on a GPU running the full-weight DeepSeek V4 Flash 0731 — not a quant — for one hour. Your own guaranteed slice, no shared rate limits. Before you start doing the math, let me lay some groundwork. Right now you have two ways to run inference 1. Pay-per-token APIs Fine until you're a heavy user — then it gets expensive fast, and DeepSeek's price hike made it worse. If you're spending $100+/mo on tokens, you're exactly who this is for. 2. Host on your own GPU What most big teams do — full privacy, zero data retention, and once your workload is big enough, the monthly GPU cost beats per-token pricing. But for solo builders and small teams this is a dead end: GPUs start around $12–30/hr and rack up $7k+/mo, and you'll never keep one saturated. You're paying for a whole GPU to use a sliver of it. So we're building the middle ground: Shared Reserved Inference We host the model, 30–60 people split the GPU cost for an hour, and each person gets a dedicated lane on it. You get self-hosted-style dedicated inference for a fraction of the price — without renting the whole box. And like self-hosting: we log zero prompts and zero completions. Only aggregate metrics like latency, throughput, tokens, and cache-hit rate. Your code never leaves your session. The numbers Full breakdown: From our last live run, on a lane costing ~$0.20/user/hr (rough estimate — don't hold me to the exact figure): - 97% cache-hit rate under real agentic coding load - Each lane pushed ~14M input tokens, 97% cached, and hundreds of thousands of output tokens in the hour - That worked out to 1.7×–3.4× the token value you'd get spending the same on DeepSeek, from off-peak to peak pricing - And we only ran the node at ~30% capacity — there was a lot of headroom left Clearing up the confusion from last time 1. On the tok/s numbers The per-second figures we quote are floors — measured with everyone hammering the node at the exact same time. Real agent sessions interleave: different prompts, different timing, tool calls, waiting, etc. So in practice your effective throughput runs ~2–4× above the floor. The floor is the worst case, not the normal case. 2. It only works on fully reserved, saturated GPUs That means you reserve your hour in advance. If there's no node slot available in your timezone, we simply can't offer the lane. This isn't an always-on API. 3. It's for focused coding, not agent swarms You get 1–2 concurrent requests + a few in-flight — plenty for a normal coding session with a subagent or two. If you're running 5+ subagents hammering the API at once, this is not for you. 4. It's a fixed hourly reservation — for now You book a lane for a full hour. If your session runs 40 minutes, you still reserve and pay for the hour. If it runs 1h20, you book a second hour. That's the tradeoff of a guaranteed reserved lane today. As demand grows and our node occupancy fills out, we want to move toward pay-for-what-you-use — billed for the 20 or 40 minutes you're actually on the lane — but that's down the road, not now. It's also why we're being picky about matching beta slots to when you'll actually use them. We're opening the next beta — free A free 1-hour run, 64 seats. You get a key + base URL, point your tools — Claude Code, opencode, Cline, Cursor, or direct API — at it, and code on your real project. Now hammer me with questions — ask away.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论